Variants of CAS12a nucleases and methods of making and use thereof

By modifying the LbCas12a polypeptide to alter its PAM recognition specificity and using a V-type CRISPR-Cas system, the limitations of existing CRISPR-Cas nucleases in genome editing are overcome, enabling broader genomic target site accessibility.

JP2025087754APending Publication Date: 2025-06-10PAIRWISE PLANTS SERVICES INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025030119
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-10-17
Filing Date
2025-02-27
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

Existing CRISPR-Cas nucleases have stringent protospacer-adjacent motif (PAM) recognition specificities, limiting the number of genomic target sites available for modification.

Method used

Development of modified Lachnospiraceae bacterium CRISPR Cas12a (LbCas12a) polypeptides with altered PAM recognition specificities, achieved through specific mutations, and the use of a V-type CRISPR-Cas system comprising a modified LbCas12a polypeptide and a guide nucleic acid to modify target nucleic acids.

Benefits of technology

The modified CRISPR-Cas system enhances PAM specificity, expanding the range of genomic target sites available for modification, thereby improving the utility of CRISPR-Cas nucleases in genome editing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025087754000001_ABST
    Figure 2025087754000001_ABST
Patent Text Reader

Abstract

To provide modified CRISPR-Cas nucleases having improved PAM specificity, and methods for designing, identifying, and selecting such CRISPR-Cas nucleases.SOLUTION: Variants of Cas12a nucleases having altered protospacer adjacent motif recognition specificity, methods of making CRISPR-CAS nuclease variants, and methods of modifying nucleic acids using the variants are provided.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] [Statement Regarding the Electronic Filing of Sequence Listing] A sequence listing in ASCII text format submitted under 37 CFR 1.821, created on October 13, 2020, and submitted via EFS-Web, entitled 1499.7WO_ST25.txt, having a size of 257,774 bytes, is provided in lieu of a paper copy. This sequence listing is incorporated herein by reference for its disclosure.

[0002] [Statement of Priority] This application claims the benefit of U.S. Provisional Application No. 62 / 916,392, filed on October 17, 2019, under 35 U.S.C. § 119(e), the entire content of which is incorporated herein by reference.

[0003] [Field of the Invention] The present invention relates to variants of Cas12a CRISPR-Cas nucleases having altered protospacer-adjacent motif recognition specificities. The present invention further relates to methods of making CRISPR-CAS nuclease variants and methods of modifying nucleic acids using the variants.

Background Art

[0004] Genome editing / modification is a process that utilizes site-specific nucleases, such as CRISPR-Cas nucleases, to introduce variations at target genomic positions. Cas9, the most widely used nuclease for genome modification, can introduce mutations in genomic regions upstream of the NGG motif (e.g., protospacer-adjacent motif (PAM)). Other Cas nucleases have different PAM recognition specificities. When the PAM specificities of these nucleases are particularly stringent, they can reduce the utility of the nucleases for genome modification by limiting the number of genomic target sites available for modification by that nuclease.

Summary of the Invention

Problems to be Solved by the Invention

[0005] To address the drawbacks in the art, the present invention provides a modified CRISPR-Cas nuclease having improved PAM specificity, and methods for designing, identifying, and selecting such CRISPR-Cas nucleases.

Means for Solving the Problems

[0006] One aspect of the present invention provides a modified Lachnospiraceae bacterium CRISPR (Clustered Regularly Interspaced Short Palindromic Repeats) Cas12a (LbCas12a) polypeptide, wherein the modified LbCas12a polypeptide has at least 80% identity with the amino acid sequence of SEQ ID NO: 1 (LbCas12a) and at the following positions with respect to the position numbering of SEQ ID NO: 1: K116, K120, K121, D122, E125, T148, T149, T152, D156, E159, Q529, G532, D535, K538, D541, Y542, L585, K591, M592, K595, V596, S599, K600, K601, Y616, Y646, and / or W649 one or more (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13 or more) in any combination of mutations (with respect to the position numbering of SEQ ID NO: 1, the following positions: K116, K120, K121, D122, E125, T152, D156, E159, G532, D535, K538, D541, and / or K595 one or more in any combination of mutations may also be), comprising, consisting essentially of, or consisting of an amino acid sequence.

[0007] A second aspect of the invention provides a V-type Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR)-associated (Cas) (CRISPR-Cas) system comprising: (a) (i) a modified LbCas12a polypeptide of the invention or a nucleic acid encoding the modified LbCas12a polypeptide of the invention, and (ii) a fusion protein comprising a polypeptide of interest or a nucleic acid encoding the polypeptide of interest; and (b) a guide nucleic acid (CRISPR RNA, CRISPR DNA, crRNA, crDNA) comprising a spacer sequence and a repeat sequence, wherein the guide nucleic acid is capable of forming a complex with the modified LbCas12a polypeptide or the fusion protein, and the spacer sequence is capable of hybridizing to a target nucleic acid, thereby guiding the modified LbCas12a polypeptide and the polypeptide of interest to the target nucleic acid, whereby the target nucleic acid is modified or regulated.

[0008] A third aspect of the invention provides a method of modifying a target nucleic acid, the method comprising contacting the target nucleic acid with (a) (i) a modified LbCas12a polypeptide of the invention, or a fusion protein comprising the modified LbCas12a polypeptide of the invention, and (ii) a guide nucleic acid (e.g., CRISPR RNA, CRISPR DNA, crRNA, crDNA); (b) a complex comprising the modified LbCas12a polypeptide of the invention and the guide nucleic acid; (c) (i) a modified lbCas12a polypeptide of the invention, or a fusion protein of the invention, and (ii) a composition comprising the guide nucleic acid; and / or (d) the system of the invention, whereby the target nucleic acid is modified.

[0009] The fourth aspect of the present invention provides a method for modifying a target nucleic acid, the method comprising contacting a cell or cell-free system containing the target nucleic acid with (a) (i) a polynucleotide encoding the modified LbCas12a polypeptide of the present invention, or an expression cassette or vector containing the same, and (ii) a guide nucleic acid, or an expression cassette or vector containing the same; and / or (b) (i) a complex or fusion protein containing the modified LbCas12a polypeptide of the present invention, and (ii) a nucleic acid construct encoding a guide nucleic acid, or an expression cassette or vector containing the same, thereby modifying the target nucleic acid.

[0010] The fifth aspect of the present invention provides a method for editing a target nucleic acid, the method comprising contacting the target nucleic acid with (a) (i) a fusion protein containing the modified LbCas12a polypeptide of the present invention and (a) (ii) a guide nucleic acid; (b) the fusion protein of the present invention and a complex containing the guide nucleic acid; (c) a composition containing the fusion protein of the present invention and the guide nucleic acid; and / or (d) the system of the present invention, thereby editing the target nucleic acid.

[0011] The sixth aspect of the present invention provides a method for editing a target nucleic acid, the method comprising contacting a cell or cell-free system containing the target nucleic acid with (a) (i) a polynucleotide encoding a fusion protein containing the modified LbCas12a polypeptide of the present invention, or an expression cassette or vector containing the same, and (a) (ii) a guide nucleic acid, or an expression cassette or vector containing the same; and / or (b) a nucleic acid construct encoding a fusion protein containing the modified LbCas12a polypeptide of the present invention and a complex containing the guide nucleic acid, or an expression cassette or vector containing the same; and / or (c) the system of the present invention, thereby editing the target nucleic acid.

[0012] A seventh aspect of the present invention provides a method for constructing a randomized DNA library comprising double-stranded nucleic acid molecules for determining the protospacer adjacent motif (PAM) requirements / specificity of a CRISPR-Cas nuclease having a PAM recognition site at the 5' end of a protospacer, the method comprising the steps of preparing two or more double-stranded nucleic acid molecules comprising the following steps: (a) synthesizing a non-target oligonucleotide (first) strand and a target oligonucleotide (second) strand for each of the two or more double-stranded nucleic acid molecules, wherein the non-target oligonucleotide strand comprises the following from 5' to 3': (i) a first sequence having about 5 to about 15 nucleotides, (ii) a second sequence having at least 4 randomized nucleotides, (iii) a protospacer sequence comprising about 16 to about 25 nucleotides, and (iv) a third sequence having about 5 to about 20 nucleotides, wherein the first sequence having about 5 to 15 nucleotides in (i) is immediately adjacent to the 5' end of the second sequence in (ii), the second sequence in (ii) is immediately adjacent to the 5' end of the protospacer sequence in (iii), and the protospacer sequence is immediately adjacent to the 5' end of the third sequence in (iv); the target oligonucleotide (second) strand is complementary to the non-target oligonucleotide strand); and, (b) annealing the non-target oligonucleotide strand to the complementary target oligonucleotide strand to produce a double-stranded nucleic acid molecule, wherein the first sequence comprises a restriction site (at its 5' end), the third sequence comprises a restriction site (at its 3' end), wherein the first sequence (i), the protospacer sequence (iii) and the third sequence (iv) of each of the two or more double-stranded nucleic acid molecules are identical), thereby constructing a randomized DNA library comprising double-stranded nucleic acid molecules.

[0013] The eighth aspect of the present invention provides a method for constructing a randomized DNA library comprising double-stranded nucleic acid molecules for determining the requirements / specificity of the protospacer adjacent motif (PAM) of a CRISPR-Cas nuclease having a PAM recognition site at the 3'-end of the protospacer, the method comprising the steps of preparing two or more double-stranded nucleic acid molecules comprising the following steps: (a) synthesizing a non-target oligonucleotide (first) strand and a target oligonucleotide (second) strand for each of the two or more double-stranded nucleic acid molecules (wherein the non-target oligonucleotide strand comprises the following from 5' to 3': (i) a first sequence having about 5 to about 20 nucleotides, (ii) a protospacer sequence comprising about 16 to about 25 nucleotides, (iii) a second sequence having at least 4 randomized nucleotides, and (iv) a third sequence having about 5 to about 15 nucleotides, wherein the first sequence having about 5 to 20 nucleotides in (i) is immediately adjacent to the 5'-end of the protospacer sequence in (ii), the second sequence in (iii) is immediately adjacent to the 3'-end of the protospacer sequence in (iii), and the third sequence in (iv) is immediately adjacent to the 3'-end of the second sequence in (iii); the target oligonucleotide (second) strand is complementary to the non-target oligonucleotide strand); and (b) annealing the non-target oligonucleotide strand to the complementary target oligonucleotide strand to produce a double-stranded nucleic acid molecule (wherein the first sequence (i) comprises a restriction site (at its 5'-end) and the third sequence (iv) comprises a restriction site (at its 3'-end), wherein the first sequence (i), the protospacer sequence (ii) and the third sequence (iv) of each of the two or more double-stranded nucleic acid molecules are the same), thereby constructing a randomized DNA library comprising double-stranded nucleic acid molecules.

[0014] A ninth aspect of the present invention provides a randomized DNA library for determining the protospacer adjacent motif (PAM) requirements / specificity of a CRISPR-Cas nuclease having a PAM recognition site at the 5' end of a protospacer, the randomized DNA library comprising two or more double-stranded nucleic acid molecules, each comprising: (a) a non-target oligonucleotide (first) strand and a target oligonucleotide (second) strand, wherein the non-target oligonucleotide strand comprises, from 5' to 3': (i) a first sequence having from about 5 to about 15 nucleotides (e.g., about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 nucleotides, and any range or value therein), (ii) a second sequence having at least 4 randomized nucleotides (e.g., at least 4, 5, 6, 7, 8, 9, 10, or more, and any range or value therein), (iii) a protospacer sequence comprising from about 16 to about 25 nucleotides, e.g., about 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 nucleotides, and (iv) a third sequence having from about 5 to about 20 nucleotides (e.g., about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 nucleotides, and any range or value therein), wherein the first sequence having from about 5 to 15 nucleotides of (i) is immediately adjacent to the 5' end of the second sequence of (ii), the second sequence of (ii) is immediately adjacent to the 5' end of the protospacer sequence of (iii), and the protospacer sequence is immediately adjacent to the 5' end of the third sequence of (iv); the target oligonucleotide (second) strand is complementary to the non-target oligonucleotide strand; and (b) the non-target oligonucleotide strand anneals to the complementary target oligonucleotide strand to produce a double-stranded nucleic acid molecule, wherein the first sequence comprises a restriction site (at its 5' end) and the third sequence comprises a restriction site (at its 3' end), wherein the first sequence (i), protospacer sequence (iii) and third sequence (iv) of each of the two or more double-stranded nucleic acid molecules are identical.

[0015] The tenth aspect of the present invention provides a randomized DNA library for determining the requirements / specificity of the protospacer adjacent motif (PAM) of a CRISPR-Cas nuclease having a PAM recognition site at the 3'-end of the protospacer. The randomized DNA library comprises two or more double-stranded nucleic acid molecules, each comprising: (a) a non-target oligonucleotide (first) strand and a target oligonucleotide (second) strand, wherein the non-target oligonucleotide strand, from 5' to 3', (i) a first sequence having about 5 to about 20 nucleotides (e.g., about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 nucleotides, and any range or value therein), (ii) a protospacer sequence comprising about 16 to about 25 nucleotides, e.g., about 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 nucleotides, (iii) a second sequence having at least 4 randomized nucleotides (e.g., at least 4, 5, 6, 7, 8, 9, 10, or more, and any range or value therein), and (iv) a third sequence having about 5 to about 15 nucleotides (e.g., about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15 nucleotides, and any range or value therein), wherein the first sequence having about 5 to 20 nucleotides in (i) is immediately adjacent to the 5'-end of the protospacer sequence in (ii), the second sequence in (iii) is immediately adjacent to the 3'-end of the protospacer sequence in (iii), and the third sequence in (iv) is immediately adjacent to the 3'-end of the second sequence in (iii); the target oligonucleotide (second) strand is complementary to the non-target oligonucleotide strand; and (b) the non-target oligonucleotide strand anneals to the complementary target oligonucleotide strand to produce a double-stranded nucleic acid molecule, wherein the first sequence contains a restriction site (at its 5'-end) and the third sequence contains a restriction site (at its 3'-end), and the first sequence (i), protospacer sequence (ii) and third sequence (iv) of each of the two or more double-stranded nucleic acid molecules are identical.

[0016] The present invention further provides an expression cassette and / or vector comprising a polynucleotide encoding the CRISPR-Cas nuclease and / or fusion protein of the present invention, and / or a cell comprising the polynucleotide, polypeptide and / or fusion protein of the present invention, and / or a kit comprising the same.

[0017] These and other aspects of the present invention are shown in more detail in the following description of the present invention.

[0018] [Brief Description of the Sequences] SEQ ID NOs: 1 to 17, 49, 50 and 51 are exemplary nucleotide sequences encoding Cas12a nuclease. SEQ ID NOs: 18 to 22 are exemplary adenosine deaminases. SEQ ID NOs: 23 to 25 and SEQ ID NOs: 42 to 48 are exemplary cytosine deaminases. SEQ ID NO: 26 is an exemplary nucleotide sequence encoding uracil-DNA glycosylase inhibitor (UGI). SEQ ID NOs: 27 to 29 provide an example of protospacer adjacent motif positions for type V CRISPR-Cas12a nuclease. SEQ ID NOs: 30 to 39 show examples of nucleotide sequences useful for producing the randomized library of the present invention for use, for example, in in vitro cleavage assays. SEQ ID NOs: 40 to 41 are exemplary regulatory sequences encoding a promoter and an intron. SEQ ID NO: 52 provides the nucleotide sequence of an exemplary expression cassette. SEQ ID NO: 53 provides the nucleotide sequence of an exemplary vector. SEQ ID NOs: 54 to 61 provide exemplary spacer sequences. SEQ ID NO: 62 provides an exemplary CRISPR RNA. [Brief Description of the Drawings]

[0019]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figures 22A - 22B

Mode for Carrying Out the Invention

[0020] The present invention will be described below with respect to the accompanying drawings and examples in which embodiments of the present invention are shown. This description is not intended to be a catalog of all the different ways in which the invention can be practiced or of all the features that can be added to the invention. For example, features illustrated with respect to one embodiment may be incorporated into other embodiments, and features illustrated with respect to a particular embodiment may be deleted from that embodiment. Thus, the present invention is considered to be able to exclude or omit any feature or combination of features shown herein in some embodiments of the present invention. Further, numerous variations and additions to the various embodiments suggested herein will be apparent to those of ordinary skill in the art in light of the present disclosure and do not depart from the present invention. Therefore, the following description is intended to illustrate some particular embodiments of the present invention and is not intended to specifically identify all permutations, combinations, and variations thereof.

[0021] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. In this specification, the technical terms used in the description of the present invention are for the purpose of describing particular embodiments only and are not intended to limit the present invention.

[0022] All publications, patent applications, patents, and other references cited in this specification are hereby incorporated by reference in their entirety for all teachings relevant to the passages and / or paragraphs in which they are indicated.

[0023] Unless the context dictates otherwise, it is clearly intended that the various features of the invention described herein can be used in any combination. Further, it is contemplated that in some embodiments of the invention, any feature or combination of features shown herein can be excluded or omitted. By way of example, if the specification describes a composition as including components A, B, and C, it is clearly intended that any one of A, B, or C, or combinations thereof, can be omitted and negated, either singly or in any combination.

[0024] The singular forms "a", "an", and "the" as used in the description of the invention and the appended claims are intended to include the plural forms as well, unless the context clearly dictates otherwise.

[0025] Also, as used herein, "and / or" includes any and all possible combinations of one or more of the associated listed items, as well as, when interpreted selectively (as "or"), the absence of a combination, and encompasses.

[0026] As used herein, the term "about", when referring to a measurable value such as an amount or concentration, means including variations of ±10%, ±5%, ±1%, ±0.5%, or even ±0.1% of the specified value, as well as the specified value. For example, "about X" means X and variations of ±10%, ±5%, ±1%, ±0.5%, or even ±0.1% of X, when X is a measurable value. The ranges provided herein for measurable values can include any other range and / or individual value therein.

[0027] As used herein, phrases such as "between X and Y" and "between about X and Y" should be construed to include X and Y. Phrases such as "between about X and Y" as used herein mean "between about X and about Y", and phrases such as "from X to Y" mean "from about X to about Y".

[0028] The recitation of a range of values herein is merely intended to serve as a shorthand method of referring individually to each separate value falling within the range, and each separate value is incorporated herein as if it were individually recited herein. For example, if the range 10-15 is disclosed, then 11, 12, 13, and 14 are also disclosed.

[0029] The terms "comprise", "comprises" and "comprising" as used herein identify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0030] The transitional phrase "consisting essentially of" as used herein means that a claim is to be construed to cover the specified materials or steps recited in the claim and those that do not materially affect the basic and novel characteristics (singular or plural) of the invention recited in the claim. Thus, the term "consisting essentially of" is not intended to be construed as equivalent to "comprising" when used in the claims of the present invention.

[0031] As used herein, the terms "increase", "increasing", "enhance", "enhancing", "improve" and "improving" (and their grammatical variations) describe an increase of at least about 25%, 50%, 75%, 100%, 150%, 200%, 300%, 400%, 500% or more as compared to a control.

[0032] As used herein, the terms "reduce", "reduced", "reducing", "reduction", "diminish" and "decrease" (and their grammatical variations) describe, for example, a reduction of at least about 5%, 10%, 15%, 20%, 25%, 35%, 50%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% as compared to a control. In certain embodiments, the reduction cannot, or essentially cannot, result in a detectable activity or amount (i.e., a trivial amount, e.g., less than about 10% or indeed 5%).

[0033] A "heterologous" or "recombinant" nucleotide sequence is a nucleotide sequence that is not naturally associated with the host cell into which it is introduced and includes multiple copies of a non-naturally occurring origin of a naturally occurring nucleotide sequence.

[0034] A "native" or "wild-type" nucleic acid, nucleotide sequence, polypeptide or amino acid sequence refers to a nucleic acid, nucleotide sequence, polypeptide or amino acid sequence of natural origin or endogenous origin. Thus, for example, a "wild-type mRNA" is an mRNA of natural origin or endogenous to an organism. A "homologous" nucleic acid sequence is a nucleotide sequence that is naturally associated with the host cell into which it is introduced.

[0035] As used herein, the terms "nucleic acid", "nucleic acid molecule", "nucleotide sequence", and "polynucleotide" refer to RNA or DNA that is linear or branched, single-stranded or double-stranded, or hybrids thereof. The term also encompasses RNA / DNA hybrids. When dsRNA is synthetically produced, less common bases such as inosine, 5-methylcytosine, 6-methyladenine, hypoxanthine, and others can be used for the pairing of antisense, dsRNA, and ribozymes. For example, polynucleotides containing C-5 propyne analogs of uridine and cytidine have been shown to bind to RNA with high affinity and to be potent antisense inhibitors of gene expression. Other modifications, such as modifications to the phosphodiester backbone of RNA or to the 2'-hydroxy in the ribose sugar moiety, can also be made.

[0036] As used herein, the term "nucleotide sequence" refers to a heteropolymer of nucleotides or the sequence of these nucleotides from the 5'-end to the 3'-end of a nucleic acid molecule, including DNA or RNA molecules such as cDNA, DNA fragments or portions, genomic DNA, synthetic (e.g., chemically synthesized) DNA, plasmid DNA, mRNA, and antisense RNA, any of which may be single-stranded or double-stranded. The terms "nucleotide sequence", "nucleic acid", "nucleic acid molecule", "nucleic acid construct", "oligonucleotide" and "polynucleotide" are also used interchangeably herein to refer to a heteropolymer of nucleotides. The nucleic acid molecules and / or nucleotide sequences provided herein are shown herein in the 5'- to 3'-direction, from left to right, using the standard code to represent the nucleotide symbols described in the U.S. sequence rules, 37 CFR §§ 1.821 - 1.825 and the World Intellectual Property Organization (WIPO) standard ST.25. As used herein, the "5'-region" can mean the region of the polynucleotide closest to the 5'-end of the polynucleotide. Thus, for example, an element within the 5'-region of a polynucleotide can be located anywhere from the first nucleotide located at the 5'-end of the polynucleotide to a nucleotide located in the middle of the polynucleotide. As used herein, the "3'-region" can mean the region of the polynucleotide closest to the 3'-end of the polynucleotide. Thus, for example, an element within the 3'-region of a polynucleotide can be located anywhere from the first nucleotide located at the 3'-end of the polynucleotide to a nucleotide located in the middle of the polynucleotide.

[0037] As used herein, the term "gene" refers to a nucleic acid molecule that can be used to produce mRNA, antisense RNA, miRNA, anti-microRNA antisense oligodeoxynucleotides (AMO), etc. A gene may or may not be capable of being used to produce a functional protein or gene product. A gene can include both coding and non-coding regions (e.g., introns, regulatory elements, promoters, enhancers, termination sequences and / or 5' and 3' untranslated regions). A gene may be "isolated", thereby meaning a nucleic acid that is substantially or essentially free from components that are normally found associated with the nucleic acid in its natural state. Such components include other cellular materials, culture media from recombinant production, and / or various chemicals used in the chemical synthesis of nucleic acids.

[0038] The term "mutation" refers to a point mutation (e.g., a missense, or nonsense, or single base pair insertion or deletion that causes a frameshift), an insertion, a deletion, and / or a truncation. When a mutation is a substitution of one residue in an amino acid sequence by another residue, or a deletion or insertion of one or more residues in the sequence, the mutation is typically described by specifying, after the original residue, the position of that residue in the sequence and the identity of the newly substituted residue.

[0039] As used herein, the terms "complementary" or "complementarity" refer to the natural binding of polynucleotides by base pairing under permissive salt and temperature conditions. For example, the sequence "A-G-T" (5' to 3') binds to the complementary sequence "T-C-A" (3' to 5'). Complementarity between two single-stranded molecules may be "partial", where only a portion of the nucleotides bind, or "complete", where there is overall complementarity between the single-stranded molecules. The degree of complementarity between nucleic acid strands has a significant effect on the efficiency and strength of hybridization between the nucleic acid strands.

[0040] As used herein, "complement" can mean 100% complementarity with a comparator nucleotide sequence, or it can mean less than 100% complementarity (e.g., complementarity such as about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%).

[0041] A "portion" or "fragment" of a nucleotide sequence of the present invention is a nucleotide sequence of reduced length compared to a reference nucleic acid or nucleotide sequence (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more nucleotides are reduced), and is identical or substantially identical (e.g., 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% identical) to the reference nucleic acid or nucleotide sequence of consecutive nucleotides, and is understood to mean consisting essentially of and / or consisting of. Such nucleic acid fragments or portions according to the present invention may, where appropriate, be contained within a larger polynucleotide of which they are a component. As an example, the repeat sequence of the guide nucleic acid of the present invention may include a portion of a wild-type CRISPR-Cas repeat sequence (e.g., wild-type Cas9 repeat, wild-type Cas12a repeat, etc.).

[0042] Different nucleic acids or proteins having identity are referred to herein as "homologs". The term "homolog" includes homologous sequences from the same and other species and orthologous sequences from the same and other species. "Identity" refers to the level of similarity between two or more nucleic acid and / or amino acid sequences from the perspective of the percentage of positional identity (i.e., sequence similarity or identity). Identity also refers to the concept of similar functional properties between different nucleic acids or proteins. Accordingly, the compositions and methods of the present invention further include homologs to the nucleotide sequences and polypeptide sequences of the present invention. As used herein, "orthologous" refers to homologous nucleotide sequences and / or amino acid sequences in different species that arose from a common ancestral gene during speciation. Homologs of the nucleotide sequences of the present invention have substantial sequence identity (e.g., at least about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5% or 100%) with the above nucleotide sequences of the present invention.

[0043] As used herein, "sequence identity" refers to the degree to which two optimally aligned polynucleotide or polypeptide sequences are invariant throughout a window of alignment of components (e.g., nucleotides or amino acids). "Identity" can be readily calculated by known methods including, but not limited to, those described in Computational Molecular Biology (Lesk, A.M., ed.) Oxford University Press, New York (1988); Biocomputing: Informatics and Genome Projects (Smith, D.W., ed.) Academic Press, New York (1993); Computer Analysis of Sequence Data, Part I (Griffin, A.M., and Griffin, H.G., eds.) Humana Press, New Jersey (1994); Sequence Analysis in Molecular Biology (von Heinje, G., ed.) Academic Press (1987); and Sequence Analysis Primer (Gribskov, M. and Devereux, J., eds.) Stockton Press, New York (1991).

[0044] As used herein, the term "percent sequence identity" or "percent identity" refers to the percentage of identical nucleotides in the linear polynucleotide sequence of a reference ("query") polynucleotide molecule (or its complementary strand) compared to a test ("subject") polynucleotide molecule (or its complementary strand) when the two sequences are optimally aligned. In some embodiments, "percent identity" can refer to the percentage of identical amino acids in an amino acid sequence compared to a reference polypeptide.

[0045] As used herein, the terms "substantially identical" or "substantial identity" in the context of two nucleic acid molecules, nucleotide sequences or protein sequences refer to two or more sequences or subsequences that have at least about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5% or 100% nucleotide or amino acid residue identity when compared and aligned for maximum correspondence using one of the following sequence comparison algorithms or by visual inspection. In some embodiments of the invention, substantial identity exists over a region of contiguous nucleotides of a nucleotide sequence of the invention that is about 10 nucleotides to about 20 nucleotides, about 10 nucleotides to about 25 nucleotides, about 10 nucleotides to about 30 nucleotides, about 15 nucleotides to about 25 nucleotides, about 30 nucleotides to about 40 nucleotides, about 50 nucleotides to about 60 nucleotides, about 70 nucleotides to about 80 nucleotides, about 90 nucleotides to about 100 nucleotides, or more nucleotides in length, and any range therein (up to the full length of the sequence). In some embodiments, the nucleotide sequence can be substantially identical over at least about 20 nucleotides (e.g., about 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40 nucleotides). In some embodiments, substantially identical nucleotide or protein sequences perform substantially the same function as the nucleotides (or encoded protein sequences) that are substantially identical.

[0046] For array comparison, typically one array serves as the reference array against which the test array is compared. When using an array comparison algorithm, the test array and the reference array are input into a computer, and sub-array coordinates are specified as needed, and the program parameters of the array algorithm are specified. Then, the array comparison algorithm calculates the percentage of array identity for the test array(s) compared to the reference array based on the specified program parameters.

[0047] Optimal array alignments for arranging comparison windows are well known to those skilled in the art and may be performed by tools such as the local homology algorithm of Smith and Waterman, the homology alignment algorithm of Needleman and Wunsch, the similarity search method of Pearson and Lipman, and may be by computerized execution of algorithms such as GAP, BESTFIT, FASTA, and TFASTA available as part of the GCG® Wisconsin Package® (Accelrys Inc., San Diego, CA). The "identity fraction" for an aligned segment of a test array and a reference array is the number of identical components shared by the two aligned arrays divided by the total number of components in the reference array segment (i.e., the entire reference array or a smaller defined portion of the reference array). The percentage of array identity is expressed as the identity fraction multiplied by 100. The comparison of one or more polynucleotide sequences may be to the full-length polynucleotide sequence or a portion thereof, or to a longer polynucleotide sequence. For purposes of the present invention, the "percentage of identity" may be determined using BLASTX version 2.0 for translated nucleotide sequences and BLASTN version 2.0 for polynucleotide sequences.

[0048] Two nucleotide sequences may be considered to be substantially complementary when the two sequences hybridize to each other under stringent conditions. In some representative embodiments, two nucleotide sequences that are considered to be substantially complementary hybridize to each other under very stringent conditions.

[0049] "Stringent hybridization conditions" and "stringent hybridization wash conditions" in the context of nucleic acid hybridization experiments such as Southern and Northern hybridizations are sequence-dependent and vary under different environmental parameters. Extensive guidance regarding nucleic acid hybridization can be found in Tijssen Laboratory Techniques in Biochemistry and Molecular Biology-Hybridization with Nucleic Acid Probes part I chapter 2 "Overview of principles of hybridization and the strategy of nucleic acid probe assays" Elsevier, New York (1993). Generally, very stringent hybridization and wash conditions are selected to be about 5 °C lower than the thermal melting point (T m ) for a specific sequence at a defined ionic strength and pH.

[0050] T m is the temperature at which 50% of the target sequence hybridizes to a perfectly matching probe (at a defined ionic strength and pH). Very stringent conditions are T mis selected to be equivalent. An example of stringent hybridization conditions for hybridization of a complementary nucleotide sequence having more than 100 complementary residues on a filter in a Southern or Northern blot is 50% formamide containing 1 mg of heparin at 42° C., with hybridization carried out overnight. An example of very stringent washing conditions is 0.15 M NaCl at 72° C. for about 15 minutes. An example of stringent washing conditions is washing with 0.2×SSC at 65° C. for 15 minutes (see Sambrook, below, for an explanation of the SSC buffer). Often, low stringency washing is carried out prior to high stringency washing to remove background probe signal. An example of medium stringency washing for a duplex of more than 100 nucleotides is 1×SSC at 45° C. for 15 minutes. An example of low stringency washing for a duplex of more than 100 nucleotides is 4-6×SSC at 40° C. for 15 minutes. For short probes (e.g., about 10-50 nucleotides), stringent conditions typically include a salt concentration of less than about 1.0 M Na ions, typically about 0.01-1.0 M Na ion concentration (or other salts) (pH 7.0-8.3), and the temperature is typically at least about 30° C. Stringent conditions can also be achieved by the addition of destabilizing agents such as formamide. Generally, in a particular hybridization assay, a signal to noise ratio that is 2-fold (or higher) than that observed for unrelated probes indicates detection of specific hybridization. Nucleotide sequences that do not hybridize to each other under stringent conditions are substantially identical if the proteins they encode are substantially identical. This can occur, for example, when copies of the nucleotide sequence are made using the maximum codon degeneracy allowed by the genetic code.

[0051] Any nucleotide sequence, polynucleotide, and / or recombinant nucleic acid construct of the present invention can be codon-optimized for expression in any organism of interest. Codon optimization is well known in the art and involves modifying the nucleotide sequence with respect to codon usage bias using a species-specific codon usage table. The codon usage table is created based on sequence analysis of the most highly expressed genes for the organism / species of interest. When the nucleotide sequence is expressed in the nucleus, the codon usage table is created based on sequence analysis of nuclear genes that are highly expressed for the species of interest. Modification of the nucleotide sequence is determined by comparing the codons present in the native polynucleotide sequence with the species-specific codon usage table. As understood in the art, codon optimization of a nucleotide sequence results in a nucleotide sequence having less than 100% (e.g., about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%, and any range or value therein) identity to the native polynucleotide sequence, but encodes a polypeptide having the same function as that encoded by the original, native nucleotide sequence. Thus, in some embodiments of the present invention, the polynucleotides, nucleic acid constructs, expression cassettes, and / or vectors of the present invention (e.g., those comprising / encoding the polypeptides, fusion proteins, complexes of the present invention, such as modified CRISPR-Cas nucleases) are codon-optimized for expression in a particular organism of interest, e.g., a particular plant species, a particular bacterial species, a particular animal species, etc.In some embodiments, the codon-optimized nucleic acid constructs, polynucleotides, expression cassettes, and / or vectors of the invention have about 70% to about 99.9% (e.g., about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.9% or 100%) or higher identity to the non-codon-optimized polynucleotides, nucleic acid constructs, expression cassettes, and / or vectors of the invention.

[0052] In any of the embodiments described herein, the polynucleotides or nucleic acid constructs of the invention can be operably associated with various promoters and / or other regulatory elements for expression in plants and / or plant cells. Thus, in some embodiments, the polynucleotides or nucleic acid constructs of the invention can further comprise one or more promoters, introns, enhancers, and / or terminators operably linked to one or more nucleotide sequences. In some embodiments, a promoter can be operably associated with an intron (e.g., the Ubi1 promoter and intron). In some embodiments, a promoter associated with an intron can be referred to as a "promoter region" (e.g., the Ubi1 promoter and intron).

[0053] As used herein in reference to polynucleotides, "operably linked" or "operably associated" means that the indicated elements are functionally related to each other and generally also physically related. Thus, the terms "operably linked" or "operably associated" as used herein refer to nucleotide sequences on a single nucleic acid molecule that are functionally related. Thus, a first nucleotide sequence operably linked to a second nucleotide sequence means that the first nucleotide sequence is arranged in a functional relationship with the second nucleotide sequence. For example, if a promoter has an effect on the transcription or expression of a nucleotide sequence, the promoter is operably associated with the nucleotide sequence. One of ordinary skill in the art understands that a control sequence (e.g., a promoter) need not be contiguous with a nucleotide sequence with which it is operably associated as long as the control sequence functions to direct expression. Thus, for example, a nucleic acid sequence that is transcribed but not translated and that is present between a promoter and a nucleotide sequence may be present, and the promoter can still be considered to be "operably linked" to the nucleotide sequence.

[0054] As used herein in reference to polypeptides, the term "linked" refers to one polypeptide being attached to another polypeptide. A polypeptide may be linked to another polypeptide directly (e.g., via a peptide bond) at the N-terminus or C-terminus or via a linker.

[0055] The term "linker" is recognized in the art and refers to a chemical group or molecule that links two molecules or moieties, such as two domains of a fusion protein, such as an LbCas12a CRISPR-Cas nuclease domain and a polypeptide of interest (e.g., a nucleic acid editing domain, a deaminase domain, adenosine deaminase, cytosine deaminase), etc. A linker may be composed of a single linking molecule or may contain more than one linking molecule. In some embodiments, the linker can be an organic molecule, group, polymer, or chemical moiety, such as a divalent organic moiety. In some embodiments, the linker can be an amino acid or a peptide.In some embodiments, the peptide linker may be from about 4 to about 100 or more amino acids in length, such as about 4, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100 or more amino acids in length (e.g., from about 4 to about 40, from about 4 to about 50, from about 4 to about 60, from about 5 to about 40, from about 5 to about 50, from about 5 to about 60, from about 9 to about 40, from about 9 to about 50, from about 9 to about 60, from about 10 to about 40, from about 10 to about 50, from about 10 to about 60, or from about 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25 amino acids to about 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100 or more amino acids in length). In some embodiments, the peptide linker may be a GS linker.

[0056] A "promoter" is a nucleotide sequence that controls or regulates the transcription of a nucleotide sequence (such as a coding sequence) that is operably linked to the promoter. The coding sequence controlled or regulated by the promoter can encode a polypeptide and / or a functional RNA. Typically, a "promoter" refers to a nucleotide sequence that contains a binding site for RNA polymerase II and directs the initiation of transcription. Generally, a promoter is found 5', i.e., upstream, of the start of the coding region of the corresponding coding sequence. A promoter may include regulatory factors for gene expression; for example, other elements that act as promoter regions. These include the TATA box consensus sequence and often the CAAT box consensus sequence (Breathnach and Chambon, (1981) Annu. Rev. Biochem. 50:349). In plants, the CAAT box can be replaced by the AGGA box (Messing et al., (1983) in Genetic Engineering of Plants, T. Kosuge, C. Meredith and A. Hollaender (eds.), Plenum Press, pp. 211-227). In some embodiments, the promoter region may include at least one intron (e.g., SEQ ID NO: 40 or SEQ ID NO: 41).

[0057] Promoters useful in the present invention can include, for example, constitutive, inducible, temporally regulated, developmentally regulated, chemically regulated, tissue-preferred and / or tissue-specific promoters for use in the preparation of recombinant nucleic acid molecules, such as "synthetic nucleic acid constructs" or "protein-RNA complexes". These various types of promoters are known in the art.

[0058] The selection of a promoter may vary according to the temporal and spatial requirements of expression and may also vary based on the host cell being transformed. Promoters for many different organisms are well known in the art. Based on the extensive knowledge present in the field, an appropriate promoter can be selected for a particular target host organism. Thus, for example, much is known about the promoters upstream of genes that are very constitutively expressed in model organisms, and such knowledge can be readily accessed and implemented in other systems as appropriate.

[0059] In some embodiments, promoters functional in plants can be used with the constructs of the invention. Non-limiting examples of promoters useful for driving expression in plants include the promoter of the ribulose bisphosphate carboxylase small subunit gene 1 (PrbcS1), the promoter of the actin gene (Pactin), the promoter of the nitrate reductase gene (Pnr), and the promoter of the duplicated carbonic anhydrase gene 1 (Pdca1) (see Walker et al. Plant Cell Rep. 23:727-735 (2005); Li et al. Gene 403:132-142 (2007); Li et al. Mol Biol Rep. 37:1143-1154 (2010)). PrbcS1 and Pactin are constitutive promoters, and Pnr and Pdca1 are inducible promoters. Pnr is induced by nitrate and repressed by ammonium (Li et al. Gene 403:132-142 (2007)), and Pdca1 is induced by salt (Li et al. Mol Biol Rep. 37:1143-1154 (2010)).

[0060] Examples of constitutive promoters useful in plants include, but are not limited to, the cestrum virus promoter (cmp) (U.S. Patent No. 7,166,770), the rice actin 1 promoter (Wang et al. (1992) Mol. Cell. Biol. 12:3399-3406; and U.S. Patent No. 5,641,876), the CaMV 35S promoter (Odell et al. (1985) Nature 313:810-812), the CaMV 19S promoter (Lawton et al. (1987) Plant Mol. Biol. 9:315-324), the nos promoter (Ebert et al. (1987) Proc. Natl. Acad. Sci USA 84:5745-5749), the Adh promoter (Walker et al. (1987) Proc. Natl. Acad. Sci. USA 84:6624-6629), the sucrose synthase promoter (Yang & Russell (1990) Proc. Natl. Acad. Sci. USA 87:4144-4148), and the ubiquitin promoter. Constitutive promoters derived from ubiquitin accumulate in many cell types. Ubiquitin promoters have been cloned from several plant species for use in transgenic plants (e.g., sunflower (Binet et al., 1991. Plant Science 79:87-94), maize (Christensen et al., 1989. Plant Molec. Biol. 12:619-632), and Arabidopsis thaliana (Norris et al. 1993. Plant Molec. Biol. 21:895-906)). The maize ubiquitin promoter (UbiP) has been developed in transgenic monocot systems, and the vector constructed for its sequence and monocot transformation is disclosed in patent publication EP0342926. The ubiquitin promoter is suitable for the expression of the nucleotide sequences of the present invention in transgenic plants, particularly monocots.Furthermore, the promoter expression cassette described by McElroy et al. (Mol. Gen. Genet. 231: 150-160 (1991)) can be readily modified for expression of the nucleotide sequences of the present invention and is particularly suitable for use in monocotyledonous hosts.

[0061] In some embodiments, a tissue-specific / tissue-preferred promoter can be used for the expression of a heterologous polynucleotide in a plant cell. Tissue-specific or preferred expression patterns include, but are not limited to, green tissue-specific or preferred, root-specific or preferred, stem-specific or preferred, flower-specific or preferred, or pollen-specific or preferred. Promoters suitable for expression in green tissue include many that regulate genes involved in photosynthesis, many of which have been cloned from both monocotyledonous and dicotyledonous plants. In one embodiment, a promoter useful in the present invention is the maize PEPC promoter from the phosphoenolpyruvate carboxylase gene (Hudspeth & Grula, Plant Molec. Biol. 12:579-589 (1989)). Non-limiting examples of tissue-specific promoters include those associated with genes encoding seed storage proteins (e.g., β-conglycinin, cruciferin, napin, and phaseolin), zein or oleosin (e.g., oleosin), or proteins involved in fatty acid biosynthesis (including acyl carrier protein, stearoyl-ACP desaturase, and fatty acid desaturase (fad2-1)), and other nucleic acids expressed during embryo development (e.g., Bce4, see, e.g., Kridl et al. (1991) Seed Sci. Res. 1:209-219; and European Patent No. 255378). Tissue-specific or tissue-preferred promoters useful for the expression of the nucleotide sequences of the present invention in plants, particularly maize, include, but are not limited to, those directing expression to roots, pith, leaves, or pollen. Such promoters are disclosed, for example, in WO93 / 07278, which is hereby incorporated by reference in its entirety.Other non-limiting examples of tissue-specific or tissue-preferred promoters useful in the present invention include the cotton rubisco promoter disclosed in U.S. Patent No. 6,040,504; the maize sucrose synthase promoter disclosed in U.S. Patent No. 5,604,121; the root-specific promoter described by de Framond (FEBS 290:103-106 (1991); EP0452269 (Ciba-Geigy)); the stem-specific promoter described in U.S. Patent No. 5,625,136 (Ciba-Geigy) that drives the expression of the maize trpA gene; the cestrum yellow leaf curling virus promoter disclosed in WO01 / 73087; and, without limitation, ProOsLPS10 and ProOsLPS11 from rice (Nguyen et al. Plant Biotechnol. Reports 9(5):297-306 (2015)), ZmSTK2_USP from maize (Wang et al. Genome 60(6):485-495 (2017)), LAT52 and LAT59 from tomato (Twell et al. Development 109(3):705-713 (1990)), Zm13 (U.S. Patent No. 10,421,972), and PLA from Arabidopsis thaliana. 2 -δ promoter (U.S. Patent No. 7,141,424), and / or a pollen-specific or preferred promoter comprising the ZmC5 promoter from maize (International PCT Publication No. WO1999 / 042587).

[0062] Further examples of plant tissue-specific / tissue-preferred promoters include, but are not limited to, the root hair-specific cis-element (RHE) (Kim et al. The Plant Cell 18:2958-2970 (2006)), the root-specific promoter RCc3 (Jeong et al. Plant Physiol. 153:185-197 (2010)) and RB7 (U.S. Patent No. 5,459,252), the lectin promoter (Lindstrom et al. (1990) Der. Genet. 11:160-167; and Vodkin (1983) Prog. Clin. Biol. Res. 138:87-98), the maize alcohol dehydrogenase 1 promoter (Dennis et al. (1984) Nucleic Acids Res. 12:3983-4000), S-adenosyl-L-methionine synthase (SAMS) (Vander Mijnsbrugge et al. (1996) Plant and Cell Physiology, 37(8):1108-1115), the maize light-harvesting complex promoter (Bansal et al. (1992) Proc. Natl. Acad. Sci. USA 89:3654-3658), the maize heat shock protein promoter (O’Dell et al. (1985) EMBO J. 5:451-458; and Rochester et al. (1986) EMBO J. 5:451-458), the pea small subunit RuBP carboxylase promoter (Cashmore, “Nuclear genes encoding the small subunit of ribulose-l,5-bisphosphate carboxylase” pp. 29-39 In: Genetic Engineering of Plants (Hollaender ed., Plenum Press 1983; and, Poulsen et al. (1986) Mol. Gen. Genet. 205:193-200), the mannopine synthase promoter of the Ti plasmid (Langridge et al. (1989) Proc. Natl. Acad. Sci. USA 86:3219-3223), the nopaline synthase promoter of the Ti plasmid (Langridge et al.(1989), supra), the petunia chalcone isomerase promoter (van Tunen et al. (1988) EMBO J. 7:1257-1263), the soybean glycine-rich protein 1 promoter (Keller et al. (1989) Genes Dev. 3:1639-1646), the truncated CaMV 35S promoter (O’Dell et al. (1985) Nature 313:810-812), the potato patatin promoter (Wenzler et al. (1989) Plant Mol. Biol. 13:347-354), the root cell promoter (Yamamoto et al. (1990) Nucleic Acids Res. 18:7449), the maize zein promoter (Kriz et al. (1987) Mol. Gen. Genet. 207:90-98; Langridge et al. (1983) Cell 34:1015-1022; Reina et al. (1990) Nucleic Acids Res. 18:6425; Reina et al. (1990) Nucleic Acids Res. 18:7449; and Wandelt et al. (1989) Nucleic Acids Res. 17:2354), the globulin-1 promoter (Belanger et al. (1991) Genetics 129:863-872), the α-tubulin cab promoter (Sullivan et al. (1989) Mol. Gen. Genet. 215:431-440), the PEPCase promoter (Hudspeth & Grula (1989) Plant Mol. Biol. 12:579-589), the R gene complex-related promoter (Chandler et al. (1989) Plant Cell 1:1175-1183), and the chalcone synthase promoter (Franken et al. (1991) EMBO J. 10:2605-2612).

[0063] Those useful for tissue-specific expression are the pea vicilin promoter (Czako et al. (1992) Mol. Gen. Genet. 235:33-40; and the tissue-specific promoter disclosed in U.S. Patent No. 5,625,136). Promoters useful for expression in mature leaves are those switched on during senescence, for example, the SAG promoter from Arabidopsis (Gan et al. (1995) Science 270:1986-1988).

[0064] Furthermore, promoters functional in chloroplasts can be used. Non-limiting examples of such promoters include the bacteriophage T3 gene 9 5’UTR and other promoters disclosed in U.S. Patent No. 7,579,516. Other promoters useful in the present invention include, but are not limited to, the S-E9 small subunit RuBP carboxylase promoter and the Kunitz trypsin inhibitor gene promoter (Kti3).

[0065] Additional regulatory elements useful in the present invention include, but are not limited to, introns, enhancers, termination sequences and / or 5’ and 3’ untranslated regions.

[0066] The introns useful in the present invention may be introns identified and isolated in plants and are inserted into an expression cassette used in plant transformation. As will be understood by those skilled in the art, introns can contain the sequences necessary for self-excision, which are incorporated in-frame within the nucleic acid construct / expression cassette. Introns can be used as spacers to separate multiple protein-coding sequences within one nucleic acid construct, or introns can be used, for example, to stabilize mRNA within one protein-coding sequence. When used within a protein-coding sequence, they are inserted "in-frame" with the excision sites included. Introns may be associated with a promoter for improving or modifying expression. As an example, promoter / intron combinations useful in the present invention include, but are not limited to, the combination of the maize Ubi1 promoter and intron.

[0067] Non-limiting examples of introns useful in the present invention include introns derived from the ADHI gene (e.g., Adh1-S introns 1, 2, and 6), introns derived from the ubiquitin gene (Ubi1), introns derived from the rubisco small subunit (rbcS) gene, introns derived from the rubisco large subunit (rbcL) gene, introns derived from the actin gene (e.g., actin-1 intron), introns derived from the pyruvate dehydrogenase kinase gene (pdk), introns derived from the nitrate reductase gene (nr), introns derived from the double carbonic anhydrase gene 1 (Tdca1), introns derived from the psbA gene, introns derived from the atpA gene, or any combination thereof. As a non-limiting example, the nucleic acid construct of the present invention may encode a base editor comprising an optimized CRISPR-Cas nuclease (e.g., SEQ ID NOs: 1-11 or 23-25) and a deaminase, wherein the nucleic acid construct further comprises a promoter comprising and / or associated with an intron. As a further non-limiting example, the nucleic acid construct of the present invention may encode a base editor comprising an optimized CRISPR-Cas nuclease (e.g., SEQ ID NOs: 1-11 or 23-25) and a deaminase, wherein the nuclease and / or deaminase comprises one or more introns, and the nucleic acid construct may further comprise a promoter comprising and / or associated with an intron.

[0068] In some embodiments, the polynucleotide and / or nucleic acid construct of the present invention may be an "expression cassette" or may be included within an expression cassette. As used herein, an "expression cassette" means, for example, a recombinant nucleic acid molecule comprising a nucleic acid construct of the present invention (e.g., encoding a modified LbCas12a of the present invention), wherein the nucleic acid construct is operably associated with at least a control sequence (e.g., a promoter). Thus, some embodiments of the present invention provide an expression cassette designed to express, for example, a nucleic acid construct of the present invention (e.g., a nucleic acid construct of the present invention encoding a modified LbCas12a of the present invention).

[0069] An expression cassette containing a nucleic acid construct of the invention may be chimeric, meaning that at least one of its components is heterologous to at least one of the other components (e.g., a promoter from a host organism operably linked to a polynucleotide of interest expressed in the host organism, where the polynucleotide of interest is from an organism different from the host or is not normally found in association with its promoter). The expression cassette may be of natural origin, but is preferably obtained recombinantly for heterologous expression.

[0070] The expression cassette can also include a transcriptional and / or translational termination region (i.e., a termination region) and / or an enhancer region that is functional in the selected host cell. A variety of transcriptional terminators and enhancers are known in the art and are available for use in expression cassettes. Transcriptional terminators are responsible for transcriptional termination and proper mRNA polyadenylation. The termination region and / or enhancer region may be native to the transcriptional start region, e.g., native to the gene encoding the LbCas12a nuclease encoded by the nucleic acid construct of the invention, native to the host cell, or native to another source (e.g., foreign or heterologous to the promoter, the gene encoding the LbCas12a nuclease encoded by the nucleic acid construct of the invention, the host cell, or any combination thereof). The enhancer region may be native to the gene encoding the LbCas12a nuclease encoded by the nucleic acid construct of the invention, native to the host cell, or of another origin (e.g., foreign or heterologous to the promoter, the gene encoding the LbCas12a nuclease encoded by the nucleic acid construct of the invention, the host cell, or any combination thereof).

[0071] The expression cassette of the present invention may contain a nucleotide sequence encoding a selectable marker that can be used to select transformed host cells. As used herein, "selectable marker" means a nucleotide sequence that, when expressed, confers a distinct phenotype on host cells expressing the marker, and thus such transformed cells can be distinguished from those that do not have the marker. Such nucleotide sequences can be selected by chemical means, for example, depending on whether the marker confers a characteristic that can be selected by using a selection agent (such as an antibiotic), or whether the marker is a simple characteristic that can be identified through observation or testing, for example, by screening (such as fluorescence), and may encode either a selectable marker or a screenable marker. Many examples of suitable selectable markers are known in the art and can be used in the expression cassettes described herein.

[0072] In addition to the expression cassette, the nucleic acid molecules / constructs and polynucleotide sequences described herein can be used with respect to vectors. The term "vector" refers to a composition for transferring, delivering, or introducing nucleic acid(s) into a cell. A vector includes a nucleic acid construct that contains the nucleotide sequence(s) to be transferred, delivered, or introduced. Vectors for use in the transformation of host organisms are well known in the art. Non-limiting examples of common classes of vectors include viral vectors, plasmid vectors, phage vectors, phagemid vectors, cosmid vectors, fosmid vectors, bacteriophages, artificial chromosomes, minicircles, or Agrobacterium binary vectors of double-stranded or single-stranded linear or circular form that may or may not be self-transmissible or motile. In some embodiments, viral vectors can include, but are not limited to, retroviral, lentiviral, adenoviral, adeno-associated, or herpes simplex virus vectors. The vectors defined herein can transform a prokaryotic or eukaryotic host by integration into the cell genome or can exist episomally (e.g., an autonomously replicating plasmid having an origin of replication). Also included are shuttle vectors, which are DNA vehicles that can replicate natively or by design in two different host organisms selected from actinomycetes and related species, bacteria, and eukaryotes (e.g., higher plants, mammals, yeast, or fungal cells). In some embodiments, the nucleic acid within the vector is under the control of and operably linked to an appropriate promoter or other regulatory element for transcription in the host cell. The vector can be a bifunctional expression vector that functions in multiple hosts. In the case of genomic DNA, this can include its own promoter and / or other regulatory elements, and in the case of cDNA, it can be under the control of an appropriate promoter and / or other regulatory elements for expression in the host cell. Thus, the nucleic acid constructs of the present invention and / or expression cassettes containing them can be included within vectors described herein and known in the art.In some embodiments, the vector may be a high-copy number vector (e.g., a high-copy number E. coli vector; e.g., pUC, pBluescript, pGEM, etc.). Thus, for example, the library of the present invention may be constructed using a high-copy number vector.

[0073] As used herein, "contact", "contacting", "contacted" and their grammatical variations refer to placing the components of a desired reaction together under conditions appropriate to effect the desired reaction (e.g., transformation, transcriptional control, genome editing, nicking, and / or cleavage). Thus, for example, a target nucleic acid can be contacted with (a) a polynucleotide and / or nucleic acid construct of the present invention encoding a modified LbCas12a nuclease of the present invention and (b) a guide nucleic acid under conditions such that the polynucleotide / nucleic acid construct is expressed to produce the modified LbCas12a nuclease, wherein the nuclease forms a complex with the guide nucleic acid and the complex hybridizes to the target nucleic acid, thereby modifying the target nucleic acid. In some embodiments, the target nucleic acid can be contacted with (a) a modified LbCas12a nuclease of the present invention and / or a fusion protein comprising the same (e.g., a modified LbCas12a nuclease of the present invention and a polypeptide of interest (e.g., a deaminase)) and (b) a guide nucleic acid, wherein the modified LbCas12a nuclease forms a complex with the guide nucleic acid and the complex hybridizes to the target nucleic acid, thereby modifying the target nucleic acid. As described herein, the target nucleic acid can be contacted with the polynucleotide / nucleic acid construct / polypeptide of the present invention before, simultaneously with, or after contact with the guide nucleic acid.

[0074] As used herein, "modifying" or "modification" in reference to a target nucleic acid includes editing (e.g., mutation) of nucleic acid / nucleotide bases, covalent modification, exchange / substitution, deletion of the target nucleic acid, cleavage, nicking, and / or transcriptional control.

[0075] In the context of the target polynucleotide, "Introducing", "introduce", "introduced" (and their grammatical variations) mean presenting the target nucleotide sequence (e.g., polynucleotide, nucleic acid construct, and / or guide nucleic acid) to a host organism or a cell of the organism (e.g., host cell; e.g., plant cell) in a manner that the nucleotide sequence has access to the inside of the cell. Thus, for example, the polynucleotides and guide nucleic acids of the present invention encoding the modified LbCas12a nuclease described herein can be introduced into the cells of an organism, thereby transforming the cells using the modified LbCas12a nuclease and the guide nucleic acid.

[0076] As used herein, the term "transformation" refers to the introduction of heterologous nucleic acid into a cell. The transformation of a cell may be stable or transient. Thus, in some embodiments, a host cell or host organism may be stably transformed by the polynucleotide / nucleic acid molecule of the present invention. In some embodiments, a host cell or host organism may be transiently transformed by the polynucleotide / nucleic acid construct of the present invention.

[0077] "Transient transformation" in the context of a polynucleotide means that the polynucleotide is introduced into the cell and not integrated into the genome of the cell.

[0078] "Stably introduce" or "stably introduced" in the context of a polynucleotide introduced into a cell is intended to mean that the introduced polynucleotide is stably integrated into the genome of the cell, and thus the cell is stably transformed by the polynucleotide.

[0079] As used herein, "stable transformation" or "stably transformed" means that a nucleic acid molecule is introduced into a cell and integrated into the genome of the cell. Thus, the integrated nucleic acid molecule can be passed on to its progeny, more specifically, to progeny over multiple generations. As used herein, "genome" includes nuclear and plastid genomes and thus includes, for example, the integration of nucleic acids into the chloroplast or mitochondrial genome. Stable transformation as used herein can also refer to a transgene that is maintained extrachromosomally, for example, as a minichromosome or plasmid.

[0080] Transient transformation can be detected, for example, by enzyme-linked immunosorbent assay (ELISA) or Western blot that can detect the presence of a peptide or polypeptide encoded by one or more transgenes introduced into an organism. Stable transformation of cells can be detected, for example, by Southern blot hybridization assay of genomic DNA of the cells using a nucleic acid sequence that specifically hybridizes to the nucleotide sequence of the transgene introduced into the organism (e.g., a plant). Stable transformation of cells can be detected, for example, by Northern blot hybridization assay of RNA of the cells using a nucleic acid sequence that specifically hybridizes to the nucleotide sequence of the transgene introduced into the host organism. Stable transformation of cells can also be detected, for example, by polymerase chain reaction (PCR) or other amplification reactions well known in the art that result in amplification of the transgene sequence that can be detected according to standard methods using specific primer sequences that hybridize to the target sequence(s) of the transgene. Transformation can also be detected by direct sequencing and / or hybridization protocols well known in the art.

[0081] Accordingly, in some embodiments, the nucleotide sequences, polynucleotides, and / or nucleic acid constructs and / or expression cassettes and / or vectors of the present invention may be transiently expressed and / or they may be stably integrated into the genome of the host organism. Accordingly, in some embodiments, the nucleic acid constructs of the present invention (e.g., the modified LbCas12a nuclease of the present invention or a fusion protein thereof; e.g., a fusion protein encoding a polynucleotide of interest, e.g., a modified LbCas12a nuclease linked to a deaminase domain) (wherein the nucleic acid construct encoding the modified LbCas12a nuclease is codon-optimized for expression in an organism (e.g., a plant, a mammal, a fungus, a bacterium, etc.)) may be transiently introduced into the cells of the organism together with a guide nucleic acid, and thus, the DNA is not maintained within the cell.

[0082] The polynucleotides / nucleic acid constructs of the present invention can be introduced into cells by any method known to those skilled in the art. In some embodiments of the present invention, the transformation of cells includes nuclear transformation. In other embodiments, the transformation of cells includes plastid transformation (e.g., chloroplast transformation). In further embodiments, the polynucleotides / nucleic acid constructs of the present invention can be introduced into cells via conventional breeding techniques.

[0083] Procedures for transforming both eukaryotes and prokaryotes are well known and routine in the art and are described throughout the literature (see, e.g., Jiang et al. 2013. Nat. Biotechnol. 31:233-239; Ran et al. Nature Protocols 8:2281-2308 (2013)).

[0084] Thus, the nucleotide sequence can be introduced into the host organism or its cells by any number of methods well known in the art. The methods of the present invention do not depend on a particular method for introducing one or more nucleotide sequences into an organism, but only on accessing the inside of at least one cell of the organism. When more than one nucleotide sequence is introduced, they can be assembled as part of a single nucleic acid construct or can be located on the same or different nucleic acid constructs as separate nucleic acid constructs. Thus, the nucleotide sequence can be introduced into the target cells in a single transformation event and / or in separate transformations, or, where relevant, the nucleotide sequence can be incorporated into a plant, for example, as part of a propagation protocol.

[0085] The present invention relates to a Cas12a nuclease modified to include a non-native PAM recognition site / sequence (e.g., a Cas12a nuclease that includes, in addition to or instead of the native PAM recognition specificity for that particular Cas12a nuclease, a non-native PAM recognition specificity). Further, the present invention relates to methods for designing, identifying, and selecting Cas12a nucleases having desirable properties including improved PAM recognition specificity.

[0086] As used herein in reference to a modified Cas12a polypeptide, "altered PAM specificity" means that the PAM specificity of the nuclease has been changed from that of the wild-type nuclease (e.g., a non-native PAM sequence is recognized in addition to and / or instead of the native PAM sequence. For example, if a modified Cas12a nuclease recognizes a PAM sequence other than and / or in addition to the native Cas12a PAM sequence of TTTV (V = A, C, or G), its PAM specificity has been altered.

[0087] The present invention relates to an LbCas12a nuclease having a modified PAM recognition specificity. In some embodiments, the present invention provides a modified Lachnospiraceae bacterium CRISPR (Clustered Regularly Interspaced Short Palindromic Repeats) Cas12a (LbCas12a) polypeptide, wherein the modified LbCas12a polypeptide has at least 80% identity (e.g., about 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100% identity; e.g., about 80% to about 100%, about 85% to about 100%, about 90% to about 100%, about 95% to about 100%) with the amino acid sequence of SEQ ID NO: 1 (LbCas12a), and at the following positions with respect to the position numbering of SEQ ID NO: 1: K116, K120, K121, D122, E125, T148, T149, T152, D156, E159, Q529, G532, D535, K538, D541, Y542, L585, K591, M592, K595, V596, S599, K600, K601, Y616, Y646, and / or W649, one or more (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13 or more) of which contain mutations, and may be mutations at one or more of the following positions with respect to the position numbering of SEQ ID NO: 1: K116, K120, K121, D122, E125, T152, D156, E159, G532, D535, K538, D541, and / or K595. In some embodiments, the mutations of the Cas12a (LbCas12a) polypeptide consist essentially of or consist of mutations at any combination of one or more of the following positions with respect to the position numbering of SEQ ID NO: 1: K116, K120, K121, D122, E125, T152, D156, E159, G532, D535, K538, D541, and / or K595.Accordingly, the modified LbCas12a polypeptide of the present invention may comprise a single mutation at any one of positions: K116, K120, K121, D122, E125, T148, T149, T152, D156, E159, Q529, G532, D535, K538, D541, Y542, L585, K591, M592, K595, V596, S599, K600, K601, Y616, Y646, and / or W649 with respect to the position numbering of SEQ ID NO: 1, or may comprise a combination of mutations at any two or more of positions: K116, K120, K121, D122, E125, T148, T149, T152, D156, E159, Q529, G532, D535, K538, D541, Y542, L585, K591, M592, K595, V596, S599, K600, K601, Y616, Y646, and / or W649 with respect to the position numbering of SEQ ID NO: 1.

[0088] In some embodiments, the mutations of the Cas12a (LbCas12a) polypeptide are the following mutations with respect to the position numbering of SEQ ID NO: 1: K116N, K116R, K120H, K120N, K120Q, K120R, K120T, K121D, K121G, K121H, K121Q, K121R, K121S, K121T, D122H, D122K, D122N, D122R, E125K, E125Q, E125R, E125Y, T148A, T148C, T148H, T148S, T149C, T149F, T149G, T149H, T149N, T149P, T149S, T149V, T152E, T152F, T152H, T152K, T152L, T152Q, T152R, T152W, T152Y, D156E, D156H, D156I, D156K, D156L, D156Q, D156R, D156W, D156Y, E159K, E159Q, E159R, E159Y, Q529A, Q529D, Q529F, Q529G, Q529H, Q529N, Q529P, Q529S, Q529T, Q529W, G532A, G532C, G532D, G532F, G532H, G532K, G532L, G532N, G532Q, G532S, D535A, D535H, D535K, D535N, D535S, D535T, D535V, K538C, K538F, K538G, K538H, K538L, K538M, K538Q, K538R, K538V, K538W, K538Y, D541A, D541E, D541H, D541I, D541N, D541R, D541Y, Y542F, Y542H, Y542K, Y542L, Y542M, Y542N, Y542R, Y542T, Y542V, L585F, L585G, L585H, K591A, K591F, K591G, K591H, K591R, K591S, K591W, K591Y, M592A, M592E, M592Q, K595H, K595L, K595M, K595Q, K595R, K595S, K595W, K595Y, V596H, V596T, S599G, S599H, S599N, K600G, K600H, K600R, K601H, K601Q, K601R, K601T, Y616E, Y616F, Y616H, Y616K, Y616R, Y646E, Y646H, Y646K, Y646N, Y646Q, Y646R, Y646W,Comprising, consisting essentially of, or consisting of one or more of W649H, W649K, W649R, W649S, and / or W649Y. As will be understood, any single Cas12a polypeptide having two or more mutations contains only a single mutation at any given position. Thus, for example, a polypeptide may have a mutation at position D535 of any one of D535A, D535H, D535K, D535N, D535S, D535T, or D535V, but the same polypeptide may further contain one or more mutations at any one of the other positions described herein. In some embodiments, the mutations of the Cas12a (LbCas12a) polypeptide are K116N, K116R, K120H, K120N, K120Q, K120R, K120T, K121D, K121G, K121H, K121Q, K121R, K121S, K121T, D122H, D122K, D122N, D122R, E125K, E125Q, E125R, E125Y, T152E, T152F, T152H, T152K, T152L, T152Q, T152R, T152W, T152Y, D156E, D156H, D156I, D156K, D156L, D156Q, D156R, D156W, D156Y, E159K, E159Q, E159R, E159Y, G532A, G532C, G532D, G532F, G532H, G532K, G532L, G532N, G532Q, G532S, D535A, D535H, D535K, D535N, D535S, D535T, D535V, K538C, K538F, K538G, K538H, K538L, K538M, K538Q, K538R, K538V, K538W, K538Y, D541A, D541E, D541H, D541I, D541N, D541R, D541Y, K595H, K595L, K595M, K595Q, K595R, K595S, K595W, and / or K595Y in any combination, consisting essentially of, or consisting of. In some embodiments, the mutations of the Cas12a (LbCas12a) polypeptide are K116R, K116N, K120Y, K121S, K121R, D122H, D122N, E125K,Comprising, consisting essentially of, or consisting of one or more of the mutations of T152R, T152K, T152Y, T152Q, T152E, T152F, D156R, D156W, D156Q, D156H, D156I, D156V, D156L, D156E, E159K, E159R, G532N, G532S, G532H, G532K, G532R, G532L, D535N, D535H, D535T, D535, S D535A, D535W, K538R K538V, K538Q, K538W, K538Y, K538F, K538H, K538L, K538M, K538C, K538G, K538A, D541E, K595R, K595Q, K595Y, K595W, K595H, K595S, and / or K595M. As understood, any single Cas12a polypeptide having two or more mutations includes a single mutation at any given position. Thus, for example, the polypeptide may have a mutation at position D535 of any one of D535A, D535H, D535K, D535N, D535S, D535T, or D535V, and may further include one or more mutations at any other position described herein.,

[0089] In some embodiments, the mutations do not include, do not consist essentially of, or do not consist of the mutations of D156R, G532R, K538R, K538V, Y542R or K595R with respect to the position numbering of SEQ ID NO: 1. In some embodiments, the mutations of the Cas12a (LbCas12a) polypeptide do not include, do not consist essentially of, or do not consist of a combination of the mutations of G532R and K595R, a combination of the mutations of G532R, K538V and Y542R, or a combination of the mutations of D156R, G532R and K532R with respect to the position numbering of SEQ ID NO: 1.,

[0090] In some embodiments, the modified LbCas12a polypeptide may include one or more amino acid mutations of SEQ ID NO: 1 described in Table 2 (in Example 2).

[0091] In some embodiments, the modified LbCas12a polypeptide may include an altered protospacer adjacent motif (PAM) specificity compared to wild-type LbCas12a (e.g., SEQ ID NO: 1). The modified LbCas12a polypeptide of the present invention may include an altered PAM specificity, which is not limited, but includes NNNG, NNNT, NNNA, NNNC, NNG, NNT, NNC, NNA, NG, NT, NC, NA, NN, NNN, NNNN, where each N in each sequence is independently selected from any of T, C, G, or A. In some embodiments, the altered PAM specificity is not limited, but may include TTTA, TTTC, TTTG, TTTT, TTCA, TTCC, TTCG, TTCT, ATTC, CTTA, CTTC, CTTG, GTTC, TATA, TATC, CTCC, TCCG, TACA, TCCG, TACA, TCCG, TCCC, TCCA, and / or TATG. In some embodiments, the altered PAM specificity may be NNNN, where each N in each sequence is independently selected from any of T, C, G, or A.

[0092] In addition to having an altered PAM recognition specificity, the modified LbCas12a nuclease may further include a mutation within the nuclease active site (e.g., the RuvC domain) (e.g., deadLbCas12a, dLbCas12a). Such modifications can result in an LbCas12a polypeptide with reduced nuclease activity (e.g., nickase activity) or an LbCas12a polypeptide with no nuclease activity.

[0093] In some embodiments, a V-type Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR)-associated (Cas) (CRISPR-Cas) system is provided, the system comprising: (a) a fusion protein comprising: (i) a modified LbCas12a nuclease of the present invention or a nucleic acid encoding the modified LbCas12a nuclease of the present invention, and (ii) a polypeptide of interest or a nucleic acid encoding the polypeptide of interest; and (b) a guide nucleic acid (CRISPR RNA, CRISPR DNA, crRNA, crDNA) comprising a spacer sequence and a repeat sequence, wherein the guide nucleic acid is capable of forming a complex with the modified LbCas12a nuclease or the fusion protein, and the spacer sequence is capable of hybridizing to a target nucleic acid, whereby the modified LbCas12a nuclease and the polypeptide of interest are guided to the target nucleic acid, whereby the system is capable of modifying (e.g., cleaving or editing) or regulating (e.g., transcriptional regulation) the target nucleic acid. In some embodiments, the system comprises a polypeptide of interest (e.g., a fusion protein) linked to the C-terminus and / or N-terminus of the modified LbCas12a nuclease, optionally via a peptide linker.

[0094] Furthermore, fusion proteins comprising the modified Cas12a nuclease of the present invention are provided herein. In some embodiments, the fusion protein may comprise a polypeptide of interest linked to the C-terminus and / or N-terminus of the modified LbCas12a. In some embodiments, the present invention provides a fusion protein comprising the modified LbCas12a (and optionally comprising an intervening linker linking the polypeptide of interest).

[0095] Any linker known in the art or later identified that does not interfere with the activity of the fusion protein may be used. A linker that does not "interfere" with the activity of the fusion protein is a linker that does not reduce or eliminate the activity of the polypeptide of the fusion protein (e.g., the nuclease and / or the polypeptide of interest); that is, nuclease activity, nucleic acid binding activity, editing activity, and / or any other activity of the nuclease or polypeptide of interest is maintained in the fusion protein in which the nuclease and polypeptide of interest are linked to each other via the linker. In some embodiments, the peptide linker may be linked to the C-terminus of the modified LbCas12a (e.g., at its N-terminus), and the fusion protein may further comprise a polypeptide of interest linked to the C-terminus of the linker. In some embodiments, the peptide linker may be linked to the N-terminus of the modified LbCas12a (e.g., at its C-terminus), and the fusion protein may further comprise a polypeptide of interest linked to the N-terminus of the linker. In some embodiments, the modified LbCas12a of the invention may be linked to a linker and / or a polypeptide of interest at both its C-terminus and N-terminus (either directly or via a linker).

[0096] In some embodiments, the linker useful in the present invention may be an amino acid or a peptide. In some embodiments, the peptide linker useful in the present invention may be of a length of about 4 to about 100 or more amino acids, for example, about 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100 or more amino acids in length (e.g., about 4 to about 40, about 4 to about 50, about 4 to about 60, about 5 to about 40, about 5 to about 50, about 5 to about 60, about 9 to about 40, about 9 to about 50, about 9 to about 60, about 10 to about 40, about 10 to about 50, about 10 to about 60, or about 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25 amino acids to about 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100 or more amino acids in length). In some embodiments, the peptide linker may be a GS linker.

[0097] Useful target polypeptides according to the present invention can include, but are not limited to, polypeptides or protein domains having deaminase (deamination) activity, nickase activity, recombinase activity, transposase activity, methylase activity, glycosylase (DNA glycosylase) activity, glycosylase inhibitor activity (e.g., uracil-DNA glycosylase inhibitor (UGI)), demethylase activity, transcriptional activation activity, transcriptional repression activity, transcription release factor activity, histone modification activity, nuclease activity, single-stranded RNA cleavage activity, double-stranded RNA cleavage activity, restriction endonuclease activity (e.g., Fok1), nucleic acid binding activity, methyltransferase activity, DNA repair activity, DNA damage activity, dismutase activity, alkylation activity, depurination activity, oxidation activity, pyrimidine dimer formation activity, integrase activity, transposase activity, polymerase activity, ligase activity, helicase activity, and / or photolyase activity.

[0098] In some embodiments, the target polypeptide may comprise at least one polypeptide or protein domain having deaminase activity. In some embodiments, the at least one polypeptide or protein domain may be an adenine deaminase domain. The adenine deaminase (or adenosine deaminase) useful in the present invention may be any adenine deaminase from any organism, whether known or later identified (see, e.g., U.S. Patent No. 10,113,163, which is incorporated herein by reference for its disclosure of adenine deaminase). Adenine deaminase can catalyze the hydrolytic deamination of adenine or adenosine. In some embodiments, the adenine deaminase can catalyze the hydrolytic deamination of adenosine or deoxyadenosine to inosine or deoxyinosine, respectively. In some embodiments, the adenosine deaminase can catalyze the hydrolytic deamination of adenine or adenosine in DNA. In some embodiments, the adenine deaminase encoded by the nucleic acid construct of the present invention can produce an A→G conversion in the sense (e.g., "+", template) strand of the target nucleic acid or a T→C conversion in the antisense (e.g., "-", complementary) strand of the target nucleic acid.

[0099] In some embodiments, adenosine deaminase may be a variant of naturally occurring adenosine deaminase. Thus, in some embodiments, the adenosine deaminase useful in the present invention may be about 70% to 100% identical to wild-type adenosine deaminase (e.g., about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to naturally occurring adenosine deaminase, and any range or value therein). In some embodiments, the deaminase or deaminases may be referred to as engineered, mutated, or evolved adenosine deaminases that do not occur naturally. Thus, for example, an engineered, mutated, or evolved adenosine deaminase polypeptide or adenosine deaminase domain may be about 70% to 99.9% identical to a naturally occurring adenosine deaminase polypeptide / domain (e.g., about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8% or 99.9% identical to a naturally occurring adenosine deaminase polypeptide or adenosine deaminase domain, and any range or value therein). In some embodiments, adenosine deaminase may be derived from bacteria (e.g., Escherichia coli, Staphylococcus aureus, Haemophilus influenzae, Caulobacter crescentus, etc.). In some embodiments, the polynucleotide encoding the adenosine deaminase polypeptide / domain may be codon-optimized for expression in an organism.

[0100] In some embodiments, the adenosine deaminase domain is a wild-type tRNA-specific adenosine deaminase domain, e.g., tRNA-specific adenosine deaminase (TadA) and / or a mutated / evolved adenosine deaminase domain, e.g., a mutated / evolved tRNA-specific adenosine deaminase domain (TadA * ). In some embodiments, the TadA domain may be from E. coli. In some embodiments, TadA may be modified, e.g., truncated, and one or more N-terminal and / or C-terminal amino acids may be deleted relative to full-length TadA (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 6, 17, 18, 19, or 20 N-terminal and / or C-terminal amino acid residues may be deleted compared to full-length TadA). In some embodiments, the TadA polypeptide or TadA domain does not contain an N-terminal methionine. In some embodiments, wild-type E. coli TadA contains the amino acid sequence of SEQ ID NO: 18. In some embodiments, the mutated / evolved E. coli TadA * contains the amino acid sequences of SEQ ID NO: 19-22. In some embodiments, the polynucleotide encoding TadA / TadA * may be codon-optimized for expression in an organism.

[0101] The cytosine deaminase (or cytidine deaminase) useful in the present invention may be any cytosine deaminase derived from any organism, whether known or later identified (see, for example, U.S. Patent No. 10,167,457, which is incorporated herein by reference for its disclosure of cytosine deaminase). In some embodiments, at least one polypeptide or protein domain may be a cytosine deaminase polypeptide or domain. In some embodiments, the cytosine deaminase polypeptide / domain may be an apolipoprotein B mRNA editing catalytic polypeptide-like (APOBEC) domain. In some embodiments, the polypeptide of interest may comprise at least one polypeptide or protein domain having glycosylase inhibitor activity. In some embodiments, the polypeptide of interest may be a uracil-DNA glycosylase inhibitor (UGI) polypeptide / domain. In some embodiments, the nucleic acid construct encoding the modified LbCas12a nuclease and the cytosine deaminase domain of the present invention (e.g., encoding a fusion protein comprising the modified LbCas12a nuclease and the cytosine deaminase domain) may further encode a uracil-DNA glycosylase inhibitor (UGI), where the UGI is codon-optimized for expression in an organism. In some embodiments, the present invention provides a fusion protein comprising a modified LbCas12a nuclease, a cytosine deaminase domain, and UGI, and / or one or more polynucleotides encoding the same, where the one or more polynucleotides may be codon-optimized for expression in an organism.

[0102] Cytosine deaminase catalyzes the hydrolytic deamination of cytidine or deoxycytidine to uridine or deoxyuridine, respectively. In some embodiments, the deaminase or deaminase domain may be a cytidine deaminase domain that catalyzes the hydrolytic deamination of cytosine to uracil. In some embodiments, the cytosine deaminase may be a variant of a cytosine deaminase of natural origin, including but not limited to primates (e.g., humans, monkeys, chimpanzees, gorillas), dogs, cows, rats or mice. Thus, in some embodiments, the cytosine deaminase useful in the present invention may be about 70% to about 100% identical to the wild-type cytosine deaminase (e.g., about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to a cytosine deaminase of natural origin, and any range or value therein). In some embodiments, the polynucleotide encoding the cytosine deaminase polypeptide / domain may be codon-optimized for expression in an organism.

[0103] In some embodiments, the cytosine deaminase useful in the present invention may be an apolipoprotein B mRNA editing complex (APOBEC) family deaminase. In some embodiments, the cytosine deaminase may be APOBEC1 deaminase, APOBEC2 deaminase, APOBEC3A deaminase, APOBEC3B deaminase, APOBEC3C deaminase, APOBEC3D deaminase, APOBEC3F deaminase, APOBEC3G deaminase, APOBEC3H deaminase, APOBEC4 deaminase, human activation-induced deaminase (hAID), rAPOBEC1, FERNY, and / or CDA1, and may also be pmCDA1, atCDA1 (e.g., At2g19570), and evolved versions thereof. In some embodiments, the cytosine deaminase may be APOBEC1 deaminase having the amino acid sequence of SEQ ID NO: 23, SEQ ID NO: 44 or SEQ ID NO: 46. In some embodiments, the cytosine deaminase may be APOBEC3A deaminase having the amino acid sequence of SEQ ID NO: 24. In some embodiments, the cytosine deaminase may be CDA1 deaminase, and may be CDA1 having the amino acid sequence of SEQ ID NO: 25 or SEQ ID NO: 43. In some embodiments, the cytosine deaminase may be FERNY deaminase, and may be FERNY having the amino acid sequence of SEQ ID NO: 42 or SEQ ID NO: 45. In some embodiments, the cytosine deaminase may be human activation-induced deaminase (hAID) having the amino acid sequence of SEQ ID NO: 47 or SEQ ID NO: 48. In some embodiments, the cytosine deaminase useful in the present invention may be about 70% to about 100% identical to the amino acid sequence of a naturally occurring cytosine deaminase (e.g., 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5% or 100% identical) (e.g., an evolved deaminase).In some embodiments, the cytosine deaminase useful in the present invention has an amino acid sequence that is about 70% to about 99.5% identical (e.g., about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 99.5% identical) (e.g., at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical) to the amino acid sequence of SEQ ID NO: 23, SEQ ID NO: 24, SEQ ID NO: 25, or SEQ ID NOs: 42-48. In some embodiments, the polynucleotide encoding the cytosine deaminase may be codon-optimized for expression in an organism, and the codon-optimized polypeptide may be about 70% to 99.5% identical to the reference polynucleotide.

[0104] The "uracil glycosylase inhibitor" (UGI) useful in the present invention may be any protein capable of inhibiting uracil-DNA glycosylase base excision repair enzyme. In some embodiments, the UGI domain comprises wild-type UGI or a fragment thereof. In some embodiments, the UGI domain useful in the present invention is about 70% to about 100% identical (e.g., about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5% or 100% identical and any range or value therein) to the amino acid sequence of a native origin UGI domain. In some embodiments, the UGI domain may comprise a polypeptide having about 70% to about 99.5% identity (e.g., at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical) to the amino acid sequence of SEQ ID NO: 26. For example, in some embodiments, the UGI domain may comprise a fragment of the amino acid sequence of SEQ ID NO: 26 that is 100% identical to a portion of the contiguous nucleotides of the amino acid sequence of SEQ ID NO: 26 (e.g., about 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80 contiguous nucleotides; e.g., about 10, 15, 20, 25, 30, 35, 40, 45, ~about 50, 55, 60, 65, 70, 75, 80 contiguous nucleotides). In some embodiments, the UGI domain may be a variant of a known UGI (e.g., SEQ ID NO: 26) having 70% to about 99.5% identity (e.g., about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5% identity, and any range or value therein) to the known UGI.In some embodiments, the polynucleotide encoding UGA may be codon-optimized for expression in an organism, and the codon-optimized polypeptide may be about 70% to about 99.5% identical to the reference polynucleotide.

[0105] In some embodiments, the modified LbCas12a nuclease may contain a mutation within its nuclease active site (e.g., RuvC). A modified LbCas12a nuclease that has a mutation within its nuclease active site(s) and no longer contains nuclease activity is generally referred to as "dead" (e.g., dLbCas12a). In some embodiments, a modified LbCas12a domain or polypeptide having a mutation within its nuclease active site(s) may have impaired or decreased activity (e.g., nickase activity) compared to the same LbCas12a nuclease without the mutation.

[0106] The modified LbCas12a nuclease of the present invention can be used in combination with a guide RNA (gRNA, CRISPR array, CRISPR RNA, crRNA) designed to function with the modified LbCas12a nuclease to modify a target nucleic acid. The guide nucleic acid useful in the present invention includes at least a spacer sequence and a repeat sequence. The guide nucleic acid can form a complex with the LbCas12a nuclease domain encoded and expressed by the polynucleotide / nucleic acid construct of the present invention encoding the modified LbCas12a nuclease. The spacer sequence can hybridize to the target nucleic acid, thereby guiding the nucleic acid construct (e.g., the modified LbCas12a nuclease (and / or the polypeptide of interest)) to the target nucleic acid, where the target nucleic acid can be modified (e.g., cleaved or edited) or regulated (e.g., transcriptional regulation) by the modified LbCas12a nuclease (and / or the encoded deaminase domain and / or the polypeptide of interest). As an example, a nucleic acid construct encoding an LbCas12a domain (e.g., a fusion protein) linked to a cytosine deaminase domain can be used in combination with an LbCas12a guide nucleic acid to modify a target nucleic acid, where the cytosine deaminase domain of the fusion protein deaminates the cytosine base in the target nucleic acid, thereby editing the target nucleic acid. In a further example, a nucleic acid construct encoding an LbCas12a domain (e.g., a fusion protein) linked to an adenine deaminase domain can be used in combination with an LbCas12a guide nucleic acid to modify a target nucleic acid, where the adenine deaminase domain of the fusion protein deaminates the adenosine base in the target nucleic acid, thereby editing the target nucleic acid.

[0107] As used herein, "guide nucleic acid", "guide RNA", "gRNA", "CRISPR RNA / DNA", "crRNA" or "crDNA" means a nucleic acid comprising at least one spacer sequence that is complementary to (and hybridizes to) a target nucleic acid (e.g., a protospacer), and at least one repeat sequence (e.g., a repeat of a type V Cas12a CRISPR-Cas system, or a fragment or portion thereof), where the repeat sequence can be linked to the 5' end and / or 3' end of the spacer sequence. The design of the gRNA of the present invention can be based on the type V Cas12a system.

[0108] In some embodiments, the Cas12a gRNA may comprise, from 5' to 3', a repeat sequence (full length or a part thereof ("handle"); e.g., a pseudoknot-like structure) and a spacer sequence.

[0109] In some embodiments, the guide nucleic acid may comprise more than one "repeat sequence - spacer" sequence (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, or more repeat - spacer sequences) (e.g., repeat - spacer - repeat, e.g., repeat - spacer - repeat - spacer - repeat - spacer - repeat - spacer - repeat - spacer). The guide nucleic acid of the present invention is synthetic, human - made, and not found in nature. The gRNA can be quite long and can be used as an aptamer (like the MS2 recruitment strategy) or as another RNA structure hanging off the spacer.

[0110] As used herein, the "repeat sequence" refers to, for example, any repeat sequence of a wild-type Cas12a locus (such as the LbCas12a locus), or the repeat sequence of a synthetic crRNA that is functional with the LbCas12a nuclease encoded by the nucleic acid construct of the present invention. The repeat sequences useful in the present invention can be any known or later-identified repeat sequences of the Cas12a locus, or can be synthetic repeats designed to function in a type V Cas12a CRISPR-Cas system. The repeat sequence may include a hairpin structure and / or a stem-loop structure. In some embodiments, the repeat sequence may form a pseudoknot-like structure at its 5' end (e.g., the "handle"). Thus, in some embodiments, the repeat sequence can be identical or substantially identical to a repeat sequence derived from a wild-type type V CRISPR-Cas locus (e.g., a wild-type Cas12a locus). The repeat sequence derived from a wild-type Cas12a locus can be determined by established algorithms, for example, using CRISPRfinder provided by CRISPRdb (see Grissa et al. Nucleic Acids Res. 35(Web Server issue):W52-7). In some embodiments, the repeat sequence or a portion thereof is ligated to the 5' end of a spacer sequence at its 3' end, thereby forming a repeat-spacer sequence (e.g., a guide RNA, crRNA).

[0111] In some embodiments, the repeat sequence comprises, consists essentially of, or consists of at least 10 nucleotides (e.g., about 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50-100 or more nucleotides, or any range or value therein; e.g., about), depending on the particular repeat and regardless of whether the guide RNA comprising the repeat is processed. In some embodiments, the repeat sequence comprises, consists essentially of, or consists of about 10 to about 20, about 10 to about 30, about 10 to about 45, about 10 to about 50, about 15 to about 30, about 15 to about 40, about 15 to about 45, about 15 to about 50, about 20 to about 30, about 20 to about 40, about 20 to about 50, about 30 to about 40, about 40 to about 80, about 50 to about 100, or more nucleotides.

[0112] The repeat sequence linked to the 5' end of the spacer sequence can include a portion of the repeat sequence (e.g., 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35 or more consecutive nucleotides of the wild-type repeat sequence). In some embodiments, the portion of the repeat sequence linked to the 5' end of the spacer sequence can be about 5 to about 10 contiguous nucleotides in length (e.g., about 5, 6, 7, 8, 9, 10 nucleotides) and can have at least 90% identity (e.g., at least about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more) to the identical region (e.g., the 5' end) of the wild-type CRISPR Cas repeat nucleotide sequence. In some embodiments, the portion of the repeat sequence can include a pseudoknot-like structure at its 5' end (e.g., a "handle").

[0113] A "spacer sequence" as used herein is a nucleotide sequence that is complementary to a target nucleic acid (e.g., a target DNA) (e.g., a protospacer). A spacer sequence can be fully complementary or substantially complementary (e.g., at least about 70% complementary (e.g., about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more, and any range or value therein)) to a target nucleic acid. Thus, in some embodiments, a spacer sequence can have 1, 2, 3, 4, or 5 mismatches compared to a target nucleic acid, and the mismatches can be contiguous or non-contiguous. In some embodiments, the spacer sequence can have about 70% complementarity to the target nucleic acid. In other embodiments, the spacer nucleotide sequence can have about 80% complementarity to the target nucleic acid. In still other embodiments, the spacer nucleotide sequence can have about 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 99.5% complementarity to the target nucleic acid (protospacer), etc. In some embodiments, the spacer sequence is 100% complementary to the target nucleic acid. The spacer sequence can have a length of about 15 nucleotides to about 30 nucleotides (e.g., about 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides, or any range or value therein). Thus, in some embodiments, the spacer sequence may have perfect or substantial complementarity over a region of the target nucleic acid (e.g., a protospacer), which may be at least about 15 nucleotides to about 30 nucleotides in length. In some embodiments, the spacer may be about 20, 21, 22, 23, 24, or 25 nucleotides in length. In some embodiments, the spacer may be 23 nucleotides in length.

[0114] In some embodiments, the 5' region of the spacer sequence of the guide RNA may be identical to the target nucleic acid while the 3' region of the spacer may be substantially complementary to the target nucleic acid (e.g., type V CRISPR-Cas), or the 3' region of the spacer sequence of the guide RNA may be identical to the target nucleic acid while the 5' region of the spacer may be substantially complementary to the target nucleic acid (e.g., type II CRISPR-Cas), and thus the overall complementarity of the spacer sequence to the target nucleic acid may be less than 100%. Thus, for example, in a guide for a type V CRISPR-Cas system, for example, the first 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 nucleotides in the 5' region of a 20 nucleotide spacer sequence (i.e., the seed region) may be 100% complementary to the target nucleic acid, and the remaining nucleotides in the 3' region of the spacer sequence may be substantially complementary (e.g., at least about 70% complementary) to the target nucleic acid. In some embodiments, the first 1-8 nucleotides (e.g., the first 1, 2, 3, 4, 5, 6, 7, 8, nucleotides, and any range therein) at the 5' end of the spacer sequence can be 100% complementary to the target nucleic acid, and the remaining nucleotides in the 3' region of the spacer sequence can be substantially complementary (e.g., at least about 50% complementary (e.g., about 50%, 55%, 60%, 65%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more)) to the target nucleic acid.

[0115] In some embodiments, the seed region of the spacer can be about 8 to about 10 nucleotides in length, about 5 to about 6 nucleotides in length, or about 6 nucleotides in length.

[0116] As used herein, a "target nucleic acid", "target DNA", "target nucleotide sequence", "target region", or "target region in a genome" refers to a region of an organism's genome that is fully complementary (100% complementary) or substantially complementary (e.g., at least 70% complementary (e.g., about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more, and any range or value therein)) to a spacer sequence in a guide RNA of the present invention. In some embodiments, a target region useful for a V-type CRISPR-Cas system (e.g., LbCas12a) is located immediately 3' to a PAM sequence in the genome of an organism (e.g., a plant genome, an animal genome, a bacterial genome). In some embodiments, the target region can be selected from any at least 15 contiguous nucleotides (e.g., 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 or more nucleotides, and any range or value therein; e.g., from about 19 to about 25 nucleotides, from about 20 to about 24 nucleotides in length, etc.) located immediately adjacent to the PAM sequence.

[0117] "Protospacer sequence" refers to a target nucleic acid, specifically a portion of a target nucleic acid (e.g., or a target region within a genome) that is completely or substantially complementary to (and hybridizes with) a spacer sequence of a CRISPR repeat-spacer sequence (e.g., guide RNA, CRISPR array, crRNA).

[0118] In the case of the Type V CRISPR-Cas Cas12a system, the protospacer sequence is flanked (immediately adjacent) to a protospacer adjacent motif (PAM) that is located at the 5' end of the non-targeted strand and the 3' end of the targeted strand (see below for an example). JPEG2025087754000002.jpg29166

[0119] A canonical Cas12a PAM is T-rich. In some embodiments, the canonical Cas12a PAM sequence may be 5'-TTN, 5'-TTTN, or 5'-TTTV.

[0120] The polypeptides, fusion proteins and / or systems of the invention may be encoded by a polynucleotide or nucleic acid construct. In some embodiments, the polynucleotide / nucleic acid construct encoding the polypeptides, fusion proteins and / or systems of the invention may be operatively associated with regulatory elements (e.g., promoters, terminators, etc.) for expression in an organism of interest and / or cells of an organism of interest as described herein. In some embodiments, the polynucleotide / nucleic acid construct encoding the polypeptides, fusion proteins and / or systems of the invention may be codon-optimized for expression in an organism.

[0121] In some embodiments, the invention provides a complex comprising (a) a modified LbCas12a polypeptide of the invention or a fusion protein of the invention and (b) a guide nucleic acid (e.g., CRISPR RNA, CRISPR DNA, crRNA, crDNA).

[0122] In some embodiments, the invention provides a composition comprising (a) a modified LbCas12a polypeptide of the invention or a fusion protein of the invention and (b) a guide nucleic acid.

[0123] In some embodiments, the present invention provides an expression cassette and / or vector comprising the polynucleotide / nucleic acid construct of the present invention. In some embodiments, an expression cassette and / or vector comprising the polynucleotide / nucleic acid construct of the present invention and / or one or more guide nucleic acids may be provided. In some embodiments, the nucleic acid construct encoding the modified CRISPR-Cas nuclease of the present invention and / or the fusion protein comprising the modified CRISPR-Cas nuclease may be included in the same or separate expression cassette or vector as that comprising the guide nucleic acid. When the nucleic acid construct is included in a separate expression cassette or vector as that comprising the guide nucleic acid, the target nucleic acid may be contacted (e.g. provided) with the expression cassette or vector comprising the nucleic acid construct of the present invention before, simultaneously with, or after the expression cassette comprising the guide nucleic acid is provided (e.g. contacted with the target nucleic acid).

[0124] In some embodiments, the invention provides expression cassettes and / or vectors encoding the compositions and / or complexes of the invention or comprising the systems of the invention.

[0125] In some embodiments, polynucleotides, nucleic acid constructs, expression cassettes and / or vectors of the invention that have been optimized for expression in an organism may be about 70% to about 100% identical (e.g., about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5% or 100%, and any value or range therein) to a polynucleotide, nucleic acid construct, expression cassette and / or vector encoding the same modified CRISPR-Cas nuclease or fusion protein of the invention but that has not been codon optimized for expression in an organism. Organisms for which a polynucleotide or nucleic acid construct may be optimized may include, but are not limited to, animals, plants, fungi, archaea, or bacteria. In some embodiments, a polynucleotide or nucleic acid construct of the invention is codon optimized for expression in a plant.

[0126] In some embodiments, the invention provides cells comprising one or more polynucleotides, guide nucleic acids, nucleic acid constructs, systems, expression cassettes and / or vectors of the invention.

[0127] The nucleic acid constructs of the invention (e.g., encoding a modified CRISPR-Cas nuclease of the invention and / or a fusion protein comprising a modified CRISPR-Cas nuclease of the invention) and expression cassettes / vectors comprising same can be used to modify target nucleic acids and / or their expression in vivo (e.g., in an organism or cells of an organism (e.g., a plant)) and in vitro (e.g., in a cell or a cell-free system).

[0128] The invention further provides methods of altering the PAM specificity of a Cas12a polypeptide. In some embodiments, methods of altering the PAM specificity are provided comprising introducing mutations into a Cas12a polypeptide, wherein the mutations are at amino acid residues K116, K120, K121, D122, E125, T148, T149, T152, D156, E159, Q529, D535, K538, D541, Y542, L585, K591, M592, K595, V596, S599, K600, K601, Y616, Y646, W649 with respect to position numbering of SEQ ID NO:1. In some embodiments, the mutations introduced into the Cas12a polypeptide are, with respect to the position numbering of SEQ ID NO:1, K116R, K116N, K120R, K120H, K120N, K120T, K120Y, K120Q, K121S, K121T, K121H, K121R, K121G, K121D, K121Q, D ... 2R, D122K, D122H, D122E, D122N, E125R, E125K, E125Q, E125Y, T148H, T148S, T148A, T148 C, T149A, T149C, T149S, T149G, T149H, T149P, T149F, T149N, T149D, T149V, T152R, T152K, T152W, T152Y, T152H, T152Q, T152E, T152L, T152F, D156R, D156K, D156Y, D156W, D156Q, D 156H, D156I, D156V, D156L, D156E, E159K, E159R, E159H, E159Y, E159Q, Q529N, Q529T, Q52 9H, Q529A, Q529F, Q529G, Q529G, Q529S, Q529P, Q529W, Q529D, G532D, G532N, G532S, G532 H, G532F, G532K, G532R, G532Q, G532A, G532L, G532C, D535N, D535H, D535V, D535T, D535, S D535A, D535W, D535K, K538RK538V, K538Q, K538W, K538Y, K538F, K538H, K538L, K538M, K538C, K538G, K538A, K538P , D541N, D541H, D541R, D541K, D541Y, D541I, D541A, D541S, D541E, Y542R, Y542K, Y542H , Y542Q, Y542F, Y542L, Y542M, Y542P, Y542V, Y542N, Y542T, L585G, L585H, L585F, K591W , K591F, K591Y, K591H, K591R, K591S, K591A, K591G, K591P, M592R, M592K, M592Q, M592E , M592A, K595R, K595Q, K595Y, K595L, K595W, K595H, K595E, K595S, K595D, K595M, V596 T, V596H, V596G, V596A, S599G, S599H, S599N, S599D, K600R, K600H, K600G, K601R, K601 H, K601Q, K601T, Y616K, Y616R, Y616E, Y616F, Y616H, Y646R, Y646E, Y646K, Y646H, Y646Q, Y646W, Y646N, W649H, W649K, W649Y, W649R, W649E, W649S, W649V, and / or W649T. In some embodiments, the mutations introduced into the Cas12a polypeptide are at amino acid residue positions: K116, K120, K121, D122, E125, T152, D156, E159, G532, D535, K538, D541, and / or K595 with respect to the position numbering of SEQ ID NO:1, where the mutations are K116R, K116N, K120Y ... 21S, K121R, D122H, D122N, E125K, T152R, T152K, T152Y, T152Q, T152E, T152F, D156R, D156W, D156Q, D156H, D156 I, D156V, D156L, D156E, E159K, E159R, G532N, G532S, G532H, G532K, G532R, G532L, D535N, D535H, D535T, D535, S D535A, D535W, K538RThe mutation may be K538V, K538Q, K538W, K538Y, K538F, K538H, K538L, K538M, K538C, K538G, K538A, D541E, K595R, K595Q, K595Y, K595W, K595H, K595S, and / or K595M. The mutation introduced may be a single mutation or a combination of two or more mutations. As will be understood, any single Cas12a polypeptide with two or more mutations contains only a single mutation at any given position. In some embodiments, the Cas12a polypeptide with altered PAM specificity by the methods of the invention is an LbCas12a polypeptide (Lachnospiraceae bacterium).

[0129] The modified Cas12a polypeptide or nuclease (e.g., LbCas12a nuclease) of the present invention can be used to modify a target nucleic acid in a cell or cell-free system (e.g., to modify a target nucleic acid, to modify the genome of a cell / organism). Thus, in some embodiments, a method of modifying a target nucleic acid is provided, the method comprising contacting the target nucleic acid with (a) (i) a modified LbCas12a polypeptide of the present invention, or a fusion protein of the present invention (e.g., a modified LbCas12a polypeptide of the present invention and a polypeptide of interest (e.g., a deaminase)) and (ii) a guide nucleic acid; (b) (i) a complex of the present invention comprising a modified LbCas12a polypeptide or fusion protein of the present invention, and (ii) a guide nucleic acid; (c) (i) a composition comprising a modified LbCas12a polypeptide of the present invention, or a fusion protein of the present invention, and (ii) a guide nucleic acid; and / or (d) a system of the present invention, thereby modifying the target nucleic acid. In some embodiments, a method for modifying / altering the genome of a cell or organism is provided, the method comprising contacting a target nucleic acid in the genome of the cell / organism with (a) (i) a modified LbCas12a polypeptide of the invention, or a fusion protein of the invention (e.g., a modified LbCas12a polypeptide of the invention and a polypeptide of interest (e.g., a deaminase)) and (ii) a guide nucleic acid; (b) a complex of the invention comprising (i) a modified LbCas12a polypeptide or fusion protein of the invention and (ii) a guide nucleic acid; (c) a composition comprising (i) a modified CRISPR-Cas nuclease of the invention (e.g., a modified LbCas12a polypeptide) or a fusion protein of the invention and (ii) a guide nucleic acid; and / or (d) a system of the invention, thereby modifying / altering the genome of the cell or organism. In some embodiments, the cell or organism is a plant cell or plant.

[0130] In some embodiments, a method of modifying a target nucleic acid is provided, the method comprising contacting a cell or cell-free system comprising the target nucleic acid with (a) (i) a polynucleotide of the invention (e.g., encoding a modified LbCas12a polypeptide of the invention, or encoding a fusion protein comprising a modified LbCas12a polypeptide of the invention and a polypeptide of interest (e.g., a deaminase), or an expression cassette or vector comprising same, and (ii) a guide nucleic acid, or an expression cassette and / or vector comprising same; and / or (b) a nucleic acid construct encoding a complex of the invention comprising a modified LbCas12a polypeptide of the invention, or a fusion protein comprising a modified LbCas12a polypeptide of the invention and a polypeptide of interest (e.g., a deaminase), or an expression cassette and / or vector comprising same, wherein the contacting is performed under conditions in which the polynucleotide and / or nucleic acid construct is expressed to produce a modified LbCas12a polypeptide and / or fusion protein that forms a complex with the guide nucleic acid, thereby modifying the target nucleic acid.In some embodiments, methods of modifying / altering the genome of a cell and / or organism are provided, comprising contacting a cell and / or a cell within the organism comprising a target nucleic acid with (a) (i) a polynucleotide of the invention (e.g., encoding a modified LbCas12a polypeptide of the invention, or encoding a fusion protein comprising a modified LbCas12a polypeptide of the invention and a polypeptide of interest (e.g., a deaminase), or an expression cassette or vector comprising same, and (ii) a guide nucleic acid, or an expression cassette and / or vector comprising same; and / or (b) a nucleic acid construct encoding a complex of the invention comprising a modified LbCas12a polypeptide of the invention, or a fusion protein comprising a modified LbCas12a polypeptide of the invention and a polypeptide of interest (e.g., a deaminase), or an expression cassette and / or vector comprising same, wherein the contacting is under conditions in which the polynucleotide and / or nucleic acid construct is expressed and a modified LbCas12a polypeptide and / or fusion protein is produced which forms a complex with the guide nucleic acid, thereby modifying the target nucleic acid.

[0131] In some embodiments, the invention provides a method of editing a target nucleic acid, the method comprising contacting the target nucleic acid with (a)(i) a fusion protein of the invention (comprising an engineered LbCas12a polypeptide of the invention and a polypeptide of interest (e.g., a deaminase)), and (a)(ii) a guide nucleic acid; (b) a complex comprising the fusion protein of the invention and a guide nucleic acid; (c) a composition comprising the fusion protein of the invention and a guide nucleic acid; and / or (d) a system of the invention, thereby editing the target nucleic acid.

[0132] In some embodiments, the invention provides a method of editing a target nucleic acid, the method comprising contacting a cell or cell-free system comprising the target nucleic acid with (a)(i) a polynucleotide encoding a fusion protein of the invention (e.g., a modified LbCas12a polypeptide of the invention and a polypeptide of interest (e.g., a deaminase)) or an expression cassette and / or vector comprising the same, and (a)(ii) a guide nucleic acid or an expression cassette and / or vector comprising the same; (b) a nucleic acid construct encoding a complex comprising the fusion protein of the invention and the guide nucleic acid or an expression cassette and / or vector comprising the same; and / or (c) a system of the invention, wherein the contacting is performed under conditions in which the polynucleotide and / or nucleic acid construct is expressed and a modified CRISPR-Cas nuclease and / or fusion protein is produced, which forms a complex with the guide nucleic acid, thereby editing the target nucleic acid.

[0133] CRISPR-Cas nucleases with engineered PAM recognition specificities can be used in many ways, including but not limited to, in creating indels (NHEJ), homology-directed repair, as genomic recognition elements without nuclease function (dead Cpf1), as genomic recognition elements with partially functional nuclease (nickase Cpf1), in fusion proteins for catalytic editing of genomic DNA (DNA base editors), in fusion proteins for catalytic editing of RNA (RNA base editors), for targeting other macromolecules to specific genomic regions; for targeting small chemicals to specific genomic regions, for labeling specific genomic regions, and / or CRISPR-directed genome recombination strategies.

[0134] When provided on a different nucleic acid construct, expression vector, and / or vector, the nucleic acid construct of the invention can be contacted with the target nucleic acid before, simultaneously with, or after contacting the guide nucleic acid with the target nucleic acid.

[0135] The modified CRISPR-Cas nucleases and polypeptides of the present invention and the nucleic acid constructs encoding them can be used to modify target nucleic acids in any organism, including, but not limited to, animals, plants, fungi, archaea, or bacteria. Animals can include, but are not limited to, mammals, insects, fish, birds, etc. Exemplary mammals for which the present invention may be useful include, but are not limited to, primates (human and non-human, e.g., chimpanzees, baboons, monkeys, gorillas, etc.), cats, dogs, mice, rats, ferrets, gerbils, hamsters, cows, pigs, horses, goats, donkeys, or sheep.

[0136] The target nucleic acid of any plant or plant part can be modified and / or edited (e.g., mutated, e.g., base edited, truncated, nicked, etc.) using the nucleic acid constructs of the present invention. Any plant (or grouping of plants into, e.g., genera or higher classifications), including angiosperms, gymnosperms, monocots, dicots, C3, C4, CAM plants, bryophytes, ferns and / or ferns other than the class Pteridophytes, microalgae, and / or macroalgae, can be modified using the nucleic acid constructs of the present invention. Plants and / or plant parts useful in the present invention can be plants and / or plant parts of any plant species / variety / cultivar. The term "plant part" as used herein includes, but is not limited to, embryos, pollen, ovules, seeds, leaves, stems, buds, flowers, branches, fruits, grains, ears, cobs, husks, petioles, roots, root tips, anthers, plant cells (including intact plant cells in plants and / or plant parts, plant protoplasts, plant tissues, plant cell tissue cultures, plant calli, plant clumps, etc.). As used herein, "buds" refers to parts above ground level, including leaves and stems. Furthermore, as used herein, "plant cells" refers to the structural and physiological units of plants, including cell walls, and may also refer to protoplasts. Plant cells can be in the form of isolated single cells, or can be cultured cells, or can be part of a higher organized unit, such as, for example, plant tissues or plant organs.

[0137] Non-limiting examples of plants useful in the present invention include turfgrasses (e.g., bluegrass, bentgrass, ryegrass, fescue), feather reed grass, tufted hair grass, miscanthus, arundo, switchgrass, vegetable crops including: artichoke, kohlrabi, arugula, leek, asparagus, lettuce (e.g., head, leaf, romaine), malanga, melons (e.g., muskmelon, watermelon, cranberry, honeydew, cantaloupe), brassica crops (e.g., Brussels sprouts, cabbage, cauliflower, broccoli, collards, kale, Chinese cabbage, bok choy), cardoni, carrots, napa, okra, tuna, and the like. Onion, celery, parsley, chickpeas, parsnips, chicory, peppers, potatoes, cucurbits (e.g., cucumber, zucchini, toadflax, pumpkin, honeydew melon, watermelon, cantaloupe), radishes, dry bulb onions, rutabaga, eggplant, burdock, endive, shallot, endive, garlic, spinach, leeks, toadflax, leafy vegetables, beets (e.g., sugar beet and fodder beet), sweet potato, chard, horseradish, tomatoes, turnips, and spices;Fruit crops, such as apples, apricots, cherries, nectarines, peaches, pears, plums, prunes, cherries, quince, figs, nuts (e.g. chestnuts, pecans, pistachios, hazelnuts, pistachios, peanuts, walnuts, macadamia nuts, almonds, etc.), citrus fruits (e.g. clementines, kumquats, oranges, grapefruit, tangerines, mandarins, lemons, limes, etc.), blueberries, black raspberries, boysenberries, cranberries, currants, gooseberries, loganberries, raspberries, strawberries, blackberries, grapes (e.g. wine and table), avocados, bananas, kiwis, persimmons, pomegranates, pineapples, tropical fruits, pome fruits, melons, mandarins, grapes, grapefruits, grape ... go, papaya, and lychee, crop plants such as clover, alfalfa, timothy grass, evening primrose, meadowfoam, corn / maize (e.g., field, sweet, popcorn), hops, jojoba, buckwheat, safflower, quinoa, wheat, rice, barley, rye, millet, sorghum, oats, lyco, sorghum, tobacco, kapok, legumes (beans (e.g., green and dry), lentils, peas, soybeans), oilseed plants (e.g., rapeseed, canola, mustard, poppy, olive, sunflower, coconut, castor oil plant, cocoa beans, peanuts, oil palm, soybean, camellia, etc.), duckweed, Arabidopsis, fiber plants (cotton, flax, hemp, jute), hemp (e.g., Cannabis sativa, sativa, Cannabis indica, and Cannabis ruderalis), Lauraceae (e.g. cinnamon, camphor), or plants, such as coffee, sugar cane, tea, and rubber plants; and / or bedding plants, such as flowering plants, cacti, succulents and / or ornamental plants (e.g. roses, tulips, violets), and trees, such as forest trees (broad-leaved and evergreen, e.g. conifers;For example, elm, ash, oak, maple, fir, spruce, cedar, pine, birch, cypress, eucalyptus, willow), as well as shrubs and other seedlings. In some embodiments, the nucleic acid constructs of the invention and / or expression cassettes and / or vectors encoding same may be used to modify corn, soybean, wheat, canola, rice, tomato, pepper, sunflower, raspberry, blackberry, black raspberry, and / or cherry.

[0138] The present invention further includes kits for carrying out the methods of the present invention. The kits of the present invention may include reagents, buffers, and / or instruments for mixing, measuring, selecting, labeling, etc., as well as instructions suitable for modifying the target nucleic acid, etc.

[0139] In some embodiments, the present invention provides a kit comprising one or more polynucleotides and / or nucleic acid constructs of the present invention, and / or expression cassettes and / or vectors comprising same, and may include instructions for their use. In some embodiments, the kit may further comprise a polypeptide of interest and / or a polynucleotide encoding same and an expression cassette and / or vector comprising same. In some embodiments, the guide nucleic acid may be provided on the same expression cassette and / or vector as the nucleic acid construct of the present invention. In some embodiments, the guide nucleic acid may be provided on a separate expression cassette or vector from the one comprising the nucleic acid construct of the present invention.

[0140] Thus, in some embodiments, a kit is provided that includes a nucleic acid construct comprising (a) a polynucleotide encoding a modified CRISPR-Cas nuclease provided herein and (b) a promoter driving expression of the polynucleotide of (a). In some embodiments, the kit may further include a nucleic acid construct encoding a guide nucleic acid, where the construct includes a cloning site for cloning a nucleic acid sequence identical or complementary to a target nucleic acid sequence into the backbone of the guide nucleic acid.

[0141] In some embodiments, the kit may include a nucleic acid construct comprising / encoding one or more nuclear localization signals, where the nuclear localization signals are fused to a CRISPR-Cas nuclease. In some embodiments, a kit is provided that includes a nucleic acid construct of the invention encoding a modified CRISPR-Cas nuclease of the invention, and / or an expression cassette and / or vector comprising the same, where the nucleic acid construct, expression cassette and / or vector may further encode one or more selectable markers useful for identifying transformants (e.g., nucleic acids encoding antibiotic resistance genes, herbicide resistance genes, etc.). In some embodiments, the nucleic acid construct may be an mRNA encoding one or more introns within the encoded CRISPR-Cas nuclease. In some embodiments, the kit may include a promoter, and a promoter and intron, for use in expressing the polypeptides and nucleic acid constructs of the invention.

[0142] Methods for modifying the PAM specificity of CRISPR-Cas nucleases and related compositions CRISPR-Cas system is directed to target nucleic acid using two main criteria: the homology of guide RNA to the target DNA sequence, and the presence of a protospacer adjacent motif (PAM) of a specific sequence. Different CRISPR-Cas nucleases have different PAM sequence requirements (e.g., NGG for SpCas9, or TTTV for LbCas12a (Cpf1), where V is any non-thymidine nucleotide). Screening new CRISPR nucleases or their mutants for their PAM requirements can be complex and unpredictable, since there are so many repeat possibilities. In vitro assays, especially PAM determination assays (PAMDA), can be used to screen the PAM specificity for any particular CRISPR nuclease or its mutants. These assays rely on randomized portions of DNA adjacent to a defined / known protospacer sequence. A guide RNA can be designed to target a known protospacer sequence, and if the randomized PAM region contains the appropriate DNA sequence (e.g., recognized by a CRISPR nuclease or a mutant thereof), the CRISPR nuclease can bind and cleave the target.

[0143] PAM recognition sites for CRISPR-Cas nucleases can be assessed using a PAM site depletion assay (e.g., PAM depletion assay) or PAM determination assay (PAMDA) (Kleinstiver et al. Nat Biotechnol 37:276-282 (2019)). For a PAM depletion assay, a library of plasmids with randomized nucleotides (base pairs) adjacent to a protospacer are tested for cleavage by a CRISPR nuclease in bacteria (e.g., E. coli). The plasmids may contain, for example, a polynucleotide that confers antibiotic resistance adjacent to the randomized PAM sequence. Sequences that are not cleaved upon exposure to CRISPR-Cas nuclease allow cells to survive in the presence of antibiotics due to the presence of an antibiotic resistance gene, while plasmids with targetable PAMs are cleaved and depleted from the library due to cell death. Sequencing of the surviving (uncut) population of plasmids allows for the calculation of a post-selection PAM depletion value, which is compared to a library not exposed to CRISPR-Cas nuclease. The sequences that are depleted from the sequence pool in the experimental library contain the PAM sequence(s) recognized by the CRISPR-Cas nuclease.

[0144] Another method that can be used to identify PAM sequences is the PAM Determination Assay (PAMDA) (Kleinstiver et al. Nat Biotechnol 37:276-282 (2019)). In this case, cleavage is performed outside of living cells. In PAMDA, a single DNA strand is synthesized with a randomized portion of nucleotides next to a defined protospacer sequence. Oligonucleotides are annealed to the 3' end of the synthesized DNA strand and extended with exonuclease minus (-exo) Klenow fragment, polymerizing across the defined and randomized sequences. This produces a duplex library, which is then cleaved with a restriction endonuclease and cloned into bacteria to amplify the total DNA. The plasmid is extracted and linearized with another restriction endonuclease to create a linear template. The template is contacted with a CRISPR-Cas nuclease-guide RNA complex. Only sequences containing the PAM recognized by the CRISPR-Cas nuclease are cleaved. Both the experimental and control libraries (not exposed to CRISPR-Cas nuclease) are then amplified by PCR. Only sequences that are not cleaved by CRISPR-Cas nuclease are amplified. The PCR amplified sequences from the control and experimental libraries (treated with CRISPR-Cas nuclease) are sequenced and compared. The PAM sequence present in the control (not exposed to CRISPR-Cas nuclease) library but not in the experimental library is the PAM sequence recognized by the CRISPR-Cas nuclease (thereby allowing the protospacer to be cleaved).

[0145] A randomized PAM library is prepared for in vitro evaluation of nuclease requirements. The steps described for this method include preparation of an unbiased randomized DNA library containing all PAM sequences to be evaluated, cloning into a plasmid, introduction of the library into bacteria to increase the total amount of starting DNA, extraction of the plasmid, linearization of the plasmid with a restriction enzyme to remove supercoils, exposure of the linearized molecule to CRIPSR-Cas nuclease, amplification of the fragments (e.g., PCR), and final sequence analysis (e.g., next generation sequencing, NGS). The initial steps of unbiased library creation and restriction digestion require at least two restriction enzymes, Klenow extension, and cleaning of the product before ligation into a vector. The use of two to three restriction enzymes typically removes some PAM sequences from the library, introducing bias into the library. Additionally, subsequent Klenow extension and washing steps may also remove PAM sequences, thereby introducing further bias into the library. To avoid sequence loss and produce more complete and unbiased libraries, the present invention provides a novel method for generating randomized PAM libraries using overlapping solid synthetic oligonucleotides (e.g., annealed oligonucleotides) with overhangs instead of restriction endonucleases and Klenow extension (see, e.g., FIG. 1b).The randomized PAM libraries produced using the methods of the present invention can then be used to test the PAM specificity of CRISPR-Cas nucleases with a higher degree of accuracy than was previously available with libraries produced by prior art methods.

[0146] Thus, in some embodiments, the present invention provides a method for constructing a randomized DNA library comprising double-stranded nucleic acid molecules for determining the protospacer adjacent motif (PAM) requirement / specificity of a CRISPR-Cas nuclease having a PAM recognition site at the 5' end of the protospacer, the method comprising the steps of preparing two or more double-stranded nucleic acid molecules: (a) synthesizing a non-target oligonucleotide (first) strand and a target oligonucleotide (second) strand for each of the two or more double-stranded nucleic acid molecules, wherein the non-target oligonucleotides are The nucleotide strand comprises from 5' to 3': (i) a first sequence having about 5 to about 15 nucleotides (e.g., about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 nucleotides, and any range therein); (ii) a second sequence having at least four randomized nucleotides (e.g., at least 4, 5, 6, 7, 8, 9, 10, or more, and any range therein); (ii) a second sequence having about 16 to about 25 nucleotides (e.g., about 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 nucleotides); and and any range therein), and (iv) a third sequence having about 5 to about 20 nucleotides (e.g., about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 nucleotides, and any range therein), wherein the first sequence having about 5 to 15 nucleotides in (i) is immediately adjacent to the 5' terminus of the second sequence in (ii), the second sequence in (ii) is immediately adjacent to the 5' terminus of the protospacer sequence in (iii), and the protospacer sequence is immediately adjacent to the 5' terminus of the third sequence in (iv). the target oligonucleotide (second) strand is complementary to the non-target oligonucleotide strand; and (b) annealing the non-target oligonucleotide strand to the complementary target oligonucleotide strand to produce double-stranded nucleic acid molecules, where the first sequence comprises a restriction site (at its 5' end) and the third sequence comprises a restriction site (at its 3' end), where the first sequence (i), the protospacer sequence (iii) and the third sequence (iv) of each of the two or more double-stranded nucleic acid molecules are identical, thereby constructing a randomized DNA library comprising double-stranded nucleic acid molecules.In some embodiments, the target strand and / or the non-target strand may be 5' phosphorylated.

[0147] In some embodiments, the present invention provides a method for constructing a randomized DNA library comprising double-stranded nucleic acid molecules for determining the protospacer adjacent motif (PAM) requirement / specificity of a CRISPR-Cas nuclease having a PAM recognition site at the 3' end of the protospacer, the method comprising: preparing two or more double-stranded nucleic acid molecules comprising: (a) synthesizing a non-target oligonucleotide (first) strand and a target oligonucleotide (second) strand for each of the two or more double-stranded nucleic acid molecules, wherein the non-target oligonucleotide strand is 5' to 3' comprising: (i) a first sequence having about 5 to about 20 nucleotides (e.g., about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 nucleotides, and any range therein); (ii) a protospacer sequence having about 16 to about 25 nucleotides (e.g., about 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 nucleotides, and any range therein); (iii) at least four randomized nucleotides (e.g., at least 4, 5, 6, 7, 8, 9, 10, or more); a second sequence having from about 5 to about 15 nucleotides (e.g., from about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15 nucleotides, and any range therein), and (iv) a third sequence having from about 5 to about 15 nucleotides (e.g., from about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15 nucleotides, and any range therein), wherein the first sequence having from about 5 to about 20 nucleotides of (i) is immediately adjacent to the 5'-terminus of the protospacer sequence of (ii), the second sequence of (iii) is immediately adjacent to the 3'-terminus of the protospacer sequence of (iii), and the third sequence of (iv) is immediately adjacent to the 3'-terminus of the second sequence of (iii); the (second) strand of nucleotides is complementary to the non-target oligonucleotide strand; and (b) annealing the non-target oligonucleotide strand to the complementary target oligonucleotide strand to produce double-stranded nucleic acid molecules, wherein the first sequence (i) comprises a restriction site (at its 5' end) and the third sequence (iv) comprises a restriction site (at its 3' end), wherein the first sequence (i), the protospacer sequence (ii) and the third sequence (iv) of each of the two or more double-stranded nucleic acid molecules are identical, thereby constructing a randomized DNA library comprising double-stranded nucleic acid molecules.In some embodiments, the target strand and / or the non-target strand may be 5' phosphorylated.

[0148] In some embodiments, the double-stranded nucleic acid molecule can be ligated into a vector to produce a vector comprising a randomized DNA library. In some embodiments, the vector can be a high copy number vector. In some embodiments, the randomized DNA library can be amplified, for example, by introducing the vector comprising the randomized DNA library into one or more bacterial cells and culturing the one or more bacterial cells. In some embodiments, the vector comprising the randomized DNA library can be isolated from one or more bacterial cells after culturing. The isolated vector can then be linearized (e.g., by contacting the vector with one or more restriction enzymes; e.g., ScaI or PfoI) for use in, for example, analyzing the PAM recognition specificity of CRISPR-Cas nuclease. In some embodiments, Pfo1 can be used to linearize the isolated vector.

[0149] In some embodiments, a randomized DNA library can be provided to determine the protospacer adjacent motif (PAM) requirement / specificity of a CRISPR-Cas nuclease having a PAM recognition site at the 5' end of the protospacer, the randomized DNA library comprising two or more double-stranded nucleic acid molecules, each of which comprises (a) a non-target oligonucleotide (first) strand and a target oligonucleotide (second) strand, wherein the non-target oligonucleotide strand comprises, from 5' to 3', (i) about 5 to about (ii) a first sequence having 15 nucleotides (e.g., about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 nucleotides, and any range therein); (iii) a second sequence having at least 4 randomized nucleotides (e.g., at least 4, 5, 6, 7, 8, 9, 10, or more, and any range therein); (iv) a second sequence having about 16 to about 25 nucleotides (e.g., about 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 nucleotides, and any range therein); and (iv) a third sequence having about 5 to about 20 nucleotides (e.g., about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 nucleotides, and any range therein), wherein the first sequence having about 5 to 15 nucleotides in (i) is immediately adjacent to the 5'-terminus of the second sequence in (ii), and the second sequence in (ii) is immediately adjacent to the 5'-terminus of the protospacer sequence in (iii), and the protospacer sequence is (iv). (a) the target oligonucleotide (second) strand is immediately adjacent to the 5' end of a third sequence of the target oligonucleotide; the target oligonucleotide (second) strand is complementary to the non-target oligonucleotide strand; and (b) the non-target oligonucleotide strand anneals to the complementary target oligonucleotide strand to produce a double-stranded nucleic acid molecule, where the first sequence comprises a restriction site (at its 5' end) and the third sequence comprises a restriction site (at its 3' end), and where the first sequence (i), the protospacer sequence (iii) and the third sequence (iv) of each of the two or more double-stranded nucleic acid molecules are identical. In some embodiments, the target strand and / or the non-target strand may be 5' phosphorylated.

[0150] In some embodiments, a randomized DNA library for determining the protospacer adjacent motif (PAM) requirement / specificity of a CRISPR-Cas nuclease having a PAM recognition site at the 3' end of the protospacer may be provided, the randomized DNA library comprising two or more double-stranded nucleic acid molecules, each of which comprises (a) a non-target oligonucleotide (first) strand and a target oligonucleotide (second) strand, wherein the non-target oligonucleotide strand comprises, from 5' to 3', (i) about 5 to about 20 nucleotides. (ii) a first sequence having at least about 16 to about 25 nucleotides (e.g., about 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 nucleotides, and any range therein); (iii) a protospacer sequence having at least four randomized nucleotides (e.g., at least 4, 5, 6, 7, 8, 9, 10, or more); and (iv) a third sequence having from about 5 to about 15 nucleotides (e.g., about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15 nucleotides, and any range therein), wherein the first sequence having from about 5 to about 20 nucleotides in (i) is immediately adjacent to the 5'-terminus of the protospacer sequence in (ii), the second sequence in (iii) is immediately adjacent to the 3'-terminus of the protospacer sequence in (iii), and the third sequence in (iv) is from about 5 to about 15 nucleotides (e.g., about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15 nucleotides, and any range therein), (i) immediately adjacent to the 3' end of the second sequence; the target oligonucleotide (second) strand is complementary to the non-target oligonucleotide strand; and (b) the non-target oligonucleotide strand anneals to the complementary target oligonucleotide strand to produce a double-stranded nucleic acid molecule, where the first sequence comprises a restriction site (at its 5' end) and the third sequence comprises a restriction site (at its 3' end), and where the first sequence (i), the protospacer sequence (ii) and the third sequence (iv) of each of the two or more double-stranded nucleic acid molecules are identical. In some embodiments, the target strand and / or the non-target strand may be 5' phosphorylated.

[0151] In some embodiments, the present invention provides a method for determining the protospacer adjacent motif (PAM) specificity of a CRISPR-Cas nuclease, the method comprising the steps of contacting the CRISPR-Cas nuclease with the randomized DNA library of the present invention; and sequencing the double-stranded nucleic acid molecules of the randomized DNA library before (e.g., as a control) and after contact with the CRISPR-Cas nuclease, wherein double-stranded nucleic acid molecules that are present in the randomized DNA library before contact with the CRISPR-Cas nuclease but not present in the randomized DNA library after contact with the CRISPR-Cas nuclease identify the PAM recognition sequence of the CRISPR-Cas nuclease, thereby determining the PAM specificity of the CRISPR-Cas nuclease.

[0152] In some embodiments, a method for determining the protospacer adjacent motif (PAM) specificity of a CRISPR-Cas nuclease comprises the following steps: contacting the CRISPR-Cas nuclease with the randomized DNA library of the present invention; sequencing the double-stranded nucleic acid molecules of the randomized DNA library before (e.g., as a control) and after contact with the CRISPR-Cas nuclease, and identifying the PAM recognition sequence of the nuclease, wherein the step of identifying comprises comparing double-stranded nucleic acid molecules present in the library before contact with the CRISPR-Cas nuclease with double-stranded nucleic acid molecules present in the library after contact with the CRISPR-Cas nuclease, wherein double-stranded nucleic acid molecules that are present in the randomized DNA library before contact with the CRISPR-Cas nuclease but not present in the randomized DNA library after contact with the CRISPR-Cas nuclease identify the PAM specificity of the CRISPR-Cas nuclease.

[0153] The result of sequencing the randomized library before contacting can serve as a control for the result of sequencing after contacting.In some embodiments, determining the PAM specificity of CRISPR-Cas nuclease can comprise performing nucleic acid sequencing.In some embodiments, sequencing can comprise next generation sequencing (NGS).

[0154] Any CRISPR-Cas nuclease can be used in the methods of the present invention to modify PAM recognition specificity. Thus, CRISPR-Cas nucleases that can be modified to have different PAM specificity compared to the wild type include, but are not limited to, Cas9, C2c1, C2c3, Cas12a (also called Cpf1), Cas12b, Cas12c, Cas12d, Cas12e, Cas13a, Cas13b, Cas13c, Cas13d, Casl, CaslB, Cas2, Cas3, Cas3', Cas3'', Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (with Csnl and Csx12), and Csx13 (with Csx13). Also known as Cas10, Csyl, Csy2, Csy3, Csel, Cse2, Cscl, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmrl, Cmr3, Cmr4, Cmr5, Cmr6, Csbl, Csb2, Csb3, Csxl7, Csxl4, Csx10, Csx16, CsaX, Csx3, Csxl, Csxl5, Csfl, Csf2, Csf3, Csf4(dinG), and / or Csf5 polypeptides or domains.

[0155] Cas12a is a type V Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR)-Cas nuclease originally identified in Prevotella and Francisella species. Cas12a (previously called Cpf1) differs in several ways from the better known type II CRISPR Cas9 nuclease. For example, Cas9 recognizes a G-rich protospacer adjacent motif (PAM) (3'-NGG) that is 3' to its guide RNA (gRNA, sgRNA) binding site (protospacer, target nucleic acid, target DNA), whereas Cas12a recognizes a T-rich PAM (5'-TTN, 5'-TTTN) that is located 5' to the binding site (protospacer, target nucleic acid, target DNA). In fact, the orientation in which Cas9 and Cas12a bind their guide RNAs is almost opposite with respect to their N- and C-termini. Furthermore, the Cas12a enzyme uses a single guide RNA (gRNA, CRISPR array, crRNA) rather than the dual guide RNAs (sgRNA (e.g., crRNA and tracrRNA)) found in the native Cas9 system, and Cas12a processes its own gRNA. In addition, Cas12a nuclease activity produces protruding DNA double-stranded breaks instead of the blunt ends generated by Cas9 nuclease activity, and Cas12a relies on a single RuvC domain to cleave both DNA strands, whereas Cas9 utilizes the HNH and RuvC domains for cleavage.

[0156] The CRISPR Cas12a polypeptide or CRISPR Cas12a domain useful in the present invention may be any known or later identified Cas12a nuclease (see, e.g., U.S. Patent No. 9,790,490, incorporated by reference for its disclosure of Cpf1 (Cas12a) sequences). The term "Cas12a", "Cas12a polypeptide" or "Cas12a domain" refers to an RNA-guided nuclease comprising a Cas12a polypeptide, or a fragment thereof, which includes the guide nucleic acid binding domain of Cas12a, and / or the active, inactive, or partially active DNA cleavage domain of Cas12a. In some embodiments, the Cas12a useful in the present invention may include a mutation in the nuclease active site (e.g., the RuvC site of the Cas12a domain). A Cas12a domain or Cas12a polypeptide that has a mutation in its nuclease active site and therefore no longer contains nuclease activity is generally referred to as a deadCas12a (e.g., dCas12a). In some embodiments, a Cas12a domain or Cas12a polypeptide that has a mutation in its nuclease active site may have impaired / reduced activity (e.g., nickase activity) compared to an identical Cas12a polypeptide that does not have the same mutation.

[0157] In some embodiments, the Cas12a domain can include, but is not limited to, the amino acid sequence of any one of SEQ ID NOs: 1-17 (e.g., SEQ ID NOs: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, and / or 17), or a polynucleotide encoding same. In some embodiments, the fusion protein of the invention can include a Cas12a domain from Lachnospiraceae bacterium ND2006 Cas12a (LbCas12a) (e.g., SEQ ID NO: 1).

[0158] A CRISPR Cas9 polypeptide or CRISPR Cas9 domain useful in the present invention may be any known or later identified Cas9 nuclease. In some embodiments, a Cas9 polypeptide useful in the present invention comprises at least 70% identity (e.g., about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, etc.) to the amino acid sequence of any known Cas9. CRISPR-Cas9 systems are well known in the art and include, but are not limited to, Cas9 polypeptides from Legionella pneumophila str. Paris, Streptococcus thermophilus CNRZ1066, Streptococcus pyogenes MI, or Neisseria lactamica 020-06.

[0159] Other nucleases that may be useful in the present invention to identify novel PAM recognition sequences include, but are not limited to, C2c1, C2c3, Cas12b, Cas12c, Cas12d, Cas12e, Cas13a, Cas13b, Cas13c, Cas13d, Casl, CaslB, Cas2, Cas3, Cas3', Cas3'', Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csnl and Csx12), Csnl, Csx13 ... including as10, Csyl, Csy2, Csy3, Csel, Cse2, Cscl, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmrl, Cmr3, Cmr4, Cmr5, Cmr6, Csbl, Csb2, Csb3, Csxl7, Csxl4, Csx10, Csx16, CsaX, Csx3, Csxl, Csxl5, Csfl, Csf2, Csf3, Csf4(dinG), and / or Csf5.

[0160] The present invention will now be described with reference to the following examples. It should be understood that these examples are not intended to limit the scope of the claimed invention, but rather are intended to be illustrative of specific embodiments. Any variations in the exemplified methods that occur to those skilled in the art are intended to be included within the scope of the present invention. EXAMPLES

[0161] Example 1. Randomized Library An example of the method of the present invention for efficient and cost-effective generation of libraries for an efficient and cost-effective in vitro cleavage assay (PAM determination assay (PAMDA)) is provided. Two libraries were generated for protospacers 1 and 2 (see Table 1). Oligonucleotides with randomized 5 nucleotide sequences at the 5' end were synthesized and verified to occupy equal molar ratios of each protospacer sequence (Integrated DNA Technologies) (Table 1). Oligonucleotides for protospacer 1 (PM0518, PM0519) and oligonucleotides for protospacer 2 (PM0520, PM0521) were annealed by placing the mixture in a thermal cycler at 95°C for 5 minutes and cooling down to 25°C / room temperature at 0.1°C / sec.

[0162] JPEG2025087754000003.jpg97166

[0163] The annealed double-stranded fragment was directly ligated into SphI and EcoRI digested pUC19 vector. The ligated protospacer construct was transformed into XL1-Blue electrocompetent E. coli cells (Agilent) and allowed to recover in 1 ml SOC medium at 37°C for 1 h. Carbenicillin plates were used to check for the presence of the ligated product in E. coli cells. The transformed E. coli cells were grown in LB broth (200 ml) supplemented with carbenicillin (50 mg / mL) for 16 h. Plasmids containing the protospacer construct were purified using Zymo midiprep kit. Plasmids / vectors were subjected to deep sequencing to calculate the frequency of A / T / G / C at each PAM position using Illumina Miseq.

[0164] This method can be used to generate libraries for PAM determination using any protospacer oligonucleotide(s) of choice, where the annealed oligonucleotides may contain any suitable restriction site selected to maintain a full complement of the PAM sequence in the library.

[0165] Example 2 Lachnospiraceae bacterium ND2006 Cpf1 (LbCpf1) requires a very specific protospacer adjacent motif (PAM). The "TTTV" sequence occurs only about 1 out of 85 bases compared to random nucleotides. This contrasts with the relative promiscuity of NGG for SpCas9, which occurs about 1 out of 16 bases in random DNA, TTN for AaC2c1, which occurs about 1 out of 16 bases in random DNA, and xCas9 / Cas9-NG, whose NG PAM requirement occurs about 1 out of 4 bases. Cpf1 PAM is much less abundant than Cas9 PAM in maize and soybean genes (Figure 2). Furthermore, adenine and cytosine (current targets for base editors) are much less accessible to LbCpf1 based on their stringent PAM requirements (Figure 3).

[0166] The stringency shown in Figure 3 for CRISPR-Cas nucleases greatly reduces the generation of potential targets and new traits. The present invention relates to the creation of CRISPR-Cas nucleases with improved ratios of accessible PAM sequences, in particular LbCpf1 (Cas12a) nucleases (e.g., nucleases with PAM recognition sites occurring at a ratio of about 1:4 or better). Such engineered Cas12a PAM mutants can be used as nucleases (for NHEJ or HDR applications) or inactivated versions can be used as genome recognition elements in genome editing tools.

[0167] PAMDA assay The PAM determination assay (PAMDA) is useful to test the PAM requirement for a CRISPR enzyme with unknown PAM recognition. There is an in vitro assay that takes advantage of the ability of CRISPR-Cas nucleases to cleave the target sequence only after successful PAM binding. Briefly, a library of DNA substrates with randomized PAM sequences is incubated with CRISPR nuclease, and then the DNA is amplified by PCR. Only intact fragments (e.g., those not recognized by the nuclease) are amplified. Cleaved fragments (those recognized by the nuclease) are not amplified. DNA from both the library exposed to the nuclease and the control library (not exposed to the nuclease) are sequenced. The two sets of sequencing results are compared to determine which sequences were cleaved and therefore not present in the assembly of the sequence after exposure to the nuclease (see Figure 4 for an example). A modified PAMDA using multiple time points was used to determine PAM binding and subsequent cleavage.

[0168] LbCpf1 mutagenesis 186 point mutations (Table 2) were designed and individually tested in the PAMDA assay as described herein. Successful engineering could alter the PAM recognition sequence to generate a novel PAM-recognizing LbCpf1, or relax the PAM stringency to produce a more promiscuous LbCpf1.

[0169] JPEG2025087754000004.jpg211166JPEG2025087754000005.jpg171166

[0170] In addition to individual mutations, combinations of mutations that alter PAM recognition will be combined and evaluated by PAMDA to provide second generation LbCpf1 mutations.

[0171] Example 3. 186 mutations were tested using three methods: (1) An in vitro method known as the PAMDA assay (Kleinstiver et al. Nat Biotechnol 37:276-282 (2019)) uses purified proteins and plasmid libraries to test each point mutation across the library. Depletion of library members was scored using next-generation sequencing (NGS). Depletion was calculated relative to the library itself (to determine absolute activity against a particular PAM) or relative to cleavage by wild-type LbCas12a (to determine whether the mutation confers new PAM recognition compared to the wild type). (2) A bacterial method known as PAM-SCANR (Leenay et al. Mol Cell 62:137-147 (2016)) uses a library in Escherichia coli to test the binding of Cas12a mutations to 256 possible PAM NNNN variants. It tests only binding, not cleavage. Since the mutations created were not anywhere near the catalytic region, binding was predicted to reflect cleavage as well (later verified in a 293T assay). The advantage of PAM-SCANR is not only the ability to rapidly test point mutations, but also the ability to test combinations of amino acid point mutations in a rapid and accurate manner. This assay can be more stringent than in vitro cleavage assays. (3) Indel assay in human HEK293T cells. This assay provides valuable eukaryotic indel data. To obtain insertions and deletions in eukaryotes, multiple criteria must be met: the CRISPR enzyme must be expressed and stable within the cell, the crRNA must be expressed and correctly processed, the protein:RNA complex must form, the complex must be stable, the complex must translocate into the nucleus in sufficient quantity, the target DNA must be accessible, the DNA must be well-targeted by a specific guide RNA design, and double-strand breaks must occur at a high enough rate to result in occasional DNA repair errors by insertion or deletion (indel). This makes the eukaryotic assay the most stringent assay in this test. Due to the low throughput of the experiment, several dozen PAMs were tested for each of the three point mutants described below, rather than all 256. Since specific guides are often ineffective due to target accessibility, three different targets were selected for each combination of PAM mutants to avoid false negatives.

[0172] 1. In vitro determination of PAM binding and cleavage Construction of a library based on the PAM plasmid A DNA library consisting of 5 random nucleotides directly 5’ of the 23-nucleotide spacer sequence was prepared. LbCas12a is known to have a 4-nucleotide protospacer adjacent motif (PAM), but the inventors decided to use 5 random nucleotides instead of 4 to allow for replication within the experiment. The spacer sequence used was 5’-GGAATCCCTTCTGCAGCACCTGG (SEQ ID NO: 30). The library contained the sequence 5’-NNNNNGGAATCCCTTCTGCAGCACCTGG (SEQ ID NO: 36). Having 5 random nucleotides results in 1024 possible PAMs assayed in this library.

[0173] The inventors used a novel method to generate this library. Instead of using a single randomized pool of PAM-spacer fusions and using polymerase to generate the complementary strand as previously described (Kleinstiver et al. Nat Biotechnol 37,:276-282 (2019)), the inventors chose a more direct method. Two 5'-phosphorylated sequences were synthesized: 5’phos / CGATGTNNNNNGGAATCCCTTCTGCAGCACCTGGGCGCAGGTCACGAGG (SEQ ID NO: 32) and AATTCCTCGTGACCTGCGCCCAGGTGCTGCAGAAGGGATTCCNNNNNACATCGCATG / 5’phos (SEQ ID NO: 35).

[0174] Upon heating and annealing, the complementary sequences between the two NNNNN sequences anneal, and the resulting ends have overhangs corresponding to the overhangs generated by SphI and EcoRI restriction endonucleases. The two oligonucleotides were annealed in a thermal cycler at 95 °C for 5 minutes at an equal molar ratio and cooled to 25 °C / room temperature at 0.1 °C / second.

[0175] The annealed double-stranded fragment was directly ligated into a pUC19 vector digested with SphI and EcoRI. XL1-Blue electrocompetent E. coli cells (Agilent) were transformed with the ligated spacer construct and recovered for 1 hour at 37 °C in 1 ml of glucose-containing Super Optimal Broth (SOC) medium. A portion of the aliquot was plated on carbenicillin plates to confirm the presence of the ligated product. The remaining transformed cells were grown for 16 hours in 200 ml of Luria Broth (LB) supplemented with 50 mg / mL carbenicillin. The spacer plasmid was purified using a plasmid midiprep kit (Zymo Research).

[0176] Verification of the PAM library The spacer vector was subjected to deep array analysis, and the frequency of A / T / G / C at each position of the PAM was calculated using an Illumina MiSeq according to the manufacturer's protocol. Briefly, 10 ng of DNA was used as a template for PCR. Phasing gene-specific forward and reverse PCR primers were designed to amplify across the target site. Amplicon libraries were produced using a two-step PCR method, where primary PCR using 5' tails enabled secondary PCR to attach Illumina i5 and i7 adapter sequences and barcodes for sorting multiplexed samples. PCR amplification was performed using the following parameters: 98 °C for 30 s; 25 cycles for PCR1 and 8 cycles for PCR2 (98 °C for 10 s, 55 °C for 20 s, 72 °C for 30 s); 72 °C for 5 min; hold at 12 °C. The PCR reaction was performed using Q5 High-Fidelity DNA Polymerase (New England BioLabs, Beverly, MA, United States). Secondary PCR amplicon samples were individually purified using AMPure XP beads according to the manufacturer's (Beckman Coulter, Brea, CA, United States) instructions; all purified samples were quantified using a plate reader, pooled at equal molar ratios, and run on an AATI Fragment Analyzer (Agilent Technologies, Palo Alto, CA, United States). The pooled amplicon libraries were sequenced on an Illumina MiSeq (2X250 paired-end) using the MiSeq Reagent Kit v2 (Illumina, San Diego, CA, United States).

[0177] Three separate reads were produced for the library and averaged. The resulting 1024 library members had an average read count of 39 reads and a standard deviation of 11.9 reads. The maximum number of average reads for any PAM sequence was 74 and the minimum was 12. The PAM counts followed a normal distribution (Figure 5).

[0178] Cloning of the LbCas12a mutant A DNA cassette consisting of the LbCas12a sequence followed by the nucleoplasmin NLS and a 6x histidine tag was synthesized (GeneWiz) (SEQ ID NO: 52) and cloned between NcoI and XhoI in the pET28a vector to generate pWISE450 (SEQ ID NO: 53). Additional glycine was added to the sequence between Met-1 and Ser-2 to facilitate cloning. Throughout this document, the numbering excludes this extra glycine. As a result, 186 different amino acid point mutations (Table 2) were generated using a similar strategy that resulted in 186 different plasmid vectors.

[0179] Expression and purification of the LbCas12a mutant Using the glycerol stock of each mutant in BL21 Star(DE3) cells (ThermoFisher Scientific), 1 mL of medium containing 50 μg / mL kanamycin was inoculated in a 24-well block. The culture was sealed with an AirPore tape sheet (Qiagen) and grown overnight with shaking at 37°C. The next morning, 100 μL of the overnight culture was inoculated into 4 mL of ZYP autoinduction medium containing kanamycin and incubated with shaking at 37°C until the OD600nm range of 0.2 - 0.5. The temperature was lowered to 18°C and the culture was grown overnight for protein expression. The cells were harvested by centrifugation and the pellet was stored at -80°C.

[0180] The following buffers were used for cell lysis and purification. A lysis buffer containing a non-ionic detergent, a solubilizer, a reducing agent, a protease inhibitor, a buffer, and salts. This solution was able to lyse bacteria, reduce viscosity, and enable downstream purification of enzymes free of interfering nucleases. Buffer A consisted of 20 mM Hepes-KOH (pH 7.5), 0.5 M NaCl, 10% glycerol, 2 mM TCEP, and 10 mM imidazole (pH 7.5). Buffer B was the same as Buffer A but also contained 20 mM imidazole. Buffer C contained 20 mM Hepes-KOH (pH 7.5), 150 mM NaCl, 10% glycerol, 0.5 mM TCEP, and 200 mM imidazole (pH 7.5).

[0181] Purification was performed using a multi-well format. Two stainless steel 5 / 32’’ BBs were added to all wells containing the cell pellet. The pellet was resuspended in 0.5 mL of chilled lysis buffer and incubated for 30 minutes at room temperature with orbital mixing. The crude lysate (0.5 mL) was added to a pre-equilibrated His MultiTrap™ plate (Cytiva LifeSciences). The plate was incubated for 5 minutes at room temperature to bind the protein. The remaining steps were carried out according to the manufacturer's instructions. Briefly, the plate was washed twice with 0.5 mL of Buffer A, then once with 0.5 mL of Buffer B, and then eluted in 0.2 mL of Buffer C. Protein concentration was determined using Pierce™ Coomassie Plus (Bradford) assay reagent. The protein eluate was stored at 4°C.

[0182] Test cleavage of the PAM library by wtLbCas12a Preliminary tests were conducted to evaluate three aspects of the experiment: confirm that the experiment was free of non-specific nucleases, confirm that there was depletion in the NTTTV PAM from the library upon addition of the crRNA guide, and look at the degree of depletion at 15 minutes for the spiked CTTTA sample.

[0183] The reaction conditions for the test depletion were as follows: 27 μL total volume containing nuclease-free water, 3 μL of NEB buffer 2.1 (New England Biolabs), 3 μL stock of 300 nM crRNA (5’-AAUUUCUACUAAGUGUAGAUGGAAUCCCUUCUGCAGCACCUGG-3’ (SEQ ID NO: 62), Synthego Corporation), and 1 μL of purified wtLbCas12a at 1 μM stock were incubated at room temperature for 10 minutes. The reaction was initiated by adding 3 μL of a 10 ng / μL stock. The library was added as is, or 1 μL of the CTTTA-containing plasmid was first added at 0.75 ng / μL. The total volume was 30 μL in all cases. The reaction was incubated at 37 °C for 15 minutes.

[0184] Table 3 provides the results of the experiment. The counts for the library regarding the TTTV sequences were from 219 to 515 counts (column 2), adding wild-type protein purified as described in the absence of crRNA did not result in depletion of library members (column 3), adding crRNA and protein resulted in depletion of all NTTTVs containing the PAM (column 4), spiking the library with CTTTA resulted in approximately 35-fold more CTTTA NGS counts (column 5), and adding wtLbCas12a and crRNA resulted in depletion of all library members (including a reduction of CTTTA from 10,776 to 193 counts). Also shown were NACGA PAM-containing library members that did not show depletion under the tested conditions, as expected since ACGA is not a PAM recognized by LbCas12a. Thus, as shown in Table 3, NTTTV PAM library members were efficiently cleaved and depleted by wtLbCas12a, while PAMs not recognized by wtLbCas12a (NACGA) were not cleaved and depleted.

[0185] JPEG2025087754000006.jpg161166

[0186] The results in Table 3 show that (1) effective and nuclease-free purification of wtLbCAs12a was achieved, (2) the library can be depleted under the conditions tested for members containing the PAM substrate, (3) depletion occurs upon addition of the spacer-targeting crRNA, and (4) the enzyme-crRNA complex outperforms individual library members since a large amount of the CTTTA substrate does not alter substrate depletion.

[0187] Cleavage of the PAM library by the LbCas12a mutant The same reaction conditions were tested for each of the 186 PAM mutants as shown in the test example of wtLbCas12a. Three time points were selected for each mutation: 75, 435, and 900 seconds at 37 °C. Multiple controls of the library alone were included. The products were subjected to Illumina HiSeq analysis (Genewiz). The data are reported in Table 4.

[0188] Processing of the absolute depletion score We, the inventors, hardly observed differences in depletion among the four possibilities for any 5-nucleotide PAM. In other words, for any 4-nucleotide sequence, ANNNN, CNNNN, GNNNN, and TNNNN had similar PAM depletion. This was consistent with what we observed in wild-type LbCAs12a experiments showing that the NTTTV sequences were all depleted in similar amounts regardless of whether N was A, C, G, or T (Table 3). Second, we observed that all three time points of 75, 435, and 900 seconds had similar depletion. This indicated that the reaction was almost complete immediately after 75 seconds at 37°C. Therefore, we were able to effectively obtain 12 data points for each PAM by averaging all four 4ntPAMs from the 5nt library and averaging all three time points. Then we divided that average by the median library value for each PAM. This gave us depletion scores for each of the 186 mutants with respect to each 4-nucleotide PAM. A depletion of 10 indicates that 90% of the four parental plasmid library members with that 4nt PAM were depleted, and a score of 20 indicates 95% depletion.

[0189] The depletion score for wild-type LbCas12a was 9.2 for the TTTV sequence. Thus, any mutant that cleaved the PAM with a score of 9.2 or higher was considered effective using wild-type as a benchmark. For example, Table 4 shows that the mutant LbCas12a-K595Y depleted 45 different PAM 4-mers from the library in vitro more than the cleavage of TTTV-containing sequences by wtLbCas12a. Each of the 186 mutants was scored using this analysis to determine in vitro PAM recognition and cleavage by each mutant. Data containing the recognition sequences are shown in Table 4, thereby showing the PAMDA depletion score of LbCas12a-K595Y that is above the wtLbCas12a score (9.2) for TTTV-containing sequences.

[0190] JPEG2025087754000007.jpg219166 JPEG2025087754000008.jpg217166JPEG2025087754000009.jpg219166JPEG2025087754000010.jpg218166JPEG2025087754000011.jpg219166JPEG2025087754000012.jpg221166JPEG2025087754000013.jpg217166JPEG2025087754000014.jpg219166JPEG2025087754000015.jpg255146

[0191] The inventors observed 36 unique PAM sequences cleaved in vitro using two LbCas12a controls. This is consistent with the observation that wtLbCas12a can recognize and cleave more sequences than just TTTV in vitro. TTCN, CTTN, TCTN, etc. have been shown to be recognized and cleaved by LbCas12a in vitro, while AsCas12a has only been shown to cleave TTTN (Zetsche et al. Cell 163:759-771 (2015)).

[0192] Some of the mutants increased the total number of PAM sequence recognition and cleavage in vitro compared to wtLbCas12a (Table 5). This suggests an overall promiscuity conferred by individual mutations rather than an absolute PAM recognition sequence. Some individual point mutants were more promiscuous than the wild type. For example, T152R recognized 57 different PAMs and K959Y recognized 45.

[0193] JPEG2025087754000016.jpg108166

[0194] Comparison of depletion against wild - type LbCas12a In vitro, wtLbCas12a can recognize and cleave more sequences than TTTV alone (Zetsche et al. Cell 163:759-771 (2015)). TTCN, CTTN, TCTN, etc. have been shown to be recognized and cleaved by LbCas12a in vitro, while AsCas12a has only been shown to cleave TTTN (Zetsche et al. Cell 163:759-771 (2015)). The aim of this study was to expand the PAM recognition of LbCas12a beyond its wild-type capabilities. To achieve this goal, the inventors used an analysis different from the library depletion score alone. These scores are important for determining absolute PAM recognition and cleavage in vitro, but do not easily highlight changes to enzyme-mediated PAM recognition due to introduced point mutations.

[0195] First, as described above, the results of 5-nucleotide depletion collapsed into 4-nucleotide PAMs. Each time point was maintained individually. The total NGS count for each mutant-time point was normalized to 100 counts per PAM to account for differences in loading on the NGS chip. Then, the overall median for each 4nt PAM was compared to each mutant-time point. This provided depletion compared to the wild type rather than depletion compared to the entire library. The results highlight which mutations altered the PAM recognition profile. The inventors took a conservative approach and selected a depletion score of 4 or greater as an indication of new PAM recognition by the mutant. A depletion score of 4 indicates that a particular PAM-containing library member was cleaved 4-fold compared to the wild-type median. For example, if 100 NGS counts remained for the wild type for a PAM with GCGC and 25 counts remained for a particular mutant-time point, a score of 4 was calculated.

[0196] A summary of each of the 186 mutations is shown in Table 6 below. Mutations provided in bold indicate that the mutation recognized and cleaved more than three new PAM sequences with a score above 4 compared to the wild type. Mutations provided in italics indicate that the mutant obtained 1 - 3 new PAM sequences with a score above 4 compared to the wild type. Mutations in normal font (not bold or italic) indicate that the point mutation did not cleave new PAM sequences with a score above 4 compared to the wild type. For a specific amino acid, such as T149, which is near the PAM recognition domain of the protein and despite testing 10 new amino acids, no new PAM recognition was obtained. Another amino acid, such as D156, was found to be a hot spot for engineering new PAM recognition motifs. When aspartic acid 156 was changed to 10 different amino acids, compared to wtLbCAs12a, 7 mutations recognized multiple new PAMs, 1 showed some new PAMs, and 2 did not obtain new PAMs. Generally, any position showing a difference in PAM recognition and cleavage compared to the wild type can be combined into double, triple, or multiple mutations to further modify PAM recognition. In total, 130 / 186 point mutations did not obtain new PAMs with a score above 4 for wtLbCas12a (not in normal font / bold or italic), 40 / 186 obtained many new PAMs (bold font), and 16 / 186 obtained 1 - 3 new PAMs (italic font). Overall, a 30% success rate (56 / 186) indicates that an effective method was used to design new PAM recognition motifs by creating point mutations for LbCas12a.

[0197] Table 6. Summary table of 186 LbCas12a point mutations (reference sequence: SEQ ID NO: 1). JPEG2025087754000017.jpg112166

[0198] Many of the point mutations conferred novel PAM recognition to LbCas12a, enabling these sequences to cleave the DNA in front of them in vitro. Some mutations resulted in increased overall promiscuity, while others that were designed and tested did not show altered recognition and cleavage of wtLbCas12a. In total, 130 / 186 point mutations did not acquire a new PAM for wtLbCas12a above score 4 (Table 6, not in normal font / italic or bold), 40 / 186 acquired many new PAMs (Table 6, bold font), and 16 / 186 acquired 1 - 3 new PAMs (Table 6, italic font). Overall, a 30% success rate (56 / 186) indicates that an effective method was used to design novel PAM recognition motifs by creating point mutations in LbCas12a.

[0199] 2. Determination of the binding of point mutants and combinations in prokaryotes Combinations of individual mutations can alter even more PAM recognition than single mutations. However, such experiments rapidly scale up and many combinations have to be tested. Using only the 40 mutations that result in LbCas12a recognizing more than three new PAMs, a library of double mutants was created, totaling 40 2 i.e., 1,600 enzymes could be tested. Creating a triple mutant library, 40 3That is, 64,000 enzymes are produced, which is not realistic. Therefore, the inventors adopted a bacterial method known as PAM-SCANR (Leenay et al. Mol Cell 62, 137-147 (2016)) to evaluate combinatorial mutations. The inventors used a library in Escherichia coli to test the binding of Cas12a mutations to 256 possible PAM NNNN variants. This assay does not test cleavage but rather tests binding in vivo. Since the mutations created were not anywhere near the catalytic region, binding was predicted to reflect cleavage as well (which was later verified in a 293T assay). The advantage of PAM-SCANR is not only the ability to rapidly test point mutations but also the ability to test combinations of amino acid point mutations in a rapid and accurate manner. It also tends to be more stringent than in vitro cleavage assays.

[0200] Reporter plasmid Plasmid pWISE1963 was used as the base vector for creating reporters each having one of 256 PAMs. The plasmid contains spectinomycin resistance, a ColE1 origin of replication, LacI, and eGFP (under the control of the lac promoter). 256 gene blocks (from just 5’ of the lacI promoter to the lacI gene) containing the fragment between the NotI and SmaI restriction sites were synthesized by Twist Bioscience. Each fragment contained a different 4-mer PAM immediately 5’ of the lacI promoter. Each gene block was cloned into pWISE1963 by restriction and ligation. Clones were selected for each variant and the PAM identity was confirmed by Sanger sequencing.

[0201] CRISPR - Cas plasmid Plasmid pWISE2031 was used as the backbone vector for the construction of all CRISPR-Cas plasmids. The plasmids contain chloramphenicol resistance, the CloDF13 origin of replication, dLbCas12a driven by promoter BbaJ23108, and LbCas12a with a crRNA targeting the lacI promoter driven by the BbaJ23119 promoter. The negative control plasmid, pWISE1961, contains the same components as pWISE2031 except that a non-targeting crRNA was used. Each point mutant and combinatorial mutant (pWISE2984 - pWISE3007) was constructed via site-directed mutagenesis of pWISE2031 (Genewiz).

[0202] Cell line The E. coli cell line JW0336, which contains a chromosomal deletion of the lacI gene, was obtained from Dharmacon Horizon Discovery. Electrocompetent cells were prepared according to the protocol described (Sambrook, J., and Russell, D. W. (2006). Transformation of E. coli by Electroporation. Cold Spring Harb Protoc 2006, pdb.prot3933). E. coli JW0336 was used in all library transformation and cell sorting experiments.

[0203] Preparation of the reporter library 10 ng of each reporter plasmid described in the above section was pooled in a single tube to generate a library for transformation and amplification. One-twentieth (approximately 0.5 ng of each reporter) of the pooled plasmid library was transformed into Super Competent XL1-Blue according to the manufacturer's instructions. After recovery for 1 hour at 37 °C with shaking at 225 rpm, the entire transformants were transferred to 1 L of LB spectinomycin and grown overnight at 37 °C with shaking at 225 rpm. The next day, plasmid DNA was extracted from the overnight culture using the ZymoPURE Plasmid Gigaprep Kit according to the manufacturer's instructions. The DNA was quantified by Nanodrop and used in all subsequent library transformations.

[0204] Library transformation and cell sorting 100 ng of the reporter plasmid library and 100 ng of the Crispr / Cas plasmid were co-transformed by electroporation into 40 μL of JW0336. The transformants were recovered at 37 °C for 1 hour with shaking at 225 rpm. At the end of recovery, 10 μL of the transformants were removed, mixed with 90 μL of LB, and plated on LB agar plates containing chloramphenicol and spectinomycin to determine the transformation efficiency. The remaining amount (990 μL) of the recovery was transferred to an overnight culture containing 29 mL of LB with spectinomycin and chloramphenicol. The culture was grown overnight at 37 °C with shaking at 225 rpm. The next morning, the colonies on the transformation plates were counted to determine the transformation efficiency; all transformations except two showed >2,000 transformants, corresponding to more than 10-fold coverage of the reporter library. The two samples that did not show more than 10-fold coverage were repeated. Then, glycerol stocks of the overnight cultures were prepared and stored at -80 °C, and 6 mL of each culture was miniprepped using the Qiagen Miniprep Kit according to the manufacturer's instructions. These minipreps were labeled "pre-sort" and stored at 4 °C.

[0205] The optical density (OD) of 1 of each overnight library culture was spun down in a tabletop microcentrifuge at 8,000 rpm and 4 °C for 5 minutes. The supernatant was removed with a pipette, and 1 mL of filter-sterilized 1×PBS buffer was added to each tube. The pellet was carefully resuspended by pipetting. Washing with 1×PBS was repeated 2 more times, and after the final resuspension, the cells (in 1×PBS, approximately 10 8 cells / mL) were placed on ice. Each sample was sorted on a Beckman-Coulter MoFlo XDP cell sorter. Gating parameters for cell sorting were set using a negative control (WT-dLbCas12a + non-targeting crRNA + reporter library) and a positive control (WT-dLbCas12a + targeting crRNA + reporter library). Samples were sorted in single-cell purity mode, voltage 425, ssc voltage 535, fsc voltage (gain) 4.0. The typical sorting rate was approximately 4000 events / second. Each sample had a minimum of 1.0×10 6 events; cell sorting was performed until 50,000 GFP-positive events were collected or the sample was depleted. If the sample was depleted, a minimum of 200 GFP-positive events were collected. GFP-positive events were collected into tubes containing 2 mL of LB with spectinomycin and chloramphenicol. After sorting, the samples were diluted to 6 mL with additional LB with spectinomycin and chloramphenicol, and then grown overnight at 37 °C with shaking at 225 rpm.

[0206] Examples of sorting are provided in FIGS. 6-11. FIG. 6 shows the cell sorting results of a negative control containing crRNA that did not target wtLbCas12a and plasmid spacers. Cells sorted from the GFP high sample do not show cells in the sorting fraction (left panel) and a single population of GFP signals indicated by a single peak (right panel). FIG. 7 shows the cell sorting results of crRNA targeting wtLbCas12a and plasmid spacers. Cells sorted from the GFP high sample show cells in the sorting fraction (left panel, GFP hi) and two populations of GFP signals indicated by two main peaks (right panel, GFP neg and GFP High). Also shown is that the cells sorted into GFP high emit fluorescence (lower right panel). FIG. 8 shows the cell sorting results of crRNA targeting LbCas12a-K595Y and plasmid spacers. Cells sorted from the GFP high sample show cells in the sorting fraction (left panel, GFP hi) and a population indicated by two main peaks (right panel, GFP neg and GFP high). FIG. 9 shows the cell sorting results of crRNA targeting the LbCas12a-G532R-K595R double mutant control (Gao et al. Nat Biotechnol 35, nbt.3900 (2017)) and plasmid spacers. Cells sorted from the GFP high sample show cells in the sorting fraction (left panel, GFP hi) and a population indicated by two main peaks (right panel). FIG. 10 shows the cell sorting results of the LbCas12a-T152R-K595Y double mutation, which is one of two combinations of point mutations used with crRNA targeting plasmid spacers. Cells sorted from the GFP high sample show cells in the sorting fraction (left panel, GFP hi) and a population indicated by two main peaks (right panel).Figure 11 shows the cell sorting results of the LbCas12a-T152R-K538W-K595Y triple mutation, which is one of three combinations of point mutations used with a crRNA targeting a plasmid spacer. Cells sorted from the GFP high sample are shown for the cells in the sorting fraction (left, GFP hi) and the indicated population (right, green line).

[0207] Next - generation sequencing The next morning, after sorting, glycerol stocks of each overnight culture were prepared and stored at -80 °C. The remaining amount of each 6 mL culture was miniprepped using the Qiagen miniprep kit according to the manufacturer's instructions. These minipreps were labeled "post-sort" and stored at 4 °C. Minipreps before and after sorting were quantified by Nanodrop, diluted 10-fold, and handed off for sequencing on an Illumina Mi-Seq.

[0208] The spacer vector was subjected to deep array analysis, and the Illumina MiSeq was used according to the manufacturer's protocol to calculate the frequency of A / T / G / C at each position of the PAM. Briefly, 10 ng of DNA was used as a template for PCR. Phasing gene-specific forward and reverse PCR primers were designed to amplify across the target site. Amplicon libraries were produced using a two-step PCR method, where primary PCR using a 5' tail enabled secondary PCR to attach Illumina i5 and i7 adapter sequences and barcodes for sorting multiplexed samples. PCR amplification was performed using the following parameters: 98 °C for 30 s; 25 cycles for PCR1 and 8 cycles for PCR2 (98 °C for 10 s, 55 °C for 20 s, 72 °C for 30 s); 72 °C for 5 min; maintained at 12 °C. The PCR reaction was performed using Q5 High-Fidelity DNA Polymerase (New England BioLabs, Beverly, MA, United States). Secondary PCR amplicon samples were individually purified using AMPure XP beads according to the manufacturer's (Beckman Coulter, Brea, CA, United States) instructions; all purified samples were quantified using a plate reader, pooled at equal molar ratios, and run on an AATI Fragment Analyzer (Agilent Technologies, Palo Alto, CA, United States). The pooled amplicon libraries were sequenced on an Illumina MiSeq (2X250 paired-end) using the MiSeq Reagent Kit v2 (Illumina, San Diego, CA, United States).

[0209] Sequencing results of cells sorted for high fluorescence Two negative control samples containing wtdLbCas12a and non-targeting crRNA were run in the presence of a 256-member reporter library. After normalizing to 1.0, the values for each member of the library were plotted as a histogram (Figure 12). Figure 12 shows the total normalized NGS counts for two separate crRNA-free controls and wild-type dLbCas12a and the reporter library. Two separate samples were analyzed and combined (a total of 512 points representing 256 PAMs x 2). The inventors selected a conservative 1.67 value (the highest count) as the cut-off for these experiments and scored anything above that as PAM binding. The standard deviation was 0.16. Instead of choosing some multiple of the standard deviation as the cut-off, the inventors chose 1.67, the absolute maximum value seen in either of the two negative controls. This gave a very stringent cut-off of more than 10 times the standard deviation of the data. In fact, only three PAM sequences were seen above 1.5.

[0210] The pool before sorting was sequenced in its entirety before sorting. The average reads / PAM were approximately 250 - 500 NGS reads depending on the sample. The pool after sorting for high fluorescence was sequenced and had similar read counts / PAM at approximately 250 - 500 reads / PAM. Then both samples were normalized against the control for small loading differences in each NGS experiment. Then the two values were subtracted and normalized to 1.0. Many PAM sequences were bound by the point mutation library above the 1.67 cut-off (Figure 13, Table 7).

[0211] JPEG2025087754000018.jpg245166JPEG2025087754000019.jpg189166

[0212] The inventors combined the three point mutations T152R, K538W, and K595Y in various combinations to generate double and triple dLbCas12a mutants (T152R+K538W, K538W+K595Y, and T152R+K538W+K595Y). These were compared to a previously described control developed in AsCas12a, known as "RR", in which the LbCas12a mutation corresponds to the G532R+K595R control (Gao et al. Nat Biotechnol 35(8):789-792(2017)). "RR" has been described as being able to generate indels within the TYCV+CCCC sequence in AsCas12a and subsequently in LbCas12a.

[0213] The same methodology was applied to score combinatorial mutations. Pools before and after sorting, each having on average approximately 250 - 500 MiSeq NGS reads per PAM library member, were sequenced. The pools before and after sorting were normalized and subtracted, and the difference was normalized against 1.0. Many PAM sequences were bound combinatorially above the 1.67 cutoff (Figure 14, Table 8).

[0214] JPEG2025087754000020.jpg247166

[0215] Overall analysis of the PAM - SCANR data Wild-type LbCas12a showed strong TTTV binding, and the LbCas12a-G532R-K595R control showed strong TYCV and CCCC binding. This mutation, called "RR", was developed in AsCas12a and shown to bind to TYCV and CCCC (Gao et al. Nat Biotechnol 35(8):789-792 (2017)). However, in vitro, wtLbCas12a was able to recognize and cleave TTCN, CTTN, TCTN, etc., while AsCas12a was only shown to cleave TTTN (Zetsche et al., Cell 163:759-771 (2015)). The inventors hypothesized that the "RR" mutation placed in the LbCas12a context was more promiscuous than when placed in the AsCas12a context, which was what was observed for this control, and LbCas12a-RR recognized 45 sequences. The results of both wild-type and LbCas12a-RR indicate the validity of the selection and sorting parameters.

[0216] The mutations tested clearly showed that novel PAMs were recognized in vivo by the individual point mutations identified in vitro. For example, K595Y bound to 13 PAMs above the 1.67 threshold, 11 of which were not recognized by wtLbCas12a and none of which contained the TTTV sequence known to be bound by Cas12a. Similarly, T152R recognized 15 distinct PAMs, in this case maintaining the wild-type TTTV. All 12 point mutations tested had novel PAM-binding sequences that were different from the dLbCas12a control, each outside of the standard TTTV motif.

[0217] Effect of combinations The inventors have found that even when multiple point mutations are combined, they do not result in a linear addition of the PAM sequences of the point mutations (Figure 15). For example, the combination of the LbCas12a point mutants K538W and K595Y results in the enzyme LbCas12a-K538W-K595Y, which in some cases shares the PAM recognition motif with K538W (vertically shaded) or K595Y (horizontally shaded), but more often results in a completely new PAM recognition sequence (hatched). Using the same example, K538W recognizes AGCT, but K538W+K595Y does not. K595Y recognizes ACGC, but K538W+K595Y does not. CCCC is not recognized by either K538W or K595Y, but the double mutant binds to it with high affinity.

[0218] Generally, combinations of mutations result in more PAM recognition than linear expansion. For example, K538W recognizes 6 PAM sequences and K595Y recognizes 13, but together they recognize 32 sequences (Figure 15). A simple additive effect would result in 19 PAMs for the double mutant, not the 32 that the inventors observed. Furthermore, only 11 of the 32 sequences recognized by the double mutant are recognized by either of the two single mutations. A similar pattern is observed when three mutations are combined (Figure 16). The combination of T152R, K538W, and K595Y results in a triple mutant with PAM recognition different from any of the three individual mutations alone. For example: GGCA, GGCC, GGGC, and GGGG are only recognized when all three mutations are made to LbCas12a. None of these PAMs are bound by either single or double mutations, but are bound only when T152R, K538W, and K595Y are all mutated together.

[0219] Comparison of point - mutant PAM recognition in PAM - SCANR vs in vitro PAMDA Considering the 12 point mutations tested as a whole, PAM-SCANR hits above 1.67 were well represented in the in vitro PAMDA depletion assay. Examples of K595Y and T152R are shown (Figure 17). Figure 17 compares all non-TTTV PAMs (gray boxes) shown by PAN-SCANR to be above the 1.67 score for K595Y (left panel) and T152R (right panel). All (except one) of the PAM-SCANR positive PAMs above the 1.67 cutoff had PAM depletion scores above the 9.2 cutoff in vitro. However, the PAM-SCANR method and analysis were more stringent than the in vitro assay and analysis. For example, 13 different PAMs were sorted, sequenced, and normalized to have values above 1.67 in PAM-SCANR. This is in contrast to the PAMDA assay which identified 45 sequences as being readily cleaved in vitro. This is thought to be a function of the relative concentration inside the cell versus in the test tube, but could also be a function of setting too stringent a cutoff for PAM-SCANR or too lenient a cutoff for the PAMDA assay.

[0220] There is a correlation between datasets indicating that the inventors' engineering of residues distant from the catalytic site affects PAM recognition and binding rather than catalysis. If mutations in these residues affect nuclease activity along with PAM binding, there would be many hits in the PAM-SCANR assay (which measures binding but not cleavage) that would not show cleavage in the PAMDA assay. The inventors did not see that pattern. The inventors have observed that mutations that affected binding (PAM-SCANR) also resulted in cleavage in vitro (PAMDA).

[0221] 3. Determination of binding, cleavage, and indel formation in eukaryotes The inventors selected three mutations, T152R, K538W, and K595Y, and tested their ability to cause insertions or deletions (indels) in eukaryotic HEK293T cells. This assay provides valuable eukaryotic indel data. To obtain insertions and deletions in eukaryotes, multiple criteria must all be met: the CRISPR enzyme must be expressed and stable within the cell, the crRNA must be expressed and correctly processed, the protein:RNA complex must form, the complex must be stable, the complex must translocate into the nucleus in sufficient quantity, the target DNA must be accessible, the DNA must be well-targeted by a particular guide RNA design, and double-strand breaks must occur at a high enough rate to sometimes cause DNA repair mistakes by insertions or deletions (indels). This makes the eukaryotic assay the most stringent assay in this test. Due to the low throughput of the experiments, several dozen PAMs were tested for each of the three point mutants described below, rather than all 256. Since specific guides are often ineffective due to target accessibility, three different targets were selected for each combination of PAM mutants to avoid false negatives.

[0222] Testing of HEK293T cells Eukaryotic HEK293T (ATCC CRL-3216) cells were cultured at 37 °C in 5% CO2 in Dulbecco's modified Eagle's medium supplemented with GlutaMax (ThermoFisher) + 10% (v / v) FBS (FBS). Wild-type and mutant LbCas12 were synthesized using solid-phase synthesis and subsequently cloned into a plasmid behind the CMV promoter. CRISPR RNA (crRNA) was cloned behind the human U6 promoter (Table 9). HEK293T cells were seeded onto 48-well collagen-coated BioCoat plates (Corning). Cells were transfected at approximately 70% confluency. 750 ng of protein plasmid and 250 ng of crRNA expression plasmid were transfected per well using 1.5 μl of Lipofectamine 3000 (ThermoFisher Scientific) according to the manufacturer's protocol. Genomic DNA from transfected cells was harvested 3 days later, indels were detected, and quantified using high-throughput Illumina amplicon sequencing.

[0223] The spacer vector was subjected to deep array analysis, and the frequencies of A / T / G / C at each position of the PAM were calculated using an Illumina MiSeq according to the manufacturer's protocol. Briefly, 10 ng of DNA was used as a template for PCR. Phasing gene-specific forward and reverse PCR primers were designed to amplify across the target site. Amplicon libraries were produced using a two-step PCR method, where primary PCR using a 5' tail enabled secondary PCR to attach Illumina i5 and i7 adapter sequences and barcodes for sorting multiplexed samples. PCR amplification was performed using the following parameters: 98 °C for 30 s; 25 cycles for PCR1 and 8 cycles for PCR2 (98 °C for 10 s, 55 °C for 20 s, 72 °C for 30 s); 72 °C for 5 min; hold at 12 °C. The PCR reaction was performed using Q5 High-Fidelity DNA Polymerase (New England BioLabs, Beverly, MA, United States). Secondary PCR amplicon samples were individually purified using AMPure XP beads according to the manufacturer's (Beckman Coulter, Brea, CA, United States) instructions; all purified samples were quantified using a plate reader, pooled at equal molar ratios, and run on an AATI Fragment Analyzer (Agilent Technologies, Palo Alto, CA, United States). The pooled amplicon libraries were sequenced on an Illumina MiSeq (2X250 paired-end) using the MiSeq Reagent Kit v2 (Illumina, San Diego, CA, United States).

[0224] Wild - type control Wild-type LbCas12a (wtLbCas12a) recognizes TTTV (TTTA, TTTC, and TTTG). The inventors tested the wild-type protein against TTTV using crRNA spacers containing 23-nucleotide spacer targets (Table 9) (Figure 18).

[0225] JPEG2025087754000021.jpg78166

[0226] Protein and target selection There were many point mutations that showed increased PAM accessibility in the PAMDA in vitro assay. Testing all of the effective PAM mutants against an endogenous 293T cell target is inefficiently impossible due to the experimental complexity, cost, and time, considering the inventors' many point mutations and 256 possible 4-nucleotide (nt) PAMs. Therefore, the inventors selected three point mutations to test against a subset of PAMs. The three tested point mutations are T152R, K538W, and K595Y.

[0227] Genomic targets in 293T cells were selected based on their PAM sequences. Genomic targets for the three point mutations were randomly selected without using specific rules, except having an appropriate 4-nt PAM and selecting 23 nucleotides downstream from that PAM. Three different spacers were selected to assay each respective PAM. This is due to the observation that the activity of CRISPR enzymes is target-specific and often unpredictable.

[0228] On average, the inventors observed that about half of the 23-nucleotide wtLbCas12a spacers tested were ineffective despite having the correct PAM TTTV sequence. With only three data points per PAM and the observation that about 50% of the targets did not produce indels, it is more beneficial to evaluate PAM recognition by visualizing the maximum indel percentage per PAM rather than the average for randomly designed spacers (Figs. 19-21). If a larger number of spacers are assayed per PAM, their average editing efficiency can be evaluated using statistical tests.

[0229] The inventors' overall transfection and assay conditions resulted in wild-type LbCAs12a that produced indels in HEK293T cells at approximately 11-26% for TTTC, 10% for TTTA, and 4-10% for TTTG (Fig. 18). These are the previously defined HEK293T target sites and guides from the literature and are thus predicted to be more effective than any randomly selected guide. Despite the random design of the crRNA guides for the mutants, many new Cas12a PAM recognition sites produced indels at a similar rate to the wild type in the TTTV sequence (Figs. 19-21). Any indel above 0.1% is above the noise of the sequencing assay read at 10,000 NGS read depth.

[0230] K595Y was able to generate indels at 25.5% in ACCG, 10.9% in CCGC, 10.1% in TCGC, 9.5% in CCCG, 8.3% in GCGC, 7.8% in CTGG, 6.3% in ACGG, 6.0% in CCCG, 5.3% in TGGC, etc. (Figure 19). All of these numbers, despite being randomly designed, are within the range of the TTTV control for wtLbCas12a. One of the main hallmarks of the Cas12a protein is the recognition of T-rich PAMs (Zetsche et al. Cell 163:759 - 771 (2015)). This limits their usefulness in genome editing technologies. K595Y clearly prefers C- and G-rich PAMs, which expands the usefulness of Cas2a to targets that were previously targets of the Cas9 CRISPR enzyme, which mainly utilized G-rich PAMs (Jinek et al. Science 337, 816 - 821 (2012)). For K595Y, only 31 out of a total of 256 possible 4-nucleotide PAMs were tested in 293T cells (or 12%). There may be many other PAMs that can be recognized by K595Y in eukaryotic cells to generate indels.

[0231] T152R was able to generate indels at 11.5% in CCTC, 10.0% in CCTG, 9.6% in CCCA, 8.4% in GCCCA, 7.2% in GCCC, 5.1% in CTGC, etc. (Figure 20). Interestingly, T152R maintained the TTTV recognition of wtLbCas12a by generating indels at 34.9% in TTTC, 10.2% in TTTA, and 6.2% in TTTG. It also picked up TTTT recognition and generated indels at 8.3%. For T152R, only 22 out of a total of 256 possible 4-nucleotide PAMs were tested in 293T cells (or 9%). There may be many other PAMs that can be recognized by T152R in eukaryotic cells to generate indels.

[0232] As shown in Figure 21, 22 out of 28 PAM targets had activity above 0.1% background, suggesting that 79% of the tested PAMs were recognized and cleaved by this enzyme, although for some applications it may be lower than desired. Six of the tested PAMs had no editing above background for the three selected targets. All three TTTV targets still had good activity at 15.6%, 6.2%, and 5.8% for TTTC, TTTG, and TTTA, respectively. Other PAM sequences with indel formation above 1% included ATTA (3.5%), TTTT (3.2%), TGTC (1.8%), AGCG (1.8%), AGTC (1.6%), AGCA (1.4%), and GGTC (1.1%). This point mutation, when used in combination with T152R and / or K595Y in the PAM-SCANR experiment, produced diverse PAM recognition, but by itself, it bound to relatively few PAMs using that assay. In the future, using double rather than single mutations may be an excellent option for producing indels in HEK293T cells. Similar to the other two point mutations, only 28 out of 256 possible 4-nucleotide PAMs were tested (11%), and this mutant may be able to recognize PAMs or targets not tested here.

[0233] Correlation between HEK293T indels and PAM - SCANR binding For T152R and K595Y, a correlation was observed between the maximum indel percentage and the PAM-SCANR score (Figures 22A - 22B). Figures 22A - 22B show the linear correlation between indel% (max) and the normalized bacterial PAM-SCANR score for LbCas12a-T152R (Figure 22A) and LbCas12a-K595Y (Figure 22B).

[0234] In particular, any point mutation with a PAM-SCANR score above the tested 1.5 produced indels at a rate higher than 5% in 293T cells. This suggests that any mutation tested in a PAM-SCANR experiment with a normalized score above 1.5 (rather than our stringent 1.67 cutoff) may be able to produce indels at a rate useful for most eukaryotic applications.

[0235] The foregoing are illustrative of the invention and are not to be construed as limiting thereof. The invention is defined by the following claims, and equivalents thereof are included herein.

Claims

1. A modified Lachnospiraceae bacterium CRISPR (Clustered Regularly Interspaced Short Palindromic Repeats) Cas12a (LbCas12a) polypeptide, comprising: The modified LbCas12a polypeptide has at least 80% identity to the amino acid sequence of SEQ ID NO:1 (LbCas12a) and the following positions with respect to the position numbering of SEQ ID NO:1: K116, K120, K121, D122, E125, T148, T149, T152, D156, E159, Q529, D535, K538, D541, Y542, L585, K591, M592, K595 , V596, S599, K600, K601, Y616, Y646, W649, and optionally mutations at one or more of the following positions with respect to the position numbering of SEQ ID NO: 1: K116, K120, K121, D122, E125, T152, D156, E159, G532, D535, K538, D541, and / or K595, Modified LbCas12a polypeptide.

2. The modified LbCas12a polypeptide is selected from the group consisting of K116R, K116N, K120R, K120H, K120N, K120T, K120Y, K120Q, K121S, K121T, K121H, K121R, K121G, K121D, K121Q, D122R, D122K, D122H, D122E, D122N, E125R, E125K, E125Q, E125Y, T148H, T148S, T148A, T148C, T149A, T149C, T149S, T149G, T149H, T149P, T149F, T149N, T149D, T149V, T152R, T152K, T152W, T152Y, T152H, T152Q, T152E, T152L, T152F, D156R, D156K, D156Y, D156W, D156Q, D156H, D156I, D156V, D156L, D156E, E159K, E159R, E159H, E159Y, E159Q, Q529N, Q529T, Q529H, Q529A, Q529F, Q529G, Q529G, Q529S, Q529P, Q529W, Q529D, G532D, G532N, G532S, G532H, G532F, G532K, G532R, G532Q, G532A, G532L, G532C, D535N, D535H, D535V, D535T, D535, S D535A, D535W, D535K, K538RK538V, K538Q, K538W, K538Y, K538F, K538H, K538L, K538M, K538C, K538G, K538A, K538P, D541N, D541H, D541R, D541K, D541Y , D541I, D541A, D541S, D541E, Y542R, Y542K, Y542H, Y542Q, Y542F, Y542L, Y542M, Y542P, Y542V, Y542N, Y542T, L585G, L585 H, L585F, K591W, K591F, K591Y, K591H, K591R, K591S, K591A, K591G, K591P, M592R, M592K, M592Q, M592E, M592A, K595R, K59 5Q, K595Y, K595L, K595W, K595H, K595E, K595S, K595D, K595M, V596T, V596H, V596G, V596A, S599G, S599H, S599N, S599D, K60 or one or more of the amino acid mutations: K600R, K600H, K600G, K601R, K601H, K601Q, K601T, Y616K, Y616R, Y616E, Y616F, Y616H, Y646R, Y646E, Y646K, Y646H, Y646Q, Y646W, Y646N, W649H, W649K, W649Y, W649R, W649E, W649S, W649V, and / or W649T, with respect to the position numbering of SEQ ID NO:

1. 116R, K116N, K120Y, K121S, K121R, D122H, D122N, E125K, T152R, T152K, T152Y, T152Q, T152E, T152F, D156R, D156W, D156Q, D156H, D156I, D156V, D156L, D156E, E159K, E159R, G532N, G532S, G532H, G532K, G532R, G532L, D535N, D535H, D535T, D535, S 2. The modified LbCasl2a polypeptide of claim 1, which may comprise one or more of the following amino acid mutations: D535A, D535W, K538R K538V, K538Q, K538W, K538Y, K538F, K538H, K538L, K538M, K538C, K538G, K538A, D541E, K595R, K595Q, K595Y, K595W, K595H, K595S, and / or K595M.

3. The modified LbCas12a polypeptide of claim 1 or 2, wherein the modified LbCas12a polypeptide has altered PAM (protospacer adjacent motif) specificity.

4. The modified LbCas12a polypeptide (e.g., deadLbCas12a, dLbCas12a) of claim 1 or 2, further comprising a mutation within a nuclease active site (e.g., the RuvC domain).

5. A fusion protein comprising a modified LbCas12a polypeptide according to any one of claims 1 to 4 and a linker, which may be a peptide linker.

6. The fusion protein of claim 5, wherein the peptide linker is linked to the C-terminus and / or N-terminus of the modified LbCas12a polypeptide.

7. A fusion protein comprising a modified LbCas12a polypeptide according to any one of claims 1 to 4 and a polypeptide of interest.

8. The fusion protein of claim 7, wherein the polypeptide of interest is linked to the modified LbCas12a polypeptide.

9. The fusion protein of claim 7 or 8, wherein the polypeptide of interest is linked to the C-terminus and / or N-terminus of the modified LbCas12a polypeptide.

10. The fusion protein of any one of claims 7 to 9, wherein the polypeptide of interest is linked to the modified LbCas12a polypeptide via a linker, which may be a peptide linker.

11. The target polypeptide has a deaminase (deamination) activity, a nickase activity, a recombinase activity, a transposase activity, a methylase activity, a glycosylase (DNA glycosylase) activity, a glycosylase inhibitor activity (e.g., uracil-DNA glycosylase inhibitor (UGI)), a demethylase activity, a transcription activation activity, a transcription repression activity, a transcription release factor (transcription release factor), a transcription inhibitor (transcription release factor), a transcription release factor ...

10. The fusion protein of any one of claims 7 to 9, comprising at least one polypeptide or protein domain having a targeting factor activity, a histone modifying activity, a nuclease activity, a single-stranded RNA cleavage activity, a double-stranded RNA cleavage activity, a restriction endonuclease activity (e.g. Fok1), a nucleic acid binding activity, a methyltransferase activity, a DNA repair activity, a DNA damaging activity, a dismutase activity, an alkylating activity, a depurinating activity, an oxidizing activity, a pyrimidine dimer forming activity, an integrase activity, a transposase activity, a polymerase activity, a ligase activity, a helicase activity, and / or a photolyase activity.

12. The fusion protein according to any one of claims 7 to 11, wherein the target polypeptide comprises at least one polypeptide or protein domain having deaminase activity.

13. The fusion protein of claim 12, wherein the at least one polypeptide or protein domain having deaminase activity is a cytosine deaminase domain or an adenine deaminase domain.

14. The fusion protein of claim 13 , wherein the cytosine deaminase domain is an apolipoprotein B mRNA editing catalytic polypeptide-like (APOBEC) domain.

15. The adenine deaminase domain is TadA (tRNA-specific adenosine deaminase) and / or TadA * 14. The fusion protein of claim 13, which is an evolved tRNA-specific adenosine deaminase.

16. The fusion protein according to any one of claims 7 to 14, wherein the at least one polypeptide has glycosylase inhibitor activity, and the at least one polypeptide may be a uracil-DNA glycosylase inhibitor (UGI).

17. A polynucleotide encoding a modified LbCas12a polypeptide according to any one of claims 1 to 4, or a fusion protein according to any one of claims 5 to 16.

18. The polynucleotide of claim 17, wherein the polynucleotide encoding the modified LbCasl2a polypeptide or the polynucleotide encoding the fusion protein is operably associated with a promoter, and the promoter may be a promoter region including an intron.

19. 19. The polynucleotide of claim 17 or 18, wherein the polynucleotide is codon-optimized for expression in an organism.

20. The polynucleotide according to any one of claims 17 to 19, wherein the organism is an animal, a plant, a fungus, an archaea, or a bacterium.

21. A complex comprising a modified LbCas12a polypeptide according to any one of claims 1 to 4, or a fusion protein according to any one of claims 5 to 16, and a guide nucleic acid (e.g., CRISPR RNA, CRISPR DNA, crRNA, crDNA).

22. A nucleic acid construct encoding the complex of claim 21.

23. A composition comprising: (a) a modified LbCas12a polypeptide according to any one of claims 1 to 4, or a fusion protein according to any one of claims 5 to 16, and (b) a guide nucleic acid.

24. An expression cassette or vector comprising a polynucleotide according to any one of claims 17 to 20, or a nucleic acid construct according to claim 22.

25. A V-shaped Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR)-associated (Cas) (CRISPR-Cas) system, Includes: (a) a fusion protein comprising: (i) a modified LbCas12a polypeptide according to any one of claims 1 to 4 or a nucleic acid encoding the modified LbCas12a polypeptide according to any one of claims 1 to 4, and (ii) a polypeptide of interest or a nucleic acid encoding said polypeptide of interest; and (b) a guide nucleic acid (CRISPR RNA, CRISPR DNA, crRNA, crDNA) comprising a spacer sequence and a repeat sequence; wherein the guide nucleic acid is capable of forming a complex with the modified LbCasl2a polypeptide or the fusion protein, and the spacer sequence is capable of hybridizing to a target nucleic acid, thereby guiding the modified LbCasl2a polypeptide and the polypeptide of interest to the target nucleic acid, thereby modifying (e.g., cleaving or editing) or modulating (e.g., transcriptionally modulating) the target nucleic acid; system.

26. The system of claim 25, wherein the modified LbCas12a polypeptide comprises a mutation within a nuclease active site (e.g., the RuvC domain) (e.g., deadCas12a, dCas12a).

27. The system of claim 25 or 26, wherein the polypeptide of interest is linked to the C-terminus and / or N-terminus of the modified LbCas12a polypeptide.

28. The system of any one of claims 25 to 27, wherein the polypeptide of interest is linked to the C-terminus and / or N-terminus of the modified LbCas12a polypeptide via a linker, which may be a peptide linker.

29. The target polypeptide has a deaminase (deamination) activity, a nickase activity, a recombinase activity, a transposase activity, a methylase activity, a glycosylase (DNA glycosylase) activity, a glycosylase inhibitor activity (eg, uracil-DNA glycosylase inhibitor (UGI)).

29. The system of any one of claims 25 to 28, comprising at least one polypeptide or protein domain having demethylase activity, transcription activation activity, transcription repression activity, transcription release factor activity, histone modification activity, nuclease activity, single-stranded RNA cleavage activity, double-stranded RNA cleavage activity, restriction endonuclease activity (e.g., Fok1), nucleic acid binding activity, methyltransferase activity, DNA repair activity, DNA damage activity, dismutase activity, alkylation activity, depurination activity, oxidation activity, pyrimidine dimer formation activity, integrase activity, transposase activity, polymerase activity, ligase activity, helicase activity, and / or photolyase activity.

30. The system according to any one of claims 25 to 29, wherein the target polypeptide comprises at least one polypeptide or protein domain having deaminase activity.

31. The system of claim 30, wherein the at least one polypeptide or protein domain having deaminase activity is a cytosine deaminase domain or an adenine deaminase domain.

32. The system of claim 31 , wherein the cytosine deaminase domain is an apolipoprotein B mRNA editing catalytic polypeptide-like (APOBEC) domain.

33. The adenine deaminase domain is TadA (tRNA-specific adenosine deaminase) and / or TadA * (an evolved tRNA-specific adenosine deaminase).

34. The system according to any one of claims 25 to 32, wherein the target polypeptide comprises a uracil-DNA glycosylase inhibitor (UGI).

35. The system of any one of claims 25 to 34, wherein one or both of (a) and (b) are contained within one or more expression cassettes and / or vectors.

36. A cell comprising a polynucleotide according to any one of claims 17 to 20, a nucleic acid construct according to claim 22, an expression cassette or vector according to claim 24, or a system according to any one of claims 25 to 35.

37. 1. A method for modifying a target nucleic acid, comprising: The target nucleic acid (a) (i) a modified LbCas12a polypeptide according to any one of claims 1 to 4, or a fusion protein according to any one of claims 5 to 16, and (ii) a guide nucleic acid (e.g., CRISPR RNA, CRISPR DNA, crRNA, crDNA); (b) the complex of claim 21 and a guide nucleic acid; (c) (i ) a composition comprising a modified LbCas12a polypeptide according to any one of claims 1 to 4, or a fusion protein according to any one of claims 5 to 16, and (ii) a guide nucleic acid; and / or (d) A system according to any one of claims 25 to 35. contacting the target nucleic acid with method.

38. 1. A method for modifying a target nucleic acid, comprising: a cell or cell-free system containing the target nucleic acid, (a) (i) a polynucleotide according to any one of claims 17 to 20, or an expression cassette or vector comprising same, and (ii) a guide nucleic acid, or an expression cassette or vector comprising same; and / or (b) a nucleic acid construct according to claim 22, or an expression cassette or vector comprising the same; contacting the target nucleic acid with method.

39. 1. A method for editing a target nucleic acid, comprising: The target nucleic acid (a)(i) a fusion protein according to any one of claims 7 to 16, and (a)(ii) a guide nucleic acid; (b) a complex comprising the fusion protein according to any one of claims 7 to 16 and a guide nucleic acid; (c) a composition comprising a fusion protein according to any one of claims 7 to 16 and a guide nucleic acid; and / or (d) a system according to any one of claims 25 to 35; contacting said target nucleic acid with method.

40. 1. A method for editing a target nucleic acid, comprising: a cell or cell-free system containing the target nucleic acid, (a)(i) a polynucleotide encoding a fusion protein according to any one of claims 7 to 16, or an expression cassette or vector comprising the same, and (a)(ii) a guide nucleic acid, or an expression cassette or vector comprising the same; and / or (b) a nucleic acid construct, or an expression cassette or vector comprising the same, encoding a complex comprising a fusion protein according to any one of claims 7 to 16 and a guide nucleic acid; and / or (c) a system according to any one of claims 25 to 35; contacting said target nucleic acid with method.

41. 1. A method for constructing a randomized DNA library comprising double-stranded nucleic acid molecules for determining the protospacer adjacent motif (PAM) requirement / specificity of a CRISPR-Cas nuclease having a PAM recognition site at the 5′ end of the protospacer, comprising: The method includes: Preparing two or more double-stranded nucleic acid molecules comprising the steps of: (a) synthesizing a non-target oligonucleotide (first) strand and a target oligonucleotide (second) strand for each of said two or more double-stranded nucleic acid molecules; wherein the non-target oligonucleotide strand comprises, from 5' to 3', the following: (i) a first sequence having from about 5 to about 15 nucleotides; (ii) a second sequence having at least four randomized nucleotides; (iii) a protospacer sequence comprising about 16 to about 25 nucleotides, and (iv) a third sequence having about 5 to about 20 nucleotides. Including, wherein (i) a first sequence having about 5-15 nucleotides is immediately adjacent to the 5'-terminus of (ii) a second sequence, which (ii) is immediately adjacent to the 5'-terminus of (iii) a protospacer sequence, which in turn is immediately adjacent to the 5'-terminus of (iv) a third sequence; said target oligonucleotide (second) strand is complementary to said non-target oligonucleotide strand; and (b) annealing the non-target oligonucleotide strand to the complementary target oligonucleotide strand to produce a double-stranded nucleic acid molecule; wherein said first sequence comprises a restriction site (at its 5' end) and said third sequence comprises a restriction site (at its 3' end), and wherein said first sequence (i), said protospacer sequence (iii) and said third sequence (iv) of each of said two or more double-stranded nucleic acid molecules are identical, thereby constructing said randomized DNA library comprising double-stranded nucleic acid molecules. method.

42. 1. A method for constructing a randomized DNA library comprising double-stranded nucleic acid molecules for determining the protospacer adjacent motif (PAM) requirement / specificity of a CRISPR-Cas nuclease having a PAM recognition site at the 3′ end of the protospacer, comprising: The method includes: Preparing two or more double-stranded nucleic acid molecules comprising the steps of: (a) synthesizing a non-target oligonucleotide (first) strand and a target oligonucleotide (second) strand for each of said two or more double-stranded nucleic acid molecules; wherein the non-target oligonucleotide strand comprises, from 5' to 3', the following: (i) a first sequence having from about 5 to about 20 nucleotides; (ii) a protospacer sequence comprising about 16 to about 25 nucleotides; (iii) a second sequence having at least four randomized nucleotides; and (iv) a third sequence having about 5 to about 15 nucleotides. Including, wherein (i) a first sequence having about 5-20 nucleotides is immediately adjacent to the 5'-terminus of the protospacer sequence of (ii), (iii) a second sequence is immediately adjacent to the 3'-terminus of the protospacer sequence of (iii), and (iv) a third sequence is immediately adjacent to the 3'-terminus of the second sequence of (iii); said target oligonucleotide (second) strand is complementary to said non-target oligonucleotide strand; and (b) annealing the non-target oligonucleotide strand to the complementary target oligonucleotide strand to produce a double-stranded nucleic acid molecule; wherein said first sequence (i) comprises a restriction site (at its 5' end) and said third sequence (iv) comprises a restriction site (at its 3' end), and wherein said first sequence (i), said protospacer sequence (ii) and said third sequence (iv) of each of said two or more double-stranded nucleic acid molecules are identical, thereby constructing said randomized DNA library comprising double-stranded nucleic acid molecules. method.

43. 43. The method of claim 41 or 42, further comprising ligating the double-stranded nucleic acid molecule into a vector to produce a vector comprising the randomized DNA library.

44. The method of any one of claims 41 to 43, further comprising the step of amplifying the randomized DNA library.

45. 45. The method of claim 44, wherein the amplifying step comprises introducing a vector containing the randomized DNA library into one or more bacterial cells and culturing the one or more bacterial cells.

46. 46. ​​The method of claim 45, further comprising isolating a vector comprising the randomized DNA library from the one or more bacterial cells.

47. 47. The method of claim 46, further comprising the step of linearizing said vector to provide said randomized DNA library.

48. 48. The method of claim 47, wherein the linearizing step comprises contacting the vector with a restriction enzyme.

49. 49. The method of claim 48, wherein the restriction enzyme is ScaI or PfoI.

50. A randomized DNA library produced by the method of any one of claims 41 to 49.

51. A randomized DNA library can be provided to determine the protospacer adjacent motif (PAM) requirement / specificity of CRISPR-Cas nucleases having a PAM recognition site at the 5' end of the protospacer, The randomized DNA library comprises two or more double-stranded nucleic acid molecules, each comprising: (a) a non-target oligonucleotide (first) strand and a target oligonucleotide (second) strand; wherein the non-target oligonucleotide strand comprises, from 5' to 3': (i) a first sequence having about 5 to about 15 nucleotides (e.g., about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 nucleotides, and any range or value therein); (ii) a second sequence having at least four randomized nucleotides (e.g., at least 4, 5, 6, 7, 8, 9, 10, or more, and any range or value therein); (iii) a protospacer sequence comprising about 16 to about 25 nucleotides, e.g., about 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 nucleotides; and (iv) a third sequence having about 5 to about 20 nucleotides (e.g., about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 nucleotides, and any range or value therein); wherein (i) a first sequence having about 5-15 nucleotides is immediately adjacent to the 5'-terminus of (ii) a second sequence, which (ii) is immediately adjacent to the 5'-terminus of (iii) a protospacer sequence, which in turn is immediately adjacent to the 5'-terminus of (iv) a third sequence; said target oligonucleotide (second) strand is complementary to said non-target oligonucleotide strand; and (b) the non-target oligonucleotide strand anneals to the complementary target oligonucleotide strand to produce a double-stranded nucleic acid molecule, wherein the first sequence comprises a restriction site (at its 5' end) and the third sequence comprises a restriction site (at its 3' end), and wherein the first sequence (i), the protospacer sequence (iii), and the third sequence (iv) of each of the two or more double-stranded nucleic acid molecules are identical.

52. A randomized DNA library can be provided to determine the protospacer adjacent motif (PAM) requirement / specificity of CRISPR-Cas nucleases having a PAM recognition site at the 3' end of the protospacer, The randomized DNA library comprises two or more double-stranded nucleic acid molecules, each comprising: (a) a non-target oligonucleotide (first) strand and a target oligonucleotide (second) strand; wherein the non-target oligonucleotide strand comprises, from 5' to 3': (i) a first sequence having about 5 to about 20 nucleotides (e.g., about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 nucleotides, and any range or value therein); (ii) a protospacer sequence comprising about 16 to about 25 nucleotides, e.g., about 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 nucleotides; (iii) a second sequence having at least four randomized nucleotides (e.g., at least 4, 5, 6, 7, 8, 9, 10, or more, and any range or value therein), and (iv) a third sequence having about 5 to about 15 nucleotides (e.g., about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15 nucleotides, and any range or value therein); wherein (i) a first sequence having about 5-20 nucleotides is immediately adjacent to the 5'-terminus of the protospacer sequence of (ii), (iii) a second sequence is immediately adjacent to the 3'-terminus of the protospacer sequence of (iii), and (iv) a third sequence is immediately adjacent to the 3'-terminus of the second sequence of (iii); said target oligonucleotide (second) strand is complementary to said non-target oligonucleotide strand; and (b) the non-target oligonucleotide strand anneals to the complementary target oligonucleotide strand to produce a double-stranded nucleic acid molecule, wherein the first sequence comprises a restriction site (at its 5' end) and the third sequence comprises a restriction site (at its 3' end), and wherein the first sequence (i), the protospacer sequence (ii) and the third sequence (iv) of each of the two or more double-stranded nucleic acid molecules are identical.

53. 1. A method for determining protospacer adjacent motif (PAM) specificity of a CRISPR-Cas nuclease, comprising: The method comprises: Contacting the CRISPR-Cas nuclease with the randomized DNA library of claim 51 or 52; and Sequencing the double-stranded nucleic acid molecules of the randomized DNA library before (control) and after contact with the nuclease; Including, wherein double-stranded nucleic acid molecules present in the randomized DNA library prior to contact with the nuclease but not present in the randomized DNA library after contact with the nuclease identify a PAM recognition sequence for the CRISPR-Cas nuclease, thereby determining the PAM specificity of the CRISPR-Cas nuclease. method.

54. 54. The method of claim 53, wherein the sequencing step comprises next generation sequencing.

55. A kit comprising a polynucleotide according to any one of claims 17 to 20 or 22 and / or an expression cassette or vector comprising same, and optionally including instructions for its use.

56. 56. The kit of claim 55, further comprising a CRISPR-Cas12a guide nucleic acid and / or an expression cassette or vector comprising same.

57. 57. The kit of claim 56, wherein the guide nucleic acid comprises a cloning site for cloning a nucleic acid sequence identical or complementary to a target nucleic acid sequence into the backbone of the guide nucleic acid.

58. The kit of any one of claims 55 to 57, wherein the polynucleotide further encodes one or more nuclear localization signals, and the one or more nuclear localization signals are fused to the CRISPR-Casl2a nuclease.

59. 59. The kit of any one of claims 55 to 58, wherein the polynucleotide, expression cassette or vector further encodes one or more selectable markers.

60. The kit of any one of claims 55 to 59, wherein the polynucleotide is an mRNA and encodes one or more introns within the encoded CRISPR-Cas12a nuclease.