Modified Nucleases

JP2024518413A5Inactive Publication Date: 2025-05-13ARTISAN DEV LABS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023568340
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-03-01
Filing Date
2022-05-06
Publication Date
2025-05-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Current CRISPR-Cas editing systems face limitations in efficiency and precision, particularly in terms of PAM specificity, editing efficiency, and off-target rates, which hinder their effectiveness in genome editing applications, especially in commercially relevant situations.

Method used

The development of novel Cas12a-based nucleases that recognize alternative PAM sequences and utilize double-stranded gRNAs, offering improved editing accuracy and reduced off-target effects, along with the integration of nuclear localization signals (NLS) for enhanced cellular localization.

Benefits of technology

The novel Cas12a-based nucleases demonstrate increased efficiency and precision in genome editing, reducing off-target effects and improving the accuracy of targeted gene modifications, making them suitable for applications in various species including humans.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Provided herein are methods and compositions utilizing modified nucleases and / or other components, such as guide nucleic acids and donor templates, for use in CRISPR systems.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Provisional Application No. 63 / 185,315, filed May 6, 2021, and U.S. Provisional Application No. 63 / 315,483, filed March 1, 2022, which applications are incorporated herein by reference. [Background technology]

[0002] Nucleic acid-guided nucleases have become important tools for research and genome engineering. The applicability of these tools can be limited by sequence-specific requirements, expression, or delivery issues.

[0003] Incorporation by Reference All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. Summary of the Invention

[0004] The novel features of the invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention are utilized, and the accompanying drawings. [Brief explanation of the drawings]

[0005] [Figure 1] A diagram of MAD7 containing one or more nuclear localization signals (NLS) is shown. [Figure 2] 1 shows the editing frequency at the DNMT1 locus and cell viability after transfection in T-cell leukemia cells following treatment with one or more guide nucleic acids complexed with MAD7 containing one or more NLSs. [Figure 3]FIG. 1 shows the editing frequency at the DNMT1 locus in T-cell leukemia cells using multiple electroporation programs in combination with SE electroporation buffer. [Figure 4] FIG. 1 shows the editing frequency at the DNMT1 locus in T-cell leukemia cells using multiple electroporation programs in combination with SF electroporation buffer. [Figure 5] FIG. 1 shows the editing frequency at the DNMT1 locus in T-cell leukemia cells using multiple electroporation programs in combination with SG electroporation buffer. [Figure 6] FIG. 1 shows the editing frequency at the DNMT1 locus in T-cell leukemia cells using multiple electroporation programs. [Figure 7] Figure 1 shows the frequency of editing by type at eight loci in T-cell leukemia cells using multiple guide nucleic acids complexed with MAD7 containing one or more NLSs. [Figure 8] 1 shows a comparison of editing efficiency of MAD7-treated T-cell leukemia cells containing one or more guide nucleic acids targeting the DNMT1 locus and control guide nucleic acids combined by editing frequency. [Figure 9] 1 shows the editing frequency per PAM motif in T-cell leukemia cells using multiple guide nucleic acids complexed with MAD7 containing one or more NLSs. [Figure 10A] 1 shows a sequence logo plot for multiple guide nucleic acids combined by compiling their frequency in T-cell leukemia cells when used in complex with MAD7 containing one or more NLSs. [Figure 10B] 1 shows the nucleotide and dinucleotide frequencies for multiple guide nucleic acids bound by editing frequencies in T-cell leukemia cells used when complexed with MAD7 containing one or more NLSs. [Figure 11] 1 shows the frequency of the trinucleotides AAA or UUU bound by editing frequencies in T-cell leukemia cells after treatment with multiple guide nucleic acids complexed with MAD7 containing one or more NLSs. [Figure 12] 1 shows the editing frequencies for both INDEL and frameshift mutations at eight loci in T-cell leukemia cells after treatment with multiple guide nucleic acids complexed with MAD7 containing one or more NLSs. [Figure 13] 1 shows the correlation between INDEL frequency in gNA validation experiments and INDEL formation in gNA screening experiments. [Figure 14] 1 shows the rate of frameshifts to INDELs at eight loci in T-cell leukemia cells after treatment with multiple guide nucleic acids complexed with MAD7 containing one or more NLSs. [Figure 15] Figure 1 shows the INDEL frequency for gNAs containing representative spacer sequences complexed with MAD7 containing one or more NLSs in T-cell leukemia cells at predicted off-target sites. [Figure 16] Figure 1 shows the INDEL frequency for gNAs containing representative spacer sequences complexed with MAD7 containing one or more NLSs in T-cell leukemia cells at predicted off-target sites. [Figure 17] INDEL frequency at the AAVS1 locus in T-cell leukemia cells after treatment with the gNA:MAD7 complex. [Figure 18] GFP insertion efficiency at the AAVS1 locus and cell viability after treatment with multiple primer constructs are shown. [Figure 19] Figure 1 shows GFP insertion efficiency at the AAVS1 locus with increasing concentrations of donor template (e.g., HDRT) and variable homology arm lengths. [Figure 20] Figure 1 shows CAR insertion efficiency at the AAVS1 locus and cell viability with variable homology arm lengths and increasing concentrations of donor template. [Figure 21] A shows the CAR insertion efficiency at the AAVS1 locus in primary T cells. B shows the cell viability at the AAVS1 locus in primary T cells. DETAILED DESCRIPTION OF THE INVENTION

[0006] CRISPR stands for Clustered Regularly Interspaced Short Palindromic Repeats. In palindromic repeats, the sequence of nucleotides is the same in both directions. Each of these palindromic repeats is followed by a short segment of spacer DNA. Small clusters of Cas (CRISPR-associated system) genes are located adjacent to the CRISPR sequences. The CRISPR / Cas system is a prokaryotic immune system that can confer resistance to foreign genetic elements, such as those present in plasmids and phages, providing prokaryotes with a form of acquired immunity. RNA carrying spacer sequences helps the Cas (CRISPR-associated) proteins recognize and cleave exogenous DNA. CRISPR sequences are found in approximately 50% of bacterial genomes, and nearly 90% of sequenced archaeal species have selected for efficient and robust metabolic and regulatory networks that prevent the biosynthesis of unwanted metabolites and optimally allocate resources to maximize overall cellular fitness. The complexity of these networks, which limits approaches for understanding their structure and function, as well as the ability to reprogram cellular networks to modify these systems for various applications, has made progress in this field difficult. Certain approaches to reprogramming cellular networks aim to modify single genes in complex pathways, but modifying a single gene can result in undesired modifications to this or other genes, hindering the identification of the changes necessary to achieve the desired endpoint and complicating the desired endpoint through modification.

[0007] Genome editing and engineering with CRISPR-Cas has had a dramatic impact on biology and biotechnology in general. The CRISPR-Cas editing system requires a polynucleotide-guided nuclease, a guide nucleic acid (gNA) (e.g., a guide RNA (gRNA) that directs the nuclease to cleave a specific region of the genome), and, optionally, a donor DNA cassette (also referred to herein as a donor template or editing sequence) that can be used to repair the cleaved dsDNA and thereby incorporate programmable edits at the desired site. The earliest demonstrations and applications of CRISPR-Cas editing used the Cas9 nuclease and associated gRNA. These systems have been used for gene editing in a variety of species, ranging from bacteria to animals and, in some cases, higher mammalian systems such as humans. However, it has been well documented that key editing parameters, such as protospacer adjacent motif (PAM) specificity, editing efficiency, and off-target rate, among others, are species-, locus-, and nuclease-dependent. There is growing interest in identifying and rapidly characterizing novel nuclease systems that can be used to extend and improve global editing capabilities.

[0008] One version of the CRISPR / Cas system, CRISPR / Cas9, has been modified to provide a useful tool for targeted genome editing. By delivering Cas9 nuclease complexed with synthetic guide RNA (gRNA) into cells, it may be possible to cut / edit the cell's genome at a predetermined location to delete an existing gene and / or add a new gene. Although these systems are useful, they have several important limitations regarding the efficiency and precision of targeted editing, problems with imprecise editing, and obstacles when used in commercially relevant situations such as gene replacement. Therefore, there is a need for an improved nucleic acid-guided nuclease system for directed and precise editing with improved efficiency.

[0009] As used herein, the terms "modulation" and "manipulation" of genome editing can refer to an increase, decrease, upregulation, downregulation, induction, change in editing activity, change in binding, change in cleavage, etc. of one or more of the targeted genes or gene clusters of certain embodiments disclosed herein.

[0010] Certain embodiments of the present disclosure may employ conventional molecular biology, microbiology, and recombinant DNA techniques within the skill of the art, which are fully explained in the literature and understood by those skilled in the art.

[0011] In other embodiments, primers used herein for preparation by conventional techniques may include sequencing primers and amplification primers. In some embodiments, plasmids and oligomers used in conventional techniques may include synthetic oligomers and oligomer cassettes.

[0012] In some embodiments disclosed herein, nucleic acid-guided nuclease systems and methods of use are provided. The nuclease system can include transcripts and other elements involved in the expression of the engineered nucleases disclosed herein, which can include the novel engineered nucleic acid-guided nuclease proteins disclosed herein and sequences encoding guide sequences (gRNAs) or novel gRNAs. In some embodiments, the nucleic acid-guided nuclease system can include at least one CRISPR-associated nucleic acid-guided nuclease construct, the disclosure of which is provided herein. In other embodiments, the nucleic acid-guided nuclease system can include at least one known guide sequence (gRNA) or at least one novel gRNA, such as a single-stranded gRNA or a double-stranded gRNA. In some embodiments, the engineered nucleic acid-guided nucleases of the present invention can be used in systems for editing genes of interest in humans or other species.

[0013] Targetable nuclease systems in bacteria and archaea have emerged as powerful tools for precise genome editing. However, naturally occurring nucleases have several limitations, such as challenges in expression and delivery due to the size of the nucleic acid sequence and protein. In certain embodiments, the novel engineered nucleic acid-guided nuclease constructs disclosed herein can be engineered to target genes in subjects and / or improve the efficiency and / or accuracy of targeted gene editing.

[0014] According to these embodiments, Cas12a is known to be a single RNA-guided CRISPR / Cas endonuclease capable of genome editing with distinct characteristics compared to Cas9. In certain embodiments, Cas12a-based systems enable rapid and reliable introduction of donor DNA into the genome. Furthermore, Cas12a extends genome editing. CRISPR / Cas12a genome editing has been evaluated not only in human cells but also in other organisms, including plants. Some features of the CRISPR / Cas12a system are different compared to CRISPR / Cas9.

[0015] It is known that Cas12a nuclease recognizes T-rich protospacer adjacent motif (PAM) sequences (e.g., 5'-TTTN-3' (AsCas12a, LbCas12a) and 5'-TTN-3' (FnCas12a)), while the comparable sequence for SpCas9 is NGG. The PAM sequence for Cas12a is located at the 5' end of the target DNA sequence, whereas for Cas9 it is at the 3' end. Furthermore, Cas12a can cleave DNA distal to its PAM, near positions +18 / +23 of the protospacer. This cleavage generates staggered DNA overhangs (e.g., sticky ends), whereas Cas9 cleaves near its PAM after the 3' position of the protospacer on both strands, creating blunt ends. In certain methods, altering the nuclease's recognition can provide improvements over Cas9 or Cas12a, improving precision. Furthermore, Cas12a is guided by a single crRNA and does not require a tracrRNA, resulting in a shorter gRNA sequence than the sgRNA used by Cas9. Surprisingly, it has been found that the modified Cas12a nuclease provided herein can function with double-stranded gRNA.

[0016] Cas12a is also known to exhibit additional ribonuclease activity that functions during crRNA processing. Cas12a can be used as an editing tool for various species (e.g., S. cerevisiae), allowing the use of alternative PAM sequences compared to those recognized by CRISPR / Cas9. The novel nucleases disclosed herein can also recognize the same or alternative PAM sequences. These novel nucleases can provide an alternative system for multiplex genome editing compared to known multiplex approaches, and can be used as an improved system for mammalian gene editing.

[0017] The well-known Cas12a protein-RNA complex recognizes the T-rich PAM and cleaves it, resulting in staggered DNA double-strand breaks. Cas12a-type nucleases interact with the pseudoknot structure formed by the 5' handle of the crRNA. The guide RNA segment, consisting of a seed region and a 3' end, has a binding sequence complementary to the target DNA sequence. Cas12a-type nucleases characterized to date have been demonstrated to function with a single gRNA and process gRNA arrays. While the Cas12a-type and Cas9 nuclease systems have proven highly impactful, neither system has been demonstrated to function as predictably as needed to enable all of the envisioned applications of gene editing technology.

[0018] Currently, various approaches are being attempted to engineer improved CRISPR editing systems with increased efficiency and precision, including manipulation of PAM specificity, stability, and the sequence of the gRNA and / or nuclease. For example, chemical modifications of CRISPR / Cas9 gRNAs predicted to increase gRNA stability have been found to result in a 3.8-fold higher indel frequency in human cells. Furthermore, other studies have included structure-guided mutagenesis of Cas12a and screened to identify variants with an increased range of recognized PAM sequences. These engineered AsCa12a recognize the TYCV and TATV PAMs in addition to the established TTTV sequence and have demonstrated enhanced activity in vitro and in tested human cells.

[0019] In certain embodiments, the Cas12a-like nucleases and engineered gRNAs disclosed herein are contemplated for use in bacteria and other prokaryotes. In certain embodiments, the engineered designer nucleases are contemplated for use in eukaryotes, such as yeast, mammals, including humans, and birds and fish, or cells derived therefrom.

[0020] In some embodiments, the off-targeting rate of the nuclease constructs disclosed herein can be reduced compared to a control, e.g., a native sequence, due to improved editing. The off-targeting rate can be easily tested.

[0021] In some embodiments, the nuclease constructs disclosed herein can share conserved coding motifs with known nucleases. In other embodiments, the nuclease constructs disclosed herein do not share conserved coding peptide motifs with known nucleases. In preferred embodiments, provided herein are compositions, methods, and / or kits wherein the CRISPR nuclease comprises a type V nuclease. In certain embodiments, provided herein are compositions, methods, and / or kits wherein the CRISPR nuclease comprises a type VA, type VB, type VC, type VD, or type VE CRISPR nuclease. In certain embodiments, provided herein are compositions, methods, and / or kits wherein the CRISPR nuclease comprises a type VA nuclease. Naturally occurring VA-type CRISPR nucleases contain a RuvC-like nuclease domain but no HNH endonuclease domain, and recognize a 5'T-rich PAM located immediately upstream of the target nucleotide sequence, with the orientation determined by the non-target strand (i.e., the strand not hybridized to the spacer sequence). These CRISPR nucleases cleave double-stranded DNA, generating staggered double-strand breaks rather than blunt ends. The cleavage site is distant from the PAM site (e.g., at least 10, 11, 12, 13, 14, or 15 nucleotides downstream of the PAM on the non-target strand and / or at least 15, 16, 17, 18, or 19 nucleotides upstream of the sequence complementary to the PAM on the target strand).

[0022] In certain embodiments, the VA-type CRISPR nuclease comprises Cpf1. Cpf1 proteins are known in the art and are described, for example, in U.S. Patent Nos. 9,790,490 and 10,113,179. Cpf1 orthologs are found in various bacterial and archaeal genomes. For example, in certain embodiments, the Cpf1 protein is derived from; Francisella novicida U112 (Fn), Acidaminococcus sp. BV3L6 (As), Lachnospiraceae bacterium ND2006 (Lb), Lachnospiraceae bacterium MA2020 (Lb2), Candidatus Methanoplasma termitum (CMt), Moraxella bovoculi 237(Mb), Porphyromonas crevioricanis(Pc), Prevotella disiens(Pd), Francisella tularensis 1, Francisella tularensis subsp.novicida, Prevotella albensis, Lachnospiraceae bacterium MC2017 1, Butyrivibrio proteoclasticus, Peregrinibacteria bacterium GW2011_GWA2_33_10, Parcubacteria bacterium GW2011_GWC2_44_17, Smithella sp, SCADC, Eubacterium eligens, Leptospira inadai, Porphyromonas macacae, Prevotella bryantii, Proteocatella sphenisci, Anaerovibrio sp.RM50, Moraxella caprae, Lachnospiraceae bacterium COE1, or Eubacterium coprostanoligenes.

[0023] In certain embodiments, the VA-type CRISPR nuclease comprises AsCpf1 or a variant thereof. In certain embodiments, the VA-type CRISPR nuclease comprises an amino acid sequence that is at least 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence set forth in SEQ ID NO: 3 of International (PCT) Application WO2021 / 158918. In certain embodiments, the VA-type CRISPR nuclease comprises the amino acid sequence set forth in SEQ ID NO: 3 of International (PCT) Application WO2021 / 158918.

[0024] In certain embodiments, the VA-type CRISPR nuclease comprises LbCpf1 or a variant thereof. In certain embodiments, the VA-type CRISPR nuclease comprises an amino acid sequence at least 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence set forth in SEQ ID NO: 4 of International (PCT) Application WO2021 / 158918. In certain embodiments, the VA-type Cas protein comprises the amino acid sequence set forth in SEQ ID NO: 4 of International (PCT) Application WO2021 / 158918.

[0025] In certain embodiments, the VA-type CRISPR nuclease comprises FnCpf1 or a variant thereof. In certain embodiments, the VA-type Cas protein comprises an amino acid sequence at least 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence set forth in SEQ ID NO: 5 of International (PCT) Application WO2021 / 158918. In certain embodiments, the VA-type Cas protein comprises the amino acid sequence set forth in SEQ ID NO: 5 of International (PCT) Application WO2021 / 158918.

[0026] In certain embodiments, the VA-type CRISPR nuclease comprises Prevotella bryantii Cpf1 (PbCpf1) or a variant thereof. In certain embodiments, the VA-type Cas protein comprises an amino acid sequence at least 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence set forth in SEQ ID NO:6 of International (PCT) Application WO2021 / 158918. In certain embodiments, the VA-type Cas protein comprises the amino acid sequence set forth in SEQ ID NO:6 of International (PCT) Application WO2021 / 158918.

[0027] In certain embodiments, the type VA CRISPR nuclease comprises Proteocatella sphenisci Cpf1 (PsCpf1) or a variant thereof. In certain embodiments, the type VA Cas protein comprises an amino acid sequence at least 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence set forth in SEQ ID NO: 7 of International (PCT) Application WO2021 / 158918. In certain embodiments, the type VA Cas protein comprises the amino acid sequence set forth in SEQ ID NO: 7 of International (PCT) Application WO2021 / 158918.

[0028] In certain embodiments, the VA-type CRISPR nuclease comprises Anaerovibrio sp. RM50 Cpf1 (As2Cpf1) or a variant thereof. In certain embodiments, the VA-type Cas protein comprises an amino acid sequence at least 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence set forth in SEQ ID NO: 8 of International (PCT) Application WO2021 / 158918. In certain embodiments, the VA-type Cas protein comprises the amino acid sequence set forth in SEQ ID NO: 8 of International (PCT) Application WO2021 / 158918.

[0029] In certain embodiments, the type VA CRISPR nuclease comprises Moraxella caprae Cpf1 (McCpf1) or a variant thereof. In certain embodiments, the type VA Cas protein comprises an amino acid sequence at least 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence set forth in SEQ ID NO:9 of International (PCT) Application WO2021 / 158918. In certain embodiments, the type VA Cas protein comprises the amino acid sequence set forth in SEQ ID NO:9 of International (PCT) Application WO2021 / 158918.

[0030] In certain embodiments, the type VA CRISPR nuclease comprises the bacterial species Lachnospiraceae COE1 Cpf1 (Lb3Cpf1) or a variant thereof. In certain embodiments, the type VA Cas protein comprises an amino acid sequence at least 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence set forth in SEQ ID NO: 10 of International (PCT) Application WO2021 / 158918. In certain embodiments, the type VA Cas protein comprises the amino acid sequence set forth in SEQ ID NO: 10 of International (PCT) Application WO2021 / 158918.

[0031] In certain embodiments, the type VA CRISPR nuclease comprises Eubacterium coprostanoligenes Cpf1 (EcCpf1) or a variant thereof. In certain embodiments, the type VA Cas protein comprises an amino acid sequence at least 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence set forth in SEQ ID NO: 11 of International (PCT) Application WO2021 / 158918. In certain embodiments, the type VA Cas protein comprises the amino acid sequence set forth in SEQ ID NO: 11 of International (PCT) Application WO2021 / 158918.

[0032] In certain embodiments, the VA-type CRISPR nuclease is not Cpf1. In certain embodiments, the VA-type CRISPR nuclease is not AsCpf1.

[0033] In certain embodiments, the type VA CRISPR nuclease comprises a type VA nuclease described in U.S. Patent No. 9,982,279.

[0034] In certain embodiments, a type VA CRISPR nuclease polypeptide used in the compositions and methods herein can be represented by a polypeptide comprising a sequence having at least 60, 70, 80, 85, 90, 95, 96, 97, 98, 99, or 100% sequence identity, preferably at least 80%, more preferably at least 90%, even more preferably at least 95%, and even more preferably at least 98% sequence identity to SEQ ID NO: 1. wherein the type VA CRISPR nuclease polypeptide further comprises at least one, two, three, four, five, or six nuclear localization sequences (NLSs), each of which can be at the amino- or carboxy-terminus of the CRISPR nuclease polypeptide, and / or one or more purification tags, and in addition, a cleavage sequence can be provided to remove a portion of the protopeptide. As used herein, the term "at or near" the N-terminus or C-terminus includes cases where the amino acid of the NLS closest to the N-terminus or C-terminus is within 300 amino acids, and in some cases within 200 amino acids, of the N-terminus or C-terminus of the polypeptide (e.g., a core polypeptide such as one of the CRISPR nucleases described herein to which the NLS(s) are attached). In certain embodiments, a Type V CRISPR nuclease polypeptide, e.g., a Type Va CRISPR polypeptide, comprises two, three, four, or five NLSs, each of which is at or near the N-terminus or C-terminus of the polypeptide, and in preferred embodiments, the NLS is at or near the N-terminus. In certain embodiments, the CRISPR nuclease polypeptide comprising one or more NLSs, and optionally a purification tag and / or cleavage site, comprises a sequence that is at least 60%, 70%, 80%, 85%, 90%, 95%, 98%, 99% or 100% identical to any one of SEQ ID NOs: 109-112, preferably at least 80%, more preferably at least 90%, even more preferably at least 95%, and even more preferably at least 98% identical.In certain embodiments, a Type V, e.g., VA CRISPR nuclease polypeptide comprises at least 1-30, 1-20, 1-15, 1-10, 1-9, 1-8, 1-7, 1-6, 1-5, 2-30, 2-20, 2-15, 2-10, 2-9, 2-8, 2-7, 2-6, 2-5, 3-30, 3-20, 3-15, 3-10, 3-9, 3-8, 3-7, 3-6, or 3-5, preferably 1-10, more preferably 2-10, and even more preferably 3-10, NLSs, each of which is at or near the N-terminus or C-terminus of the polypeptide, and in preferred embodiments, at or near the N-terminus. In certain embodiments, at least two or at least three of the NLSs have different mechanisms, i.e., different mechanisms, for localizing the attached polypeptide to the nucleus. Such mechanisms are well known in the art, see, e.g., Lu et al. Cell Commun Signal (2021) 19:60 https: / / doi.org / 10.1186 / s12964-021-00741-y. Suitable NLS, purification tag, and cleavage site sequences can be as described elsewhere herein, for example, in the sections designated Nuclear Localization Signals, Purification Tags, and Cleavage Sites.

[0035] SEQ ID NO: 1

[0036] A nucleotide sequence encoding SEQ ID NO:1 can include a sequence having less than 99%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, or 40% sequence identity to SEQ ID NO:22, and in preferred embodiments, has less than 75% sequence identity. In certain embodiments, a nucleotide sequence encoding SEQ ID NO:1 can also include a nucleic acid sequence encoding one or more NLSs at the N-terminus and / or C-terminus described herein, and / or a tag, such as a purification tag at the N-terminus described herein. In certain embodiments, provided herein are compositions comprising a first polynucleotide encoding a polypeptide comprising a nucleic acid-guided nuclease, including a CRISPR type V nuclease polypeptide, wherein the polynucleotide has less than 75% sequence identity to SEQ ID NO:22, e.g., the nuclease polypeptide comprises at least one, two, three, four, or five NLSs, each of which is at or near the N-terminus or C-terminus of the nuclease polypeptide. The NLS can be any of those described herein. The first polynucleotide can include a sequence encoding a purification tag, e.g., a purification tag described herein, and / or a cleavage site, e.g., a cleavage site described herein. In certain embodiments, the first polynucleotide encodes a polypeptide comprising a sequence at least 60%, 70%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% identical, preferably at least 80%, more preferably at least 90%, even more preferably at least 95%, and even more preferably at least 98% identical, to any one of SEQ ID NOs: 109-112, e.g., SEQ ID NO: 109, or SEQ ID NO: 110, or SEQ ID NO: 111, or SEQ ID NO: 112. The first polynucleotide comprises a sequence at least 50%, 60%, 70%, 80%, 90%, 95%, 97% or 99% identical or 100% identical to SEQ ID NO:113, preferably at least 80%, more preferably at least 90%, even more preferably at least 95%, even more preferably at least 98% identical.In certain embodiments, the composition further comprises a second polynucleotide encoding a gNA or a portion thereof, wherein the gNA, e.g., gRNA, comprises a target nucleotide sequence within the polynucleotide or a spacer sequence that targets a target nucleotide sequence within the polynucleotide encoding the gNA, e.g., gRNA, wherein the gNA, e.g., gRNA, is compatible with a type V CRISPR nuclease. In certain embodiments, the first polynucleotide and the second polynucleotide are the same. The composition may further comprise a third polynucleotide comprising a donor template. In certain embodiments, a vector is provided comprising one of the polynucleotide compositions described in this paragraph. In certain embodiments, a cell, such as a human cell, e.g., an immune cell, e.g., a T cell, or a stem cell, e.g., an iPSC, is provided comprising one of the polynucleotide compositions described in this paragraph. In certain embodiments, a method is provided comprising inserting any one of the polynucleotide compositions described in this paragraph into a cell. In certain embodiments, inserting the composition comprises electroporation.

[0037] SEQ ID NO:22:

[0038] Exemplary nucleotide sequences encoding SEQ ID NO:1 can include, for example, SEQ ID NOs:23-42.

[0039] SEQ ID NO: 23

[0040] SEQ ID NO: 24

[0041] SEQ ID NO: 25

[0042] SEQ ID NO: 26

[0043] SEQ ID NO: 27

[0044] SEQ ID NO: 28

[0045] SEQ ID NO: 29

[0046] SEQ ID NO: 30

[0047] SEQ ID NO: 31

[0048] SEQ ID NO: 32

[0049] SEQ ID NO: 33

[0050] SEQ ID NO: 34

[0051] SEQ ID NO: 35

[0052] SEQ ID NO: 36

[0053] SEQ ID NO: 37

[0054] SEQ ID NO: 38

[0055] SEQ ID NO: 39

[0056] SEQ ID NO: 40

[0057] SEQ ID NO: 41

[0058] SEQ ID NO: 42

[0059] Nucleic acid-guided nucleases can encompass engineered nucleotide sequences of natural sequences, engineered sequences, or synthetic variants. Non-limiting examples of the types of engineering that can be performed to obtain a non-natural nuclease system are as follows: Engineering can include codon optimization to facilitate or improve expression in host cells, such as heterologous host cells. Engineering can reduce the size or molecular weight of the nuclease to facilitate expression or delivery. Engineering can alter the PAM selection to alter PAM specificity or broaden the range of PAMs recognized. Engineering can alter, increase, or decrease the stability, processivity, specificity, or efficiency of the targetable nuclease system. Engineering can alter, increase, or decrease protein stability. Engineering can alter, increase, or decrease nucleic acid scanning processivity. Engineering can alter, increase, or decrease target sequence specificity. Engineering can alter, increase, or decrease nuclease activity. Engineering can alter, increase, or decrease editing efficiency. Engineering can alter, increase, or decrease transformation efficiency. Manipulation can alter, increase, or decrease nuclease activity or induce nucleic acid expression. As used herein, a non-naturally occurring nucleic acid sequence can be a synthetic variant engineered sequence or an engineered nucleotide sequence. Such non-naturally occurring nucleic acid sequences can be amplified, cloned, assembled, synthesized, generated with synthetic oligonucleotides or dNTPs, or obtained using methods known to those of skill in the art. In certain embodiments, examples of non-naturally occurring nucleic acid-derived nucleases disclosed herein can include nucleic acid-derived nucleases having engineered polypeptide sequences (e.g., SEQ ID NOS: 2-4) and synthetic variant nucleotide sequences (e.g., SEQ ID NOS: 43-63).

[0060] SEQ ID NO: 2

[0061] SEQ ID NO: 3

[0062] SEQ ID NO:4

[0063] SEQ ID NO:109:

[0064] SEQ ID NO:110:

[0065] SEQ ID NO: 111

[0066] SEQ ID NO: 112

[0067] SEQ ID NO: 43

[0068] SEQ ID NO: 44

[0069] SEQ ID NO: 45

[0070] SEQ ID NO: 46

[0071] SEQ ID NO: 47

[0072] SEQ ID NO: 48

[0073] SEQ ID NO: 49

[0074] SEQ ID NO:50

[0075] SEQ ID NO:51

[0076] SEQ ID NO:52

[0077] SEQ ID NO:53

[0078] SEQ ID NO:54

[0079] SEQ ID NO: 55

[0080] SEQ ID NO:56

[0081] SEQ ID NO:57

[0082] SEQ ID NO:58

[0083] SEQ ID NO:59

[0084] SEQ ID NO: 60

[0085] SEQ ID NO: 61

[0086] SEQ ID NO: 62

[0087] SEQ ID NO: 63

[0088] SEQ ID NO: 64

[0089] SEQ ID NO: 65

[0090] SEQ ID NO: 66

[0091] SEQ ID NO: 67

[0092] SEQ ID NO: 68

[0093] SEQ ID NO: 69

[0094] SEQ ID NO: 70

[0095] SEQ ID NO:71

[0096] SEQ ID NO:72

[0097] SEQ ID NO: 73

[0098] SEQ ID NO:74

[0099] SEQ ID NO: 75

[0100] SEQ ID NO:76

[0101] SEQ ID NO:77

[0102] SEQ ID NO:78

[0103] SEQ ID NO:79

[0104] SEQ ID NO: 80

[0105] SEQ ID NO: 81

[0106] SEQ ID NO:82

[0107] SEQ ID NO: 83

[0108] SEQ ID NO:84

[0109] SEQ ID NO: 85

[0110] SEQ ID NO:86

[0111] SEQ ID NO:87

[0112] SEQ ID NO: 88

[0113] SEQ ID NO:89

[0114] SEQ ID NO: 90

[0115] SEQ ID NO: 91

[0116] SEQ ID NO:92

[0117] SEQ ID NO: 93

[0118] SEQ ID NO:94

[0119] SEQ ID NO: 95

[0120] SEQ ID NO:96

[0121] SEQ ID NO:97

[0122] SEQ ID NO: 98

[0123] SEQ ID NO: 99

[0124] SEQ ID NO: 100

[0125] SEQ ID NO: 101

[0126] SEQ ID NO: 102

[0127] SEQ ID NO: 103

[0128] SEQ ID NO: 104

[0129] SEQ ID NO: 105

[0130] SEQ ID NO: 113

[0131] In certain embodiments, the nucleic acid-guided nuclease disclosed herein, e.g., Type V, preferably Type VA CRISPR nuclease polypeptide, comprises a polypeptide having an amino acid sequence at least 50% identical to SEQ ID NO: 2. In certain embodiments, the nucleic acid-guided nuclease, e.g., Type V, preferably Type VA CRISPR nuclease polypeptide disclosed herein, comprises a polypeptide having an amino acid sequence that is at least about 60%, 65%, 75%, 85%, 95%, 99% or about 100% identical, preferably at least 80%, more preferably at least 90%, even more preferably at least 95%, and even more preferably at least 98% identical to the amino acid sequence of SEQ ID NO: 2. In certain embodiments, the nucleic acid-guided nuclease, e.g., Type V, preferably Type VA CRISPR nuclease polypeptide disclosed herein, comprises a polypeptide having an amino acid sequence that is at least about 50% identical to SEQ ID NO: 3. In certain embodiments, the nucleic acid-guided nuclease, e.g., a Type V, preferably a Type VA CRISPR nuclease polypeptide disclosed herein, comprises a polypeptide having an amino acid sequence that is at least about 60%, 65%, 75%, 85%, 95%, 99%, or about 100% identical, preferably at least 80%, more preferably at least 90%, even more preferably at least 95%, and even more preferably at least 98% identical to the amino acid sequence of SEQ ID NO: 3. In certain embodiments, the nucleic acid-guided nuclease, e.g., a Type V, preferably a Type VA CRISPR nuclease polypeptide disclosed herein, comprises a polypeptide having an amino acid sequence that is at least about 50% identical to SEQ ID NO: 4. In certain embodiments, the nucleic acid-guided nucleases disclosed herein comprise polypeptides having an amino acid sequence that is at least about 60%, 65%, 75%, 85%, 95%, 99% or about 100% identical to the amino acid sequence of SEQ ID NO:4, preferably at least 80%, more preferably at least 90%, even more preferably at least 95%, and even more preferably at least 98% identical.In certain embodiments, a nucleic acid-guided nuclease, e.g., a Type V, preferably a Type VA CRISPR nuclease polypeptide disclosed herein, comprises a polypeptide having an amino acid sequence at least about 50% identical to any one of SEQ ID NOs: 109-112. In certain embodiments, a nucleic acid-guided nuclease disclosed herein comprises a polypeptide having an amino acid sequence at least 60%, 70%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% identical, preferably at least 80%, more preferably at least 90%, even more preferably at least 95%, even more preferably at least 98% identical to the amino acid sequence of any one of SEQ ID NOs: 109-112. In certain embodiments, a nucleic acid-guided nuclease, e.g., a Type V, preferably a Type VA CRISPR nuclease polypeptide disclosed herein, comprises a polypeptide having an amino acid sequence at least about 50% identical to SEQ ID NO: 109. In certain embodiments, a nucleic acid-guided nuclease disclosed herein comprises a polypeptide having an amino acid sequence at least 60%, 70%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% identical, preferably at least 80%, more preferably at least 90%, even more preferably at least 95%, and even more preferably at least 98% identical to SEQ ID NO: 109. In certain embodiments, a nucleic acid-guided nuclease, such as a Type V, preferably a Type VA CRISPR nuclease polypeptide disclosed herein, comprises a polypeptide having an amino acid sequence at least about 50% identical to SEQ ID NO: 110. In certain embodiments, the nucleic acid-guided nucleases disclosed herein comprise polypeptides having an amino acid sequence that is at least 60%, 70%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% identical to SEQ ID NO:110, preferably at least 80%, more preferably at least 90%, even more preferably at least 95%, and even more preferably at least 98% identical.In certain embodiments, a nucleic acid-guided nuclease, such as a Type V, preferably a Type VA CRISPR nuclease polypeptide disclosed herein, comprises a polypeptide having an amino acid sequence that is at least about 50% identical to SEQ ID NO: 111. In certain embodiments, a nucleic acid-guided nuclease disclosed herein comprises a polypeptide having an amino acid sequence that is at least 60%, 70%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% identical, preferably at least 80%, more preferably at least 90%, even more preferably at least 95%, even more preferably at least 98% identical to SEQ ID NO: 111. In certain embodiments, a nucleic acid-guided nuclease, such as a Type V, preferably a Type VA CRISPR nuclease polypeptide disclosed herein, comprises a polypeptide having an amino acid sequence that is at least about 50% identical to SEQ ID NO: 112. In certain embodiments, a nucleic acid-guided nuclease disclosed herein comprises a polypeptide having an amino acid sequence that is at least 60%, 70%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% identical to SEQ ID NO:112, preferably at least 80%, more preferably at least 90%, even more preferably at least 95%, and even more preferably at least 98% identical.

[0132] Nuclear localization signal (NLS) In certain embodiments, compositions, e.g., nucleases, disclosed herein include one or more nuclear localization sequences (NLSs), e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more NLSs. In some embodiments, compositions, e.g., engineered nucleases, include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more NLSs at or near the amino terminus, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more NLSs at or near the carboxy terminus, or combinations thereof (e.g., one or more NLSs at the amino terminus and one or more NLSs at the carboxy terminus). When two or more NLSs are present, each can be selected independently of the others, such that a single NLS can be present in two or more copies and / or in combination with one or more other NLSs present in one or more copies. In certain embodiments, an engineered nuclease includes four NLSs.

[0133] Non-limiting examples of NLSs include NLS sequences derived from: the SV40 virus large T antigen NLS having the amino acid sequence PKKKRKV (SEQ ID NO: 5); an NLS from nucleoplasmin (e.g., the nucleoplasmin bipartite NLS having the sequence KRPAATKKAGQAKKKK (SEQ ID NO: 6)); the c-myc NLS having the amino acid sequence PAAKRVKLD (SEQ ID NO: 7) or RQRRNELKRSP (SEQ ID NO: 8); the hRNPA1 M9 NLS having the sequence NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY (SEQ ID NO: 9); the IBB domain from importin alpha with the sequence RMRIZFKNKGKDTAELRRRRVEVSVELRKAKKDEQILKRRNV (SEQ ID NO: 10); the fibroid T protein with the sequences VSRKRPRP (SEQ ID NO: 11) and PPKKARED (SEQ ID NO: 12); the human p53 with the sequence PQPKKKPL (SEQ ID NO: 13); and the mouse c-abl The sequence of IV, SALIKKKKKMAP (SEQ ID NO: 14); the sequences of influenza virus NS1, DRLRR (SEQ ID NO: 15) and PKQKKRK (SEQ ID NO: 16); the sequence of hepatitis virus delta antigen, RKLKKKIKKL (SEQ ID NO: 17); the sequence of mouse Mx1 protein, REKKKFLKRR (SEQ ID NO: 18); the sequence of human poly(ADP-ribose) polymerase, KRKGDEVDGVDEVAKKKSKK (SEQ ID NO: 19); the sequence of steroid hormone receptor (human) glucocorticoid, RKCLQAGMNLEARKTKK (SEQ ID NO: 20); and EGL-13, MSRRRKANPTKLSENAKKLAKEVEN, SEQ ID NO: 107.

[0134] In certain embodiments, nucleases provided herein comprise at least one myc-associated NLS comprising the sequence PAAKKKKLD (SEQ ID NO: 21); in certain embodiments, the myc-associated NLS is at the N-terminus of the nuclease. In certain embodiments, nucleases provided herein comprise at least one nucleoplasmin NLS comprising the sequence KRPAATKKAGQAKKKK (SEQ ID NO: 6); in certain embodiments, the nucleoplasmin NLS is at the C-terminus of the nuclease. In certain embodiments, nucleases provided herein comprise at least one or at least two SV40 NLS sequences comprising the sequence PKKKRKV (SEQ ID NO: 5); in certain embodiments, the SV40 NLS is at the C-terminus of the nuclease. In certain embodiments, nucleases provided herein comprise one NLS at the N-terminus and three NLSs at the C-terminus, e.g., one myc-associated NLS at the N-terminus, one nucleoplasmin NLS, and two SV40 NLSs at the C-terminus. In certain embodiments, the nucleases provided herein comprise one myc-related NLS having the sequence PAAKKKKLD (SEQ ID NO: 21) at the N-terminus, and one nucleoplasmin NLS comprising the sequence KRPAATKKAGQAKKKK (SEQ ID NO: 6) and two SV40 NLSs comprising the sequence PKKKRKV (SEQ ID NO: 5) at the C-terminus.

[0135] Generally, one or more NLSs have sufficient strength to drive the accumulation of detectable amounts of nucleic acid-guided nucleases in the nuclei of eukaryotic cells. Generally, the strength of nuclear localization activity can depend on the number of NLSs, the specific NLS(s) used, or a combination of these factors. Detection of nuclear accumulation can be performed by any suitable technique. For example, a detectable marker can be fused to the nucleic acid-guided nuclease in combination with a means for detecting its nuclear location (e.g., a nuclear-specific stain such as DAPI) so that its location within the cell can be visualized. Cell nuclei can be isolated from the cells, and their contents can then be analyzed by any suitable process for detecting proteins, such as immunohistochemistry, Western blot, or enzyme activity assays. Nuclear accumulation can also be determined indirectly, such as by assaying the effect of nucleic acid-guided nuclease complex formation (e.g., assaying for DNA cleavage or mutation in the target sequence, or assaying for changes in gene expression activity affected by targetable nuclease complex formation and / or nucleic acid-guided nuclease activity) compared to a control that was not exposed to the nucleic acid-guided nuclease or targetable nuclease complex, or a control that was exposed to a nucleic acid-guided nuclease that does not have one or more NLSs.

[0136] In certain embodiments, a nucleic acid-guided nuclease disclosed herein comprises a polypeptide having an amino acid sequence at least about 60%, 65%, 75%, 85%, 95%, 99%, or about 100% identical to the amino acid sequence of SEQ ID NO:4 and at least one myc-associated NLS comprising the sequence PAAKKKKLD (SEQ ID NO:21); in certain embodiments, the myc-associated NLS is at the N-terminus of the nuclease. In certain embodiments, a nucleic acid-guided nuclease disclosed herein comprises a polypeptide having an amino acid sequence at least about 60%, 65%, 75%, 85%, 95%, 99%, or about 100% identical to the amino acid sequence of SEQ ID NO:4 and at least one nucleoplasmin NLS comprising the sequence KRPAATKKAGQAKKKK (SEQ ID NO:6); in certain embodiments, the nucleoplasmin NLS is at the C-terminus of the nuclease. In certain embodiments, the nucleic acid-guided nuclease disclosed herein comprises a polypeptide having an amino acid sequence at least about 60%, 65%, 75%, 85%, 95%, 99%, or about 100% identical to the amino acid sequence of SEQ ID NO: 4 and at least one or at least two SV40 NLS sequences comprising the sequence PKKKRKV; in certain embodiments, the SV40 NLS is at the C-terminus of the nuclease. In certain embodiments, the nucleic acid-guided nuclease disclosed herein comprises a polypeptide having an amino acid sequence at least about 60%, 65%, 75%, 85%, 95%, 99%, or about 100% identical to the amino acid sequence of SEQ ID NO: 4, an NLS at the N-terminus, and three NLSs at the C-terminus, e.g., a myc-related NLS at the N-terminus, and a nucleoplasmin NLS and two SV40 NLSs at the C-terminus. In certain embodiments, a nucleic acid-guided nuclease disclosed herein comprises a polypeptide having an amino acid sequence at least about 60%, 65%, 75%, 85%, 95%, 99% or about 100% identical to the amino acid sequence of SEQ ID NO: 4, one myc-related NLS having the sequence PAAKKKKLD (SEQ ID NO: 21) at its N-terminus, one nucleoplasmin NLS comprising the sequence KRPAATKKAGQAKKKK (SEQ ID NO: 6), and two SV40 NLSs comprising the sequence PKKKRKV (SEQ ID NO: 5) at its C-terminus.In certain embodiments, a nucleic acid-guided nuclease disclosed herein comprises a polypeptide having an amino acid sequence at least about 60%, 65%, 75%, 85%, 95%, 99%, or about 100% identical to the amino acid sequence of SEQ ID NO: 1, one, two, or three NLSs at the N-terminus, and one, two, or three NLSs at the C-terminus. In certain embodiments, a nucleic acid-guided nuclease disclosed herein comprises a polypeptide having an amino acid sequence at least about 60%, 65%, 75%, 85%, 95%, 99%, or about 100% identical to the amino acid sequence of SEQ ID NO: 1, one myc-related NLS having the sequence PAAKKKKLD (SEQ ID NO: 21) at the N-terminus, one nucleoplasmin NLS comprising the sequence KRPAATKKAGQAKKKK (SEQ ID NO: 6), and two SV40 NLSs comprising the sequence PKKKRKV (SEQ ID NO: 5) at the C-terminus.

[0137] Purification Tag In certain embodiments, the nucleic acid-guided nucleases provided herein can include a tag, e.g., a purification tag, e.g., at the N-terminus. Examples of tags include poly-His tags, e.g., Gly-6xHis tags or Gly-8xHis tags, short epitope tags, e.g., FLAG, hemagglutinin (HA), c-myc, T7, and Glu-Glu; maltose-binding protein (mbp); N-terminal glutathione S-transferase (GST); and calmodulin-binding peptide (CBP). In certain embodiments, the nucleic acid-guided nucleases provided herein can include a poly-His tag, e.g., a Gly-6xHis tag, at the N-terminus. These Gly-6xHis tags are applicable for several reasons, including: 1) the 6xHis tag can be used in protein purification to achieve binding to a chromatography column for purification, and 2) the N-terminal glycine allows for site-specific chemical modification, which further enables advanced protein engineering. Furthermore, the Gly-6xHis is designed to be easily removed by digestion with tobacco etch virus (TEV) protease, if necessary. In these constructs, the Gly-6xHis tag is placed at the N-terminus. The Gly-6xHis tag is described in more detail in Martos-Maldonado et al., Nat Commun. (2018) 17;9(1):3307, the disclosure of which is incorporated herein by reference. Thus, in certain embodiments, provided herein is a nucleic acid-guided nuclease disclosed herein, comprising a polypeptide having an amino acid sequence at least about 60%, 65%, 75%, 85%, 95%, 99%, or about 100% identical to the amino acid sequence of SEQ ID NO: 4, and a poly-His tag, e.g., a Gly-6xHis tag, at the N-terminus. In certain embodiments, provided herein is a nucleic acid-guided nuclease as disclosed herein, comprising a polypeptide having an amino acid sequence at least about 60%, 65%, 75%, 85%, 95%, 99% or about 100% identical to the amino acid sequence of SEQ ID NO: 4, a poly-His tag, e.g., a Gly-6xHis tag, at the N-terminus, and / or a TEV cleavage site at the N-terminus.In certain embodiments, provided are nucleic acid-guided nucleases having an N-terminal poly-His tag, e.g., a Gly-6xHis tag, and an N-terminal TEV cleavage site, such as a polypeptide having at least about 60%, 65%, 75%, 85%, 95%, 99%, or about 100% identity to the amino acid sequence of SEQ ID NO: 2. In certain embodiments, provided herein are nucleic acid-guided nucleases disclosed herein that include a polypeptide having an amino acid sequence at least about 60%, 65%, 75%, 85%, 95%, 99%, or about 100% identity to the amino acid sequence of SEQ ID NO: 1, an N-terminal poly-His tag, e.g., a Gly-6xHis tag, and / or an N-terminal TEV cleavage site. Additionally or alternatively, the nuclease may include one or more NLSs described herein.

[0138] Cutting site In addition to, or instead of, including one or more NLSs, purification tags, and / or other additional amino acid sequences described herein, the engineered nuclease polypeptides disclosed herein can include one or more cleavage sequences, which may be at or near the N-terminus or C-terminus. Any suitable cleavage site can be used, and if multiple cleavage sites are used, they may be the same or different. In certain embodiments, the cleavage site comprises a tobacco etch virus protease cleavage sequence, referred to herein as the "TEV sequence" (SEQ ID NO: 108). The TEV sequence may be at or near the amino terminus. Generally, the cleavage sequence, e.g., the TEV sequence, is positioned such that cleavage at the cleavage sequence leaves intact any other additional amino acid sequences, particularly any NLSs, added to the original nuclease polypeptide. The TEV cleavage site may have the amino acid sequence ENLYFQS (SEQ ID NO: 108).

[0139] In certain embodiments, provided herein is a nucleic acid sequence encoding a polypeptide having at least 50% nucleic acid identity to the polypeptide set forth in SEQ ID NO: 2. In certain embodiments, provided herein is a nucleic acid sequence encoding a polypeptide having at least about 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, more than 95%, or 100% to the polypeptide set forth in SEQ ID NO: 2. In certain embodiments, provided herein is a nucleic acid sequence encoding a polypeptide having at least 50% nucleic acid identity to the polypeptide set forth in SEQ ID NO: 3. In certain embodiments, provided herein is a nucleic acid sequence encoding a polypeptide having at least about 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, more than 95%, or 100% to the polypeptide set forth in SEQ ID NO: 3. In certain embodiments, provided herein is a nucleic acid sequence encoding a polypeptide having at least 50% nucleic acid identity to the polypeptide set forth in SEQ ID NO: 4. In certain embodiments, provided herein are nucleic acid sequences encoding polypeptides having at least about 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, greater than 95%, or 100% polynucleotide identity to the polypeptide set forth in SEQ ID NO: 4. In certain embodiments, provided herein are nucleic acids having at least about 50%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, or 100% polynucleotide identity to any one of SEQ ID NOs: 23-105. In certain embodiments, provided herein are nucleic acids having at least about 50%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, or 100% polynucleotide identity to any one of SEQ ID NOs: 23-42. In certain embodiments, provided herein are nucleic acids having at least about 50%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, or 100% polynucleotide identity to any one of SEQ ID NOs: 43-65.In certain embodiments, provided herein are nucleic acids having at least about 50%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, or 100% polynucleotide identity to any one of SEQ ID NOs: 43-53. In certain embodiments, provided herein are nucleic acids having at least about 50%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, or 100% polynucleotide identity to any one of SEQ ID NOs: 54-58. In certain embodiments, provided herein are nucleic acids having at least about 50%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, or 100% polynucleotide identity to any one of SEQ ID NOs: 59-63. In certain embodiments, provided herein are nucleic acids having at least about 50%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, or 100% polynucleotide identity to any one of SEQ ID NOs: 43. In certain embodiments, provided herein are nucleic acids having at least about 50%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, or 100% polynucleotide identity to any one of SEQ ID NOs: 64-84. In certain embodiments, provided herein are nucleic acids having at least about 50%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, or 100% polynucleotide identity to any one of SEQ ID NOs: 64. In certain embodiments, provided herein are nucleic acids having at least about 50%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, or 100% polynucleotide identity to any one of SEQ ID NOs: 64-74. In certain embodiments, provided herein are nucleic acids having at least about 50%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, or 100% polynucleotide identity to any one of SEQ ID NOs: 75-79.In certain embodiments, provided herein are nucleic acids having at least about 50%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, or 100% polynucleotide identity to any one of SEQ ID NOs: 80-84. In certain embodiments, provided herein are nucleic acids having at least about 50%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, or 100% polynucleotide identity to any one of SEQ ID NOs: 85-105. In certain embodiments, provided herein are nucleic acids having at least about 50%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, or 100% polynucleotide identity to any one of SEQ ID NOs: 85. In certain embodiments, provided herein are nucleic acids having at least about 50%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, or 100% polynucleotide identity to any one of SEQ ID NOs: 85-95. In certain embodiments, provided herein are nucleic acids having at least about 50%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, or 100% polynucleotide identity to any one of SEQ ID NOs: 96-100. In certain embodiments, provided herein are nucleic acids having at least about 50%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, or 100% polynucleotide identity to any one of SEQ ID NOs: 101-105.

[0140] The nucleic acid sequence encoding the nucleic acid-guided nuclease can be operably linked to a promoter. Such nucleic acid sequences can be linear or circular. The nucleic acid sequence can be contained within a larger linear or circular nucleic acid sequence that includes additional elements such as an origin of replication, a selectable or screenable marker, a terminator, other components of the targetable nuclease system, such as a guide nucleic acid, and / or an editing or recorder cassette as disclosed herein. In some embodiments, the nucleic acid sequence can include sequences encoding at least one glycine, at least one poly-histidine tag, such as a 6X histidine tag, and / or at least one, two, three, four, or five nuclear localization signal tags, some or all of which can be located on the amino side of the polypeptide, the carboxy side of the polypeptide, or a combination thereof. As described in more detail below, the larger nucleic acid sequence can be a recombinant expression vector.

[0141] Guide nucleic acid In certain embodiments, the compositions and methods disclosed herein comprise a guide nucleic acid (gNA), e.g., a gRNA.

[0142] Generally, a guide polynucleotide, also referred to as a guide nucleic acid (gNA), can form a complex with a compatible nucleic acid-guided nuclease, such as those disclosed herein, and can hybridize with a target nucleic acid sequence, thereby directing the nuclease to the target nucleic acid sequence. A subject nucleic acid-guided nuclease that can form a complex with a guide polynucleotide can be referred to as a nucleic acid-guided nuclease that is compatible with the guide polynucleotide. Furthermore, a guide polynucleotide that can form a complex with a nucleic acid-guided nuclease can be referred to as a guide polynucleotide or guide nucleic acid that is compatible with the nucleic acid-guided nuclease. In some embodiments, the polynucleotides (gRNAs) disclosed herein can be split into fragments, e.g., two separate polynucleotides, comprising a synthetic tracrRNA and crRNA. Such gNAs, e.g., gRNAs, can be referred to as double-stranded gNAs or split gNAs, e.g., gRNAs.

[0143] The guide polynucleotide can be DNA. The guide polynucleotide can be RNA. The guide polynucleotide can comprise both DNA and RNA. The guide polynucleotide can comprise modified nucleotides or non-naturally occurring nucleotides. When the guide polynucleotide comprises RNA, the RNA guide polynucleotide can be encoded by a DNA sequence on a polynucleotide molecule such as a plasmid, linear construct, or editing cassette disclosed herein.

[0144] The guide polynucleotide can include a guide sequence, also referred to herein as a spacer sequence. A guide (spacer) sequence is a polynucleotide sequence that has sufficient complementarity with a target polynucleotide sequence (also referred to herein as a target nucleic acid sequence) to hybridize with the target sequence and induce sequence-specific binding of a complexed nucleic acid-guided nuclease to the target sequence. The degree of complementarity between a guide sequence and its corresponding target sequence, when optimally aligned using a suitable alignment algorithm, is about 50%, 60%, 75%, 80%, 85%, 90%, 95%, 97.5%, 99%, or more, or about more than 50%, 60%, 75%, 80%, 85%, 90%, 95%, 97.5%, 99%, or more. Any suitable algorithm for aligning sequences can be used to determine optimal alignment. In some embodiments, a guide sequence can be about 5, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, 75, or more, or more than about 5, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, 75, or more nucleotides in length. In other embodiments, a guide sequence can be less than about 75, 50, 45, 40, 35, 30, 25, 20 nucleotides in length. Preferably, the guide sequence is 10 to 30 nucleotides in length. The guide sequence may be 15 to 20 nucleotides in length. The guide sequence may be 15 nucleotides in length. The guide sequence may be 16 nucleotides in length. The guide sequence may be 17 nucleotides in length. The guide sequence may be 18 nucleotides in length. The guide sequence may be 19 nucleotides in length. The guide sequence may be 20 nucleotides in length.

[0145] The guide polynucleotide can include a scaffold sequence. Generally, a "scaffold sequence" can include any sequence having sufficient sequence to promote the formation of a targetable nuclease complex, including, but not limited to, a nucleic acid-guided nuclease, and the guide polynucleotide can include a scaffold sequence and a guide sequence. A sufficient sequence within the scaffold sequence to promote the formation of a targetable nuclease complex can include a degree of complementarity along the length of two sequence regions within the scaffold sequence, such as one or two sequence regions involved in the formation of a secondary structure. In some cases, the one or two sequence regions are included on the same polynucleotide or encoded on the same polynucleotide. In some cases, the one or two sequence regions are included on different polynucleotides or encoded on different polynucleotides. Optimal alignment can be determined by any suitable alignment algorithm and can further account for secondary structures, such as self-complementarity within either one or two sequence regions. In some embodiments, the degree of complementarity between one or two sequence regions along the length of the shorter of the two, when optimally aligned, is about 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 97.5%, 99% or more, or greater than about 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 97.5%, 99% or more. In some embodiments, at least one of the two sequence regions can be about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 40, 50, or more nucleotides in length, or more than about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 40, 50, or more nucleotides in length.

[0146] The scaffold sequence of the subject guide polynucleotide may comprise a secondary structure. The secondary structure may comprise a pseudoknot region. In some cases, the binding kinetics of the guide polynucleotide to the nucleic acid-guided nuclease is determined in part by the secondary structure within the scaffold sequence. In some cases, the binding kinetics of the guide polynucleotide to the nucleic acid-guided nuclease is determined in part by the nucleic acid sequence comprising the scaffold sequence. In some aspects, the present invention provides nucleases that bind to guide polynucleotides that may comprise a conserved scaffold sequence. For example, the nucleic acid-guided nucleases for use in the present disclosure may bind to a conserved pseudoknot region.

[0147] In certain embodiments, the engineered polynucleotide (gRNA) can be split into fragments comprising synthetic tracrRNA and crRNA.

[0148] As used herein, "guide nucleic acid" or "guide polynucleotide" can refer to one or more polynucleotides and can include 1) a guide (spacer) sequence capable of hybridizing to a target sequence, and 2) a scaffold sequence capable of interacting with or complexing with a nucleic acid-guided nuclease described herein. A guide nucleic acid can be provided as one or more nucleic acids. In certain embodiments, the guide sequence and scaffold sequence are provided as a single polynucleotide. In other aspects, a guide nucleic acid can include at least one amplicon-targeting fragment.

[0149] A guide nucleic acid can be compatible with a nucleic acid-guided nuclease if the two elements can form a functional targetable nuclease complex that can cleave the target sequence.In certain methods, the compatible scaffold sequence of a compatible guide nucleic acid can be found by scanning the sequence adjacent to the natural nucleic acid-guided nuclease locus.For example, a natural nucleic acid-guided nuclease can be encoded on the genome adjacent to the corresponding compatible guide nucleic acid or scaffold sequence.

[0150] Nucleic acid-guided nucleases can be compatible with guide nucleic acids not found in the nuclease's endogenous host. Such orthogonal guide nucleic acids can be determined by empirical testing. Orthogonal guide nucleic acids can be derived from different bacterial species, synthesized, or otherwise engineered to be non-natural.

[0151] Orthogonal guide nucleic acids compatible with common nucleic acid-guided nucleases can include one or more common features. The common features can include sequences outside the pseudoknot region. The common features can include the pseudoknot region. The common features can include primary sequence or secondary structure.

[0152] A guide nucleic acid can be engineered to target a desired target sequence by modifying the guide (spacer) sequence so that the guide sequence is complementary to the target sequence, thereby allowing hybridization between the guide sequence and the target sequence. A guide nucleic acid with an engineered guide sequence can be referred to as an engineered guide nucleic acid. Engineered guide nucleic acids are often non-natural and not found in nature.

[0153] Engineered guide nucleic acids can be formed using the Synthetic Tracr RNA (STAR) system. When combined with the Cas12a protein, STAR can form at least one ribonucleoprotein (RNP) complex that targets specific genomic loci. STAR utilizes the natural properties of CRISPR (clustered regularly interspaced short palindromic repeats), which function like an immune system against invading viruses and plasmid DNA. Short DNA sequences (spacers) from invading viruses are integrated into CRISPR loci within the bacterial genome, serving as a "memory" of the previous infection. Upon reinfection, complementary mature CRISPR RNA (crRNA) is triggered to detect matching viral sequences. Together, the crRNA and trans-activating crRNA (tracrRNA) guide CRISPR-associated (Cas) nucleases to cleave double-strand breaks in the "foreign" DNA sequence. This prokaryotic CRISPR "immune system" has been engineered to function as a simple, easy, and rapidly implementable RNA-guided mammalian genome editing tool. STARs (including synthetic crRNAs and tracrRNAs) can form ribonucleoprotein (RNP) complexes that target specific genomic loci when combined with Cas12a proteins. Engineered guide nucleic acids formed using the RNA (STAR) system can result in split gRNAs. Split gRNAs, i.e., double-stranded guide RNAs, are described in further detail in WO2021067788A1.

[0154] In certain embodiments, provided herein are ribonucleoprotein (RNP) complexes comprising at least one nuclease disclosed herein. In certain embodiments, the RNP complex can comprise at least one nuclease having an amino acid sequence at least 50% identical to SEQ ID NO:2. In certain embodiments, the RNP complex can comprise at least one nuclease having an amino acid sequence at least about 60%, 65%, 75%, 85%, 95%, 99%, or about 100% identical to the amino acid sequence of SEQ ID NO:2. In certain embodiments, the RNP complex can comprise at least one nuclease having an amino acid sequence at least 50% identical to SEQ ID NO:3. In certain embodiments, the RNP complex can comprise at least one nuclease having an amino acid sequence at least about 60%, 65%, 75%, 85%, 95%, 99%, or about 100% identical to the amino acid sequence of SEQ ID NO:3. In certain embodiments, the RNP complex can comprise at least one nuclease having an amino acid sequence at least 50% identical to SEQ ID NO:4. In certain embodiments, an RNP complex can comprise at least one nuclease having an amino acid sequence at least about 60%, 65%, 75%, 85%, 95%, 99%, or about 100% identical to the amino acid sequence of SEQ ID NO: 4. In certain embodiments, an RNP complex comprising a nuclease disclosed herein can further comprise at least one STAR gRNA (double-stranded guide RNA). In certain embodiments, an RNP complex comprising a nuclease disclosed herein can further comprise at least one non-STAR gRNA (e.g., single-stranded guide RNA). In certain embodiments, an RNP complex comprising a nuclease disclosed herein can further comprise at least one polynucleotide. In certain embodiments, a polynucleotide comprised in an RNP complex disclosed herein can be greater than about 50 nucleotides in length. In certain embodiments, a polynucleotide comprised in an RNP complex disclosed herein can be about 50, about 150, about 500, about 1000, or more than 1000 nucleotides in length.In certain embodiments, two or more nucleases can be added to RNP complex to affect overall editing efficiency.In certain embodiments, two or more gRNAs can be added to RNP complex, which allows multiple editing of two or more sites in one transfection, improving efficiency.In other embodiments, two or more DNA templates can be added to RNP to enable multiple editing at one or more sites based on specific desired repair results.

[0155] In certain embodiments, a composition comprising a type V, e.g., type VA, CRISPR nuclease polypeptide as described herein further comprises a guide nucleic acid (gNA), e.g., gRNA, or a polynucleotide encoding a gNA, e.g., gRNA, comprising a spacer sequence that targets a target nucleotide sequence (also referred to herein as a target nucleic acid sequence) within a polynucleotide (as the context may dictate), wherein the gNA, e.g., gRNA, is compatible with a type V, e.g., type VA, CRISPR nuclease. Generally, as used herein, a polynucleotide containing a target nucleotide sequence (target nucleic acid sequence) includes a polynucleotide containing the target nucleotide sequence (target nucleic acid sequence). Such a polynucleotide may be any suitable polynucleotide, such as a cellular genome or a portion of a cellular genome. In certain embodiments, the target nucleotide sequence (target nucleic acid sequence) is within 50 nucleotides of a protospacer adjacent motif (PAM) sequence specific for a V-type CRISPR nuclease, such as a PAM comprising the sequence of YTTN, where Y is T or C, and N is A, T, G, or C, or a sequence of YTTV or TTTV, where V is A, G, or C. In certain embodiments, the PAM comprises the sequence of YTTV or TTTV, where V is A, G, or C. In certain embodiments, the gRNA is a gRNA, such as a double-stranded (split) gRNA. A gNA, e.g., a gRNA, can include one or more chemical modifications, such as 2'-O-alkyl, 2'-O-methyl, phosphorothioate, phosphonoacetate, thiophosphonoacetate, 2'-O-methyl-3'-phosphorothioate, 2'-O-methyl-3'-phosphonoacetate, 2'-O-methyl-3'-thiophosphonoacetate, 2'-deoxy-3'-phosphonoacetate, 2'-deoxy-3'-thiophosphonoacetate, suitable alternatives, or a combination thereof.In certain embodiments, the guanine:uracil ratio in the gRNA is at least 51:49, 52:48, 53:47, 54:46, 55:45, 56:44, 57:43, 58:42, 59:42, or 60:40, preferably at least 53:47, more preferably at least 54:46, and even more preferably at least 55:45. See Example 12 and Figure 10. In certain embodiments, the molar ratio of gNA (e.g., gRNA) to Type V CRISPR nuclease is at least 1.1:1, 1.2:1, 1.3:1, 1.4:1, 1.5:1, 1.6:1, 1.7:1, 1.8:1, 2:1, 2.2:1, 2.5:1, or 3:1 and / or is 1.2:1, 1.3:1, 1.4:1, 1.5:1, 1.6:1, 1.7:1, 1.8:1, 2:1, 2.2:1, 2.5:1, 3:1, or 4:1 or less, preferably 1.1:1 to 2.5:1, more preferably 1.2:1 to 2:1, even more preferably 1.2:1 to 1.7:1. See, e.g., Example 13. In certain embodiments, the molar amount of gNA (e.g., gRNA) is at least 10, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 170, 190, or 200 pmol, and / or no more than 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 170, 190, 200, 250, or 300 pmol, preferably 25-200 pmol, more preferably 50-100 pmol, and even more preferably 65-85 pmol. See Example 13.

[0156] In certain embodiments, the composition comprising a type V, e.g., type VA, CRISPR nuclease polypeptide described herein further comprises a donor template, also referred to herein as an editing template. The donor template can comprise homology arms, i.e., nucleotide sequences complementary to the polynucleotide sequences on either side of the cleavage site where the donor template is inserted. The donor template can be added in any suitable amount, for example, in certain embodiments, at least 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 1.1, 1.2, 1.3, 1.4, 1.5, 1.7, 2, 2.5, 3, 4, or 5 μg μL. -1 and / or 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 1.1, 1.2, 1.3, 1.4, 1.5, 1.7, 2, 2.5, 3, 4, 5, 7, or 10 μg μL -1 Less than 0.3 to 2 μg μL -1 , more preferably 0.5 to 1.5 μg μL -1 , and even more preferably 0.8 to 1.2 μg μL -1 It can exist in

[0157] In certain embodiments, a composition comprising a Type V, e.g., Type VA, CRISPR nuclease polypeptide described herein further comprises an anionic polymer. Any suitable anionic polymer can be used. Examples of anionic polymers include 1,2,3-heptanetriol, 2-amino-2-(hydroxymethyl)-1,3-propanediol (Tris), 3-(1-pyridino)-1-propanesulfonate (NDSB201), 3-[(3-cholamidopropyl)dimethylammonio]-1-propanesulfonate (CHAPS), 6-aminocaproic acid, adenosine diphosphate (ADP), adenosine triphosphate (ATP), α-cyclodextrin, amidosulfobetaine-14 (ASB-14), ammonium acetate, ammonium nitrate, ammonium sulfate, arginine, arginine ethyl ester, barium chloride, barium iodide, benzamidine HCl, β-cyclodextrin, beta-mercaptoethanol (BME), biotin, calcium chloride, cesium chloride. Calcium, cesium sulfate, cetyltrimethylammonium bromide (CTAB), choline chloride, citric acid, cobalt chloride, copper(II) chloride, cyclohexanol, D-sorbitol, dimethylethylammonium propanesulfonate (NDSB195), dithiothreitol (DTT), erythritol, ethanol, ethylene glycol, ethylene glycol-bis(ββ-aminoethyl ether)-N,N,N',N'-tetraacetic acid (EGTA), ethylenediaminetetraacetic acid (EDTA), formamide, gadolinium bromide, gamma-butyrolactone, glucose, glutamic acid, glutamine, glycerol, glycine, glycine betaine, glycine-glycine-glycine, guanidine HCl, guanosine triphosphate (GTP), holmium chloride, imidazole, iron(III) chloride, JeffamineM-600, lanthanum acetate, lauryl sulfobetaine, lauryldimethylamine N-oxide (LDAO), lithium sulfate, magnesium chloride, magnesium sulfate, manganese chloride, mannitol, N-(2-hydroxyethyl)piperazine-N'-(3-propanesulfonic acid) (EPPS), N-dodecyl beta-D-maltoside (DDM), N-ethyl urea, n-hexanol, N-lauryl sarcoside, N-lauryl sarcosine, N-methylformamide, N-methyl urea, n-octyl-b-D-glucoside (OG: octyl glucoside), n-pentanol, nickel chloride, non-detergent sulfobetaine (NDSB), Nonidet P40 (NP40), octyl beta-D-glucopyranoside, poly-L-glutamic acid, polyethylene glycol (e.g., PEG300, PEG3350, PEG4000), polyethylene glycol lauryl ether (Brij35), polyoxyethylene (2) oleyl ether (Brij93), polyoxyethylene cetyl ether (Brij56), polyvinylpyrrolidone 40 (PVP40), potassium chloride, potassium citrate, potassium nitrate, proline, putrescine, spermidine, spermine, ribonucleic acid, PEG-100, PEG-1200, PEG-1400, PEG-1500, PEG-1600, PEG-1700, PEG-1800, PEG-1900, PEG-2000, PEG-2100, PEG-2200, PEG-2300, PEG-2400, PEG-2500, PEG-2600, PEG-2700, PEG-2800, PEG-2900, PEG-3000, PEG-3100, PEG-3200, PEG-3350, PEG-4000, PEG-4100, PEG-4200, PEG-4300, PEG-4400, PEG-4500, PEG-4600, PEG-4700, PEG-4800, PEG-49 ... Flavin, samarium bromide, sarcosine, sodium acetate, sodium chloride, sodium dodecyl sulfate (SDS), sodium fluoride, sodium iodide, sodium lauroyl sarcosinate (sarcosyl), sodium malonate, sodium molybdate, sodium selenite, sodium sulfate, sodium thiocyanate, sucrose, taurine, trehalose, tricine, triethylamine, trimethylamine N-oxide (TMAO), tris(2-carboxyethyl)phosphine (TCEP), TritonX-100, Tween 20, Tween 60, Tween 80, urea, vitamin B12, xylitol, yttrium chloride, yttrium nitrate, zinc chloride, Zwittergent 3-08, Zwittergent 3-14, or combinations thereof. In certain embodiments, the anionic polymer comprises polyglutamic acid. In certain embodiments, the anionic polymer, e.g., PGA, is at least 20, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 170, 200, 250, 300, 400, or 500 μg μL -1 , and / or 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 170, 200, 250, 300, 400, 500, 700, or 1000 μg μL -1 Below 20-200 μg μL, preferably -1 , more preferably 50 to 150 μg μL -1 , and even more preferably 80 to 120 μg μL -1 (PGA).

[0158] In certain embodiments, the present disclosure provides a cell comprising one or more of the compositions described herein, for example, a composition comprising a V-type, for example, a VA-type CRISPR nuclease polypeptide, which comprises one or more NLSs, and in certain embodiments, a purification tag and / or cleavage site.Any suitable cell can be used.In certain embodiments, the cell is a human cell, such as an immune cell, for example, a T cell, or a stem cell, for example, an induced pluripotent stem cell (iPSC).

[0159] In certain embodiments, the present disclosure provides a method for inserting one or more of the compositions described herein, for example, compositions comprising one or more NLSs, and in certain embodiments, a CRISPR nuclease polypeptide of type V, for example, type VA, which comprises a purification tag and / or cleavage site, into cells.Any suitable method can be used for insertion.In certain embodiments, electroporation is used.Electroporation conditions can be optimized.See, for example, Examples.

[0160] In certain embodiments, methods are provided that involve modifying a target polynucleotide, comprising contacting the target polynucleotide with a composition(s) described herein (e.g., a composition comprising one or more NLSs and a suitable gRNA, e.g., a type V, e.g., type VA, CRISPR nuclease polypeptide), thereby enabling the composition to modify the target polynucleotide, and in some cases a genomic region, e.g., a genome or portion of a genome in a human cell, such as a cell, e.g., an immune cell (e.g., a T cell), or a stem cell, e.g., an iPSC. In certain cases, the composition(s) comprise a donor template, such as a donor template comprising a polynucleotide encoding a polypeptide to be expressed by the cell, and in certain embodiments, the polypeptide comprises a chimeric antigen receptor (CAR) or portion thereof; see, e.g., the Examples. In certain embodiments, the cell is a human cell, e.g., an immune cell, such as a T cell, or a stem cell, such as an iPSC.

[0161] Nuclease System Certain embodiments disclosed herein are targetable nuclease systems. In certain embodiments, the targetable nuclease system can comprise a nucleic acid-guided nuclease and a compatible guide nucleic acid (also referred to interchangeably herein as a "guide polynucleotide" and a "gRNA"). The targetable nuclease system can comprise a nucleic acid-guided nuclease or a polynucleotide sequence encoding the nucleic acid-guided nuclease. The targetable nuclease system can comprise a guide nucleic acid or a polynucleotide sequence encoding the guide nucleic acid.

[0162] In general, the targetable nuclease systems disclosed herein can be characterized by elements that promote the formation of a targetable nuclease complex at the site of a target sequence, where the targetable nuclease complex comprises a nucleic acid-guided nuclease and a guide nucleic acid.

[0163] The guide nucleic acid, together with the nucleic acid-guided nuclease, forms a targetable nuclease complex that is capable of binding to a target sequence within a target polynucleotide as determined by the guide sequence of the guide nucleic acid.

[0164] Generally, generating a double-stranded break requires that in most cases a targetable nuclease complex binds to a target sequence as determined by a guide nucleic acid, and the nuclease recognizes a protospacer adjacent motif (PAM) sequence adjacent to the target sequence.

[0165] A targetable nuclease complex can include a nucleic acid-guided nuclease and a compatible guide nucleic acid having an amino acid sequence at least 50% identical to SEQ ID NO:2. A targetable nuclease complex can include a nucleic acid-guided nuclease and a compatible guide nucleic acid having an amino acid sequence at least about 60%, 65%, 75%, 85%, 95%, 99%, or about 100% identical to the amino acid sequence of SEQ ID NO:2. A protospacer adjacent motif (PAM) sequence flanks the target sequence. A targetable nuclease complex can include a nucleic acid-guided nuclease and a compatible guide nucleic acid having an amino acid sequence at least 50% identical to SEQ ID NO:3. A targetable nuclease complex can include a nucleic acid-guided nuclease and a compatible guide nucleic acid having a nuclease having an amino acid sequence at least about 60%, 65%, 75%, 85%, 95%, 99%, or about 100% identical to the amino acid sequence of SEQ ID NO:3. The targetable nuclease complex can include a nucleic acid-guided nuclease having an amino acid sequence at least 50% identical to SEQ ID NO:4 and a compatible guide nucleic acid. The targetable nuclease complex can include a nucleic acid-guided nuclease and a compatible guide nucleic acid, with the nuclease having an amino acid sequence at least about 60%, 65%, 75%, 85%, 95%, 99%, or about 100% identical to the amino acid sequence of SEQ ID NO:4. In certain embodiments, the guide nucleic acid can include a scaffold sequence that is compatible with the selected nucleic acid-guided nuclease. In any of these embodiments, the guide sequence can be engineered to be complementary to any desired target sequence. The selected guide sequence can be engineered to hybridize to any desired target sequence. In certain embodiments, the guide sequence is a double-stranded guide RNA.

[0166] The target sequence of a targetable nuclease complex can be any polynucleotide, endogenous or exogenous to a prokaryotic or eukaryotic cell, or in vitro. For example, the target sequence can be a polynucleotide present in the nucleus of a eukaryotic cell. The target sequence can be a sequence encoding a gene product (e.g., a protein) or a non-coding sequence (e.g., a regulatory polynucleotide or junk DNA). As used herein, the target sequence is considered to be associated with a PAM, i.e., a short sequence recognized by the targetable nuclease complex. While the exact sequence and length requirements of the PAM vary depending on the nucleic acid-guided nuclease used, the PAM can be a 2- to 5-base pair sequence adjacent to the target sequence. Examples of PAM sequences are provided in the Examples section below, and those skilled in the art can identify additional PAM sequences for use with a given nucleic acid-guided nuclease. Furthermore, engineering the PAM-interacting (PI) domain can enable programming of PAM specificity, improving the fidelity of target site recognition and enhancing the versatility of the nucleic acid-guided nuclease genome engineering platform. Nucleic acid-guided nucleases may be engineered to alter their PAM specificity, for example, as described in Kleinstiver et al., Nature. 2015 Jul. 23;523(7561):481-5, the disclosure of which is incorporated herein in its entirety.

[0167] A PAM site is a nucleotide sequence adjacent to a target sequence. In most cases, a nucleic acid-guided nuclease can cleave a target sequence only if the appropriate PAM is present. The PAM is nucleic acid-guided nuclease-specific and can vary between two different nucleic acid-guided nucleases. The PAM can be 5' or 3' of the target sequence. The PAM can be upstream or downstream of the target sequence. The PAM can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more nucleotides in length. Most often, the PAM is 2-6 nucleotides in length.

[0168] In some embodiments disclosed herein, the PAM may be provided on a separate oligonucleotide, in which case providing the PAM on the oligonucleotide allows for cleavage of an otherwise uncleavable target sequence, since there is no adjacent PAM on the same polynucleotide as the target sequence.

[0169] The polynucleotide sequence encoding the components of the targetable nuclease system can comprise one or more vectors. Generally, as used herein, the term "vector" can refer to a nucleic acid molecule capable of transporting another nucleic acid to which it is linked. Vectors include, but are not limited to, nucleic acid molecules that are single-stranded, double-stranded, or partially double-stranded; nucleic acid molecules that contain one or more free ends or no free ends (e.g., circular); nucleic acid molecules that contain DNA, RNA, or both; and other types of polynucleotides known in the art. One type of vector is a "plasmid," which refers to a circular double-stranded DNA loop into which additional DNA segments can be inserted, such as by standard molecular cloning techniques. Another type of vector is a viral vector, in which viral-derived DNA or RNA sequences are present within the vector and packaged into a virus (e.g., retrovirus, replication-deficient retrovirus, adenovirus, replication-deficient adenovirus, and adeno-associated virus). Other vectors (e.g., non-episomal mammalian vectors) can be integrated into the genome of a host cell upon introduction into the host cell. A recombinant expression vector can contain a nucleic acid of the invention in a form suitable for expression of the nucleic acid in a host cell, which can mean that the recombinant expression vector contains one or more regulatory elements, which can be selected based on the host cell to be used for expression, operably linked to the nucleic acid sequence to be expressed.

[0170] In some embodiments, a regulatory element can be operably linked to one or more elements of the targetable nuclease system to drive expression of one or more components of the targetable nuclease system.

[0171] In some embodiments, the vector can include a regulatory element operably linked to the polynucleotide sequence encoding the nucleic acid-guided nuclease. The polynucleotide sequence encoding the nucleic acid-guided nuclease can be codon-optimized for expression in a target cell, such as a prokaryotic cell or a eukaryotic cell. The eukaryotic cell can be a yeast, fungus, algae, plant, animal, or human cell. The eukaryotic cell can be derived from an organism such as a mammal, including, but not limited to, a human, mouse, rat, rabbit, dog, or non-human mammal, including a non-human primate.

[0172] Generally, codon optimization can refer to the process of modifying a nucleic acid sequence to enhance expression in a target host cell by replacing at least one codon of the native sequence with a codon that is more frequently or most frequently used in the host cell gene while maintaining the native amino acid sequence. Various species exhibit specific biases toward codons for specific amino acids. As contemplated herein, genes can be adjusted for optimal gene expression in a given organism based on codon optimization. Codon usage tables are readily available, such as the "Codon Usage Database" available at www.kazusa.or.jp, and these tables can be adapted in many ways. See Nakamura, Y., et al., "Codon usage tabulated from the international DNA sequence databases: status for the year 2000," Nucl. Acids Res. 28:292 (2000).

[0173] The nucleic acid-guided nuclease and one or more guide nucleic acids can be delivered as either DNA or RNA. Delivering both the nucleic acid-guided nuclease and the guide nucleic acid as RNA (unmodified or containing base or backbone modifications) molecules can be used to shorten the time the nucleic acid-guided nuclease remains in cells. This can reduce the level of off-target cleavage activity in target cells. Because nucleic acid-guided nucleases as mRNAs take time to be translated into proteins, it can be advantageous to deliver the guide nucleic acid several hours after delivery of the nucleic acid-guided nuclease mRNA to maximize the level of guide nucleic acid available for interaction with the nucleic acid-guided nuclease protein. In other cases, the nucleic acid-guided nuclease mRNA and the guide nucleic acid are delivered simultaneously. In other examples, the guide nucleic acid is delivered sequentially, for example, 0.5, 1, 2, 3, 4 hours or more after the nucleic acid-guided nuclease mRNA.

[0174] The guide nucleic acid in the form of RNA or encoded on a DNA expression cassette can be introduced into a host cell, which can contain a nucleic acid-guided nuclease encoded on a vector or chromosome. The guide nucleic acid can be provided in the cassette as one or more polynucleotides, which can be contiguous or discontinuous within the cassette. In certain embodiments, the guide nucleic acid is provided in the cassette as a single contiguous polynucleotide.

[0175] Various delivery systems can be used to introduce the nucleic acid-guided nuclease (DNA or RNA) and guide nucleic acid (DNA or RNA) into host cells. According to these embodiments, useful systems include, but are not limited to, yeast systems, lipofection systems, microinjection systems, biolistics systems, virosomes, liposomes, immunoliposomes, polycations, lipid:nucleic acid conjugates, virions, artificial virions, viral vectors, electroporation, cell-penetrating peptides, nanoparticles, nanowires (Shalek et al., Nano Letters, 2012), and exosomes. Molecular Trojan horse liposomes (Pardridge et al., Cold Spring Harb Protoc; 2010; doi:10.1101 / pdb.prot5407) can be used to deliver the engineered nuclease and guide nuclease across the blood-brain barrier.

[0176] In some embodiments, an editing template, also referred to herein as a donor template, is also provided. The editing template may be a component of the vector described herein, may be included in a separate vector, or may be provided as a separate polynucleotide, such as an oligonucleotide, linear polynucleotide, or synthetic polynucleotide. In some cases, the editing template is located on the same polynucleotide as the guide nucleic acid. In some embodiments, the editing template is designed to function as a template in homologous recombination, such as within or near the target sequence nicked or cleaved by the nucleic acid-guided nuclease as part of the complex disclosed herein. The editing template polynucleotide may be of any suitable length, such as about 10, 15, 20, 25, 50, 75, 100, 150, 200, 500, 1000, or more, or more than about 10, 15, 20, 25, 50, 75, 100, 150, 200, 500, 1000, or more nucleotides in length. In some embodiments, the editing template polynucleotide is complementary to a portion of a polynucleotide that may comprise a target sequence. When optimally aligned, the editing template polynucleotide may overlap one or more nucleotides of the target sequence (e.g., about 1, 5, 10, 15, 20, 25, 30, 35, 40, or more, or more than about 1, 5, 10, 15, 20, 25, 30, 35, 40, or more nucleotides). In some embodiments, when the editing template sequence and a polynucleotide that may comprise the target sequence are optimally aligned, the nearest nucleotide of the template polynucleotide is within about 1, 5, 10, 15, 20, 25, 50, 75, 100, 200, 300, 400, 500, 1000, 5000, 10000, or more nucleotides of the target sequence.

[0177] In some embodiments, methods are provided for delivering one or more polynucleotides, such as one or more vectors or linear polynucleotides described herein, one or more transcription products thereof, and / or one or more proteins transcribed therefrom, to a host cell. In some aspects, the invention further provides cells produced by such methods, and organisms can comprise or be produced from such cells. In some embodiments, an engineered nuclease in combination (and optionally complexed) with a guide nucleic acid is delivered to the cell.

[0178] Conventional viral and non-viral gene transfer methods can be used to introduce nucleic acids into cells, such as prokaryotic cells, eukaryotic cells, mammalian cells, or target tissues. These methods can be used to administer nucleic acids encoding components of engineered nucleic acid-guided nuclease systems to cells in culture or host organisms. Non-viral vector delivery systems include DNA plasmids, RNA (e.g., transcripts of the vectors described herein), naked nucleic acids, and nucleic acids complexed with delivery vehicles such as liposomes. Viral vector delivery systems include DNA and RNA viruses that have either episomal or integrated genomes after delivery to cells. Any gene therapy method known in the art is considered useful herein. Non-viral delivery methods of nucleic acids are contemplated herein. Adeno-associated virus ("AAV") vectors can also be used to transduce target nucleic acids into cells, for example, in in vitro production of nucleic acids and peptides, and for in vivo and ex vivo gene therapy procedures.

[0179] In some embodiments, host cells are transiently or non-transiently transfected with one or more vectors, linear polynucleotides, polypeptides, nucleic acid-protein complexes, or any combination thereof described herein. In some embodiments, cells are transfected in vitro, in culture, or ex vivo. In some embodiments, cells are transfected as they naturally occur in a subject. In some embodiments, transfected cells are harvested from a subject. In some embodiments, cells are derived from cells harvested from a subject, such as a cell line.

[0180] In some embodiments, cells transfected with one or more vectors, linear polynucleotides, polypeptides, nucleic acid-protein complexes, or any combination thereof described herein are used to establish new cell lines that can contain one or more transfection-derived sequences. In some embodiments, cells transiently transfected with components of the engineered nucleic acid-guided nuclease systems described herein (such as by transient transfection of one or more vectors, or transfection with RNA) and modified through the activity of the engineered nuclease complex are used to establish new cell lines, which can include cells that contain the modifications but lack any other exogenous sequences.

[0181] In some embodiments, one or more vectors described herein are used to produce non-human transgenic cells, organisms, animals, or plants. In some embodiments, the transgenic animal is a mammal, such as a mouse, rat, or rabbit. Methods for producing transgenic cells, organisms, plants, and animals are known in the art and generally begin with a method of cell transformation or transfection, such as those described herein.

[0182] In certain embodiments, in an engineered nuclease complex, a "target sequence" can refer to a sequence to which a guide sequence is designed to have complementarity, where hybridization between the target sequence and the guide sequence promotes the formation of an engineered nuclease complex. The target sequence can comprise any polynucleotide, such as DNA, RNA, or a DNA-RNA hybrid. The target sequence can be located in the nucleus or cytoplasm of a cell. The target sequence can be placed in vitro or in a cell-free environment.

[0183] In some embodiments, the formation of an engineered nuclease complex can include a guide nucleic acid hybridized to a target sequence and complexed with one or more novel engineered nucleases as disclosed herein, resulting in cleavage of one or both strands within or near the target sequence (e.g., within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, or more base pairs therefrom). Cleavage can occur within the target sequence, 5' of the target sequence, upstream of the target sequence, 3' of the target sequence, or downstream of the target sequence.

[0184] In some embodiments, one or more vectors driving the expression of one or more components of a targetable nuclease system are introduced into a host cell or in vitro, resulting in the formation of a targetable nuclease complex at one or more target sites. For example, the nucleic acid-guided nuclease and guide nucleic acid can each be operably linked to separate regulatory elements on separate vectors. Alternatively, two or more elements expressed from the same or different regulatory elements can be combined into a single vector, along with one or more additional vectors providing any components of the targetable nuclease system not included in the first vector. Targetable nuclease system elements combined in a single vector can be arranged in any suitable orientation, such as one element being located 5' ("upstream") or 3' ("downstream") relative to the second element. The coding sequence of one element can be located on the same or opposite strand of the coding sequence of the second element and oriented in the same or opposite direction. In some embodiments, a single promoter drives the expression of transcripts encoding the nucleic acid-guided nuclease and one or more guide nucleic acids. In some embodiments, the nucleic acid-guided nuclease and one or more guide nucleic acids are operably linked to and expressed from the same promoter. In other embodiments, one or more guide nucleic acids, or polynucleotides encoding one or more guide nucleic acids, are introduced into a cell or in vitro environment that may already contain the nucleic acid-guided nuclease or a polynucleotide sequence encoding the nucleic acid-guided nuclease.

[0185] In some embodiments, when multiple different guide sequences are used, a single expression construct can be used to target nuclease activity to multiple different corresponding target sequences in cells or in vitro. For example, a single vector can contain about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, or more, or more than about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, or more guide sequences. In other embodiments, about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more, or more than about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more, or more than about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more such guide sequence-containing vectors can be provided and optionally delivered to cells in vivo or in vitro.

[0186] In some embodiments, the methods and compositions disclosed herein can include two or more guide nucleic acids, such that each guide nucleic acid has a different guide sequence and thereby targets a different target sequence. According to these embodiments, multiple guide nucleic acids can be used for multiplexing, where multiple targets are targeted simultaneously. Additionally or alternatively, multiple guide nucleic acids can be introduced into a cell population, such that each cell in the population receives a different or random guide nucleic acid, thereby targeting multiple different target sequences across the cell population. In such cases, the resulting collection of altered cells can then be referred to as a library.

[0187] In other embodiments, the methods and compositions disclosed herein can include multiple different nucleic acid-guided nucleases, each with one or more different corresponding guide nucleic acids, thereby enabling targeting of different target sequences by the different nucleic acid-guided nucleases. In some such cases, each nucleic acid-guided nuclease can correspond to a separate plurality of guide nucleic acids, enabling two or more non-overlapping, partially overlapping, or fully overlapping multiplexing events.

[0188] In some embodiments, the nucleic acid-guided nuclease has DNA cleavage activity or RNA cleavage activity. In some embodiments, the nucleic acid-guided nuclease directs cleavage of one or both strands at the location of a target sequence, such as within the target sequence and / or within the complement of the target sequence. In some embodiments, the nucleic acid-guided nuclease directs cleavage of one or both strands within about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 50, 100, 200, 500, or more base pairs from the first or last nucleotide of the target sequence.

[0189] In certain embodiments, the present invention provides methods for modifying a target sequence in vitro or in a prokaryotic or eukaryotic cell, which may be in vivo, ex vivo, or in vitro. In some embodiments, the method includes extracting a cell or cell population, such as a prokaryotic cell, or a cell or cell population from a human or non-human animal or plant (including microalgae or other organisms), and modifying the cell(s). Culturing can be performed at any stage, in vitro or ex vivo. The cell(s) can even be reintroduced into a host, such as a non-human animal or plant (including microalgae). In the case of reintroduced cells, they can be stem cells.

[0190] In some embodiments, the method comprises binding a targetable nuclease complex to a target sequence to cause cleavage of the target sequence, thereby modifying the target sequence, wherein the targetable nuclease complex comprises a nucleic acid-guided nuclease complexed with a guide nucleic acid, wherein the guide sequence of the guide nucleic acid hybridizes to a target sequence within the target polynucleotide. In some aspects, the invention provides methods of modifying expression of a target polynucleotide in vitro or in a prokaryotic or eukaryotic cell. In some embodiments, the method comprises binding a targetable nuclease complex to a target sequence having a target polynucleotide, such that binding can result in increased or decreased expression of the target polynucleotide, wherein the targetable nuclease complex comprises a nucleic acid-guided nuclease complexed with a guide nucleic acid, wherein the guide sequence of the guide nucleic acid hybridizes to a target sequence within the target polynucleotide.

[0191] In certain embodiments, the present invention provides kits comprising any one or more of the elements disclosed in the methods and compositions above. The elements can be provided individually or in combination and can be provided in any suitable container, such as a vial, bottle, or tube. In some embodiments, the kit includes instructions in one or more languages, for example, in two or more languages.

[0192] In some embodiments, the kit includes one or more reagents for use in a process utilizing one or more of the components described herein. The reagents may be provided in any suitable container. For example, the kit may provide one or more reaction or storage buffers. The reagents may be provided in a form that is ready for use in an assay or that requires the addition of one or more other components prior to use (e.g., a concentrate or lyophilized form). The buffer may be any buffer, including, but not limited to, sodium carbonate buffer, sodium bicarbonate buffer, borate buffer, Tris buffer, MOPS buffer, HEPES buffer, and combinations thereof. In some embodiments, the buffer is alkaline. In some embodiments, the buffer has a pH of about 7 to about 10. In some embodiments, the kit includes one or more oligonucleotides corresponding to a guide sequence for insertion into a vector to operably link the guide sequence and the regulatory element. In some embodiments, the kit includes an editing template.

[0193] In some embodiments, the targetable nuclease complexes have a wide variety of utilities, including modification (e.g., deletion, insertion, translocation, inactivation, activation) of target sequences in various cell types. Such targetable nuclease complexes of the present invention have broad applications, for example, in biochemical pathway optimization, genome-wide studies, genome engineering, gene therapy, drug screening, disease diagnosis, and prognosis. An exemplary targetable nuclease complex comprises a nucleic acid-guided nuclease disclosed herein complexed with a guide nucleic acid, wherein the guide sequence of the guide nucleic acid can hybridize to a target sequence within a target polynucleotide. The guide nucleic acid can comprise a guide sequence linked to a scaffold sequence. The scaffold sequence can comprise one or more sequence regions that have a degree of complementarity such that they together form a secondary structure.

[0194] The editing template polynucleotide can include a sequence to be integrated (e.g., a mutant gene). The sequence for integration can be an endogenous or exogenous sequence to the cell. Examples of sequences to be integrated include protein-coding polynucleotides or non-coding RNAs (e.g., microRNAs). Thus, the sequence for integration can be operably linked to a suitable control sequence(s). Alternatively, the sequence to be integrated may provide a regulatory function. The sequence to be integrated can be a mutation or variant of an endogenous wild-type sequence. Alternatively, the sequence to be integrated can be a wild-type version of an endogenous mutant sequence. Additionally or alternatively, the sequence to be integrated can be a variant or mutant form of an endogenous mutant or variant sequence.

[0195] In certain embodiments, the upstream or downstream sequence can comprise from about 20 bp to about 2500 bp, e.g., about 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 2100, 2200, 2300, 2400, or about 2500 bp. In some embodiments, exemplary upstream or downstream sequences have from about 15 bp to about 2000 bp, from about 30 bp to about 1000 bp, from about 50 bp to about 750 bp, from about 600 bp to about 1000 bp, or from about 700 bp to about 1000 bp.

[0196] In some embodiments, the editing template polynucleotide can further comprise a marker. In certain embodiments, some markers can facilitate screening for targeted integration. Examples of suitable markers can include, but are not limited to, restriction sites, fluorescent proteins, or selection markers. In certain embodiments, the exogenous polynucleotide template can be constructed using recombinant techniques.

[0197] In one embodiment of an exemplary method for modifying a target polynucleotide by incorporating an editing template polynucleotide, a double-stranded break is introduced into a genomic sequence by an engineered nuclease complex, and the break can be repaired by homologous recombination using the editing template, resulting in the incorporation of the template into the target polynucleotide. The presence of the double-stranded break can increase the efficiency of incorporation of the editing template.

[0198] Disclosed herein are methods for modifying expression of a polynucleotide in a cell. Some methods involve increasing or decreasing expression of a target polynucleotide by using a targetable nuclease complex that binds to the target polynucleotide.

[0199] Detection of gene expression levels can be performed in real time in amplification assays. In one embodiment, amplified products can be directly visualized using fluorescent DNA-binding agents, including, but not limited to, DNA intercalators and DNA groove binders. Because the amount of intercalator incorporated into double-stranded DNA molecules can be proportional to the amount of amplified DNA product, the amount of amplified product can be easily determined by quantifying the fluorescence of the intercalated dye using conventional optical systems in the art. DNA-binding dyes suitable for this purpose include, but are not limited to, SYBR Green, SYBR Blue, DAPI, propidium iodide, Hoeste, SYBR Gold, ethidium bromide, acridine, proflavine, acridine orange, acriflavine, fluorocoumaine, ellipticine, daunomycin, chloroquine, distamycin D, chromomycin, homidium, mithramycin, ruthenium polypyridyl, anthramycin, and others known to those skilled in the art.

[0200] In some embodiments, other fluorescent labels, such as sequence-specific probes, can be used in the amplification reaction to facilitate detection and quantification of the amplified products. Probe-based quantitative amplification relies on sequence-specific detection of the desired amplification product. It utilizes fluorescent target-specific probes (e.g., TaqMan™ probes) to improve specificity and sensitivity. Methods for performing probe-based quantitative amplification are well established in the art.

[0201] In some embodiments, drug-induced changes in the expression of sequences associated with a signal transduction biochemical pathway can also be determined by examining the corresponding gene product. Determining protein levels can involve (a) contacting proteins contained in a biological sample with an agent that specifically binds to a protein associated with a signal transduction biochemical pathway, and (b) identifying any agent:protein complexes thus formed. In one aspect of this embodiment, the agent that specifically binds to a protein associated with a signal transduction biochemical pathway is an antibody, preferably a monoclonal antibody.

[0202] In some embodiments, the amount of drug:polypeptide complex formed during the binding reaction can be quantified by a standard quantitative assay. As shown above, the formation of the drug:polypeptide complex can be measured directly by the amount of label remaining at the binding site. In another method, proteins associated with a signal transduction biochemical pathway are tested for their ability to compete with a labeled analog for the binding site of a particular drug. In this competitive assay, the amount of captured label is inversely proportional to the amount of protein sequence associated with a signal transduction biochemical pathway present in the test sample.

[0203] In some embodiments, many techniques for protein analysis based on the general principles outlined above are known in the art and are contemplated herein, including, but not limited to, radioimmunoassays, ELISAs (enzyme-linked immunoradiometric assays), "sandwich" immunoassays, immunoradiometric assays, in situ immunoassays (e.g., using colloidal gold, enzyme, or radioisotope labels), Western blot analysis, immunoprecipitation assays, immunofluorescence assays, and SDS-PAGE.

[0204] In some embodiments, when practicing the subject methods, it may be desirable to identify expression patterns of proteins associated with signaling biochemical pathways in different body tissues, different cell types, and / or different subcellular structures. These studies can be performed using tissue-, cell-, or subcellular structure-specific antibodies that can bind to protein markers that are selectively expressed in particular tissues, cell types, or subcellular structures.

[0205] In other embodiments, altered expression of genes associated with a signal transduction biochemical pathway can also be determined by examining changes in the activity of the gene product compared to control cells. Assays for drug-induced changes in the activity of proteins associated with a signal transduction biochemical pathway depend on the biological activity and / or signal transduction pathway under investigation. For example, if the protein is a kinase, changes in its ability to phosphorylate downstream substrate(s) can be determined by various assays known in the art. Representative assays include, but are not limited to, immunoblotting and immunoprecipitation using antibodies, such as anti-phosphotyrosine antibodies, that recognize phosphorylated proteins. Furthermore, kinase activity can be detected by high-throughput chemiluminescence assays.

[0206] In certain embodiments, where a protein associated with a signaling biochemical pathway is part of a signal transduction cascade that leads to fluctuations in intracellular pH conditions, a pH-sensitive molecule, such as a fluorescent pH dye, can be used as a reporter molecule. In another example, where the protein associated with a signaling biochemical pathway is an ion channel, fluctuations in membrane potential and / or intracellular ion concentration can be monitored. Many commercially available kits and high-throughput devices are suitable for rapid and robust screening of ion channel modulators. Representative devices include FLIPR™ (Molecular Devices, Inc.) and VIPR (Aurora Biosciences). These devices can simultaneously detect responses from over 1,000 sample wells in a microplate and provide measurement and functional data in real time within one second or even one millisecond.

[0207] When performing any of the methods disclosed herein, a suitable vector can be introduced into a cell, tissue, organism, or embryo via one or more methods known in the art, including, but not limited to, microinjection, electroporation, sonoporation, biolistics, calcium phosphate-mediated transfection, cationic transfection, liposome transfection, dendrimer transfection, heat shock transfection, nucleofection transfection, magnetofection, lipofection, impale infection, phototransfection, proprietary agent-enhanced nucleic acid uptake, and delivery via liposomes, immunoliposomes, virosomes, or artificial virions. In some methods, the vector is introduced into the embryo by microinjection. The vector(s) can be microinjected into the nucleus or cytoplasm of the embryo. In some methods, the vector(s) can be introduced into the cell by nucleofection.

[0208] The target polynucleotide of targetable nuclease complex can be any polynucleotide that is endogenous or exogenous to host cell.For example, the target polynucleotide can be a polynucleotide that exists in the nucleus of eukaryotic cell, the genome of prokaryotic cell, or an extrachromosomal vector of host cell.The target polynucleotide can be a sequence that codes for a gene product (e.g., protein) or a non-coding sequence (e.g., regulatory polynucleotide or junk DNA).

[0209] Some embodiments disclosed herein relate to the use of the engineered nucleic acid-guided nuclease system disclosed herein, for example, to target and knock out genes, amplify genes, and / or repair specific mutations associated with DNA repeat instability and medical disorders. The nuclease system can be used to suppress and correct these genomic instability defects. In other embodiments, the engineered nucleic acid-guided nuclease system disclosed herein can be used to correct genetic defects associated with Lafora disease. Lafora disease is an autosomal recessive disorder characterized by progressive myoclonic epilepsy that can begin in adolescence as epileptic seizures. The condition leads to seizures, muscle spasms, difficulty walking, dementia, and ultimately death.

[0210] In yet another aspect of the present invention, the engineered / novel nucleic acid-guided nuclease system can be used to correct genetic eye diseases resulting from several gene mutations.

[0211] In certain embodiments disclosed herein, the engineered nucleic acid-guided nuclease constructs can recognize protospacer adjacent motif (PAM) sequences other than, or in addition to, TTTN. In other embodiments, the engineered nucleic acid-guided nuclease constructs disclosed herein can be further mutated to improve targeting efficiency or selected from a library for specific targeted characteristics. Other embodiments disclosed herein relate to vectors containing the constructs disclosed herein, useful for further analysis and selection of improved genome editing characteristics.

[0212] Other embodiments disclosed herein include kits for packaging and shipping nucleic acid-guided nuclease constructs and / or novel gRNAs disclosed herein, or known gRNAs disclosed herein, and further comprising at least one container. In certain embodiments, some reagents required for the kit can be included for convenience, ease of shipping, and efficiency. [Example]

[0213] Example 1: Culture of Jurkat human T-cell leukemia cell line and primary human T cells Human Jurkat T-cell leukemia cells (Leibniz Institute DSMZ-German Collection of Microorganisms and Cell Cultures GmbH (ACC282)) were grown in RPMI 1640 medium (ThermoFisher Scientific) containing 10% heat-inactivated fetal bovine serum (FBS) (ThermoFisher Scientific) supplemented with 1% penicillin-streptomycin antibiotic mixture (ThermoFisher Scientific). Cells were cultured at 37°C in a 5% CO2 incubator at a density of 0.5–1.5 x 10. 6 cells mL -1 Twenty-four hours before transfection, cells were diluted to 0.1x10 6 cells mL -1Cell culture medium supernatants were periodically tested for mycoplasma contamination using the MycoAlert PLUS Mycoplasma Detection Kit (Lonza).

[0214] Example 2: Isolation and culture of primary T cells T cells were isolated from human peripheral blood obtained from healthy adults by immunomagnetic negative selection using the EasySep Human T Cell Isolation Kit (STEMCELL Technologies). After isolation, T cells were collected at a concentration of 12.5 ng mL -1 Human recombinant IL-2, 5ng mL -1 IL-7, and 5ng mL -1 25 μL mL of ImmunoCult-XF T Cell Expansion Medium (STEMCELL Technologies) containing IL-15 (STEMCELL Technologies) -1 Activated with ImmunoCult Human CD3 / CD28 / CD2 T-Cell Activator (STEMCELL Technologies), 1.0x10 6 cells mL -1 The cells were cultured at 37°C in a 5% CO2 incubator until transfection 48 hours later.

[0215] Example 3: RNP formulation Ribonucleoprotein complexes (RNPs) were generated immediately before transfection by incubating each guide nucleic acid (gNA) with MAD7 at a molar ratio of 3:2 (gNA:MAD7) for 15 min at room temperature. For Jurkat experiments, unless otherwise stated, RNP complexes were generated by mixing each gNA (150 pmol), MAD7 (100 pmol), and nuclease-free water. For T cell experiments, 15–50 kDa poly-L-glutamic acid (PGA, 100 μg μL) was used. -1 1.6 μL of an aqueous solution of gNA (Alamanda Polymers) was added to the gNA, followed by the addition of MAD7 and nuclease-free water.

[0216] Example 4: Generation of donor templates by PCR amplification Donor templates containing site-specific homology arms, their respective promoters, and the respective genes (GFP or Hu19 scFv-CD8α-CD28-CD3ζ CAR) were amplified from the corresponding pTwist ampicillin high-copy plasmids (Twist Bioscience) using homology arm-specific PCR primers. The donor templates were amplified using a two-step PCR program: initial denaturation at 98°C for 30 seconds, followed by a denaturation cycle at 98°C for 10 seconds, and extension at 72°C for 30 seconds, with 40 cycles per 1 kb amplicon, followed by a 10-minute hold at 72°C. Each 50 μL PCR reaction contained 10 ng of amplification template (plasmid DNA), 0.5 μM homology arm-specific forward and reverse primers, nuclease-free water (IDT), 3% DMSO, and 1x Phusion High-Fidelity PCR Master Mix (ThermoFisher Scientific) containing HF buffer. The PCR product was purified using a NucleoSpin Gel and PCR Clean-up Kit (Macherey-Nagel) with two 20 μL eluates. The purified HDR template was collected and quantified using a NanoDrop One Microvolume UV-Vis Spectrophotometer (ThermoFisher Scientific). The template was concentrated using an Amicon Ultra 0.5 mL 30K centrifugal filter: 100 μg of DNA was transferred per unit, filled with nuclease-free water to 500 μL, and centrifuged at 10,000 g for 10 minutes to reduce the volume to 50 μL. The DNA was washed twice with nuclease-free water and collected in a new tube by inversion and centrifugation at 10,000 g for 15 seconds. The HDR template was collected, diluted, and the concentration quantified using a Qubit dsDNA HS Assay Kit (ThermoFisher Scientific). For cell testing, 0.5–1 μg μL was used. -1 The HDR template was used.

[0217] Example 5: Transfection of Jurkat cells For transfection, a Lonza 4D Nucleofector with a shuttle unit (V4SC-2960 Nucleocuvette Strips) was used according to the manufacturer's instructions. For transfection, cells were harvested by centrifugation (200 g, RT, 5 min) and transfected at 10 × 10 in SF cell line Nucleofector X kit buffer (Lonza) unless otherwise noted. 6 cells mL -1 The cells were resuspended in 20 μL of PBS. The cell suspension was mixed with RNP and immediately transferred to a nucleocuvette for transfection. After transfection, the cells were immediately resuspended in pre-warmed growth medium and plated into 96-well, flat-bottom, non-cell culture-treated plates (Falcon). The plates were cultured at 37°C in a 5% CO2 incubator at a density of 0.5–1.0 x 10 cells. 6 cells mL -1 After 48 hours, cells were harvested for viability assays and genomic DNA as described below. For insertion of homologous recombination repair templates, HDR templates were added to cells immediately before transfection, and the suspension was transferred to RNP. Transfection parameters, cell harvesting steps, and growth conditions were as described in Example 1. Cells were harvested 48 hours after transfection for viability assessment, 7 days for CAR insertion efficiency, or 7, 14, and 21 days for GFP insertion efficiency.

[0218] Example 6: Transfection of primary T cells 48 hours after isolation, cells were harvested by centrifugation (300 g, RT, 5 min) and plated at 50 × 10 in supplemented P3 Primary Cell Nucleofector Kit buffer (Lonza). 6 cells mL -1The cells were resuspended in 20 μL of medium. The cells were mixed with the HDR template, and the suspension was transferred to RNP immediately before transfection (nucleofection program EH-115). After transfection, 80 μL of pre-warmed cell culture medium without IL-2 was added to the electroporation cuvette. When M3814 (Selleckchem) was used, 80 μL of pre-warmed culture medium without IL-2 and containing M3814 at a final concentration of 2 μM was added to the electroporation cuvette. After 10 minutes of incubation at 37°C, the T cells were incubated with M3814 at a final concentration of 2 μM and 12.5 ng mL of medium. -1 Cells were transferred to 96-well flat-bottom non-cell culture-treated plates (Falcon) containing pre-warmed culture medium pretreated with IL-2. For experiments using M3814, cells were cultured at 0.25 x 10 6 cells mL -1 or 1.3x10 6 cells mL -1 The cells were seeded at a density of 1000 μg / ml and maintained at 37°C in a 5% CO2 incubator. Viability assays were performed 24 hours after transfection, after which the cells were replated in fresh growth medium containing IL-2. CAR insertion efficiency was measured 7, 11, or 13 days after transfection.

[0219] Example 7: Flow Cytometry Flow cytometry evaluation was performed on a CytoFLEXS instrument (Beckmen Coulter) using a 96-well plate format. Measurements of cell viability, PDCD1 expression, GFP expression, and CAR expression were performed on 10,000 or 20,000 single-cell events in Jurkat cells or primary T cells, respectively.

[0220] For cell viability and GFP knock-in assays, approximately 250,000 cells per sample were transferred to a 96-well V-bottom cell culture plate and evaluated after a series of consecutive washing and staining steps. The first step involved centrifuging the cells at 300 g for 5 minutes at room temperature, discarding the supernatant, washing the cells with 150 μL of Dulbecco's PBS / 2% FBS (STEMCELL Technologies) or cell staining buffer (Biolegend), respectively, followed by a second centrifugation and removal of the supernatant. The final step involved cell viability staining using 150 μL of Dulbecco's PBS / 2% FBS containing 7-aminoactinomycin D (7-AAD, 1:1,000; ThermoFisher) or 50 μL of cell staining buffer containing Zombie Violet dye (1:200; Biolegend), respectively. Measurements of cell viability and GFP expression were collected simultaneously for 7-AAD (excitation: yellow-green laser; emission: 561 nm), Zombie Violet (excitation: violet laser; emission: 405 nm), and GFP (excitation: blue laser; emission: 488 nm) as needed.

[0221] For CAR knock-in efficiency detection, approximately 250,000 cells per sample were transferred to a 96-well V-bottom well, washed as described above with cell staining buffer, and resuspended in 50 μL of cell staining buffer containing PE Anti-Myc tag antibody [9E10] (1:50; Abcam) and Zombie Violet Dye (1:200; Biolegend) for 30 minutes. Cells were then washed twice with 150 μL of cell staining buffer and finally resuspended in 100 μL of cell staining buffer for flow cytometry (excitation: yellow-green laser, emission: 561 nm).

[0222] For the detection of PDCD1 knockout efficiency, approximately 250,000 Jurkat cells per sample were transferred to a 96-well V-bottom cell culture plate and evaluated after a series of consecutive washing and staining steps. The first step involved centrifuging the cells at 300 g for 5 min at 4°C and discarding the supernatant. The cells were then stained using 100 μL of cell staining buffer (Biolegend) containing APC / Cyanine7 anti-human CD279 (PD-1) antibody (1:100; Biolegend) and incubated in the dark at 4°C for 30 min. The cells were then centrifuged at 300 g for 5 min at 4°C and the supernatant was discarded. The next step involved repeating centrifugation at 300 g for 5 min at 4°C twice, discarding the supernatant, and washing the cells with 150 μL of ice-cold cell staining buffer (Biolegend). In the final step, cells were resuspended in 100 μL of cell staining buffer for flow cytometry measurements (excitation: red laser, emission: 633 nm).

[0223] Example 8: DNA extraction Forty-eight hours after transfection, cells were harvested by centrifugation (1,000 g, 10 min) in 96-well V-bottom plates (Greiner), washed with PBS (Sigma-Aldrich), and lysed in 20 μL QuickExtract DNA extraction solution (Epicentre, Lucigen). DNA was extracted according to the manufacturer's protocol: 15 min at 65°C, 15 min at 68°C, 10 min at 95°C, cooled to 4°C, and stored at 4°C. Prior to amplicon PCR, genomic DNA was diluted 20-fold with nuclease-free water.

[0224] Example 9: Amplicon Sequencing Extracted genomic DNA was quantified using a NanoDrop (ThermoFisher Scientific). Amplicons were constructed in two PCR steps: in the first PCR, the region of interest (150–400 bp) was amplified from 10–30 ng of genomic DNA using Phusion High-Fidelity PCR Master Mix (ThermoFisher Scientific) with primers containing Illumina forward and reverse adapters containing the appropriate locus-specific complementary sequences. The amplified product was purified using Agencourt AMPure XP beads (Ramcon) at a sample-to-bead ratio of 1:1.8. DNA was eluted from the beads with nuclease-free water, and the size of the purified amplicons was analyzed on a 2% agarose E-gel using an E-gel electrophoresis system (ThermoFisher Scientific). In the second PCR, a unique pair of Illumina-compatible indexes (Nextera XT Index Kit v2) was added to the amplicons using KAPA HiFi HotStart Ready Mix (Roche). The amplified products were purified using Agencourt AMPure XP beads (Ramcon) at a sample-to-bead ratio of 1:1.8. DNA was eluted from the beads with 10 mM Tris-HCl pH 8.5, 0.1% Tween 20. The size of the purified DNA fragments was verified on a 2% agarose gel using an E-gel electrophoresis system (ThermoFisher Scientific), quantified using the Qubit dsDNA HS Assay Kit (ThermoFisher Scientific), and then pooled at equimolar concentrations. The quality of the amplicon library was verified using an analytical High Sensitivity DNA Kit (Agilent) prior to sequencing. The final library was sequenced on an Illumina MiSeq System using the MiSeq Reagent Kit v.2 (300 cycles, 2x250bp, paired-end reads).De-multiplexed FASTQ files were obtained from BaseSpace (Illumina).

[0225] Example 10: NGS Data Analysis An initial quality assessment of the resulting reads was performed using FastQC36. Sequencing data were aligned and analyzed with CRISPResso2 software. For INDEL frequency analysis, the CRISPRessoBatch command and the following parameters were used: -cleavage_offset1--quantification_window_size10----quantification_window_center1--expand_ambiguous_alignments. For ORF disruption analysis, the CRISPRessoBatch command and the following parameters were used: -cleavage_offset1-coding_seq.<EXON_SEQ> --quantification_window_size0--quantification_window_center1--expand_ambiguous_alignments were used. The percentage change of CRISPResso2 software output was analyzed in Excel.

[0226] Example 11: CRISPR-MAD7 platform for human genome editing using the Jurkat T-cell leukemia line MAD7 nuclease containing a His6 tag and either one (MAD7-1NLS) or four (MAD7-4NLS) nuclear localization signals (NLSs) was used (Figure 1). RNPs were generated as described in Example 3. The editing frequency of MAD7 nuclease complexed with one or more guide nucleic acids containing spacer sequences of SEQ ID NOS: 86-384 shown in Table 1 was determined by nucleofection of the RNPs in Jurkat T cells using the Lonza-recommended nucleofection program SE-CL-120 (Example 5), followed by genomic DNA extraction (Example 8), amplification of the edited locus, and targeted next-generation sequencing to identify editing (Example 9), and finally computational analysis of modification frequency using the CRISPResso2 algorithm (Example 10).

[0227] [Table 1-1]

[0228] [Table 1-2]

[0229] [Table 1-3]

[0230] [Table 1-4]

[0231] [Table 1-5]

[0232] [Table 1-6]

[0233] [Table 1-7]

[0234] [Table 1-8]

[0235] [Table 1-9]

[0236] First, we compared the editing frequency of MAD7 containing one or four NLSs complexed with each gNA using gNAs targeting the DNMT1 locus. RNP concentration-dependent modification efficiency was observed, as indicated by the increasing proportion of modified amplicons (Figure 2, left axis; MAD7-1NLS is shown in dark gray, MAD7-4NLS is shown in light gray). Error bars represent one standard deviation of sample 3 (n = 3). In this experiment, Jurkat cells showed improved editing frequency when treated with RNPs containing MAD-4NLS, indicating that NLS optimization can improve editing efficiency. A slight decrease in cell viability was observed with higher concentrations of RNPs containing four NLSs compared to one NLS (Figure 2, right axis). Specifically, Figure 2 shows the editing frequency (n = 3, mean ± SD) at the DNMT1 locus and cell viability of T-cell leukemia cells in response to the amount of MAD7 and MAD7-RNP containing one or four nuclear localization signals (NLSs) (pmol; constant MAD7:gNA ratio of 1:1.5). Dark gray bars and circles represent the average modification frequency and viability using MAD7-1NLS, respectively. Light gray bars and circles represent the average modification frequency and viability using MAD7-4NLS, respectively.

[0237] To optimize editing activity, we tested 93 different transfection conditions and combined 31 nucleofection programs with three buffers in a Lonza Nucleofection 96-well shuttle system (Figures 3-5). Figures 3, 4, and 5 show the editing frequency (bars; x-axis) for each electroporation condition (buffers SE, SF, and SG, respectively) compared to the control (y-axis, upper control). While most buffer program transfection combinations resulted in suboptimal viability (dots; x-axis) and editing frequency, analysis revealed several conditions that supported substantial rates of both cell viability and editing. Two improved conditions observed in this screen, namely SF-CA-137 and SG-CA-138, were then validated and compared with Lonza-recommended nucleofection programs for T-cell leukemia, namely SE-CL-120 and SE-CK-116 (Figure 6). Specifically, Figure 6 shows the editing frequencies at the DNMT1 locus (n = 4; mean ± SD) in T-cell leukemia cell lines achieved using the transfection conditions identified in Figure 2 (100 pmol of MAD7-4NLS) and the Lonza-recommended nucleofection programs SE-CK-116 and SE-CL-120, as well as the two best nucleofection programs observed in this study, SF-CA-137 and SG-CA-138 (Figures 3-5). Dark gray bars represent the average modification frequency using crDNMT1. Light gray bars represent the average modification frequency using crIDTneg (Integrated DNA Technologies, IDT).

[0238] Example 12: Scalable high-level MAD7-RNP editing of immunologically relevant genes in Jurkate T-cell leukemia cell lines Using the Jurkat T-cell leukemia cell line as a model system, we screened for GNAs that exhibit high editing efficiency. This screen included 298 unique gNAs containing one or more spacer sequences (SEQ ID NOS: 86-384) in Table 1 that target immune checkpoint receptors PDCD1, TIM3, LAG3, TIGIT, and CTLA4, checkpoint phosphatases PTPN6 (SHP-1) and PTPN11 (SHP-2), and the TCR signaling subunit CD247 (CD3ζ). RNPs were generated as described in Example 3, nucleofected (Example 5), genomic DNA extracted (Example 8), edited loci amplified and sequenced (Example 9), and sequencing data were computationally analyzed using the CRISPResso2 algorithm (Example 10).

[0239] The CRISPResso2 software reports the frequency of modifications (insertions, deletions, and substitutions) within a quantification window adjacent to the location of the MAD7-induced cleavage in the amplicon sequence. To better understand the detection of editing events, we compared the types of modifications detected in 230 amplicons sequenced in both gNA-treated and mock samples (without MAD7). The relatively high modification frequency (median 1%) in mock reactions was observed as a result of the high frequency of substitutions (Figure 7, light gray bars). Substitutions were detected at a median frequency of 0.96% (presumably due to errors in NGS base calling or substitutions occurring during DNA amplification), whereas insertions and deletions were found at much lower median frequencies of 0.003% and 0.042%, respectively. Specifically, Figure 7 shows the editing frequencies at eight different loci using 298 gNAs (n = 3; mean ± SD) in T-cell leukemia cell lines, along with various editing types (all modifications, insertions only, deletions only, substitutions only, or insertions and deletions (INDELs)). Editing was achieved using the transfection conditions shown in Example 11, Figure 2 (100 pmol MAD7-4NLS) and one of the Lonza nucleofection programs tested (Figure 6; SF-CA-137). Dark gray box plots represent the average modification frequencies using gNAs. Light gray box plots represent the average modification frequencies using crIDTneg (IDT). Therefore, both insertion and deletion frequencies (INDELs) were used as a means of quantifying the editing activity of the CRISPR-MAD7 system to minimize low-end noise. Furthermore, the low INDEL frequency in the MOCK reaction allowed for the sensitive detection of editing events at a significantly higher proportion of sites (Fisher's exact test, P = 3 x 10 -12 (Figure 8). Analysis of gNAs with low INDEL frequencies showed statistically significant editing in gNA-treated samples compared to mock samples with INDEL frequencies as low as 0.5% (Fisher's exact test, P = 4x10 -8(Figure 8). This demonstrates the sensitivity of the assay to detect modifications in the sub-1% range. Specifically, Figure 8 shows the INDEL frequencies at eight different loci, corresponding to two modification types, using 298 gNAs (n=3; mean ± SD) in T-cell leukemia cell lines. All modifications were <1%, and INDELs were <1%, <0.5%, or <0.1%, with <1% INDELs (Fisher's exact test, P=3x10). -12 ) and less than 0.5% (Fisher's exact test, P = 4x10 -8 The INDEL frequency is lower in the MOCK compared to the gNA reaction. The dark grey box plot represents the average INDEL frequency using gNA. The light grey box plot represents the INDEL frequency using crIDTneg (IDT).

[0240] Because MAD7 can target a wide range of PAMs, we analyzed the editing specificity of MAD7 in Jurkat cells by screening gNAs flanking all YTTN PAM variants. MAD7 showed editing with all eight combinations of YTTN PAMs, but in this experiment, editing was only observed with the YTTV and TTV consensus sequences (Fisher's exact test; P = 2x10). -3 and P = 2x10 -4) were higher. The majority of highly active (>50% INDEL frequency) gNAs were found at sites with YTTV and TTV PAMs, whereas moderately active (>10% INDEL frequency) gNAs were found to target all PAM sequences except CTTT. This indicates that MAD7 can edit a wide range of target PAMs, even at low frequencies (Figure 9). Specifically, Figure 9 shows the unique INDEL frequencies at eight different loci using 298 gNAs (n = 3, mean ± SD) in T-cell leukemia cell lines, corresponding to eight YTTN PAM combinations and TTTV, YTTN, and YTTV PAM motifs. The gray zone of the plot represents moderately active gNAs (10-50% INDELs), the upper zone represents highly active gNAs (>50% INDELs), and the lower zone represents active gNAs (1-10% INDELs). The INDEL frequencies in the YTTV and TTTV PAM motifs are significantly higher than those in the YTTN motif (Fisher's exact test, P = 2x10 -3 , and P=2x10 -4 ).

[0241] Given the analysis of a large number of gNAs, we determined biases in editing efficiency of target DNA sequences. Sequence logos were compared to DNA-complementary gNA sequences for inactive (less than 1% INDELs), active (1-10% INDELs), moderately active (10-50% INDELs), and highly active (more than 50% INDELs) gNAs (Figure 10A). This experiment did not identify a strong bias for ribonucleotides at specific positions, but guanine was overrepresented and uracil was underrepresented in moderately and highly active gNAs. Next, the frequency of ribonucleotide bases was analyzed within the same four classes of gNAs (Figure 10B). The analysis confirmed a significant increase in guanine and a decrease in uracil in highly active gNAs. Specifically, Figure 10 shows sequence logos comparing DNA-complementary gNA sequences; (A) High-activity (>50% INDELs), moderate-activity (10-50% INDELs), active (1-10% INDELs), and inactive (<1% INDELs) gNAs do not show a strong bias for ribonucleotides at specific positions, but high-activity and moderate-activity gNAs appear to overrepresent guanine and underrepresent uracil; (B) Nucleotide frequencies of inactive (<1% INDELs; dark gray boxes), active (1-10% INDELs; medium-dark gray boxes), moderate-activity (10-50% INDELs; light gray boxes), and high-activity (>50% INDELs; white boxes) gNAs. Compared to inactive gNAs, high-activity gNAs have a significantly increased guanine and decreased uracil (Fisher's exact test, P = 4x10, respectively). -3 and P = 3 x 10 -4 ) Furthermore, a significant increase in the guanine-cytosine content and a decrease in the adenine-uracil content were observed in moderately active gNAs compared with inactive gNAs (Fisher's exact test, P = 1 × 10 -2). Furthermore, the data showed that nearly 40% of inactive gNAs carried three or more adenine or uracil ribonucleotides, while none of the highly active and less than 20% of moderately active gNAs contained such runs (Figure 11). These sequence features can act as an algorithm for selecting putatively highly active gNAs during the first round of screening, potentially reducing the overall cost of identifying gNAs for various genes of interest. Specifically, Figure 11 shows the proportion of gNAs in AAA and / or UUU runs, corresponding to gNAs with highly active (>50% INDELs), moderately active (10-50% INDELs), active (1-10% INDELs), and inactive (<1% INDELs) INDEL frequencies. The proportions of inactive (<1% INDELs) and active (1-10% INDELs) gNAs containing such runs are higher compared to highly active (>50% INDELs) gNAs (Fisher's exact test, P = 1x10). -3 , and P=4x10 -4 ).

[0242] Example 13: Validation of gNAs for gene editing and disruption of immune-related genes using T-cell leukemia lines The highly efficient gNAs identified in our initial analysis were validated by assaying the INDEL frequency for the top three or five gNAs for each of the selected immunologically relevant genes (Figure 12). Specifically, Figure 12 shows the frequencies of INDELs (dark gray bars) and frameshifts (light gray bars) (n = 3; mean ± SD) for 38 highly efficient gNAs in T-cell leukemia cell lines. Alternating gray and white zones on the plot represent groups of 3-5 highly efficient gNAs per location. In validation experiments, INDEL frequency significantly correlated with measurements from the initial screening, highlighting the reproducibility of the INDEL assay (Figure 13). Specifically, Figure 13 shows the correlation between INDEL frequency in the gNA validation experiment and INDEL formation in the gNA screening experiment (Spearman correlation = 0.91; P = 9 × 10). -14), highlighting the reproducibility of the INDEL assay. Using CRISPresso2 software, we estimated the degree of open reading frame (ORF) disruption for each of the validated gNAs (Figure 12). Furthermore, for four highly efficient gNAs targeting three different exons in the PDCD1 locus, surface expression of PDCD1 protein was measured by flow cytometry at days 4, 7, and 11 posttransfection (data not shown). These data revealed that protein surface expression after transfection with crPDCD1_2, a gNA targeting the PDCD1 gene at the extracellular domain of the protein, was low at 10% at day 4 posttransfection and remained at this level even at day 11 posttransfection. Surface expression after transfection with the remaining three gNAs was significantly higher, at 35% and 85%, after transfection with both crPDCD1_3, crPDCD1_4, and crPDCD1_5, respectively. This is consistent with the ORF data analysis, showing that for most of the gNAs, including the highly efficient crPDCD1, the number of predicted INDELs resulting in frameshifts was comparable to that expected from an unbiased DNA repair process, resulting in frameshifts at two-thirds of the edited loci (Figure 14). However, some gNAs exhibited significantly different degrees of ORF disruption, with crCD247_4 resulting in frameshifts at a frequency of 97%, while crTIM3_1 and crTIM3_3 resulted in frameshifts at frequencies of 23% and 44%, respectively (Figure 14). Specifically, Figure 14 shows the ratio of frameshift to INDEL frequency (dark gray bars) in T-cell leukemia cell lines, corresponding to 38 highly efficient gNAs. The average proportion of INDELs resulting in frameshifts (dashed line) is approximately 66%. Alternating gray and white zones on the plot represent groups of 3–5 highly efficient gNAs per site. Analysis of repair products revealed that in crTIM3_1 and to some extent crTIM3_3, bias arises from sequences directly repeated at the DNA break site, which may have promoted microhomology-mediated end-joining (MMEJ) repair after DNA breakage.These data are useful for the informed selection of gNAs for gene KO, as some gNAs, such as crTIM3_1, have a much lower frequency of gene disruption than would be predicted based on the frequency of INDEL formation.

[0243] Another consideration for selecting gNAs is the possibility of off-target cleavage events. The validated list of gNAs was analyzed using CasOFFinder software to predict potential off-target editing sites within the genome with up to four mismatches between the gNA and the target DNA sequence. Using the Bioconductor R package, the predicted off-target sites were matched to a human gene database to extract sites targeting exons and introns within genes. The extent of editing activity at these sites was then examined by targeted next-generation sequencing, more specifically, at 25 predicted off-target sites for the top two PDCD1 gNAs, namely, crPDCD1_1 and crPDCD1_2. Analysis revealed low levels of off-target activity at the crPDCD1_2_13 and crPDCD1_2_15 sites; however, INDEL formation at these two sites was not statistically significant compared to the mock sample (non-targeting gNA) (pairwise T-test, P≥0.05; Figures 15 and 16). For the top two gNAs targeting the remaining seven genes (i.e., TIM3, LAG3, TIGIT, CTLA4, PTPN6, PTPN11, and CD247; spacer sequences in Table 1), we assayed the INDEL frequency at 43 putative off-target sites with up to three mismatches between the gNA and the target DNA sequence. This analysis revealed no detectable activity at any of the putative off-target sites (Figures 15 and 16), thereby confirming the high cleavage fidelity of the MAD7-gNA complex. Specifically, Figures 15-16 show the INDEL frequency (n = 3; mean ± SD) of MAD7 in T-cell leukemia cell lines at predicted off-target sites analyzed by targeted deep sequencing. For crPDCD1, we analyzed the INDEL frequency at putative off-target editing sites with four or fewer mismatches between the gNA and the target DNA sequence and three or fewer mismatches on the remaining gNAs. PAM and spacer sequences with mismatches marked in red are displayed next to their respective measured INDEL frequencies.No significant INDEL frequency was detected at any of the off-target sites (pairwise T-test, P≥0.05).

[0244] Insertion of exogenous transgenes is an important aspect of mammalian cell engineering. CRISPR-Cas-mediated gene insertion is achieved by homologous recombination repair of CRISPR-induced DNA breaks using an HDR donor template to copy the exogenous gene sequence into the target DNA locus. Several studies have shown that HDR templates composed of linear double-stranded DNA provide the most robust and efficient method for transgene insertion using the CRISPR-Cas genome editing system.

[0245] The Jurkat T-cell leukemia cell line was used to evaluate the efficiency of transgene insertion and expression using the CRISPR-MAD7 RNP complex. 0.5 μL of highly active gNA targeting the AAVS1 (spacer sequence in Table 1) safe harbor locus (Figure 17) was added. -1The AAVS1 gene was used in combination with eight different HDR repair templates flanked by 500-base-pair symmetric homology arms (HA). Specifically, Figure 17 shows the INDEL frequency (n=3; mean ± SD) at the AAVS1 locus in T-cell leukemia cell lines, correlated with MAD7-RNP amount (pmol; constant MAD7:gNA ratio of 1:1.5). Dark gray bars represent the average INDEL frequency using crAAVS1. Light gray bars represent the average modification frequency using crIDTneg (IDT). The HDR inserts included eight promoters (Table 2) that varied in both size and promoter strength driving GFP expression (Figure 18). When transient GFP expression decreased 14 days after transfection, comparable insertion efficiencies were observed with four of the eight promoters (JET, PGK, EF1a, and CAG) with stable GFP expression of up to 30% (Figure 18), suggesting that insert size does not affect integration efficiency in the human T-cell leukemia cell line AAVS1. Specifically, Figure 18 shows the GFP insertion efficiency in AAVS1 (n=3, mean ± SD) and cell viability of the T-cell leukemia cell line measured 14 days after transfection. Consisting of eight different promoters, 0.5 μL -1 HDR templates flanked by 500 base pair symmetric homology arms were used. Promoter sizes in base pairs: CMV, 1400; SCP, 970; CMVe-SCP, 1270; CMVmax, 1830; JET, 1100; CAG, 2600; PGK, 1410; EF-1α, 2090. Dark gray bars and circles represent the average insertion frequency and cell viability using crAAVS1. Light gray bars represent the average insertion frequency and cell viability using crIDTneg (IDT).

[0246] [Table 2-1]

[0247] [Table 2-2]

[0248] [Table 2-3]

[0249] Subsequently, while keeping the amount of MAD7-RNP constant, various homology arm lengths (100 μg for 500 bp) and HDR template amounts (0.125 μL) were investigated for insertion efficiency. -1 , 0.25 μL -1 , 0.5 μL -1 , and 1 μL -1 The effect of the GFP insertion site was evaluated using the JET and EF1a promoters. A higher integration efficiency of up to 30% was observed with HDR templates flanked by 500 HA residues compared with 100 base pairs. Furthermore, the data showed improved insertion efficiency with increasing amounts of HDR templates flanked by either 100 or 500 HA residues, but at the same time, cell viability slightly decreased (Figure 19). Specifically, Figure 19 shows the GFP insertion efficiency in the T-cell leukemia cell line AAVS1 (n = 3; mean ± SD) measured at days 2, 7, 14, and 21 post-transfection, relative to the amount of donor template. No transient GFP expression was observed at day 21 post-transfection. Cell viability (filled circles) was measured at day 2 post-transfection. The top panel shows the GFP insertion efficiency using a donor template flanked by short homology arms (100 bp HA), while the bottom panel shows that using a donor template flanked by long homology arms (500 bp HA). The left panel shows the GFP insertion efficiency using a donor template (long, approximately 2000 bp) containing the EF-1α promoter, and the right panel's donor template (short, approximately 1000 bp) containing the JET promoter. The amounts of donor template, represented by the upper slope of the bar graph, are 0.125, 0.25, 0.5–1 μg μL. -1 Dark grey bars represent the average insertion frequency using crAAVS1. Light grey bars represent the average insertion frequency using crIDTneg(IDT).

[0250] Next, primary T cells isolated from human peripheral blood of three donors were incubated with 100 μL of poly-L-glutamic acid (PGA) in combination with the protocol selected from the experiments described above, i.e., 150:100 pmol gNA:MAD7RNP complex and 1 μg μL -1 HDR template, 100 μg μL -1 We analyzed the integration efficiency of clinically relevant CAR transgenes containing the JET or EF1a promoter and a bovine growth hormone-derived polyadenylation sequence flanked by 100 or 500 base pairs of HA using a combination of poly-L-glutamic acid (PGA). We used an anti-CD19 CAR with a fully human variable region (Hu19CAR), CD8α hinge and transmembrane domain, CD28 costimulatory domain, and CD3ζ activation domain. Using HDR templates flanked by 100 and 500 base pairs of HA, we observed moderate insertion efficiency with AAVS1, but stable CAR expression of up to 14% and 16%, respectively. Normalized cell viability measured 24 hours posttransfection was relatively low: 22% for JET-500-CAR, 35% for JET-100-CAR, 43% for EF1a-100-CAR, and 55% or less for EF1a-500-CAR (Figure 20). It is important to emphasize that treatment with PGA resulted in both higher CAR insertion efficiency and cell viability compared to treatment without PGA (P≦0.05; data not shown). Specifically, Figure 20 shows the CAR insertion efficiency in AAVS1 primary pan T cells measured on days 7 and 11 after transfection (D=3; n=3; mean±SD). Cell viability was measured 24 hours after transfection. Individual panels show the CAR insertion efficiency using the donor template structure described in Figure 19. The amounts of donor template MAD7-RNP and PGA were 1 μg / μL, in that order. -1 , 100:150 pmol MAD7:gNA, and 100 μL -1For transfection of primary T cells, the nucleofection program P3-EH-115 was used. D represents the number of biological replicas per D and the number of technical replicas per D. Dark gray bars represent the average insertion frequency using crAAVS1. Light gray bars represent the average insertion frequency using crIDTneg(IDT).

[0251] Multiple parameters were reevaluated to further optimize primary T cell viability and CAR insertion efficiency in AAVS1. Pan T cells isolated from the blood of two donors were used to generate 100 μL of T cells. -1 The effect of RNP amount with PGA and EF1a-500-CAR template amount on CAR insertion efficiency and cell viability was examined (data not shown). The RNP amount was reduced to 75:50 pmol of gNA:MAD7 RNP complex, while the donor template amount was increased to 1.5 μL. -1 Increasing the concentration of CRISPR-MAD7 to 1 μL resulted in improved CAR insertion efficiency without significantly affecting cell viability (P≧0.05; data not shown). Furthermore, using the above transfection conditions in combination with cell recovery in post-transfection culture medium pretreated with 2 μM M3814 improved CAR insertion efficiency by nearly 5-fold compared to other experiments (Figure 21). The optimized CRISPR-MAD7 transfection protocol resulted in CAR insertion efficiencies of up to 85% (median 65%) at 13 days post-transfection and a high normalized cell viability of 62% at 24 hours post-transfection. Specifically, Figure 21 shows CAR insertion efficiency in primary pan T cells (AAVS1) measured at 7 days post-transfection and re-measured in two biological replicates at 13 days post-transfection (D=2; n=3). Cell viability was measured 24 hours post-transfection (D=5; n=3; mean ± SD). The amount or concentration of the donor template, MAD7-RNP, PGA, and M3814 was 1.5 μg / μL, respectively. -1 , 50:75pmol MAD7:gNA, 100μg / μL -1, and 2 μM. For transfection of primary T cells, the nucleofection program P3-EH-115 was used. D represents the number of biological replicas per D and the number of technical replicas per D. Dark gray bars represent the average insertion frequency using crAAVS1. Light gray bars represent the average insertion frequency using crIDTneg(IDT).

[0252] equivalent Throughout this specification, when compositions are described as having, including, or comprising certain components, or when processes and methods are described as having, including, or comprising certain steps, it is contemplated that there are also compositions of the invention that consist essentially of, or consist of, the recited components, and that there are processes and methods of the invention that consist essentially of, or consist of, the recited process steps.

[0253] In this application, when an element or component is said to be included in and / or selected from a list of enumerated elements or components, it is understood that the element or component may be any one of the listed elements or components, or the element or component may be selected from a group consisting of two or more of the listed elements or components.

[0254] Furthermore, it should be understood that elements and / or features of the compositions or methods described herein, whether expressly or implicitly stated herein, may be combined in various ways without departing from the spirit and scope of the present invention. For example, when reference is made to a particular compound, that compound may be used in various embodiments of the compositions of the present invention and / or methods of the present invention, unless the context otherwise indicates. In other words, although embodiments are described and illustrated herein so that a clear and concise application may be described and depicted, it is intended and will be understood that the embodiments can be combined or separated in various ways without departing from the teachings and invention(s) herein. For example, it should be understood that all features described and illustrated herein may be applicable to all aspects of the invention(s) described and illustrated herein.

[0255] The terms "a," "an," and "the," and similar references in the context of describing the present invention (particularly in the context of the claims below) should be construed to include both the singular and the plural, unless otherwise indicated herein or clearly contradicted by context. For example, the term "a cell" includes a plurality of cells, including mixtures thereof. When the plural is used for compounds, salts, etc., this is construed to mean a single compound, salt, etc.

[0256] The phrase "at least one of" should be understood to include each of the listed items following the phrase individually, and various combinations of two or more of the listed items, unless otherwise understood from context and usage. The phrase "and / or" in the context of more than two listed items should be understood to have the same meaning, unless otherwise understood from context.

[0257] Use of the terms "include," "includes," "including," "have," "has," "having," "contain," "contains," or "containing," including grammatical equivalents thereof, should be understood to be generally open-ended and non-limiting, e.g., not excluding additional, unrecited elements or steps, unless otherwise specifically stated or otherwise understood by context.

[0258] When the term "about" is used before a quantitative value, the present invention also includes the specific quantitative value itself, unless otherwise specified. As used herein, the term "about" refers to a ±10% variation from the nominal value, unless otherwise indicated or estimated.

[0259] It should be understood that the order of steps or order for performing certain actions is immaterial so long as the invention remains operable. Moreover, two or more steps or actions may be conducted simultaneously.

[0260] The use of any and all examples or exemplary language herein, such as "such as" or "including," is intended merely to better illustrate the invention and does not impose limitations on the scope of the invention unless otherwise claimed. No language in the specification should be construed as indicating any non-claimed element as essential to the practice of the invention.

[0261] The present invention may be embodied in other specific forms without departing from its spirit or essential characteristics. Accordingly, the foregoing embodiments are to be considered in all respects as illustrative and not limiting of the invention described herein. The scope of the invention is, therefore, indicated by the appended claims rather than by the foregoing description, and all changes that come within the meaning and range of equivalency of the claims are intended to be embraced therein.

[0262] Embodiment In embodiment 1, provided herein is a composition comprising a nucleic acid-guided nuclease comprising a type V CRISPR nuclease polypeptide comprising at least one nuclear localization signal (NLS) at or near the N-terminus or C-terminus of the polypeptide. In embodiment 2, provided herein is a composition described in embodiment 1, wherein the nuclease is a type Va nuclease. In embodiment 3, provided herein is a composition described in embodiment 1 or embodiment 2, wherein the type V CRISPR nuclease polypeptide has at least 60%, 70%, 80, 85, 90, 95, 96, 97%, 98, 99, or 100% sequence identity, preferably at least 80%, more preferably at least 90%, even more preferably at least 95%, and even more preferably at least 98% sequence identity to SEQ ID NO: 1. In embodiment 4, provided herein is a composition described in any of the preceding embodiments, wherein the type V CRISPR nuclease polypeptide comprises two NLSs, one or both of which are at or near the N-terminus or C-terminus of the polypeptide. In embodiment 5, provided herein is a composition according to any of the preceding embodiments, wherein the type V CRISPR nuclease polypeptide comprises three NLSs, each of which is located at or near the N-terminus or C-terminus of the polypeptide. In embodiment 6, provided herein is a composition according to any of the preceding embodiments, wherein the type V CRISPR nuclease polypeptide comprises four NLSs, each of which is located at or near the N-terminus or C-terminus of the polypeptide. In embodiment 7, provided herein is a composition according to any of the preceding embodiments, wherein the type V CRISPR nuclease polypeptide comprises at least five NLSs, each of which is located at or near the N-terminus or C-terminus of the polypeptide. In embodiment 8, provided herein is a composition according to any one of embodiments 4-7, wherein at least two of the NLSs are located at or near the N-terminus of the polypeptide. In embodiment 9, provided herein is a composition according to any one of embodiments 5-7, wherein at least three of the NLSs are located at or near the N-terminus of the polypeptide.In embodiment 10, provided herein is a composition of embodiment 6 or 7, wherein at least four of the NLSs are at or near the N-terminus of the polypeptide. In embodiment 11, provided herein is a composition of embodiment 7, wherein five NLSs are at or near the N-terminus of the polypeptide. In embodiment 12, provided herein is a composition of embodiment 11, comprising a sequence at least 60%, 70%, 80%, 85%, 90%, 95%, 98%, 99% or 100% identical, preferably at least 80%, more preferably at least 90%, even more preferably at least 95%, even more preferably at least 98% identical to any one of SEQ ID NOs: 109-112. In embodiment 13, provided herein is a composition of any one of embodiments 1-3, wherein the Type V CRISPR nuclease polypeptide comprises at least 1-30, 1-20, 1-15, 1-10, 1-9, 1-8, 1-7, 1-6, 1-5, 2-30, 2-20, 2-15, 2-10, 2-9, 2-8, 2-7, 2-6, 2-5, 3-30, 3-20, 3-15, 3-10, 3-9, 3-8, 3-7, 3-6, or 3-5, preferably 1-10, more preferably 2-10, even more preferably 3-10, NLSs, each of which is at or near the N-terminus or C-terminus of the polypeptide. In embodiment 14, provided herein is a composition of any one of embodiments 4-11, wherein at least two of the NLSs have different nuclear localization mechanisms. In embodiment 15, provided herein is a composition according to any one of embodiments 5-7, or 9-11, wherein at least three of the NLSs have different nuclear localization mechanisms.In embodiment 16, provided herein is a composition according to any of the preceding embodiments, wherein one or more of the NLSs comprise an NLS from SV40 large T antigen, an NLS from nucleoplasmin, e.g., the nucleoplasmin bipartite NLS, a c-myc NLS, an hRNPA1 M9 NLS, an IBB domain of importin alpha NLS, a sarcoma T protein NLS, a sequence derived from human p53 NLS, a sequence derived from mouse c-abl IV NLS, an influenza virus NS1 NLS, a hepatitis virus delta antigen NLS, a mouse Mx1 protein NLS, a human poly(ADP-ribose) polymerase NLS, a steroid hormone receptor (human) glucocorticoid NLS, and / or an EGL-13 NLS. In embodiment 17, provided herein is a composition according to embodiment 16, wherein one or more of the NLSs comprise an NLS from SV40 large T antigen. In embodiment 18, provided herein is a composition according to embodiment 16, wherein two or more of the NLSs comprise an NLS from the large T antigen of the SV40 virus. In embodiment 19, provided herein is a composition according to embodiment 17 or 18, wherein the NLS(s) comprise the sequence of SEQ ID NO:5. In embodiment 20, provided herein is a composition according to any one of embodiments 16 to 19, wherein one or more of the NLSs comprise an NLS from nucleoplasmin. In embodiment 21, provided herein is a composition according to embodiment 20, wherein the nucleoplasmin NLS comprises the sequence of SEQ ID NO:6. In embodiment 22, provided herein is a composition according to any one of embodiments 16 to 21, wherein one or more of the NLSs comprise a c-myc NLS. In embodiment 23, provided herein is a composition according to embodiment 22, wherein the c-myc NLS comprises the sequence of SEQ ID NO:7, SEQ ID NO:8, or SEQ ID NO:21. In embodiment 24, provided herein is a composition according to embodiment 23, wherein the c-myc NLS comprises the sequence of SEQ ID NO:21. In embodiment 25, provided herein is a composition according to any one of embodiments 16 to 24, wherein one or more of the NLSs comprises an EGL-13 NLS. In embodiment 26, provided herein is a composition according to embodiment 25, wherein the EGL-13 NLS comprises the sequence of SEQ ID NO: 107.In embodiment 27, provided herein is a composition according to any of the preceding embodiments, wherein the Type V CRISPR nuclease polypeptide further comprises a purification tag. In embodiment 28, provided herein is a composition according to embodiment 27, wherein the purification tag is located at or near the N-terminus of the nuclease polypeptide. In embodiment 29, provided herein is a composition according to embodiment 27 or embodiment 28, wherein the purification tag comprises a poly-His tag, e.g., a Gly-6xHis tag or a Gly-8xHis tag; a short epitope tag, e.g., FLAG, hemagglutinin (HA), c-myc, T7, Glu-Glu; maltose-binding protein (mbp); N-terminal glutathione S-transferase (GST); or calmodulin-binding peptide (CBP). In embodiment 30, provided herein is a composition according to embodiment 29, wherein the purification tag comprises a poly-His tag. In embodiment 31, provided herein is a composition according to embodiment 30, wherein the purification tag comprises a gly-6xHis tag. In embodiment 32, provided herein is a composition according to embodiment 30, wherein the purification tag comprises a gly-8x His tag. In embodiment 33, provided herein is a composition according to any of the preceding embodiments, wherein the Type V CRISPR nuclease polypeptide comprises a cleavage site. In embodiment 34, provided herein is a composition according to embodiment 33, wherein the cleavage site is at or near the N-terminus of the nuclease polypeptide. In embodiment 35, provided herein is a composition according to embodiment 33 or embodiment 34, wherein the cleavage site comprises a tobacco etch virus (TEV) cleavage site. In embodiment 36, provided herein is a composition according to embodiment 35, wherein the cleavage site comprises the sequence of SEQ ID NO: 108. In embodiment 37, provided herein is a composition according to embodiment 36, wherein the cleavage site comprises five NLSs, a purification tag, and a cleavage site at or near the N-terminus of the polypeptide, wherein the cleavage site is after the purification tag. In embodiment 38, provided herein is a composition according to embodiment 37, comprising a sequence at least 60%, 70%, 80%, 85%, 90%, 95%, 98%, 99% or 100% identical, preferably at least 8%, more preferably at least 90%, even more preferably at least 95%, even more preferably at least 98% identical to SEQ ID NO: 111 or 112. In embodiment 39, provided herein is a composition according to embodiment 37, comprising a sequence at least 60%, 70%, 80%, 85%, 90%, 95%, 98%, 99% or 100% identical, preferably at least 8%, more preferably at least 90%, even more preferably at least 95%, even more preferably at least 98% identical to SEQ ID NO: 112. In embodiment 40, provided herein is a composition according to any of the preceding embodiments, further comprising a guide nucleic acid (gNA), e.g., gRNA, comprising a spacer sequence that targets a target nucleotide sequence within a polynucleotide, or a polynucleotide encoding the gNA, e.g., gRNA, wherein the gNA, e.g., gRNA, is compatible with a Type V CRISPR nuclease. In embodiment 41, provided is a composition according to embodiment 40, wherein the target nucleotide is within 50 nucleotides of a protospacer adjacent motif (PAM) sequence specific for a Type V CRISPR nuclease.In embodiment 42, provided herein is a composition according to embodiment 41, wherein the PAM comprises the sequence YTTN, where Y is T or C, and N is A, T, G, or C. In embodiment 43, provided herein is a composition according to embodiment 42, wherein the PAM comprises the sequence YTTV or TTTV, where V is A, G, or C. In embodiment 44, provided herein is a composition according to embodiment 40, wherein the gNA is a gRNA. In embodiment 45, provided herein is a composition according to embodiment 44, wherein the gRNA is a double-stranded gRNA. In embodiment 46, provided herein is a composition according to embodiment 44 or embodiment 45, wherein the composition comprises a gRNA, and the gRNA comprises one or more chemical modifications. In embodiment 47, provided herein is a composition of embodiment 46, wherein the chemical modification comprises, for example, 2'-O-alkyl, 2'-O-methyl, phosphorothioate, phosphonoacetate, thiophosphonoacetate, 2'-O-methyl-3'-phosphorothioate, 2'-O-methyl-3'-phosphonoacetate, 2'-O-methyl-3'-thiophosphonoacetate, 2'-deoxy-3'-phosphonoacetate, 2'-deoxy-3'-thiophosphonoacetate, suitable alternatives, or a combination thereof. In embodiment 48, provided herein is a composition according to any one of embodiments 44 to 47, wherein the ratio of guanine:uracil in the gRNA is at least 51:49, 52:48, 53:47, 54:46, 55:45, 56:44, 57:43, 58:42, 59:42, or 60:40, preferably at least 53:47, more preferably at least 54:46, and even more preferably at least 55:45.In embodiment 49, provided herein is a composition according to any one of embodiments 40 to 48, wherein the molar ratio of gRNA, e.g., gRNA to Type V CRISPR nuclease is at least 1.1:1, 1.2:1, 1.3:1, 1.4:1, 1.5:1, 1.6:1, 1.7:1, 1.8:1, 2:1, 2.2:1, 2.5:1, or 3:1 and / or 1.2:1, 1.3:1, 1.4:1, 1.5:1, 1.6:1, 1.7:1, 1.8:1, 2:1, 2.2:1, 2.5:1, 3:1, or 4:1 or less, preferably between 1.1:1 and 2.5:1, more preferably between 1.2:1 and 2:1, even more preferably between 1.2:1 and 1.7:1. In embodiment 50, provided herein is a composition according to any one of embodiments 40 to 49, wherein the molar amount of gNA, e.g., gRNA, is at least 10, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 170, 190, or 200 pmol and / or no more than 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 170, 190, 200, 250, or 300 pmol, preferably between 25 and 200 pmol, more preferably between 50 and 100 pmol, and even more preferably between 65 and 85 pmol. In embodiment 51, provided herein is a composition according to any one of embodiments 40 to 50, further comprising a donor template. In embodiment 52, provided herein is a composition according to embodiment 51, wherein the donor template comprises homology arms.In embodiment 53, provided herein is a composition according to embodiment 51 or embodiment 52, wherein the donor template is present in an amount of at least 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 1.1, 1.2, 1.3, 1.4, 1.5, 1.7, 2, 2.5, 3, 4, or 5 μg μL and / or no more than 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 1.1, 1.2, 1.3, 1.4, 1.5, 1.7, 2, 2.5, 3, 4, 5, 7, or 10 μg μL, preferably between 0.3 and 2 μg μL, more preferably between 0.5 and 1.5 μg μL, and even more preferably between 0.8 and 1.2 μg μL. In embodiment 54, there is provided a composition according to any one of embodiments 40 to 53, further comprising an anionic polymer. In embodiment 55, there is provided herein a composition according to embodiment 54, wherein the anionic polymer comprises polyglutamic acid (PGA). In embodiment 56, provided herein is a composition according to embodiment 54 or embodiment 55, wherein the anionic polymer is present in a concentration of at least 20, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 170, 200, 250, 300, 400, or 500 μg μL and / or no more than 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 170, 200, 250, 300, 400, 500, 700, or 1000 μg μL, preferably between 20 and 200 μg μL, more preferably between 50 and 150 μg μL, and even more preferably between 80 and 120 μg μL.

[0263] In embodiment 57, provided herein is a cell comprising the composition of any of the preceding embodiments. In embodiment 58, provided herein is a cell of embodiment 56, wherein the cell is a human cell. In embodiment 59, provided herein is a cell of embodiment 58, wherein the cell is an immune cell or a stem cell. In embodiment 60, provided herein is a cell of embodiment 59, wherein the cell is an immune cell. In embodiment 61, provided herein is a cell of embodiment 60, wherein the cell is a T cell. In embodiment 62, provided herein is a cell of embodiment 59, wherein the cell is a stem cell. In embodiment 63, provided herein is a cell of embodiment 62, wherein the cell is an induced pluripotent stem cell (iPSC).

[0264] In embodiment 64, provided herein is a method comprising inserting into a cell the composition of any one of embodiments 1 to 56. In embodiment 65, provided is the method of embodiment 64, wherein inserting the composition into a cell comprises electroporation.

[0265] In embodiment 66, provided herein is a method for modifying a target polynucleotide, comprising (i) contacting with the composition of any one of embodiments 40 to 56, and (ii) allowing the nuclease and guide nucleic acid to modify the target genomic region. In embodiment 67, provided herein is the method of embodiment 66, wherein the composition is the composition of any one of embodiments 51 to 56. In embodiment 68, provided herein is the method of embodiment 66 or 67, wherein the target polynucleotide is a genome or part of a genome in a cell. In embodiment 69, provided herein is the method of embodiment 68, wherein the cell is a human cell. In embodiment 70, provided herein is the method of embodiment 69, wherein the cell is an immune cell or stem cell. In embodiment 71, provided herein is the method of embodiment 70, wherein the cell is an immune cell. In embodiment 72, provided herein is the method of embodiment 71, wherein the cell is a T cell. In embodiment 73, provided herein is the method of embodiment 70, wherein the cell is a stem cell. In embodiment 74, provided herein is a method according to embodiment 73, wherein the stem cell is an iPSC. In embodiment 75, provided herein is a method according to any one of embodiments 67 to 74, wherein the donor template comprises a mutation in a PAM within 50 nucleotides of the target nucleotide sequence in the target polynucleotide. In embodiment 76, provided herein is a method according to any one of embodiments 68 to 74, wherein the composition is the composition according to embodiment 67, and the donor template comprises a polynucleotide encoding a polypeptide expressed by the cell. In embodiment 77, provided herein is a method according to embodiment 76, wherein the polypeptide expressed by the cell comprises a chimeric antigen receptor (CAR) or a portion thereof. In embodiment 78, provided herein is a method according to embodiment 77, wherein the cell is a human T cell or a human iPSC. In embodiment 79, provided herein is a method according to embodiment 77, wherein the cell is a human T cell. In embodiment 80, provided herein is a method according to embodiment 77, wherein the cell is a human iPSC.

[0266] In embodiment 81, provided herein is a composition comprising a first polynucleotide encoding a polypeptide comprising a nucleic acid-guided nuclease, comprising a CRISPR Type V nuclease polypeptide, wherein the polynucleotide has less than 75% sequence identity to SEQ ID NO: 22. In embodiment 82, provided herein is a composition of embodiment 81, wherein the nuclease polypeptide comprises at least one, two, three, four, or five NLSs, each of the NLSs being at or near the N-terminus or C-terminus of the nuclease polypeptide. In embodiment 83, provided herein is a composition according to embodiment 82, wherein one or more of the NLSs comprise an NLS from SV40 large T antigen, an NLS from nucleoplasmin, e.g., the nucleoplasmin bipartite NLS, a c-myc NLS, an hRNPA1 M9 NLS, an IBB domain of importin alpha NLS, a sarcoma T protein NLS, a sequence derived from human p53 NLS, a sequence derived from mouse c-abl IV NLS, an influenza virus NS1 NLS, a hepatitis virus delta antigen NLS, a mouse Mx1 protein NLS, a human poly(ADP-ribose) polymerase NLS, a steroid hormone receptor (human) glucocorticoid NLS, and / or an EGL-13 NLS. In embodiment 84, provided herein is a composition according to embodiment 83, wherein one or more of the NLSs comprise an NLS from SV40 large T antigen. In embodiment 85, provided herein is the composition of embodiment 84, wherein the NLS(s) comprise the sequence of SEQ ID NO: 5. In embodiment 86, provided herein is a composition according to any one of embodiments 83 to 85, wherein one or more of the NLSs comprise an NLS from nucleoplasmin. In embodiment 87, provided herein is a composition according to embodiment 86, wherein the nucleoplasmin NLS comprises the sequence of SEQ ID NO: 6. In embodiment 88, provided herein is a composition according to any one of embodiments 83 to 87, wherein one or more of the NLSs comprises a c-myc NLS. In embodiment 89, provided herein is a composition according to embodiment 88, wherein the c-myc NLS comprises the sequence of SEQ ID NO: 7, SEQ ID NO: 8, or SEQ ID NO: 21.In embodiment 90, provided herein is a composition described in embodiment 88, wherein the c-myc NLS comprises the sequence of SEQ ID NO: 21. In embodiment 91, provided herein is a composition described in any one of embodiments 83 to 90, wherein one or more of the NLSs comprises an EGL-13 NLS. In embodiment 92, provided herein is a composition described in embodiment 91, wherein the EGL-13 NLS comprises the sequence of SEQ ID NO: 107. In embodiment 93, provided herein is a composition described in any one of embodiments 82 to 92, wherein the NLS(s) are at or near the N-terminus of the polypeptide. In embodiment 94, provided herein is a composition described in any one of embodiments 81 to 93, wherein the first polynucleotide comprises a polynucleotide encoding a purification tag. In embodiment 95, provided herein is a composition described in embodiment 94, wherein the purification tag is at or near the N-terminus of the nuclease polypeptide. In embodiment 96, provided herein is a composition according to embodiment 94 or embodiment 95, wherein the purification tag comprises a poly-His tag, such as a Gly-6xHis tag or a Gly-8xHis tag; a short epitope tag, such as FLAG, hemagglutinin (HA), c-myc, T7, Glu-Glu; maltose-binding protein (mbp); N-terminal glutathione S-transferase (GST); or calmodulin-binding peptide (CBP). In embodiment 97, provided herein is a composition according to embodiment 96, wherein the purification tag comprises a poly-His tag. In embodiment 98, provided herein is a composition according to embodiment 97, wherein the purification tag comprises a gly-6xHis tag. In embodiment 99, provided herein is a composition according to embodiment 97, wherein the purification tag comprises a gly-8xHis tag. In embodiment 100, provided herein is any one of the compositions according to embodiments 81 to 99, wherein the Type V CRISPR nuclease polypeptide comprises a cleavage site. In embodiment 101, provided herein is a composition according to embodiment 100, wherein the cleavage site is at or near the N-terminus of the nuclease polypeptide.In embodiment 102, provided herein is a composition according to embodiment 100 or embodiment 101, wherein the cleavage site comprises a tobacco etch virus (TEV) cleavage site. In embodiment 103, provided herein is a composition according to embodiment 102, wherein the cleavage site comprises the sequence of SEQ ID NO: 108. In embodiment 104, provided herein is a composition according to embodiment 103, comprising five NLSs, a purification tag, and a cleavage site at or near the N-terminus of the polypeptide, wherein the cleavage site follows the purification tag. In embodiment 105, provided herein is a composition according to any one of embodiments 81 to 104, wherein the polynucleotide encodes a polypeptide comprising a sequence at least 60, 70, 80, 85, 90, 95, 98, 99%, or 100% identical, preferably at least 80%, more preferably at least 90%, even more preferably at least 95%, and even more preferably at least 98% identical to any one of SEQ ID NOs: 109 to 112. In embodiment 106, provided herein is a composition according to any one of embodiments 81 to 105, wherein the polynucleotide encodes a polypeptide comprising a sequence at least 60, 70, 80, 85, 90, 95, 98, 99%, or 100%, preferably at least 80%, more preferably at least 90%, even more preferably at least 95%, and even more preferably at least 98% identical to any one of SEQ ID NO: 112. In embodiment 107, provided herein is a composition according to any one of embodiments 81 to 105, wherein the first polynucleotide comprises a sequence at least 50, 60, 70, 80, 90, 95, 97, or 99% identical, or 100% identical, preferably at least 80%, more preferably at least 90%, even more preferably at least 95%, and even more preferably at least 98% identical to SEQ ID NO: 113.In embodiment 108, provided herein is a composition according to any one of embodiments 81-107, further comprising a second polynucleotide encoding a gNA or a portion thereof, wherein the gNA, e.g., gRNA, comprises a target nucleotide sequence within the polynucleotide or a spacer sequence that targets a target nucleotide sequence within the polynucleotide encoding the gNA, e.g., gRNA, wherein the gNA, e.g., gRNA, is compatible with a Type V CRISPR nuclease. In embodiment 109, provided herein is a composition according to embodiment 108, wherein the first and second polynucleotides are the same. In embodiment 110, provided herein is a composition according to any one of embodiments 81-109, further comprising a third polynucleotide comprising a donor template.

[0267] In embodiment 111, provided herein is a vector comprising a polynucleotide or a polynucleotide according to any one of embodiments 81 to 110.

[0268] In embodiment 112, provided herein is a cell comprising the composition of any one of embodiments 81 to 110. In embodiment 113, provided herein is a composition of embodiment 112, wherein the cell is a human cell. In embodiment 114, provided herein is a composition of embodiment 113, wherein the cell is an immune cell or a stem cell. In embodiment 115, provided herein is a composition of embodiment 113, wherein the cell is an immune cell. In embodiment 116, provided herein is a composition of embodiment 115, wherein the cell is a T cell. In embodiment 117, provided herein is a composition of embodiment 113, wherein the cell is a stem cell. In embodiment 118, provided herein is a composition of embodiment 117, wherein the cell is an iPSC.

[0269] In embodiment 119, provided herein is a method comprising inserting into a cell the composition of any one of embodiments 81 to 111. In embodiment 120, provided is the method of embodiment 119, wherein inserting the composition into a cell comprises electroporation.

[0270] In embodiment 121, there is provided herein a method comprising: (i) inserting into a cell the composition of any one of embodiments 81 to 107; and (ii) inserting into the cell a gRNA, e.g., gRNA, compatible with the Type V CRISPR nuclease encoded by the composition. In embodiment 122, there is provided a method of embodiment 121, wherein steps (i) and (ii) comprise electroporation.

[0271] While preferred embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Those skilled in the art will now envision numerous variations, changes, and substitutions without departing from the invention. It is understood that various alternatives to the embodiments of the invention described herein may be employed in practicing the invention. It is intended that the following claims define the scope of the invention and that methods and structures within the scope of these claims and their equivalents be covered thereby. [Sequence List Free Text]

[0272] SEQ ID NO: 1: Description of artificial sequence: synthetic polypeptide SEQ ID NO: 2: Description of artificial sequence: synthetic polypeptide SEQ ID NO: 3: Description of artificial sequence: synthetic polypeptide SEQ ID NO: 4: Description of artificial sequence: synthetic polypeptide SEQ ID NO: 5: Description of artificial sequence: synthetic peptide SEQ ID NO: 6: Description of artificial sequence: synthetic peptide SEQ ID NO: 7: Description of artificial sequence: synthetic peptide SEQ ID NO: 8: Description of artificial sequence: synthetic peptide SEQ ID NO: 9: Description of artificial sequence: synthetic polypeptide SEQ ID NO: 10: Description of artificial sequence: synthetic polypeptide SEQ ID NO: 11: Description of artificial sequence: synthetic peptide SEQ ID NO: 12: Description of artificial sequence: synthetic peptide SEQ ID NO: 13: Description of artificial sequence: synthetic peptide SEQ ID NO: 14: Description of artificial sequence: synthetic peptide SEQ ID NO: 15: Description of artificial sequence: synthetic peptide SEQ ID NO: 16: Description of artificial sequence: synthetic peptide SEQ ID NO: 17: Description of artificial sequence: synthetic peptide SEQ ID NO: 18: Description of artificial sequence: synthetic peptide SEQ ID NO: 19: Description of artificial sequence: synthetic peptide SEQ ID NO: 20: Description of artificial sequence: synthetic peptide SEQ ID NO: 21: Description of artificial sequence: synthetic peptide SEQ ID NO: 22: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 23: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 24: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 25: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 26: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 27: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 28: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 29: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 30: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 31: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 32: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 33: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 34: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 35: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 36: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 37: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 38: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 39: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 40: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 41: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 42: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 43: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 44: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 45: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 46: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 47: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 48: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 49: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 50: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 51: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 52: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 53: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 54: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 55: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 56: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 57: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 58: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 59: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 60: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 61: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 62: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 63: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 64: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 65: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 66: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 67: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 69: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 70: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 71: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 72: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 73: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 74: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 75: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 76: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 77: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 78: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 79: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 80: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 81: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 82: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 83: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 84: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 85: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 86: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 87: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 88: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 89: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 90: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 91: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 92: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 93: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 94: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 95: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 96: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 97: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 98: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 99: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 100: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 101: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 102: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 103: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 104: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 105: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 107: Description of artificial sequence: synthetic peptide SEQ ID NO: 108: Description of artificial sequence: synthetic peptide SEQ ID NO: 109: Description of artificial sequence: synthetic polypeptide SEQ ID NO: 110: Description of artificial sequence: synthetic polypeptide SEQ ID NO: 111: Description of artificial sequence: synthetic polypeptide SEQ ID NO: 112: Description of artificial sequence: synthetic polypeptide SEQ ID NO: 113: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 114: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 115: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 116: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 117: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 118: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 119: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 120: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 121: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 122: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 123: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 124: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 125: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 126: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 127: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 128: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 129: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 130: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 131: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 132: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 133: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 134: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 135: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 136: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 137: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 138: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 139: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 140: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 141: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 142: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 143: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 144: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 145: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 146: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 147: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 148: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 149: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 150: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 151: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 152: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 153: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 154: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 155: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 156: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 157: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 158: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 159: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 160: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 161: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 162: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 163: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 164: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 165: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 166: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 167: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 168: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 169: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 170: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 171: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 172: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 173: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 174: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 175: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 176: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 177: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 178: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 179: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 180: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 181: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 182: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 183: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 184: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 185: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 186: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 187: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 188: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 189: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 190: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 191: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 192: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 193: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 194: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 195: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 196: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 197: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 198: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 199: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 200: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 201: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 202: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 203: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 204: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 205: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 206: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 207: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 208: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 209: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 210: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 211: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 212: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 213: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 214: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 215: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 216: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 217: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 218: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 219: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 220: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 221; Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 222: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 223: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 224: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 225: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 226: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 227: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 228: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 229: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 230: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 231: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 232: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 233: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 234: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 235: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 236: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 237: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 238: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 239: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 240: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 241: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 242: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 243: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 244: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 245: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 246: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 247: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 248: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 249: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 250: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 251: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 252: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 253: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 254: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 255: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 256: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 257: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 258: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 259: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 260: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 261: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 262: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 263: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 264: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 265: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 266: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 267: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 268: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 269: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 270: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 271: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 272: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 273: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 274: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 275: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 276: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 277: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 278: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 279: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 280: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 281: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 282: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 283: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 284: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 285: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 286: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 287: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 288: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 289: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 290: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 291: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 292: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 293: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 294: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 295: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 296: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 297: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 298: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 299: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 300: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 301: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 302: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 303: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 304: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 305: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 306: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 307: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 308: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 309: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 310: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 311: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 312: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 313: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 314: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 315: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 316: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 317: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 318: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 319: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 320: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 321: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 322: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 323: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 324: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 325: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 326: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 327: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 328: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 329: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 330: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 331: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 332: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 333: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 334: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 335: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 336: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 337: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 338: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 339: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 340: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 341: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 342: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 343: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 344: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 345: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 346: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 347: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 348: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 349: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 350: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 351: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 352: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 353: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 354: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 355: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 356: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 357: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 358: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 359: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 360: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 361: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 362: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 363: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 364: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 365: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 366: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 367: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 368: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 369: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 370: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 371: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 372: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 373: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 374: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 375: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 376: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 377: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 378: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 379: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 380: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 381: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 382: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 383: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 384: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 385: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 386: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 387: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 388: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 389: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 390: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 391: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 392: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 393: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 394: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 395: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 396: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 397: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 398: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 399: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 400: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 401: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 402: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 403: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 404: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 405: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 406: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 407: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 408: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 409: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 410: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 411: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 412: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 413: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 414: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 415: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 416: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 417: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 418: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 419: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 420: Description of artificial sequence: synthetic polynucleotide SEQ ID NO: 421: Description of artificial sequence: synthetic peptide SEQ ID NO: 422: Description of artificial sequence: synthetic peptide SEQ ID NO: 423: Description of artificial sequence: 6xHis tag SEQ ID NO: 424: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 425: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 426: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 427: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 428: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 429: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 430: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 431: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 432: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 433: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 434: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 435: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 436: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 437: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 438: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 439: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 440: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 441: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 442: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 443: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 444: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 445: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 446: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 447: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 448: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 449: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 450: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 451: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 452: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 453: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 454: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 455: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 456: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 457: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 458: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 459: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 460: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 461: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 462: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 463: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 464: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 465: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 466: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 467: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 468: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 469: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 470: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 471: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 472: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 473: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 474: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 475: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 476: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 477: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 478: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 479: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 480: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 481: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 482: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 483: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 484: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 485: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 486: Description of artificial sequence: synthetic oligonucleotide SEQ ID NO: 487: Description of artificial sequence: synthetic oligonucleotide

Claims

1. A composition comprising a nucleic acid-guided nuclease, wherein the nucleic acid-guided nuclease comprises a V-type CRISPR nuclease polypeptide and comprises at least one nuclear localization signal (NLS) at or near the N-terminus or C-terminus of the polypeptide.

2. The composition of claim 1 , wherein the nuclease is a Type Va nuclease.

3. The composition of claim 1 or 2, wherein the V-type CRISPR nuclease polypeptide has at least 60% sequence identity to SEQ ID NO:

1.

4. The composition of any one of claims 1-3, wherein the Type V CRISPR nuclease polypeptide comprises two NLSs, one or both of which are at or near the N-terminus or C-terminus of the polypeptide.

5. The composition of any one of claims 1-4, wherein the Type V CRISPR nuclease polypeptide comprises three NLSs, each of which is at or near the N-terminus or C-terminus of the polypeptide.

6. The composition of any one of claims 1-5, wherein the Type V CRISPR nuclease polypeptide comprises four NLSs, each of which is at or near the N-terminus or C-terminus of the polypeptide.

7. The composition of any one of claims 1-6, wherein the Type V CRISPR nuclease polypeptide comprises at least five NLSs, each of which is at or near the N-terminus or C-terminus of the polypeptide.

8. The composition of any one of claims 4 to 7, wherein at least two of said NLSs are at or near the N-terminus of said polypeptide.

9. The composition of any one of claims 5 to 7, wherein at least three of the NLSs are at or near the N-terminus of the polypeptide.

10. The composition of any one of claims 6 to 7, wherein at least four of the NLSs are at or near the N-terminus of the polypeptide.

11. The composition of claim 7, wherein the five NLS are at or near the N-terminus of the polypeptide.

12. The composition of claim 11, comprising a sequence at least 60% identical to any one of SEQ ID NOs: 109-112.

13. The composition of any one of claims 1 to 3, wherein the V-type CRISPR nuclease polypeptide comprises at least 1 to 30 NLSs, each of which is at or near the N-terminus or C-terminus of the polypeptide.

14. The composition according to any one of claims 4 to 11, wherein at least two of the NLSs have different nuclear localization mechanisms.

15. The composition of any one of claims 5 to 7 or 9 to 11, wherein at least three of the NLSs have different nuclear localization mechanisms.

16. 16. The composition of any one of claims 1 to 15, wherein one or more of the NLSs comprises an NLS of the large T antigen of the SV40 virus, an NLS from nucleoplasmin, a nucleoplasmin bipartite NLS, a c-myc NLS, a hRNPA1 M9 NLS, an IBB domain of the importin alpha NLS, a sarcoma T protein NLS, a sequence derived from a human p53 NLS, a sequence derived from a mouse c-abl IV NLS, a sequence of an influenza virus NS1 NLS, a sequence of a hepatitis virus delta antigen NLS, a sequence of a mouse Mx1 protein NLS, a sequence of a human poly(ADP-ribose)polymerase NLS, a sequence of a steroid hormone receptor (human) glucocorticoid NLS, and / or a sequence of an EGL-13 NLS.

17. 17. The composition of claim 16, wherein one or more of the NLSs comprises an NLS of the large T antigen of the SV40 virus, an NLS from nucleoplasmin, a c-myc NLS, or an EGL-13 NLS.

18. 17. The composition of claim 16, wherein two or more of the NLSs comprise the NLS of the SV40 virus large T antigen.

19. 19. The composition of claim 17 or 18, wherein in one or more of the NLSs, the NLS of the large T antigen of the SV40 virus comprises the sequence of SEQ ID NO:5, the NLS from nucleoplasmin comprises the sequence of SEQ ID NO:6, the c-myc NLS comprises the sequence of SEQ ID NO:7, SEQ ID NO:8 or SEQ ID NO:21, and / or the EGL-13 NLS comprises the sequence of SEQ ID NO:

107. 。

20. The composition of any one of claims 1 to 19, wherein the V-type CRISPR nuclease polypeptide further comprises a purification tag.

21. 21. The composition of claim 20, wherein the purification tag is at or near the N-terminus of the nuclease polypeptide.

22. 22. The composition of claim 20 or 21, wherein the purification tag comprises a poly-his tag, maltose binding protein (mbp), N-terminal glutathione S-transferase (GST), or calmodulin binding peptide (CBP).

23. 23. The composition of claim 22, wherein the poly-his tag comprises a gly-6x His tag or a gly-8x His tag.

24. The composition of any one of claims 1 to 23, wherein the V-type CRISPR nuclease polypeptide further comprises a cleavage site.

25. 25. The composition of claim 24, wherein the cleavage site is at or near the N-terminus of the nuclease polypeptide.

26. 26. The composition of claim 24 or 25, wherein the cleavage site comprises a tobacco etch virus (TEV) cleavage site.

27. 27. The composition of claim 26, wherein the cleavage site comprises the sequence of SEQ ID NO:

108.

28. 28. The composition of claim 27, comprising five NLSs, a purification tag, and the cleavage site at or near the N-terminus of the polypeptide, wherein the cleavage site is after the purification tag.

29. 29. The composition of claim 28, comprising a sequence at least 60% identical to SEQ ID NO: 111 or 112.

30. The composition of any one of claims 1 to 29, further comprising a guide nucleic acid (gNA) comprising a spacer sequence that targets a target nucleotide sequence within a polynucleotide or a polynucleotide encoding a gNA, wherein the gNA is compatible with the V-type CRISPR nuclease.

31. 31. The composition of claim 30, wherein the target nucleotide is within 50 nucleotides of a protospacer adjacent motif (PAM) sequence specific for the V-type CRISPR nuclease.

32. 32. The composition of claim 31 , wherein the PAM comprises the sequence YTTN, where Y is T or C and N is A, T, G or C.

33. 33. The composition of claim 32, wherein the PAM comprises the sequence YTTV or TTTV, where V is A, G or C.

34. The composition of claim 30, wherein the gNA is a gRNA.

35. 35. The composition of claim 34, wherein the gRNA is a double-stranded gRNA.

36. 36. The composition of Claims 34 or 35, wherein the composition comprises the gRNA, and the gRNA comprises one or more chemical modifications.

37. 37. The composition of claim 36, wherein the chemical modification comprises 2'-O-alkyl, 2'-O-methyl, phosphorothioate, phosphonoacetate, thiophosphonoacetate, 2'-O-methyl-3'-phosphorothioate, 2'-O-methyl-3'-phosphonoacetate, 2'-O-methyl-3'-thiophosphonoacetate, 2'-deoxy-3'-phosphonoacetate, 2'-deoxy-3'-thiophosphonoacetate, suitable alternatives, or combinations thereof.

38. 38. The composition of any one of claims 34 to 37, wherein the ratio of guanine:uracil in the gRNA is at least 51:

49.

39. The composition of any one of claims 30 to 38, wherein the molar ratio of gNA to V-type CRISPR nuclease is at least 1.1:

1.

40. The composition of any one of claims 30 to 39, wherein the molar amount of gNA is at least 10 pmol and / or not more than 300 pmol.

41. The composition of any one of claims 30 to 40, further comprising a donor template.

42. The composition of claim 41 , wherein the donor template comprises homology arms.

43. The donor template is at least 0.1 μg μL -1 and / or 10 μg μL -1 43. The composition of claim 41 or 42, present in the following amounts:

44. The composition of any one of claims 30 to 43, further comprising an anionic polymer.

45. 45. The composition of claim 44, wherein the anionic polymer comprises polyglutamic acid (PGA).

46. The anionic polymer is at least 20 μg μL -1 , and / or 1000 μg μL -1 46. ​​The composition of claim 44 or 45, present in the following concentrations:

47. A cell comprising a composition described in any one of claims 1 to 46.

48. 48. The cell of claim 47, wherein the cell is a human cell.

49. 49. The cell of claim 48, wherein the cell is an immune cell or a stem cell.

50. 50. The cell of claim 49, wherein the cell is a T cell or an induced pluripotent stem cell (iPSC).

51. 11. A method for modifying a target polynucleotide comprising: (i) contacting with a composition of any one of claims 30 to 46; and (ii) allowing the nuclease and the guide nucleic acid to modify a target genomic region.

52. The method of claim 51, wherein the composition is a composition according to any one of claims 41 to 46.

53. 53. The method of claim 51 or 52, wherein the target polynucleotide is a genome or part of a genome in a cell.

54. 54. The method of claim 53, wherein the cell is a human cell.

55. 55. The method of claim 54, wherein the cell is an immune cell or a stem cell.

56. 56. The method of claim 55, wherein the cell is a T cell or an iPSC.

57. 57. The method of any one of claims 52-56, wherein the donor template comprises a mutation in a PAM within 50 nucleotides of the target nucleotide sequence in the target polynucleotide.

58. 57. The method of any one of claims 53 to 56, wherein the composition is the composition of claim 52 and the donor template comprises a polynucleotide encoding a polypeptide expressed by the cell.

59. 59. The method of claim 58, wherein the polypeptide expressed by the cell comprises a chimeric antigen receptor (CAR) or a portion thereof.

60. 60. The method of claim 59, wherein the cell is a human T cell or a human iPSC.

61. A composition comprising a first polynucleotide encoding a polypeptide comprising a nucleic acid-induced nuclease comprising a V-type CRISPR nuclease polypeptide, wherein the polynucleotide has less than 75% sequence identity to SEQ ID NO:

22.

62. 62. The composition of claim 61, wherein the nuclease polypeptide comprises at least one, two, three, four, or five NLSs, each of which is at or near the N-terminus or C-terminus of the nuclease polypeptide.

63. 63. The composition of claim 62, wherein one or more of the NLSs comprise an NLS of the SV40 virus large T antigen, an NLS from nucleoplasmin, a nucleoplasmin bipartite NLS, a c-myc NLS, a hRNPA1 M9 NLS, an IBB domain of the importin alpha NLS, a sarcoma T protein NLS, a sequence derived from a human p53 NLS, a sequence derived from a mouse c-abl IV NLS, a sequence of an influenza virus NS1 NLS, a sequence of a hepatitis virus delta antigen NLS, a sequence of a mouse Mx1 protein NLS, a sequence of a human poly(ADP-ribose)polymerase NLS, a sequence of a steroid hormone receptor (human) glucocorticoid NLS, and / or a sequence of an EGL-13 NLS.

64. 64. The composition of claim 63, wherein one or more of the NLSs comprises an NLS of the large T antigen of the SV40 virus, an NLS from nucleoplasmin, a c-myc NLS, and / or an EGL-13 NLS.

65. 65. The composition of claim 64, wherein in one or more of the NLSs, the NLS of the SV40 virus large T antigen comprises the sequence of SEQ ID NO:5, the NLS from nucleoplasmin comprises the sequence of SEQ ID NO:6, the c-myc NLS comprises the sequence of SEQ ID NO:7, SEQ ID NO:8 or SEQ ID NO:21, and / or the EGL-13 NLS comprises the sequence of SEQ ID NO:

107.

66. 66. The composition of any one of claims 62-65, wherein one or more of the NLSs are at or near the N-terminus of the polypeptide.

67. 67. The composition of any one of claims 61 to 66, wherein the first polynucleotide comprises a polynucleotide encoding a purification tag.

68. 68. The composition of claim 67, wherein the purification tag is at or near the N-terminus of the nuclease polypeptide.

69. 69. The composition of claim 67 or 68, wherein the purification tag comprises a poly-his tag, a short epitope tag, maltose binding protein (mbp), N-terminal glutathione S-transferase (GST), or calmodulin binding peptide (CBP).

70. 70. The composition of claim 69, wherein the poly-his tag comprises a gly-6xHis tag or a gly-8xHis tag.

71. 71. The composition of any one of claims 61 to 70, wherein the V-type CRISPR nuclease polypeptide comprises a cleavage site.

72. 72. The composition of claim 71, wherein the cleavage site is at or near the N-terminus of the nuclease polypeptide.

73. 73. The composition of claim 71 or 72, wherein the cleavage site comprises a tobacco etch virus (TEV) cleavage site.

74. 74. The composition of claim 73, wherein the cleavage site comprises the sequence of SEQ ID NO:

108.

75. 75. The composition of claim 74, comprising five NLSs, a purification tag, and the cleavage site at or near the N-terminus of the polypeptide, wherein the cleavage site is after the purification tag.

76. The composition of any one of claims 61 to 75, wherein the polynucleotide encodes a polypeptide comprising a sequence that is at least 60% identical to any one of SEQ ID NOs: 109 to 112.

77. 77. The composition of any one of claims 61 to 76, wherein the first polynucleotide comprises a sequence at least 50% identical to SEQ ID NO:

113.

78. The composition of any one of claims 61 to 77, further comprising a second polynucleotide encoding a gNA or a portion thereof, wherein the gNA comprises a target nucleotide sequence within a polynucleotide or a spacer sequence that targets a target nucleotide sequence within a polynucleotide encoding the gNA, wherein the gNA is compatible with the V-type CRISPR nuclease.

79. 79. The composition of claim 78, wherein the first polynucleotide and the second polynucleotide are the same.

80. 80. The composition of any one of claims 61 to 79, further comprising a third polynucleotide comprising a donor template.

81. A vector comprising a first polynucleotide encoding a polypeptide comprising a nucleic acid-induced nuclease comprising a V-type CRISPR nuclease polypeptide, wherein the polynucleotide has less than 75% sequence identity to SEQ ID NO:

22.

82. A cell comprising a composition described in any one of claims 61 to 80 or a vector described in claim 81.

83. The cell of claim 82, wherein the cell is a human cell.

84. 84. The cell of claim 83, wherein the cell is an immune cell or a stem cell.

85. 85. The cell of claim 84, wherein the cell is a T cell or an iPSC.

86. A genome editing method comprising inserting a composition described in any one of claims 61 to 80 or a vector described in claim 81 into a cell.

87. The genome editing method of claim 86, wherein inserting the composition into the cell comprises electroporation.

88. A genome editing method comprising: (i) inserting into a cell a composition described in any one of claims 61 to 77; and (ii) inserting into the cell a gNA that is compatible with the V-type CRISPR nuclease encoded by the composition.

89. The genome editing method of claim 88, wherein steps (i) and (ii) comprise electroporation.