Improved gene editing system, guides, and methods

AU2025209734A1Pending Publication Date: 2026-07-30RENAGADE THERAPEUTICS MANAGEMENT INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
AU · AU
Patent Type
Applications
Current Assignee / Owner
RENAGADE THERAPEUTICS MANAGEMENT INC
Filing Date
2025-01-16
Publication Date
2026-07-30

AI Technical Summary

Technical Problem

Current CRISPR-Cas type V systems face challenges in achieving sufficient editing efficiency, precision, deliverability, scalability, and affordability for treating genetic disorders and complex diseases.

Method used

Development of Cas Type V-based gene editing systems comprising a Type V polypeptide and guide RNA that form a complex to target and edit genomic sequences, with optional accessory proteins, delivered via DNA, RNA, protein, or protein-nucleic acid complexes using vectors like AAV and LNPs, employing homology-dependent repair or non-homologous end joining for cellular repair.

Benefits of technology

Enhances editing efficiency, precision, and deliverability, making it suitable for treating genetic disorders and complex diseases while being scalable and cost-effective.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The present disclosure provides methods and compositions comprising novel Cas Type V programmable nucleases and lipid nanoparticles capable of delivering the Cas Type V programmable nucleases and genome editing systems comprising same. For therapeutic applications, as well as plants and industrial biotechnology.
Need to check novelty before this filing date? Find Prior Art

Description

IMPROVED GENE EDITING SYSTEM, GUIDES, AND METHODSRELATED APPLICATIONS

[0001] This application refers to and incorporates by reference the contents of tire following U.S. Provisional Application Nos: 63 / 621,975 [Ref. No. CSG018-P1], filed January 17, 2024. 63 / 552,429 [Ref. No. CSG018- P2], filed February 12, 2024, 63 / 648,416 [Ref. No. CSG018-P3]. filed May 16. 2024. and 63 / 706,484 [Ref. No. CSG018-P4] filed October 11, 2024, in their entireties.

[0002] The foregoing applications, and all documents cited therein and all documents cited or referenced herein, together with any manufacturer’s instructions, descriptions, product specifications, and product sheets for any products mentioned herein or in any document incorporated by reference herein, are hereby incorporated herein by reference, and may be employed in the practice of the disclosed subject matter. More specifically, all referenced documents are incorporated by reference to the same extent as if each individual document was specifically and individually indicated to be incorporated by reference.SEQUENCE LISTING

[0003] The instant application contains a Sequence Listing which has been submitted via Patent Center and is hereby incorporated by reference in its entirety. Said .xml copy, created on 16 January 2025, is named J0356-99006.xml, and is 4,258.691 bytes in size.TECHNICAL FIELD

[0004] The present disclosure generally relates to gene editing systems, nucleases, guides, methods and compositions used for precise genome editing, including nucleic acid insertions, replacements, and deletions at targeted and precise genome sites, wherein said systems, methods, and compositions are based on engineered class II / type V Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR)-Cas systems and engineered guides.BACKGROUND

[0005] Genome editing tools encompass a diverse set of technologies that can make many types of genomic alterations in various contexts. These technologies have evolved over the last couple of decades to provide a range of user-programmable editing tools that include ZFN (zinc finger) nuclease editing systems, meganuclease editing systems, and TALENS (transcription activator-like effector nucleases). The past decade has seen an explosive growth in a new generation of genome editing systems based on components from bacterial immune pathways, including CRISPR (clustered regularly interspaced short palindromic repeats) and the associated CRISPR-associated proteins (e g., CRISPR-Cas9) (Jinek et al., “A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity,” Science, Vol. 337 (6096), pp. 816- 821), meganuclease editors (Boissel et al., “megaTALs: a rare-cleaving nuclease architecture for therapeuticgenome engineering,” Nucleic Acids Research 42: pp. 2591-2601) and bacterial retron systems (Schubert et al., '‘High-throughput functional variant screens via in vivo production of single-stranded DNA,” PNAS, April 27, 2021, Vol. 118(18), pp. 1-10). In particular, CRISPR-Cas9 has been derivatized in numerous ways to expand upon its guide RNA-based programmable double-strand cutting activity to fonn systems ranging from finding alternative CRISPR Cas nuclease enzymes having different PAM requirements and cutting properties (e.g.. Casl2a, Casl2f, Casl3a. and Casl3b) to base editing (Komor et al., “Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage,” Nature, May 19, 2016, 533 (7603); pp. 420-424 [cytosine base editors or CBEs] and Gaudelli et al., “Programmable base editing of A-T to G-C in genomic DNA without DNA cleavage,” Nature, Vol. 551, pp. 464-471 [adenine base editors or ABEs]) to prime editing (Anzalone et al., “Search-and-replace genome editing without double-strand breaks or donor DNA,” Nature, Dec 2019, 576 (7789): pp. 149-157) to twin prime editing (Anzalone et al., “Programmable deletion, replacement, integration and inversion of large DNA sequences with twin prime editing,” Nature Biotechnology, Dec 9, 2021, vol. 40, pp. 731-740) to epigenetic editing (Kungulovski and Jeltsch, “Epigenome Editing: State of the Art, Concepts, and Perspective,” Trends in Genetics, Vol.32, 206, pp. 101-113) to CRISPR-directed integrase editing (Y amell et al., “Drag-and-drop genome insertion of large sequences without double-stranded DNA cleavage using CRISPR-directed integrases,” Nature Biotechnology, Nov 24. 2022, doi.org / 10.1038 / s41587-022-01527-4 (“PASTE”)).

[0006] In particular, application of CRISPR-associated systems (“CRISPR-Cas systems”) in human therapeutics is anticipated to be curative in ameliorating various monogenic diseases and disorders. Current clinical trials are underway to treat, for instance, Transfusion-dependent 0-thalassemia (TDT) and sickle cell disease (SCD) by tire autologous transfusion of CRISPR / Cas9-edited CD34+ hematopoietic stem cells Frangoul, Haydar et al. “CRISPR-Cas9 Gene Editing for Sickle Cell Disease and |3- Thalassemia.” The New England journal of medicine vol. 384,3 (2021): 252-260. doi:10.1056 / NEJMoa2031054 and ATTR amyloidosis Gillmore, Julian D et al. “CRISPR- Cas9 In Vivo Gene Editing for Transthyretin Amyloidosis.” The New England journal of medicine vol. 385,6 (2021): 493-502. doi: 10.1056 / NEJMoa2107454, which is incorporated herein by reference.

[0007] The potential of such CRISPR-Cas systems has sparked the discover}' of many novel CRISPR-Cas variants where such systems have been classified into 2 classes (i.e., class I and II) and 6 types and 33 subtypes based on their genes, protein subunits and the structure of their gRNAs. Makarova, K.S.. Wolf, Y.I., Iranzo, J. et al. Evolutionary classification of CRISPR-Cas systems: a burst of class 2 and derived variants. Nat Rev Microbiol 18, 67-83 (2020). doi: 10.1038 / s41579-019-0299-x, which is incorporated herein by reference.

[0008] Among the diverse CRISPR-Cas systems, class II has the most extensive applications in gene editing due to its earlier discover}' and by virtue of it having only one effector protein. By contrast, the effectornucleases of the type V family are diverse due to extensive diversity over the N-tenninus of the protein, as evident by comparing the cry stal structures of Casl2a, Casl2b, and Casl2e type V nucleases (Tong et al., “The Versatile Type V CRISPR Effectors and Their Application Prospects,” Front. Cell Dev. Biol., 2021, vol. 8). The C-terminus regions of the type V effector nucleases are more highly conserved, however, which comprise a conserved RuvC-like endonuclease (RuvC) domain. It is reported that the RuvC domain of type V effectors is derived from the TnpB protein encoded by autonomous or non-autonomous transposons (Shmakov et al., “Diversity and evolution of class 2 CRISPR-Cas systems,” 2017, Nat. Rev. Microbiol. 15, 169-182. doi: 10.1038 / nrmicro.2016.184). The type V systems are further subdivided into many subty pes, including types V-A to V-I, type V-K, type V-U, and CRISPR Cas® (Hajizadeh et al., “The expanding class 2 CRISPR toolbox: diversity, applicability, and targeting drawbacks,” 2019, BioDrugs 33, 503-513. doi: 10.1007 / s40259-019-00369-y). The corresponding effector nucleases in these various subtypes have shown a range of different substrates, including some that act only on double-stranded DNA (dsDNA), but also those that act on both dsDNA as well as single-stranded DNA (ssDNA), and those that act on single-stranded RNA (ssRNA). This multifunctionality has put the type V CRISPR-Cas system into the focus of recent studies.

[0009] While a number of CRISPR-Cas type V systems have been used for various applications, including gene editing, reported drawbacks have been published to indicate the need for improved CRISPR-Cas type V systems for suitability of desired applications. Therefore, there remains much room for improvement and design to achieve an effective type V CRISPR-Cas system for gene editing that bears sufficient editing efficiency, improved precision, better deliverability, and which remains affordable, easy to scale, and has improved ability to treat various genetic disorders and complex diseases.SUMMARY

[0010] Hie present disclosure provides Cas TypeV-based gene editing systems for use in various applications, including precision gene editing in cells, tissues, organs, or organisms. In various embodiments, the Cas TypeV-based gene editing systems comprise (a) a Type V polypeptide and (b) a Type V guide RNA which is capable of associating with a Type V polypeptide to form a complex such that the complex localizes to a target nucleic acid sequence (e.g., a genomic or plasmid target sequence) and binds thereto. In various embodiments, the Type V polypeptide has a nuclease activity which results in the cutting of both strands of DNA

[0011] In various embodiments, the Cas Type V polypeptide is a polypeptide selected from Table S15A, or a polypeptide having at least 70%, 75%, 80%, 85%, 90%, 95%, 99%. or 100% sequence identity with a polypeptide from Table S 15A. In various other embodiments, the Cas Type V polypeptide is encoded by a polynucleotide sequence selected from Table S15B, or a polynucleotide having at least 70%, 75%, 80%, 85%, 90%, 95%, 99%, or 100% sequence identity with a polynucleotide of Table S15B. In various otherembodiments, the Casl2a guide RNA is selected from any Cas Type V guide sequence disclosed in Table S15C, or a nucleic acid molecule having at least 70%, 75%, 80%, 85%, 90%, 95%, 99%, or 100% sequence identity with a Casl2a guide sequence of Table S15C.

[0012] In various embodiments, the Cas Type V guide RNA may comprise (a) a portion that binds or associates with a Cas Type V polypeptide and (b) a region that comprises a targeting sequence, i.e.. a sequence which is complementary to target nucleic acid sequence. For Cas Type V guide RNA designs, just like for Cas9 guide RNA, the target sequence is typically next to a PAM sequence. But for Cas Type V, the PAM sequence in various embodiments is typically TTTV, where V typically represents A, C, or G. In various embodiments, the “V” of the TTTV is immediately adjacent to the most 5’ base of the non-targeted strand side of the protospacer element. As for Cas9 guide RNA designs, the PAM sequence is typically not included in the guide RNA design.

[0013] In various embodiments, the guide RNA for Cas Type V is relatively short at only approximately 40- 44 bases long. The part that base pairs to the protospacer in the target sequence is 20-24 bases in length, and there is also a constant about 20-base section that binds to Cas Type V.

[0014] In various embodiments, nomenclature for a Cas Type V guide RNA is referred to as a “crRNA” and there is no Cas9-like “tracrRNA” component.

[0015] In other aspects, the Cas Type V -based gene editing systems may comprise one or more additional accessory proteins having genome modifying functions, including recombinases, invertases, nucleases, polymerases, ligases, deaminases, reverse transcriptases, or epigenetic modifying functions. In various embodiments, the accessory' proteins may be provided separately. In other embodiments, tire accessory' proteins may be fused to a Cas Type V nuclease, optionally with a linker.

[0016] In still another aspect, the disclosure provides delivery systems for introducing the Cas Type V - based gene editing systems or components thereof into cells, tissues, organs, or organisms. Depending on the chosen format, the Cas Type V -based gene editing systems and / or the individual or combined components thereof may be delivered as DNA molecules (e.g., encoded on one or more plasmids), RNA molecules (e.g., guide RNAs for targeting the Cas Type V protein or linear or circular mRNAs coding for the Cas Type V protein or accessory' protein components of the Cas Type V -based gene editing systems), proteins (e.g., Casl2a polypeptides, accessory proteins having other functions (e.g., recombinases, nucleases, polymerases, ligases, deaminases, or reverse transcriptases), or protein-nucleic acid complexes (e.g., complexes between a guide RNA and a Cas Type V protein or fusion protein comprising a Cas Type V protein).

[0017] In another aspect, the present disclosure provides nucleic acid molecules encoding the Cas Type V - based gene editing systems or components thereof. In yet another aspect, the disclosure provides vectors for transferring and / or expressing said Cas Type V -based gene editing systems, e.g., under in vitro, ex vivo, andin vivo conditions. In still another aspect, the disclosure provides cell-delivery compositions and methods, including compositions for passive and / or active transport to cells (e.g., plasmids), delivery by virus-based recombinant vectors (e.g., AAV and / or lentivirus vectors), delivery by non-virus-based systems (e.g., liposomes and LNPs), and delivery by vims-like particles of the Cas Type V -based gene editing systems described herein. Depending on the delivery system employed, the Cas Type V -based gene editing systems described herein may be delivered in the form of DNA (e.g., plasmids or DNA-based vims vectors), RNA (e.g., guide RNA and mRNA delivered by LNPs), a mixture of DNA and RNA, protein (e.g., vims-like particles), and ribonucleoprotein (RNP) complexes. Any suitable combinations of approaches for delivering the components of the herein disclosed Cas Type V -based gene editing systems may be employed.

[0018] In other embodiments, the Cas Type V -based gene editing systems may comprise a template DNA comprising an edit, e.g., a single strand or double strand donor molecule (linear or circular) which may be used by the cell to repair a single or double cut lesion introduced by a Cas Type V -based gene editing systems by way of cellular repair processes, including homology-dependent repair (HDR) (e.g., in dividing cells) or non-homologous end joining (NHEJ) (in non-dividing cells).

[0019] In one embodiment, each of the components of the Cas Type V -based gene editing systems is delivered by an all-RNA system, e.g., the delivery of one or more RNA molecules (e.g., mRNA and / or guide RNA) by one or more LNPs, wherein the one or more RNA molecules form the guide RNA and / or are translated into the polypeptide components (e.g., the Cas Type V polypeptides and / or any accessory proteins), and a DNA or RNA-encoded template DNA molecule (e.g., donor template), as appropriate or desired.

[0020] In yet another aspect, the disclosure provides methods for genome editing by introducing a Cas Type V -based gene editing system described herein into a cell (e.g., under in vitro, in vivo, or ex vivo conditions) comprising a target edit site, thereby resulting in an edit at the target edit. In other aspects, the disclosure provides formulations comprising any of the aforementioned components for delivery to cells and / or tissues, including in vitro, in vivo, and ex vivo delivery, recombinant cells and / or tissues modified by the recombinant Cas Type V -based gene editing systems and methods described herein, and methods of modifying cells by conducting genome editing using the herein disclosed Cas Type V -based gene editing systems.

[0021] The disclosure also provides methods of making the Cas Type V -based gene editing systems, their protein and nucleic acid molecule components, vectors, compositions and fonnulations described herein, as well as to pharmaceutical compositions and kits for modifying cells under in vitro, in vivo, and ex vivo conditions that comprise the herein disclosed genome editing and / or modification systems.

[0022] In various aspects, the invention provides an isolated or recombinant polynucleotide comprising or consisting of a nucleic acid sequence selected from tire group consisting of:(a) a nucleic acid sequence that encodes a Cas Type V polypeptide having the amino acid sequence of SEQ ID NO: 334 (No. ID405), SEQ ID NO: 58 (No. ID414), or SEQ ID NO: 564 (No. ID418), SEQ ID NO: 335 (No. ID406), SEQ ID NO: 331 (No. ID411), SEQ ID NO: 20 (No. ID415), or SEQ ID NO: 445 (No. ID4I9);(b) a nucleic acid sequence that encodes a polypeptide at least 50%. at least 55%. at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8% or at least 99.9% identical to a Cas Type V polypeptide of SEQ ID NO: 334 (No. ID405), SEQ ID NO: 58 (No. 1D414), or SEQ ID NO: 564 (No. 1D418), SEQ ID NO: 335 (No. 1D406), SEQ ID NO: 331 (No. ID41 1), SEQ ID NO: 20 (No. ID415), or SEQ ID NO: 445 (No. ID419);(c) a nucleic acid sequence that is a degenerate variant of the nucleic acid sequence in (a) or (b); and(d) a nucleic acid sequence that hybridizes under stringent conditions to the nucleic acid sequence in in (a) or (b).

[0023] In related aspects, the invention provides an isolated or recombinant guide RNA comprising or consisting of a nucleic acid sequence selected from the group consisting of:(a) one or more crRNA direct repeat sequences or a reverse complement selected from (Group 1) SEQ ID NO:7-12; (Group 2) SEQ ID NO:24-27; (Group 3) SEQ ID NO:36-39; (Group 4) SEQ ID NO:49-52; (Group 5) SEQ ID NO:63-68; (Group 6) SEQ ID NO:84-91; (Group 7) SEQ ID NO: 106-1 1 1 ; (Group 8) SEQ ID NO: 122-125; (Group 9) SEQ ID Nos:21 1- 290; (Group 10) SEQ ID NO:343-354; (Group 11) SEQ ID NO:374-379; (Group 12) SEQ ID NO:390-393; (Group 13) SEQ ID NO:411-422; and (Group 14) SEQ ID N0:500-541;(b) 20 to 35 nucleotides or up to the length of the crRNA from the 3' end of tire crRNA direct repeat sequences or a reverse complement (a) linked to a targeting guide attached to the 3’ end of the direct repeat sequence that is of 16-30 nucleotides in length;(c) (Group 1) SEQ ID NO: 13-15; (Group 2) SEQ ID NO:28-29; (Group 3) SEQ ID NO:40- 41; (Group 4) SEQ ID NO:53-54; (Group 5) SEQ ID NO:69-71; (Group 6) SEQ ID NO:92-95; (Group 7) SEQ ID NO: 112-114; (Group 8) SEQ ID NO: 126-127; (Group 9) SEQ ID NO:291- 330; (Group 10) SEQ ID NO:355-360; (Group 11) SEQ ID NO:380-382; (Group 12) SEQ ID NO:394-395; (Group 13) SEQ ID NO:423-428; and (Group 14) SEQ ID NO:542-563;(d) a nucleic acid sequence that is a degenerate variant of (Group 1 ) SEQ ID NO:13-15;(Group 2) SEQ ID NO:28-29; (Group 3) SEQ ID NO:40-41; (Group 4) SEQ ID NO:53-54;(Group 5) SEQ ID NO:69-71; (Group 6) SEQ ID NO:92-95; (Group 7) SEQ ID NO: 112-114; (Group 8) SEQ ID NO: 126-127; (Group 9) SEQ ID NO:291-330; (Group 10) SEQ ID NO:355- 360; (Group 11) SEQ ID NO:380-382; (Group 12) SEQ ID NO:394-395; (Group 13) SEQ ID NO:423-428; and (Group 14) SEQ ID NO:542-563;(e) a nucleic acid sequence at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99% or at least 99.9% identical to : (Group 1) SEQ ID NO: 13- 15; (Group 2) SEQ ID NO:28-29; (Group 3) SEQ ID NO:40-41; (Group 4) SEQ ID NO:53-54; (Group 5) SEQ ID NO:69-71; (Group 6) SEQ ID NO:92-95; (Group 7) SEQ ID NO: 112-114; (Group 8) SEQ ID NO: 126-127; (Group 9) SEQ ID NO:291-330; (Group 10) SEQ ID NO:355- 360; (Group 11) SEQ ID NO:380-382; (Group 12) SEQ ID NO:394-395; (Group 13) SEQ ID NO:423-428; and (Group 14) SEQ ID NO:542-563; and(f) a nucleic acid sequence that hybridizes under stringent conditions to : (Group 1) SEQ ID NO: 13-15; (Group 2) SEQ ID NO:28-29; (Group 3) SEQ ID NO:40-41; (Group 4) SEQ ID NO:53-54; (Group 5) SEQ ID NO:69-71; (Group 6) SEQ ID NO:92-95; (Group 7) SEQ ID NO: 1 12-1 14; (Group 8) SEQ ID NO: 126-127; (Group 9) SEQ ID NO:291-330; (Group 10) SEQ ID NO:355-360; (Group 11) SEQ ID NO:380-382; (Group 12) SEQ ID NO:394-395;(Group 13) SEQ ID NO:423-428; and (Group 14) SEQ ID NO:542-563.

[0024] In some embodiments, the isolated or recombinant polynucleotide comprising or consisting of a nucleic acid sequence encoding one or more Cas Type V polypeptides of the disclosure is paired with one or more cognate guide RNA of the disclosure.

[0025] In certain exemplary aspects, provided herein is a Cas Type V gene editing system comprising:(a) one or more polypeptide sequences comprising at least 45%, 46%, 47%, 48%, 49%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, or 65% sequence identity to any one of sequences selected from SEQ ID NO: 334 (No. ID405), SEQ ID NO: 58 (No. ID414), or SEQ ID NO: 564 (No. ID418). SEQ ID NO: 335 (No. ID406), SEQ ID NO: 331 (No. ID411), SEQ ID NO: 20 (No.ID415), and SEQ ID NO: 445 (No. ID419); and(b) one or more polynucleotide sequences comprising a guide RNA, wherein the guide RNA comprises a complementary sequence to that of a targeted polynucleotide sequence.

[0026] In various embodiments, disclosed is a method of modifying a targeted polynucleotide sequence, said method comprising:(a) one or more polypeptide sequences comprising at least 45%, 46%, 47%, 48%, 49%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, or 65% sequence identity to any one of sequences selected from SEQ ID NO: 334 (No. ID405), SEQ ID NO: 58 (No. ID414), or SEQ ID NO: 564 (No. ID418), SEQ ID NO: 335 (No. ID406), SEQ ID NO: 331 (No. ID411), SEQ ID NO: 20 (No.ID415), and SEQ ID NO: 445 (No. ID419); and(b) one or more polynucleotide sequences comprising a guide RNA, wherein the guide RNA comprises a complementary' sequence to that of a targeted polynucleotide sequence; and(c) introducing into a host cell the one or more polypeptide sequences of (a) and the one or more polynucleotide sequences of (b) in a delivery’ vector; wherein the polypeptide sequence is configured to form a ribonucleoprotein complex with the guide RNA, and wherein the ribonucleoprotein complex modifies a targeted polynucleotide sequence.

[0027] In certain preferred embodiments, the method comprises contacting the host cell with a guide RNA, wherein the guide RNA optionally forms a ribonucleoprotein complex with the polypeptide and the guide RNA.

[0028] In various aspects, the present disclosure provides delivery of a Casl2a-based gene editing system described herein Cast 2a in various viral and non-viral vectors. In certain preferred embodiments, the LNP comprises: a) one or more ionizable lipids; b) one or more structural lipids; c) one or more PEGylated lipids; and d) one or more phospholipids.

[0029] In certain embodiments, the LNP comprises one or more ionizable lipids selected from the group consisting of those disclosed herein.

[0030] Also provided herein are pharmaceutical compositions comprising a site-specific modification of a target region of a host cell genome comprising a Cas Type V -based gene editing system described herein Cas Type V comprising one or more Cas Type V polypeptides; one or more cognate guide RNA; and LNP suitable for therapeutic administration.

[0031] In various aspects, provided herein is a method of treating a subject in need thereof, comprising administering to the subject a pharmaceutical composition described herein. In some embodiments, the subject is ameliorated from a diseases or disorders including but not limited to various monogenic diseases or disorders.

[0032] In various embodiments, the disclosure relates to the following numbered paragraphs:1. A genome editing system comprising:a Cas Type V polypeptide or variant thereof, or a nucleic acid sequence encoding a Cas Type V polypeptide or variant thereof; a second nucleic acid sequence encoding a guide RNA; wherein the Cas Type V polypeptide and tire guide RNA form an RNA-protein complex; wherein the genome editing system optionally further comprises a donor nucleic acid sequence capable of modifying a target sequence.2. The genome editing system of paragraph 1, wherein the Cas Type V polypeptide or variant thereof is a polypeptide selected from Table S15A (SEQ ID NO: 334 (No. ID405), SEQ ID NO: 58 (No. ID414), or SEQ ID NO: 564 (No. ID418), SEQ ID NO: 335 (No. ID406), SEQ ID NO: 331 (No. ID411), SEQ ID NO: 20 (No. ID415), and SEQ ID NO: 445 (No. ID419)), or a polypeptide having at least 70%, 75%, 80%, 85%, 90%, 95%, 99%, or 100% sequence identity’ with a polypeptide from Table S15A (SEQ ID NO: 334 (No. ID405), SEQ ID NO: 58 (No. ID414), or SEQ ID NO: 564 (No. ID418). SEQ ID NO: 335 (No. ID406), SEQ ID NO: 331 (No. ID411), SEQ ID NO: 20 (No. ID415), and SEQ ID NO: 445 (No. ID419)).3. The genome editing system of paragraph 1, wherein the Cas 12a polypeptide is encoded by a polynucleotide sequence selected from Table S15B (SEQ ID NO: 365 (No. ID405), SEQ ID NO: 75 (No. ID414). or SEQ ID NO: 565 (No. ID418), SEQ ID NO: 366 (No. ID406), SEQ ID NO: 331 (No. ID411). SEQ ID NO: 30 (No. ID415). or SEQ ID NO: 445 (No. 1D419)), or a polynucleotide having at least 70%, 75%, 80%, 85%, 90%, 95%, 99%, or 100% sequence identity with a polypeptide from Table S15B (SEQ ID NO: 365 (No. ID405), SEQ ID NO: 75 (No. ID414), or SEQ ID NO:565 (No. ID418), SEQ ID NO: 366 (No. ID406), SEQ ID NO: 331 (No. ID411), SEQ ID NO: 30 (No. ID415), or SEQ ID NO: 445 (No. ID419)).4. The genome editing system of paragraph 1, wherein the Cas Type V guide RNA is selected from any Cas Type V guide sequence disclosed in Table S 15C (SEQ ID NO:28-29. 69-71, 355-360, 542-563), or a nucleic acid molecule having at least 70%, 75%, 80%, 85%, 90%, 95%, 99%, or 100% sequence identity with a Cas Type V guide sequence of Table S15C.5. The genome editing system of paragraph 1, wherein the Cas Type V polypeptide or variant thereof is operably fused to an accessory domain.6. The genome editing system of paragraph 5, wherein the accessory domain is a deaminase domain, nuclease domain, reverse transcriptase domain, integrase domain, recombinase domain, transposase domain, endonuclease domain, or exonuclease domain.7. The genome editing system of paragraph 1, wherein the Cas Type V polypeptide or variant thereof is operably fused to a deaminase domain.8. The genome editing system of paragraph 1, wherein the Cas Type V polypeptide or variant thereof is operably fused to a reverse transcriptase domain.9. The genome editing system of paragraph 1, wherein the Cas Type V polypeptide or variant thereof is operably fused to a recombinase domain.10. The genome editing system of paragraph 1, wherein the Cas Type V polypeptide or variant thereof is operably fused to an integrase domain.11. Tire genome editing system of paragraph 1, wherein the Cas Type V polypeptide or variant thereof is operably fused to atransposase domain.12. The genome editing system of paragraph 1, wherein the Cas Type V polypeptide or variant thereof is engineered to have an enhanced genome editing efficiency relative to a wildtype SpCas9.13. The genome editing system of paragraph 12, wherein the enhanced genome editing efficiency comprises at least two to fivefold increase in editing efficiency relative to a wildtype SpCas9.14. Tire genome editing system of any one of the above paragraphs wherein the donor nucleic acid sequence repairs the target region of the genome editing system genome cleaved by the RNA-protein complex.15. The genome editing system of any one of the above paragraphs wherein the nucleic acid sequence encoding the Cas Type V polypeptide and tire guide RNA are transiently expressed in tire host cell genome.16. The genome editing system of any one of the above paragraphs wherein the nucleic acid sequence encoding the Cas Type V polypeptide and the guide RNA are integrated into and expressed from the host cell genome.17. Hie genome editing system of any one of the above paragraphs wherein the nucleic acid sequence encoding the Cas Type V polypeptide and tire guide RNA are integrated into and expressed from a plasmid.1 . The genome editing system of any one of the above paragraphs wherein the genome editing system further comprises a donor nucleic acid sequence to modify a target region of the host cell genome.19. The genome editing system of any one of the above paragraphs wherein administering the system to a host cell results in one or more edits.20. The genome editing system of claim 19, wherein the one or more edits comprises an insertion, deletion, base change / substitution, or inversion, or a combination thereof.21 . The genome editing system of claim 19, wherein the one or more edits comprises a modification in the nucleobase sequence of a target nucleic acid molecule.22. The genome editing system of claim 19, wherein the one or more edits comprises a whole-exon insertion, deletion, or substitution.23. The genome editing system of claim 19, wherein the one or more edits comprises a whole-intron insertion, deletion, or substitution.24. The genome editing system of claim 19, wherein the one or more edits comprises a whole-gene insertion, deletion, or substitution.25. Tire genome editing system of claim 19, wherein the one or more edits comprises an edit to the sequence of a gene or to a region of a gene, e.g., an exon or intron.26. The genome editing system of any one of the above paragraphs wherein the Cas Type V polypeptide recognizes a protospacer-adjacent motif (PAM).27. Tire genome editing system of any one of the above paragraphs wherein the genome editing system installs one or more desired sequence modifications of one or more monogenic disorders or diseases.28. Tire genome editing system of any one of the above paragraphs wherein the genome editing system installs one or more desired epigenetic modifications of one or more monogenic disorders or diseases.29. The genome editing system of any one of the above paragraphs wherein the Cas Type V polypeptide comprises one or more modifications in one or more domains selected from (a) a nuclease domain (e g., RuvC domain) and (b) a PAM-interacting domain.30. Tire genome editing system of any one of the above paragraphs further comprising a delivery vector.31. The genome editing system of paragraph 30, wherein the delivery’ vector is selected from viral vector is selected from a retroviral vector, a lentiviral vector, an adenoviral, an adeno-associated viral vector, vaccinia viral vector, poxviral vector, and herpes simplex viral vector.32. Tire genome editing system of paragraph 30, wherein the delivery vector comprises a non-viral vector selected from cationic liposomes, lipid nanoparticles (LNPs). cationic polymers, vesicles, and gold nanoparticles.33. The genome editing system of any one of the above paragraphs wherein the modification of the target sequence of the host cell genome comprises binding activity, cleavage activity, nickase activity, deaminase activity, reverse transcriptase activity, transcriptional activation activity, transcriptional inhibitory activity, or transcriptional epigenetic activity.34. Tire genome editing system of any one of the above paragraphs wherein any of the nucleic acid molecules — including any guide RNA or donor DNA — comprises one or more chemical modifications selected from 2'-O-Me, 2'-F, and 2'F-ANA at 2'OH; 2'F-4'-Ca-OMe and 2',4'-di-Ca-OMe at 2' and 4' carbons; phosphodiester modifications comprising sulfide-based Phosphorothioate (PS) or acetate-based phosphonoacetate alterations; combinations of the ribose and phosphodiester modifications; lockednucleic acid (LNA), bridged nucleic acids (BNA), S-constrained ethyl (cEt), and unlocked nucleic acid (UNA); modifications to produce a phosphodiester bond between the 2' and 5' carbons (2',5'-RNA) of adjacent RNAs; and a butane 4-carbon chain link between adjacent RNAs.35. The genome editing system of any one of the above paragraphs wherein any guide RNA comprises one or more chemical modifications selected from 2'-0-Me, 2'-F. and 2'F-ANA at 2'OH; 2'F-4'-Ca-OMe and 2’,4'-di-Ca-OMe at 2' and 4' carbons; phosphodiester modifications comprising sulfide-based Phosphorothioate (PS) or acetate-based phosphonoacetate alterations; combinations of the ribose and phosphodiester modifications; locked nucleic acid (LNA), bridged nucleic acids (BNA), S-constrained ethyl (cEt), and unlocked nucleic acid (UNA); modifications to produce a phosphodiester bond between the 2' and 5' carbons (2',5'-RNA) of adjacent RNAs; and a butane 4-carbon chain link between adjacent RNAs.36. Tire genome editing system of any one of the above paragraphs wherein any donor or template DNA comprises one or more chemical modifications selected from 2'-0-Me, 2'-F, and 2'F-ANA at 2'OH; 2'F- 4'-Ca-0Me and 2',4'-di-Ca-OMe at 2' and 4' carbons; phosphodiester modifications comprising sulfide- based Phosphorothioate (PS) or acetate-based phosphonoacetate alterations; combinations of the ribose and phosphodiester modifications; locked nucleic acid (LNA), bridged nucleic acids (BNA), S- constrained ethyl (cEt), and unlocked nucleic acid (UNA); modifications to produce a phosphodiester bond between the 2' and 5' carbons (2',5'-RNA) of adjacent RNAs; and a butane 4-carbon chain link between adjacent nucleotides.37. A method for editing the DNA of a host cell, producing one or more compositions comprising: a Cas Type V polypeptide or a nucleic acid sequence encoding a Cas Type V polypeptide; a second nucleic acid sequence encoding a guide RNA. wherein the second nucleic acid sequence and the Cas Type V polypeptide form an RNA-protein complex; wherein the genome editing system optionally further comprises a donor nucleic acid sequence capable of modifying a target sequence; and introducing tire composition into a host cell; optionally selecting for the host cell comprising the modification or the donor nucleic acid sequence into the host cell genome; and optionally culturing the edited host cells under conditions sufficient for growth.38. The method of paragraph 37, wherein the Cas Type V polypeptide is: operably fused to a nuclease;operably fused to a deaminase; operably fused to a reverse transcriptase; operably fused to a recombinase; operably fused to a transposase; operably fused to a epigenetic effector; or operably fused to any combination of a, b, c. d, e and / or f.39. The method of paragraph 37, further comprising quantifying or characterizing the editing of the target region.40. The method of paragraph 37, wherein tire method provides editing efficiency of greater than 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 99% relative to SpCas9.41. Tire method of paragraph 37, further comprising introducing into the host cell a second donor nucleic acid sequence paired with a second guide RNA to modify the second target region of the host cell genome.42. The method of paragraph 37 further comprising introducing into the host cell at least two desired modification sequences for multiplexing.43. The method of paragraph 37 wherein the method comprises insertion or stable integration of the one or more desired modification sequence into the host cell genome.44. The method of paragraph 37 wherein the host cell genome comprises a chromosome or chromosome and plasmid.45. The method of paragraph 37 wherein the target region is modified by an insertion, deletion or alteration of one or more base pairs at the target region in the host cell genome.46. The method of paragraph 37 wherein the one or more desired modification sequence is selected from one or more sequences associated with one or more monogenic disorders or diseases.47. The method of paragraph 37 wherein the host cell is a primary human cell.48. The method of paragraph 37 wherein the step of introducing into the host cell comprises a delivery vector operably linked to the genome editing system.49. Hie method of paragraph 48 wherein the delivery vector is selected from viral vector is selected from a retroviral vector, a lentiviral vector, an adenoviral, an adeno-associated viral vector, vaccinia viral vector, poxviral vector, and herpes simplex viral vector.50. The method of paragraph 48 wherein the delivery vector comprises a non-viral vectors selected from cationic liposomes, lipid nanoparticles (LNPs), cationic polymers, vesicles, and gold nanoparticles.51. The method of paragraph 37 wherein the editing method results in enhanced editing efficiency and / or low cytotoxicity.52. A gene editing construct comprising: an Cas Type V domain; (b) a reverse transcriptase domain; (c) a transcriptional modulating polypeptide; (d) a recombinase domain; (e) a transposose domain; or (f) any combination of a, b. c, d, e, or f.53. Tire gene editing construct of claim 52, further comprising a donor nucleic acid sequence capable of modifying a target sequence; and

[0033] In various aspects, the target region is modified by an insertion, deletion or alteration of one or more base pairs at the target region in the host cell genome.

[0034] In various embodiments, one or more desired modification sequence is selected from one or more sequences associated with one or more monogenic disorders or diseases.

[0035] In certain preferred embodiments, the methods and compositions provide editing efficiency of greater than 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 99% relative to SpCas9.

[0036] Related aspects provide for the use of a Cas Type V-based gene editing system described herein Cas Type Vin the application for plants, yeast, bacteria, and fungi and desired bioindustrial applications for producing value-added components in such systems in a recombinant manner.

[0037] Accordingly, it is an object of the invention not to encompass within the invention any previously known product, process of making the product, or method of using the product such that Applicants reserve the right and hereby disclose a disclaimer of any previously known product, process, or method. It is further noted that tire invention does not intend to encompass within the scope of tire invention any product, process, or making of the product or method of using tire product, which does not meet the written description and enablement requirements of tire USPTO (35 U.S.C. § 112, first paragraph) or the EPO (Article 83 of tire EPC), such that Applicants reserve the right and hereby disclose a disclaimer of any previously described product, process of making the product, or method of using the product. It may be advantageous in the practice of the invention to be in compliance with Art. 53(c) EPC and Rule 28(b) and (c) EPC. All rights to explicitly disclaim any embodiments that are the subject of any granted patent(s) of applicant in the lineage of this application or in any other lineage or in any prior filed application of any third party is explicitly reserved. Nothing herein is to be constmed as a promise.DESCRIPTION OF THE DRAWINGS

[0038] FIGs. 1 A-l C are schemes depicting the predicted stem loop structures of crRNA sequences of the present disclosure corresponding to Group 1 sequences.

[0039] FIGs. 2A-2B are schemes depicting the predicted stem loop structures of crRNA sequences of the present disclosure corresponding to Group 2 sequences.

[0040] FIGs. 3A-3B are schemes depicting the predicted stem loop structures of crRNA sequences of the present disclosure corresponding to Group 3 sequences.

[0041] FIGs. 4A-4B are schemes depicting the predicted stem loop structures of crRNA sequences of the present disclosure corresponding to Group 4 sequences.

[0042] FIGs. 5A-5C are schemes depicting the predicted stem loop structures of crRNA sequences of the present disclosure corresponding to Group 5 sequences.

[0043] FIGs. 6A-6D are schemes depicting the predicted stem loop structures of crRNA sequences of the present disclosure corresponding to Group 6 sequences.

[0044] FIGs. 7A-7C are schemes depicting the predicted stem loop structures of crRNA sequences of the present disclosure corresponding to Group 7 sequences.

[0045] FIGs. 8A-8B are schemes depicting the predicted stem loop structures of crRNA sequences of the present disclosure corresponding to Group 8 sequences.

[0046] FIGs. 9A-9NN are schemes depicting tire predicted stem loop structures of crRNA sequences of the present disclosure corresponding to Group 9 sequences.

[0047] FIGs. 10A-10F are schemes depicting tire predicted stem loop structures of crRNA sequences of the present disclosure corresponding to Group 10 sequences.

[0048] FIGs. 11A-11C are schemes depicting the predicted stem loop structures of crRNA sequences of the present disclosure corresponding to Group 11 sequences.

[0049] FIGs. 12A-12B are schemes depicting the predicted stem loop structures of crRNA sequences of the present disclosure corresponding to Group 12 sequences.

[0050] FIGs. 13A-13F are schemes depicting tire predicted stem loop structures of crRNA sequences of the present disclosure corresponding to Group 13 sequences.

[0051] FIGs. 14A- 14V are schemes depicting the predicted stem loop structures of crRNA sequences of the present disclosure corresponding to Group 14 sequences.

[0052] FIG. 15 is a phylogenetic tree illustrating the location of PAM sequences added at each protein, as described in Example 9. The phylogenetic tree was generated using Geneious Prime 2022.1.1 implementation of FastTree on Muscle multiple sequence alignment of selected protein sequences. PAM sequence weblogos were generated using WebLogo 3 web application from PFMs.

[0053] FIG. 16 is a blot showing cleavage products of genomic target DNMT] visualized on 2 % agarose gel. Editing efficiency values for LbaCasl2a and each ortholog were calculated using Image J software, in accordance with Example 10.

[0054] FIG. 17 is a blot showing cleavage products of genomic target RUNX 1 visualized on 2 % agarose gel. Editing efficiency values for LbaCasl2a and each ortholog were calculated using ImageJ software, in accordance with Example 10.

[0055] FIG. 18 is a blot showing cleavage products of genomic target SCN1A visualized on 2 % agarose gel. Editing efficiency values for LbaCasl2a and each ortholog were calculated using ImageJ software, in accordance with Example 10.

[0056] FIG. 19 is a blot showing cleavage products of genomic target FANCF site 2 visualized on 2 % agarose gel. Editing efficiency values for LbaCasl2a and each ortholog were calculated using ImageJ software, in accordance with Example 10.

[0057] FIG. 20 is a blot showing cleavage products of genomic target FANCF site 1 visualized on 2 % agarose gel. Editing efficiency values for LbaCasl2a and each ortholog were calculated using ImageJ software, in accordance with Example 10.

[0058] FIG. 21 is a graph showing a comparison of Casl2a orthologs activity on different targets (n > 3). Results are calculated from T7 endonuclease assay, in accordance with Example 11.

[0059] FIG. 22A is a graph showing genome editing efficiency results for ID405, ID414, ID418, LbaCasl2a depicted as indels frequency at RUNX1 and SCN1A target sites as detennined by deep-sequencing in accordance with Example 12.

[0060] FIG. 22B is a graph showing genome editing efficiency results for 1D405, 1D414, 1D418. LbaCasI2a depicted as indels frequency at RUNX1, SCN1A, DNMT1, FANCF site 1, and FANCF site 2, as determined by deep-sequencing in accordance with Example 12.

[0061] FIGs. 23A-23E are tables showing the top five most common editing outcomes observed in deep sequencing data of ID405, ID414, ID418 and LbaCas 12a genomic targets in RUNX1 (FIG. 23A), SCN1A (FIG. 23B). DNMT1 (FIG. 23C), FANCF Site 1 (FIG. 23D), and FANCF Site 2 (FIG. 23E) genes as compared to reference sequences.

[0062] FIG. 24 is a graph showing genome editing efficiency results depicted as indels frequency as determined by deep-sequencing as described in Example 12.

[0063] FIG. 25 is a set of tables showing the top 5 most common editing outcomes observed in deep sequencing data of ID428 and ID433 genomic targets exhibiting low but observable editing as compared to reference sequences as described in Example 12.

[0064] FIG. 26 is a set of blots showing endonuclease activity comparison between SpyCas9, LbaCas 12a. ID405, and ID414. Blue arrows mark cleavage products of LbaCasl2a, ID405, and ID414 nucleases; black arrows mark cleavage products of SpyCas9 nucleases. Percentages above each gel well show the editing number detennined from the gel using ImageJ software. See Example 13 for further details.

[0065] FIG. 27 is a set of blots showing endonuclease activity comparison between SpyCas9. LbaCasl2a, and ID418. Blue arrows mark cleavage products of LbaCasl2a and ID418 nucleases; black arrows mark cleavage products of SpyCas9 nucleases. Percentages above each gel well show the editing number determined from the gel using Image J software. See Example 13 for further details.

[0066] FIG. 28A is a set of blots showing cleavage products of genomic target PCSK9 visualized on 2 % agarose gel in the presence of various ID405 mutants. Editing efficiency values indicated above gel wells for LbaCasl2a and each ID405 mutant were calculated using ImageJ software. See Example 13 for further details.

[0067] FIG. 28B is a set of blots showing cleavage products of genomic target CISH visualized on 2 % agarose gel in the presence of various ID405 mutants. Editing efficiency values indicated above gel wells for LbaCasl2a and each ID405 mutant were calculated using ImageJ software. See Example 13 for further details.

[0068] FIG. 28C is a set of blots showing cleavage products of genomic target TTR visualized on 2 % agarose gel in the presence of various ID405 mutants. Editing efficiency values indicated above gel wells for LbaCasl2a and each ID405 mutant were calculated using ImageJ software. See Example 13 for further details.

[0069] FIG. 28D is a set of blots showing cleavage products of genomic target PCSK9 visualized on 2 % agarose gel in the presence of various ID414 mutants. Editing efficiency values indicated above gel wells for LbaCasl2a and each ID414 mutant were calculated using ImageJ software. See Example 13 for further details.

[0070] FIG. 28E is a set of blots showing cleavage products of genomic target CISH visualized on 2 % agarose gel in the presence of various ID414 mutants. Editing efficiency values indicated above gel wells for LbaCasl2a and each ID414 mutant were calculated using ImageJ software. See Example 13 for further details.

[0071] FIG. 28F is a set of blots showing cleavage products of genomic target TTR visualized on 2 % agarose gel in the presence of various ID414 mutants. Editing efficiency values indicated above gel wells for LbaCasl2a and each ID414 mutant were calculated using ImageJ software. See Example 13 for further details.

[0072] FIG. 28G is a set of blots showing cleavage products of genomic target BCLlla visualized on 2% agarose gel in the presence of various ID405 mutants. Editing efficiency values for LbaCasl2a and each ID405 mutant were calculated using ImageJ software.

[0073] FIG. 28H is a set of blots showing cleavage products of genomic target HBG1 visualized on 2 % agarose gel in the presence of various ID405 mutants. Editing efficiency values for LbaCasl2a and each ID405 mutant were calculated using Image J software.

[0074] FIG. 281 is a set of blots showing cleavage products of genomic target BCL1 la visualized on 2 % agarose gel in the presence of various ID414 mutants. Editing efficiency values for LbaCasl2a and each ID414 mutant were calculated using Image J software.

[0075] FIG. 28 J is a set of blots showing cleavage products of genomic target HBG1 visualized on 2 % agarose gel in the presence of various ID414 mutants. Editing efficiency values for LbaCasl2a and each ID414 mutant were calculated using Image J software.

[0076] FIG. 28K is a set of blots showing cleavage products of genomic target PCSK9 visualized on 2 % agarose gel. Editing efficiency values for LbaCasl2a and each ID418 mutant were calculated using ImageJ software.

[0077] FIG. 28L is a set of blots showing cleavage products of genomic target CISH visualized on 2 % agarose gel in the presence of various ID418 mutants. Editing efficiency values for LbaCasl2a and each ID418 mutant were calculated using ImageJ software.

[0078] FIG. 28M is a set of blots showing cleavage products of genomic target CISH visualized on 2 % agarose gel in the presence of various ID418 mutants. Editing efficiency values for LbaCasl2a and each ID418 mutant were calculated using ImageJ software.

[0079] FIG. 28N is a set of blots showing cleavage products of genomic target BCLlla visualized on 2 % agarose gel in the presence of various ID418 mutants. Editing efficiency values for LbaCasl2a and each ID418 mutant were calculated using ImageJ software.

[0080] FIG. 280 is a set of blots showing cleavage products of genomic target HBG1 visualized on 2 % agarose gel in the presence of various ID418 mutants. Editing efficiency values for LbaCasl2a and each ID418 mutant were calculated using ImageJ software.

[0081] FIG. 29A is a graph showing a comparison of ID405 wild-type and ID405-1 mutant editing efficiency on different targets (n > 3). Targets are BCL1 la, CISFI, HBG1, PCSK9, and TTR. Results were calculated from T7 endonuclease assay data. For each gene target in the cluster of bar graphs, beginning on the left-most side of each clusture, the bars correspond to ID405, ID405-1, LbaCasl2a, and AsCasl2a Ultra. Note that no editing activity was observed for the BCL1 la and TTR targets (no left-most bar corresponding to 1D405 activity). See Example 13 for further details.

[0082] FIG. 29B is a graph showing a comparison of ID414 wild-type and ID414- 1 mutant activity on different targets (n > 3). Targets are BCL1 la, CISH, HBG1, PCSK9, and TTR. Results were calculated from T7 endonuclease assay data. For each gene target in the cluster of bar graphs, beginning on the left-most sideof each clusture, the bars correspond to ID414, ID414-1, LbaCasl2a, and AsCasl2a Ultra. Note that no editing activity was observed for the BCL1 la and HBG1 targets (no left-most bar corresponding to ID414 activity). See Example 13 for further details.

[0083] FIG. 30 depicts a phylogenetic tree of relationships among each of the Cast 2a ortholog sequences presented in Table S15A versus the canonical LbCasl2a sequence of SEQ ID NO: 1382 (provided in Appendix A, Section Q). The phylogenetic tree was calculated using the Clustal Omega multiple sequence alignment online tools available at EMBL (the European Molecular Biology Laboratory).

[0084] FIG. 31 shows a sequence alignment among each of the Casl2a orthologs provided in Table S15A and tire canonical LbCasl2a sequence of SEQ ID NO: 1382. Bolded-underlined residues are marked with an asterisk (“*”) and denote a fully conserved amino acid residue present in all of the aligned sequences at that alignment position. The amino acid residue positions marked with a colon (“:”) denote aligned amino acid residues which are highly similar although not identically conserved. The highly similar residues are those where the substitutions among the sequences have strongly similar properties. The amino acid residue positions marked with a period (“.”) denote aligned amino acid residues which are moderately similar. The highly similar residues are those where the substitutions among tire sequences have strongly similar properties. The moderately similar residues are those where the substitutions among the sequence have weakly similar properties. The underlined regions are referred to as “highly conserved regions" and include (a) at least one fully conserved residue, and (b) at least one highly similar or moderatly similar residue. See Appendix A, Section Q for further description.

[0085] FIG. 32 is a graph showing genome editing efficiency results depicted as indels frequency as determined by deep-sequencing (n > 4), as described in Example 15.

[0086] FIG. 33 is a set of tables showing the top 5 most common editing outcomes observed in deep sequencing data of ID405, ID414 and ID418 wild type and mutant proteins as well as LbCasl2a and AsCasl2a Ultra genomic targets as compared to reference sequence in HBG1 gene.

[0087] FIG. 34 is a set of blots showing cleavage products of genomic target HBG1 visualized on 2 % agarose gel. Editing efficiency values for LbaCasl2a, AsCasl2a Ultra and each ID405 mutant were calculated using Image J software, wt ID405 - wild-type ID405 control, ID405-1 - mutant from the first targeted mutation wave (D169R).

[0088] FIG. 35 is a set of blots showing cleavage products of genomic target TTR visualized on 2 % agarose gel. Editing efficiency values for LbaCasl2a. AsCasl2a Ultra and each 1D405 mutant were calculated using ImageJ software, wt ID405 - wild-type ID405 control, ID405-1 - mutant from the first targeted mutation wave (D169R).

[0089] FIG. 36 is a graph comparing ID405-1 mutants from the first targeted mutation wave, wild type ID405, and ID405 second targeted mutation wave nucleases. NHEJ percentage calculated from T7 Endonuclease I assay results, generated using Image J software (n = 6).

[0090] FIG. 37 is a set of blots showing cleavage products of genomic target HBG1 visualized on 2 % agarose gel. Editing efficiency values for LbaCasl2a. AsCasl2a Ultra and each ID418 mutant were calculated using Image J software, wt ID418 - wild-type ID418 control.

[0091] FIG. 38 is a set of blots showing cleavage products of genomic target TTR visualized on 2 % agarose gel. Editing efficiency values for LbaCasl2a, AsCasl2a Ultra and each ID418 mutant were calculated using Image J software, wt ID418 - wild-type ID418 control.

[0092] FIG. 39 is a set of blots showing cleavage products of genomic target HBG1 visualized on 2 % agarose gel. Editing efficiency values for LbaCasl2a. AsCasl2a Ultra and each ID414 mutant were calculated using ImageJ software, wt ID448 - wild-type ID414 control: ID414-T154R and ID414-S802L - mutants from the first mutation round (ID414-1 and ID414-6, respectively).

[0093] FIG. 40 is a set of blots showing cleavage products of genomic target TTR visualized on 2 % agarose gel. Editing efficiency values for LbaCasl2a, AsCasl2a Ultra and each ID414 mutant were calculated using ImageJ software, wt ID448 - wild-type ID414 control; ID414-T154R and ID414-S802L - mutants from the first mutation round (ID414-1 and ID414-6, respectively).

[0094] FIG. 41 is a graph showing genome editing efficiency results depicted as indels frequency as determined by deep-sequencing (n > 6).

[0095] FIG. 42 is a set of tables showing the top 5 most common editing outcomes observed in deep sequencing data of ID405 wild type and mutant proteins as well as AsCasl2a Ultra genomic targets as compared to reference sequence in HBG1 gene.

[0096] FIG. 43 is a set of blots showing cleavage products of genomic target HBG1 visualized on 2 % agarose gel. Editing efficiency values for Alt-R A.s. Casl2a (Cpfl) and each purified Casl2a were calculated using ImageJ software.

[0097] FIG. 44 is a set of blots showing cleavage products of genomic target TTR visualized on 2 % agarose gel. Editing efficiency values for each purified Casl2a were calculated using ImageJ software.

[0098] FIG. 45 is a graph comparing purified mutant nucleases and Cast 2a variant editing efficiency in HEK293T on HBG1 genomic target. Results are calculated from T7 Endonuclease I assay.

[0099] FIG. 46 is a graph comparing purified mutant nucleases and Casl2a variant editing efficiency in HEK293T on TTR genomic target. Results are calculated from T7 Endonuclease I assay.

[0100] FIG. 47 represents a process for improving a programmable nuclease, such as, but not limited to a type V ortholog (e.g., ID405), by introducing one or more arginine residues in place of a naturally-occurringnon-arginine residues at positions predicted to have an impact on the interaction between the ID405 amino acid sequence and a guide RNA (prediction based on computational methodology of Example 16). The substitutable positions are then ranked in accordance w ith the magnitude of stabilizing energetics associated with arginine replacement at any particular position as outlined in Example 16. Variant type V editors are then engineered and tested to determine which variants provide for any improvement in editing efficiency. Results are shown in Example 16.

[0101] FIG. 48 compares the interaction between Lys292 in ID405-1 with a replaced Arg292 with the guide RNA showing a more favorable interaction with the Arg292 residue, as described in Example 16.

[0102] FIG. 49 compares the interaction betw een Gln492 in ID405-1 with a replaced Arg492 w ith the guide RNA showing a more favorable interaction with the Arg492 residue, as described in Example 16.

[0103] FIG. 50 compares the interaction between Asn770 in ID405-1 with a replaced Arg770 with the guide RNA showing a more favorable interaction with the Arg770 residue, as described in Example 16.

[0104] FIG. 51 provides a graph showing the editing efficiency in primary HSPCs of select ID405-1 variants with arginine substitution mutations. WT designated ID405-1. Each variant enzyme, e.g., S972R, designates ID405-1 with a particular arginine substitution (e.g., S972R designates that serine at position 972 relative to ID405-1 has been substituted with an arginine). See Example 16 for details.

[0105] FIG. 52 is a schematic of ID405-1 variants comprising one or more NLS and optional linkers to improve 1D405-1 editing activity, in accordance with Example 17.

[0106] FIG. 53 is a schematic that shows that modified ID405-1 variant containing various linkers and NLSs in various configurations at the N- and / or C-termini improves editing as compared to baseline Construct 1 in primary HSPCs by up to a 5-fold increase. See Example 17 for details.

[0107] FIG. 54A is a schematic representation of the selection scheme for ID405 used in Example 18. Hie ID405 plasmid library was transformed into E. coll cells. Hie most active variants cut ccdB encoding plasmid which led to its subsequent degradation. This allowed for E. colt to grow on selective media promoting ccdB induction. Cells that receive plasmid with WT or non-fiinctional / less active ID405 mutant died due to activity of ccdB.

[0108] FIG. 54B is a schematic of the ID405 containing plasmid pACY CDuet_ID405-HBGl-crRNA-TTR- HDV used for directed evolution.

[0109] FIG. 54C is a schematic of the domain structure of ID405. Domains are labeled in the protein schematic. Marked segments above the protein denote regions selected for mutagenesis.

[0110] FIG. 55 is a set of photos of test assays evaluating the performance of ID405 constructs against ccdB containing selection plasmids with TCTT and TTGG PAM sequences upstream of the HBG1 target.

[0111] FIG. 56A is a scatter plot summarizing variant frequency in N-terminus library under the selection against TCTT PAM. Nucleotide positions denote the amplified sequence used for library preparation.

[0112] FIG. 56B is a scatter plot summarizing variant frequency in N-terminus library under the selection against TTGG PAM. Nucleotide positions denote the amplified sequence used for library preparation.

[0113] FIG. 56C is a scatter plot summarizing variant frequency in N-terminus library under the selection against TTGG PAM under the presence of 0. ImM IPTG. Nucleotide positions denote the amplified sequence used for library preparation.

[0114] FIG. 56D is a scatter plot summarizing variant frequency in C-terminus library under the selection against TCTT PAM. Nucleotide positions denote the amplified sequence used for library’ preparation.

[0115] FIG. 56E is a scatter plot summarizing variant frequency in C-terminus library under the selection against TTGG PAM. Nucleotide positions denote the amplified sequence used for library preparation.

[0116] FIG. 56F is a scatter plot summarizing variant frequency in C-terminus library under the selection against TTGG PAM under the presence of 0. 1 mM IPTG. Nucleotide positions denote the amplified sequence used for library? preparation.

[0117] FIG. 56G is a scatter plot summarizing variant frequency in N-2 library’ under the selection against TCTT PAM. Nucleotide positions denote the amplified sequence used for library preparation.

[0118] FIG. 56H is a scatter plot summarizing variant frequency in N-2 library under the selection against TTGG PAM. Nucleotide positions denote the amplified sequence used for library? preparation.

[0119] FIG. 561 is a scatter plot summarizing variant frequency in N-2 library under the selection against TTGG PAM under the presence of 0.1 mM IPTG. Nucleotide positions denote the amplified sequence used for library’ preparation.

[0120] FIG. 56J is a scatter plot summarizing variant frequency in C-2 library under the selection against TCTT PAM. Nucleotide positions denote the amplified sequence used for library preparation.

[0121] FIG. 56K is a scatter plot summarizing variant frequency in C-2 library under the selection against TTGG PAM. Nucleotide positions denote the amplified sequence used for library? preparation.

[0122] FIG. 56L is a scatter plot summarizing variant frequency in C-2 library? under the selection against TTGG PAM under the presence of 0.1 mM IPTG. Nucleotide positions denote the amplified sequence used for library preparation.

[0123] FIG. 56M is a scatter plot summarizing variant frequency in a repeated selection of N-2 library under the selection against TTGG PAM using a greater number of E. colt transformants in round 1. Nucleotide positions denote the amplified sequence used for library preparation.

[0124] FIG. 56N is a scatter plot summarizing variant frequency in a repeated selection of N-2 library under the selection against TTGG PAM under the presence of 0. 1 mM IPTG using a greater number of E. coli transformants in round 1. Nucleotide positions denote the amplified sequence used for library preparation.

[0125] FIG. 57A, FIG. 57B, and FIG. 57C show the computationally predicted optimal 2’-0Me guide modification sites in guide RNA for lbCasl2a over the invariant region of nucleotides 1-20 (FIG. 57A), in guide RNA for asCasl9 over the invariant region of nucleotides 1-19 (FIG. 57B). and in guide RNA for fhCasl2 over the scaffold region of nucleotides 1-19 (FIG. 57C). Lower case letters (boxes with white background fill) denote sites selected by the algorithm that have 2’-0Me modifications. Upper case letters (boxes with grey background fill) denote nucleotides sites that are not modified. Scores next to each sequence (one per row) denote the potential loss of hydrogen bonding interactions were that position to be modified with a 2’-OMe modification. Scores over each nucleotide (one per column) denote tire averaged hydrogen bond interaction calculated for that guide nucleotide position. Three exemplary modified guides with varying scores (high; med; low) are also given for each Type V family member (indicated by black stars). The top row guide sequence is modified at every position with 2’-0Me modifications.

[0126] FIG. 58A, FIG. 58B shows optimal 2-OMe guide modification sites in spacer region at nucleotides 20-39 of the asCasl2 guide sequence (FIG. 58A) and at nucleotides 20-39 of the fhCasl2 guide sequence (FIG. 58B). Lower case letters (boxes with white background fill) denote sites selected by the algorithm that may have 2’-OMc modifications. Upper case letters (boxes with grey background fill) denote nucleotides sites that are not modified. Scores next to each sequence (one per row) denote the potential loss of hydrogen bonding interactions due modifications. Scores over each nucleotide (one per column) denote the averaged hydrogen bond interaction calculated for that guide nucleotide. Three-four exemplary modified guides with varying scores are also given for each Type V family member (indicated by black stars).

[0127] FIGs. 59A-59B illustrate results of MOE analysis performed with Cas 12a guide bound protein structure. Based on MOE structural protocol, nucleotide positions are identified where the 2’-OH of gRNA nucleotide is making contact with Cas 12 protein. Dots indicate sum of all interactions at those positions between the 2 ’OH group of a given nucleotide and either the guide itself or with the protein and represent positions that ought not to be modified.

[0128] FIGs. 60A-60C illustrate results of MOE analy sis performed with Cas 12a guide bound protein structure. Based on MOE structural protocol, nucleotide positions are identified where the 2 ’-OH of gRNA nucleotide is making contact with Cas 12 protein. Dots indicate sum of all interactions at those positions between the 2’OH group of a given nucleotide and either the guide itself or with the protein and represent positions that ought not to be modified.

[0129] FIG. 61 illustrates, based on the computation methodology of Example 19, nucleotide sites along the length of a type V guide RNA which may permissibly be modified with 2’-OMe modifications.

[0130] FIG. 62, seven different chemical modification patterns are designed for gRNA of three different targets (target 1= PCSK9 gene; target 2 = B2M gene; target 3 = BCL11A gene). The BCL11A target is the BCL11A binding site in tire HBG1 / 2 promoters. Targeting this site with a CRISPR-Cas editor and guide RNA can inactivate the binding site of the BCL11A transcriptional repressor, thereby increasing transcription of HBG1 / 2 genes, which results in increased production of HBG1 / 2 subunits, and consequently, increased production of HbF. Each of these three guides have same direct repeat sequence (UGAAUUUCUACUGUUGUAGAU; SEQ ID NO: 1744) but have different spacer length for unique DNA targets. In every guide, first two and last two phosphates have been converted to phosphorothioate. According to MOE, position 3 may not be suitable for phosphorothioate modification (FIG. 63 and 64) and hence kept unmodified. For this work, it was decided to work with 405-1 gRNA which has UG dinucleotide added before Casl2 guide sequence. For this work, these two nucleotides are assigned with -2 and -1 number. Following nucleotide sequence is identical with Cas 12a gRNA direct repeat (AAUUUCUA . .. .) and they are assigned with 1, 2, 3 ... .as nucleotide identification number.

[0131] FIG. 63. Shows the permissible sites to be modified based on detecting the strength of phosphate interactions for FnCasl2a / guide complexes.

[0132] FIG. 64. Shows the permissible sites to be modified based on detecting the strength of phosphate interactions for Casl2a / guide complexes.

[0133] FIG. 65. Summary of modified guides tested in vitro in Example 20. The modified guides Modi, Mod4, Mod5, Mod6, and Mod7 were tested in accordance with the methodology depicted in FIG. 66.

[0134] FIG. 66. Methodology fortesting modified guide RNAs as detailed in Example 20.

[0135] FIG. 67. Results showing high editing rates in primary HSCs at a clinically relevant locus (up to 77% editing at hB2M locus) as outlined in Example 20.

[0136] FIG. 68 Schematic showing an RNP complex (based on Cas9) wherein the nuclease component comprises an NLS (PKKKRKV)(SEQ ID NO:584). The complex forms in the nuclease and relocates to the nuclease as facilitated by the NLS.

[0137] FIG. 69. Depicts the concept of a PNA-NLS probe to couple an NLS directly to a guide RNA. Here, the exemplary’ PNA is 9 residues in length which is joined through an optional linker to one or more NLS.

[0138] FIG. 70. Depicts the concept of a PNA-NLS probe to couple an NLS directly to a guide RNA at an additional sequence element added to the guide RNA referred to as the PNA binder element. This may be configured at either end of the guide RNA. Here, the exemplary PNA is 9 residues in length which is joined through an optional linker to one or more NLS.

[0139] FIG. 71 A. An exemplary guide RNA targeting TTR gene and having the sequence from 5’ to 3’ of SEQ ID NO: 2585, wherein the bolded sequence as shown in Appendix B, SEQ ID NO:2585, hybridizes to the PNA having the nucleic acid sequence AGCCACGAAAA (SEQ ID NO:2305) fused to the NLS peptide PKKKRKV (SEQ ID NO:584).

[0140] FIG. 71B. An exemplary guide RNA targeting TTR gene and having the sequence from 5 ' to 3’ of SEQ ID NO: 2585, wherein the bolded sequence as shown in Appendix B, SEQ ID NO:2585, hybridizes to the PNA having the nucleic acid sequence CCACGAAAA fused to the NLS peptide PKKKRKV (SEQ ID NO:584).

[0141] FIG. 71C. An exemplary guide RNA targeting any gene and having the sequence from 5’ to 3' of SEQ ID NO: 2586, wherein the bolded sequence as shown in Appendix B, SEQ ID NO:2586, hybridizes to the PNA having tire nucleic acid sequence AGCCACGAAAA (SEQ ID NO:2305) fused to the NLS peptide PKKKRKV (SEQ ID NO:584). where in “N” designates any nucleotide and will depend upon the target sequence being targeted by the guide RNA.

[0142] FIG. 72. A graph of % indel formation by 405-1 nuclease with various modified gRNAs in a human HSPC donor. These results show that RNA extensions and DNA extensions can improve editing at the HBG locus 3-10 fold in an all RNA format relative to the mod6 benchmark using the 405-1 nuclease.

[0143] FIG. 73. A graph of % indel formation by 405-1 nuclease with various modified gRNAs in a human HSPC donor. These results show that RNA extensions and DNA extensions do not improve editing at the B2M locus in an all RNA format relative to the mod6 benchmark using the 405-1 nuclease.

[0144] FIG. 74. Shows a graph of % indel formation by 405-1 nuclease with various modified gRNAs in a human HSPC donor at the HBG locus. These results show that RNA and DNA extensions are not compatible with the Type V mod6 schema but guides modified with mod7 schema can tolerate RNA extensions.

[0145] FIG. 75. Modification 6 schema. This figure outlines the positions of 2’-0Me modifications and phosphothiorate modifications in gRNAs with the modification 6 schema.Light grey and uppercase letters = Unmodifed; Dark grey and lowercase letters = modified with phosphothiorate (or in the alternative, indicated below with a post-nt asterisk); White and lowercase letters = modified with 2’-0Me. The sequences of FIG. 75 are: gRNA0260: u*g*AAUuuCUacUguuGuagaUcCUUGUCaagGcUAUUGGU*c*a (SEQ ID NO: 2232) (aka “mod 6” guide); gRNA0390: u*g*AAUuuCUacUguuGuaGAucCUUGUCaaggcuAUUGGu*c*a (SEQ ID NO: 2233); gRNA0391: u*g*AAUUuCUacUguuGuaGAUcCUUGUCaaggCUAUUGGu*c*a (SEQ ID NO:2234); gRNA0392: u*g*AAUUuCUacUguuGUAGAUcCUUGUCaagGCUAUUGGu*c*a (SEQ ID NO:2235), wherein A, G, U, C = unmodified RNA nucleotide; a, g, u, c = 2'-O-Methyl-nucleotide; * = Phosphorothioate linkage

[0146] FIG. 76. Shows a graph of % indel formation by 405-1 nuclease with various modified gRNAs in a human HSPC donor at the HBG locus. These results show that decreasing the proportion of modification in a gRNA can increase editing 4-5 fold in vitro in CD34+ human HSPCs.

[0147] FIG. 77. Shows a graph of % indel formation at the HBG locus in human CD34+ HSPCs by gRNAs with 3’ extensions that create a hairpin with the spacer region of the gRNA. The results show 3’ extensions creating hairpins with a type V spacer (backfolds) increased editing above the mod6 benchmark gRNA at the HBG locus in human HSPCs. gRNA0260 is “mod6” having the sequence and modification scheme of u*g*AAUuuCUacUguuGuagaUcCUUGUCaagGcUAUUGGU*c*a (SEQ ID NO: 680) (aka “mod 6” guide), wherein A, G, U, C = unmodified RNA nucleotide; a, g, u, c = 2'-O-Methyl-nucleotide; * = Phosphorothioate linkage.

[0148] FIG. 78. Shows a graph of % indel formation by 405-1 with modified gRNAs across three different human HSPC donors at the HBG locus. The results showed that the effect of modified gRNAs on indel formation is consistent across HSPC donors using 405-1 3x NLS nuclease. For each cluster of graphs, the samples appear in the same ordering of (a) Donor 310, (b) Donor 309, and (c) Donor 308.

[0149] FIG. 79. Shows a graph of % indel formation at the HBG locus by modified gRNAs across 405-1 protein variants in human hematopoietic stem cells. The results show that the addition of the point mutant K292R and a 3X NLS increased indel fonnation across all modified gRNAs tested. gRNAs containing a 5‘ DNA extension of a 15bp randomized sequence improved indel formation by up to ~8 fold. Other modified gRNAs tested show more modest increases in indel formation. For each cluster of graphs, the samples appear in the same ordering of (a) 405-1, (b) 405-1 3x NLS, (c) 4051- E1014R + 3xNLS, and (d) 405-1 K292R + 3xNLS.

[0150] FIG. 80 (A) A schematic representation of luciferase mRNA with miRNA target sites (ts) inserted in the 3’ UTR. Luciferase mRNA consists of 5’ cap. 5' UTR. luciferase coding sequence, 3’ UTR followed by a poly A tail. A single copy (xl) or three copies (x3) of miRNA ts were inserted at the 3’ end of the 3’ UTR. An alternative (alt) insertion site is located at the 5’ end of the 3’ UTR, near the luciferase coding sequence. When combining two miRNA target sites in the same UTR, the first miRNA target site is inserted at the ‘alt’ location, and the second miRNA target site is inserted at the 3’ end of the UTR. (B) Incorporation of miR-122 ts into the 3’ UTR of luciferase mRNA results in the suppression of encoded protein in the Huh7 hepatocyte cell line, where miR-122 is expressed. A single copy of the target site at the end of 3’ UTRnear the poly A tail is sufficient to suppress luciferase expression by 5-fold compared to the control. Increasing the number of copies to three does not further enhance suppression. However, when the target site is inserted at the 5’ end of the UTR, near the coding sequence, suppression is enhanced to 10-fold. In contrast, insertion of the target site of miR-142, which is not expressed in hepatocyte cell line, has no impact on luciferase expression by itselfand does not influence miR-122 mediated suppression when both target sites are inserted together on the same UTR.

[0151] FIG. 81 depicts (A) A schematic structure of luciferase mRNA with miRNA target sites (ts) inserted in the 3’ UTR. Luciferase mRNA consists of 5' cap, 5 ’UTR, luciferase coding sequence, 3' UTR followed by poly A tail. In this structure, three copies of miRNA target site are inserted at the 5‘ end of 3‘ UTR, near luciferase coding sequence. When combining two miRNA target sites in the same UTR, the first miRNA target site is inserted near the luciferase coding sequence, and the second miRNA target site is inserted at the 3’ end of UTR. (B) Cell type-specific suppression of luciferase expression was achieved by incorporating appropriate miRNA ts in the 3’ UTR. miR-122 ts inserted in the 3’ UTR of luciferase mRNA led to the suppression of protein expression in a hepatocyte cell line, where miR-122 is expressed (left), while no suppression was observed in a monocyte cell line, where miR-122 is not expressed (right). Similarly, miR- 142 ts suppressed luciferase expression exclusively in the monocyte cell line, where miR-142 is expressed. Insertion of miR-122 and miR-142 target sites in the same UTR inhibits luciferase expression in both hepatocyte and monocyte cell line.

[0152] FIG. 82 depicts four bar graphs showing the result of the miRNAs in Table 22 of Example 23 evaluated for their target-site mediated suppression in various immune cell lines and CD34+CD38- CD90+CD45RA-Ein- long-term hematopoietic stem cells. While the insertion of liver specific miR-122 or epithelial specific miR-200b, 200c. 203a and 205 target sites in the 3 ’ UTR of cargo does not affect its expression in immune cells, hematopoietic miRNAs suppress cargo expression in one or more cell types. miR-142 and let7e effectively suppress cargo expression in all cell lines. miR-223 exhibits mild activity in all cell lines, with the least effect observed in long-term hematopoietic cells. miR-155 demonstrates suppression in the monocyte cell line and in long-term hematopoietic cells, albeit weaker in the latter. miR-342 suppresses expression in both monocyte and T-cell lines. Both miR-126 5p and 3p, which are abundantly expressed in endothelial cells, show a substantial suppressive effect in long-term hematopoietic stem cells.

[0153] FIG. 83 is a set of bar graphs comparisng genome editing efficiency on the hB2M target, as reported in Example 18. T7E1 = T7 endonuclease I assay; CE = capillary electrophoresis. Results reported as the average of three biological replicates each consisting of three technical replicates in the case of the T7EI assay and two biological replicates, each consisting of three technical replicates in the case of the capillary electrophoresis. Dots represent average values of each biological replicate.

[0154] FIG. 84 is a set of bar graphs comparing the T7Endonuclease 1 assay results reported in Example 18. The graph shows the data from the 13 best performing mutants selected from the study. Results reported as the average of three biological replicates each consisting of three technical replicates. 405-1 = the D169Rmutant standard (SEQ ID NO: 585). Each additional mutant indicated in the graph is a cumulative mutant based off of the standard 405-1 sequence.

[0155] FIG. 85 is a set of bar graphs comparing TapeStation DNA assay results reported in Example 18. The graph shows tire data from the 13 best performing mutants selected from the study. Data is representaed as two biological replicates, each consisting of three technical replicates. 405-1 = the D169R mutant standard (SEQ ID NO: 585). Each additional mutant is a cumulative mutant based off of the standard 405-1 sequence. Each additional mutant indicated in the graph is a cumulative mutant based off of the standard 405-1 sequence.

[0156] FIG. 86 is a bar graph demonstrating the effect of various guide RNA extension strategies and the resulting effect in efficacy in editing at the HBG1 / 2 locus, as reported in Example 22. Extension of gRNA at the 5‘ end improved editing in primary human HSPCs and increased in vitro editing to up to 80% at tire HBG1 / 2 locus.

[0157] FIG. 87A and FIG. 87B are bar graphs reporting the efficacy of different guide RNA sequences for use with Type V mutant 405-1 K292R, in vitro, at 9.375 ng and 150 ng dosages. As shown in FIG. 86A, gRNAs 1-3 demonstrated the highest editing efficiency at tire B2M locus.

[0158] FIG. 88 is a bar graph showing the efficacy of different guide RNA sequences for use with Type V mutant 405-1 K292R. in vivo, at two dose levels. Utilizing these gRNAs, Type V nuclease 405-1 K292R demonstrated high editing efficiency at the B2M locus, nearly comparable with a literature standard Cas9 system.DETAILED DESCRIPTION

[0159] Hie present disclosure provides Cas TypeV-based gene editing systems for use in various applications, including precision gene editing in cells, tissues, organs, or organisms. In various embodiments, the Cas TypeV -based gene editing systems comprise (a) a Cas TypeV polypeptide and (b) a Cas TypeV guide RNA which is capable of associating with a Cas TypeV polypeptide to form a complex such that the complex localizes to a target nucleic acid sequence (e.g., a genomic or plasmid target sequence) and binds thereto. In various embodiments, the Cas TypeV polypeptide has a nuclease activity which results in the cutting of at least one strand of DNA.

[0160] In exemplary embodiments, the Cas TypeV systems and / or components thereof described herein are formulated as part of a lipid nanoparticle (LNP). In some embodiments, a lipid nanoparticle comprises an ionizable lipid, a structural lipid, a PEGylated lipid, and a phospholipid.

[0161] In various embodiments, the Cas 12a polypeptide is apolypeptide selected from Table S15A (SEQ ID NO: 334 (No. ID405), SEQ ID NO: 58 (No. ID414), or SEQ ID NO: 564 (No. ID418), SEQ ID NO: 335 (No.ID406), SEQ ID NO: 331 (No. ID411), SEQ ID NO: 20 (No. ID415), and SEQ ID NO: 445 (No. ID419)), or a polypeptide having at least 70%, 75%, 80%, 85%, 90%, 95%, 99%, or 100% sequence identity with a polypeptide from Table S15A (SEQ ID NO: 334 (No. ID405), SEQ ID NO: 58 (No. ID414), or SEQ ID NO: 564 (No. ID418), SEQ ID NO: 335 (No. ID406), SEQ ID NO: 331 (No. ID411), SEQ ID NO: 20 (No. ID415). and SEQ ID NO: 445 (No. ID419)).

[0162] In various embodiments, the Cas Type V polypeptide is encoded by a polynucleotide sequence selected from Table S15B (SEQ ID NO: 365 (No. ID405), SEQ ID NO: 75 (No. ID414), or SEQ ID NO: 565 (No. ID418), SEQ ID NO: 366 (No. ID406), SEQ ID NO: 331 (No. ID411), SEQ ID NO: 30 (No. ID415), or SEQ ID NO: 445 (No. ID419)), or a polynucleotide having at least 70%, 75%, 80%, 85%, 90%, 95%, 99%, or 100% sequence identity with a polypeptide from Table S 15B (SEQ ID NO: 365 (No. ID405), SEQ ID NO: 75 (No. ID414), or SEQ ID NO:565 (No. ID418). SEQ ID NO: 366 (No. ID406), SEQ ID NO: 331 (No.ID411). SEQ ID NO: 30 (No. ID415), or SEQ ID NO: 445 (No. ID419)).

[0163] In various embodiments, the Cas Type V guide RNA is selected from any Cas Type V guide sequence disclosed in Table S15C (SEQ ID NO:28-29, 69-71, 355-360, 542-563), or a nucleic acid molecule having at least 70%, 75%, 80%, 85%, 90%, 95%, 99%, or 100% sequence identity with a Cas Type V guide sequence of Table S15C (SEQ ID NO:28-29, 69-71, 355-360, 542-563).

[0164] In various embodiments, the Cas Type V guide RNA may comprise (a) a portion that binds or associates with a Cas Type V polypeptide and (b) a region that comprises a targeting sequence, i.e.. a sequence which is complementary to target nucleic acid sequence. For Cas Type V guide RNA designs, just like for Cas9 guide RNA, the target sequence is typically next to a PAM sequence. But for Cas Type V, the PAM sequence in various embodiments is ty pically TTTV, where V ty pically represents A, C, or G. In various embodiments, the “V” of the TTTV is immediately adjacent to the most 5’ base of tire non-targeted strand side of the protospacer element. As for Cas9 guide RNA designs, the PAM sequence is typically not included in the guide RNA design.

[0165] In various embodiments, the guide RNA for Cas Type V is relatively short at only approximately 40- 44 bases long. The part that base pairs to the protospacer in the target sequence is 20-24 bases in length, and there is also a constant about 20-base section that binds to Cas Type V.

[0166] In various embodiments, nomenclature for a Cas Type V guide RNA is referred to as a “crRNA” and there is no Cas9-like ’tracrRNA" component.

[0167] In other aspects, the Cas Ty pe V-based gene editing systems may comprise one or more additional accessory? proteins having genome modifying functions, including recombinases, invertases, nucleases, polymerases, ligases, deaminases, reverse transcriptases, or epigenetic modifying functions. In variousembodiments, the accessory proteins may be provided separately. In other embodiments, the accessory proteins may be fused to Cas Type V, optionally with a linker.

[0168] In still another aspect, the disclosure provides delivery systems for introducing the Cas Type V-based gene editing systems or components thereof into cells, tissues, organs, or organisms. Depending on the chosen format, the Cas Type V-based gene editing systems and / or the individual or combined components thereof may be delivered as DNA molecules (e.g., encoded on one or more plasmids), RNA molecules (e.g., guide RNAs for targeting the Cas Type V protein or linear or circular mRNAs coding for the Cas Type V protein or accessory' protein components of the Cas Type V-based gene editing systems), proteins (e.g., Cas Type V polypeptides, accessory proteins having other functions (e.g., recombinases, nucleases, polymerases, ligases, deaminases, or reverse transcriptases), or protein-nucleic acid complexes (e.g., complexes between a guide RNA and a Cas Type V protein or fusion protein comprising a Cas Type V protein).

[0169] In another aspect, the present disclosure provides nucleic acid molecules encoding the Cas Type V- based gene editing systems or components thereof. In yet another aspect, the disclosure provides vectors for transferring and / or expressing said Cas Type V-based gene editing systems, e.g., under in vitro, ex vivo, and in vivo conditions. In still another aspect, tire disclosure provides cell-delivery compositions and methods, including compositions for passive and / or active transport to cells (e.g., plasmids), delivery by virus-based recombinant vectors (e.g., AAV and / or lentivirus vectors), delivery by non-virus-based systems (e.g.. liposomes and LNPs), and delivery by virus-like particles of the Cas Type V-based gene editing systems described herein. Depending on the delivery system employed, the Cas Type V-based gene editing systems described herein may be delivered in the form of DNA (e.g., plasmids or DNA-based virus vectors), RNA (e.g., guide RNA and rnRNA delivered by LNPs), a mixture of DNA and RNA, protein (e.g., virus-like particles), and ribonucleoprotein (RNP) complexes. Any suitable combinations of approaches for delivering the components of the herein disclosed Cas Type V-based gene editing systems may be employed.

[0170] In other embodiments, the Cas Type V-based gene editing systems may comprise a template DNA comprising an edit, e.g., a single strand or double strand donor molecule (linear or circular) which may be used by the cell to repair a single or double cut lesion introduced by a Cas Type V-based gene editing systems by way of cellular repair processes, including homology-dependent repair (HDR) (e.g., in dividing cells) or non-homologous end joining (NHEJ) (in non-dividing cells).

[0171] In one embodiment, each of the components of the Cas Type V-based gene editing systems is delivered by an all-RNA system, e.g.. the delivery of one or more RNA molecules (e.g., mRNA and / or guide RNA) by one or more LNPs, wherein the one or more RNA molecules form the guide RNA and / or are translated into the polypeptide components (e.g., the Cas Type V polypeptides and / or any accessoryproteins), and a DNA or RNA-encoded template DNA molecule (e.g., donor template), as appropriate or desired.

[0172] In yet another aspect, the disclosure provides methods for genome editing by introducing a Cas Type V-based gene editing system described herein into a cell (e.g., under in vitro, in vivo, or ex vivo conditions) comprising a target edit site, thereby resulting in an edit at the target edit. In other aspects, the disclosure provides formulations comprising any of the aforementioned components for delivery to cells and / or tissues, including in vitro, in vivo, and ex vivo delivery, recombinant cells and / or tissues modified by the recombinant Cas Type V-based gene editing systems and methods described herein, and methods of modifying cells by conducting genome editing using the herein disclosed Cas Type V-based gene editing systems.

[0173] The disclosure also provides methods of making the Cas Type V-based gene editing systems, their protein and nucleic acid molecule components, vectors, compositions and fonnulations described herein, as well as to pharmaceutical compositions and kits for modifying cells under in vitro, in vivo, and ex vivo conditions that comprise the herein disclosed genome editing and / or modification systems.A. General Definitions

[0174] Unless otherwise defined, all terms of art, notations and other scientific terminology used herein are intended to have the meanings commonly understood by those of skill in the art to which this disclosure pertains. In some cases, terms with commonly understood meanings are defined herein for clarity and / or for ready reference, and the inclusion of such definitions herein should not necessarily be construed to represent a difference over what is generally understood in the art. The techniques and procedures described or referenced herein are generally well understood and commonly employed using conventional methodologies by those skilled in the art, such as, for example, tire widely utilized molecular cloning methodologies described in Sambrook et al., Molecular Cloning: A Laboratory Manual 4th ed. (2012) Cold Spring Harbor Laboratory Press. Cold Spring Harbor, NY. As appropriate, procedures involving the use of commercially available kits and reagents are generally carried out in accordance with manufacturer-defined protocols and conditions unless otherwise noted.An

[0175] The articles “a” and “an” are used herein to refer to one or to more than one (i.e., to at least one) of the grammatical object of the article. By way of example, “an element” means one element or more than one element.About

[0176] ‘ ‘About” as used herein when referring to a measurable value such as an amount, a temporal duration, and the like, is meant to encompass variations of ±10%, as such variations are appropriate to perform the disclosed methods.Biologically active

[0177] As used herein, the term “biologically active” refers to a characteristic of an agent (e.g., DNA, RNA, or protein) that has activity in a biological system (including in vitro and in vivo biological system), and particularly in a living organism, such as in a mammal, including human and non-human mammals. For instance, an agent when administered to an organism has a biological effect on that organism, is considered to be biologically active.Bulge

[0178] As used herein, the term “bulge” refers to a small region of unpaired base(s) that interrupts a “stem” of base-paired nucleotides. The bulge may comprise one or two single-stranded or unbase-paired nucleotides joined at both ends by base-paired nucleotides of the stem. The bulge can be symmetrical (viz., the two unbase-paired single-stranded regions have the same number of nucleotides), or asymmetrical (viz., the unbase-paired single stranded region(s) have different or unequal numbers of nucleotides), or there is only one unbase-paired nucleotide on one strand. A bulge can be described as A / B (such as a “2 / 2 bulge,” or a “1 / 0 bulge”) wherein A represents the number of impaired nucleotides on the upstream strand of the stem, and B represents the number of unpaired nucleotides on the downstream strand of the stem. An upstream strand of a bulge is more 5’ to a downstream strand of the bulge in the primary nucleotide sequence. cDNA

[0179] As used hereing, the term “cDNA” refers to a strand of DNA copied from an RNA template, e.g., by a reverse transcriptase.Casl2a or Casl2a polypeptide

[0180] As used herein, tire “Casl2a polypeptide”, “Casl2a protein” or “Casl2a nuclease” refers to a RNA- binding site-directed CRISPR Cas TypeV polypeptide that recognizes and / or binds RNA and is targeted to a specific DNA sequence. An Cas 12a system as described herein refers to a specific DNA sequence by the RNA molecule to which the Casl2a polypeptide or Casl2a protein is bound. The RNA molecule comprises a sequence that binds, hybridizes to, or is complementary' to a target sequence within the targeted polynucleotide sequence, thus targeting the bound polypeptide to a specific location within the targeted polynucleotide sequence (the target sequence). “Cas 12a” is a type of CRISPR Class II Type V nuclease. The specification may describe the polypeptides contemplated in the scope of this application as Cas 12a polypeptides or alternatively as Cas TypeV polypeptides, or the like.Cleavage

[0181] As used herein, the term “cleavage” refers to the breakage of the covalent backbone of a DNA molecule. Cleavage can be initiated by a variety of methods including, but not limited to, enzymatic or chemical hydrolysis of a phosphodiester bond. Both single -stranded cleavage and double-stranded cleavageare possible, and double-stranded cleavage can occur as a result of two distinct single-stranded cleavage events. DNA cleavage can result in the production of either blunt ends or staggered ends.Cognate

[0182] The term “cognate” refers to two biomolecules that normally interact or co-exist in nature. Complementary

[0183] As used herein, the terms “complementary” or “substantially complementary" are meant to refer to a nucleic acid (e.g., RNA, DNA) that comprises a sequence of nucleotides that enables it to non-covalently bind, i.e., form Watson-Crick base pairs and / or G / U base pairs, “anneal”, or “hybridize,” to another nucleic acid in a sequence -specific, antiparallel, manner (i.e., a nucleic acid specifically binds to a complementary nucleic acid) under the appropriate in vitro and / or in vivo conditions of temperature and solution ionic strength. Standard Watson-Crick base-pairing includes: adenine (A) pairing with thymidine (T), adenine (A) pairing with uracil (U), and guanine (G) pairing with cytosine (C) [DNA. RNA], In addition, for hybridization between two RNA molecules (e.g., dsRNA), and for hybridization of a DNA molecule with an RNA molecule (e.g., when a DNA target nucleic acid base pairs with a guide RNA, etc.): guanine (G) can also base pair with uracil (U). For example, G / U base-pairing is at least partially responsible for the degeneracy (i.e., redundancy) of the genetic code in the context of tRNA anti -codon base-pairing with codons in mRNA. Thus, in the context of this disclosure, a guanine (G) is considered complementary to both a uracil (U) and to an adenine (A). For example, when a G / U base-pair can be made at a given nucleotide position of a dsRNA duplex of a guide RNA molecule, the position is not considered to be non-complementary, but is instead considered to be complementary.

[0184] It is understood that the sequence of a polynucleotide need not be 100% complementary to that of its target nucleic acid to be specifically hybridizable or hybridizable. Moreover, a polynucleotide may hybridize over one or more segments such that intervening or adjacent segments are not involved in the hybridization event (e.g., a bulge, a loop structure or hairpin structure, etc.). A polynucleotide can comprise 60% or more, 65% or more, 70% or more, 75% or more, 80% or more, 85% or more, 90% or more, 95% or more, 98% or more, 99% or more, 99.5% or more, or 100% sequence complementarity to a target region within the target nucleic acid sequence to which it will hybridize. For example, an antisense nucleic acid in which 18 of 20 nucleotides of the antisense compound are complementary to a target region, and would therefore specifically hybridize, would represent 90 percent complementarity. In this example, the remaining noncomplementary nucleotides may be clustered or interspersed with complementary nucleotides and need not be contiguous to each other or to complementary nucleotides. Percent complementarity between particular stretches of nucleic acid sequences within nucleic acids can be determined using any convenient method. Example methods include BEAST programs (basic local alignment search tools) and PowerBLAST programs (Altschul et al., J.Mol. Biol., 1990, 215, 403-410; Zhang and Madden, Genome Res., 1997, 7, 649-656), the Gap program (Wisconsin Sequence Analysis Package, Version 8 for Unix, Genetics Computer Group, University Research Park, Madison Wis.), e.g., using default settings, which uses the algorithm of Smith and Waterman (Adv. Appl. Math., 1981, 2, 482-489), and the like. Consisting essentially of Consisting essentially of

[0185] The phrase “consisting essentially of’ is meant herein to exclude anything that is not the specified active component or components of a system, or that is not the specified active portion or portions of a molecule.Control sequences

[0186] The term “control sequences” is intended to include, at a minimum, all components whose presence is essential for expression, and can also include additional components whose presence is advantageous, for example, leader sequences and fusion partner sequences. Suitable expression vectors include, without limitation, plasmids and viral vectors derived from, for example, bacteriophage, baculoviruses, tobacco mosaic virus, herpes viruses, cytomegalovirus, retroviruses, vaccinia viruses, adenoviruses, and adeno- associated viruses. Numerous vectors and expression systems are commercially available, such as from Novagen (Madison. WI), Clontech (Palo Alto, CA), Stratagene (La Jolla, CA), and Invitrogen / Life Technologies (Carlsbad, CA). The present invention comprehends recombinant vectors that may include viral vectors, bacterial vectors, protozoan vectors, DNA vectors, or recombinants thereof.Degenerate variant

[0187] As used herein, the phrase “degenerate variant” of a reference nucleic acid sequence encompasses nucleic acid sequences that can be translated, according to the standard genetic code, to provide an amino acid sequence identical to that translated from the reference nucleic acid sequence. Hie term “degenerate oligonucleotide” or “degenerate primer” is used to signify an oligonucleotide capable of hybridizing with target nucleic acid sequences that are not necessarily identical in sequence butthat are homologous to one another within one or more particular segments.

[0188] Engineered nucleic acid constructs of the present disclosure may be encoded by a single molecule (e.g. , encoded by or present on the same plasmid or other suitable vector) or by multiple different molecules (e.g., multiple independently-replicating vectors).DNA-guided nuclease

[0189] As used herein, an “DNA-guided nuclease” is a type of “programmable nuclease,” and a specific type of “nucleic acid-guided nuclease.” An example of a DNA-guided nuclease is reported in Varshney et al ., DNA-guided genome editing using structure-guided endonucleases, Genome Biology, 2016, 17(1), 187, which may be used in the context of the present disclosure and is incorporated herein by reference. As usedherein, the term “DNA-guided nuclease” or “DNA-guided endonuclease” refers to a nuclease that associates covalently or non-covalently with a guide RNA thereby forming a complex between the guide RNA and the DNA-guided nuclease. The guide RNA comprises a spacer sequence which comprises a nucleotide sequence having complementarity’ with a strand of a target DNA sequence. Thus, the DNA-guided nuclease is indirectly guided or programmed to localize to a specific site in a DNA molecule through its association with the guide RNA. which directly binds or anneals to a strand of the target DNA through its complementarity region via Watson-Crick base-pairing.DNA regulatory sequences

[0190] As used herein, the terms “DNA regulatory sequences,” “control elements,” and '■regulatory elements,” can be used interchangeably herein to refer to transcriptional and translational control sequences, such as promoters, enhancers, polyadenylation signals, terminators, protein degradation signals, and the like, that provide for and / or regulate transcription of a non-coding sequence (e.g., guide RNA) or a coding sequence and / or regulate translation of a mRNA into an encoded polypeptide.Domain

[0191] The term “domain” as used herein refers to a structure of a biomolecule that contributes to a known or suspected function of the biomolecule. Domains may be co- extensive with regions or portions thereof; domains may also include distinct, non-contiguous regions of a biomolecule. Examples of protein domains include, but are not limited to, an Ig domain, an extracellular domain, a transmembrane domain, and a cytoplasmic domain.

[0062] As used herein, the term “molecule” means any compound, including, but not limited to, a small molecule, peptide, protein, sugar, nucleotide, nucleic acid, lipid, etc., and such a compound can be natural or synthetic.Donor nucleic acid

[0192] By a “donor nucleic acid” or “donor polynucleotide” or “donor DNA” or “HDR donor DNA” it is meant a single-stranded DNA to be inserted at a site cleaved by a programmable nuclease (e.g.. a CRISPR / Cas effector protein; a TALEN; a ZFN; a meganuclease) (e.g., after dsDNA cleavage, after nicking a target DNA, after dual nicking a target DNA, and the like). Tire donor polynucleotide can contain sufficient homology to a genomic sequence at the target site, e.g. 70%, 80%, 85%, 90%, 95%, or 100% homology with the nucleotide sequences flanking the target site, e.g., within about 200 bases or less of the target site, e.g., within about 190 bases or less of the target site, e.g., within about 180 bases or less of the target site, e.g.. within about 170 bases or less of the target site, e.g., within about 160 bases or less of the target site, e.g.. within about 150 bases or less of the target site, e.g., within about 140 bases or less of the target site, e.g., within about 130 bases or less of the target site, e.g., within about 120 bases or less of the target site, e.g., within about 110 bases or less of the target site, e.g., within about 100 bases or less of the target site, e.g.,within about 90 bases or less of the target site, e.g., within about 80 bases or less of the target site, e.g., within about 70 bases or less of the target site, e.g., within about 60 bases or less of the target site, e.g., 50 bases or less of the target site, e.g., within about 30 bases, within about 15 bases, within about 10 bases, within about 5 bases, or immediately flanking the target site, to support homology-directed repair between it and the genomic sequence to which it bears homology.Effective amount

[0193] An ■‘effective amount” as used herein, means an amount which provides a therapeutic or prophylactic benefit under the conditions of administration.Encapsulation efficiency

[0194] As used herein, “encapsulation efficiency” refers to the amount of a therapeutic and / or prophylactic that becomes part of a nanoparticle composition, relative to the initial total amount of therapeutic and / or prophylactic used in the preparation of a nanoparticle composition. For example, if 97 mg of a polynucleotide are encapsulated in a nanoparticle composition out of a total 100 mg of therapeutic and / or prophylactic initially provided to the composition, the encapsulation efficiency may be given as 97%. As used herein, “encapsulation” may refer to complete, substantial, or partial enclosure, confinement, surrounding, or encasement.Encodes

[0195] As used herein, a DNA sequence that “encodes” a particular RNA is a DNA nucleotide sequence that is transcribed into RNA. A DNA polynucleotide may encode an RNA (mRNA) that is translated into protein (and therefore the DNA and the mRNA both encode the protein), or a DNA polynucleotide may encode an RNA that is not translated into protein (e.g. tRNA, rRNA, microRNA (miRNA), a “non-coding” RNA (ncRNA). a guide RNA, etc.).Exosomes

[0196] As used herein, the term “exosomes” refer to small membrane bound vesicles with an endocytic origin. Without wishing to be bound by theory, exosomes are generally released into an extracellular environment from host / progenitor cells post fusion of multivesicular bodies the cellular plasma membrane. As such, exosomes can include components of the progenitor membrane in addition to designed components. Exosome membranes are generally lamellar, composed of a bilayer of lipids, with an aqueous inter- nanoparticle space.Expression vector

[0197] As used herein, the term “expression vector” or “expression construct” refers to a vector that includes one or more expression control sequences, and an “expression control sequence” is a DNA sequence that controls and regulates the transcription and / or translation of another DNA sequence. Expression controlsequences are sequences which control the transcription, post-transcriptional events and translation of nucleic acid sequences.

[0198] Expression control sequences include appropriate transcription initiation, termination, promoter and enhancer sequences; efficient RNA processing signals such as splicing and polyadenylation signals; sequences that stabilize cytoplasmic mRNA; sequences that enhance translation efficiency (e.g., ribosome binding sites); sequences that enhance protein stability; and when desired, sequences that enhance protein secretion. The nature of such control sequences differs depending upon the host organism; in prokary otes, such control sequences generally include promoter, ribosomal binding site, and transcription termination sequence.Fusion protein

[0199] Tire tenn “fusion protein" refers to a polypeptide comprising a polypeptide or fragment coupled to heterologous amino acid sequences optionally via an amino acid linker. Fusion proteins are useful because they can be constructed to contain two or more desired functional elements from two or more different proteins. A fusion protein comprises at least 10 contiguous amino acids from a polypeptide of interest, more preferably at least 20 or 30 amino acids, even more preferably at least 40, 50 or 60 amino acids, yet more preferably at least 75, 100 or 125 amino acids.

[0200] Fusions that include tire entirety of the proteins of the present invention have particular utility. The heterologous polypeptide included within the fusion protein of the present invention is at least 6 amino acids in length, often at least 8 amino acids in length, and usefully at least 15, 20, and 25 amino acids in length. Fusions that include larger polypeptides, such as an IgG Fc region, and even entire proteins, such as the green fluorescent protein (“GFP”) chromophore -containing proteins, have particular utility. Fusion proteins can be produced recombinantly by constructing a nucleic acid sequence which encodes the polypeptide or a fragment thereof in frame with a nucleic acid sequence encoding a different protein or peptide and then expressing the fusion protein. Alternatively, a fusion protein can be produced chemically by crosslinking the polypeptide or a fragment thereof to another protein.Guide RNA

[0201] The RNA molecule that binds to the Casl2a polypeptide and targets the polypeptide to a specific location within the targeted polynucleotide sequence is referred to herein as the “guide RNA" or “guide RNA polynucleotide" (also referred to herein as a “guide RNA" or “gRNA" or “crRNA”). A guide RNA comprises two segments, a “DNA-targeting segment” and a “protein-binding segment.” By “segment” it is meant a segment / section / region of a molecule, e.g., a contiguous stretch of nucleotides in an RNA. As an illustrative, non-limiting example, a protein-binding segment of a guide RNA can comprise base pairs 5-20 of the RNA molecule that is 40 base pairs in length; and the DNA-targeting segment can comprise base pairs 21-40 of theRNA molecule that is 40 base pairs in length. The definition of '‘segment,” unless otherwise specifically defined in a particular context, is not limited to a specific number of total base pairs, is not limited to any particular number of base pairs from a given RNA molecule, is not limited to a particular number of separate molecules within a complex, and may include regions of RNA molecules that are of any total length and may or may not include regions with complementarity to other molecules.

[0202] The DNA-targeting segment (or “DNA-targeting sequence”) comprises a nucleotide sequence that is complementary to a specific sequence within a targeted polynucleotide sequence (the complementary strand of the targeted polynucleotide sequence) designated the “protospacer- like” sequence herein. The proteinbinding segment (or “protein-binding sequence”) interacts with a site-directed modifying polypeptide. When the site-directed modifying polypeptide is an Casl2a polypeptide, site-specific cleavage of the targeted polynucleotide sequence may occur at locations detennined by both (i) base-pairing complementarity between the guide RNA and the targeted polynucleotide sequence; and (ii) a short motif (referred to as the protospacer adjacent motif (PAM)) in the targeted polynucleotide sequence.Heterologous nucleic acid

[0203] As used herein, the term “heterologous nucleic acid” refers to a genoty pically distinct entity from that of the rest of the entity to which it is compared or into which it is introduced or incorporated. For example, a polynucleotide introduced by genetic engineering techniques into a different cell type is a heterologous polynucleotide (e.g., DNA or RNA) and, if expressed, can encode a heterologous polypeptide. Similarly, a cellular sequence (e.g., a gene or portion thereof) that is incorporated into a viral vector is a heterologous nucleotide sequence with respect to the vector.Homology

[0204] A protein has “homology” or is “homologous” to a second protein if the nucleic acid sequence that encodes the protein has a similar sequence to the nucleic acid sequence that encodes the second protein. Alternatively, a protein has homology to a second protein if the two proteins have "similar" amino acid sequences. (Thus, the term “homologous proteins” is defined to mean that the two proteins have similar amino acid sequences.) As used herein, homology between two regions of amino acid sequence (especially with respect to predicted structural similarities) is interpreted as implying similarity in function.

[0205] Sequence homology for polypeptides, which is also referred to as percent sequence identity, is typically measured using sequence analysis software. See, e.g.. the Sequence Analysis Software Package of the Genetics Computer Group (GCG), University of Wisconsin Biotechnology Center, 910 University Avenue, Madison, Wis. 53705. Protein analysis software matches similar sequences using a measure of homology assigned to various substitutions, deletions and other modifications, including conservative amino acid substitutions. For instance, GCG contains programs such as “Gap” and “Bestfit” which can be used withdefault parameters to determine sequence homology or sequence identity between closely related polypeptides, such as homologous polypeptides from different species of organisms or between a wild-type protein and a mutein thereof. See, e.g. , GCG Version 6.1.

[0206] A preferred algorithm when comparing a particular polypeptide sequence to a database containing a large number of sequences from different organisms is the computer program BLAST (Altschul et al., J. Mol. Biol. 215:403-410 (1990); Gish and States, Nature Genet. 3:266-272 (1993); Madden etal.,Meth. Enzymol. 266: 131-141 (1996); Altschul et al.. Nucleic Acids Res. 25:3389-3402 (1997); Zhang and Madden, Genome Res. 7:649-656 (1997)), especially blastp ortblastn (Altschul et al., Nucleic Acids Res. 25:3389-3402 (1997)). Preferred parameters for BLASTp are: Expectation value: 10 (default); Filter: seg (default); Cost to open a gap: 11 (default); Cost to extend a gap: 1 (default); Max. alignments: 100 (default); Word size: 11 (default);No. of descriptions: 100 (default); Penalty Matrix: BLOWSUM62.

[0207] The length of polypeptide sequences compared for homology will generally be at least about 16 amino acid residues, usually at least about 20 residues, more usually at least about 24 residues, typically at least about 28 residues, and preferably more than about 35 residues. When searching a database containing sequences from a large number of different organisms, it is preferable to compare amino acid sequences. Database searching using amino acid sequences can be measured by algorithms other than blastp known in the art. For instance, polypeptide sequences can be compared using FASTA, a program in GCG Version 6.1.

[0208] FASTA provides alignments and percent sequence identity of the regions of the best overlap between the query and search sequences. Pearson, Methods Enzymol. 183:63-98 (1990) (incorporated by reference herein). For example, percent sequence identity between amino acid sequences can be determined using FASTA with its default parameters (a word size of 2 and the PAM250 scoring matrix), as provided in GCG Version 6.1, herein incorporated by reference.Homology-directed repair

[0209] As used herein, “homology-directed repair (HDR)” refers to the specialized form DNA repair that takes place, for example, during repair of double-strand breaks in cells. This process requires nucleotide sequence homology, uses a “donor” molecule to template repair of a “target” molecule (i.e., the one that experienced the double-strand break), and leads to the transfer of genetic infonnation from the donor to the target. Homology-directed repair may result in an alteration of tire sequence of the target molecule (e g ., insertion, deletion, mutation), if the donor polynucleotide differs from tire target molecule and part or all of the sequence of the donor polynucleotide is incorporated into the targeted polynucleotide sequence.Identical

[0210] As used herein, the term “identical” refers to two or more sequences or subsequences which are the same. In addition, the term “substantially identical,” as used herein, refers to two or more sequences whichhave a percentage of sequential units which are the same when compared and aligned for maximum correspondence over a comparison window, or designated region as measured using a comparison algorithm or by manual alignment and visual inspection. By way of example only, two or more sequences may be “substantially identical'’ if the sequential units are about 60% identical, about 65% identical, about 70% identical, about 75% identical, about 80% identical, about 85% identical, about 90% identical, or about 95% identical over a specified region. Such percentages to describe the “percent identity” of two or more sequences. The identity of a sequence can exist over a region that is at least about 75-100 sequential units in length, over a region that is about 50 sequential units in length, or, where not specified, across the entire sequence. This definition also refers to the complement of a test sequence.

[0211] Alternatively, substantially identical or similarity exists when a nucleic acid or fragment thereof hybridizes to another nucleic acid, to a strand of another nucleic acid, or to the complementary strand thereof, under stringent hybridization conditions. “Stringent hybridization conditions” and “stringent wash conditions” in the context of nucleic acid hybridization experiments depend upon a number of different physical parameters. Nucleic acid hybridization will be affected by such conditions as salt concentration, temperature, solvents, the base composition of the hybridizing species, length of the complementary regions, and tire number of nucleotide base mismatches betw een the hybridizing nucleic acids, as will be readily appreciated by those skilled in the art. One having ordinary skill in the art knows how to vary these parameters to achieve a particular stringency of hybridization.Isolated

[0212] ‘ ‘Isolated” means altered or removed from the natural state. For example, a nucleic acid or a peptide naturally present in a living animal is not “isolated,” but the same nucleic acid or peptide partially or completely separated from the coexisting materials of its natural state is “isolated.” An isolated nucleic acid or protein can exist in substantially purified fonn. or can exist in a non-native environment such as. for example, a host cell. An “isolated nucleic acid” refers to a nucleic acid segment or fragment, which has been separated from sequences which flank it in a naturally occurring state, i.e., a DNA fragment, which has been removed from the sequences which are normally adjacent to the fragment, i.e., the sequences adjacent to the fragment in a genome in which it naturally occurs. Tire term also applies to nucleic acids which have been substantially purified from other components, which naturally accompany the nucleic acid, i.e., RNA or DNA or proteins, which naturally accompany it in the cell. The term therefore includes, for example, a recombinant DNA or RNA, which is incorporated into a vector, into an autonomously replicating plasmid or virus, or into the genomic DNA or RNA of a prokaryote or eukaryote, or which exists as a separate molecule (i.e., as a cDNA or a genomic or cDNA fragment produced by PCR or restriction enzyme digestion) independent ofother sequences. It also includes a recombinant DNA or RNA. which is part of a hybrid gene encoding additional polypeptide sequence.Isolated protein

[0213] The term “isolated protein” or “isolated polypeptide” is a protein or polypeptide that by virtue of its origin or source of derivation (1) is not associated with naturally associated components that accompany it in its native state, (2) exists in a purity not found in nature, where purity can be adjudged with respect to tire presence of other cellular material (e.g. , is free of other proteins from the same species) (3) is expressed by a cell from a different species, or (4) does not occur in nature (e.g., it is a fragment of a polypeptide found in nature or it includes amino acid analogs or derivatives not found in nature or linkages other than standard peptide bonds). Thus, a polypeptide that is chemically synthesized or synthesized in a cellular system different from tire cell from which it naturally originates will be “isolated” from its naturally associated components. A polypeptide or protein may also be rendered substantially free of naturally associated components by isolation, using protein purification techniques well known in the art. As thus defined, “isolated” does not necessarily require that the protein, polypeptide, peptide or oligopeptide so described has been physically removed from its native environment.Lipid nanoparticle (LNP)

[0214] As used herein, the term “lipid nanoparticle” or LNP refers to a type of lipid particle delivery system formed of small solid or semi-solid particles possessing an exterior lipid layer with a hydrophilic exterior surface that is exposed to the non-LNP environment, an interior space which may aqueous (vesicle like) or non-aqueous (micelle like), and at least one hydrophobic inter-membrane space. LNP membranes may be lamellar or non-lamellar and may be comprised of 1, 2, 3, 4, 5 or more layers. In some embodiments, LNPs may comprise a nucleic acid (e.g. Casl2a editing system) into their interior space, into the inter membrane space, onto their exterior surface, or any combination thereof. In some embodiments, an LNP of the present disclosure comprises an ionizable lipid, a structural lipid, a PEGylated lipid (aka PEG lipid), and a phospholipid. In alternative embodiments, an LNP comprises an ionizable lipid, a structural lipid, a PEGylated lipid (aka PEG lipid), and a zwitterionic amino acid lipid.

[0215] Further discuss of liposomes can be found, for example, in Tenchov et al., “Lipid Nanoparticles - From Liposomes to mRNA Vaccine Delivery, a Landscape of Diversity and Advancement,” ACS Nano, 2021. 15, pp. 16982-17015 (the contents of which are incorporated by reference).Linker

[0216] As used herein, the term “linker” refers to a molecule linking or joining two other molecules or moieties. Tire linker can be an amino acid sequence in the case of a linker joining Evo fusion proteins. For example, an RNA-guided nuclease (e.g., Casl2a) can be fused to a reverse transcriptase or deaminase by anamino acid linker sequence. The linker can also be a nucleotide sequence in the case of joining two nucleotide sequences together. For example, in the instant case, a guide RNA at its 5' and / or 3' ends may be linked by a nucleotide sequence linker to one or more nucleotide sequences (e.g., a RT template in the case of a prime editor guide RNA). In other embodiments, the linker is an organic molecule, group, polymer, or chemical moiety. In some embodiments, the linker is 5-100 amino acids in length, for example. 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15. 16. 17. 18, 19, 20, 21, 22, 23, 24. 25. 26. 27, 28, 29, 30, 30-35. 35-40, 40- 45, 45-50. 50-60, 60-70, 70-80, 80-90, 90-100, 100-150, or 150-200 amino acids in length. Longer or shorter linkers are also contemplated.Liposomes

[0217] As used herein, the tenn “liposomes” refer to small vesicles that contain at least one lipid bilayer membrane surrounding an aqueous inner-nanoparticle space that is generally not derived from a progenitor / host cell.Micelles

[0218] As used herein, the term “micelles” refer to small particles which do not have an aqueous intraparticle space.Modified derivative

[0219] A “modified derivative” refers to polypeptides or fragments thereof that are substantially homologous in primary structural sequence but which include, e.g., in vivo or in vitro chemical and biochemical modifications or which incorporate amino acids that are not found in the native polypeptide. Such modifications include, for example, acetylation, carboxylation, phosphorylation, glycosylation, ubiquitination, labeling, e.g., with radionuclides, and various enzymatic modifications, as will be readily appreciated by those skilled in the art. A variety of methods for labeling polypeptides and of substituents or labels useful for such purposes are well known in the art, and include radioactive isotopes such as125I,32P,35S. and3H. ligands which bind to labeled antiligands (e.g.. antibodies), fluorophores. chemiluminescent agents, enzymes, and antiligands which can serve as specific binding pair members for a labeled ligand. The choice of label depends on the sensitivity required, ease of conjugation with the primer, stability requirements, and available instrumentation. Methods for labeling polypeptides are well known in the art. See, e.g, Ausubel et al., Current Protocols in Molecular Biology. Greene Publishing Associates (1992, and Supplements to 2002) (hereby incorporated by reference).Modulating

[0220] By the term “modulating,” as used herein, is meant mediating a detectable increase or decrease in the level of a response in a subject compared with the level of a response in the subject in tire absence of a treatment or compound, and / or compared with the level of a response in an otherwise identical but untreatedsubject. The term encompasses perturbing and / or affecting a native signal or response thereby mediating a beneficial therapeutic response in a subject, preferably, a human.Mutated

[0221] The term “mutated" when applied to nucleic acid sequences means that nucleotides in a nucleic acid sequence may be inserted, deleted or changed compared to a reference nucleic acid sequence. A single alteration may be made at a locus (a point mutation) or multiple nucleotides may be inserted, deleted or changed at a single locus. In addition, one or more alterations may be made at any number of loci within a nucleic acid sequence. A nucleic acid sequence may be mutated by any method known in the art including but not limited to mutagenesis techniques such as “error-prone PCR” (a process for performing PCR under conditions where the copying fidelity of the DNA polymerase is low, such that a high rate of point mutations is obtained along the entire length of the PCR product; see, e.g., Leung etal., Technique, 1: 11-15 (1989) and Caldwell and Joyce, PCR Methods Applic. 2:28-33 (1992)); and “oligonucleotide -directed mutagenesis” (a process which enables the generation of site- specific mutations in any cloned DNA segment of interest: see, e.g., Reidhaar-Olson and Sauer, Science 241:53-57 (1988)).Nanoparticle

[0222] As used herein, the term “nanoparticle” refers to any particle ranging in size from 10- 1,000 nm.Non-homologous end joining

[0223] As used herein, “non-homologous end joining (NHEJ)” refers to the repair of double-strand breaks in DNA by direct ligation of the break ends to one another without the need for a homologous template (in contrast to homology -directed repair, which requires a homologous sequence to guide repair). NHEJ often results in the loss (deletion) of nucleotide sequence near the site of the double-strand break.Non-peptide analog

[0224] Tire term “non-peptide analog” refers to a compound with properties that are analogous to those of a reference polypeptide. A non-peptide compound may also be termed a "peptide mimetic” or a “peptidomimetic.” See, e.g., Jones, Amino Acid and Peptide Synthesis, Oxford University Press ( 1992); Jung, Combinatorial Peptide and Nonpeptide Libraries: A Handbook, John Wiley (1997); Bodanszky et al., Peptide Chemistry’— A Practical Textbook, Springer Verlag (1993); Synthetic Peptides: A Users Guide, (Grant, ed., W. H. Freeman and Co., 1992); Evans et al., J. Med. Chem. 30: 1229 (1987); Fauchere, J. Adv. Drug Res. 15:29 (1986); Veber and Freidinger, Trends Neurosci.. 8:392-396 (1985); and references sited in each of the above, which are incorporated herein by reference. Such compounds are often developed with the aid of computerized molecular modeling. Peptide mimetics that are structurally similar to useful peptides of the present invention may be used to produce an equivalent effect and are therefore envisioned to be part of the present invention.Nuclear localization sequence (NLS)

[0225] As used herein, the term “nuclear localization sequence” or“NLS” refers to an amino acid sequence that promotes import of a protein (e.g., a RNA-guided nuclease) into the cell nucleus, for example, by nuclear transport. Nuclear localization sequences are known in the art. For example, NLS sequences are described in Plank et al., international PCT application, PCT / EP2000 / 011690, filed November 23. 2000, published as WO / 2001 / 038547 on May 31, 2001, the contents of which are incorporated herein by reference for its disclosure of exemplary nuclear localization sequences.Nucleic acid

[0226] As used herein, tire term “nucleic acid” or “nucleic acid molecule” or “nucleic acid sequence” or “polynucleotide” generally refer to deoxyribonucleic or ribonucleic oligonucleotides in either single- or double -stranded fonn. The term may (or may not) encompass oligonucleotides containing known analogues of natural nucleotides. The term also may (or may not) encompass nucleic acid-like structures with synthetic backbones, see, e.g., Eckstein, 1991; Baserga et ah, 1992; Milligan, 1993; WO 97 / 03211; WO 96 / 39154; Mata, 1997; Strauss-Soukup, 1997; and Samstag, 1996. The term encompasses both ribonucleic acid (RNA) and DNA, including cDNA, genomic DNA, synthetic, synthesized (e.g., chemically synthesized) DNA, and / or DNA (or RNA) containing nucleic acid analogs. The nucleotides Adenine (A). Thymine (T), Guanine (G) and Cytosine (C) also may (or may not) encompass nucleotide modifications, e.g. , methylated and / or hydroxylated nucleotides, e.g., Cytosine (C) encompasses 5-methylcytosine and 5- hydroxymethylcytosine. The nucleic acid can be in any topological conformation. For instance, the nucleic acid can be singlestranded, double-stranded, triple-stranded, quadruplexed, partially double-stranded, branched, hairpinned, circular, or in a padlocked conformation.Nucleic acid-guided nuclease

[0227] As used herein, the term “nucleic acid-guided nuclease” or “nucleic acid-guided endonuclease” refers to a nuclease (e.g., Casl2a) that associates covalently or non-covalently with a guide nucleic acid (e.g.. a guide RNA or a guide DNA) thereby forming a complex between the guide nucleic acid and the nucleic acid- guided nuclease. The guide nucleic acid comprises a spacer sequence which comprises a nucleotide sequence having complementarity with a strand of a target DNA sequence. Thus, the nucleic acid-guided nuclease is indirectly guided or programmed to localize to a specific site in a DNA molecule through its association with the guide nucleic acid, which directly binds or anneals to a strand of the target DNA through its complementarity region via Watson-Crick base-pairing. In some embodiments, the nucleic acid-guided nuclease will include a DNA-binding activity (e.g., as in the case for CRISPR Casl2a). Most commonly, the nucleic acid-guided nuclease is programmed by associating with a guide RNA molecule and in such cases the nuclease may be called “RNA-guided nuclease.” When programmed by a guide DNA, the nuclease may becalled a “DNA-guided nuclease.” Nucleic acid-guided, RNA-guided, or DNA-guided nucleases may also be referred to as '‘programmable nucleases,” which also include other classes of programmable nucleases which associate with specific DNA sequences through amino acid / nucleotide sequence recognition (e.g., zinc fingers nucleases (ZFN) and transcription activator like effector nucleases (TALEN)) rather than through guide RNAs. In addition, any nuclease contemplated herein may also be engineered to remove, inactivate, or otherwise eliminate one or more nuclease activities (e.g., by introducing a nuclease-inactivating mutation in the active site(s) of a nuclease, e.g., in the RuvC domain of a Casl2a). A nuclease that has been modified to remove, inactivate, or otherwise eliminate all nuclease activity may be referred to as a ‘'dead” nuclease. A dead nuclease is not able to cut either strand of a double-stranded DNA molecule. A nuclease that has been modified to remove, inactivate, or otherwise eliminate at least one nuclease activity but which still retains at least one nuclease activity may be referred to as a “nickase” nuclease. A nickase nuclease cuts one strand of a double -stranded DNA molecule, but not both strands. For example, a CRISPR Cas9 naturally comprises two distinct nuclease activity domains, namely, the HNH domain and the RuvC domain. The HNH domain cuts the strand of DNA bound to the guide RNA and the RuvC domain cuts the protospacer strand. One can obtain a nickase Cas9 by inactivating either the HNH domain or the RuvC domain. One can obtain a dead Cas9 by inactivating both the HNH domain and the RuvC domain. Other RNA-guided nuclease may be similarly converted to nickases and / or dead nucleases by inactivating one or more of the existing nuclease domains.Off-target effects

[0228] “Off-target effects” refer to non-specific genetic modifications that can occur when the CRISPR nuclease binds at a different genomic site than its intended target due to mismatch tolerance Hsu, P., Scott, D., Weinstein, J. et al. DNA targeting specificity of RNA- guided Cas9 nucleases. Nat Biotechnol 31, 827- 832 (2013). https: / / doi.org / 10.1038 / nbt.2647.Operably linked

[0229] As used herein, the term “operably linked” or “under transcriptional control,” when used in conjunction with the description of a promoter, refers to the correct location and orientation in relation to a polynucleotide e.g., a coding sequence) to control the initiation of transcription by RNA polymerase and expression of the coding sequence, such as one for the msr gene, msd gene, and / or the ret gene. Other transcriptional control regulatory elements (e.g., enhancer sequences, transcription factor binding sites) may also be operably linked to a gene if their location relative to a gene controls or regulates tire expression of the gene.PEG lipid

[0230] As used herein, a '‘PEG lipid” or “PEGylated lipid” refers to a lipid comprising a polyethylene glycol component.Peptide

[0231] As used herein, the tenns “peptide,” “polypeptide,” and “protein” are used interchangeably, and refer to a compound comprised of amino acid residues covalently linked by peptide bonds. A protein or peptide must contain at least two amino acids, and no limitation is placed on the maximum number of amino acids that can comprise a protein’s or peptide’s sequence. Polypeptides include any peptide or protein comprising two or more amino acids joined to each other by peptide bonds. As used herein, the term refers to both short chains, which also commonly are referred to in the art as peptides, oligopeptides and oligomers, for example, and to longer chains, which generally are referred to in the art as proteins, of which there are many types. “Polypeptides” include, for example, biologically active fragments, substantially homologous polypeptides, oligopeptides, homodimers, heterodimers, variants of polypeptides, modified polypeptides, derivatives, analogs, fusion proteins, among others. The polypeptides include natural peptides, recombinant peptides, synthetic peptides, or a combination thereof.Promoter

[0232] As used herein, the tenn “promoter” is art-recognized and refers to a nucleic acid molecule with a sequence recognized by the cellular transcription machinery and which is able to initiate transcription of a downstream gene. A promoter can be constitutively active, meaning that the promoter is always active in a given cellular context, or conditionally active, meaning that the promoter is only active in the presence of a specific condition. For example, a conditional promoter may only be active in the presence of a specific protein that connects a protein associated with a regulatory element in the promoter to the basic transcriptional machinery, or only in the absence of an inhibitory molecule. Within the promoter sequence will be found a transcription initiation site, as well as protein binding domains responsible for the binding of RNA polymerase. Eukaryotic promoters will often, but not always, contain “TATA” boxes and “CAT” boxes. Various promoters, including inducible promoters, may be used to drive expression by the various vectors of the present disclosure.Programmable nuclease

[0233] As used herein, the term “programmable nuclease” is meant to refer to a polypeptide that has the property of selective localization to a specific desired nucleotide sequence target in a nucleic acid molecule (e.g., to a specific gene target) due to one or more targeting functions. Such targeting functions can include one or more DNA-binding domains, such as zinc finger domains characteristic of many different types of DNA binding proteins or TALE domains characteristic of TALEN proteins. Such targeting function may also include tire ability to associate and / or form a complex with a guide RNA, which then localizes to aspecific site on the DNA which bears a sequence that is complementary to a portion of the guide RNA (i.e., the spacer of the guide RNA). In some embodiments, the programmable nuclease may be a single protein which comprises both a domain that binds directly (e.g., a ZF protein) or indirectly (e.g., an RNA-guided protein) to a target DNA site, as well as a nuclease domain. In other embodiments, the programmable nuclease may be a composite of two or more separate proteins or domains (from different proteins) which together provide the necessary functions of selective DNA binding and nuclease activity. For example, the programmable nuclease may comprise a (a) nuclease -inactive RNA-guided nuclease (which still is capable of binding a guide RNA, localizing to a target DNA, and binding to the target DNA, but not capable of cutting or nicking the strands) fused to a (b) nuclease protein or domain, such as a FokI nuclease.Polypeptide

[0234] Tire tenn “polypeptide" encompasses both naturally-occurring and non-naturally- occurring proteins, and fragments, mutants, derivatives and analogs thereof. A polypeptide may be monomeric or polymeric. Further, a polypeptide may comprise a number of different domains each of which has one or more distinct activities.Polypeptide fragment

[0235] The term “polypeptide fragment” as used herein refers to a polypeptide that has a deletion, e.g., an amino-tenninal and / or carboxy-terminal deletion compared to a full-length polypeptide. In a preferred embodiment, the polypeptide fragment is a contiguous sequence in which the amino acid sequence of the fragment is identical to the corresponding positions in the naturally-occurring sequence. Fragments typically are at least 5, 6, 7, 8, 9 or 10 amino acids long, preferably at least 12, 14, 16 or 18 amino acids long, more preferably at least 20 amino acids long, more preferably at least 25, 30, 35, 40 or 45, amino acids, even more preferably at least 50 or 60 amino acids long, and even more preferably at least 70 amino acids long.Polypeptide mutant

[0236] A “polypeptide mutant" or “mutein" refers to a polypeptide whose sequence contains an insertion, duplication, deletion, rearrangement or substitution of one or more amino acids compared to the amino acid sequence of a native or wild-type protein. A mutein may have one or more amino acid point substitutions, in which a single amino acid at a position has been changed to another amino acid, one or more insertions and / or deletions, in which one or more amino acids are inserted or deleted, respectively, in the sequence of the naturally- occurring protein, and / or truncations of the amino acid sequence at either or both the amino or carboxy termini. A mutein may have the same but preferably has a different biological activity compared to the naturally-occurring protein. A mutein has at least 85% overall sequence homology to its wild-type counterpart. Even more preferred are muteins having at least 90% overall sequence homology to the wild-type protein. In an even more preferred embodiment, a mutein exhibits at least 95% sequence identity, even more preferably 98%, even more preferably 99% and even more preferably 99.9% overall sequence identity.

[0237] Sequence homology may be measured by any common sequence analysis algorithm, such as Gap or Bestfit.

[0238] Amino acid substitutions can include those which: (1) reduce susceptibility to proteolysis, (2) reduce susceptibility to oxidation, (3) alter binding affinity for forming protein complexes. (4) alter binding affinity or enzymatic activity, and (5) confer or modify other physicochemical or functional properties of such analogs.

[0239] As used herein, tire twenty conventional amino acids and their abbreviations follow conventional usage. See Immunology-A Synthesis (Golub and Gren eds., Sinauer Associates, Sunderland, Mass., 2nded. 1991), which is incorporated herein by reference. Stereoisomers (e.g., D-amino acids) of the twenty conventional amino acids, unnatural amino acids such as a-, a-disubstituted amino acids. N-alkyl amino acids, and other unconventional amino acids may also be suitable components for polypeptides of the present invention. Examples of unconventional amino acids include: 4-hydroxyproline, y-carboxyglutamate, e- N,N,N- trimethyllysine, e-N-acetyllysine, O-phosphoserine, N-acetylserine, N-fonnylmethionine, 3- methylhistidine, 5-hydroxylysine. N-methylarginine, and other similar amino acids and imino acids (e.g., 4- hydroxyproline). In the polypeptide notation used herein, the left-hand end corresponds to tire amino terminal end and the right-hand end corresponds to the carboxy- terminal end, in accordance with standard usage and convention.Recombinant

[0240] The term “recombinant” refers to a biomolecule, e.g., a gene or protein, that (1) has been removed from its naturally occurring environment, (2) is not associated with all or a portion of a polynucleotide in which the gene is found in nature. (3) is operatively linked to a polynucleotide which it is not linked to in nature, or (4) does not occur in nature. The term “recombinant” can be used in reference to cloned DNA isolates, chemically synthesized polynucleotide analogs, or polynucleotide analogs that are biologically synthesized by heterologous systems, as well as proteins and / or mRNAs encoded by such nucleic acids.

[0241] As used herein, an endogenous nucleic acid sequence in the genome of an organism (or the encoded protein product of that sequence) is deemed “recombinant” herein if a heterologous sequence is placed adjacent to the endogenous nucleic acid sequence, such that tire expression of this endogenous nucleic acid sequence is altered. In this context, a heterologous sequence is a sequence that is not naturally adjacent to the endogenous nucleic acid sequence, whether or not the heterologous sequence is itself endogenous (originating from the same host cell or progeny thereof) or exogenous (originating from a different host cell or progeny thereof). By way of example, a promoter sequence can be substituted (e.g., by homologousrecombination) for the native promoter of a gene in the genome of a host cell, such that this gene has an altered expression pattern. This gene would now become “recombinant” because it is separated from at least some of the sequences that naturally flank it.

[0242] A nucleic acid is also considered “recombinant” if it contains any modifications that do not naturally occur to the corresponding nucleic acid in a genome. For instance, an endogenous coding sequence is considered “recombinant” if it contains an insertion, deletion or a point mutation introduced artificially, e.g., by human intervention. A “recombinant nucleic acid” also includes a nucleic acid integrated into a host cell chromosome at a heterologous site and a nucleic acid construct present as an episome.Recombinant host cell

[0243] The term “recombinant host cell” (or simply “host cell”), as used herein, is intended to refer to a cell into which a recombinant vector has been introduced. It should be understood that such terms are intended to refer not only to the particular subject cell but to the progeny of such a cell. Because certain modifications may occur in succeeding generations due to either mutation or environmental influences, such progeny may not, in fact, be identical to the parent cell, but are still included within the scope of the term “host cell” as used herein. A recombinant host cell may be an isolated cell or cell line grown in culture or may be a cell which resides in a living tissue or organism.

[0244] Suitable methods of genetic modification such as “transformation” include e.g.. viral or bacteriophage infection, transfection, conjugation, protoplast fusion, lipofection, electroporation, calcium phosphate precipitation, polyethyleneimine (PEI)-mediated transfection, DEAE-dextran mediated transfection, liposome-mediated transfection, particle gun technology, calcium phosphate precipitation, direct micro injection, nanoparticle- mediated nucleic acid delivery (see, e.g., Panyam et al., Adv Drug Deliv Rev. 2012 Sep. 13. pii: S0169-409X(12)00283-9. doi: 10.1016 / j.addr.2012.09.023), and the like. The choice of method of genetic modification is generally dependent on the type of cell being transformed and the circumstances under which the transformation is taking place (e.g., in vitro, ex vivo, or in vivo). A general discussion of these methods can be found in Ausubel, et al.. Short Protocols in Molecular Biology, 3rd ed., Wiley & Sons, 1995.Recombinant nucleic acid

[0245] A ‘ ‘recombinant nucleic acid” or “recombinant nucleotide” refers to a molecule that is constructed by joining nucleic acid molecules, which optionally may self-replicate in a live cell. Recombinant nucleic acids and synthetic nucleic acids also include those molecules that result from the replication of either of the foregoing.Region

[0246] The term "region" as used herein refers to a physically contiguous portion of the primary structure of a biomolecule. In the case of proteins, a region is defined by a contiguous portion of the amino acid sequence of that protein.RNA-guided nuclease

[0247] As used herein, an “RNA-guided nuclease" is a type of “programmable nuclease.” and a specific type of “nucleic acid-guided nuclease.” As used herein, the term “RNA-guided nuclease” or “RNA-guided endonuclease” refers to a nuclease that associates covalently or non-covalently with a guide RNA thereby forming a complex between the guide RNA and the RNA-guided nuclease. The guide RNA comprises a spacer sequence which comprises a nucleotide sequence having complementarity with a strand of a target DNA sequence. Thus, the RNA-guided nuclease is indirectly guided or programmed to localize to a specific site in a DNA molecule through its association with the guide RNA, which directly binds or anneals to a strand of the target DNA through its complementarity region via Watson-Crick base-pairing.Sequence identity

[0248] As used herein, the term “sequence identity" refers to the overall relatedness between polymeric molecules, e.g., between polynucleotide molecules (<?.g, DNA molecules and / or RNA molecules) and / or between polypeptide molecules. Calculation of the percent identity of two polynucleotide sequences, for example, can be performed by aligning the two sequences for optimal comparison purposes (e.g.. gaps can be introduced in one or both of a first and a second nucleic acid sequences for optimal alignment and nonidentical sequences can be disregarded for comparison purposes). For example, the length of a sequence aligned for comparison purposes is at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, or 100% of the length of tire reference sequence. The nucleotides at corresponding nucleotide positions are then compared. In other examples, the length of sequence identity comparison may be over a stretch of at least about nine nucleotides, usually at least about 20 nucleotides, more usually at least about 24 nucleotides, typically at least about 28 nucleotides, more typically at least about 32 nucleotides, and preferably at least about 36 or more nucleotides. When a position in the first sequence is occupied by the same nucleotide as the corresponding position in the second sequence, then the molecules are identical at that position. The percent identity between the two sequences is a function of the number of identical positions shared by the sequences, taking into account the number of gaps, and the length of each gap, which needs to be introduced for optimal alignment of the two sequences. The comparison of sequences and determination of percent identity between two sequences can be accomplished using a mathematical algorithm. For example, the percent identity between two nucleotide sequences can be determined using methods such as those described in Computational Molecular Biology, Lesk, A. M., ed., Oxford University Press, New' York, 1988; Biocomputing: Informatics and Genome Projects, Smith, D. W.,ed., Academic Press, New York, 1993; Sequence Analysis in Molecular Biology, von Heinje, G., Academic Press, 1987; Computer Analysis of Sequence Data, Part I, Griffin, A. M., and Griffin, H. G., eds., Humana Press, New Jersey, 1994; and Sequence Analysis Primer, Gribskov, M. and Devereux, J., eds., M Stockton Press, New York, 1991; each of which is incorporated herein by reference. For example, the percent identity between two nucleotide sequences can be determined using the algorithm of Meyers and Miller (CABIOS, 1989. 4: 11-17). which has been incorporated into the ALIGN program (version 2.0) using a PAM120 weight residue table, a gap length penalty of 12 and a gap penalty of 4. The percent identity between two nucleotide sequences can, alternatively, be determined using the GAP program in the GCG software package using an NWSgapdna. CMP matrix. Methods commonly employed to determine percent identity betw een sequences include, but are not limited to those disclosed in Carillo, H. and Lipman, D., SIAM J Applied Math., 48: 1073 (1988); incorporated herein by reference. Techniques for determining identity are codified in publicly available computer programs. Exemplary computer software to determine homology between two sequences include, but are not limited to, GCG program package, Devereux, J., et al., Nucleic Acids Research, 12(1), 387 (1984)), BLASTP, BLASTN, and FASTA (Altschul etal.,J. Mol. Biol. 215:403-410 (1990);

[0249] Gish and States, Nature Genet. 3:266-272 (1993); Madden et al., Meth. Enzymol. 266: 131- 141 (1996); Altschul et al., Nucleic Acids Res. 25:3389-3402 (1997); Zhang and Madden, Genome Res. 7:649-656 (1997)), (Altschul et al., Nucleic Acids Res. 25:3389-3402 (1997)). Polynucleotide sequences, for instance, can be compared using FASTA, Gap or Bestfit. which are programs in Wisconsin Package Version 10.0, Genetics Computer Group (GCG), Madison, Wis. FASTA provides alignments and percent sequence identity of tire regions of the best overlap between the query' and search sequences. Pearson, Methods Enzymol. 183:63-98 (1990) (hereby incorporated by reference in its entirety). Percent sequence identity between nucleic acid sequences, for instance, can be determined using FASTA with its default parameters (a word size of 6 and the NOPAM factor for the scoring matrix) or using Gap with its default parameters as provided in GCG Version 6.1.Specific binding

[0250] “Specific binding” refers to the ability of two molecules to bind to each other in preference to binding to other molecules in the environment. Typically, “specific binding” discriminates over adventitious binding in a reaction by at least tw o-fold, more typically by at least 10-fold, often at least 100-fold. Typically, the affinity or avidity of a specific binding reaction, as quantified by a dissociation constant, is about 107M or stronger (e.g., about 10'8M, 10'9M or even stronger).Stem and loop

[0251] As used herein, the term “stem” refers to two or more base pairs, such as 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more base pairs, fomied by inverted repeat sequences connected at a “tip,”where the more 5 ’ or ’ upstream” strand of the stem bends to allows the more 3 ' or “downstream” strand to base-pair with the upstream strand. The number of base pairs in a stem is the “length” of the stem. The tip of the stem is typically at least 3 nucleotides, but can be 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15 or more nucleotides.

[0252] Larger tips with more than 5 nucleotides are also referred to as a “loop.” An otherwise continuous stem may be interrupted by one or more bulges as defined herein. The number of unpaired nucleotides in the bulge(s) are not included in the length of the stem. The position of a bulge closest to the tip can be described by the number of base pairs between the bulge and the tip (e.g., the bulge is 4 bps from the tip). The position of tire other bulges (if any) further away from the tip can be described by the number of base pairs in the stem between the bulge in question and the tip, excluding any unpaired bases of other bulges in between. As used herein, the tenn “loop” in the polynucleotide refers to a single stranded stretch of one or more nucleotides, such as 2. 3, 4, 5, 6. 7, 8, 9, or 10 nucleotides, wherein the most 5?nucleotide and the most 3’ nucleotide of the loop are each linked to a base-paired nucleotide in a stem.

[0253] A “stem-loop structure” or a “hairpin” refers to a nucleic acid having a secondary structure that includes a region of nucleotides which are known or predicted to form a double strand (stem portion) that is linked on one side by a region of predominantly single -stranded nucleotides (loop portion). Such structures are well known in the art and these terms are used consistently with their known meanings in the art. As is known in the art, a stem-loop structure does not require exact base-pairing. Tirus, the stem may include one or more base mismatches. Alternatively, the base-pairing may be exact, i.e., not include any mismatches.

[0077] As used herein, the term “operably linked” or “under transcriptional control,” when used in conjunction with the description of a promoter, refers to the correct location and orientation in relation to a polynucleotide (e.g. , a coding sequence) to control the initiation of transcription by RNA polymerase and expression of tire coding sequence.Stringent hybridization

[0254] In general, “stringent hybridization” is performed at about 25°C below the thermal melting point (Tm) for the specific DNA hybrid under a particular set of conditions. “Stringent washing” is performed at temperatures about 5°C lower than the Tm for the specific DNA hybrid under a particular set of conditions. Tire Tm is the temperature at which 50% of the target sequence hybridizes to a perfectly matched probe. See Sambrook et al., Molecular Cloning: A Laboratory Manual, 2d ed.. Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y. (1989). page 9.51, hereby incorporated by reference. For purposes herein, “stringent conditions” are defined for solution phase hybridization as aqueous hybridization (i.e., free of formamide) in 6xSSC (where 20xSSC contains 3.0 M NaCl and 0.3 M sodium citrate), 1% SDS at 65°C for 8-12 hours, followed by two washes in 0.2xSSC, 0.1% SDS at 65°C for 20 minutes. It will be appreciated bythe skilled worker that hybridization at 65 °C will occur at different rates depending on a number of factors including the length and percent identity of the sequences which are hybridizing. Hybridization does not require the sequence of the polynucleotide to be 100% com piemen tan to the target polynucleotide. Hybridization also includes one or more segments such that intervening or adjacent segments that are not involved in the hybridization event (e.g., a loop structure or hairpin structure).

[0255] The nucleic acids (also referred to as polynucleotides) of this present invention may include both sense and antisense strands of RNA, cDNA, genomic DNA, and synthetic forms and mixed polymers of the above. They may be modified chemically or biochemically or may contain non -natural or derivatized nucleotide bases, as will be readily appreciated by those of skill in the art. Such modifications include, for example, labels, methylation, substitution of one or more of the naturally occurring nucleotides with an analog, intemucleotide modifications such as uncharged linkages (e.g., methyl phosphonates. phosphotriesters, phosphoramidates, carbamates, etc.), charged linkages (e.g., phosphorothioates. phosphorodithioates, etc.), pendent moieties (e.g., polypeptides), intercalators (e.g., acridine, psoralen, etc.), chelators, alkylators, and modified linkages (e.g., alpha anomeric nucleic acids, etc.) Also included are synthetic molecules that mimic polynucleotides in their ability to bind to a designated sequence via hydrogen bonding and other chemical interactions. Such molecules are known in tire art and include, for example, those in which peptide linkages substitute for phosphate linkages in the backbone of the molecule. Other modifications can include, for example, analogs in which the ribose ring contains a bridging moiety or other structure such as the modifications found in “locked” nucleic acids.Subject

[0256] As used herein, the term “subject” refers to an individual organism, for example, an individual mammal. In some embodiments, tire subject is a human. In some embodiments, the subject is anon-human mammal. In some embodiments, the subject is a non-human primate. In some embodiments, the subject is a rodent. In some embodiments, the subject is a sheep, a goat, a cattle, a cat, or a dog. In some embodiments, the subject is a vertebrate, an amphibian, a reptile, a fish, an insect, a fly, or a nematode. In some embodiments, the subject is a research animal. In some embodiments, the subject is genetically engineered, e.g., a genetically engineered non-human subject. The subject may be of either sex and at any stage of development. The terms “individual,” “subject,” “host,” and “patient,” used interchangeably herein.Synthetic nucleic acid

[0257] A “synthetic or artificial nucleic acid” refers nucleic acids that are non-naturally occuring sequences. Such sequences do not originate from, or are not known to be present in any living organism (e.g., based on sequence search in existing sequence databases).Targeted polynucleotide sequence

[0258] As used herein '‘targeted polynucleotide sequence” refers to a DNA polynucleotide that comprises a “target site” or “target sequence.” The terms “target site,” “target sequence,” “target protospacer DNA,” or “protospacer-like sequence” are used interchangeably herein to refer to a nucleic acid sequence present in a targeted polynucleotide sequence to which a DNA-targeting segment of a guide RNA will recognize and / or bind, provided sufficient conditions for binding exist. For example, the target site (or target sequence) 5'- GAGCATATC-3' within a targeted polynucleotide sequence is targeted by (or is bound by, or hybridizes with, or is complementary to) the RNA sequence 5'-GAUAUGCUC-3'. Suitable DNA / RNA binding conditions include physiological conditions normally present in a cell.

[0259] Other suitable DNA / RNA binding conditions (e.g., conditions in a cell -free system) are known in the art; see, e g., Sambrook, supra. Tire strand of the targeted polynucleotide sequence that is complementary to and hybridizes with the guide RNA is referred to as the “complementary strand” and the strand of the targeted polynucleotide sequence that is complementary to the “complementary strand” (and is therefore not complementary to the guide RNA) is referred to as the “noncomplementary strand” or “non-complementary strand.”Target site

[0260] As used herein, a “target site” as used herein is a polynucleotide (e.g., DNA such as genomic DNA) that includes a site or specific locus (“target site” or “target sequence”) targeted by a Cast 2a gene editing system disclosed herein. In the context of a Casl2a gene editing system disclosed herein that comprise an RNA-guided nuclease, a target sequence is the sequence to which the guide sequence of a guide nucleic acid (e.g., guide RNA) will hybridize. For example, the target site (or target sequence) 5'-GTCAATGGACC-3' (SEQ ID NO:2663) within a target nucleic acid is targeted by (or is bound by, or hybridizes with, or is complementary to) the sequence 5'-GGTCCATTGAC-3' (SEQ ID NO:2664). Suitable hybridization conditions include physiological conditions normally present in a cell. For a double stranded target nucleic acid, the strand of the target nucleic acid that is complementary to and hybridizes with the guide RNA is referred to as the “complementary strand” or “target strand”; while the strand of the target nucleic acid that is complementary to the “target strand” (and is therefore not complementary to the guide RNA) is referred to as the “non-target strand” or “non-complementary strand.”Therapeutic

[0261] Tire term “therapeutic” as used herein means a treatment and / or prophylaxis. A therapeutic effect is obtained by suppression, diminution, remission, or eradication of at least one sign or symptom of a disease or disorder state.Therapeutically effective amount

[0262] The term '‘therapeutically effective amount’’ refers to the amount of the subject compound that will elicit the biological or medical response of a tissue, system, or subject that is being sought by the researcher, veterinarian, medical doctor or other clinician. The term “therapeutically effective amount” includes that amount of a compound that, when administered, is sufficient to prevent development of, or alleviate to some extent, one or more of the signs or symptoms of the disorder or disease being treated. Tire therapeutically effective amount will vary depending on the compound, the disease and its severity and the age, weight, etc., of the subject to be treated.Treat or treatment

[0263] To “treat” a disease as the term is used herein, means to reduce the frequency or severity of at least one sign or symptom of a disease or disorder experienced by a subject. Treatment

[0264] As used herein, the terms “treatment,” “treat,” and “treating,” refer to a clinical intervention aimed to reverse, alleviate, delay the onset of, or inhibit the progress of a disease or disorder, or one or more symptoms thereof, as described herein. In some embodiments, treatment may be administered after one or more symptoms have developed and / or after a disease has been diagnosed. In other embodiments, treatment may be administered in the absence of symptoms, e.g., to prevent or delay onset of a symptom or inhibit onset or progression of a disease. For example, treatment may be administered to a susceptible individual prior to the onset of symptoms (e g ., in light of a history of symptoms and / or in light of genetic or other susceptibility factors). Treatment may also be continued after symptoms have resolved, for example, to prevent or delay their recurrence.Upstream and downstream

[0265] As used herein, the terms “upstream” and “downstream” are terms of relativity that define the linear position of at least two elements located in a nucleic acid molecule (whether single or double -stranded) that is orientated in a 5'-to-3' direction. A first element is said to be upstream of a second element in a nucleic acid molecule where the first element is positioned somewhere that is 5' to the second element. Conversely, a first element is downstream of a second element in a nucleic acid molecule where the first element is positioned somewhere that is 3' to the second element.Variant

[0266] As used herein the tenn “variant” should be taken to mean the exhibition of qualities that have a pattern that deviates from what occurs in nature, e.g., a variant retron RT is retron RT comprising one or more changes in amino acid residues as compared to a wild type retron RT amino acid sequence. The term “variant” encompasses homologous proteins having at least 75%, or at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 99% percent identity with a reference sequence and having the same or substantially the same functional activity or activities as the reference sequence. The term also encompassesmutants, truncations, or domains of a reference sequence, and which display the same or substantially the same functional activity or activities as the reference sequence.Vector

[0267] As used herein, tire term “vector” permits or facilitates the transfer of a polynucleotide from one environment to another. It is a replicon such as a plasmid, phage, or cosmid into which another DNA segment may be inserted so as to bring about the replication of the inserted segment. Generally, a vector is capable of replication when associated with the proper control elements. The term “vector” may include cloning and expression vectors, as well as viral vectors and integrating vectors.Wild type

[0268] As used herein the term “wild type” is a term of the art understood by skilled persons and means the typical form of an organism, strain, gene, protein, or characteristic as it occurs in nature as distinguished from mutant or variant formsB. Chemical DefinitionsAlkyl

[0269] “Alkyl” refers to a straight or branched hydrocarbon chain radical consisting solely of carbon and hydrogen atoms, which is saturated or unsaturated (i.e., contains one or more double and / or triple bonds), having from one to thirty or more carbon atoms (e.g.. C1-C24 alkyl), one to twelve carbon atoms (C1-C12 alkyl), one to eight carbon atoms (C1-C8 alkyl) or one to six carbon atoms (C1-C6 alkyl) and which is attached to the rest of the molecule by a single bond, e.g., methyl, ethyl, n propyl, 1 -methylethyl (iso propyl), n butyl, n pentyl, 1,1 dimcthylcthyl (t butyl), 3 methylhexyl, 2 methylhexyl, cthcnyl, propyl cnyl, but-l-cnyl, pent- 1-enyl, penta-1, 4-dienyl, ethynyl, propynyl, butynyl, pentynyl, hexynyl, and the like. Alkyl groups that include one or more units of unsaturation (one or more double and / or triple bond) can be C2-C24, C2-C12, C2-C8 or C2-C6 groups, for example. Unless specifically stated otherwise, an alkyl group is optionally substituted. The term “alkyl,” by itself or as part of another substituent means, unless otherwise stated, a straight or branched chain hydrocarbon having the number of carbon atoms designated (i.e., Cl-6 means one to six carbon atoms) and includes straight, branched chain, or cyclic substituent groups.Alkoxy

[0270] For example, the term “alkoxy” employed alone or in combination with other terms means, unless otherwise stated, an alkyl group having the designated number of carbon atoms, as defined above, connected to the rest of the molecule via an oxygen atom, such as, for example, methoxy, ethoxy, 1 -propoxy, 2 -propoxy (isopropoxy) and the higher homologs and isomers. Preferred are (C1-C3) alkoxy, particularly ethoxy and methoxy.Alkylamino

[0271] As used herein, the terms '‘alkoxy,” “alkylamino” and “alkylthio” are used in their conventional sense, and refer to alkyl groups linked to molecules via an oxygen atom, an amino group, a sulfur atom, respectively.Alkylene

[0272] “Alkylene” or “alkylene chain” refers to a straight or branched divalent hydrocarbon chain consisting solely of carbon and hydrogen, which is saturated or unsaturated (i.e., contains one or more double (alkenylene) and / or triple bonds (alkynylcnc)). and having, for example, from one to thirty or more carbon atoms (e.g., C1-C24 alkylene), one to fifteen carbon atoms (C1-C15 alkylene), one to twelve carbon atoms (C1-C12 alkylene), one to eight carbon atoms (C1-C8 alkylene), one to six carbon atoms (C1-C6 alkylene), two to four carbon atoms (C2-C4 alkylene), one to two carbon atoms (C1-C2 alkylene), e.g., methylene, ethylene, propylene, n-butylene, ethenylene, propenylene, n-butenylene, propynylene, n-butynylene, and the like. Alkylene groups that include one or more units of unsaturation (one or more double and / or triple bond) can be C2-C24, C2-C12, C2-C8 or C2-C6 groups, for example. Tire alkylene chain is attached to the rest of the molecule through a single or double bond and to the radical group through a single or double bond. The points of attachment of the alkylene chain to the rest of the molecule and to the radical group can be through one carbon or any two carbons within the chain. Unless stated otherwise specifically in the specification, an alkylene chain may be optionally substituted.Amino aryl

[0273] As used herein, the term “amino aryl" refers to an aryl moiety which contains an amino moiety. Such amino moieties may include, but are not limited to primary amines, secondary amines, tertiary amines, quaternary amines, masked amines, or protected amines. Such tertiary amines, masked amines, or protected amines may be converted to primary amine or secondary amine moieties. Additionally, the amine moiety may include an amine- like moiety which has similar chemical characteristics as amine moieties, including but not limited to chemical reactivity.Aromatic

[0274] As used herein, the term “aromatic” refers to a carbocycle or heterocycle with one or more polyunsaturated rings and having aromatic character, i.e. having (4n + 2) delocalized p (pi) electrons, where n is an integer.Aryl

[0275] As used herein, tire term “aryl,” employed alone or in combination with other terms, means, unless otherwise stated, a carbocyclic aromatic system containing one or more rings (typically one, two or three rings) wherein such rings may be attached together in a pendent manner, such as a biphenyl, or may be fused,such as naphthalene. Examples include phenyl, anthracyL and naphthyl. Preferred are phenyl and naphthyl, most preferred is phenyl.Cvcloakylene

[0276] “Cycloalkylene” is a divalent cycloalkyl group. Unless otherwise stated specifically in the specification, a cycloalkylene group may be optionally substituted.Cvcloalkyl

[0277] ‘ ‘Cycloalkyl” or “carbocyclic ring” refers to a stable non aromatic monocyclic or polycyclic hydrocarbon radical consisting solely of carbon and hydrogen atoms, which may include fused or bridged ring systems, having from three to fifteen carbon atoms, preferably having from three to ten carbon atoms, and which is saturated or unsaturated and attached to the rest of the molecule by a single bond. Monocyclic radicals include, for example, cyclopropyl, cyclobutyl, cyclopentyl, cyclohexyl, cycloheptyl, and cyclooctyl. Polycyclic radicals include, for example, adamantyl, norbomyl. decalinyl, 7,7 dimethyl bicyclo[2.2. l]heptanyl, and the like. Unless specifically stated otherwise, a cycloalkyl group is optionally substituted.Halo

[0278] As used herein, tire tenn “halo” or “halogen” alone or as part of another substituent means, unless otherwise stated, a fluorine, chlorine, bromine, or iodine atom, preferably, fluorine, chlorine, or bromine, more preferably, fluorine or chlorine.Heteroalkyl

[0279] As used herein, the term “heteroalkyl” by itself or in combination with another term means, unless otherwise stated, a stable straight or branched chain alkyl group consisting of the stated number of carbon atoms and one or tw o or more heteroatoms typically selected from the group consisting of O, N, Si, P, and S, and wherein the nitrogen and sulfur atoms may be optionally oxidized and the nitrogen heteroatom may be a primary, secondary, tertiary or quaternary nitrogen. The heteroatom(s) may be placed at any position of the heteroalkyl group, including between the rest of the heteroalkyl group and the fragment to which it is attached, as well as attached to the most distal carbon atom in the heteroalkyl group. Examples of heteroalkyl groups include: -O-CH2-CH2-CH3, -CH2-CH2-CH2-OH, - CH2-CH2-NH-CH3, -CH2-S-CH2-CH3, and - CH2CH2-S(=O)-CH3. Up to tw o heteroatoms may be consecutive, such as. for example, -CH2-NH-OCH3. or -CH2-CH2-S-S-CH3.Heteroaryl

[0280] As used herein, the term "hcteroaty l" or “heteroaromatic” refers to aryl groups which contain at least one heteroatom typically selected from N, O, Si, P, and S; wherein the nitrogen and sulfur atoms may be optionally oxidized, and the nitrogen atom(s) may be optionally teriatry or quatemized. Heteroaryl groupsmay be substituted or unsubstituted. A heteroaryl group may be attached to the remainder of the molecule through a heteroatom. A polycyclic heteroaryl may include one or more rings that are partially saturated. Examples include tetrahydroquinoline, 2,3-dihydrobenzofuryl, 1 -pyrrolyl, 2-pyrrolyl, 3-pyrrolyl, 3- pyrazolyl, 2 -imidazolyl, 4-imidazolyl, pyrazinyl, 2-oxazolyl, 4-oxazolyl, 2-phenyl-4- oxazolyl, 5 -oxazolyl, 3- isoxazolyl, 4-isoxazolyl, 5-isoxazolyl, 2-thiazolyl, 4-thiazolyl, 5- thiazolyl, 2-furyl, 3-furyl, 2-thienyl, 3- thienyl. 2-pyridyl, 3-pyridyl, 4-pyridyl, 2-pyrimidyl, 4- pyrimidyl, 5 -benzothiazolyl, purinyl, 2- benzimidazolyl, 5-indolyl, 1 -isoquinolyl, 5- isoquinolyl, 2-quinoxalinyl, 5-quinoxalinyl, 3-quinolyl, and 6- quinolyl. Examples of non-aromatic heterocycles include monocyclic groups such as aziridine, oxirane, thiirane, azetidine, oxetane, thietane, pyrrolidine, pyrroline, imidazoline, pyrazolidine, dioxolane, sulfolane,2.3 -dihydrofuran, 2,5-dihydrofuran. tetrahydrofuran, thiophane, piperidine, 1,2, 3, 6- tetrahydropyridine, 1,4- dihydropyridine, piperazine, morpholine, thiomorpholine, pyran, 2,3- dihydropyran, tetrahydropyran. 1,4- dioxane, 1,3-dioxane. homopiperazine, homopiperidine, 1,3-dioxepane. 4,7-dihydro-l,3-dioxepin and hexamethyleneoxide. Examples of heteroaryl groups include pyridyl, pyrazinyl, pyrimidinyl (particularly 2- and 4-pyrimidinyl), pyridazinyl, thienyl, furyl, pyrrolyl (particularly 2-pyrrolyl), imidazolyl, thiazolyl, oxazolyl, pyrazolyl (particularly 3- and 5-pyrazolyl), isothiazolyl, 1,2,3-triazolyl, 1,2,4-triazolyl, 1,3,4- triazolyl, tetrazolyl, 1,2,3-thiadiazolyl, 1,2,3-oxadiazolyl, 1,3,4-thiadiazolyl and 1,3,4- oxadiazolyl. Examples of polycyclic heterocycles include indolyl (particularly 3-, 4-, 5-. 6- and 7-indolyl), indolinyl, quinolyl, tetrahydroquinolyl, isoquinolyl (particularly 1- and 5- isoquinolyl), 1.2.3.4-tetrahydroisoquinolyl. cinnolinyl, quinoxalinyl (particularly 2- and 5- quinoxalinyl), quinazolinyl, phthalazinyl, 1,8-naphthyridinyl,1.4-benzodioxanyl, coumarin, dihydrocoumarin, 1,5-naphthyridinyl, bcnzof'iiryl (particularly 3-, 4-, 5-, 6- and 7-benzofuiy 1), 2,3-dihydrobenzofuryl, 1,2-benzisoxazolyl, benzothienyl (particularly 3-, 4-, 5-, 6-, and 7- benzothienyl), benzoxazolyl, benzothiazolyl (particularly 2-benzothiazolyl and 5- benzothiazolyl), purinyl, benzimidazolyl (particularly 2-benzimidazolyl), benztriazolyl, thioxanthinyl, carbazolyl, carbolinyl, acridinyl, pyrrolizidinyl, and quinolizidinyl. The aforementioned listing of heterocyclyl and heteroaryl moieties is intended to be representative and not limiting.Heterocyclyl

[0281] As used herein, tire tenn “heterocyclyl” or “heterocyclic ring” refers to a stable 3- to 18-membered non-aromatic ring radical which consists of two to twelve carbon atoms and from one to six heteroatoms typically selected from the group consisting of N. 0, Si, P, and S. Unless stated otherwise specifically in the specification, the heterocyclyl radical may be a monocyclic, bicyclic, tricyclic or tetracyclic ring system, which may include fused or bridged ring systems: and the nitrogen, carbon or sulfur atoms in the heterocyclyl radical may be optionally oxidized; the nitrogen atom may be optionally quatemized; and the heterocyclyl radical may be partially or fully saturated. Examples of such heterocyclyl radicals include, but are not limitedto, dioxolanyl, thienyl[l,3]dithianyl, decahydroisoquinolyl, imidazolinyl, imidazolidinyl, isothiazolidinyl, isoxazolidinyl, morpholinyl, octahydroindolyl, octahydroisoindolyl, 2-oxopiperazinyl, 2-oxopiperidinyl, 2- oxopyrrolidinyl, oxazolidinyl, piperidinyl, piperazinyl, 4-piperidonyl, pyrrolidinyl, pyrazolidinyl, quinuclidinyl, thiazolidinyl. tetrahydrofuryl, trithianyl, tetrahydropyranyl, thiomorpholinyl, thiamorpholinyl, 1-oxo-thiomorpholinyl, and 1,1-dioxo-thiomorpholinyl. Unless specifically stated otherwise, a heterocyclyl group may be optionally substituted.Substituents

[0282] As described herein, compounds of the present disclosure may contain '‘optionally substituted” moieties. In general, the term “substituted”, whether preceded by the term “optionally” or not, means that one or more hydrogens of the designated moiety are replaced with a suitable substituent. Unless otherwise indicated, an “optionally substituted” group may have a suitable substituent at each substitutable position of the group, and when more than one position in any given structure may be substituted with more than one substituent selected from a specified group, the substituent may be either the same or different at every position. Combinations of substituents envisioned by this disclosure are preferably those that result in the fonnation of stable or chemically feasible compounds. Tire term “stable”, as used herein, refers to compounds that are not substantially altered when subjected to conditions to allow for their production, detection, and, in certain embodiments, their recovery, purification, and use for one or more of the purposes disclosed herein.

[0283] Suitable monovalent substituents on a substitutable carbon atom of an “optionally substituted” group are independently halogen; — (CH2)0-4R°; — (CH2)0-4OR°; — O(CH2)0-4R°, — O— (CH2)0-4C(O)OR°;— (CH2)0-4CH(OR°)2; — (CH2)0-4SR°; — (CH2)0-4Ph, which may be substituted with R°; — (CH2)0- 4O(CH2)0-lPh which may be substituted with R°; — CH=CHPh, which may be substituted with R°; — (CH2)0-40(CH2)0-1 -pyridyl which may be substituted with R°; — NO2; — CN; — N3; — (CH2)0-4N(R°)2; — (CH2)0-4N(R°)C(O)R°; — N(R°)C(S)R°; — (CH2)0-4N(R°)C(O)NR° 2; — N(R°)C(S)NR° 2: — (CH2)0- 4N(R°)C(0)0R°; — N(R°)N(R°)C(O)R°; — N(R°)N(R°)C(0)NR° 2; — N(R°)N(R°)C(0)0R°; — (CH2)0- 4C(O)R°; — C(S)R°; — (CH2)0-4C(O)OR°: — (CH2)0-4C(O)SR°; — (CH2)0-4C(O)OSiR° 3; — (CH2)0- 4OC(O)R°; — OC(O)(CH2)0-4SR°, SC(S)SR°; — (CH2)0-4SC(O)R°; — (CH2)0-4C(O)NR° 2; — C(S)NR° 2; — C(S)SR°; — SC(S)SR°, — (CH2)0-4OC(O)NR° 2; — C(0)N(0R°)R°; — C(O)C(O)R°; — C(O)CH2C(O)R°; — C(NOR°)R°; — (CH2)0-4SSR°: — (CH2)0-4S(O)2R°; — (CH2)0-4S(O)2OR°; — (CH2)0-4OS(O)2R°; — S(O)2NR° 2; — (CH2)0-4S(O)R°; — N(R°)S(O)2NR° 2; — N(R°)S(O)2R°; — N(0R°)R°; — C(NH)NR° 2; — P(O)2R°; — P(O)R° 2; — OP(O)R° 2; — OP(O)(OR°)2; SiR° 3; — (Cl-4 straight or branched alkylenejO — N(R°)2; or — (Cl-4 straight or branched alkylene)C(O)O — N(R°)2,wherein each R° may be substituted as defined below and is independently hydrogen, Cl -6 aliphatic, — CH2Ph. — O(CH2)0-lPh. — CH2-(5-6 membered heteroaryl ring), or a 5-6-membered saturated, partially unsaturated, or aryl ring having 0-4 heteroatoms independently selected from nitrogen, oxygen, or sulfur, or, notwithstanding the definition above, two independent occurrences of R°, taken together with their intervening atom(s), form a 3-12-membered saturated, partially unsaturated, or aryl mono- or bicyclic ring having 0-4 heteroatoms independently selected from nitrogen, oxygen, or sulfur, which may be substituted as defined below.

[0284] Suitable monovalent substituents on R° (or tire ring formed by taking two independent occurrences of R° together with their intervening atoms), are independently halogen, — (CH2)0-2R*, -(haloR*). — (CH2)0-2OH, — (CH2)0-2OR*, — (CH2)0-2CH(OR*)2: — O(haloR’), — CN, — N3, — (CH2)0-2C(O)R*, — (CH2)0-2C(O)OH, — (CH2)0-2C(O)OR*, — (CH2)0-2SR*, — (CH2)0-2SH, — (CH2)0-2NH2, — (CH2)0- 2NHR*, — (CH2)0-2NR* 2, — NO2, — SiR* 3, — OSiR* 3, — C(O)SR*, — (Cl-4 straight or branched alkylene)C(O)OR*. or — SSR* wherein each R* is unsubstituted or where preceded by "halo" is substituted only with one or more halogens, and is independently selected from Cl-4 aliphatic, — CH2Ph, — O(CH2)0- IPh, or a 5-6-membered saturated, partially unsaturated, or aryl ring having 0-4 heteroatoms independently selected from nitrogen, oxygen, or sulfur. Suitable divalent substituents on a saturated carbon atom of R° include =0 and =S.

[0285] Suitable divalent substituents on a saturated carbon atom of an “optionally substituted” group include the following: =0, =S, =NNR*2, =NNHC(0)R*,

[0286] =NNHC(O)OR*. =NNHS(O)2R*, =NR*, =N0R*, — O(C(R*2))2-3O— , or — S(C(R*2))2-3S— , wherein each independent occurrence of R* is selected from hydrogen, Cl -6 aliphatic which may be substituted as defined below, or an unsubstituted 5-6-membered saturated, partially unsaturated, or aryl ring having 0-4 heteroatoms independently selected from nitrogen, oxygen, or sulfur. Suitable divalent substituents that are bound to vicinal substitutable carbons of an “optionally substituted” group include: — O(CR*2)2-3O — , wherein each independent occurrence of R* is selected from hydrogen, Cl-6 aliphatic which may be substituted as defined below, or an unsubstituted 5-6-membered saturated, partially unsaturated, or aryl ring having 0-4 heteroatoms independently selected from nitrogen, oxygen, or sulfur.

[0287] Suitable substituents on the aliphatic group of R* include halogen, — R*, - (haloR*), — OH, — OR*, — O(haloR*), — CN, — C(O)OH, — C(O)OR*, — NH2, — NHR*, —NR* 2, or — NO2, wherein each R* is unsubstituted or where preceded by “halo” is substituted only with one or more halogens, and is independently Cl-4 aliphatic, — CH2Ph, — O(CH2)0- IPh, or a 5-6-membered saturated, partially unsaturated, or aryl ring having 0-4 heteroatoms independently selected from nitrogen, oxygen, or sulfur.

[0288] Suitable substituents on a substitutable nitrogen of an “optionally substituted” group include — R . — NRf2, — C(O)Rf, — C(O)ORt, — C(O)C(O)Rt, — C(O)CH2C(O)Rt,— S(O)2Rt, — S(O)2NR:2, — C(S)NRt2, — C(NH)NRf2, or — N(R')S(O)2R'; wherein each R' is independently hydrogen, Cl -6 aliphatic which may be substituted as defined below, unsubstituted — OPh, or an unsubstituted 5-6-membered saturated, partially unsaturated, or aryl ring having 0-4 heteroatoms independently selected from nitrogen, oxygen, or sulfur, or. notwithstanding the definition above, two independent occurrences of R:. taken together with their intervening atom(s) form an unsubstituted 3-12-membered saturated, partially unsaturated, or aryl mono- or bicyclic ring having 0-4 heteroatoms independently selected from nitrogen, oxygen, or sulfur.

[0289] Suitable substituents on the aliphatic group of R ' are independently halogen, — R*, -(haloR*), — OH, —OR*, — O(haloR*), — CN, — C(O)OH, — C(O)OR*, — NH2, — NHR*, —NR* 2, or — NO2, wherein each R* is unsubstituted or where preceded by “halo” is substituted only with one or more halogens, and is independently Cl-4 aliphatic. — CH2Ph, — O(CH2)0-lPh. or a 5-6-membered saturated, partially unsaturated, or aryl ring having 0-4 heteroatoms independently selected from nitrogen, oxygen, or sulfur.

[0290] Heteroatoms such as nitrogen may have hydrogen substituents and / or any permissible substituents of organic compounds described herein which satisfy the valences of the heteroatoms. It is understood that "substitution" or "substituted" includes the implicit proviso that such substitution is in accordance with permitted valence of the substituted atom and the substituent, and that the substitution results in a stable compound, i.e.. a compound that does not spontaneously undergo transformation, for example, by rearrangement, cyclization, or elimination.

[0291] In a broad aspect, the permissible substituents include acyclic and cyclic, branched and unbranched, carbocyclic and heterocyclic, aromatic and nonaromatic substituents of organic compounds. Illustrative substituents include, for example, those described herein. Hie permissible substituents can be one or more and the same or different for appropriate organic compounds. Hie heteroatoms such as nitrogen may have hydrogen substituents and / or any permissible substituents of organic compounds described herein which satisfy the valencies of the heteroatoms.

[0292] In various embodiments, the substituent is selected from alkoxy, aryloxy, alkyl, alkenyl, alkynyl, amide, amino, aryl, arylalkyl, carbamate, carboxy, cyano, cycloalkyl, ester, ether, formyl, halogen, haloalkyl, heteroaryl, heterocyclyl, hydroxyl, ketone, nitro, phosphate, sulfide, sulfinyl, sulfonyl, sulfonic acid, sulfonamide, and thioketone, each of which optionally is substituted with one or more suitable substituents. In some embodiments, the substituent is selected from alkoxy, aryloxy, alkyl, alkenyl, alkynyl, amide, amino, aryl, arylalkyl, carbamate, carboxy, cycloalkyl, ester, ether, formyl, haloalkyl, heteroaryl, heterocyclyl, ketone, phosphate, sulfide, sulfinyl, sulfonyl, sulfonic acid, sulfonamide, and thioketone, wherein each of the alkoxy, aryloxy, alkyl, alkenyl, alkynyl, amide, amino, aryl, arylalkyl, carbamate, carboxy, cycloalkyl, ester,ether, formyl, haloalkyl, heteroaryl, heterocyclyl, ketone, phosphate, sulfide, sulfinyl, sulfonyl, sulfonic acid, sulfonamide, and thioketone can be further substituted with one or more suitable substituents.

[0293] Examples of substituents include, but are not limited to, halogen, azide, alkyl, aralkyl, alkenyl, alkynyl, cycloalkyl, hydroxyl, alkoxyl, amino, nitro, sulfhydryl, imino, amido, phosphonate, phosphinate, carbonyl, carboxyl, silyl, ether, alkylthio, sulfonyl, sulfonamide, ketone, aldehyde, thioketone, ester, heterocyclyl, -CN, aryl, aryloxy, perhaloalkoxy, aralkoxy, heteroaryl, heteroaryloxy. heteroarylalkyl, heteroaralkoxy, azido, alkylthio, oxo, acylalkyl, carboxy esters, carboxamido, acyloxy, aminoalkyl, alkylaminoaryl, alkylaryl, alkylaminoalkyl, alkoxyaryl, arylanuno. aralkylamino, alkylsulfonyl, carboxamidoalkylaryl, carboxamidoaryl, hydroxyalkyl, haloalkyl, alkylaminoalkylcarboxy, aminocarboxamidoalkyl, cyano, alkoxyalkyl, perhaloalkyl, arylalkyloxyalkyl, and the like. In some embodiments, the substituent is selected from cyano, halogen, hydroxyl, and nitro.

[0294] Throughout the disclosure, chemical substituents described in Markush structures are represented by variables. Where a variable is given multiple definitions as applied to different Markush formulas in different sections of the disclosure, it is to be understood that each definition should only apply to the applicable fonnula in the appropriate section of the disclosure.Abbreviations

[0295] As used herein, the following abbreviations and initialisms have the indicated meanings:

[0296] The details of one or more embodiments of the disclosure are set forth in the accompanying description below. Although any materials and methods similar or equivalent to those described herein can be used in the practice or testing of the present disclosure, the preferred materials and methods are now described. Other features, objects and advantages of the disclosure will be apparent from the description. In the description, the singular forms also include the plural unless the context clearly dictates otherwise. Unless defined otherwise, all technical and scientific tenns used herein have tire same meaning as commonly- understood by one of ordinary skill in the art to which this disclosure belongs. In the case of conflict, the present description will control.C. Casl2a (or Cas Type V) Sequences

[0297] Hie present disclosure provides a Cas Type V polypeptide comprising an amino acid sequence having at least 80%, 85%, 90%, 91%. 92%. 93%. 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity to any one of the amino acid sequence sequences selected from SEQ ID NO: 334 (No.ID405), SEQ ID NO: 58 (No. TD414), SEQ ID NO: 564 (No. ID418), SEQ ID NO: 335 (No. ID406), SEQ ID NO: 331 (No. ID411), SEQ ID NO: 20 (No. ID415), and SEQ ID NO: 445 (No. ID419), wherein the poly-peptide comprises one or more substitutions at positions selected from: Q492, T306, L748, A739, F297, N536, D851, N770. D675, K955, S972. N798, K298, H853, E377, V295. K852, F854, N1015, T845. E1019, K745, K292, Q565. K745, 1723, K821, K721, A810, K768. S16, 128. V312, V929, L474. H470, E402, K207, V294, R212, F79. Q299, Q465. F537, K1284, V929, DI 192. Q1203. T838. N735, 11063. K933, N1249, 197. R212, A1251, N1264, R1138, N588, E584, V453, Q475, N76, R304, 1471, E544, V254, D231, H496, V534, D577, 128, A183, D341, K595, V501, S568, H470, V254, D231 and F197, based on the amino acid sequence provided in SEQ ID NO: 334 (No. ID405), or one or more corresponding substitutions in any of the amino acid sequences selected from: SEQ ID NO: 58 (No. ID414), SEQ ID NO: 564 (No. ID418), SEQ ID NO: 335 (No. ID406), SEQ ID NO: 331 (No. ID411), SEQ ID NO: 20 (No. ID415), and SEQ ID NO: 445 (No.ID419).

[0298] Particular substitutions and combinations of substitutions are disclosed elsewhere herein. One of skill in the art will be able to identify the corresponding residue in any homologous protein and in the respective encoding nucleic acid by methods well known in the art, e.g., by- sequence alignment anddetermination of homologous residues. The Cas Type V polypeptide may be isolated, engineered, and / or non-naturally occuring.

[0299] In some embodiments, Cas Type V polypeptides of the present disclosure are capable of inducing indel formation when contacted with a population of target sequences and a guide RNA targeting the Cas Type V polypeptide to the target sequence under conditions suitable for inducing indel formation. The target sequence may be a DNA sequence. In some embodiments, the target sequence may be associated with a disease or disorder. In some embodiments, the substitutions in the Cas Type V polypeptides disclosed herein enhance tire ability of the Cas Type V polypeptide to induce indels in comparison with Cas Type V polypeptides that do not comprise the substitutions. For example, in some embodiments the Cas Type V polypeptide induces indel formation when contacted with a population of target sequences and a guide RNA targeting the Cas Type V polypeptide to the target sequence under conditions suitable for inducing indel formation, wherein the contacting results in an increase in the percentage of target polynucleotide sequences comprising an indel in the population of target sequences of at least 20%. at least 30%, at least 40%, or at least 50% as compared to the percentage of target polynucleotide sequences comprising an indel in the population when the population is contacted with a reference polypeptide comprising the sequence of SEQ ID NO: 1554 (No. ID405-1) and a guide RNA that targets the reference polypeptide to the target polynucleotide sequence.

[0300] The Cas Type V polypeptide may be associated with one or more additional accessory proteins having genome modifying functions, including recombinases, invertases, nucleases, polymerases, ligases, deaminases, reverse transcriptases, or epigenetic modifying functions. In various embodiments, the accessory proteins may be provided separately. In other embodiments, the accessory proteins may be fused to the Cas Type V polypeptide, optionally with a linker. The disclosure provides a fusion protein comprising the Cas Type V polypeptide fused to a heterologous amino acid sequence, such as an accessory protein as described herein.

[0301] The disclosure provides a gene editing system comprising: (a) one or more Cas Type V polypeptides of tire disclosure or one or more polynucleotides encoding the Cas Type V polypeptide(s); and (b) one or more polynucleotide sequences comprising a guide RNA (gRNA) or one or more polynucleotides encoding the gRNA, wherein the gRNA comprises a complementary sequence to that of a targeted polynucleotide sequence. The guide RNA may be a guide RNA as disclosed herein.

[0302] Tire present disclosure also provides polypeptides comprising (i) a Cas Type V polypeptide, (ii) one, two, three or more Nuclear Localization Sequences (NLS); and optionally (iii) one or more linker sequences.The Cas Type V polypeptide may be any of the Cas Type V polypeptides disclosed herein. The polypeptides may comprise any of the NLS sequences, linker sequences, and / or structures as described elsewhere herein.

[0303] In some embodiments, polypeptides of the present disclosure are capable of inducing indel formation when contacted with a population of target sequences and a guide RNA targeting the polypeptide to the target sequence under conditions suitable for inducing indel formation. Tire target sequence may be a DNA sequence. In some embodiments, tire target sequence may be associated with a disease or disorder. In some embodiments, the NLS sequences, linker sequences, and / or structures of the polypeptides disclosed herein enhance tire ability of the polypeptide to induce indels in comparison with polypeptides that do not comprise the structures, NLS sequences, and linker sequences. For example, in some embodiments the polypeptide induces indel formation when contacted with a population of target sequences and a guide RNA targeting the polypeptide to the target sequence under conditions suitable for inducing indel formation, wherein the contacting results in an increase in the percentage of target poly nucleotide sequences comprising an indel in the population of target sequences of at least 40%, at least 50%, at least 60%. or at least 70% as compared to the percentage of target polynucleotide sequences comprising an indel in the population when the population is contacted with a reference polypeptide comprising the structure N-[SEQ ID NO: 1554 (ID405-l)]-[SEQ ID NO: SEQ ID NO: 1556]-C.

[0304] The polypeptide may be associated with one or more additional accessory proteins having genome modifying functions, including recombinases, invertases, nucleases, polymerases, ligases, deaminases, reverse transcriptases, or epigenetic modifying functions. In various embodiments, the accessory proteins may be provided separately. In other embodiments, the accessory proteins may be fused to the polypeptide, optionally with a linker. The disclosure provides a fusion protein comprising the polypeptide fused to a heterologous amino acid sequence, such as an accessory protein as described herein.

[0305] The disclosure provides a gene editing system comprising: (a) one or more polypeptides of the disclosure or one or more polynucleotides encoding the polypeptide(s); and (b) one or more polynucleotide sequences comprising a guide RNA (gRNA) or one or more polynucleotides encoding the gRNA, wherein the gRNA comprises a complementary sequence to that of a targeted polynucleotide sequence. The guide RNA may be a guide RNA as disclosed herein.

[0306] The disclosure provides a method of modifying a targeted polynucleotide sequence, the method comprising contacting the targeted polynucleotide sequence with a gene editing system as disclosed herein. Hie method may be carried out in vitro, ex vivo or in vivo. In some embodiments, the disclosure provides a method of modifying a targeted polynucleotide sequence, the method comprising introducing into a cell the gene editing system of the disclosure.

[0307] The disclosure provides one or more polynucleotides encoding: a Cas Type V polypeptide or polypeptide as disclosed herein, a fusion protein as disclosed herein, or a gene editing system as disclosed herein. The disclosure further provides one or more vectors comprising the one or more polynucleotides. The disclosure further provides a cell comprising the one or more vectors or the one or more polynucleotides.

[0308] Any of the Cas Type V polypeptides, polypeptides, fusion proteins, gene editing systems, polynucleotides, vectors, or pharmacal compositions disclosed herein can be used in medicine, for example for use in treating or preventing a disease or disorder as disclosed elsewhere herein.

[0309] Exemplary amino acid and protein coding sequences of Cas Type V polypeptides of tire present disclosure are provided in Appendix B.

[0310] Hie present disclosure provides Casl2a (or Cas Type V) polypeptides and nucleic acid molecules encoding same for use in the Casl2a-based gene editing systems described herein for use in various applications, including precision gene editing in cells, tissues, organs, or organisms. In various embodiments, the Cas 12a-based gene editing systems comprise (a) a Cas 12a (or Cas Type V) polypeptide (or a nucleic acid molecule encoding a Casl2a (or Cas Type V) polypeptide) and (b) a Casl2a (or Cas Type V) guide RNA which is capable of associating with a Cas 12a (or Cas Type V) polypeptide to form a complex such that the complex localizes to a target nucleic acid sequence (e.g., a genomic or plasmid target sequence) and binds thereto. In various embodiments, the Cas 12a (or Cas Type V) polypeptide has a nuclease activity which results in the cutting of both strands of DNA.

[0311] As outlined in B. Paul. Biomedical Journal, Vol. 43, No. 1, February 2020 pages 8-17. the CRISPR- Cas systems are classified into two classes (Classes 1 and 2) that are subdivided into six types (types I through VI). Class 1 (types I, III and IV) systems use multiple Cas proteins in their CRISPR ribonucleoprotein effector nucleases and Class 2 systems (types II, V and VI) use a single Cas protein. Class 1 CRISPR-Cas systems are most commonly found in bacteria and archaea, and comprise ~90% of all identified CRISPR-Cas loci. The Class 2 CRISPR-Cas systems, comprising the remaining ~10%, exists almost exclusively in bacteria, and assemble a ribonucleoprotein complex, consisting of a CRISPR RNA (crRNA) and a Cas protein. The crRNA contains information to target a specific DNA sequence. These multidomain effector proteins achieve interference by complementarity between the crRNA and the target sequence after recognition of the PAM (Protospacer Adjacent Motif) sequence, which is adjacent to the target DNA. These ribonucleoprotein complexes have been redesigned for precise genome editing by providing a crRNA with a redesigned guide sequence, which is complementary to the sequence of the targeted DNA. Tire most widely characterized CRISPR-Cas system is tire type 11 subtype Il-A that is found in Streptococcus pyogenes (Sp), which uses the protein SpCas9, Cas9 was the first Cas-protein engineered for use in geneediting. Class 2 type V is further classified into 4 subtypes (V-A, V-B, V-C, V-U). At present, V-C and V- U remain widely uncharacterised and no structural information on these systems is available. V-A encodes the protein Casl2a (also known as Cpfl) and recently several high resolution structures of Cast 2a have provided an insight into its working mechanism.

[0312] Type II (e.g., Cas9) and type V (e.g., Casl2a) CRISPR-Cas systems possess a characteristic Ruv-C like nuclease domain, which has been shown to be related to IS605 family transposon encoded TnpB proteins. Crystallographic and cryo-EM data reveal that Cast 2a adopts a bilobed structure formed by the REC and Nuc lobes. The REC lobe is comprised of REC 1 and REC2 domains, and the Nuc lobe is comprised of tire RuvC, the PAM-interacting (PI) and the WED domains, and additionally, the bridge helix (BH). The RuvC endonuclease domain of this effector protein is made up of three discontinuous parts (RuvC I -III) . The RNase site for processing its own crRNA is situated in the WED-III subdomain, and the DNase site is located in the interface between the RuvC and the Nuc domains. These structural studies have also shown that the only the 5 ’ repeat region of the crRNA is involved in the assembly of the binary complex. The 19 / 20 nt repeat region forms a pseudoknot structure through intramolecular base pairing. Tire crRNA is stabilized through interactions with the WED, RuvC and REC2 domains of the endonuclease, as well as two hydrated Mg2+ ions. This binary interference complex is then responsible for recognizing and degrading foreign DNA.

[0313] PAM recognition is a critical initial step in identifying a prospective DNA molecule for degradation since the PAM allows the CRISPR-Cas systems to distinguish their own genomic DNA from invading nucleic acids. Casl2a employs a multistep quality control mechanism to ensure the accurate and precise recognition of target spacer sequences. The WED II-III. RECI and PAM-interacting domains are responsible for PAM recognition and for initiating the hybridization of the DNA target with the crRNA. After recognition of the dsDNA by WED and REC 1 domains, the conserved loop-lysine helix-loop (LKL) region in the PI domain, containing three conserved lysines (K667, K671, K677 in FnCasl2a), inserts the helix into the PAM duplex with assistance from two conserved prolines in the LKL region. Structural studies show the helix is inserted at an angle of 45° with respect to the dsDNA longitudinal axis, promoting the unwinding of the helical dsDNA. The critical positioning of the three conserved lysines on the dsDNA initiates the uncoupling of the Watson-Crick interaction between the base pairs of the dsDNA after the PAM. The target dsDNA unzipping allows the hybridization of the crRNA w ith the strand containing the PAM, the ‘target strand (TS), w hile the uncoupled DNA strand, non-target strand (NTS), is conducted towards the DNase site by the PAM-interacting domain. Casl2a has been shown to efficiently target spacer sequences following 5'T-rich PAM sequence. The PAM for LbCasl2a and AsCasl2a has a sequence of 5'-TTTN-3' and forFnCasl2a a sequence of 5'-TTN-3' and is situated upstream of the 5’end of the non-target strand. It has also been shown that in addition to the canonical 5'-TTTN-3' PAM, Casl2a also exhibits relaxed PAM recognition for suboptimal C-containing PAM sequences by forming altered interactions with the targeted DNA duplex.

[0314] Thus, Casl2a is another class II CRISPR / Cas system RNA-guided nuclease with similarities to Cas9 and may be used analogously. Unlike Cas9. Casl2a does not require a tracrRNA and only depends on a crRNA in its guide RNA, which provides the advantage that shorter guide RNAs can be used with Cast 2a for targeting than Cas9. Casl2a is capable of cleaving either DNA or RNA. The PAM sites recognized by Casl2a have the sequences 5'-YTN-3' (where “Y’’ is a pyrimidine andC'N” is any nucleobase) or 5'-TTN-3', in contrast to the G-rich PAM site recognized by Cas9. Casl2a cleavage of DNA produces double-stranded breaks with a sticky-ends having a 4 or 5 nucleotide overhang. For further discussion of Casl2a, see, e.g., Ledford et al. (2015) Nature. 526 (7571): 17-17. Zetsche et al. (2015) Cell. 163 (3):759-771, Murovec et al. (2017) Plant Biotechnol. J. 15(8):917-926, Zhang et al. (2017) Front. Plant Sci. 8: 177, Fernandes et al. (2016) Postepy Biochem. 62(3):315-326; herein incorporated by reference.

[0315] Any Casl2a (or Cas Type V) polypeptide or variant thereof may be used in the present disclosure, including those described in the herein tables and provided in the accompanying sequence listing.

[0316] In various embodiments, the Cas 12a (or Cas Type V) polypeptide is a polypeptide selected from Table S15A (SEQ ID NO: 334 (No. ID405), SEQ ID NO: 58 (No. ID414), or SEQ ID NO: 564 (No. ID418), SEQ ID NO: 335 (No. ID406), SEQ ID NO: 331 (No. ID411), SEQ ID NO: 20 (No. ID415), and SEQ ID NO: 445 (No. ID419)), or a polypeptide having at least 70%, 75%, 80%, 85%, 90%, 95%, 99%, or 100% sequence identity with a polypeptide from Table S15A (SEQ ID NO: 334 (No. ID405), SEQ ID NO: 58 (No. ID414). or SEQ ID NO: 564 (No. ID418), SEQ ID NO: 335 (No. ID406), SEQ ID NO: 331 (No. ID411), SEQ ID NO: 20 (No. ID415). and SEQ ID NO: 445 (No. ID419)).

[0317] In various embodiments, the Cas 12a (or Cas Type V) polypeptide is encoded by a polynucleotide sequence selected from Table S15B (SEQ ID NO: 365 (No. ID405), SEQ ID NO: 75 (No. ID414), or SEQ ID NO: 565 (No. ID418), SEQ ID NO: 366 (No. ID406), SEQ ID NO: 331 (No. ID411), SEQ ID NO: 30 (No. ID415), or SEQ ID NO: 445 (No. ID419)), or a polynucleotide having at least 70%, 75%, 80%, 85%, 90%, 95%, 99%, or 100% sequence identity with a polypeptide from Table S15B (SEQ ID NO: 365 (No. ID405), SEQ ID NO: 75 (No. ID414), or SEQ ID NO:565 (No. ID418), SEQ ID NO: 366 (No. ID406), SEQ ID NO: 331 (No. ID411), SEQ ID NO: 30 (No. ID415). or SEQ ID NO: 445 (No. ID419)).

[0318] Any Cas 12a (or Cas Type V) polypeptide may be utilized with the compositions described herein. Hie Cas 12a editing systems contemplated herein are not meant to be limiting in any way. The Cas 12a editingsystems disclosed herein may comprise a canonical or naturally-occurring Casl2a, or any ortholog Casl2a protein, or any variant Cast 2a protein — including any naturally occurring variant, mutant, or otherwise engineered version of Cast 2a — that is known or which can be made or evolved through a directed evolutionary or otherwise mutagenic process. In various embodiments, the Casl2a or Casl2a variants can have a nickase activity, i.e.. only cleave of strand of the target DNA sequence. In other embodiments, the Casl2a or Cast 2a variants have inactive nucleases, i.e.. are “dead” Cast 2a proteins. Other variant Cast 2a proteins that may be used are those having a smaller molecular weight than the canonical Casl2a (e.g., for easier delivery) or having modified amino acid sequences or substitutions.

[0319] In various aspects, the present invention provides one or more modifications of Casl2a (or Cas Type V) polypeptides, including, for example, mutations to increase sufficiency and / or efficiency and modification of the Casl2a. In some embodiments, one or more domains of the Casl2a are modified, e.g., RuvC, REC, WED, BH, PI and NUC domains. In certain preferred embodiments, tire modifications provide editing efficiency of greater than 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%. 90%. 95%, 99% relative to SpCas9. Even more preferably, the methods and compositions provide enhanced transduction efficiency and / or low cytotoxicity.

[0320] The Cas 12a (or Cas Type V) gene editing systems and therapeutics described herein may comprise one or more nucleic acid components (e.g., a guide RNA or a coding RNA that encodes a component of the Casl2a system) which may be codon optimized.

[0321] For example, a nucleotide sequence (e.g., as part of an RNA payload) encoding a nucleobase editing system of the disclosure is codon optimized. Codon optimization methods are known in the art. For example, a protein encoding sequence of any one or more of the sequences provided herein may be codon optimized. Codon optimization, in some embodiments, may be used to match codon frequencies in target and host organisms to ensure proper folding: bias GC content to increase mRNA stability or reduce secondary structures; minimize tandem repeat codons or base runs that may impair gene construction or expression; customize transcriptional and translational control regions; insert or remove protein trafficking sequences; remove / add post translation modification sites in encoded protein (e.g., glycosylation sites); add, remove or shuffle protein domains; insert or delete restriction sites; modify ribosome binding sites and mRNA degradation sites; adjust translational rates to allow the various domains of the protein to fold properly; or reduce or eliminate problem secondary structures within the polynucleotide. Codon optimization tools, algorithms and services are known in the art — non-limiting examples include services from GeneArt (Life Technologies), DNA2.0 (Menlo Park Calif.) and / or proprietary methods. In some embodiments, the protein encoding sequence is optimized using optimization algorithms. In some embodiments, a codon optimizedsequence shares less than 95% sequence identity to a naturally-occurring or wild-type sequence (e.g., a naturally-occurring or wild-type mRNA sequence encoding anucleobase editing enzyme). In some embodiments, a codon optimized sequence shares less than 90% sequence identity to a naturally-occurring or wild-type sequence (e.g., a naturally-occurring or wild-type mRNA sequence encoding a nucleobase editing enzyme). In some embodiments, a codon optimized sequence shares less than 85% sequence identity to a naturally-occurring or wild-type sequence (e.g.. a naturally-occurring or wild-type mRNA sequence encoding a nucleobase editing enzyme). In some embodiments, a codon optimized sequence shares less than 80% sequence identity to a naturally-occurring or wild-type sequence (e.g., a naturally-occurring or wild-type mRNA sequence encoding a nucleobase editing enzyme). In some embodiments, a codon optimized sequence shares less than 75% sequence identity to a naturally-occurring or wild-type sequence (e.g., a naturally- occurring or wild-type mRNA sequence encoding a nucleobase editing enzyme).

[0322] In some embodiments, a codon optimized sequence shares between 65% and 85% (e.g., between about 67% and about 85% or between about 67% and about 80%) sequence identity to a naturally-occurring or wild-type sequence (e.g.. a naturally-occurring or wild-type mRNA sequence encoding a nucleobase editing enzyme). In some embodiments, a codon optimized sequence shares between 65% and 75% or about 80% sequence identity to a naturally-occurring or wild-type sequence (e.g., a naturally-occurring or wild-type mRNA sequence encoding a nucleobase editing enzyme).

[0323] When transfected into mammalian cells, the modified mRNA payloads have a stability of between 12-18 hours, or greater than 18 hours, e.g., 24, 36, 48, 60, 72, or greater than 72 hours.

[0324] In some embodiments, a codon optimized RNA may be one in which the levels of G / C are enhanced. The G / C-content of nucleic acid molecules (e.g., mRNA) may influence the stability of the RNA. RNA having an increased amount of guanine (G) and / or cytosine (C) residues may be functionally more stable than RNA containing a large amount of adenine (A) and thymine (T) or uracil (U) nucleotides. As an example, WO02 / 098443 discloses a pharmacal composition containing an mRNA stabilized by sequence modifications in the translated region. Due to the degeneracy of the genetic code, the modifications work by substituting existing codons forthose that promote greater RNA stability without changing the resulting amino acid. The approach is limited to coding regions of the RNA.

[0325] In some embodiments, the disclosure provides engineered Casl2a variants or mutants which have been modified by introducing one or more amino acid substitutions into a baseline sequence (e.g., a wildtype sequence).

[0326] Any available methods may be utilized to obtain or construct a variant or mutant Cas 12a protein. The term ‘'mutation,” as used herein, refers to a substitution of a residue within a sequence, e.g., a nucleic acid or amino acid sequence, with another residue, or a deletion or insertion of one or more residues w ithin a sequence. Mutations are typically described herein by identifying tire original residue followed by the position of the residue within the sequence and by the identity of the newly substituted residue. Various methods for making the amino acid substitutions (mutations) provided herein are well known in the art (e.g., site-directed mutagenesis or directed evolution engineering), and are provided by, for example, Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y. (2012)). Mutations can include a variety of categories, such as single base polymorphisms, microduplication regions, indel, and inversions, and is not meant to be limiting in any way. Mutations can include “loss-of-function” mutations w hich is the normal result of a mutation that reduces or abolishes a protein activity. Most loss-of-function mutations are recessive, because in a heterozygote the second chromosome copy carries an unmutated version of the gene coding for a fully functional protein w hose presence compensates for the effect of the mutation. Mutations also embrace “gain-of-function” mutations, which confer an abnormal activity on a protein or cell that is otherwise not present in a normal condition. Many gain-of-function mutations are in regulatory sequences rather than in coding regions, and can therefore have a number of consequences. For example, a mutation might lead to one or more genes being expressed in the wrong tissues, these tissues gaining functions that they normally lack. Because of their nature, gain-of-function mutations are usually dominant.

[0327] Mutations can be introduced into a reference Cas 12a protein using site-directed mutagenesis. Older methods of site-directed mutagenesis know n in the art rely on sub-cloning of the sequence to be mutated into a vector, such as an M13 bacteriophage vector, that allows the isolation of single-stranded DNA template. In these methods, one anneals a mutagenic primer (i.e., a primer capable of annealing to the site to be mutated but bearing one or more mismatched nucleotides at the site to be mutated) to the single-stranded template and then polymerizes the complement of the template starting from the 3' end of the mutagenic primer. The resulting duplexes are then transformed into host bacteria and plaques are screened for the desired mutation. More recently, site-directed mutagenesis has employed PCR methodologies, w hich have the advantage of not requiring a single-stranded template. In addition, methods have been developed that do not require sub-cloning. Several issues must be considered when PCR-based site-directed mutagenesis is performed. First, in these methods it is desirable to reduce the number of PCR cycles to prevent expansion of undesired mutations introduced by the polymerase. Second, a selection must be employed in order to reduce the number of non-mutated parental moleculespersisting in the reaction. Third, an extended-length PCR method is preferred in order to allow the use of a single PCR primer set. And fourth, because of the non-template-dependent terminal extension activity of some thermostable polymerases it is often necessary’ to incorporate an end-polishing step into the procedure prior to blunt-end ligation of the PCR-generated mutant product.

[0328] Mutations may also be introduced by directed evolution processes, such as phage-assisted continuous evolution (PACE) or phage-assisted noncontinuous evolution (PANCE). Tire term “phage- assisted continuous evolution (PACE)," as used herein, refers to continuous evolution that employs phage as viral vectors. The general concept of PACE technology has been described, for example, in International PCT Application, PCT / US2009 / 056194, filed Sep. 8, 2009, published as WO 2010 / 028347 on Mar. 11, 2010; International PCT Application, PCT / US2011 / 066747, filed Dec. 22, 2011, published as WO 2012 / 088381 on Jun. 28, 2012; U.S. Pat. No. 9,023,594, issued May 5, 2015, International PCT Application, PCT / US2015 / 012022, filed Jan. 20, 2015, published as WO 2015 / 134121 on Sep. 11, 2015, and International PCT Application, PCT / US2016 / 027795, filed Apr. 15, 2016. published as WO 2016 / 168631 on Oct. 20, 2016. the entire contents of each of which are incorporated herein by reference. Variant Casl2as may also be obtain by phage-assisted non-continuous evolution (PANCE),” which as used herein, refers to non-continuous evolution that employs phage as viral vectors. PANCE is a simplified technique for rapid in vivo directed evolution using serial flask transfers of evolving "selection phage’ (SP), which contain a gene of interest to be evolved, across fresh E. coli host cells, thereby allowing genes inside the host E. coli to be held constant while genes contained in the SP continuously evolve. Serial flask transfers have long served as a widely-accessible approach for laboratory evolution of microbes, and, more recently, analogous approaches have been developed for bacteriophage evolution. The PANCE system features lower stringency than the PACE system.

[0329] The disclosure contemplates any engineered Casl2a variants or mutants which have been modified by introducing one or more amino acid substitutions into a baseline sequence, including conservative substitutions of one amino acid for another. For example, mutation of an amino acid with a hydrophobic side chain (e.g.. alanine, valine, isoleucine, leucine, methionine, phenylalanine, tyrosine, or tryptophan) may be changed to a second amino acid with a different hydrophobic side chain (e.g., alanine, valine, isoleucine, leucine, methionine, phenylalanine, tyrosine, or tryptophan). For example, a mutation of an alanine to a threonine (e.g., a A262T mutation) may also include a mutation from an alanine to an amino acid that is similar in size and chemical properties to a threonine, for example, serine. As another example, mutation of an amino acid with a positively charged side chain (e g., arginine, histidine, or lysine) may include a mutationto a second amino acid with a different positively charged side chain (e.g., arginine, histidine, or lysine). As another example, mutation of an amino acid with a polar side chain (e.g., serine, threonine, asparagine, or glutamine) may also include a mutation to a second amino acid with a different polar side chain (e.g., serine, threonine, asparagine, or glutamine). Additional similar amino acid pairs include, but are not limited to, the following: phenylalanine and tyrosine; asparagine and glutamine: methionine and cysteine; aspartic acid and glutamic acid; and arginine and lysine. The skilled artisan would recognize that such conservative amino acid substitutions may only have minor effects on protein structure and may be well tolerated without compromising function. In some embodiments, any amino acid mutations provided herein from one amino acid to a threonine may be an amino acid mutation to a serine. In some embodiments, any amino acid mutations provided herein from one amino acid to an arginine may be an amino acid mutation to a lysine. In some embodiments, any amino acid mutations provided herein from one amino acid to an isoleucine, may be an amino acid mutation to an alanine, valine, methionine, or leucine. In some embodiments, any amino acid mutations provided herein from one amino acid to a lysine may be an amino acid mutation to an arginine. In some embodiments, any amino acid mutations provided herein from one amino acid to an aspartic acid may be an amino acid mutation to a glutamic acid or asparagine. In some embodiments, any amino acid mutations provided herein from one amino acid to a valine may be an amino acid mutation to an alanine, isoleucine, methionine, or leucine. In some embodiments, any amino acid mutations provided herein from one amino acid to a glycine may be an amino acid mutation to an alanine. It should be appreciated, however, that additional conserved amino acid residues would be recognized by the skilled artisan and any of the amino acid mutations to other conserved amino acid residues are also within the scope of this disclosure. Tire amino acid substitutions may also be non-conservative amino acid substitutions.

[0330] In various embodiments, an Alanine (A) residue of a Casl2a protein may be substituted with any one of tire following amino acids: Arginine (R); Asparagine (N); Aspartic Acid (D); Cysteine (C); Glutamic acid (E) ; Glutamine (N); Glycine (G): Histidine (H); Isoleucine (I); Leucine (L); Lysine (K); Methionine (M); Phenylalanine (F); Proline (P); Serine (S); Threonine (T); Tryptophan (W); Tyrosine (Y): or Valine (V).

[0331] In another embodiment, an Arginine (R) residue of a Casl2a protein may be substituted with any one of the following amino acids: Alanine (A); Asparagine (N); Aspartic Acid (D); Cysteine (C); Glutamic acid (E) ; Glutamine (N); Glycine (G); Histidine (H); Isoleucine (I); Leucine (L); Lysine (K); Methionine (M); Phenylalanine (F); Proline (P); Serine (S): Threonine (T); Try ptophan (W); Tyrosine (Y); or Valine (V).

[0332] In another embodiment, an Asparagine (N) residue of a Casl2a protein may be substituted with any one of the following amino acids: Alanine (A); Arginine (R); Aspartic Acid (D); Cysteine (C); Glutamic acid(E) ; Glutamine (N): Glycine (G); Histidine (H); Isoleucine (I); Leucine (L): Lysine (K); Methionine (M); Phenylalanine (F); Proline (P); Serine (S); Threonine (T); Tryptophan (W); Tyrosine (Y); or Valine (V).

[0333] In another embodiment, an Aspartic Acid (D) residue of a Casl2a protein may be substituted with any one of the following amino acids: Alanine (A); Arginine (R); Asparagine (N); Cysteine (C); Glutamic acid (E) ; Glutamine (N); Glycine (G); Histidine (H); Isoleucine (I); Leucine (L); Lysine (K); Methionine (M); Phenylalanine (F); Proline (P); Serine (S); Threonine (T); Tryptophan (W); Tyrosine (Y); or Valine (V).

[0334] In another embodiment, an Cysteine (C) residue of a Casl2a protein may be substituted with any one of the following amino acids: Alanine (A); Arginine (R); Asparagine (N); Aspartic Acid (D); Glutamic acid (E) ; Glutamine (N); Glycine (G): Histidine (H); Isoleucine (I); Leucine (L); Lysine (K); Methionine (M); Phenylalanine (F); Proline (P); Serine (S); Threonine (T): Tryptophan (W): Tyrosine (Y): or Valine (V).

[0335] In another embodiment, an Glutamic acid (E) residue of a Casl2a protein may be substituted with any one of the following amino acids: Alanine (A): Arginine (R): Asparagine (N); Aspartic Acid (D);Cysteine (C); Glutamine (N); Glycine (G); Histidine (H); Isoleucine (I): Leucine (L); Lysine (K); Methionine (M); Phenylalanine (F); Proline (P); Serine (S); Threonine (T); Tryptophan (W); Tyrosine (Y); or Valine (V).

[0336] In another embodiment, an Glutamine (N) residue of a Cast 2a protein may be substituted with any one of the following amino acids: Alanine (A); Arginine (R); Asparagine (N); Aspartic Acid (D); Cysteine (C); Glutamic acid (E); Glycine (G); Histidine (H); Isolcucinc (I); Leucine (L); Lysine (K); Methionine (M); Phenylalanine (F); Proline (P); Serine (S); Threonine (T); Tryptophan (W); Tyrosine (Y): or Valine (V).

[0337] In another embodiment, an Glycine (G) residue of a Casl2a protein may be substituted with any one of the following amino acids: Alanine (A); Arginine (R); Asparagine (N); Aspartic Acid (D); Cysteine (C); Glutamic acid (E); Glutamine (N); Histidine (H); Isoleucine (I); Leucine (L): Lysine (K): Methionine (M); Phenylalanine (F); Proline (P); Serine (S): Threonine (T); Tryptophan (W); Tyrosine (Y); or Valine (V).

[0338] In another embodiment, an Histidine (H) residue of a Cas 12a protein may be substituted with any one of the following amino acids: Alanine (A); Arginine (R); Asparagine (N); Aspartic Acid (D); Cysteine (C); Glutamic acid (E); Glutamine (N); Glycine (G); Isoleucine (I); Leucine (L); Lysine (K); Methionine (M); Phenylalanine (F); Proline (P); Serine (S); Threonine (T); Tryptophan (W); Tyrosine (Y); or Valine (V).

[0339] In another embodiment, an Isoleucine (I) residue of a Cas 12a protein may be substituted with any one of tire following amino acids: Alanine (A); Arginine (R); Asparagine (N); Aspartic Acid (D); Cysteine (C); Glutamic acid (E); Glutamine (N); Glycine (G); Histidine (H); Leucine (L); Lysine (K); Methionine (M); Phenylalanine (F); Proline (P); Serine (S); Threonine (T): Tryptophan (W); Tyrosine (Y): or Valine (V).

[0340] In another embodiment, an Leucine (L) residue of a Cast 2a protein may be substituted with any one of tire following amino acids: Alanine (A); Arginine (R); Asparagine (N); Aspartic Acid (D); Cysteine (C); Glutamic acid (E); Glutamine (N); Glycine (G); Histidine (H); Isoleucine (I); Lysine (K); Methionine (M); Phenylalanine (F); Proline (P); Serine (S); Threonine (T); Tryptophan (W); Tyrosine (Y); or Valine (V).

[0341] In another embodiment, an Lysine (K) residue of a Cas 12a protein may be substituted with any one of tire following amino acids: Alanine (A); Arginine (R); Asparagine (N); Aspartic Acid (D); Cysteine (C); Glutamic acid (E); Glutamine (N); Glycine (G); Histidine (H); Isoleucine (I); Leucine (L); Methionine (M); Phenylalanine (F); Proline (P); Serine (S); Threonine CI ): Trvptophan (W): Tyrosine (Y): or Valine (V).

[0342] In another embodiment, an Methionine (M) residue of a Cas 12a protein may be substituted with any one of the following amino acids: Alanine (A); Arginine (R): Asparagine (N); Aspartic Acid (D): Cysteine (C); Glutamic acid (E); Glutamine (N); Glycine (G); Histidine (H); Isoleucine (I): Leucine (L); Lysine (K); Phenylalanine (F); Proline (P); Serine (S); Threonine (T); Tryptophan (W); Tyrosine (Y); or Valine (V).

[0343] In another embodiment, an Phenylalanine (F) residue of a Cas 12a protein may be substituted with any one of the following amino acids: Alanine (A); Arginine (R); Asparagine (N); Aspartic Acid (D); Cysteine (C); Glutamic acid (E); Glutamine (N); Glycine (G); Histidine (H); Isoleucine (I); Leucine (L); Lysine (K); Methionine (M); Proline (P); Serine (S); Threonine (T); Tryptophan (W); Tyrosine (Y); or Valine (V).

[0344] In another embodiment, an Proline (P) residue of a Cas 12a protein may be substituted with any one of the following amino acids: Alanine (A); Arginine (R); Asparagine (N); Aspartic Acid ( D): Cysteine (C): Glutamic acid (E); Glutamine (N); Glycine (G); Histidine (H); Isoleucine (I); Leucine (L); Lysine (K); Methionine (M); Phenylalanine (F); Serine (S); Threonine (T); Tryptophan (W); Tyrosine (Y); or Valine (V).

[0345] In another embodiment, an Serine (S) residue of a Cas 12a protein may be substituted with any one of the following amino acids: Alanine (A); Arginine (R); Asparagine (N); Aspartic Acid (D); Cysteine (C); Glutamic acid (E); Glutamine (N); Glycine (G); Histidine (H); Isoleucine (I); Leucine (L); Lysine (K); Methionine (M); Phenylalanine (F); Proline (P); Threonine (T); Tryptophan (W); Tyrosine (Y): or Valine (V).

[0346] In another embodiment, an Threonine (T) residue of a Cas 12a protein may be substituted with any one of the following amino acids: Alanine (A); Arginine (R); Asparagine (N); Aspartic Acid (D): Cysteine (C): Glutamic acid (E); Glutamine (N); Glycine (G); Histidine (H ): Isoleucine (I); Leucine (L): Lysine (K): Methionine (M); Phenylalanine (F); Proline (P): Serine (S); Try ptophan (W); Tyrosine (Y); or Valine (V).

[0347] In another embodiment, an Tryptophan (W) residue of a Cas 12a protein may be substituted with any one of the following amino acids: Alanine (A); Arginine (R); Asparagine (N); Aspartic Acid (D); Cysteine (C); Glutamic acid (E); Glutamine (N); Glycine (G); Histidine (H); Isoleucine (I); Leucine (L); Lysine (K); Methionine (M); Phenylalanine (F); Proline (P); Serine (S); Threonine (T); Tyrosine (Y); or Valine (V).

[0348] In another embodiment, an Tyrosine (Y) residue of a Cas 12a protein may be substituted with any one of tire following amino acids: Alanine (A); Arginine (R); Asparagine (N); Aspartic Acid (D); Cysteine (C); Glutamic acid (E); Glutamine (N); Glycine (G); Histidine (H); Isoleucine (I); Leucine (L); Lysine (K); Methionine (M); Phenylalanine (F); Proline (P); Serine (S): Threonine (T); Tryptophan (W); or Valine (V).

[0349] In another embodiment, an Valine (V) residue of a Cas 12a protein may be substituted with any one of the following amino acids: Alanine (A); Arginine (R); Asparagine (N); Aspartic Acid (D); Cysteine (C); Glutamic acid (E); Glutamine (N); Glycine (G): Histidine (H); Isoleucine (I); Leucine (L); Lysine (K); Methionine (M); Phenylalanine (F); Proline (P); Serine (S); Threonine (T); or Tryptophan (W).

[0350] In addition, the amino acid subsitutions may include that of any non-naturally occurring amino acid analog or amino acid derivative that are known in the art.

[0351] While not intending to be limiting, the following are exemplary' embodiments of mutant variants contemplated by the instant specification and Examples and which are based on Casl2a ID405 (SEQ ID NO: 334), Casl2a ID414 (SEQ ID NO: 58), and Casl2a ID418 (SEQ ID NO: 564). It will be appreciated that any of the following specific substitutions and / or combinations of specific substitutions may be introduced into the corresponding amino acid residues (as determined by a sequence alignment) of any other Type V nuclease enzyme disclosed herein.Variants based on ID405 (SEO ID NO: 334)

[0352] In various embodiments, the Casl2a may be a Casl2a variant based on ID405 (SEQ ID NO: 334), and may include any of the following substitutions and in any combination (or any amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 99%, or up to 100% sequence identity with SEQ ID NO: 334 having any of the following substitutions):• a D 169 substitution;• a C554 substitution;• a N559 substitution;• a Q565 substitution;• a L860 substitution;• a R950 substitution; and / or• a R954 substitution.

[0353] In various embodiments, the Casl2a may be a Casl2a variant based on ID405 (SEQ ID NO: 334), and may include any of the following substitutions and in any combination (or any amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 99%, or up to 100% sequence identity with SEQ ID NO: 334 having any of the following substitutions):• a D 169R substitution;• a C554N substitution;• a C554R substitution;• a N559R substitution:• a Q565R substitution;• a L860Q substitution;• a R950K substitution; and / or• a R954A substitution.

[0354] In various embodiments, the Casl2a may be a Casl2a variant based on ID405 (SEQ ID NO: 334), and may include any of the following substitutions and in any combination (or any amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 99%, or up to 100% sequence identity with SEQ ID NO: 334 having any of the following substitutions):• a D 169 substitution;• a D 169 / R950 / R954 substitution set;• a D169 / N559 / Q565 substitution set;• a C554 substitution;• a C554 substitution; and / or• a L860 substitution.

[0355] In various embodiments, the Casl2a may be a Casl2a variant based on ID405 (SEQ ID NO: 334), and may include any of the following substitutions and in any combination (or any amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 99%, or up to 100% sequence identity with SEQ ID NO: 334 having any of the following substitutions):• a D 169R substitution;• a D169R / R950K / R954A substitution set;• a D169R / N559R / Q565R substitution set;• a C554R substitution;• a C554N substitution; and / or• a L860Q substitution.

[0356] The full amino acid and protein coding sequences of these mutant nucleases are provided in Appendix A at Section P (Casl2a Mutant Type V nuclease and associated sequences).Variants based on ID414 (SEQ ID NO: 58)

[0357] In various embodiments, the Casl2a may be a Casl2a variant based on ID414 (SEQ ID NO: 58), and may include any of the following substitutions and in any combination (or any amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 99%, or up to 100% sequence identity with SEQ ID NO: 58 having any of the following substitutions):• a T154 substitution;• a N531 substitution;• a G546 substitution;• a K542 substitution;• a S802 substitution;• a R887 substitution; and / or• a R891 substitution.

[0358] In various embodiments, the Casl2a may be a Casl2a variant based on ID414 (SEQ ID NO: 58), and may include any of the following substitutions and in any combination (or any amino acid sequence having at least 70%, 75%, 80%, 85%. 90%. 95%. 99%. or up to 100% sequence identity with SEQ ID NO: 58 having any of the following substitutions):• a T154R substitution;• a N531R substitution;• a G546R substitution;• a K542R substitution;• a S802L substitution;• a R887K substitution; and / or• a R891A substitution.

[0359] In various embodiments, the Casl2a may be a Casl2a variant based on ID414 (SEQ ID NO: 58), and may include any of the following substitutions and in any combination (or any amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 99%, or up to 100% sequence identity with SEQ ID NO: 58 having any of the following substitutions):• a T154 substitution;• a T154 / R887 / R891 substitutions;• a T154 / G536 / K542 substitutions;• a N531 / S802 substitutions;• a N531 substitution; and / or• a S802 substitution.

[0360] In various embodiments, the Casl2a may be a Casl2a variant based on ID414 (SEQ ID NO: 58), and may include any of the following substitutions and in any combination (or any amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 99%, or up to 100% sequence identity with SEQ ID NO: 58 having any of the following substitutions):• a T 154R substitution;• a T154R / R887K / R891A substitutions;• a T154R / G536R / K542R substitutions;• a N531R / S802L substitutions;• a N531R substitution; and / or• a S802L substitution.

[0361] Hie full amino acid and protein coding sequences of these mutant nucleases are provided in Appendix A at Section P (Casl2a Mutant Type V nuclease and associated sequences).Variants based on ID418 (SEQ ID NO: 564)

[0362] In various embodiments, the Casl2a may be a Casl2a variant based on ID418 (SEQ ID NO: 564), and may include any of the following substitutions and in any combination (or any amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 99%, or up to 100% sequence identity with SEQ ID NO: 564 having any of the following substitutions):• a D 161 substitution;• a N527 substitution;• a T532 substitution;• a K538 substitution;• a Q799 substitution;• a R888 substitution; and / or• a R892 substitution.

[0363] In various embodiments, the Casl2a may be a Casl2a variant based on ID418 (SEQ ID NO: 564), and may include any of the following substitutions and in any combination (or any amino acid sequence having at least 70%, 75%, 80%, 85%. 90%. 95%. 99%, or up to 100% sequence identity with SEQ ID NO: 564 having any of tire following substitutions):• a D161R substitution;• a N527R substitution;• a T532R substitution;• a K538R substitution;• a Q799L substitution;• a R888K substitution; and / or• a R892A substitution.

[0364] In various embodiments, the Casl2a may be a Casl2a variant based on ID418 (SEQ ID NO: 564), and may include any of the following substitutions and in any combination (or any amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 99%, or up to 100% sequence identity with SEQ ID NO: 564 having any of the following substitutions): a D161 substitution;• a D 161 / R888 / R892 substitution;• a D161 / T532 / K538 substitution;• a N527 / Q799 substitution;• a N527 substitution; and / or• a Q799 substitution.

[0365] In various embodiments, the Casl2a may be a Cast 2a variant based on ID418 (SEQ ID NO: 564), and may include any of the following substitutions and in any combination (or any amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 99%, or up to 100% sequence identity with SEQ ID NO: 564 having any of the following substitutions):• a D 161 R substitution;• a D161R / R888K / R892A substitution;• a D161R / T532R / K538R substitution;• a N527R / Q799L substitution;• a N527R substitution; and / or• a Q799L substitution.

[0366] The full amino acid and protein coding sequences of these mutant nucleases are provided in Appendix A at Section P (Casl2a Mutant Type V nuclease and associated sequences).

[0367] In addition, various embodiments of variant Casl2a orthologs are described in Section Q of Appendix A.

[0368] In addition, embodiments of Cast 2a mutant variants based on ID405, ID414, and ID418 are described in the computational approach to directed mutagenesis described in Example 14.

[0369] It is noted that when this disclosure speaks to a polypeptide (including anywhere in this specification, including in tire Appendix A and the Examples) having a percent identity with respect to another amino acid sequence (a reference amino acid sequence), such as a polypeptide at least 50%, at least 55%, at least 60%, at least 65%, at least 70%. at least 75%, at least 80%, at least 85%, at least 90%, at least 91%. at least 92%, at least 93%, at least 94%. at least 95. at least 96%, at least 97%, at least 98%, at least 99%%. at least 99.1%. at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8% or at least 99.9% identical to another amino acid sequence (a reference amino acid sequence), such as one of SEQ ID NO: 334 (No. ID405), SEQ ID NO: 58 (No. ID414), or SEQ ID NO: 564 (No. ID418), SEQ ID NO: 335 (No. ID406), SEQ ID NO: 331 (No. ID411), SEQ ID NO: 20 (No. ID415), or SEQ ID NO: 445 (No. ID419), it is advantageous that in the polypeptide having a percent identity to the reference amino acid sequence conserved regions of the reference amino acid sequence (e.g., conserved when compared with other Casl2as, such as those identified herein, such as described in the multi-sequences alignment of FIG. 31) be preserved and / or that the polypeptide has at least one activity selected from endonuclease activity; endoribonucleaseactivity, or RNA-guided DNase activity and / or that the polypeptide of which comprises: a. one or more a- helical recognition lobe (REC) and a nuclease lobe (NUC); b. a Wedge (WED), a-helical recognition lobe (REC), PAM-interacting (PI), RuvC nuclease, Bridge Helix (BH) and NUC domains; or c. one or more domains selected from RuvC, REC. WED, BH, PI and NUC domains and / or that the polypeptide recognizes or binds crRNA(s) or is bound to crRNA(s), such as a crRNA sequence from Table S15C. Likewise, when this disclosure speaks to a nucleic acid sequence or molecule having a percent identity with respect to a nucleic acid sequence having a percent identity with respect to another nucleic acid sequence or molecule (a reference nucleic acid sequence), such as a nucleic acid sequence at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%. at least 95%, at least 96%, at least 97%, at least 98%, at least 99%%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%. at least 99.8% or at least 99.9% identical to another nucleic acid sequence (a reference nucleic acid sequence, such as a sequence selected from SEQ ID NO: 365 (No. ID405), SEQ ID NO: 74 (No. ID414), or SEQ ID NO: 565 (No. ID418), SEQ ID NO: 366 (No. ID406), SEQ ID NO: 331 (No. ID411), SEQ ID NO: 30 (No. ID415), or SEQ ID NO: 445 (No. ID419), it is advantageous that in the nucleic acid sequence that has a percent identity to the reference nucleic acid sequence that conserved regions of the reference nucleic acid sequence (e.g., conserved when compared with other Casl2as, such as those identified herein) be preserved and / or that in the polypeptide that is expressed from the nucleic acid sequence that has a percent identity to the reference nucleic acid sequence that the polypeptide contain conserved region(s) (e.g., conserved when compared with other Casl2as, such as those identified herein) and / or that the polypeptide has at least one activity selected from endonuclease activity; endoribonuclease activity, or RNA-guided DNase activity and / or that the polypeptide of which comprises: a. one or more a-helical recognition lobe (REC) and a nuclease lobe (NUC); b. a Wedge (WED), a-helical recognition lobe (REC), PAM-interacting (PI), RuvC nuclease. Bridge Helix (BH) and NUC domains; or c. one or more domains selected from RuvC. REC. WED, BH. Pl and NUC domains and / or that the polypeptide recognizes or binds crRNA(s) or is bound to crRNA(s), such as a crRNA sequence from Table S15C.D. Casl2a (or Cas Type V) Guide RNA SequencesCasl2a (Cas Type V) guide sequences

[0370] The present disclosure further provides guide RNAs for use in accordance with the disclosed nucleic acid programmable DNA binding proteins (e.g., Casl2a) for use in methods of editing. The disclosure provides guide RNAs that are designed to recognize target sequences. Such gRNAs may be designed to have guide sequences (or “spacers”) having complementarity to a target sequence. Such gRNAs may be designedto have not only a guide sequences having complementarity to a target sequence to be edited, but also to have a backbone sequence that interacts specifically with the nucleic acid programmable DNA binding protein.

[0371] In various aspects, provided are one or more guide RNA sequences. In preferred embodiments, the gRNA is cleaved and processed into one or more intermediate crRNAs, which are subsequently processed into one or more mature crRNAs. In some embodiments, the gRNA comprises a precursor CRISPR RNAs (pre-crRNA) encoding one or more crRNAs or one or more intermediate or mature crRNAs, each guide RNA comprising at a minimum a repeat-spacer in the 5’ to 3‘ direction, wherein the repeat comprises a stem -loop structure and the spacer comprises a DNA-targeting segment complementary to a target sequence in the targeted polynucleotide sequence. In certain embodiments, the gRNA is cleaved by a RNase activity of the Casl2a polypeptide into one or more mature crRNAs, each comprising at least one repeat and at least one spacer.

[0372] In other embodiments, one or more repeat-spacer directs the Casl2a (or Cas Type V) polypeptides to two or more distinct sites in the targeted polynucleotide sequence. Preferably, the gRNA is cleaved and processed into one or more intermediate crRNAs, which are subsequently processed into one or more mature crRNAs. More preferably, the pre- crRNA or intermediate crRNA are processed into mature crRNA by an Cas 12a (or Cas Type V) polypeptide, and the mature crRNA becomes available for directing the Cas 12a (or Cas Type V) endonuclease activity. In alternative embodiments, the gRNA is linked to a single or double strand DNA donor template, and the donor template is cleaved from the gRNA by the Cas 12a (or Cas Type V) polypeptide. The donor polynucleotide template remains linked to gRNA while the Cas 12a (or Cas Type V) polypeptide cleaves gRNA to liberate intermediate or mature crRNAs.

[0373] In exemplary embodiments, the Cas 12a (or Cas Type V) system comprises one or more guide RNA comprising:(a) one or more crRNA direct repeat sequences or a reverse complement selected from (Group 1) SEQ ID NO:7-12; (Group 2) SEQ ID NO:24-27: (Group 3) SEQ ID NO:36-39; (Group 4) SEQ ID NO:49-52; (Group 5) SEQ ID NO:63-68; (Group 6) SEQ ID NO:84-91; (Group 7) SEQ ID NO: 106- 111: (Group 8) SEQ ID NO: 122-125; (Group 9) SEQ ID Nos:211-290; (Group 10) SEQ ID NO:343-354; (Group 11) SEQ ID NO:374-379; (Group 12) SEQ ID NO:390-393; (Group 13) SEQ ID NO:411-422; and (Group 14) SEQ ID N0:500-541;(b) 20 to 35 nucleotides or up to the length of the crRNA from the 3' end of the crRNA direct repeat sequences or a reverse complement (a) linked to a targeting guide attached to the 3 ’ end of the direct repeat sequence that is of 16-30 nucleotides in length;(c) (Group 1) SEQ ID NO:13-15; (Group 2) SEQ ID NO:28-29; (Group 3) SEQ ID NO:40-41; (Group 4) SEQ ID NO:53-54; (Group 5) SEQ ID NO:69-71; (Group 6) SEQ ID NO:92-95; (Group 7) SEQ ID NO: 112-114; (Group 8) SEQ ID NO: 126-127; (Group 9) SEQ ID NO:291-330; (Group 10) SEQ ID NO:355-360; (Group 11) SEQ ID NO:380-382; (Group 12) SEQ ID NO:394-395; (Group 13) SEQ ID NO:423-428; and (Group 14) SEQ ID NO:542-563;(d) a nucleic acid sequence that is a degenerate variant of (Group 1) SEQ ID NO: 13-15; (Group 2) SEQ ID NO:28-29; (Group 3) SEQ ID NO:40-41; (Group 4) SEQ ID NO:53-54; (Group 5) SEQ ID NO:69-71; (Group 6) SEQ ID NO:92-95; (Group 7) SEQ ID NO: 112-114; (Group 8) SEQ ID NO: 126-127; (Group 9) SEQ ID NO:291-330; (Group 10) SEQ ID NO:355-360; (Group 11) SEQ ID NO:380-382; (Group 12) SEQ ID NO:394-395; (Group 13) SEQ ID NO:423-428; and (Group 14) SEQ ID NO:542-563;(e) anucleic acid sequence at least 70%, at least 71%, at least 72%. at least 73%. at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81 %, at least 82%, at least 83%, at least 84%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99% or at least 99.9% identical to : (Group 1) SEQ ID NO: 13-15; (Group 2) SEQ ID NO:28-29; (Group 3) SEQ ID NO:40-41; (Group 4) SEQ ID NO:53-54; (Group 5) SEQ ID NO:69-71; (Group 6) SEQ ID NO:92-95; (Group 7) SEQ ID NO: 112-114; (Group 8) SEQ ID NO: 126-127; (Group 9) SEQ ID NO:291- 330; (Group 10) SEQ ID NO:355-360; (Group 11) SEQ ID NO:380-382; (Group 12) SEQ ID NO:394- 395; (Group 13) SEQ ID NO:423-428; and (Group 14) SEQ ID NO:542-563; and(f) a nucleic acid sequence that hybridizes under stringent conditions to : (Group 1) SEQ ID NO: 13-15; (Group 2) SEQ ID NO:28-29; (Group 3) SEQ ID NO:40-41; (Group 4) SEQ ID NO:53-54; (Group 5) SEQ ID NO:69-71; (Group 6) SEQ ID NO:92-95; (Group 7) SEQ ID NO: 112-114; (Group 8) SEQ ID NO: 126-127; (Group 9) SEQ ID NO:291-330; (Group 10) SEQ ID NO:355-360; (Group 11) SEQ ID NO:380-382; (Group 12) SEQ ID NO:394-395; (Group 13) SEQ ID NO:423-428; and (Group 14) SEQ ID NO:542-563.

[0374] In preferred embodiments, the Casl2a (or Cas Type V) proteins target and cleave targeted polynucleotides that is complementary to a cognate guide RNA. In certain embodiments, the guide RNA comprises crRNA, which includes the natural CRISPR array. Such variants are derived from the first direct repeat, a “leader” sequence and involved in signaling or the direct repeat retains genetic diversity that doesn’t affect functionality. The direct repeat is degenerate, generally near the 3 ’ end of the repeat array.

[0375] In various embodiments, the crRNA comprises about 15-40 nucleotides or direct repeat sequences comprising about 20-30 nucleotides. In exemplary embodiments, the direct repeat is selected from (Group 1) SEQ ID NO:7-12; (Group 2) SEQ ID NO:24-27; (Group 3) SEQ ID NO:36-39; (Group 4) SEQ ID NO:49- 52; (Group 5) SEQ ID NO:63-68; (Group 6) SEQ ID NO:84-91; (Group 7) SEQ ID NO: 106-111; (Group 8)SEQ ID NO: 122-125; (Group 9) SEQ ID Nos:211-290; (Group 10) SEQ ID NO:343-354; (Group 11) SEQ ID NO:374-379; (Group 12) SEQ ID NO:390-393; (Group 13) SEQ ID NO:411-422; and (Group 14) SEQ ID N0:500-541. More preferably, the crRNA comprises a guide segment of 16-26 nucleotides or 20-24 nucleotides. Accordingly, in various embodiments, tire crRNA of the Casl2a genome editing systems hybridizes to one or more targeted polynucleotide sequence. In certain preferred embodiments, the crRNA is 43- nucleotides. In other embodiments, the crRNA is made up of a 20-nucleotide 5 '-handle and a 23- nucleotide leader sequence. In certain embodiments, the leader sequence comprises a seed region and 3' termini, both of which are complementary to the target region in the genome Li, Bin et al. “Engineering CRISPR-Cpfl crRNAs and rnRNAs to maximize genome editing efficiency.” Nature biomedical engineering vol. 1,5 (2017): 0066. doi: 10.1038 / s41551-017-0066.

[0376] A single crRNA-guided endonuclease and has the ribonuclease activity to process its pre-crRNA into mature crRNA Zetsche, Bernd et al. “A Survey of Genome Editing Activity for 16 Casl2a Orthologs.” The Keio journal of medicine vol. 69,3 (2020): 59-65. doi: 10.2302 / kjm.2019-0009-OA; Fonfara, Ines et al. “The CRISPR-associated DNA- cleaving enzyme Cpfl also processes precursor CRISPR RNA.” Nature vol. 532,7600 (2016): 517-21. doi: 10. 1038 / nature 17945, which enables multiplex editing in a single crRNA transcript. Campa, Carlo C et al. “Multiplexed genome engineering by Casl2a and CRISPR arrays encoded on single transcripts.” Natu re methods vol. 16,9 (2019): 887-893. doi: 10.1038 / s41592-019-0508-6; Zetsche, Bernd et al. “Multiplex gene editing by CRISPR- Cpfl using a single crRNA array.” Nature biotechnology vol. 35,1 (2017): 31-34. doi: 10.1038 / nbt.3737

[0377] Preferably, the crRNA-guided endonuclease provides alteration of numerous loci in host cell genomes.

[0378] More preferably, the Cast 2a (or Cas Type V) comprises multiplexing performed using two methods. One method involves expressing many single gRNAs under different small RNA promoters either in same vector or in different vectors. Another method, multiple single gRNAs are fused with a tRNA recognition sequence, which are expressed as a single transcript under one promoter.

[0379] In some embodiments, the guide RNA may be 15-100 nucleotides in length and comprise a sequence of at least 10, at least 15, or at least 20 contiguous nucleotides that is complementary to atarget nucleotide sequence. The guide RNA may comprise a spacer sequence of 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40 contiguous nucleotides that is complementary to a target nucleotide sequence. In some cases, the guide sequence has a length in a range of from 17-30 nucleotides (nt) (e.g., from 17-25, 17-22, 17-20, 19-30, 19-25, 19-22. 19-20, 20-30. 20-25, or 20-22 nt). In some cases, the guide sequence has a length in a range of from 17-25 nucleotides (nt) (e.g., from 17-22, 17-20, 19-25, 19-22, 19-20, 20-25, or 20-22 nt). In some cases, tire guide sequence has a length of 17 or more nt (e.g., 18 or more, 19 or more, 20 or more, 21 or more, or 22 or more nt; 19 nt, 20 nt, 21 nt, 22 nt, 23 nt, 24 nt, 25 nt, etc.). In some cases, the guide sequence has a length of 19 or more nt (e.g., 20 or more, 21 or more, or 22 or more nt; 19 nt, 20 nt, 21 nt, 22 nt, 23 nt, 24 nt, 25 nt, etc.). In some cases, tire guide sequence has a length of 17 nt. In some cases, the guide sequence has a length of 18 nt. In some cases, the guide sequence has a length of 19 nt. In some cases, the guide sequence has a length of 20 nt. In some cases, the guide sequence has a length of 21 nt. In some cases, the guide sequence has a length of 22 nt. In some cases, the guide sequence has a length of 23 nt.

[0380] In some cases, the spacer sequence has a length of from 15 to 50 nucleotides (e.g., from 15 nucleotides (nt) to 20 nt, from 20 nt to 25 nt, from 25 nt to 30 nt, from 30 nt to 35 nt, from 35 nt to 40 nt, from 40 nt to 45 nt, or from 45 nt to 50 nt).

[0381] A subject guide RNA can interact with a target nucleic acid (e.g., double stranded DNA (dsDNA), single stranded DNA (ssDNA), single stranded RNA (ssRNA), or double stranded RNA (dsRNA)) in a sequence-specific manner via hybridization (i.e., base pairing).

[0382] The guide RNA can be modified to hybridize to any desired target sequence (e.g., while taking the PAM into account, e.g.. when targeting a dsDNA target) within a target nucleic acid (e.g., a eukaryotic target nucleic acid such as genomic DNA). In some cases, the percent complementarity between tire spacer sequence of the guide and the target site of the target nucleic acid is 60% or more (e.g., 65% or more, 70% or more, 75% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%). In some cases, the percent complementarity between the spacer and the target site of the target nucleic acid is 80% or more (e.g., 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%). In some cases, the percent complementarity between the spacer and the target site of the target nucleic acid is 90% or more (e.g., 95% or more, 97% or more, 98% or more. 99% or more, or 100%). In some cases, the percent complementarity between the spacer and the target site of the target nucleic acid is 100%.

[0383] In some cases, the percent complementarity between the spacer sequence and the target site of the target nucleic acid is 100% over an at least 5-nucleotide contiguous region of the spacer. In some cases, the percent complementarity between tire guide sequence and the target site of the target nucleic acid over an at least 6-nucleotide contiguous region of the spacer is 60% or more (e.g., 70% or more, 75% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%). In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid over an at least 7-nucleotide contiguous region of the spacer is 60% or more (e.g., 70% or more, 75% ormore, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%). In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid over an at least 8-nucleotide contiguous region of the spacer is 60% or more (e.g., 70% or more, 75% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%). In some cases, the percent complementarity between tire guide sequence and the target site of the target nucleic acid over an at least 9-nucleotide contiguous region of the spacer is 60% or more (e.g., 70% or more, 75% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%). In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid over an at least 10-nucleotide contiguous region of the spacer is 60% or more (e.g., 70% or more, 75% or more, 80% or more, 85% or more, 90% or more, 95% or more. 97% or more, 98% or more, 99% or more, or 100%). In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid over an at least 11 -nucleotide contiguous region of the spacer is 60% or more (e.g., 70% or more, 75% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%). In some cases, the percent complementarity between tire guide sequence and the target site of the target nucleic acid over an at least 12-nucleotide contiguous region of the spacer is 60% or more (e.g., 70% or more, 75% or more, 80% or more. 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%). In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid over an at least 13-nucleotide contiguous region of the spacer is 60% or more (e.g., 70% or more, 75% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%). In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid over an at least 14-nucleotide contiguous region of the spacer is 60% or more (e.g., 70% or more. 75% or more, 80% or more, 85% or more, 90% or more, 95% or more. 97% or more, 98% or more, 99% or more, or 100%). In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid over an at least 15-nucleotide contiguous region of the spacer is 60% or more (e.g., 70% or more, 75% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%). In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid over an at least 16-nucleotide contiguous region of the spacer is 60% or more (e.g., 70% or more, 75% or more, 80% or more, 85% or more. 90% or more, 95% or more. 97% or more, 98% or more, 99% or more, or 100%). In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid over an at least 17-nucleotide contiguous region of the spacer is 60% or more (e.g., 70% or more, 75% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%). In some cases, thepercent complementarity between the guide sequence and the target site of the target nucleic acid over an at least 18-nucleotide contiguous region of the spacer is 60% or more (e.g., 70% or more, 75% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%). In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid over an at least 19-nucleotide contiguous region of the spacer is 60% or more (e.g., 70% or more, 75% or more, 80% or more, 85% or more, 90% or more. 95% or more, 97% or more, 98% or more, 99% or more, or 100%). In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid over an at least 20-nucleotide contiguous region of the spacer is 60% or more (e.g., 70% or more, 75% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%). In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid over an at least 21 -nucleotide contiguous region of the spacer is 60% or more (e.g., 70% or more, 75% or more. 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%). In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid over an at least 22 -nucleotide contiguous region of the spacer is 60% or more (e.g., 70% or more, 75% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%).

[0384] In some cases, the percent complementarity between the spacer sequence and the target site of the target nucleic acid is 100% over an at least 5-10 nucleotide contiguous region of tire spacer. In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid over an at least 6-11 nucleotide contiguous region of the spacer is 60% or more (e.g.. 70% or more. 75% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%). In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid over an at least 7-12 nucleotide contiguous region of the spacer is 60% or more (e.g., 70% or more, 75% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%). In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid over an at least 8-13 nucleotide contiguous region of the spacer is 60% or more (e.g., 70% or more, 75% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%). In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid over an at least 9-14 nucleotide contiguous region of the spacer is 60% or more (e.g., 70% or more, 75% or more, 80% or more, 85% or more, 90% or more, 95% or more. 97% or more, 98% or more, 99% or more, or 100%). In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid over an at least 10-15 nucleotide contiguous region of the spacer is 60% or more (e.g., 70% or more, 75% or more, 80% or more, 85% ormore, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%). In some cases, the percent complementarity between tire guide sequence and the target site of the target nucleic acid over an at least 11-16 nucleotide contiguous region of the spacer is 60% or more (e.g., 70% or more, 75% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%). In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid over an at least 12-17 nucleotide contiguous region of the spacer is 60% or more (e.g.. 70% or more. 75% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%). In some cases, the percent complementarity between the guide sequence and the target site of tire target nucleic acid over an at least 13-18 nucleotide contiguous region of the spacer is 60% or more (e.g., 70% or more, 75% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%). In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid over an at least 14-19 nucleotide contiguous region of the spacer is 60% or more (e.g., 70% or more, 75% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%). In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid over an at least 15-20 nucleotide contiguous region of the spacer is 60% or more (e.g., 70% or more, 75% or more, 80% or more, 85% or more. 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%). In some cases, tire percent complementarity between the guide sequence and the target site of the target nucleic acid over an at least 16-21 nucleotide contiguous region of the spacer is 60% or more (e.g., 70% or more, 75% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%). In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid over an at least 17-22 nucleotide contiguous region of the spacer is 60% or more (e.g., 70% or more. 75% or more. 80% or more, 85% or more, 90% or more, 95% or more, 97% or more. 98% or more, 99% or more, or 100%). In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid over an at least 18-23 nucleotide contiguous region of the spacer is 60% or more (e.g., 70% or more, 75% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%). In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid over an at least 19-24 nucleotide contiguous region of the spacer is 60% or more (e.g.. 70% or more, 75% or more, 80% or more, 85% or more, 90% or more. 95% or more, 97% or more, 98% or more. 99% or more, or 100%). In some cases, tire percent complementarity between the guide sequence and the target site of the target nucleic acid over an at least 20-25 nucleotide contiguous region of the spacer is 60% or more (e.g., 70% or more, 75% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%). In some cases, thepercent complementarity between the guide sequence and the target site of the target nucleic acid over an at least 21-26 nucleotide contiguous region of the spacer is 60% or more (e.g., 70% or more, 75% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%). In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid over an at least 22-27 nucleotide contiguous region of the spacer is 60% or more (e.g., 70% or more. 75% or more. 80% or more, 85% or more, 90% or more, 95% or more, 97% or more. 98% or more, 99% or more, or 100%).

[0385] In various embodiments, the guide RNAs may have a scaffold or core region that complexes with a cognate nucleic acid programmable DNA binding protein (e.g., CRISPR Cas9 or Casl2a). In some cases, a guide scaffold can have two stretches of nucleotides that are complementary to one another and hybridize to form a double stranded RNA duplex (dsRNA duplex). Thus, in some cases, the protein binding segment of a guide RNA includes a dsRNA duplex. In some embodiments, the dsRNA duplex region includes a range of from 5-25 base pairs (bp) (e.g., from 5-22, 5-20, 5-18, 5-15, 5-12, 5-10. 5-8, 8-25, 8-22, 8-18, 8-15, 8-12. 12- 25, 12-22, 12-18. 12- 15, 13-25, 13-22. 13-18, 13-15. 14-25, 14-22. 14-18, 14-15, 15-25, 15-22, 15-18. 17-25, 17-22, or 17-18 bp, e.g., 5 bp, 6 bp, 7 bp, 8 bp, 9 bp, 10 bp, etc.). In some cases, the dsRNA duplex region includes a range of from 6-15 base pairs (bp) (e.g., from 6-12, 6-10, or 6-8 bp, e.g., 6 bp, 7 bp, 8 bp, 9 bp, 10 bp, etc.). In some cases, the duplex region includes 5 or more bp (e.g., 6 or more, 7 or more, or 8 or more bp). In some cases, tire duplex region includes 6 or more bp (e.g., 7 or more, or 8 or more bp). In some cases, not all nucleotides of the duplex region are paired, and therefore the duplex fonning region can include a bulge. Tire term "bulge” herein is used to mean a stretch of nucleotides (which can be one nucleotide) that do not contribute to a double stranded duplex, but which are surround 5’ and 3’ by nucleotides that do contribute, and as such a bulge is considered part of the duplex region. In some cases, the dsRNA includes 1 or more bulges (e.g., 2 or more, 3 or more, 4 or more bulges). In some cases, the dsRNA duplex includes 2 or more bulges (e.g., 3 or more, 4 or more bulges). In some cases, the dsRNA duplex includes 1-5 bulges (e.g., 1-4, 1- 3. 2-5, 2-4, or 2-3 bulges).

[0386] Thus, in some cases, the stretches of nucleotides that hybridize to one another to form the dsRNA duplex in a guide scaffold region have 70%-100% complementarity (e.g.. 75%-100%, 80%-10%. 85%-100%, 90%- 100%, 95%- 100% complementarity) with one another. In some cases, the stretches of nucleotides that hybridize to one another to form the dsRNA duplex have 70%-100% complementarity (e.g., 75%-100%, 80%-10%, 85%- 100%, 90%-100%, 95%-100% complementarity) with one another. In some cases, the stretches of nucleotides that hybridize to one another to form the dsRNA duplex have 85%-100% complementarity (e.g., 90%-100%, 95%-100% complementarity) with one another. In some cases, thestretches of nucleotides that hybridize to one another to form the dsRNA duplex have 70%-95% complementarity (e.g., 75%-95%, 80%-95%, 85%-95%, 90%-95% complementarity) with one another. In other words, in some cases, the dsRNA duplex includes two stretches of nucleotides that have 70%-100% complementarity (e.g., 75%-100%, 80%-10%, 85%-100%, 90%-100%, 95 %- 100% complementarity) with one another. In some cases, the dsRNA duplex includes two stretches of nucleotides that have 85%-100% complementarity (e.g., 90%-100%, 95%-100% complementarity) with one another. In some cases, the dsRNA duplex includes two stretches of nucleotides that have 70%-95% complementarity (e.g., 75%-95%, 80%-95%, 85%-95%, 90%-95% complementarity) with one another.

[0387] In various embodiments, the scaffold region of a guide RNA can also include one or more (1, 2, 3, 4, 5, etc.) mutations relative to a naturally occurring scaffold region. For example, in some cases a base pair can be maintained while the nucleotides contributing to the base pair from each segment can be different. In some cases, the duplex region of a subject guide RNA includes more paired bases, less paired bases, a smaller bulge, a larger bulge, fewer bulges, more bulges, or any convenient combination thereof, as compared to a naturally occurring duplex region (of a naturally occurring guide RNA).

[0388] Examples of various guide RNAs can be found in tire art. and in some cases variations similar to those introduced into Cas9 guide RNAs can also be introduced into guide RNAs of the present disclosure (e.g., mutations to the dsRNA duplex region, extension of the 5’ or 3’ end for added stability for to provide for interaction with another protein, and the like). For example, see Jinek et aL, Science. 2012 Aug 17;337(6096): 816-21 ; Chylinski ct al., RNA Biol. 2013 May;10(5):726- 37; Ma ct al., Biomcd Res Int. 2013;2013:270805; Hou et al., Proc Natl Acad Sci U S A. 2013 Sep 24;110(39): 15644-9; Jinek et al., Elife. 2013;2:e00471; Pattanayak et al., Nat Biotechnol. 2013 Sep;31(9):839-43; Qi et al. Cell. 2013 Feb 28 ;152(5): 1173-83 ; Wang et al., Cell. 2013 May 9; 153(4):910-8; Auer et al.. Genome Res. 2013 Oct 31; Chen et al., Nucleic Acids Res. 2013 Nov 1 ;41 (20):el9; Cheng et al., Cell Res. 2013 Oct;23(10): 1163-71; Cho et al., Genetics. 2013 Nov;195(3): 1177-80; DiCarlo et al., Nucleic Acids Res. 2013 Apr;41(7):4336-43;Dickinson ct al., Nat Methods. 2013 Oct; 10(10): 1028-34; Ebina ct al., Sci Rep. 2013;3:2510; Fujii ct. al, Nucleic Acids Res. 2013 Nov I;4 l(20):el87; Hu et al., Cell Res. 2013 Nov;23(ll): 1322-5; Jiang et al., Nucleic Acids Res. 2013 Nov l;41(20):el88; Larson et al.. Nat Protoc. 2013 Nov;8(l 1):2180-96; Mali et. at., Nat Methods. 2013 Oct; 10(10):957-63; Nakayama et al., Genesis. 2013 Dec;51(12):835-43; Ran et al., Nat Protoc. 2013 Nov;8(l 1):2281 -308; Ran et al., Cell. 2013 Sep 12; 154(6): 1380-9; Upadhyay et al., G3 (Bethesda). 2013 Dec 9;3(12):2233-8; Walsh et aL, Proc Natl Acad Sci U S A. 2013 Sep 24; 110(39): 15514- 5; Xie et al., Mol Plant. 2013 Oct 9; Yang et al., Cell. 2013 Sep 12; 154(6): 1370-9; Briner et al., Mol Cell. 2014 Oct 23;56(2):333-9; and U.S. patents and patent applications: 8.906,616; 8,895,308; 8,889,418;8,889,356: 8,871,445; 8,865,406; 8,795,965; 8,771,945; 8,697,359; 20140068797; 20140170753; 20140179006; 20140179770; 20140186843; 20140186919; 20140186958; 20140189896; 20140227787; 20140234972; 20140242664; 20140242699; 20140242700; 20140242702; 20140248702; 20140256046; 20140273037; 20140273226; 20140273230; 20140273231: 20140273232; 20140273233; 20140273234; 20140273235; 20140287938; 20140295556; 20140295557: 20140298547; 20140304853; 20140309487; 20140310828; 20140310830; 20140315985; 20140335063: 20140335620; 20140342456; 20140342457; 20140342458; 20140349400; 20140349405: 20140356867; 20140356956; 20140356958; 20140356959; 20140357523; 20140357530; 20140364333; and 20140377868; all of which are hereby incorporated by reference in their entirety.Guide RNA modifications

[0389] In one embodiment, the guide RNAs contemplated herein comprise non-naturally occurring nucleic acids and / or non-naturally occurring nucleotides and / or nucleotide analogs, and / or chemical modifications. Non-naturally occurring nucleic acids can include, for example, mixtures of naturally and non-naturally occurring nucleotides. Non-naturally occurring nucleotides and / or nucleotide analogs may be modified at the ribose, phosphate, and / or base moiety. In an embodiment of the invention, a guide RNA component nucleic acid comprises ribonucleotides and non-ribonucleotides. In one such embodiment, a guide RNA component comprises one or more ribonucleotides and one or more deoxyribonucleotides. In an embodiment of the invention, the guide RNA (including pegRNA) component comprises one or more non-naturally occurring nucleotide or nucleotide analog such as a nucleotide with phosphorothioate linkage, a locked nucleic acid (LNA) nucleotides comprising a methylene bridge between the 2' and 4' carbons of the ribose ring, or bridged nucleic acids (BNA).

[0390] Other examples of modified nucleotides include 2'-O-methyl analogs, 2'-deoxy analogs, or 2'-fluoro analogs. Further examples of modified bases include, but are not limited to, 2-aminopurine, 5-bromo-uridine, pseudouridine, inosine, 7-methylguanosine. Examples of coRNA chemical modifications include, without limitation, incorporation of 2'-O-mcthyl (M), 2'-O-mcthyl 3 'phosphorothioate (MS), S-constraincd ethyl(cEt), or 2'-O-methyl 3 'thioPACE (MSP) at one or more terminal nucleotides. Such chemically modified oRNA components can comprise increased stability and increased activity as compared to unmodified oRNA components, though on-target vs. off-target specificity is not predictable. (See, Hendel, 2015, Nat Biotechnol. 33(9):985-9, doi: 10.1038 / nbt.3290, published online 29 June 2015 Ragdarm et al., 0215, PNAS, E7110-E7111; Allerson et al., J. Med. Chem. 2005, 48:901-904; Bramsen et al., Front. Genet., 2012, 3: 154; Deng et al., PNAS, 2015, 112: 11870-11875; Sharma et al., MedChemComm., 2014, 5: 1454- 1471; Hendel et al., Nat. Biotechnol. (2015) 33(9): 985-989; Li et al., Nature Biomedical Engineering, 2017,1, 0066 DOI: 10.1038 / s41551-017-0066). In one embodiment, the 5’ and / or 3’ end of a guide RNA (including pegRNA) component is modified by a variety of functional moieties including fluorescent dyes, polyethylene glycol, cholesterol, proteins, or detection tags. (See Kelly et al., 2016, J. Biotech. 233:74-83). In one embodiment, a guide RNA (including pegRNA) component comprises ribonucleotides in a region that binds to a target sequence and one or more deoxyribonucletides and / or nucleotide analogs in a region that binds to a nucleic acid programmable DNA binding protein (e.g.. Cas9 nickase).

[0391] In an embodiment, deoxyribonucleotides and / or nucleotide analogs are incorporated in engineered guide RNA component structures. In one embodiment, 3-5 nucleotides at either the 3?or the 5’ end of a guide RNA component is chemically modified. In one embodiment, only minor modifications are introduced in the seed region, such as 2’-F modifications. In one embodiment, 2’-F modification is introduced at the 3’ end of a guide RNA component. In one embodiment, three to five nucleotides at the 5’ and / or tire 3’ end of the reRNA component are chemically modified with 2' -O-methyl (M), 2’-O-methyl 3’ phosphorothioate (MS), S- constrained ethyl(cEt), or 2’ -O-methyl 3’ thioPACE (MSP). Such modification can enhance genome editing efficiency (see Hendel et al., Nat. Biotechnol. (2015) 33(9): 985-989). In one embodiment, all of the phosphodiester bonds of a guide RNA (including pegRNA) component are substituted with phosphorothioates (PS) for enhancing levels of gene disruption. In other embodiments, one or more of the phosphodiester bonds of a guide RNA are substituted with phosphorothioate (PS), i.e., a phosphorothioate bond / linkage. For example, at least four of the phosphodiester bonds of a guide RNA may be substituted with phosphorothioates (PS). A phosphorothioate bond may be provided between the first and second nucleotides from the 5 ’ end of the guide RNA, and / or between the second and third nucleotides from the 5 ’ end of the guide RNA, and / or between the first and second nucleotides from the 3’ end of the guide RNA, and / or between the second and third nucleotides from the 3’ end of the guide RNA.

[0392] In one embodiment, more than five nucleotides at the 5’ and / or the 3’ end of the guide RNA (including pegRNA) component are chemically modified with 2’-0-Me, 2’-F or S-constrained ethyl(cEt). Such chemically modified guide RNA (including pegRNA) component can mediate enhanced levels of gene disruption (see Ragdami et al., 0215, PNAS, E7110-E7111). In an embodiment of the invention, a guide RNA (including pegRNA) component is modified to comprise a chemical moiety at its 3’ and / or 5’ end. Such moieties include, but are not limited to amine, azide, alkyne, thio, dibenzocyclooctyne (DBCO), or Rhodamine. In certain embodiment, the chemical moiety is conjugated to the guide RNA (including pegRNA) component by a linker, such as an alkyl chain. In one embodiment, the chemical moiety of the modified nucleic acid component can be used to attach the guide RNA (including pegRNA) component to another molecule, such as DNA, RNA, protein, or nanoparticles. Such chemically modified guide RNA(including pegRNA) component can be used to identify or enrich cells generically edited by a gene editing system described herein.

[0393] Other guide RNA modifications are described in Kim, D.Y., Lee, J.M., Moon, S.B. et al. Efficient CRISPR editing with a hypercompact Casl2fl and engineered guide RNAs delivered by adeno-associated virus. Nat Biotechnol 40, 94-102 (2022).

[0394] Accordingly, in various aspects of the invention, the guide RNA are modified in one or more locations within the molecule. MSI, an internal penta(uridinylate) (UUUUU) sequence in the tracrRNA; MS2, the 3' terminus of the crRNA; MS3, the ‘stem 1’ region of the tracrRNA; MS4, the tracrRNA-crRNA complementary region; and MS5. the ‘stem 2’ region of the tracrRNA.

[0395] Various aspects of the invention provide methods and compositions for improved guide RNA stability via chemical modifications. Braasch, D. A., Jensen, S., Liu, Y ., Kaur, K., Arar, K., White. M. A., et al. (2003). RNA interference in mammalian cells by chemically-modified RNA. Biochemistry 42, 7967- 7975. doi: 10.102 l / bi0343774. Chiu, Y. L , and Rana, T. M. (2003). siRNA function in RNAi: a chemical modification analysis. RNA 9, 1034-1048. doi: 10.1261 / ma.5103703. Behlke, M. A. (2008). Chemical modification of siRNAs for in vivo use. Oligonucleotides , 305-319. doi: 10.1089 / oli.2008.0164. Bennett, C. F., and Swayze, E. E. (2010). RNA targeting therapeutics: molecular mechanisms of antisense oligonucleotides as a therapeutic platform. Annu. Rev. Pharmacol. Toxicol. 50, 259-293. doi:10.1146 / annurev.pharmtox.010909. 105654. Deleavey, G. F., and Damha, M. J. (2012). Designing chemically modified oligonucleotides for targeted gene silencing. Chem. Biol. 19, 937-954. doi: 10.1016 / j.chembiol.2012.07.011. Lennox, K. A., and Behlke, M. A. (2020). Chemical modifications in RNA interference and CRISPR / Cas genome editing reagents. Methods Mol. Biol. 2115, 23-55. doi: 10.1007 / 978-1- 0716-0290-4_2.

[0396] For instance, Hendel et al. improved guide RNA stability by chemically modifying gRNA ends to reduce degradation by exonucleases, RNA nuclease. Hendel. A., Bak, R. O., Clark, J. T.. Kennedy, A. B., Ryan. D. E., Roy. S., et al. (2015a). Chemically modified guide RNAs enhance CRISPR-Cas genome editing in human primary' cells. Nat. Biotechnol. 33, 985-989. doi: 10.1038 / nbt.3290. Chemical modifications of gRNAs may enable more efficient and safer gene-editing in primary cells suitable for clinical applications.

[0397] A review of types of chemical modifications are provided in the table below. Allen, Daniel et al. ‘‘Using Synthetically Engineered Guide RNAs to Enhance CRISPR Genome Editing Systems in MammalianCells.” Frontie rs in genome editing vol. 2 617910. 28 Jan. 2021, doi: 10.3389 / fgeed.2020.617910.

[0398] Accordingly, in various embodiments of the present invention, the genome editing system comprising a guide RNA and further comprises one or more chemical modifications selected from, but not limited to tire modifications in the above table.

[0399] In exemplary embodiments, chemical modifications to the guide RNA (including pegRNA) include modifications on the ribose rings and phosphate backbone of guide RNA (including pegRNA) and modifications at the 2'OH include 2'-0-Me. 2'-F, and 2'F-ANA. More extensive ribose modifications include 2'F-4'-Ca-OMe and 2',4'-di-Ca-OMe combine modification at both the 2' and 4' carbons. Phosphodiester modifications include sulfide-based Phosphorothioate (PS) or acetate-based phosphonoacetate alterations. Combinations of the ribose and phosphodiester modifications have given way to formulations such as 2'-O- methyl 3 'phosphorothioate (MS), or 2'-O-methyl-3 '-thioPACE (MSP), and 2 '-O-methyl-3 '-phosphonoacetate (MP) RNAs. Locked and unlocked nucleotides such as locked nucleic acid (LNA), bridged nucleic acids (BNA), S-constrained ethyl (cEt), and unlocked nucleic acid (UNA) are examples of sterically hindered nucleotide modifications. Modifications to make a phosphodiester bond between the 2' and 5' carbons (2',5'- RNA) of adjacent RNAs as well as a butane 4-carbon chain link between adjacent RNAs have been described.

[0400] In certain embodiments, the guide RNA comprises one or more hairpins as depicted in the appended Drawings. Preferably, the guide RNA comprises 0 -10 hairpins. In some embodiments, the guide RNA comprises 1-3 hairpins. In some embodiments, the guide RNA comprises 2 hairpins. More preferably, a hairpin comprises 6-20 ribonucleotides.

[0401] Modification of the sgRNA is also an efficient way of enhancing the efficiency of the CRISPR-Cas systems. Kim, Daesik et al. “Evaluating and Enhancing Target Specificity of Gene-Editing Nucleases and Deaminases.’’ A mual review of biochemistry vol. 88 (2019): 191-220. doi: 10. 1146 / annurev-biochem- 013118-111730. For instance, adding a “U4AU6” (SEQ ID NO:2662) motif at the end of the crRNA Bin Moon, Su et al. “Highly efficient genome editing by CRISPR-Cpfl using CRISPR RNA with a uridinylate- rich 3’-overhang.” Nature communications vol. 9,1 3651. 7 Sep. 2018, doi: 10.1038 / s41467-018-06129-w or using a pol- Il-driven truncated pre-tRNA Zhang, Xuhua et al. “Genetic editing and interrogation with Cpfl and caged truncated pre-tRNA-like crRNA in mammalian cells.” Cell discovery vol. 4 36. 10 Jul. 2018, doi: 10.1038 / s41421-018-0035-0 have been demonstrated.

[0402] Accordingly, various embodiments provide for the modification of the sgRNA to enhance the efficiency of the CRISPR-Cas 12a systems and modifications to express the crRNA to improve the activity of the CRISPR- Casl2a system.

[0403] Additional embodiments provide guide RNA modifications including but not limited to one or more chemical modifications selected from 2'-0-Me, 2'-F, and 2'F-ANA at 2'OH; 2'F-4'-Ca-OMe and 2',4'-di-Ca- OMe at 2' and 4' carbons; phosphodiester modifications comprising sulfide-based Phosphorothioate (PS) or acetate-based phosphonoacetate alterations; combinations of the ribose and phosphodiester modifications; locked nucleic acid (LNA), bridged nucleic acids (BN A). S-constrained ethyl (cEt), and unlocked nucleic acid (UNA); modifications to produce a phosphodiester bond between the 2' and 5' carbons (2'.5'-RNA) of adjacent RNAs; and a butane 4-carbon chain link between adjacent RNAs.

[0404] The present disclosure provides guide RNAs comprising one or more nucleotides having a 2'-0-Me modification and one or more phosphorothioate linkages. In some embodiments, the guide RNA comprises at least four phosphorothioate linkages, i.e., at least four phosphates of the guide RNA have been converted to phosphorothioate. In some embodiments, four phosphates of the guide RNA have been converted to phosphorothioate. For example, in certain embodiments the first two and last two phosphates of tire guide RNA have been converted to phosphorothioate.

[0405] Hie guide RNA disclosed herein may comprise a direct repeat sequence and a spacer sequence contiguous with the direct repeat sequence. Direct repeat sequences and spacer sequences are described elsewhere herein. In some embodiments, the direct repeat sequence consists of between 10 and 30 nucleotides. For example, the direct repeat sequence may consist of 10, 1 1, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides. In certain embodiments, the direct repeat sequence consists of 19 nucleotides. For example, the direct repeat sequence may consist of AAUUUCUACUGUUGUAGAU (SEQ ID NO: 2584).

[0406] The guide RNA disclosed herein may comprise a modified nucleotide, such as a 2'-0-Me modified nucleotide, at certain positions of the direct repeat sequence, as described elsewhere herein. The guide RNA disclosed herein may comprise a modified nucleotide, such as a 2'-0-Me modified nucleotide, at certain positions of the spacer sequence, as described elsewhere herein. In certain embodiments, one or more nucleotides are added to the 5’ end of the direct repeat sequence. For example, the guide RNA may comprise a UG or UUUU immediately 5’ to the 5’ end of tire direct repeat sequence. In some embodiments, the position of the modified nucleotide is identified by reference to tire direct repeat sequence of the guide RNA, wherein position 1 is the first nucleotide of the direct repeat sequence. Accordingly, for example, where the direct repeat sequence is 19 nucleotides in length and a modified nucleotide is indicated to be at position 5. the modified nucleotide will be at position 5 of the direct repeat sequence. Where the position of a modification is at aposition beyond the 3' end of the direct repeat sequence, the modified nucleotide will be in another portion of the guide RNA. For example, in some embodiments, the direct repeat sequence iscontiguous with the spacer sequence. In such embodiments, modified nucleotides indicated to be at positions beyond the 3' end of the direct repeat sequence may be located in the spacer sequence. For example, in some embodiments, if the direct repeat sequence is 19 nucleotides in length and immediately followed by the spacer sequence, a modified nucleotide indicated to be at position 27 is the 8th nucleotide in the spacer sequence. In some embodiments, one or more nucleotides are added to the 5' end of the direct repeat sequence, and the one or more nucleotides are modified. In such embodiments, the modified nucleotides will be at positions preceding the 5’ end of the direct repeat sequence. A first nucleotide immediately preceding the 5’ end of the direct repeat sequence is indicated to be at position -1. A second nucleotide preceding tire first nucleotide is indicated to be at position -2. Accordingly, for example, where the guide RNA comprises a UG immediately 5’ to the 5’ end of the direct repeat sequence, and U and G are modified nucleotides, the modified nucleotides are indicated to be at positions -2 and -1, respectively. In another example, where the guide RNA comprises UUUU immediately 5’ to the 5’ end of the direct repeat sequence, and UUUU are all modified nucleotides, the modified nucleotides are indicated to be at positions -4, -3, -2 and -1.

[0407] In some embodiments, the guide RNAs disclosed herein target a Cas Type V polypeptide to a target sequence, such as a Cas Type V polypeptide disclosed herein or a polypeptide comprising a Cas Type V polypeptide as disclosed herein, and complexes comprising the guide RNA and the Cas Type V polypeptide are capable of inducing indel formation in a target polynucleotide sequence. Tire target sequence may be a DNA sequence. In some embodiments, the target sequence may be associated with a disease or disorder, hi some embodiments, 2'-0-Me modifications and phosphorothioate linkages as disclosed herein enhance the ability of the guide RNA to induce indels in comparison with guide RNAs that do not comprise 2'-0-Me modifications or phosphorothioate linkages. For example, in some embodiments the guide RNA targets a Cas Type V polypeptide to a target sequence, and a complex comprising the guide RNA and the Cas Type V polypeptide induces indel formation when contacted with a population of the target sequences under conditions suitable for inducing indel formation, wherein the contacting results in an increase in the percentage of target polynucleotide sequences comprising an indel in the population of target sequences of at least 20%, at least 30%. at least 40%. at least 50%, at least 60%, at least 70%, at least 80%. at least 90%. at least 100%, at least 110%, or at least 120% as compared to the percentage of target polynucleotide sequences comprising an indel in the population when the population is contacted with a complex comprising a reference guide RNA and the Cas Type V polypeptide, wherein the reference guide RNA consists of the same nucleotide sequence as the guide RNA and wherein the nucleotides are not modified (i.e., the guide RNA does not comprise 2'-0-Me modifications or phosphorothioate linkages).

[0408] In still other embodiments, the guide RNAs disclosed herein may be modified by introducing additional RNA motifs into the guide RNAs, e.g., at the 5' and 3' termini of the guide RNAs. Such structures may include, but are not limited to RNA hairpins, RNA step-loops, RNA quadruplexes, cap structures, and poly(A) tails, or ribozyme functions and the like. Also, guide RNAs could also be modified to include one or more nuclear localization sequences.

[0409] The present disclosure provides guide RNAs comprising (i) one or more NLS sequences, optionally wherein the NLS is chemically conjugated to the guide RNA: (ii) one or more RNA-NLS sequences, optionally wherein the gRNA is fused directly or indirectly to the RNA-NLS: and / or (iii) a PNA probe comprising a polynucleotide sequence complementary with one or more portions of the gRNA, wherein the PNA probe is conjugated to one or more NLS sequences; and optionally wherein the gRNA comprises one or more linkers.

[0410] In some embodiments, the PNA binds to the 3’ end of the gRNA. In some embodiments, the guide RNA comprises a PNA binding sequence at the 5’ end or 3’ end, and the PNA probe binds to the PNA binding sequence. In certain embodiments, the PNA may range from 2-30 nucleotides, such as 9 nucleotides.

[0411] The PNA probes may be modified to enhance solubility. For example, the PNA probe may comprise a solubility enhancing group such as ethylene glycol. Tire PNA probes may comprise a fluorophore such as Cy5. In various embodiments, the PNA probe comprises certain sequences and / or structures as described elsewhere herein.

[0412] The present disclosure further provides guide RNAs comprising (i) a 5’ DNA extension sequence and / or a 5’ RNA extension sequence; and / or (ii) a 3’ DNA extension sequence and / or a 3’ RNA extension sequence. The guide RNA may be a guide RNA as described herein. The extension sequence may comprise 4 or more nucleotides. In certain embodiments, the extension sequence consists of 4-25 nucleotides. In certain embodiments, the extension sequence consists of 4 nucleotides. In certain embodiments, the extension sequence consists of 9 nucleotides. In certain embodiments, the extension sequence consists of 15 nucleotides. In certain embodiments, the extension sequence consists of 25 nucleotides. Exemplary extension sequences are described elsewhere herein.

[0413] In some embodiments, the 3’ DNA or 3’ RNA extension sequence anneals to a spacer sequence in the guide RNA to form a hairpin structure. Hairpin structures are described elsewhere herein. In certain embodiments, the 3’DNA or 3’RNA extension sequence anneals to the 3’ end of the spacer sequence, thereby forming a hairpin structure. In certain embodiments, the last nucleotide at the 3’ end of tire extension sequence anneals to the last nucleotide at the 3" end of the spacer sequence. In certain embodiments, the last nucleotide at the 3' end of the extension sequence anneals to a nucleotide within 1, 2. 3, 4, 5, 6. 7, 8, 9, 10,11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 nucleotides of the last nucleotide at the 3’ end of the spacer sequence. As described elsewhere herein, the guide RNA may comprise one or more hairpin structures.

[0414] Guide RNAs disclosed herein may further comprise one or more linkers. In certain embodiments, the guide RNA comprises one or more linkers between the 5 ’ DNA and / or RNA extension sequences, and / or between the 3’ DNA and / or RNA extension sequences. Linkers are described elsewhere herein.

[0415] Any of the features of guide RNAs disclosed herein may be combined.

[0416] In some embodiments, the guide RNAs disclosed herein targets a Cas Type V polypeptide to a target sequence, such as a Cas Type V polypeptide disclosed herein or a polypeptide comprising a Cas Type V polypeptide as disclosed herein, and complexes comprising the guide RNA and the Cas Type V polypeptide are capable of inducing indel fomiation in a target polynucleotide sequence. The target sequence may be a DNA sequence. In some embodiments, the target sequence may be associated with a disease or disorder. In some embodiments, the 5' DNA extension sequences, 5’ RNA extension sequences, 3’ DNA extension sequences, and / or 3’ RNA extension sequences enhance the ability of the guide RNA to induce indels in comparison with guide RNAs that do not comprise 5’ DNA extension sequences, 5’ RNA extension sequences, 3’ DNA extension sequences, or 3’ RNA extension sequences. For example, in some embodiments the guide RNA targets a Cas Type V polypeptide to a target sequence and a complex comprising the guide RNA and the Cas Type V polypeptide induces indel fonnation when contacted with a population of the target sequences under conditions suitable for inducing indel formation, wherein the contacting results in an increase in the percentage of target polynucleotide sequences comprising an indel in the population of target sequences of at least 1.5-fold, at least 2-fold, at least 2.5-fold, at least 3-fold, at least 3.5-fold, at least 4-fold, at least 4.5-fold, at least 5-fold, at least 5.5-fold, at least 6-fold, at least 6.5-fold, at least 7-fold, at least 7.5-fold, at least 8-fold, at least 8.5-fold, at least 9-fold, at least 9.5-fold, or at least 10- fold as compared to the percentage of target polynucleotide sequences comprising an indel in the population when the population is contacted with a reference guide RNA and the Cas Type V polypeptide, wherein the reference guide RNA consists of the same nucleotide sequence as the guide RNA and wherein the guide RNA does not comprise a DNA or RNA extension sequence.

[0417] The disclosure provides a gene editing system comprising: (a) one or more Cas Type V polypeptides or one or more polynucleotides encoding the Cas Type V polypeptide(s); and (b) one or more polynucleotide sequences comprising a guide RNA (gRNA) of the disclosure or one or more polynucleotides encoding the gRNA, wherein the gRNA comprises a complementary sequence to that of a targeted polynucleotide sequence. The Cas Type V polypeptide may be a Cas Type V polypeptide as disclosed herein.

[0418] The disclosure provides a method of modifying a targeted polynucleotide sequence, the method comprising contacting the targeted polynucleotide sequence with a gene editing system as disclosed herein. The method may be carried out in vitro, ex vivo or in vivo. In some embodiments, the disclosure provides a method of modifying a targeted polynucleotide sequence, tire method comprising introducing into a cell the gene editing system of the disclosure.

[0419] The disclosure provides one or more polynucleotides encoding a guide RNA as disclosed herein or a gene editing system as disclosed herein. The disclosure further provides one or more vectors comprising the one or more polynucleotides. The disclosure further provides a cell comprising the one or more vectors or the one or more polynucleotides.

[0420] Any of the guide RNAs; gene editing systems; polynucleotides encoding the guide RNAs or gene editing systems: vectors; cells; or pharmaceutical compositions disclosed herein can be used in medicine, for example for use in treating or preventing a disease or disorder as disclosed elsewhere herein.

[0421] Additional RNA motifs could also improve function or stability of the guide RNAs. Addition of dimerization motifs - such as kissing loops or a GNRA tetraloop / tetraloop receptor pair - at the 5' and 3' termini of tire guide RNAs could also result in effective circularization of the guide RNAs, improving stability. Additionally, it is envisioned that addition of these motifs could enable the physical separation of guide RNA components, e.g.. separation of the Casl2a binding region from the spacer sequence. Short 5' extensions or 3 ’ extensions to the guide RNAs that form a small toehold hairpin at either or both ends of the guide RNAs could also compete favorably against the annealing of intracomplcmcntary regions along the length of the guide RNAs. Finally, kissing loops could also be used to recruit other RNAs or proteins to the genomic site targeted by the guide RNA.

[0422] Guide RNAs could be further improved via directed evolution, in an analogous fashion to how7protein function can be improved. Directed evolution could enhance guide RNA function and / or reduce offsite targeting and / or indels and / or improve precise editing efficiency.

[0423] The present disclosure contemplates any such w ays to further improve tire stability and / or functionality of the guide RNAs disclosed here.

[0424] In some embodiments, the RNAs (including the guide RNAs) used in the compositions of the disclosure have undergone a chemical or biological modification to render them more stable. Exemplary modifications to an RNA include the depletion of a base (e.g., by deletion or by the substitution of one nucleotide for another) or modification of a base, for example, the chemical modification of a base. The phrase "chemical modifications" as used herein, includes modifications which introduce chemistries whichdiffer from those seen in naturally occurring RNA, for example, covalent modifications such as the introduction of modified nucleotides, (e.g., nucleotide analogs, or the inclusion of pendant groups which are not naturally found in such mRNA molecules).

[0425] Other suitable polynucleotide modifications that may be incorporated into the RNAs used in the compositions of the disclosure include, but are not limited to, 4'- thio-modified bases: 4'-thio-adenosine, 4'- thio-guanosine, 4'-thio-cytidine, 4'-thio-uridine, 4'- thio-5-mcthyl-cytidine. 4'-thio-pseudouridine, and 4'-thio- 2-thiouridine. pyridin-4-one ribonucleoside, 5 -aza-uridine, 2-thio-5 -aza-uridine, 2-thiouridine, 4-thio- pseudouridine, 2- thio-pseudouridine. 5-hydroxyuridine, 3 -methyluridine, 5-carboxymethyl-uridine, 1- carboxymethyl-pseudouridine, 5-propynyl-uridine, 1-propynyl-pseudouridine, 5- taurinomethyluridine, 1- taurinomethyl-pseudouridine, 5-taurinomethyl-2 -thio-uridine, 1- taurinomethyl-4-thio-uridine, 5-methy 1- uridine, 1 -methyl -pseudouridine, 4-thio-l -methyl- pseudouridine, 2-thio-l -methyl -pseudouridine, 1-methyl- 1-deaza-pseudouridine, 2-thio-l - methyl- 1-deaza-pseudouridine, dihydrouridine, dihydropseudouridine, 2- thio-dihydrouridine, 2-thio-dihydropseudouridine, 2-methoxyuridine. 2-methoxy-4-thio-uridine, 4-methoxy- pseudouridine, 4-methoxy-2 -thio-pseudouridine, 5 -aza-cytidine, pseudoisocytidine, 3-methyl- cytidine. N4- acetylcytidine, 5 -formylcytidine, N4-methylcytidine, 5 -hydroxymethyl cytidine, 1-methyl-pseudoisocytidine, pyrrolo-cytidine, pyrrolo-pseudoisocytidine, 2-thio-cytidine, 2- thio-5 -methyl -cytidine, 4-thio- pseudoisocytidine, 4-thio-l -methyl -pseudoisocytidine, 4-thio- 1 -methy l- 1-deaza-pseudoisocytidine, 1- methyl-l-deaza-pseudoisocytidine, zebularine, 5-aza- zebularine, 5-methyl-zebularine, 5-aza-2-thio- zebularine, 2-thio-zebularine. 2-methoxy- cytidine, 2-methoxy-5-methyl-cytidine, 4-methoxy- pseudoisocytidine, 4-methoxy-l -methyl- pseudoisocytidine, 2-aminopurine, 2,6-diaminopurine, 7-deaza- adenine, 7-deaza-8-aza- adenine, 7-deaza-2-aminopurine, 7-deaza-8-aza-2 -aminopurine, 7-deaza-2,6- diaminopurine, 7-deaza-8-aza-2,6-diaminopurine, 1 -methyladenosine, N6-methyladenosine, N6- isopentenyladenosine, N6-(cis-hydroxyisopentenyl)adenosine, 2-methylthio-N6-(cis- hy droxy isopentenyl)adenosine, N6-glycinylcarbamoyladenosine, N6- threonylcarbamoyladenosine, 2- methy lthio-N6-threonyl carbamoyladenosine. N6,N6- dimethyladenosine, 7-methyladenine, 2-methylthio- adenine, and 2-methoxy-adenine. inosine. 1-methyl-inosine, wyosine, wybutosine, 7-deaza-guanosine. 7- deaza-8-aza-guanosine, 6-thio- guanosine, 6-thio-7-deaza-guanosine, 6-thio-7-deaza-8-aza-guanosine, 7- methyl-guanosine, 6-thio-7-methyl-guanosine, 7-m ethy linosine, 6-methoxy-guanosine, 1 -methylguanosine, N2- mcthylguanosinc, N2,N2-dimcthy Iguanosinc, 8-oxo-guanosinc, 7-mcthyl-8-oxo-guanosinc, l-mcthyl-6- thio-guanosine, N2-methy 1-6-thio-guanosine, and N2,N2-dimethyl-6-thio- guanosine, and combinations thereof. The tenn modification also includes, for example, the incorporation of non-nucleotide linkages or modified nucleotides into the mRNA sequences of the present invention (e.g., modifications to one or both of the 3' and 5' ends of an mRNA molecule encoding a functional protein or enzyme). Such modificationsinclude the addition of bases to an mRNA sequence (e.g., the inclusion of a poly A tail or a longer poly A tail), the alteration of the 3' UTR or the 5' UTR, complexing the mRNA with an agent (e.g., a protein or a complementary nucleic acid molecule), and inclusion of elements which change the structure of an RNA molecule (e.g., which form secondary structures).

[0426] In some embodiments, RNAs (e.g., guide RNAs) include a 5' cap structure. A 5' cap is typically added as follows: first, an RNA terminal phosphatase removes one of tire terminal phosphate groups from the 5' nucleotide, leaving two terminal phosphates; guanosine triphosphate (GTP) is then added to the tenninal phosphates via a guanylyl transferase, producing a 5'5'5 triphosphate linkage: and the 7-nitrogen of guanine is then methylated by a methyltransferase. Examples of cap structures include, but are not limited to, m7G(5')ppp (5'(A,G(5')ppp(5')A and G(5')ppp(5')G. Naturally occurring cap structures comprise a 7- methyl guanosine that is linked via a triphosphate bridge to the 5 '-end of the first transcribed nucleotide, resulting in a dinucleotide cap of m7G(5')ppp(5')N, where N is any nucleoside. In vivo, the cap is added enzymatically. Hie cap is added in the nucleus and is catalyzed by the enzyme guanylyl transferase. The addition of the cap to the 5' terminal end of RNA occurs immediately after initiation of transcription. The terminal nucleoside is typically a guanosine, and is in the reverse orientation to all the other nucleotides, i.e., G(5')ppp(5')GpNpNp.

[0427] Additional cap analogs include, but are not limited to, a chemical structures selected from the group consisting of m7GpppG, m7GpppA, m7GpppC: unmethylated cap analogs (e.g., GpppG): dimethylated cap analog (e.g., m2,7GpppG), trimethylated cap analog (e.g., m2,2,7GpppG), dimethylated symmetrical cap analogs (e.g., m7Gpppm7G), or anti reverse cap analogs (e.g., ARCA; m7,2'OmcGpppG, m72'dGpppG, m7,3'OmeGpppG, m7,3'dGpppG and their tetraphosphate derivatives) (see. e.g., Jemielity, J. et al., "Novel 'anti-reverse' cap analogs with superior translational properties", RNA, 9: 1108-1122 (2003)).

[0428] Typically, tire presence of a "tail" serves to protect the RNA (e.g., guide RNAs) from exonuclease degradation. A poly A or poly U tail is thought to stabilize natural messengers and synthetic sense RNA. Therefore, in certain embodiments a long poly A or poly U tail can be added to an RNA molecule thus rendering the RNA more stable. Poly A or poly U tails can be added using a variety of art-recognized techniques. For example, long poly A tails can be added to synthetic or in vitro transcribed RNA using poly A polymerase (Yokoe, et al. Nature Biotechnology.1996; 14: 1252-1256). A transcription vector can also encode long poly A tails. In addition, poly A tails can be added by transcription directly from PCR products. Poly A may also be ligated to the 3' end of a sense RNA with RNA ligase (see, e.g., Molecular Cloning A Laboratory Manual, 2nd Ed., ed. by Sambrook, Fritsch and Maniatis (Cold Spring Harbor Laboratory Press: 1991 edition)).

[0429] Typically, the length of a poly A (SEQ ID NO:2652) or poly U (SEQ ID NO:2653) tail can be at least about 10, 50, 100, 200, 300, 400 at least 500 nucleotides. In some embodiments, a poly-A tail on the 3' terminus of mRNA typically includes about 10 to 300 adenosine nucleotides (e.g., about 10 to 200 adenosine nucleotides, about 10 to 150 adenosine nucleotides, about 10 to 100 adenosine nucleotides, about 20 to 70 adenosine nucleotides, or about 20 to 60 adenosine nucleotides). In some embodiments, mRNAs include a 3' poly(C) tail structure. A suitable poly-C (SEQ ID NO:2654) tail on the 3' terminus of mRNA typically include about 10 to 200 cytosine nucleotides (e.g., about 10 to 150 cytosine nucleotides, about 10 to 100 cytosine nucleotides, about 20 to 70 cytosine nucleotides, about 20 to 60 cytosine nucleotides, or about 10 to 40 cytosine nucleotides). Tire poly-C tail may be added to the poly-A or poly U tail or may substitute the poly-A or poly U tail.

[0430] RNAs according to tire present disclosure (e.g., Casl2a guide RNAs) may be synthesized according to any of a variety of known methods. For example, RNAs according to the present invention may be synthesized via in vitro transcription (IVT). Briefly, IVT is typically perfonned with a linear or circular DNA template containing a promoter, a pool of ribonucleotide triphosphates, a buffer system ...

Claims

CLAIMS1. A Cas Type V polypeptide comprising an amino acid sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity to any one of the amino acid sequence sequences selected from SEQ ID NO: 334 (No. ID405), SEQ ID NO: 58 (No. ID414), SEQ ID NO: 564 (No. ID418), SEQ ID NO: 335 (No. ID406), SEQ ID NO: 331 (No. ID411), SEQ ID NO: 20 (No. ID415), and SEQ ID NO: 445 (No. ID419), wherein the polypeptide comprises one or more substitutions at positions selected from: K292, Q492, T306. L748, A739. F297. N536, D851. N770, D675, K955. S972, N798, K298, H853, E377. V295, K852, F854. N1015, T845, E1019, K745, Q565, K745, 1723, K821, K721, A810, K768, S16, 128, V312, V929, L474, H470, E402, K207, V294, R212, F79, Q299, Q465, F537, K1284, V929, DI 192, Q1203, T838, N735, 11063, K933, N1249, 197, R212, A1251, N1264, R1138, N588, 1584, V453, Q475, N76, R304, 1471, E544, V254, D231, H496, V534, D577, 128, A183, D341, K595, V501, S568, H470, V254, D231 and F197, based on the amino acid sequence provided in SEQ ID NO: 334 (No. ID405), or one or more corresponding substitutions in any of the amino acid sequences selected from: SEQ ID NO: 58 (No. ID414), SEQ ID NO: 564 (No. ID418), SEQ ID NO: 335 (No. ID406), SEQ ID NO: 331 (No. ID411), SEQ ID NO: 20 (No. ID4L5), and SEQ ID NO: 445 (No. ID419).

2. Hie Cas Type V polypeptide of claim 1, comprising one or more substitutions selected from: Q492X, I L748X. A739X, F297X, N536X, D851X. N770X, D675X, K955X. S972X, N798X, K298X, H853X, E37 V295X, K852X, F854X, N1015X, T845X, E1019X, K745X and K292X, based on the ammo acid sequence provided in SEQ ID NO: 334 (No. ID405), or one or more corresponding substitutions in any of the amino acid sequences selected from: SEQ ID NO: 58 (No. ID414), SEQ ID NO: 564 (No. ID418), SEQ ID NO: 335 (No.D406), SEQ ID NO: 331 (No. ID411), SEQ ID NO: 20 (No. ID415), and SEQ ID NO: 445 (No. ID419), wherein < is arginine, histidine or lysine. i . The Cas Type V polypeptide of claim 1 or 2, wherein, excluding the one or more substitutions, the Cas Type V polypeptide is at least 95% identical to any one of: SEQ ID NO: 334 (No. ID405), SEQ ID NO: 58 (No. ID414), SEQ ID NO: 564 (No. ID418), SEQ ID NO: 335 (No. ID406), SEQ ID NO: 331 (No. ID411), SEQ ID NO: 20 (No. ID415), or SEQ ID NO: 445 (No. ID419).

4. The Cas Type V polypeptide of any preceding claim, comprising one or more substitutions selected from: K955X, F297X, T306X, E1018X, E377X, K852X, K298X, T845X, V295X, N536X, Q492X, N770X, N101 and K292X, based on the amino acid sequence provided in SEQ ID NO: 334 (No. ID405), or one or morecorresponding substitutions in any of the amino acid sequences selected from: SEQ ID NO: 58 (No. ID414), SEQ ID NO: 564 (No. ID418), SEQ ID NO: 335 (No. ID406), SEQ ID NO: 331 (No. ID411), SEQ ID NO: 20 (No. ID415), and SEQ ID NO: 445 (No. ID419), wherein X is arginine, histidine or lysine.

5. The Cas Type V polypeptide of any preceding claim, comprising one or more substitutions selected from: K292X. Q492X, N536X, N770X and N 1015X. based on the amino acid sequence provided in SEQ ID NO: 334 (No. ID405), or one or more corresponding substitutions in any of the amino acid sequences selected from: SEQ ID NO: 58 (No. ID414), SEQ ID NO: 564 (No. ID4I8), SEQ ID NO: 335 (No. ID406), SEQ ID NO: 331 (No. ID411), SEQ ID NO: 20 (No. ID415), and SEQ ID NO: 445 (No. ID419), wherein X is arginine, histidine or lysine.

6. The Cas Type V polypeptide of any preceding claim, comprising the substitutions K292X and D169X. based on the amino acid sequence provided in SEQ ID NO: 334 (No. ID405), or one or more corresponding substitutions in any of the amino acid sequences selected from: SEQ ID NO: 58 (No. ID414), SEQ ID NO: 564 (No. ID418), SEQ ID NO: 335 (No. ID406), SEQ ID NO: 331 (No. ID411), SEQ ID NO: 20 (No. ID415), and SEQ ID NO: 445 (No. ID419), wherein X is arginine, histidine or lysine.

7. The Cas Type V polypeptide of any preceding claim, comprising one or more combinations of substitut selected from:N536X and N770X;N536X and N1015X;K292X, N536X and NI015X;K292X and N770X;K292X, N536X and N770X;K292X, N770X and N1015X;K292X and N1015X;K292X and Q492X;K292X, Q492X and N536X;K292X, Q492X and N770X;K292X, Q492X and NI0I5X;N536X, N770X and N1015X;Q492X, N536X and N770X;Q492X, N536X and N1015X;Q492X and N770X;Q492X, N770X and N1015X;Q492X and N536X;K292X and N536X;N770X and N1015X:Q492X and N1015X:K292X, Q492X, N536X, N770X and N1015X;K292X, N536X, N770X and N1015X;K292X, Q492X, N536X and N770X;K292X, Q492X, N536X and N1015X;K292X, Q492X, N770X and N1015X; and / orQ492X, N536X, N770X and N1015X; based on the amino acid sequence provided in SEQ ID NO: 334 (No. ID405), or one or more corresponding combinations of substitutions in any of the amino acid sequences selected from: SEQ ID NO: 58 (No. ID414), SEQ ID NO: 564 (No. ID418), SEQ ID NO: 335 (No. ID406), SEQ ID NO: 331 (No. ID411), SEQ ID NO: 20 (No. ID415), and SEQ ID NO: 445 (No. ID419), wherein X is arginine, histidine or lysine.

8. The Cas Type V polypeptide of any preceding claim, comprising the substitution D169X, based on the acid sequence provided in SEQ ID NO: 334 (No. ID405), or a corresponding substitution in any of the am sequences selected from: SEQ ID NO: 58 (No. ID414), SEQ ID NO: 564 (No. ID418), SEQ ID NO: 335 (No. ID406), SEQ ID NO: 331 (No. ID411), SEQ ID NO: 20 (No. ID415), and SEQ ID NO: 445 (No. ID419), wherein X is arginine, histidine or lysine.>. The Cas Type V polypeptide of any preceding claim, wherein X is arginine.

0. Tire Cas Type V polypeptide of any preceding claim, comprising a D169R substitution, based on the amino acid sequence provided in SEQ ID NO: 334 (No. ID405), or a corresponding substitution in any of the amino acid sequences selected from: SEQ ID NO: 58 (No. ID414), SEQ ID NO: 564 (No. ID418), SEQ ID NO: 335 (No. ID406). SEQ ID NO: 331 (No. ID411), SEQ ID NO: 20 (No. ID415), and SEQ ID NO: 445 (No. ID419).1 1 . The Cas Type V polypeptide of any preceding claim, comprising one or more substitutions at positions selected from: Q565, K745, 1723, K821, K721, A810, K768, S16, 128, V312, V929, L474, H470, E402, K2i V294, R212, F79, Q299, Q465, F537, K1284, V929, DI 192, Q1203, T838, N735, 11063, K933, N1249, 197R212, A1251, N1264, R1138, N588, 1584, V453, Q475, N76, R304, 1471, E544, V254, D231, H496, V534, D577, 128, A183, D341, K595, V501, S568, H470, V254, D231 and F197, based on the amino acid sequence provided in SEQ ID NO: 334 (No. ID405), or one or more corresponding substitutions in any of the amino acid sequences selected from: SEQ ID NO: 58 (No. ID414), SEQ ID NO: 564 (No. ID418), SEQ ID NO: 335 (No. ID406). SEQ ID NO: 331 (No. ID411), SEQ ID NO: 20 (No. ID415), and SEQ ID NO: 445 (No. ID419).

12. The Cas Type V polypeptide of any preceding claim, comprising:(a) one or more substitutions at positions selected from: Q565, K745, 1723, K821, K721, A810, K768, S16, 128, V312, V929; and / or(b) one or more combinations of substitutions at positions:L474, H470 and E402;K207, V294, R212. F79 and Q299;Q465 and F537;K1284, V929, DI 192 and Q1203;T838 and N735;11063, K933 and N1249;197 and R212;A1251. N1264 and R1138;N588, 1584, V453 and Q475;N76 and R304;1471 and E544;V254 and D231;H496, V534 and D577;128, A183 and D341;K595, V501, S568 and H470;V254, D231 and Fl 97;V254 and D341; and / orV254, D341 and F197; based on the amino acid sequence provided in SEQ ID NO: 334 (No. ID405), or one or more corresponding substitutions or combinations of substitutions in any of the amino acid sequences selected from: SEQ ID NO: 58 (No. ID414), SEQ ID NO: 564 (No. ID418), SEQ ID NO: 335 (No. ID406), SEQ ID NO: 331 (No. ID41 1), SEO ID NO: 20 (No. ID415), and SEQ ID NO: 445 (No. ID419).

13. The Cas Type V polypeptide of any preceding claim, wherein comprising one or more substitutions selected from: Q565K, K745M, I723V, K821I, K721M, A810V, K768M, S16T, I28F, V312E, V929I, L474Q, H470N, E402V K207N, V294M, R212Q, F79S, Q299E, Q465R, F537I, K1284R, V929D, D1192E, Q1203H, T838S, N735I, I1063N, K933M, N1249S, I97T, R212L, A1251V, N1264I, R1138S, N588I, I584S, V453E, Q475H, N76D, R304S, I471T, E544D, V254M, D231V, H496N, V534F, D577G. I28F. A183T, D341V, K595T, V501D, S568I, H470L, V254M, D23 IV and F197V, based on the amino acid sequence provided in SEQ ID NO: 334 (No. ID405), or one or more corresponding substitutions in any of the amino acid sequences selected from: SEQ ID NO: 58 (No. ID414), SEQ ID NO: 564 (No. ID418), SEQ ID NO: 335 (No. ID406), SEQ ID NO: 331 (No. ID411), SEQ ID NO: 20 (No. ID415), and SEQ ID NO: 445 (No. ID419).

14. The Cas Type V polypeptide of any preceding claim, comprising:(a) one or more substitutions selected from: Q565K, K745M. I723V, K821I, K721M, A810V, K768M. S16T, I28F, V312E, V929I; and / or(b) one or more combinations of substitutions selected from:L474Q, H470N and E402V;K207N, V294M, R212Q, F79S and Q299E;Q465R and F537I;K1284R. V929D. DI 192E and Q1203H;T838S and N735I;I1063N, K933M and N1249S;I97T and R212L;A1251V, N1264I and R1138S;N588I, I584S, V453E and Q475H;N76D and R304S;147 IT and E544D;V254M and D231V;H496N, V534F and D577G;I28F, A183T and D341V;K595T. V501D, S568I and H470L:V254M, D231V and F197V:V254M and D341V; and / orV254M, D341V and Fl 97V, based on the amino acid sequence provided in SEQ ID NO: 334 (No. ID405), or one or more correspondingsubstitutions or combinations of substitutions in any of the amino acid sequences selected from: SEQ ID NO: 58 (No. ID414), SEQ ID NO: 564 (No. ID418), SEQ ID NO: 335 (No. ID406), SEQ ID NO: 331 (No. ID411), SEQ ID NO: 20 (No. ID415), and SEQ ID NO: 445 (No. ID419).

15. The Cas Type V polypeptide of any preceding claim, comprising the following substitutions: K207N. V294M, R212Q. F79S and Q299F, based on the amino acid sequence provided in SEQ ID NO: 334 (No. ID405), or one or more corresponding substitutions or combinations of substitutions in any of the amino acid sequences selected from: SEQ ID NO: 58 (No. ID414), SEQ ID NO: 564 (No. ID418), SEQ ID NO: 335 (No. ID406), SEQ ID NO: 331 (No. ID411), SEQ ID NO: 20 (No. ID415), and SEQ ID NO: 445 (No. ID419).

16. The Cas Type V polypeptide of any preceding claim, wherein the substitutions are based on the amino acid sequence provided in SEQ ID NO: 334 (No. ID405).

17. Tire Cas Type V polypeptide of any preceding claim, comprising or consisting of:(a) an amino acid sequence having at least 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity to one or more regions of any one of SEQ ID NOs: 2306-2351 corresponding to positions: 18-23; 49-57; 100-105; 128-132; 158-182; 252-311; 359-405; 440-450; 485-499; 528-569; 609-628; 672-677; 734-750; 768-776;812; 840-864; 926-977; and 996, based on the numbering of SEQ ID NO: 334 (No. 1D405); or(b) an amino acid sequence having at least 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence to all of the following regions of any one of SEQ ID NOs: 2306-2351 corresponding to positions: 18-23; 49-57; 100-105; 128-132; 158-182; 252-311; 359-405; 440-450; 485-499; 528-569; 609-628; 672-677; 734-750; 768- 776; 790-812; 840-864; 926-977; and 996, based on the numbering of SEQ ID NO: 334 (No. ID405); or(c) an amino acid sequence having at least 95%, 96%, 97%, 98%, 99%, 99.5%. or 100% sequence identity o any one of SEQ ID NOs: 2306-2351. 1570-1600, 1632-1670, or 2354-2380.

8. Tire Cas Type V polypeptide of any preceding claim, wherein the Cas Type V polypeptide induces indel rbnnation when contacted with a population of target sequences and a guide RNA targeting the Cas Type V polypeptide to the target sequence under conditions suitable for inducing indel formation, and wherein the contacting results in an increase in the percentage of target polynucleotide sequences comprising an indel in the population of target sequences of at least 20%, at least 30%, at least 40%. or at least 50% as compared to the percentage of target polynucleotide sequences comprising an indel in the population when the population is contacted with a reference polypeptide comprising the sequence of SEQ ID NO: 1554 (No. ID405-1) and a j RNA that targets the reference polypeptide to the target polynucleotide sequence.

19. A fusion protein comprising the Cas Type V polypeptide of any preceding claim fused to a heterologous amino acid sequence.

20. A gene editing system comprising:(a) one or more Cas Type V polypeptides of any preceding claim or one or more polynucleotides encoding the Cas Type V polypeptide (s); and(b) one or more polynucleotide sequences comprising a guide RNA (gRNA) or one or more polynucleotides encoding tire gRNA, wherein the gRNA comprises a complementary sequence to that of a targeted polynucleotide sequence.

21. One or more polynucleotides encoding the Cas Type V polypeptide, fusion protein, or gene editing system of any preceding claim.

22. One or more vectors comprising the one or more polynucleotides of claim 21.

23. A cell comprising the one or more vectors of claim 22 or the one or more polynucleotides of claim 21.

24. A pharmaceutical composition comprising the Cas Type V polypeptide of any one of claims 1-18, the protein of claim 19, the gene editing system of claim 20, the one or more polynucleotides of claim 21, the one or more vectors of claim 22, or the cell of claim 23.

15. A method of modifying a targeted polynucleotide sequence, the method comprising introducing into a cell a ;cnc editing system according to claim 20.