Gene editing components, systems, and methods of use
The Cas V-type-based gene editing systems address inefficiencies in existing CRISPR-Cas V-type systems by using a V-type polypeptide and guide RNA complex for precise gene editing, enhancing editing efficiency and treatment of genetic disorders.
Patent Information
- Application Number
- JP2025502584
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-04-10
- Filing Date
- 2023-07-17
- Publication Date
- 2025-08-05
AI Technical Summary
Existing CRISPR-Cas V-type systems face challenges in achieving sufficient editing efficiency, accuracy, affordability, scalability, and effectiveness for treating genetic disorders and complex diseases, necessitating improvements for better delivery and performance.
Development of Cas V-type-based gene editing systems comprising a V-type polypeptide and a V-type guide RNA that form a complex to target and bind to specific nucleic acid sequences, with optional accessory proteins for enhanced functionality, delivered via various methods including DNA, RNA, protein, and ribonucleoprotein complexes, utilizing lipid nanoparticles for efficient delivery.
The system achieves precise gene editing with improved efficiency, accuracy, and scalability, enabling effective treatment of genetic disorders and complex diseases.
Smart Images

Figure 2025525569000001_ABST
Abstract
Description
[Technical Field]
[0001] Related Applications This application is a continuation of U.S. Provisional Application Serial No. 63 / 368,722, filed July 18, 2022 (Attorney Docket No. CSG001-P1); U.S. Provisional Application Serial No. 63 / 368,724, filed July 18, 2022 (Attorney Docket No. CSG001-P2); U.S. Provisional Application Serial No. 63 / 368,726, filed July 18, 2022 (Attorney Docket No. CSG001-P3); U.S. Provisional Application Serial No. 63 / 368,728, filed July 18, 2022 (Attorney Docket No. CSG001-P4); U.S. Provisional Application Serial No. 63 / 368,730 filed on July 18, 2022 (Attorney Docket No. CSG001-P5); U.S. Provisional Application Serial No. 63 / 368,731 filed on July 18, 2022 (Attorney Docket No. CSG001-P6); U.S. Provisional Application Serial No. 63 / 368,734 filed on July 18, 2022 (Attorney Docket No. CSG001-P7); U.S. Provisional Application Serial No. 63 / 368,735 filed on July 18, 2022 (Attorney Docket No. CSG001-P8); U.S. Provisional Application Serial No. 63 / 368,736 (Attorney Docket No. CSG001-P9); U.S. Provisional Application Serial No. 63 / 368,737 filed July 18, 2022 (Attorney Docket No. CSG001-P10); U.S. Provisional Application Serial No. 63 / 368,738 filed July 18, 2022 (Attorney Docket No. CSG001-P11); U.S. Provisional Application Serial No. 63 / 368,741 filed July 18, 2022 (Attorney Docket No. CSG001-P12); U.S. Provisional Application Serial No. 63 / 368,742 filed July 18, 2022 (Attorney Docket No. No. CSG001-P13; U.S. Provisional Application Serial No. 63 / 368,744, filed July 18, 2022 (Attorney Docket No. CSG001-P14); U.S. Application Serial No. 18 / 297,346, filed April 7, 2023 (Attorney Docket No. CSG002-T1); and U.S. Provisional Application Serial No. 63 / 495,198, filed April 10, 2023 (Attorney Docket No. CSG001-P15), each of which is incorporated herein by reference in its entirety.
[0002] The foregoing applications, and all documents cited therein or during their prosecution ("Application Citations"), and all documents cited or referenced in the Application Citations, and all documents cited or referenced herein ("Prescription Citations"), and all documents cited or referenced in the Prescription Citations, together with any manufacturer's instructions, manuals, product specifications, and product sheets for any products mentioned herein or in any document incorporated herein by reference, are hereby incorporated by reference herein and may be used in the practice of this invention. More specifically, all references are incorporated by reference to the same extent as if each individual document was specifically and individually indicated to be incorporated by reference.
[0003] The present disclosure generally relates to systems, methods, and compositions used for precise genome editing, including nucleic acid insertion, replacement, and deletion at targeted and precise genomic sites, which systems, methods, and compositions are based on novel and / or engineered class II / V clustered regularly interspaced short palindromic repeats (CRISPR)-Cas systems. [Background technology]
[0004] Genome editing tools encompass a diverse set of techniques that can produce many types of genome modifications in a variety of contexts. These techniques have evolved over the past two decades to provide a variety of user-programmable editing tools, including ZFN (zinc finger) nuclease editing systems, meganuclease editing systems, and TALEN (transcription activator-like effector nucleases). The past decade has seen explosive growth of a new generation of genome editing systems based on components from bacterial immune pathways, including CRISPR (clustered regularly interspaced short palindromic repeats) and related CRISPR-associated proteins (e.g., CRISPR-Cas9) (Jinek et al., “A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity,” Science, Vol. 337(6096), pp. 816-821), meganuclease editing agents (Boissel et al., “megaTALs: a rare-cleaving nuclease architecture for therapeutic genome engineering,” Nucleic Acids Research 42: pp. 2591-2601), and bacterial retron systems (Schubert et al., “High-throughput functional variant screens via in vivo production of single-stranded DNA,” PNAS, April 27, 2021, Vol. 118(18), pp. 1-10). In particular, CRISPR-Cas9 has evolved from its guide RNA-based programmable double-strand break activity to discover alternative CRISPR Cas nuclease enzymes (e.g., Cas12a, Cas12f, Cas13a, and Cas13b) with different PAM requirements and cleavage properties, leading to the development of base editing (Komor et al., “Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage,” Nature,May 19,2016,533(7603);pp.420-424 [cytosine base editors or CBEs] and Gaudelli et al., “Programmable base editing of AT to GC in genomic DNA without DNA cleavage,” Nature,Vol.551,pp.464-471 [adenine base editors or ABEs]), prime editing (Anzalone et al.,“Search-and-replace genome editing without double-strand breaks or donor DNA,” Nature,Dec 2019,576(7789):pp.149-157), twin prime editing (Anzalone et al.,“Programmable deletion,replacement,integration and inversion of large DNA sequences with twin prime editing,” Nature Biotechnology,Dec 9, 2021, vol. 40, pp. 731-740), epigenetic editing (Kungulovski and Jeltsch, “Epigenome Editing: State of the Art, Concepts, and Perspective,” Trends in Genetics, Vol. 32, 206, pp. 101-113), CRISPR-induced integrase editing (Yarnell et al., “Drag-and-drop genome insertion of large sequences without double-stranded DNA cleavage using CRISPR-directed integrases,” Nature Biotechnology, Nov 24, 2022, doi.org / 10.1038 / s41587-022-01527-4 (“PASTE”)) have been derivatized in numerous ways to form systems ranging from
[0005] In particular, the application of CRISPR-associated systems ("CRISPR-Cas systems") in human therapeutics is expected to be curative in ameliorating various monogenic diseases and disorders. For example, clinical trials are currently underway to treat transfusion-dependent β-thalassemia (TDT) and sickle cell disease (SCD) by autologous transfusion of CRISPR / Cas9-edited CD34+ hematopoietic stem cells. Frangoul, Haydar et al. "CRISPR-Cas9 Gene Editing for Sickle Cell Disease and β-Thalassemia." The New England journal of medicine vol. 384, 3 (2021): 252-260. doi: 10.1056 / NEJMoa2031054 and ATTR amyloidosis Gillmore, Julian D et al. "CRISPR-Cas9 In Vivo Gene Editing for Transthyretin Amyloidosis." The New England journal of medicine vol. 385, 6 (2021): 493-502. doi: 10.1056 / NEJMoa2107454 (incorporated herein by reference).
[0006] The potential of such CRISPR-Cas systems has sparked the discovery of many novel CRISPR-Cas variants, which have been classified into two classes (i.e., classes I and II), six types, and 33 subtypes based on the structure of their genes, protein subunits, and their gRNAs. Makarova, K.S., Wolf, Y.I., Iranzo, J. et al. Evolutionary classification of CRISPR-Cas systems: a burst of class 2 and derived variants. Nat Rev Microbiol 18, 67-83 (2020). doi:10.1038 / s41579-019-0299-x (incorporated herein by reference).
[0007] Among the diverse CRISPR-Cas systems, class II has the broadest applications in gene editing due to its earlier discovery and the resulting single effector protein. Meanwhile, type V effector nucleases are diverse due to the extensive diversity across the N-terminus of the protein, as revealed by comparing the crystal structures of Cas12a, Cas12b, and Cas12e type V nucleases (Tong et al., “The Versatile Type V CRISPR Effectors and Their Application Prospects,” Front. Cell Dev. Biol., 2021, vol. 8). The C-terminal region of type V effector nucleases is more highly conserved, but contains a conserved RuvC-like endonuclease (RuvC) domain. The RuvC domain of type V effectors has been reported to be derived from the TnpB protein, which is encoded by autonomous or non-autonomous transposons (Shmakov et al., “Diversity and evolution of class 2 CRISPR-Cas systems,” 2017, Nat. Rev. Microbiol. 15, 169-182. doi:10.1038 / nrmicro.2016.184). Type V systems are further subdivided into many subtypes, including types VA-VI, VK, VU, and CRISPR-CasΦ (Hajizadeh et al., “The expanding class 2 CRISPR toolbox: diversity, applicability, and targeting drawbacks,” 2019, BioDrugs 33, 503-513. doi:10.1007 / s40259-019-00369-y). The corresponding effector nucleases in these various subtypes exhibit a variety of different substrates, including those that act only on double-stranded DNA (dsDNA), those that act on both dsDNA and single-stranded DNA (ssDNA), and those that act on single-stranded RNA (ssRNA). Due to this versatility, the V-type CRISPR-Cas system has become the focus of recent research. While numerous CRISPR-Cas V-type systems have been used for various applications, including gene editing, reported drawbacks have been published indicating the need for improved CRISPR-Cas V-type systems for suitability for desired applications. Therefore, there remains much room for improvement and design to achieve effective V-type CRISPR-Cas systems for gene editing that have sufficient editing efficiency, improved accuracy, better delivery, remain affordable, are easily scalable, and have improved capabilities for treating various genetic disorders and complex diseases. [Prior art documents] [Non-patent literature]
[0008] [Non-Patent Document 1] Jinek et al., “A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity,” Science, Vol. 337(6096), pp. 816-821 [Non-patent document 2] Boissel et al., “megaTALs: a rare-cleaving nuclease architecture for therapeutic genome engineering,” Nucleic Acids Research 42:pp.2591-2601 [Non-patent document 3] Schubert et al., “High-throughput functional variant screens via in vivo production of single-stranded DNA,” PNAS,April 27,2021,Vol.118(18),pp.1-10 [Non-patent document 4] Komor et al., “Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage,” Nature, May 19, 2016, 533(7603);pp.420-424 [cytosine base editors or CBEs] [Non-patent document 5] Gaudelli et al., “Programmable base editing of AT to GC in genomic DNA without DNA cleavage,” Nature, Vol. 551, pp. 464-471 [adenine base editors or ABEs] [Non-patent document 6] Anzalone et al., “Search-and-replace genome editing without double-strand breaks or donor DNA,” Nature, Dec 2019, 576(7789):pp.149-157 [Non-Patent Document 7] Anzalone et al., “Programmable deletion, replacement, integration and inversion of large DNA sequences with twin prime editing,” Nature Biotechnology,Dec 9,2021,vol.40,pp.731-740 [Non-patent document 8] Kungulovski and Jeltsch, “Epigenome Editing:State of the Art,Concepts,and Perspective,” Trends in Genetics,Vol.32,206,pp.101-113 [Non-Patent Document 9] Yarnell et al.,“Drag-and-drop genome insertion of large sequences without double-stranded DNA cleavage using CRISPR-directed integrases,”
Outdoor Tools 10
Outdoor Content11
Outdoor Tools 12
Outdoor Content13
[0009] The present disclosure provides Cas V-type-based gene editing systems for use in a variety of applications, including precise gene editing in cells, tissues, organs, or organisms. In various embodiments, the Cas V-type-based gene editing systems include (a) a V-type polypeptide and (b) a V-type guide RNA capable of associating with the V-type polypeptide to form a complex such that the complex localizes to and binds to a target nucleic acid sequence (e.g., a genomic or plasmid target sequence). In various embodiments, the V-type polypeptide has nuclease activity that results in cleavage of both strands of DNA.
[0010] In various embodiments, the Cas V type polypeptide is a polypeptide selected from Table S15A or a polypeptide having at least 70%, 75%, 80%, 85%, 90%, 95%, 99%, or 100% sequence identity to a polypeptide from Table S15A. In various other embodiments, the Cas V type polypeptide is encoded by a polynucleotide sequence selected from Table S15B or a polynucleotide having at least 70%, 75%, 80%, 85%, 90%, 95%, 99%, or 100% sequence identity to a polynucleotide of Table S15B. In various other embodiments, the Cas12a guide RNA is selected from any Cas V type guide sequence disclosed in Table S15C or a nucleic acid molecule having at least 70%, 75%, 80%, 85%, 90%, 95%, 99%, or 100% sequence identity to a Cas12a guide sequence of Table S15C.
[0011] In various embodiments, a Cas V-type guide RNA can include (a) a portion that binds to or associates with a Cas V-type polypeptide and (b) a targeting sequence, i.e., a region containing a sequence complementary to a target nucleic acid sequence. For Cas V-type guide RNA design, just as with Cas9 guide RNAs, the target sequence is typically adjacent to a PAM sequence. However, for Cas V-types, the PAM sequence is typically TTTV (where V typically represents A, C, or G) in various embodiments. In various embodiments, the "V" in TTTV is immediately adjacent to the 5'-most base on the non-target strand of the protospacer element. For Cas9 guide RNA design, the PAM sequence is typically not included in the guide RNA design.
[0012] In various embodiments, guide RNAs for Cas V types are relatively short, approximately 40-44 bases in length. The portion of the target sequence that base pairs with the protospacer is 20-24 bases in length, with a defined section of approximately 20 bases that binds to the Cas V type.
[0013] In various embodiments, the nomenclature for the Cas V-type guide RNA is referred to as a "crRNA," and there is no Cas9-like "tracrRNA" component.
[0014] In other aspects, the Cas V-based gene editing system can include one or more additional accessory proteins with genome-modifying functions, including recombinases, invertases, nucleases, polymerases, ligases, deaminases, reverse transcriptases, or epigenetic modification functions. In various embodiments, the accessory proteins can be provided separately. In other embodiments, the accessory proteins can be fused to the Cas V-type nuclease, optionally with a linker.
[0015] In yet another embodiment, the present disclosure provides a delivery system for introducing a Cas V-based gene editing system or components thereof into a cell, tissue, organ, or organism. Depending on the format selected, the Cas V-based gene editing system and / or its individual or combined components can be delivered as a DNA molecule (e.g., encoded on one or more plasmids), an RNA molecule (e.g., a guide RNA for targeting the Cas V protein of the Cas V-based gene editing system, or a linear or circular mRNA encoding the Cas V protein or accessory protein components), a protein (e.g., a Cas12a polypeptide, an accessory protein with other function (e.g., a recombinase, nuclease, polymerase, ligase, deaminase, or reverse transcriptase), or a protein-nucleic acid complex (e.g., a complex between a guide RNA and a Cas V protein or a fusion protein containing a Cas V protein).
[0016] In another aspect, the present disclosure provides nucleic acid molecules encoding a Cas V-type-based gene editing system or components thereof. In yet another aspect, the present disclosure provides vectors for transferring and / or expressing the Cas V-type-based gene editing system, e.g., under in vitro, ex vivo, and in vivo conditions. In yet another aspect, the present disclosure provides cell delivery compositions and methods, including compositions for passive and / or active transport into cells (e.g., plasmids), delivery by viral-based recombinant vectors (e.g., AAV and / or lentiviral vectors), delivery by non-viral-based systems (e.g., liposomes and LNPs), and delivery by virus-like particles of the Cas V-type-based gene editing systems described herein. Depending on the delivery system used, the Cas V-type-based gene editing systems described herein can be delivered in the form of DNA (e.g., plasmids or DNA-based viral vectors), RNA (e.g., guide RNA and mRNA delivered by LNPs), a mixture of DNA and RNA, protein (e.g., virus-like particles), and ribonucleoprotein (RNP) complexes. Any suitable combination of approaches for delivering the components of the Cas V-based gene editing systems disclosed herein may be used.
[0017] In other embodiments, the Cas V-based gene editing system can include template DNA, e.g., a single- or double-stranded donor molecule (linear or circular), containing edits that can be used by cells to repair single or double break lesions introduced by the Cas V-based gene editing system using cellular repair processes including homology-dependent repair (HDR) (e.g., in dividing cells) or non-homologous end joining (NHEJ) (in non-dividing cells).
[0018] In one embodiment, each of the components of the Cas V-based gene editing system is delivered by an all-RNA system, e.g., delivery of one or more RNA molecules (e.g., mRNA and / or guide RNA) by one or more LNPs (wherein the one or more RNA molecules form guide RNAs and / or are translated into polypeptide components (e.g., Cas V polypeptides and / or any accessory proteins)), and, if appropriate or desired, a DNA or RNA-encoding template DNA molecule (e.g., donor template).
[0019] In yet another aspect, the disclosure provides methods for genome editing by introducing the Cas V-based gene editing system described herein into a cell containing a target editing site (e.g., under in vitro, in vivo, or ex vivo conditions), thereby resulting in editing at the target site. In other aspects, the disclosure provides formulations comprising any of the aforementioned components for delivery to cells and / or tissues, including in vitro, in vivo, and ex vivo delivery, recombinant cells and / or tissues modified by the recombinant Cas V-based gene editing system and methods described herein, and methods for modifying cells by genome editing using the Cas V-based gene editing system disclosed herein.
[0020] The present disclosure also provides methods for making the Cas V-based gene editing systems described herein, their protein and nucleic acid molecular components, vectors, compositions and formulations, as well as pharmaceutical compositions and kits for modifying cells under in vitro, in vivo, and ex vivo conditions that comprise the genome editing and / or modification systems disclosed herein.
[0021] In various aspects, the present invention provides a method for producing a medicament for the treatment of a pulmonary arthritis, comprising: (a) a nucleic acid sequence encoding a Cas V-type polypeptide having the amino acid sequence of SEQ ID NO: 334 (No. ID405), SEQ ID NO: 58 (No. ID414), or SEQ ID NO: 564 (No. ID418), SEQ ID NO: 335 (No. ID406), SEQ ID NO: 331 (No. ID411), SEQ ID NO: 20 (No. ID415), or SEQ ID NO: 445 (No. ID419); (b) Cas of SEQ ID NO: 334 (ID405), SEQ ID NO: 58 (ID414), or SEQ ID NO: 564 (ID418), SEQ ID NO: 335 (ID406), SEQ ID NO: 331 (ID411), SEQ ID NO: 20 (ID415), or SEQ ID NO: 445 (ID419) a nucleic acid sequence encoding a polypeptide at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8% or at least 99.9% identical to a type V polypeptide; (c) a nucleic acid sequence that is a degenerate variant of the nucleic acid sequence in (a) or (b); and (d) a nucleic acid sequence that hybridizes under stringent conditions with the nucleic acid sequence in (a) or (b). The present invention provides an isolated or recombinant polynucleotide comprising or consisting of a nucleic acid sequence selected from the group consisting of:
[0022] In a related aspect, the present invention provides a method for producing a medicament for the treatment of a pulmonary arthritis, comprising: (a) One or more crRNA direct repeat sequences or reverse complements selected from (Group 1) SEQ ID NOs: 7-12; (Group 2) SEQ ID NOs: 24-27; (Group 3) SEQ ID NOs: 36-39; (Group 4) SEQ ID NOs: 49-52; (Group 5) SEQ ID NOs: 63-68; (Group 6) SEQ ID NOs: 84-91; (Group 7) SEQ ID NOs: 106-111; (Group 8) SEQ ID NOs: 122-125; (Group 9) SEQ ID NOs: 211-290; (Group 10) SEQ ID NOs: 343-354; (Group 11) SEQ ID NOs: 374-379; (Group 12) SEQ ID NOs: 390-393; (Group 13) SEQ ID NOs: 411-422; and (Group 14) SEQ ID NOs: 500-541; (b) 20-35 nucleotides from the 3' end of the crRNA direct repeat sequence or reverse complement (a) linked to a targeting guide attached to the 3' end of the direct repeat sequence, which is 16-30 nucleotides in length, or up to the length of said crRNA; (c) (Group 1) SEQ ID NOs: 13 to 15; (Group 2) SEQ ID NOs: 28 to 29; (Group 3) SEQ ID NOs: 40 to 41; (Group 4) SEQ ID NOs: 53 to 54; (Group 5) SEQ ID NOs: 69 to 71; (Group 6) SEQ ID NOs: 92 to 95; (Group 7) SEQ ID NOs: 112 to 114; (Group 8) SEQ ID NOs: 126 to 127; (Group 9) SEQ ID NOs: 291 to 330; (Group 10) SEQ ID NOs: 355 to 360; (Group 11) SEQ ID NOs: 380 to 382; (Group 12) SEQ ID NOs: 394 to 395; (Group 13) SEQ ID NOs: 423 to 428; and (Group 14) SEQ ID NOs: 542 to 563; (d) Nucleic acid sequences that are degenerate variants of (Group 1) SEQ ID NOs: 13 to 15; (Group 2) SEQ ID NOs: 28 to 29; (Group 3) SEQ ID NOs: 40 to 41; (Group 4) SEQ ID NOs: 53 to 54; (Group 5) SEQ ID NOs: 69 to 71; (Group 6) SEQ ID NOs: 92 to 95; (Group 7) SEQ ID NOs: 112 to 114; (Group 8) SEQ ID NOs: 126 to 127; (Group 9) SEQ ID NOs: 291 to 330; (Group 10) SEQ ID NOs: 355 to 360; (Group 11) SEQ ID NOs: 380 to 382; (Group 12) SEQ ID NOs: 394 to 395; (Group 13) SEQ ID NOs: 423 to 428; and (Group 14) SEQ ID NOs: 542 to 563; (e) (Group 1) SEQ ID NOs: 13 to 15; (Group 2) SEQ ID NOs: 28 to 29; (Group 3) SEQ ID NOs: 40 to 41; (Group 4) SEQ ID NOs: 53 to 54; (Group 5) SEQ ID NOs: 69 to 71; (Group 6) SEQ ID NOs: 92 to 95; (Group 7) SEQ ID NOs: 112 to 114; (Group 8) SEQ ID NOs: 126 to 127; (Group 9) SEQ ID NOs: 291 to 330; (Group 10) SEQ ID NOs: 355 to 360; (Group 11) SEQ ID NOs: 380 to 382; (Group 12) SEQ ID NOs: 394 to 395; (Group 13) SEQ ID NOs: Nos. 423-428; and (Group 14) nucleic acid sequences at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or at least 99.9% identical to SEQ ID NOs: 542-563; and (f) Nucleic acid sequences that hybridize under stringent conditions to (Group 1) SEQ ID NOs: 13 to 15; (Group 2) SEQ ID NOs: 28 to 29; (Group 3) SEQ ID NOs: 40 to 41; (Group 4) SEQ ID NOs: 53 to 54; (Group 5) SEQ ID NOs: 69 to 71; (Group 6) SEQ ID NOs: 92 to 95; (Group 7) SEQ ID NOs: 112 to 114; (Group 8) SEQ ID NOs: 126 to 127; (Group 9) SEQ ID NOs: 291 to 330; (Group 10) SEQ ID NOs: 355 to 360; (Group 11) SEQ ID NOs: 380 to 382; (Group 12) SEQ ID NOs: 394 to 395; (Group 13) SEQ ID NOs: 423 to 428; and (Group 14) SEQ ID NOs: 542 to 563. The present invention provides an isolated or recombinant guide RNA comprising or consisting of a nucleic acid sequence selected from the group consisting of:
[0023] In some embodiments, an isolated or recombinant polynucleotide comprising or consisting of a nucleic acid sequence encoding one or more Cas V-type polypeptides of the present disclosure is paired with one or more cognate guide RNAs of the present disclosure.
[0024] In certain exemplary aspects, the present disclosure provides: (a) one or more polypeptide sequences comprising at least 45%, 46%, 47%, 48%, 49%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, or 65% sequence identity to any one of the sequences selected from SEQ ID NO: 334 (No. ID405), SEQ ID NO: 58 (No. ID414), or SEQ ID NO: 564 (No. ID418), SEQ ID NO: 335 (No. ID406), SEQ ID NO: 331 (No. ID411), SEQ ID NO: 20 (No. ID415), and SEQ ID NO: 445 (No. ID419); and (b) one or more polynucleotide sequences comprising a guide RNA, wherein the guide RNA comprises a sequence complementary to that of the targeted polynucleotide sequence; A Cas V-type gene editing system comprising:
[0025] In various embodiments, there is provided a method of modifying a targeted polynucleotide sequence, the method comprising: (a) one or more polypeptide sequences comprising at least 45%, 46%, 47%, 48%, 49%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, or 65% sequence identity to any one of the sequences selected from SEQ ID NO: 334 (No. ID405), SEQ ID NO: 58 (No. ID414), or SEQ ID NO: 564 (No. ID418), SEQ ID NO: 335 (No. ID406), SEQ ID NO: 331 (No. ID411), SEQ ID NO: 20 (No. ID415), and SEQ ID NO: 445 (No. ID419); and (b) one or more polynucleotide sequences comprising a guide RNA, wherein the guide RNA comprises a sequence complementary to that of a targeted polynucleotide sequence; and (c) introducing into a host cell one or more polypeptide sequences of (a) and one or more polynucleotide sequences of (b) in a delivery vector; wherein the polypeptide sequence is configured to form a ribonucleoprotein complex with the guide RNA, and the ribonucleoprotein complex modifies a targeted polynucleotide sequence.
[0026] In certain preferred embodiments, the methods comprise contacting the host cell with a guide RNA, wherein the guide RNA optionally forms a ribonucleoprotein complex with the polypeptide and the guide RNA.
[0027] In various aspects, the present disclosure provides for delivery of the Cas12a described herein in a variety of viral and non-viral vectors for the Cas12a-based gene editing systems. In certain preferred embodiments, the LNP comprises: a) one or more ionizable lipids; b) one or more structural lipids; c) one or more PEGylated lipids; and d) one or more phospholipids Includes.
[0028] In certain embodiments, the LNP comprises one or more ionizable lipids selected from the group consisting of those disclosed in Table X.
[0029] Also provided herein are pharmaceutical compositions comprising site-specific modification of a target region of a host cell genome comprising a Cas V-based gene editing system described herein, the Cas V comprising one or more Cas V polypeptides; one or more cognate guide RNAs; and a LNP suitable for therapeutic administration.
[0030] In various aspects, provided herein are methods of treating a subject in need thereof, comprising administering to the subject a pharmaceutical composition described herein. In some embodiments, the subject is improved from a disease or disorder, including, but not limited to, various monogenic diseases or disorders.
[0031] In various embodiments, the present disclosure relates to the following numbered paragraphs:
[0032] 1. A genome editing system comprising: (a) a Cas V-type polypeptide or a variant thereof, or a nucleic acid sequence encoding a Cas V-type polypeptide or a variant thereof; (b) a second nucleic acid sequence encoding a guide RNA; wherein the Cas V polypeptide and the guide RNA form an RNA-protein complex; The genome editing system optionally further comprises a donor nucleic acid sequence capable of modifying a target sequence.
[0033] 2. Cas 2. The genome editing system of paragraph 1, wherein the Type V polypeptide or variant thereof is a polypeptide selected from Table S15A (SEQ ID NO: 334 (No. ID405), SEQ ID NO: 58 (No. ID414), or SEQ ID NO: 564 (No. ID418), SEQ ID NO: 335 (No. ID406), SEQ ID NO: 331 (No. ID411), SEQ ID NO: 20 (No. ID415), and SEQ ID NO: 445 (No. ID419)), or a polypeptide having at least 70%, 75%, 80%, 85%, 90%, 95%, 99%, or 100% sequence identity to a polypeptide from Table S15A (SEQ ID NO: 334 (No. ID405), SEQ ID NO: 58 (No. ID414), or SEQ ID NO: 564 (No. ID418), SEQ ID NO: 335 (No. ID406), SEQ ID NO: 331 (No. ID411), SEQ ID NO: 20 (No. ID415), and SEQ ID NO: 445 (No. ID419)).
[0034] 3. The Cas12a polypeptide is a polynucleotide sequence selected from Table S15B (SEQ ID NO: 365 (No. ID405), SEQ ID NO: 75 (No. ID414), or SEQ ID NO: 565 (No. ID418), SEQ ID NO: 366 (No. ID406), SEQ ID NO: 331 (No. ID411), SEQ ID NO: 30 (No. ID415), or SEQ ID NO: 445 (No. ID419)), or a polynucleotide sequence selected from Table S15B (SEQ ID NO: 365 (No. ID405), SEQ ID NO: 75 (No. ID414), or SEQ ID NO: 565 (No. ID418), SEQ ID NO: 366 (No. ID406), SEQ ID NO: 331 (No. ID411), SEQ ID NO: 30 (No. ID415), or SEQ ID NO: 445 (No. ID419)). 5 (No. ID414), or SEQ ID NO: 565 (No. ID418), SEQ ID NO: 366 (No. ID406), SEQ ID NO: 331 (No. ID411), SEQ ID NO: 30 (No. ID415), or SEQ ID NO: 445 (No. ID419)).
[0035] 4. The genome editing system of paragraph 1, wherein the Cas V type guide RNA is selected from any Cas V type guide sequence disclosed in Table S15C (SEQ ID NOs: 28-29, 69-71, 355-360, 542-563), or a nucleic acid molecule having at least 70%, 75%, 80%, 85%, 90%, 95%, 99%, or 100% sequence identity to a Cas V type guide sequence of Table S15C.
[0036] 5. The genome editing system of paragraph 1, wherein the Cas V polypeptide or variant thereof is operably fused to an accessory domain.
[0037] 6. The genome editing system of paragraph 5, wherein the accessory domain is a deaminase domain, a nuclease domain, a reverse transcriptase domain, an integrase domain, a recombinase domain, a transposase domain, an endonuclease domain, or an exonuclease domain.
[0038] 7. The genome editing system of paragraph 1, wherein the Cas V polypeptide or variant thereof is operably fused to a deaminase domain.
[0039] 8. The genome editing system of paragraph 1, wherein the Cas V polypeptide or variant thereof is operably fused to a reverse transcriptase domain.
[0040] 9. The genome editing system of paragraph 1, wherein the Cas V-type polypeptide or variant thereof is operably fused to a recombinase domain.
[0041] 10. The genome editing system of paragraph 1, wherein the Cas V polypeptide or variant thereof is operably fused to an integrase domain.
[0042] 11. The genome editing system of paragraph 1, wherein the Cas V-type polypeptide or variant thereof is operably fused to a transposase domain.
[0043] 12. The genome editing system of paragraph 1, wherein the Cas V polypeptide or variant thereof is engineered to have improved genome editing efficiency compared to wild-type SpCas9.
[0044] 13. The genome editing system described in paragraph 12, wherein the improved genome editing efficiency comprises at least a 2- to 5-fold increase in editing efficiency compared to wild-type SpCas9.
[0045] 14. The genome editing system of any one of the preceding paragraphs, wherein the donor nucleic acid sequence repairs the target region of the genome editing system genome that was cleaved by the RNA-protein complex.
[0046] 15. The genome editing system of any one of the preceding paragraphs, wherein the nucleic acid sequences encoding the Cas V-type polypeptide and the guide RNA are transiently expressed in the host cell genome.
[0047] 16. The genome editing system of any one of the preceding paragraphs, wherein the nucleic acid sequences encoding the Cas V-type polypeptide and the guide RNA are integrated into and expressed from the host cell genome.
[0048] 17. The genome editing system of any one of the preceding paragraphs, wherein the nucleic acid sequences encoding the Cas V polypeptide and the guide RNA are incorporated into and expressed from a plasmid.
[0049] 18. The genome editing system of any one of the preceding paragraphs, wherein the genome editing system further comprises a donor nucleic acid sequence for modifying a target region of the host cell genome.
[0050] 19. A genome editing system described in any one of the above paragraphs, wherein administering the system to a host cell results in one or more edits.
[0051] 20. The genome editing system of claim 19, wherein the one or more edits include an insertion, deletion, base change / substitution, or inversion, or a combination thereof.
[0052] 21. The genome editing system of claim 19, wherein the one or more edits comprise a modification in the nucleic acid base sequence of the target nucleic acid molecule.
[0053] 22. The genome editing system of claim 19, wherein the one or more edits comprise a whole-exon insertion, deletion, or substitution.
[0054] 23. The genome editing system of claim 19, wherein the one or more edits include a whole intron insertion, deletion, or substitution.
[0055] 24. The genome editing system of claim 19, wherein the one or more edits include a whole gene insertion, deletion, or substitution.
[0056] 25. The genome editing system of claim 19, wherein the one or more edits include edits to the sequence of a gene or to a region of a gene, such as an exon or intron.
[0057] 26. The genome editing system of any one of the preceding paragraphs, wherein the Cas V-type polypeptide recognizes a protospacer adjacent motif (PAM).
[0058] 27. The genome editing system of any one of the above paragraphs, wherein the genome editing system introduces one or more desired sequence modifications for one or more monogenic disorders or diseases.
[0059] 28. The genome editing system of any one of the preceding paragraphs, wherein the genome editing system introduces one or more desired epigenetic modifications for one or more monogenic disorders or diseases.
[0060] 29. The genome editing system of any one of the preceding paragraphs, wherein the Cas V-type polypeptide comprises one or more modifications in one or more domains selected from (a) a nuclease domain (e.g., a RuvC domain) and (b) a PAM-interaction domain.
[0061] 30. The genome editing system of any one of the above paragraphs, further comprising a delivery vector.
[0062] 31. The genome editing system described in paragraph 30, wherein the delivery vector is selected from a viral vector selected from a retroviral vector, a lentiviral vector, an adenovirus, an adeno-associated virus vector, a vaccinia virus vector, a poxvirus vector, and a herpes simplex virus vector.
[0063] 32. The genome editing system described in paragraph 30, wherein the delivery vector comprises a non-viral vector selected from cationic liposomes, lipid nanoparticles (LNPs), cationic polymers, vesicles, and gold nanoparticles.
[0064] 33. The genome editing system of any one of the preceding paragraphs, wherein the modification of the target sequence of the host cell genome comprises binding activity, cleavage activity, nickase activity, deaminase activity, reverse transcriptase activity, transcriptional activation activity, transcriptional inhibition activity, or transcriptional epigenetic activity.
[0065] 34. The genome editing system of any one of the preceding paragraphs, wherein any of the nucleic acid molecules (including any guide RNA or donor DNA) comprises one or more chemical modifications selected from 2'-O-Me, 2'-F, and 2'F-ANA at the 2'OH; 2'F-4'-Cα-OMe and 2',4'-di-Cα-OMe at the 2' and 4' carbons; phosphodiester modifications including sulfide-based phosphorothioate (PS) or acetate-based phosphonoacetate modifications; combinations of ribose and phosphodiester modifications; locked nucleic acids (LNA), bridged nucleic acids (BNA), S-constrained ethyl (cEt), and non-locked nucleic acids (UNA); modifications to generate phosphodiester bonds between the 2' and 5' carbons of adjacent RNAs (2',5'-RNA); and butane 4-carbon chain linkages between adjacent RNAs.
[0066] 35. The genome editing system of any one of the preceding paragraphs, wherein the optional guide RNA comprises one or more chemical modifications selected from 2'-O-Me, 2'-F, and 2'F-ANA at the 2'OH; 2'F-4'-Cα-OMe and 2',4'-di-Cα-OMe at the 2' and 4' carbons; phosphodiester modifications including sulfide-based phosphorothioate (PS) or acetate-based phosphonoacetate modifications; combinations of ribose and phosphodiester modifications; locked nucleic acids (LNA), bridged nucleic acids (BNA), S-constrained ethyl (cEt), and non-locked nucleic acids (UNA); modifications to generate phosphodiester bonds between the 2' and 5' carbons of adjacent RNAs (2',5'-RNA); and butane 4-carbon chain linkages between adjacent RNAs.
[0067] 36. The genome editing system of any one of the preceding paragraphs, wherein any donor or template DNA comprises one or more chemical modifications selected from 2'-O-Me, 2'-F, and 2'F-ANA at the 2'OH; 2'F-4'-Cα-OMe and 2',4'-di-Cα-OMe at the 2' and 4' carbons; phosphodiester modifications including sulfide-based phosphorothioate (PS) or acetate-based phosphonoacetate modifications; combinations of ribose and phosphodiester modifications; locked nucleic acids (LNA), bridged nucleic acids (BNA), S-constrained ethyl (cEt), and non-locked nucleic acids (UNA); modifications to generate phosphodiester bonds between the 2' and 5' carbons of adjacent nucleotides (2',5'-RNA); and butane 4-carbon chain linkages between adjacent RNAs.
[0068] 37. A method for editing the DNA of a host cell, comprising: a) 1. A Cas V-type polypeptide or a nucleic acid sequence encoding a Cas V-type polypeptide; 2. A second nucleic acid sequence encoding a guide RNA, wherein the second nucleic acid sequence and the Cas V-type polypeptide form an RNA-protein complex. wherein the genome editing system optionally further comprises a donor nucleic acid sequence capable of modifying a target sequence; and b) introducing said composition into a host cell; c) selecting for the host cells that contain the modification to the host cell genome or the donor nucleic acid sequence; and d) culturing the optionally edited host cells under conditions sufficient for growth; The method.
[0069] 38. The Cas V type polypeptide is a. operably fused to a nuclease; b. operably fused to a deaminase; c. operably fused to a reverse transcriptase; d. operably fused to a recombinase; e. operably fused to a transposase; f. operably fused to an epigenetic effector; or operably fused to any combination of g a, b, c, d, e and / or f; 38. The method according to paragraph 37.
[0070] 39. The method of paragraph 37, further comprising quantifying or characterizing said editing of said target region.
[0071] 40. The method of paragraph 37, wherein the method provides an editing efficiency of greater than 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, or 99% compared to SpCas9.
[0072] 41. The method of paragraph 37, further comprising introducing into the host cell a second donor nucleic acid sequence paired with a second guide RNA to modify a second target region of the host cell genome.
[0073] 42. The method of paragraph 37, further comprising introducing into said host cell at least two desired modified sequences for multiplexing.
[0074] 43. The method of paragraph 37, wherein the method comprises the insertion or stable integration of one or more desired modified sequences into the host cell genome.
[0075] 44. The method of paragraph 37, wherein the host cell genome comprises a chromosome or a chromosome and a plasmid.
[0076] 45. The method of paragraph 37, wherein the target region is modified by an insertion, deletion or alteration of one or more base pairs at the target region in the host cell genome.
[0077] 46. The method of paragraph 37, wherein the one or more desired modified sequences are selected from one or more sequences associated with one or more monogenic disorders or diseases.
[0078] 47. The method of paragraph 37, wherein the host cell is a primary human cell.
[0079] 48. The method of paragraph 37, wherein the step of introducing into the host cell comprises a delivery vector operably linked to the genome editing system.
[0080] 49. The method of paragraph 48, wherein the delivery vector is selected from a viral vector selected from a retroviral vector, a lentiviral vector, an adenovirus, an adeno-associated virus vector, a vaccinia virus vector, a poxvirus vector, and a herpes simplex virus vector.
[0081] 50. The method of paragraph 48, wherein the delivery vector comprises a non-viral vector selected from cationic liposomes, lipid nanoparticles (LNPs), cationic polymers, vesicles, and gold nanoparticles.
[0082] 51. The method described in paragraph 37, wherein the editing method results in improved editing efficiency and / or reduced cytotoxicity.
[0083] 52. (a) a Cas V-type domain; (b) a reverse transcriptase domain; (c) a transcriptional regulatory polypeptide; (d) a recombinase domain; (e) a transposose domain; or (f) any combination of a, b, c, d, e, or f. A gene editing construct comprising:
[0084] 53. The gene editing construct of claim 52, further comprising a donor nucleic acid sequence capable of modifying a target sequence; and
[0085] In various embodiments, the target region is modified by the insertion, deletion or alteration of one or more base pairs at the target region in the host cell genome.
[0086] In various embodiments, the one or more desired modified sequences are selected from one or more sequences associated with one or more monogenic disorders or diseases.
[0087] In certain preferred embodiments, the methods and compositions provide editing efficiencies of greater than 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, or 99% compared to SpCas9.
[0088] Related aspects provide for the use of the Cas V-based gene editing systems described herein in applications for plants, yeast, bacteria, and fungi, as well as in desired bioindustrial applications to generate valued components in such systems in a recombinant manner.
[0089] Accordingly, it is the object of the present invention not to encompass within its scope any previously known products, processes for making products, or methods for using products, and Applicant hereby reserves the right to disclose disclaimers of any previously known products, processes, or methods. It is further noted that the present invention does not intend to encompass within its scope any products, processes, or methods for making products or methods for using products that do not meet the written description and enablement requirements of the USPTO (35 U.S.C. §112, first paragraph) or the EPO (Article 83 EPC), and Applicant hereby reserves the right to disclose disclaimers of previously antedated products, processes for making products, or methods for using products. Compliance with Art. 53(c) EPC and Rules 28(b) and (c) EPC may be advantageous in the practice of the present invention. All rights to expressly disclaim any embodiment that is the subject of any granted patent(s) of Applicant in this or any other line or in any prior third-party application are expressly reserved. Nothing herein should be construed as a covenant. [Brief explanation of the drawings]
[0090] [Figure 1] A-C are schemes showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 1 sequences.
[0091] [Figure 2] A-B are schemes showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 2 sequences.
[0092] [Figure 3] A-B are schemes showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 3 sequences.
[0093] [Figure 4]A-B are schemes showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 4 sequences.
[0094] [Figure 5] A-B are schemes showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 5 sequences.
[0095] [Figure 6A] 1 is a scheme showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 6 sequences. [Figure 6B] 1 is a scheme showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 6 sequences. [Figure 6C] 1 is a scheme showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 6 sequences. [Figure 6D] 1 is a scheme showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 6 sequences.
[0096] [Figure 7] A-C are schemes showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 7 sequences.
[0097] [Figure 8] A-B are schemes showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 8 sequences.
[0098] [Figure 9A] 1 is a scheme showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 9 sequences. [Figure 9B] 1 is a scheme showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 9 sequences. [Figure 9C]1 is a scheme showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 9 sequences. [Figure 9D] 1 is a scheme showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 9 sequences. [Figure 9E] 1 is a scheme showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 9 sequences. [Figure 9F] 1 is a scheme showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 9 sequences. [Figure 9G] 1 is a scheme showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 9 sequences. [Figure 9H] 1 is a scheme showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 9 sequences. [Figure 9I] 1 is a scheme showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 9 sequences. [Figure 9J] 1 is a scheme showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 9 sequences. [Figure 9K] 1 is a scheme showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 9 sequences. [Figure 9L] 1 is a scheme showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 9 sequences. [Figure 9M] 1 is a scheme showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 9 sequences. [Figure 9N] 1 is a scheme showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 9 sequences. [Figure 9O] 1 is a scheme showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 9 sequences. [Figure 9P]1 is a scheme showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 9 sequences. [Figure 9Q] 1 is a scheme showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 9 sequences. [Figure 9R] 1 is a scheme showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 9 sequences. [Figure 9S] 1 is a scheme showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 9 sequences. [Figure 9T] 1 is a scheme showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 9 sequences. [Figure 9U] 1 is a scheme showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 9 sequences. [Figure 9V] 1 is a scheme showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 9 sequences. [Figure 9W] 1 is a scheme showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 9 sequences. [Figure 9X] 1 is a scheme showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 9 sequences. [Figure 9Y] 1 is a scheme showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 9 sequences. [Figure 9Z] 1 is a scheme showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 9 sequences. [Figure 9AA] 1 is a scheme showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 9 sequences. [Figure 9BB] 1 is a scheme showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 9 sequences. [Figure 9CC]1 is a scheme showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 9 sequences. [Figure 9DD] 1 is a scheme showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 9 sequences. [Figure 9EE] 1 is a scheme showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 9 sequences. [Figure 9FF] 1 is a scheme showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 9 sequences. [Figure 9GG] 1 is a scheme showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 9 sequences. [Figure 9HH] 1 is a scheme showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 9 sequences. [Figure 9II] 1 is a scheme showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 9 sequences. [Figure 9JJ] 1 is a scheme showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 9 sequences. [Figure 9KK] 1 is a scheme showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 9 sequences. [Figure 9LL] 1 is a scheme showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 9 sequences. [Figure 9MM] 1 is a scheme showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 9 sequences. [Figure 9NN] 1 is a scheme showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 9 sequences.
[0099] [Figure 10A] 1 is a scheme showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 10 sequences. [Figure 10B]1 is a scheme showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 10 sequences. [Figure 10C] 1 is a scheme showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 10 sequences. [Figure 10D] 1 is a scheme showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 10 sequences. [Figure 10E] 1 is a scheme showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 10 sequences. [Figure 10F] 1 is a scheme showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 10 sequences.
[0100] [Figure 11] A-C are schemes showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 11 sequences.
[0101] [Figure 12] A-B are schemes showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 12 sequences.
[0102] [Figure 13A] 1 is a scheme showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 13 sequences. [Figure 13B] 1 is a scheme showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 13 sequences. [Figure 13C] 1 is a scheme showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 13 sequences. [Figure 13D] 1 is a scheme showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 13 sequences. [Figure 13E] 1 is a scheme showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 13 sequences. [Figure 13F] 1 is a scheme showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 13 sequences.
[0103] [Figure 14A] 1 is a scheme showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 14 sequences. [Figure 14B] 1 is a scheme showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 14 sequences. [Figure 14C] 1 is a scheme showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 14 sequences. [Figure 14D] 1 is a scheme showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 14 sequences. [Figure 14E] 1 is a scheme showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 14 sequences. [Figure 14F] 1 is a scheme showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 14 sequences. [Figure 14G] 1 is a scheme showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 14 sequences. [Figure 14H] 1 is a scheme showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 14 sequences. [Figure 14I] 1 is a scheme showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 14 sequences. [Figure 14J] 1 is a scheme showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 14 sequences. [Figure 14K] 1 is a scheme showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 14 sequences. [Figure 14L]1 is a scheme showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 14 sequences. [Figure 14M] 1 is a scheme showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 14 sequences. [Figure 14N] 1 is a scheme showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 14 sequences. [Figure 14O] 1 is a scheme showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 14 sequences. [Figure 14P] 1 is a scheme showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 14 sequences. [Figure 14Q] 1 is a scheme showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 14 sequences. [Figure 14R] 1 is a scheme showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 14 sequences. [Figure 14S] 1 is a scheme showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 14 sequences. [Figure 14T] 1 is a scheme showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 14 sequences. [Figure 14U] 1 is a scheme showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 14 sequences. [Figure 14V] 1 is a scheme showing the predicted stem-loop structures of the crRNA sequences of the present disclosure corresponding to Group 14 sequences.
[0104] [Figure 15]The figure shows that the PAM sequence added to each protein in the phylogenetic tree was determined as described in Example 9. Phylogenetic tree generated using the Geneious Prime 2022.1.1 implementation of FastTree with Muscle multiple sequence alignment of selected protein sequences. PAM sequence web logo generated from PFM using the WebLogo3 web application.
[0105] [Figure 16] Cleavage products of genome-targeted DNMT1 visualized on a 2% agarose gel. Editing efficiency values for LbaCas12a and each ortholog were calculated using ImageJ software according to Example 10.
[0106] [Figure 17] Cleavage products of the genomic target RUNX1 visualized on a 2% agarose gel. Editing efficiency values for LbaCas12a and each ortholog were calculated using ImageJ software according to Example 10.
[0107] [Figure 18] Cleavage products of the genomic target SCN1A visualized on a 2% agarose gel. Editing efficiency values for LbaCas12a and each ortholog were calculated using ImageJ software according to Example 10.
[0108] [Figure 19] Cleavage products of the genomic target FANCF site 2 visualized on a 2% agarose gel. Editing efficiency values for LbaCas12a and each ortholog were calculated using ImageJ software according to Example 10.
[0109] [Figure 20] Cleavage products of the genomic target FANCF site 1 visualized on a 2% agarose gel. Editing efficiency values for LbaCas12a and each ortholog were calculated using ImageJ software according to Example 10.
[0110] [Figure 21] Comparison of Cas12a ortholog activity on different targets (n≧3). Results are calculated from T7 endonuclease assays according to Example 11. The bars for each ortholog correspond to the legend from top to bottom: DNMT1, RUNX1, SCN1A, FANCF site 1, FANCF site 2.
[0111] [Figure 22] A. Genome editing efficiency results for ID405, ID414, ID418, and LbaCas12a shown as indel frequencies at RUNX1 and SCN1A target sites as determined by deep sequencing according to Example 12. The bars for each ortholog correspond to the legend from top to bottom: RUNX1 site 1, SCN1A site 1.
[0112] B Genome editing efficiency results for ID405, ID414, ID418, LbaCas12a shown as indel frequencies at RUNX1, SCN1A, DNMT1, FANCF site 1, and FANCF site 2 (left to right for each ortholog) as determined by deep sequencing according to Example 12.
[0113] [Figure 23A-1] Top 5 most common editing outcomes observed in deep sequencing data of ID405, ID414, ID418 and LbaCas12a genomic targets in the RUNX1 gene compared to the reference sequence. [Figure 23A-2] Top 5 most common editing outcomes observed in deep sequencing data of ID405, ID414, ID418 and LbaCas12a genomic targets in the RUNX1 gene compared to the reference sequence. [Figure 23B-1]Top 5 most common editing outcomes observed in deep sequencing data of ID405, ID414, ID418 and LbaCas12a genomic targets in the SCN1A gene compared to the reference sequence. [Figure 23B-2] Top 5 most common editing outcomes observed in deep sequencing data of ID405, ID414, ID418 and LbaCas12a genomic targets in the SCN1A gene compared to the reference sequence. [Figure 23C-1] Top 5 most common editing outcomes observed in deep sequencing data of ID405, ID414, ID418 and LbaCas12a genomic targets in the DNMT1 gene when compared to the reference sequence. [Figure 23C-2] Top 5 most common editing outcomes observed in deep sequencing data of ID405, ID414, ID418 and LbaCas12a genomic targets in the DNMT1 gene when compared to the reference sequence. [Figure 23D-1] The top five most common editing outcomes observed in deep sequencing data of ID405, ID414, ID418 and LbaCas12a genomic targets in the FANCF site 1 gene when compared to the reference sequence. [Figure 23D-2] The top five most common editing outcomes observed in deep sequencing data of ID405, ID414, ID418 and LbaCas12a genomic targets in the FANCF site 1 gene when compared to the reference sequence. [Figure 23E-1] The top five most common editing outcomes observed in deep sequencing data of ID405, ID414, ID418 and LbaCas12a genomic targets in the FANCF site 2 gene when compared to the reference sequence. [Figure 23E-2] The top five most common editing outcomes observed in deep sequencing data of ID405, ID414, ID418 and LbaCas12a genomic targets in the FANCF site 2 gene when compared to the reference sequence.
[0114] [Figure 24] Genome editing efficiency results shown as indel frequencies determined by deep sequencing as described in Example 12. The bars for each ortholog correspond to the legend from top to bottom: FANCF site 2, SCN1A site 1, SCN1A site 2, DNMT1 site 2, DNMT1 site 3.
[0115] [Figure 25] The top five most common editing outcomes observed in deep sequencing data of ID428 and ID433 genomic targets showing low but observable editing when compared to the reference sequence described in Example 12.
[0116] [Figure 26-1] Comparison of endonuclease activity among SpyCas9, LbaCas12a, ID405, and ID414. Cas9 TriLink mRNA was synthesized by TriLink; Cas9 IVT, LbaCas12a, ID405, and ID414 mRNA were synthesized in-house using in vitro transcription reactions. Blue arrows indicate the cleavage products of LbaCas12a, ID405, and ID414 nucleases; black arrows indicate the cleavage products of SpyCas9 nuclease. The percentages above each gel well indicate the number of edits determined from the gel using ImageJ software. For further details, see Example 13. [Figure 26-2]Comparison of endonuclease activity among SpyCas9, LbaCas12a, ID405, and ID414. Cas9 TriLink mRNA was synthesized by TriLink; Cas9 IVT, LbaCas12a, ID405, and ID414 mRNA were synthesized in-house using in vitro transcription reactions. Blue arrows indicate the cleavage products of LbaCas12a, ID405, and ID414 nucleases; black arrows indicate the cleavage products of SpyCas9 nuclease. The percentages above each gel well indicate the number of edits determined from the gel using ImageJ software. For further details, see Example 13.
[0117] [Figure 27-1] Comparison of endonuclease activity between SpyCas9, LbaCas12a, and ID418. Cas9 TriLink mRNA was synthesized by TriLink; Cas9 IVT, LbaCas12a, and ID418 mRNA were synthesized in-house using in vitro transcription reactions. Blue arrows indicate the cleavage products of LbaCas12a and ID418 nucleases; black arrows indicate the cleavage products of SpyCas9 nuclease. The percentages above each gel well indicate the number of edits determined from the gel using ImageJ software. For further details, see Example 13. [Figure 27-2] Comparison of endonuclease activity between SpyCas9, LbaCas12a, and ID418. Cas9 TriLink mRNA was synthesized by TriLink; Cas9 IVT, LbaCas12a, and ID418 mRNA were synthesized in-house using in vitro transcription reactions. Blue arrows indicate the cleavage products of LbaCas12a and ID418 nucleases; black arrows indicate the cleavage products of SpyCas9 nuclease. The percentages above each gel well indicate the number of edits determined from the gel using ImageJ software. For further details, see Example 13.
[0118] [Figure 28A] Cleavage products of the genome-targeted PCSK9 visualized on a 2% agarose gel in the presence of various ID405 mutants. Editing efficiency values shown above the gel wells for LbaCas12a and each ID405 mutant were calculated using ImageJ software. For further details, see Example 13.
[0119] [Figure 28B] Cleavage products of genome-targeted CISH visualized on a 2% agarose gel in the presence of various ID405 mutants. Editing efficiency values shown above the gel wells for LbaCas12a and each ID405 mutant were calculated using ImageJ software. For further details, see Example 13.
[0120] [Figure 28C] Cleavage products of the genome-targeted TTR visualized on a 2% agarose gel in the presence of various ID405 mutants. Editing efficiency values shown above the gel wells for LbaCas12a and each ID405 mutant were calculated using ImageJ software. For further details, see Example 13.
[0121] [Figure 28D] Cleavage products of the genome-targeted PCSK9 visualized on a 2% agarose gel in the presence of various ID414 mutants. Editing efficiency values shown above the gel wells for LbaCas12a and each ID414 mutant were calculated using ImageJ software. For further details, see Example 13.
[0122] [Figure 28E] Cleavage products of genome-targeted CISH visualized on a 2% agarose gel in the presence of various ID414 mutants. Editing efficiency values shown above the gel wells for LbaCas12a and each ID414 mutant were calculated using ImageJ software. For further details, see Example 13.
[0123] [Figure 28F] Cleavage products of the genome-targeted TTR visualized on a 2% agarose gel in the presence of various ID414 mutants. Editing efficiency values shown above the gel wells for LbaCas12a and each ID414 mutant were calculated using ImageJ software. For further details, see Example 13.
[0124] [Figure 28G] Cleavage products of the genome-targeted BCL11a visualized on a 2% agarose gel in the presence of various ID405 mutants. Editing efficiency values for LbaCas12a and each ID405 mutant were calculated using ImageJ software.
[0125] [Figure 28H] Cleavage products of the genome-targeted HBG1 visualized on a 2% agarose gel in the presence of various ID405 mutants. Editing efficiency values for LbaCas12a and each ID405 mutant were calculated using ImageJ software.
[0126] [Figure 28I] Cleavage products of the genome-targeted BCL11a visualized on a 2% agarose gel in the presence of various ID414 mutants. Editing efficiency values for LbaCas12a and each ID414 mutant were calculated using ImageJ software.
[0127] [Figure 28J] Cleavage products of the genome-targeted HBG1 visualized on a 2% agarose gel in the presence of various ID414 mutants. Editing efficiency values for LbaCas12a and each ID414 mutant were calculated using ImageJ software.
[0128] [Figure 28K] Cleavage products of the genome-targeted PCSK9 visualized on a 2% agarose gel. Editing efficiency values for LbaCas12a and each ID418 variant were calculated using ImageJ software.
[0129] [Figure 28L] Cleavage products of genome-targeted CISH visualized on a 2% agarose gel in the presence of various ID418 mutants. Editing efficiency values for LbaCas12a and each ID418 mutant were calculated using ImageJ software.
[0130] [Figure 28M] Cleavage products of genome-targeted CISH visualized on a 2% agarose gel in the presence of various ID418 mutants. Editing efficiency values for LbaCas12a and each ID418 mutant were calculated using ImageJ software.
[0131] [Figure 28N] Cleavage products of the genome-targeted BCL11a visualized on a 2% agarose gel in the presence of various ID418 mutants. Editing efficiency values for LbaCas12a and each ID418 mutant were calculated using ImageJ software.
[0132] [Figure 28O] Cleavage products of the genome-targeted HBG1 visualized on a 2% agarose gel in the presence of various ID418 mutants. Editing efficiency values for LbaCas12a and each ID418 mutant were calculated using ImageJ software.
[0133] [Figure 29]A. Comparison of ID405 wild-type and ID405-1 mutant editing efficiencies at different targets (n≧3). Targets are BCL11a, CISH, HBG1, PCSK9, and TTR. Results were calculated from T7 endonuclease assay data. For each gene target in the cluster of bar graphs, from the leftmost bar of each cluster, the bar corresponds to ID405, ID405-1, LbaCas12a, and AsCas12a Ultra. Note that editing activity is not observed for BCL11a and TTR targets (no leftmost bar corresponds to ID405 activity). For further details, see Example 13.
[0134] B. Comparison of ID414 wild-type and ID414-1 mutant activity at different targets (n≧3). Targets are BCL11a, CISH, HBG1, PCSK9, and TTR. Results are calculated from T7 endonuclease assay data. For each gene target in the cluster of bar graphs, from the leftmost bar of each cluster, the bar corresponds to ID414, ID414-1, LbaCas12a, and AsCas12a Ultra. Note that no editing activity is observed for BCL11a and HBG1 targets (no leftmost bar corresponds to ID414 activity). For further details, see Example 13.
[0135] [Figure 30] A phylogenetic tree of the relationships between each of the Cas12a orthologous sequences presented in Table S15A versus the canonical LbCas12a sequence of SEQ ID NO: 1368 (provided in subsection Q of Section K) is shown. The phylogenetic tree was calculated using the Clustal Omega multiple sequence alignment online tool available at EMBL (European Molecular Biology Laboratory).
[0136] [Figure 31-1]Table S15A shows a sequence alignment between each of the Cas12a orthologs provided and the canonical LbCas12a sequence of SEQ ID NO: 1385. Bold underlined residues are marked with an asterisk ("*") and indicate fully conserved amino acid residues present in all aligned sequences at that alignment position. Amino acid residue positions marked with a colon (":") indicate aligned amino acid residues that are highly similar but not identically conserved. Highly similar residues are those in which the substitution between sequences is strong and has similar properties. Amino acid residue positions marked with a period (".") indicate aligned amino acid residues that are moderately similar. Highly similar residues are those in which the substitution between sequences is strong and has similar properties. Moderately similar residues are those in which the substitution between sequences is weak and has similar properties. The underlined regions are referred to as "highly conserved regions" and contain (a) at least one well-conserved residue and (b) at least one highly similar or moderately similar residue. See Section K, Subsection Q for further explanation. [Figure 31-2] Same as above [Figure 31-3] Same as above [Figure 31-4] Same as above [Figure 31-5] Same as above [Figure 31-6] Same as above [Figure 31-7] Same as above [Figure 31-8] Same as above [Figure 31-9] Same as above [Figure 31-10] Same as above [Figure 31-11] Same as above [Figure 31-12] Same as above [Figure 31-13] Same as above [Figure 31-14] Same as above [Figure 31-15] Same as above [Figure 31-16] Same as above [Figure 31-17] Same as above [Figure 31-18] Same as above [Figure 31-19]Same as above [Figure 31-20] Same as above [Figure 31-21] Same as above [Figure 31-22] Same as above [Figure 31-23] Same as above [Figure 31-24] Same as above [Figure 31-25] Same as above [Figure 31-26] Same as above [Figure 31-27] Same as above [Figure 31-28] Same as above [Figure 31-29] Same as above [Figure 31-30] Same as above [Figure 31-31] Same as above [Figure 31-32] Same as above DETAILED DESCRIPTION OF THE INVENTION
[0137] The present disclosure provides Cas V-type based gene editing systems for use in a variety of applications, including precise gene editing in cells, tissues, organs, or organisms. In various embodiments, the Cas V-type based gene editing systems include (a) a Cas V-type polypeptide and (b) a Cas V-type guide RNA capable of associating with the Cas V-type polypeptide to form a complex such that the complex localizes to and binds to a target nucleic acid sequence (e.g., a genomic or plasmid target sequence). In various embodiments, the Cas V-type polypeptide has nuclease activity that results in cleavage of at least one strand of DNA.
[0138] In exemplary embodiments, the Cas V systems and / or components thereof described herein are formulated as part of a lipid nanoparticle (LNP). In some embodiments, the lipid nanoparticle comprises an ionizable lipid, a structured lipid, a PEGylated lipid, and a phospholipid.
[0139] In various embodiments, the Cas12a polypeptide is a polypeptide selected from Table S15A (SEQ ID NO:334 (No. ID405), SEQ ID NO:58 (No. ID414), or SEQ ID NO:564 (No. ID418), SEQ ID NO:335 (No. ID406), SEQ ID NO:331 (No. ID411), SEQ ID NO:20 (No. ID415), and SEQ ID NO:445 (No. ID419)), or a polypeptide having at least 70%, 75%, 80%, 85%, 90%, 95%, 99%, or 100% sequence identity to a polypeptide from Table S15A (SEQ ID NO:334 (No. ID405), SEQ ID NO:58 (No. ID414), or SEQ ID NO:564 (No. ID418), SEQ ID NO:335 (No. ID406), SEQ ID NO:331 (No. ID411), SEQ ID NO:20 (No. ID415), and SEQ ID NO:445 (No. ID419)).
[0140] In various embodiments, the Cas V type polypeptide is encoded by a polynucleotide sequence selected from Table S15B (SEQ ID NO:365 (No. ID405), SEQ ID NO:75 (No. ID414), or SEQ ID NO:565 (No. ID418), SEQ ID NO:366 (No. ID406), SEQ ID NO:331 (No. ID411), SEQ ID NO:30 (No. ID415), or SEQ ID NO:445 (No. ID419)), or a polynucleotide having at least 70%, 75%, 80%, 85%, 90%, 95%, 99%, or 100% sequence identity to a polypeptide from Table S15B (SEQ ID NO:365 (No. ID405), SEQ ID NO:75 (No. ID414), or SEQ ID NO:565 (No. ID418), SEQ ID NO:366 (No. ID406), SEQ ID NO:331 (No. ID411), SEQ ID NO:30 (No. ID415), or SEQ ID NO:445 (No. ID419)).
[0141] In various embodiments, the Cas V type guide RNA is selected from any Cas V type guide sequence disclosed in Table S15C (SEQ ID NOs: 28-29, 69-71, 355-360, 542-563) or a nucleic acid molecule having at least 70%, 75%, 80%, 85%, 90%, 95%, 99%, or 100% sequence identity to a Cas V type guide sequence in Table S15C (SEQ ID NOs: 28-29, 69-71, 355-360, 542-563).
[0142] In various embodiments, a Cas V-type guide RNA can include (a) a portion that binds to or associates with a Cas V-type polypeptide and (b) a targeting sequence, i.e., a region containing a sequence complementary to a target nucleic acid sequence. For Cas V-type guide RNA design, just as with Cas9 guide RNAs, the target sequence is typically adjacent to a PAM sequence. However, for Cas V-types, the PAM sequence is typically TTTV (where V typically represents A, C, or G) in various embodiments. In various embodiments, the "V" in TTTV is immediately adjacent to the 5'-most base on the non-target strand of the protospacer element. For Cas9 guide RNA design, the PAM sequence is typically not included in the guide RNA design.
[0143] In various embodiments, the guide RNA for the Cas V type is relatively short, approximately 40-44 bases in length. The portion of the target sequence that base pairs with the protospacer is 20-24 bases in length, with a fixed section of approximately 20 bases that binds to the Cas V type.
[0144] In various embodiments, the nomenclature for Cas V-type guide RNAs is referred to as "crRNAs," and there is no Cas9-like "tracrRNA" component.
[0145] In other aspects, the Cas V-based gene editing system can include one or more additional accessory proteins with genome-modifying functions, including recombinases, invertases, nucleases, polymerases, ligases, deaminases, reverse transcriptases, or epigenetic modification functions. In various embodiments, the accessory proteins can be provided separately. In other embodiments, the accessory proteins can be fused to the Cas V, optionally using a linker.
[0146] In yet another embodiment, the present disclosure provides a delivery system for introducing a Cas V-based gene editing system or components thereof into a cell, tissue, organ, or organism. Depending on the format selected, the Cas V-based gene editing system and / or its individual or combined components can be delivered as a DNA molecule (e.g., encoded on one or more plasmids), an RNA molecule (e.g., a guide RNA for targeting the Cas V protein of the Cas V-based gene editing system, or a linear or circular mRNA encoding the Cas V protein or accessory protein components), a protein (e.g., a Cas V polypeptide, an accessory protein with other function (e.g., a recombinase, nuclease, polymerase, ligase, deaminase, or reverse transcriptase), or a protein-nucleic acid complex (e.g., a complex between a guide RNA and a Cas V protein or a fusion protein containing a Cas V protein).
[0147] In another aspect, the present disclosure provides nucleic acid molecules encoding a Cas V-type-based gene editing system or components thereof. In yet another aspect, the present disclosure provides vectors for transferring and / or expressing the Cas V-type-based gene editing system, e.g., under in vitro, ex vivo, and in vivo conditions. In yet another aspect, the present disclosure provides cell delivery compositions and methods, including compositions for passive and / or active transport into cells (e.g., plasmids), delivery by viral-based recombinant vectors (e.g., AAV and / or lentiviral vectors), delivery by non-viral-based systems (e.g., liposomes and LNPs), and delivery by virus-like particles of the Cas V-type-based gene editing systems described herein. Depending on the delivery system used, the Cas V-type-based gene editing systems described herein can be delivered in the form of DNA (e.g., plasmids or DNA-based viral vectors), RNA (e.g., guide RNA and mRNA delivered by LNPs), a mixture of DNA and RNA, protein (e.g., virus-like particles), and ribonucleoprotein (RNP) complexes. Any suitable combination of approaches for delivering the components of the Cas V-based gene editing systems disclosed herein may be used.
[0148] In other embodiments, the Cas V-based gene editing system can include template DNA, e.g., a single- or double-stranded donor molecule (linear or circular), containing edits that can be used by cells to repair single or double break lesions introduced by the Cas V-based gene editing system using cellular repair processes including homology-dependent repair (HDR) (e.g., in dividing cells) or non-homologous end joining (NHEJ) (in non-dividing cells).
[0149] In one embodiment, each of the components of the Cas V-based gene editing system is delivered by an all-RNA system, e.g., delivery of one or more RNA molecules (e.g., mRNA and / or guide RNA) by one or more LNPs (wherein the one or more RNA molecules form guide RNAs and / or are translated into polypeptide components (e.g., Cas V polypeptides and / or any accessory proteins)), and, if appropriate or desired, a DNA or RNA-encoding template DNA molecule (e.g., donor template).
[0150] In yet another aspect, the disclosure provides methods for genome editing by introducing the Cas V-based gene editing system described herein into a cell containing a target editing site (e.g., under in vitro, in vivo, or ex vivo conditions), thereby resulting in editing at the target site. In other aspects, the disclosure provides formulations comprising any of the aforementioned components for delivery to cells and / or tissues, including in vitro, in vivo, and ex vivo delivery, recombinant cells and / or tissues modified by the recombinant Cas V-based gene editing system and methods described herein, and methods for modifying cells by genome editing using the Cas V-based gene editing system disclosed herein.
[0151] The present disclosure also provides methods for making the Cas V-based gene editing systems described herein, their protein and nucleic acid molecular components, vectors, compositions and formulations, as well as pharmaceutical compositions and kits for modifying cells under in vitro, in vivo, and ex vivo conditions that comprise the genome editing and / or modification systems disclosed herein.
[0152] A. General definition Unless otherwise defined, all technical terms, notations, and other scientific terminology used herein are intended to have the meaning commonly understood by one of ordinary skill in the art to which this disclosure pertains. In some cases, terms having commonly understood meanings are defined herein for clarity and / or ease of reference, and the inclusion of such definitions herein should not necessarily be construed as representing a difference to that commonly understood in the art. The techniques and procedures described or referenced herein are generally well known and commonly employed using conventional methodologies by those of ordinary skill in the art, such as the widely used molecular cloning methodology described in Sambrook et al., Molecular Cloning: A Laboratory Manual, 4th ed. (2012), Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY. Where appropriate, procedures involving the use of commercially available kits and reagents are typically performed according to protocols and conditions defined by the manufacturer, unless otherwise noted.
[0153] An The articles "a" and "an" are used herein to refer to one or to more than one (i.e., to at least one) of the grammatical object of the article. By way of example, "an element" means one element or more than one element.
[0154] about "About," as used herein when referring to a measurable value, e.g., amount, time duration, etc., is meant to encompass a ±10% variation, such variation being appropriate for practicing the disclosed methods.
[0155] Biologically active As used herein, the term "biologically active" refers to a characteristic of an agent (e.g., DNA, RNA, or protein) that is active in biological systems (including in vitro and in vivo biological systems), and particularly in living animals such as mammals, including human and non-human mammals. For example, an agent is considered to be biologically active if, when administered to an organism, it has a biological effect on that organism.
[0156] bulge As used herein, the term "bulge" refers to a small region of unpaired base(s) interrupting a "stem" of base-paired nucleotides. A bulge can contain one or two single-stranded or unpaired nucleotides connected at both ends by base-paired nucleotides of the stem. A bulge can be symmetric (i.e., the two unpaired single-stranded regions have the same number of nucleotides) or asymmetric (i.e., the unpaired single-stranded regions have different or unequal numbers of nucleotides), or there is only one unpaired nucleotide on one strand. A bulge can be described as A / B (e.g., a "2 / 2 bulge" or a "1 / 0 bulge"), where A represents the number of unpaired nucleotides on the upstream strand of the stem and B represents the number of unpaired nucleotides on the downstream strand of the stem. The upstream strand of the bulge is 5' from the downstream strand of the bulge in the primary nucleotide sequence.
[0157] Cas12a or Cas12a polypeptide As used herein, "Cas12a polypeptide," "Cas12a protein," or "Cas12a nuclease" refers to a CRISPR Cas V-type polypeptide that recognizes and / or binds to RNA and is targeted to a specific DNA sequence. The Cas12a system described herein refers to a specific DNA sequence by an RNA molecule bound by a Cas12a polypeptide or Cas12a protein. The RNA molecule contains a sequence that binds to, hybridizes with, or is complementary to a target sequence within a targeted polynucleotide sequence, thereby targeting the bound polypeptide to a specific location within the targeted polynucleotide sequence (target sequence). "Cas12a" is a type of CRISPR class II type V nuclease. This specification may refer to polypeptides contemplated within the scope of this application as Cas12a polypeptides or, alternatively, as Cas V-type polypeptides, etc.
[0158] cDNA As used herein, the term "cDNA" refers to a strand of DNA copied from an RNA template, for example, by the enzyme reverse transcriptase.
[0159] Disconnect As used herein, the term "cleavage" refers to the breakage of the covalent backbone of a DNA molecule. Cleavage can be initiated by a variety of methods, including, but not limited to, enzymatic or chemical hydrolysis of phosphodiester bonds. Both single-strand and double-strand breaks are possible, and double-strand breaks can occur as a result of two distinct single-strand break events. DNA cleavage can result in the generation of either blunt ends or cohesive ends.
[0160] Cognate The term "cognate" refers to two biomolecules that normally interact or coexist in nature.
[0161] complementary As used herein, the terms "complementary" or "substantially complementary" refer to a nucleic acid (e.g., RNA, DNA) that contains a sequence of nucleotides that, under appropriate in vitro and / or in vivo conditions of temperature and solution ionic strength, can non-covalently bind, i.e., form Watson-Crick base pairs and / or G / U base pairs, to another nucleic acid and "anneal" or "hybridize" in a sequence-specific antiparallel manner (i.e., the nucleic acid specifically binds to a complementary nucleic acid). Standard Watson-Crick base pairings include adenine (A) and thymidine (T), adenine (A) and uracil (U), and guanine (G) and cytosine (C) pairings [DNA, RNA]. Also, for hybridization between two RNA molecules (e.g., dsRNA), and for hybridization between a DNA molecule and an RNA molecule (e.g., when a DNA target nucleic acid base pairs with a guide RNA): guanine (G) can also base pair with uracil (U). For example, G / U base pairing accounts for at least part of the degeneracy (i.e., redundancy) of the genetic code in relation to tRNA anticodon base pairing with codons in mRNA. Thus, in the context of the present disclosure, guanine (G) is considered complementary to both uracil (U) and adenine (A). For example, if a G / U base pair can be made at a given nucleotide position of the dsRNA duplex of a guide RNA molecule, that position is not considered non-complementary, but instead is considered complementary.
[0162] It is understood that the sequence of a polynucleotide does not need to be 100% complementary to that of its target nucleic acid to be specifically hybridizable or hybridizable. Moreover, a polynucleotide can hybridize across one or more segments such that intervening or adjacent segments are not involved in the hybridization event (e.g., bulge, loop structure, or hairpin structure). A polynucleotide can contain 60% or more, 65% or more, 70% or more, 75% or more, 80% or more, 85% or more, 90% or more, 95% or more, 98% or more, 99% or more, 99.5% or more, or 100% sequence complementarity with the target region within the target nucleic acid sequence to which it will hybridize. For example, an antisense nucleic acid in which 18 of 20 nucleotides of an antisense compound are complementary to the target region and therefore will specifically hybridize will exhibit 90% complementarity. In this example, the remaining non-complementary nucleotides may be clustered or interspersed with complementary nucleotides and need not be contiguous with each other or with complementary nucleotides. The percentage of complementarity between specific stretches of nucleic acid sequences within a nucleic acid can be determined using any convenient method. Exemplary methods include the BLAST program (Basic Local Alignment Search Tool) and PowerBLAST program (Altschul et al., J. Mol. Biol., 1990, 215, 403-410; Zhang and Madden, Genome Res., 1997, 7, 649-656), the Gap program (Wisconsin Sequence Analysis Package, Version 8 for Unix, Genetics Computer Group, University Research Park, Madison Wis.) (e.g., using default settings using the Smith and Waterman algorithm (Adv. Appl. Math., 1981, 2, 482-489)), and the like.
[0163] consists essentially of The phrase "consisting essentially of" is used herein to mean excluding anything that is not the specified active ingredient(s) of the system or is not the specified active portion(s) of the molecule.
[0164] Control Arrays The term "control sequences" is intended to include, at a minimum, all components whose presence is essential for expression, and may also include additional components whose presence is advantageous, such as leader sequences and fusion partner sequences. Suitable expression vectors include, but are not limited to, plasmids and viral vectors (e.g., derived from bacteriophage, baculovirus, tobacco mosaic virus, herpes virus, cytomegalovirus, retrovirus, vaccinia virus, adenovirus, and adeno-associated virus). Numerous vectors and expression systems are commercially available from, for example, Novagen (Madison, WI), Clontech (Palo Alto, CA), Stratagene (La Jolla, CA), and Invitrogen / Lifetechnologies (Carlsbad, CA). The present invention encompasses recombinant vectors, which may include viral vectors, bacterial vectors, protozoan vectors, DNA vectors, or recombinant forms thereof.
[0165] Degenerate variants As used herein, the phrase "degenerate variant" of a reference nucleic acid sequence encompasses nucleic acid sequences that can be translated according to the standard genetic code to provide the same amino acid sequence as that translated from the reference nucleic acid sequence. The terms "degenerate oligonucleotide" or "degenerate primer" are used to refer to oligonucleotides that are capable of hybridizing to target nucleic acid sequences that are not necessarily identical in sequence but are homologous to each other within one or more specific segments.
[0166] The engineered nucleic acid constructs of the present disclosure can be encoded by a single molecule (e.g., encoded by or present on the same plasmid or other suitable vector) or by multiple distinct molecules (e.g., multiple independently replicating vectors).
[0167] DNA-guided nucleases As used herein, a "DNA-guided nuclease" is a type of "programmable nuclease" and a specific type of "nucleic acid-guided nuclease." Examples of DNA-guided nucleases are reported in Varshney et al., DNA-guided genome editing using structure-guided endonucleases, Genome Biology, 2016, 17(1), 187 (which may be used in the context of the present disclosure and are incorporated herein by reference). As used herein, the term "DNA-guided nuclease" or "DNA-guided endonuclease" refers to a nuclease that covalently or non-covalently associates with a guide RNA, thereby forming a complex between the guide RNA and the DNA-guided nuclease. The guide RNA comprises a spacer sequence comprising a nucleotide sequence having complementarity to a strand of a target DNA sequence. Thus, a DNA-guided nuclease is indirectly guided or programmed to localize to a specific site in a DNA molecule through its association with a guide RNA that directly binds or anneals to a strand of target DNA through its complementary region via Watson-Crick base pairing.
[0168] DNA regulatory sequences As used herein, the terms "DNA regulatory sequence," "control element," and "regulatory element" may be used interchangeably herein to refer to transcriptional and translational control sequences, such as promoters, enhancers, polyadenylation signals, terminators, proteolysis signals, etc., that provide and / or control the transcription of a non-coding sequence (e.g., guide RNA) or a coding sequence and / or control the translation of mRNA into an encoded polypeptide.
[0169] domain The term "domain," as used herein, refers to a structure of a biomolecule that contributes to the known or suspected function of the biomolecule. A domain can be coextensive with a region or portion thereof; a domain can also include distinct, non-contiguous regions of a biomolecule. Examples of protein domains include, but are not limited to, Ig domains, extracellular domains, transmembrane domains, and cytoplasmic domains.
[0062] As used herein, the term "molecule" means any compound, including, but not limited to, small molecules, peptides, proteins, sugars, nucleotides, nucleic acids, lipids, etc.; such compounds can be natural or synthetic.
[0170] Donor nucleic acid By "donor nucleic acid" or "donor polynucleotide" or "donor DNA" or "HDR donor DNA" is meant a single-stranded DNA that is inserted at the site cleaved by a programmable nuclease (e.g., a CRISPR / Cas effector protein; a TALEN; a ZFN; a meganuclease) (e.g., after dsDNA cleavage, after nicking the target DNA, after double-nicking the target DNA, etc.). The donor polynucleotide has sufficient homology to the genomic sequence at the target site, e.g., to flank the target site (e.g., within about 200 bases or less of the target site, e.g., within about 190 bases or less of the target site, e.g., within about 180 bases or less of the target site, e.g., within about 170 bases or less of the target site, e.g., within about 160 bases or less of the target site, e.g., within about 150 bases or less of the target site, e.g., within about 140 bases or less of the target site, e.g., within about 130 bases or less of the target site, e.g., within about 120 bases or less of the target site, e.g., within about 110 bases or less of the target site, e.g., within about 140 bases or less of the target site, e.g., within about 150 bases or less of the target site, e.g., within about 160 bases or less of the target site, e.g., within about 170 bases or less of the target site, e.g., within about 180 bases or less of the target site, e.g., within about 190 bases or less of the target site, e.g., within about 200 bases or less of the target site, e.g., within about 210 bases or less of the target site, e.g., within about 220 bases or less of the target site, e.g., within about 230 bases or less of the target site, e.g., within about 240 bases or less of the target site, e.g., within about 250 bases or less of the target site, e.g., within about 260 bases or less of the target site, e.g., within about 270 bases or less of the target The nucleotide sequence may contain 70%, 80%, 85%, 90%, 95%, or 100% homology to nucleotide sequences within about 100 bases or less of the target site, e.g., within about 90 bases or less of the target site, e.g., within about 80 bases or less of the target site, e.g., within about 70 bases or less of the target site, e.g., within about 60 bases or less of the target site, e.g., within about 50 bases or less of the target site, e.g., within about 30 bases, within about 15 bases, within about 10 bases, within about 5 bases), or 70%, 80%, 85%, 90%, 95%, or 100% homology to nucleotide sequences immediately flanking the target site, thereby supporting homologous repair between it and the genomic sequence to which it has homology.
[0171] Effective dose "Effective amount," as used herein, means an amount that provides a therapeutic or prophylactic benefit under the conditions of administration.
[0172] Encapsulation efficiency As used herein, "encapsulation efficiency" refers to the amount of therapeutic and / or prophylactic agent that becomes part of a nanoparticle composition relative to the initial total amount of therapeutic and / or prophylactic agent used in preparing the nanoparticle composition. For example, if 97 mg of polynucleotide is encapsulated in the nanoparticle composition out of a total of 100 mg of therapeutic and / or prophylactic agent initially provided in the composition, the encapsulation efficiency can be given as 97%. As used herein, "encapsulation" can refer to complete, substantial, or partial inclusion, entrapment, surrounding, or packaging.
[0173] Code As used herein, a DNA sequence that "encodes" a particular RNA is a DNA nucleotide sequence that is transcribed into RNA. A DNA polynucleotide can encode an RNA (mRNA) that is translated into a protein (thus, both DNA and mRNA encode proteins), or a DNA polynucleotide can encode an RNA that is not translated into a protein (e.g., tRNA, rRNA, microRNA (miRNA), "non-coding" RNA (ncRNA), guide RNA, etc.).
[0174] Exosomes As used herein, the term "exosome" refers to small membrane-bound vesicles of endocytic origin. Without wishing to be bound by theory, exosomes are typically released from host / progenitor cells into the extracellular environment after fusion of multivesicular bodies with the plasma membrane. Thus, exosomes may contain components of precursor membranes in addition to engineered components. The exosome membrane is typically lamellar, consisting of a lipid bilayer with an aqueous internal nanoparticle space.
[0175] Expression vector As used herein, the term "expression vector" or "expression construct" refers to a vector that contains one or more expression control sequences, where an "expression control sequence" is a DNA sequence that controls and regulates the transcription and / or translation of another DNA sequence. Expression control sequences are sequences that control the transcription, post-transcriptional events, and translation of a nucleic acid sequence.
[0176] Expression control sequences include appropriate transcription initiation, termination, promoter, and enhancer sequences; efficient RNA processing signals, such as splicing and polyadenylation signals; sequences that stabilize cytoplasmic mRNA; sequences that improve translation efficiency (e.g., ribosome binding sites); sequences that improve protein stability; and, if desired, sequences that improve protein secretion. The nature of such control sequences varies depending on the host organism; in prokaryotes, such control sequences usually include promoters, ribosome binding sites, and transcription termination sequences.
[0177] Fusion proteins The term "fusion protein" refers to a polypeptide comprising a polypeptide or fragment coupled to a heterologous amino acid sequence, optionally via an amino acid linker. Fusion proteins are useful because they can be constructed to contain two or more desired functional elements from two or more different proteins. A fusion protein contains at least 10 contiguous amino acids from the polypeptide of interest, more preferably at least 20 or 30 amino acids, even more preferably at least 40, 50, or 60 amino acids, and even more preferably at least 75, 100, or 125 amino acids.
[0178] Fusions comprising the entire protein of the invention have particular utility. The heterologous polypeptide comprised in the fusion protein of the invention is at least 6 amino acids in length, often at least 8 amino acids in length, and usually at least 15, 20, and 25 amino acids in length. Fusions comprising larger polypeptides, such as IgG Fc regions, and even entire proteins, such as green fluorescent protein ("GFP") chromophore-containing proteins, have particular utility. Fusion proteins can be produced recombinantly by constructing a nucleic acid sequence encoding a polypeptide or fragment thereof in frame with a nucleic acid sequence encoding a different protein or peptide, and then expressing the fusion protein. Alternatively, fusion proteins can be produced chemically by crosslinking a polypeptide or fragment thereof to another protein.
[0179] guide RNA An RNA molecule that binds to a Cas12a polypeptide and targets the polypeptide to a specific location within a targeted polynucleotide sequence is referred to herein as a "guide RNA" or "guide RNA polynucleotide" (also referred to herein as a "guide RNA" or "gRNA" or "crRNA"). A guide RNA comprises two segments: a "DNA-targeting segment" and a "protein-binding segment." A "segment" refers to a segment / section / region of a molecule, e.g., a contiguous stretch of nucleotides in an RNA. As illustrative, non-limiting examples, the protein-binding segment of a guide RNA may comprise base pairs 5-20 of an RNA molecule that is 40 base pairs in length; the DNA-targeting segment may comprise base pairs 21-40 of an RNA molecule that is 40 base pairs in length. The definition of "segment" is not limited to a particular number of total base pairs, unless specifically defined otherwise in a particular context, nor is it limited to any particular number of base pairs from a given RNA molecule, nor is it limited to a particular number of separate molecules within a complex, and can include regions of an RNA molecule that are of any total length, and may or may not include regions having complementarity to other molecules.
[0180] The DNA targeting segment (or "DNA targeting sequence") comprises a nucleotide sequence complementary to a specific sequence within the targeted polynucleotide sequence (the complementary strand of the targeted polynucleotide sequence), herein designated a "protospacer-like" sequence. The protein-binding segment (or "protein-binding sequence") interacts with the site-directed modifying polypeptide. When the site-directed modifying polypeptide is a Cas12a polypeptide, site-specific cleavage of the targeted polynucleotide sequence can occur at a position determined by both (i) base-pairing complementarity between the guide RNA and the targeted polynucleotide sequence; and (ii) a short motif (termed a protospacer adjacent motif (PAM)) in the targeted polynucleotide sequence.
[0181] heterologous nucleic acid As used herein, the term "heterologous nucleic acid" refers to an entity that is genotypically distinguishable from that of the rest of the entity to which it is compared or into which it is introduced or incorporated. For example, a polynucleotide introduced into a different cell type by genetic engineering techniques is a heterologous polynucleotide (e.g., DNA or RNA) that, when expressed, may encode a heterologous polypeptide. Similarly, a cellular sequence (e.g., a gene or portion thereof) incorporated into a viral vector is a heterologous nucleotide sequence relative to the vector.
[0182] homology A protein has "homology" or is "homologous" to a second protein if the nucleic acid sequence encoding the protein has a similar sequence as the nucleic acid sequence encoding the second protein. Alternatively, a protein has homology to a second protein if the two proteins have "similar" amino acid sequences. (Thus, the term "homologous proteins" is defined to mean that two proteins have similar amino acid sequences.) As used herein, homology between two regions of amino acid sequence (especially with respect to predicted structural similarities) is interpreted as implying similarity of function.
[0183] Sequence homology for polypeptides, also referred to as percent sequence identity, is typically measured using sequence analysis software. See, e.g., Sequence Analysis Software Package of the Genetics Computer Group (GCG), University of Wisconsin Biotechnology Center, 910 University Avenue, Madison, Wis. 53705. Protein analysis software matches similar sequences using homology indices assigned to various substitutions, deletions, and other modifications, including conservative amino acid substitutions. For example, GCG contains programs such as "Gap" and "Bestfit," which can be used with default parameters to determine sequence homology or sequence identity between closely related polypeptides, such as homologous polypeptides from different species of organisms, or between a wild-type protein and its mutant protein. See, e.g., GCG version 6.1.
[0184] A preferred algorithm for comparing a particular polypeptide sequence to a database containing a large number of sequences from different organisms is the computer program BLAST (Altschul et al., J. Mol. Biol. 215:403-410 (1990); Gish and States, Nature Genet. 3:266-272 (1993); Madden et al., Meth. Enzymol. 266:131-141 (1996); Altschul et al., Nucleic Acids Res. 25:3389-3402 (1997); Zhang and Madden, Genome Res. 7:649-656 (1997)), particularly blastp or tblastn (Altschul et al., Nucleic Acids Res. 25:3389-3402 (1997)). Preferred parameters for BLASTp are expectation: 10 (default); filter: seg (default); gap opening cost: 11 (default); gap extension cost: 1 (default); maximum alignments: 100 (default); word size: 11 (default); number of descriptions: 100 (default); penalty matrix: BLOWSUM62.
[0185] The length of polypeptide sequences compared for homology will usually be at least about 16 amino acid residues, usually at least about 20 residues, more usually at least about 24 residues, typically at least about 28 residues, and preferably greater than about 35 residues. When searching a database containing sequences from many different organisms, it is preferable to compare amino acid sequences. Database searches using amino acid sequences can be performed by algorithms other than blastp known in the art. For example, polypeptide sequences are compared using the FASTA GCG version 6.1 program.
[0186] FASTA provides alignments and percent sequence identity of the regions of the best overlap between the query and search sequences. Pearson, Methods Enzymol. 183:63-98 (1990) (incorporated herein by reference). For example, percent sequence identity between amino acid sequences can be determined using FASTA with its default parameters (word size of 2 and PAM250 scoring matrix) provided in GCG version 6.1 (incorporated herein by reference).
[0187] The present disclosure provides a polypeptide having a percent identity with respect to another amino acid sequence (reference amino acid sequence), for example, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.9%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9 ... When referring to a polypeptide that is 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% identical, it means that in a polypeptide having a percent identity to a reference amino acid sequence, conserved regions of the reference amino acid sequence (e.g., conserved compared to other Cas12a, e.g., those identified herein) are conserved, and / or the polypeptide has at least one activity selected from endonuclease activity; endoribonuclease activity, or RNA-guided DNase activity, and / or the polypeptide has one or more of: a. one or more α-helix recognition lobes (REC) and nuclease lobes (NUC); b. wedge (WED), α-helix recognition lobes (REC), PAM interaction (PI), RuvC nuclease, bridging helix (BH), and NUC domains; or c.Advantageously, the polypeptide comprises one or more domains selected from the RuvC, REC, WED, BH, PI and NUC domains and / or recognizes or binds to the crRNA(s) or is bound to the crRNA(s), such as the crRNA sequences from Table S15C. Similarly, the present disclosure relates to a nucleic acid sequence or molecule that has a percent identity with respect to another nucleic acid sequence or molecule (reference nucleic acid sequence), for example, a sequence selected from SEQ ID NO: 365 (No. ID405), SEQ ID NO: 74 (No. ID414), or SEQ ID NO: 565 (No. ID418), SEQ ID NO: 366 (No. ID406), SEQ ID NO: 331 (No. ID411), SEQ ID NO: 30 (No. ID415), or SEQ ID NO: 445 (No. ID419)). When referring to a nucleic acid sequence that is 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% identical, it means that in a nucleic acid sequence having a percent identity to a reference nucleic acid sequence, the conserved region of the reference nucleic acid sequence (e.g., conserved compared to other Cas12a, e.g., those identified herein) is conserved, and / or in a polypeptide expressed from a nucleic acid sequence having a percent identity to a reference nucleic acid sequence, the polypeptide contains conserved region(s) (e.g., conserved compared to other Cas12a, e.g., those identified herein), and / or the polypeptide has at least one activity selected from endonuclease activity; endoribonuclease activity, or RNA-guided DNase activity, and / or the polypeptide comprises: a. one or more α-helical recognition lobes (REC) and nuclease lobes (NUC); b.Advantageously, the polypeptide comprises one or more domains selected from the wedge (WED), α-helix recognition lobe (REC), PAM interaction (PI), RuvC nuclease, bridging helix (BH), and NUC domains; or c. RuvC, REC, WED, BH, PI, and NUC domains and / or recognizes or binds to crRNA(s) or is bound to crRNA(s), such as the crRNA sequences from Table S15C.
[0188] Homologous recombination repair As used herein, " homology-directed repair (HDR) " refers to a specialized form of DNA repair that occurs, for example, during the repair of double-strand breaks in cells. This process requires nucleotide sequence homology and uses a "donor" molecule to template the repair of a "target" molecule (i.e., the one that has experienced a double-strand break), resulting in the transfer of genetic information from the donor to the target. If the donor polynucleotide is different from the target molecule, and some or all of the sequence of the donor polynucleotide is incorporated into the targeted polynucleotide sequence, homology-directed repair can result in the alteration (e.g., insertion, deletion, mutation) of the sequence of the target molecule.
[0189] Same As used herein, the term "identical" refers to two or more sequences or subsequences that are the same. The term "substantially identical," as used herein, refers to two or more sequences that have the same percentage of sequence units when compared and aligned for maximum correspondence over a comparison window, or designated region, as measured using a comparison algorithm or by manual alignment and visual inspection. By way of example only, two or more sequences may be "substantially identical" if the sequence units are about 60% identical, about 65% identical, about 70% identical, about 75% identical, about 80% identical, about 85% identical, about 90% identical, or about 95% identical over a designated region. Such percentages represent the "percent identity" of two or more sequences. Sequence identity can exist over a region at least about 75-100 sequence units in length, over a region about 50 sequence units in length, or, if not specified, over the entire sequence. This definition also refers to the complement of a test sequence.
[0190] Alternatively, substantial identity or similarity exists when a nucleic acid or a fragment thereof hybridizes to another nucleic acid, to a strand of another nucleic acid, or to its complementary strand under stringent hybridization conditions. "Stringent hybridization conditions" and "stringent wash conditions" in the context of nucleic acid hybridization experiments depend on many different physical parameters. Nucleic acid hybridization will be affected by conditions such as salt concentration, temperature, solvent, base composition of the hybridizing species, length of the complementary region, and the number of nucleotide base mismatches between the hybridizing nucleic acids, as will be readily understood by those skilled in the art. Those skilled in the art know how to vary these parameters to achieve a particular stringency of hybridization.
[0191] isolated "Isolated" means altered or removed from the natural state. For example, a nucleic acid or peptide naturally occurring in a living animal is not "isolated," but the same nucleic acid or peptide partially or completely separated from its coexisting materials is. An isolated nucleic acid or protein may exist in a substantially purified form or may exist in a non-native environment, such as a host cell. An "isolated nucleic acid" refers to a nucleic acid segment or fragment separated from adjacent sequences in its naturally occurring state, i.e., a DNA fragment removed from sequences normally adjacent to the fragment, i.e., sequences adjacent to the fragment in the naturally occurring genome. The term also applies to nucleic acids that have been substantially purified from other components naturally associated with the nucleic acid, i.e., RNA or DNA or proteins naturally associated with the cell. Thus, the term includes recombinant DNA or RNA that exists, for example, in a vector, an autonomously replicating plasmid or virus, or incorporated into prokaryotic or eukaryotic genomic DNA or RNA, or exists as a separate molecule independent of other sequences (i.e., as cDNA or genomic or cDNA fragments generated by PCR or restriction enzyme digestion). It also includes recombinant DNA or RNA that is part of a hybrid gene that encodes additional polypeptide sequences.
[0192] Isolated proteins The term "isolated protein" or "isolated polypeptide" refers to a protein or polypeptide that, depending on its origin or source, (1) is not associated with naturally associated components that accompany it in its native state; (2) exists in a purity not found in nature (purity can be adjusted for the presence of other cellular material) (e.g., free of other proteins from the same species); (3) is expressed by cells from a different species; or (4) does not occur in nature (e.g., it is a fragment of a naturally occurring polypeptide, or it contains an amino acid analog or derivative not found in nature, or a bond other than a standard peptide bond). Thus, a polypeptide that is chemically synthesized or synthesized in a cellular system different from the cell of natural origin would be "isolated" from its naturally associated components. A polypeptide or protein can also be substantially freed from naturally associated components by isolation, using protein purification techniques well known in the art. As defined, "isolated" does not necessarily require that the protein, polypeptide, peptide, or oligopeptide so described has been physically removed from its native environment.
[0193] Lipid nanoparticles (LNPs) As used herein, the term "lipid nanoparticle" or LNP refers to a type of lipid particle delivery system formed from small solid or semi-solid particles possessing an outer lipid layer with a hydrophilic exterior surface exposed to the non-LNP environment, an internal space that may be aqueous (vesicle-like) or non-aqueous (micelle-like), and at least one hydrophobic internal membrane space. The LNP membrane may be lamellar or non-lamellar and may contain one, two, three, four, five, or more layers. In some embodiments, LNPs may contain nucleic acids (e.g., a Cas12a editing system) within their internal space, within the internal membrane space, on their outer surface, or any combination thereof. In some embodiments, LNPs of the present disclosure contain ionizable lipids, structured lipids, PEGylated lipids (also known as PEG lipids), and phospholipids. In alternative embodiments, LNPs contain ionizable lipids, structured lipids, PEGylated lipids (also known as PEG lipids), and zwitterionic amino acid lipids.
[0194] Further discussion of liposomes can be found, for example, in Tenchov et al., “Lipid Nanoparticles—From Liposomes to mRNA Vaccine Delivery, a Landscape of Diversity and Advancement,” ACS Nano, 2021, 15, pp. 16982-17015 (the contents of which are incorporated by reference).
[0195] Linker As used herein, the term "linker" refers to a molecule that links or connects two other molecules or moieties. A linker can be an amino acid sequence, in the case of a linker connecting two fusion proteins. For example, an RNA-guided nuclease (e.g., Cas12a) can be fused to a reverse transcriptase or deaminase by an amino acid linker sequence. A linker can also be a nucleotide sequence when connecting two nucleotide sequences together; for example, in this example, the guide RNA at its 5' and / or 3' end can be linked to one or more nucleotide sequences (e.g., the RT template in the case of a prime editor guide RNA) by a nucleotide sequence linker. In other embodiments, the linker is an organic molecule, group, polymer, or chemical moiety. In some embodiments, the linker is 5 to 100 amino acids in length, e.g., 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 30-35, 35-40, 40-45, 45-50, 50-60, 60-70, 70-80, 80-90, 90-100, 100-150, or 150-200 amino acids in length. Longer and shorter linkers are also contemplated.
[0196] Liposomes As used herein, the term "liposome" refers to small vesicles containing at least one lipid bilayer membrane that contains an aqueous interior nanoparticulate space that is usually not derived from a precursor / host cell.
[0197] Micelle As used herein, the term "micelle" refers to small particles that do not have an aqueous interparticle space.
[0198] Modified derivatives "Modified derivatives" refers to polypeptides or fragments thereof that are substantially homologous in primary structural sequence, but which include, for example, in vivo or in vitro chemical and biochemical modifications or incorporate amino acids not found in the native polypeptide. Such modifications include, for example, acetylation, carboxylation, phosphorylation, glycosylation, ubiquitination, labeling with, for example, radionuclides, and various enzymatic modifications, as will be readily understood by those of skill in the art. A variety of methods for labeling polypeptides and a variety of substituents or labels useful for such purposes are well known in the art and include radioisotopes, e.g., 125 I, 32 P, 35 S, and 3 Labeling agents include ligands that bind to H, labeled antiligands (e.g., antibodies), fluorophores, chemiluminescent agents, enzymes, and antiligands that can function as specific binding members for labeled ligands. The choice of label depends on the required sensitivity, ease of conjugation with the primer, stability requirements, and available instrumentation. Methods for labeling polypeptides are well known in the art. See, for example, Ausubel et al., Current Protocols in Molecular Biology, Greene Publishing Associates (1992, and Supplements to 2002) (incorporated herein by reference).
[0199] To adjust The term "modulating," as used herein, means mediating a detectable increase or decrease in the level of a response in a subject compared to the level of the response in the subject in the absence of a treatment or compound, and / or compared to the level of the response in an otherwise identical, but untreated, subject. The term encompasses perturbing and / or affecting a native signal or response, thereby mediating a beneficial therapeutic response in a subject, preferably a human.
[0200] Mutated The term "mutated," when applied to a nucleic acid sequence, means that nucleotides in a nucleic acid sequence can be inserted, deleted, or changed compared to a reference nucleic acid sequence. A single modification can occur at a locus (point mutation), or multiple nucleotides can be inserted, deleted, or changed at a single locus. Also, one or more modifications can occur at any number of loci within a nucleic acid sequence. Nucleic acid sequences can be mutated by any method known in the art, including, but not limited to, mutagenesis techniques, such as "error-prone PCR" (a process in which PCR is performed under conditions in which the copying fidelity of the DNA polymerase is low, so that a high rate of point mutations is obtained along the entire length of the PCR product; see, e.g., Leung et al., Technique, 1:11-15 (1989) and Caldwell and Joyce, PCR Methods Applic. 2:28-33 (1992)); and "oligonucleotide-directed mutagenesis" (a process that allows for the generation of site-specific mutations in any cloned DNA segment of interest; see, e.g., Reidhaar-Olson and Sauer, Science 241:53-57 (1988)).
[0201] nanoparticles As used herein, the term "nanoparticle" refers to any particle ranging in size from 10 to 1,000 nm.
[0202] Non-homologous end joining As used herein, "non-homologous end joining (NHEJ)" refers to the repair of double-strand breaks in DNA by direct ligation of the broken ends to each other without the need for a homologous template (as opposed to homology-directed repair, which requires a homologous sequence to guide the repair). NHEJ often results in the loss (deletion) of nucleotide sequences near the site of the double-strand break.
[0203] Non-peptide analogues The term "non-peptide analog" refers to a compound that has properties similar to those of a reference polypeptide. Non-peptide compounds may also be referred to as "peptidomimetics" or "peptidomimetics." See, e.g., Jones, Amino Acid and Peptide Synthesis, Oxford University Press (1992); Jung, Combinatorial Peptide and Nonpeptide Libraries: A Handbook, John Wiley (1997); Bodanszky et al., Peptide Chemistry--A Practical Textbook, Springer Verlag (1993); Synthetic Peptides: A Users Guide, (Grant, ed., W.H. Freeman and Co., 1992); Evans et al., J. Med. Chem. 30:1229 (1987); Fauchere, J. Adv. Drug Res. 15:29 (1986); Veber and Freidinger, Trends Neurosci., 8:392-396 (1985); and the references cited therein, each of which is incorporated by reference. Such compounds are often developed with the aid of computerized molecular modeling. Peptide mimetics that are structurally similar to the useful peptides of the invention may be used to produce an equivalent effect and are envisioned as part of the present invention.
[0204] Nuclear localization sequence (NLS) As used herein, the term "nuclear localization sequence" or "NLS" refers to an amino acid sequence that facilitates the transport of a protein (e.g., an RNA-guided nuclease) into the cell nucleus, for example, by nuclear transport. Nuclear localization sequences are known in the art. For example, NLS sequences are described in International PCT Application PCT / EP2000 / 011690 to Plank et al., filed November 23, 2000, and published as WO / 2001 / 038547 on May 31, 2001 (the contents of which are incorporated herein by reference for their disclosure of exemplary nuclear localization sequences).
[0205] nucleic acid As used herein, the terms "nucleic acid" or "nucleic acid molecule" or "nucleic acid sequence" or "polynucleotide" generally refer to deoxyribonucleic acid or ribonucleic acid oligonucleotides in either single- or double-stranded form. The term may (or may not) include oligonucleotides containing known analogs of natural nucleotides. The term may also (or may not) include nucleic acid-like structures having synthetic backbones. See, e.g., Eckstein, 1991; Baserga et al., 1992; Milligan, 1993; WO 97 / 03211; WO 96 / 39154; Mata, 1997; Strauss-Soukup, 1997; and Samstag, 1996. The term encompasses both ribonucleic acid (RNA) and DNA, including cDNA, genomic DNA, synthetic, synthetically (e.g., chemically synthesized) DNA, and / or DNA (or RNA) containing nucleic acid analogs. The nucleotides adenine (A), thymine (T), guanine (G), and cytosine (C) may also include (or not include) nucleotide modifications, such as methylated and / or hydroxylated nucleotides; for example, cytosine (C) includes 5-methylcytosine and 5-hydroxymethylcytosine. Nucleic acids can be in any topological conformation. For example, nucleic acids can be single-stranded, double-stranded, triple-stranded, quadruplexed, partially double-stranded, branched, hairpinned, circular, or in a locking conformation.
[0206] Nucleic Acid-Guided Nucleases As used herein, the term "nucleic acid-guided nuclease" or "nucleic acid-guided endonuclease" refers to a nuclease (e.g., Cas12a) that covalently or non-covalently associates with a guide nucleic acid (e.g., a guide RNA or guide DNA), thereby forming a complex between the guide nucleic acid and the nucleic acid-guided nuclease. The guide nucleic acid includes a spacer sequence that includes a nucleotide sequence that has complementarity to a strand of a target DNA sequence. Thus, the nucleic acid-guided nuclease is indirectly guided or programmed to localize to a specific site in a DNA molecule through its association with a guide nucleic acid that directly binds or anneals to a strand of the target DNA through its complementary region via Watson-Crick base pairing. In some embodiments, the nucleic acid-guided nuclease will include DNA-binding activity (e.g., as in the case of CRISPR Cas12a). Most commonly, the nucleic acid-guided nuclease is programmed by association with a guide RNA molecule; in such cases, the nuclease may be referred to as an "RNA-guided nuclease." When programmed by a guide DNA, a nuclease may be referred to as a "DNA-guided nuclease." Nucleic acid-guided, RNA-guided, or DNA-guided nucleases may also be referred to as "programmable nucleases," including other classes of programmable nucleases that associate with specific DNA sequences not via guide RNA but via amino acid / nucleotide sequence recognition (e.g., zinc finger nucleases (ZFNs) and transcription activator-like effector nucleases (TALENs)). Any nuclease contemplated herein may also be engineered to remove, inactivate, or otherwise eliminate one or more nuclease activities (e.g., by introducing nuclease-inactivating mutations in the active site(s) of the nuclease, e.g., in the RuvC domain of Cas12a). A nuclease modified to remove, inactivate, or otherwise eliminate all nuclease activity may be referred to as a "death" nuclease. A death nuclease is unable to cleave either strand of a double-stranded DNA molecule.A nuclease that has been modified to remove, inactivate, or otherwise eliminate at least one nuclease activity but still retains at least one nuclease activity may be referred to as a "nickase" nuclease. Nickase nucleases cleave one strand of a double-stranded DNA molecule, but not both strands. For example, CRISPR Cas9 naturally contains two distinct nuclease activity domains: the HNH domain and the RuvC domain. The HNH domain cleaves the strand of DNA bound to the guide RNA, and the RuvC domain cleaves the protospacer strand. Nickase Cas9 can be obtained by inactivating either the HNH domain or the RuvC domain. Death Cas9 can be obtained by inactivating both the HNH domain and the RuvC domain. Other RNA-guided nucleases can be similarly converted to nickases and / or death nucleases by inactivating one or more of their existing nuclease domains.
[0207] Off-target effects "Off-target effects" refer to nonspecific gene modifications that can occur when CRISPR nucleases bind at genomic sites other than their intended targets due to mismatch tolerance. Hsu, P., Scott, D., Weinstein, J. et al. DNA targeting specificity of RNA-guided Cas9 nucleases. Nat Biotechnol 31, 827-832 (2013). https: / / doi.org / 10.1038 / nbt.2647.
[0208] operably linked As used herein, the terms "operably linked" or "under transcriptional control," when used in conjunction with describing a promoter, refer to the correct location and orientation relative to a polynucleotide (e.g., a coding sequence) to control transcription by RNA polymerase and the initiation of expression of a coding sequence, e.g., for the msr, msd, and / or ret genes. Other transcriptional control regulatory elements (e.g., enhancer sequences, transcription factor binding sites) can also be operably linked to a gene if their location relative to the gene controls or regulates expression of the gene.
[0209] PEG lipids As used herein, "PEG lipid" or "PEGylated lipid" refers to a lipid that includes a polyethylene glycol moiety.
[0210] peptide As used herein, the terms "peptide," "polypeptide," and "protein" are used interchangeably and refer to compounds comprising amino acid residues covalently linked by peptide bonds. A protein or peptide must contain at least two amino acids, and there is no limit on the maximum number of amino acids that a protein's or peptide's sequence may contain. A polypeptide includes any peptide or protein comprising two or more amino acids connected to each other by peptide bonds. As used herein, the term refers to both short chains, commonly referred to in the art as peptides, oligopeptides, and oligomers, and longer chains, commonly referred to as proteins, of which many types exist. "Polypeptide" includes, inter alia, biologically active fragments, substantially homologous polypeptides, oligopeptides, homodimers, heterodimers, polypeptide variants, modified polypeptides, derivatives, analogs, and fusion proteins. A polypeptide includes natural peptides, recombinant peptides, synthetic peptides, or combinations thereof.
[0211] promoter As used herein, the term "promoter" is art-recognized and refers to a nucleic acid molecule having a sequence that is recognized by the cellular transcription machinery and capable of initiating transcription of a downstream gene. Promoters can be constitutively active (meaning the promoter is always active in a given cellular environment) or conditionally active (meaning the promoter is active only in the presence of certain conditions). For example, a conditional promoter may be active only in the presence of a specific protein that connects proteins associated with regulatory elements in the promoter to the basal transcription machinery, or in the absence of inhibitory molecules. Within the promoter sequence, a transcription initiation site will be found, as well as a protein binding domain responsible for binding RNA polymerase. Eukaryotic promoters often, but not always, contain "TATA" boxes and "CAT" boxes. A variety of promoters, including inducible promoters, can be used to induce expression by the various vectors of the present disclosure.
[0212] Programmable nucleases As used herein, the term "programmable nuclease" refers to a polypeptide that has the property of selective localization to a specific desired nucleotide sequence (e.g., a specific gene target) in a nucleic acid molecule due to one or more targeting functions. Such targeting functions may include one or more DNA-binding domains, such as zinc finger domains characteristic of many different types of DNA-binding proteins or TALE domains characteristic of TALEN proteins. Such targeting functions may also include the ability to associate with and / or form a complex with a guide RNA, which then localizes to a specific site on DNA bearing a sequence complementary to a portion of the guide RNA (i.e., the spacer of the guide RNA). In some embodiments, a programmable nuclease can be a single protein that contains both a domain that binds to a target DNA site directly (e.g., a ZF protein) or indirectly (e.g., an RNA-guided protein) and a nuclease domain. In other embodiments, a programmable nuclease can be composed of two or more separate proteins or domains (from different proteins) that together provide the necessary functions of selective DNA binding and nuclease activity. For example, a programmable nuclease can include (a) a nuclease-inactive RNA-guided nuclease (still capable of binding guide RNA, localizing to, and binding to target DNA, but not strand cleaving or nicking) fused to (b) a nuclease protein or domain, e.g., a FokI nuclease.
[0213] Polypeptides The term "polypeptide" encompasses both naturally occurring and non-naturally occurring proteins, as well as fragments, variants, derivatives, and analogs thereof. Polypeptides can be monomeric or polymeric. Furthermore, polypeptides can contain multiple different domains, each of which has one or more distinguishable activities.
[0214] Polypeptide fragments The term "polypeptide fragment," as used herein, refers to a polypeptide that has a deletion, e.g., an amino- and / or carboxy-terminal deletion, compared to a full-length polypeptide. In preferred embodiments, a polypeptide fragment is a contiguous sequence in which the amino acid sequence of the fragment is identical at corresponding positions in a naturally occurring sequence. Fragments are typically at least 5, 6, 7, 8, 9, or 10 amino acids in length, preferably at least 12, 14, 16, or 18 amino acids in length, more preferably at least 20 amino acids in length, more preferably at least 25, 30, 35, 40, or 45 amino acids in length, even more preferably at least 50 or 60 amino acids in length, and even more preferably at least 70 amino acids in length.
[0215] Polypeptide variants A "polypeptide variant" or "mutant protein" refers to a polypeptide whose sequence contains one or more amino acid insertions, duplications, deletions, rearrangements, or substitutions compared to the amino acid sequence of a native or wild-type protein. Mutant proteins can have one or more amino acid point substitutions, in which a single amino acid is changed to another amino acid at a certain position, one or more insertions and / or deletions, in which one or more amino acids are inserted or deleted, respectively, in the sequence of the naturally occurring protein, and / or truncations of the amino acid sequence at either the amino or carboxy terminus, or both. Mutant proteins can have the same biological activity as the naturally occurring protein, but preferably have different biological activity. Mutant proteins have at least 85% overall sequence identity with their wild-type counterparts. Even more preferred are mutant proteins having at least 90% overall sequence identity with the wild-type protein. In even more preferred embodiments, mutant proteins exhibit at least 95% overall sequence identity, even more preferably 98%, even more preferably 99%, and even more preferably 99.9% overall sequence identity.
[0216] Sequence homology can be measured by any common sequence analysis algorithm, such as Gap or Bestfit.
[0217] Amino acid substitutions may include those that (1) reduce susceptibility to proteolysis, (2) reduce susceptibility to oxidation, (3) alter binding affinity for forming protein complexes, (4) alter binding affinity or enzymatic activity, and (5) confer or modify other physicochemical or functional properties of such analogs.
[0218] As used herein, the twenty conventional amino acids and their abbreviations follow conventional usage. nd ed. 1991) (incorporated by reference). "X" denotes any amino acid. Unless otherwise indicated, "B" denotes Asx (aspartic acid or asparagine) and "Z" denotes Glx (glutamic acid or glutamine). In certain concretely depicted consensus sequences, a string of adjacent Z's, e.g., "ZZZZZZZ," indicates an amino acid, which may be any amino acid or absent (as distinguished from "any amino acid"). Stereoisomers of the 20 conventional amino acids, unnatural amino acids, e.g., α-, α-disubstituted amino acids, N-alkyl amino acids, and other unconventional amino acids (e.g., D-amino acids) may also be suitable components for the polypeptides of the invention. Examples of unconventional amino acids include 4-hydroxyproline, γ-carboxyglutamate, ε-N,N,N-trimethyllysine, ε-N-acetyllysine, O-phosphoserine, N-acetylserine, N-formylmethionine, 3-methylhistidine, 5-hydroxylysine, N-methylarginine, and other similar amino acids and imino acids (e.g., 4-hydroxyproline). In the polypeptide notation used herein, the left-hand end corresponds to the amino terminus and the left-hand end corresponds to the carboxy terminus, in accordance with standard usage and convention.
[0219] Recombination The term "recombinant" refers to a biological molecule, e.g., a gene or protein, that is (1) removed from its naturally occurring environment, (2) not associated with all or a portion of a polynucleotide in which the gene is found in nature, (3) operably linked to a polynucleotide to which it is not naturally linked, or (4) not found in nature. The term "recombinant" can be used in reference to cloned DNA isolates, chemically synthesized polynucleotide analogs, or polynucleotide analogs biologically synthesized by heterologous systems, as well as proteins and / or mRNAs encoded by such nucleic acids.
[0220] As used herein, an endogenous nucleic acid sequence (or its encoded protein product) in the genome of an organism is considered "recombinant" herein when a heterologous sequence is placed adjacent to the endogenous nucleic acid sequence such that expression of the endogenous nucleic acid sequence is altered. In this context, a heterologous sequence is a sequence that is not naturally adjacent to the endogenous nucleic acid sequence, regardless of whether the heterologous sequence is itself endogenous (originating from the same host cell or its progeny) or exogenous (originating from a different host cell or its progeny). As an example, a promoter sequence can be replaced (e.g., by homologous recombination) with the native promoter of a gene in the genome of a host cell such that the gene has an altered expression pattern. The gene would be "recombinant" because it is separated from at least some of the sequences that naturally flank it.
[0221] A nucleic acid is also considered "recombinant" if it contains any modification that does not naturally occur in the corresponding nucleic acid in a genome. For example, an endogenous coding sequence is considered "recombinant" if it contains an insertion, deletion, or point mutation that has been artificially introduced, e.g., by human intervention. A "recombinant nucleic acid" also includes a nucleic acid integrated into a host cell chromosome at a heterologous site and a nucleic acid construct that exists as an episome.
[0222] Recombinant host cells The term "recombinant host cell" (or simply "host cell"), as used herein, is intended to refer to a cell into which a recombinant vector has been introduced. It should be understood that such terms are intended to refer not only to the particular subject cell but to the progeny of such a cell. Because certain modifications may occur in successive generations due to either mutation or environmental influences, such progeny may not, in fact, be identical to the parent cell, but are still included within the scope of the term "host cell" as used herein. A recombinant host cell can be an isolated cell or cell line grown in culture, or can be a cell that is present in a living tissue or organism.
[0223] Suitable methods of genetic modification, e.g., "transformation," include, for example, viral or bacteriophage infection, transfection, conjugation, protoplast fusion, lipofection, electroporation, calcium phosphate precipitation, polyethyleneimine (PEI)-mediated transfection, DEAE-dextran-mediated transfection, liposome-mediated transfection, particle gun technology, calcium phosphate precipitation, direct microinjection, nanoparticle-mediated nucleic acid delivery (see, e.g., Panyam et al., Adv Drug Deliv Rev. 2012 Sep. 13. pii:S0169-409X(12)00283-9. doi:10.1016 / j.addr.2012.09.023), etc. The choice of method of genetic modification typically depends on the type of cell being transformed and the context in which the transformation is occurring (e.g., in vitro, ex vivo, or in vivo). A general discussion of these methods can be found in Ausubel, et al., Short Protocols in Molecular Biology, 3rd ed., Wiley & Sons, 1995.
[0224] Recombinant nucleic acids "Recombinant nucleic acid" or "recombinant nucleotide" refers to a molecule constructed by joining nucleic acid molecules, which are optionally capable of autonomous replication in living cells. Recombinant nucleic acids and synthetic nucleic acids also include those molecules resulting from replication of any of the foregoing.
[0225] region The term "region," as used herein, refers to a physically contiguous portion of the primary structure of a biomolecule. In the case of a protein, a region is defined by a contiguous portion of the amino acid sequence of that protein.
[0226] RNA-guided nucleases As used herein, an "RNA-guided nuclease" is a type of "programmable nuclease" and a specific type of "nucleic acid-guided nuclease." As used herein, the term "RNA-guided nuclease" or "RNA-guided endonuclease" refers to a nuclease that covalently or non-covalently associates with a guide RNA, thereby forming a complex between the guide RNA and the RNA-guided nuclease. The guide RNA contains a spacer sequence that includes a nucleotide sequence that has complementarity to a strand of a target DNA sequence. Thus, the RNA-guided nuclease is indirectly guided or programmed to localize to a specific site in a DNA molecule through its association with a guide RNA, which directly binds or anneals to a strand of target DNA through its complementary region via Watson-Crick base pairing.
[0227] Sequence identity As used herein, the term "sequence identity" refers to the overall relatedness between polymeric molecules, e.g., between polynucleotide molecules (e.g., DNA molecules and / or RNA molecules) and / or between polypeptide molecules. Calculation of the percent identity of two polynucleotide sequences can be performed, for example, by aligning the two sequences for optimal comparison purposes (e.g., gaps can be introduced in one or both of the first and second nucleic acid sequences for optimal alignment, and non-identical sequences can be ignored for comparison purposes). For example, the length of an aligned sequence for comparison purposes is at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, or 100% of the length of the reference sequence. Nucleotides at corresponding nucleotide positions are then compared. In other examples, the length of sequence identity comparison can be over a stretch of at least about 9 nucleotides, usually at least about 20 nucleotides, more usually at least about 24 nucleotides, typically at least about 28 nucleotides, more typically at least about 32 nucleotides, and preferably at least about 36 or more nucleotides. If a position in the first sequence is occupied by the same nucleotide as the corresponding position in the second sequence, then the molecules are identical at that position. The percent identity between two sequences is a function of the number of identical positions shared by the sequences, taking into account the number of gaps, and the length of each gap that needs to be introduced for optimal alignment of the two sequences. Comparing the sequences and determining the percent identity between two sequences can be accomplished using a mathematical algorithm.For example, the percent identity between two nucleotide sequences can be determined using methods such as those described in Computational Molecular Biology, Lesk, AM, ed., Oxford University Press, New York, 1988; Biocomputing: Informatics and Genome Projects, Smith, DW, ed., Academic Press, New York, 1993; Sequence Analysis in Molecular Biology, von Heinje, G., Academic Press, 1987; Computer Analysis of Sequence Data, Part I, Griffin, AM, and Griffin, HG, eds., Humana Press, New Jersey, 1994; and Sequence Analysis Primer, Gribskov, M. and Devereux, J., eds., M Stockton Press, New York, 1991 (each of which is incorporated herein by reference). For example, percent identity between two nucleotide sequences can be determined using the algorithm of Meyers and Miller (CABIOS, 1989, 4:11-17) incorporated into the ALIGN program (version 2.0) using a PAM120 weight residue table, a gap length penalty of 12, and a gap penalty of 4. Alternatively, percent identity between two nucleotide sequences can be determined using the GAP program in the GCG software package using the NWSgapdna.CMP matrix. Commonly used methods for determining percent identity between sequences include, but are not limited to, those disclosed in Carillo, H. and Lipman, D., SIAM J Applied Math., 48:1073 (1988) (incorporated herein by reference). Techniques for determining identity are codified in publicly available computer programs.Exemplary computer software for determining homology between two sequences includes the GCG program package, Devereux, J., et al., Nucleic Acids Research, 12(1), 387 (1984), BLASTP, BLASTN, and FASTA (Altschul et al., J. Mol. Biol. 215:403-410 (1990);
[0228] These include, but are not limited to, Gish and States, Nature Genet. 3:266-272 (1993); Madden et al., Meth. Enzymol. 266:131-141 (1996); Altschul et al., Nucleic Acids Res. 25:3389-3402 (1997); Zhang and Madden, Genome Res. 7:649-656 (1997)), (Altschul et al., Nucleic Acids Res. 25:3389-3402 (1997)). For example, polynucleotide sequences can be compared using FASTA, Gap, or Bestfit, programs in the Wisconsin Package Version 10.0 (Genetics Computer Group (GCG), Madison, Wis.). FASTA provides alignments and percent sequence identity of the regions of the best overlap between the query and search sequences. Pearson, Methods Enzymol. 183:63-98 (1990), which is incorporated herein by reference in its entirety. Percent sequence identity between nucleic acid sequences can be determined, for example, using FASTA with its default parameters (word size 6 and NOPAM factor for scoring matrix) or Gap with its default parameters, as provided in GCG version 6.1.
[0229] specific binding "Specific binding" refers to the ability of two molecules to bind to each other in preference to binding to other molecules in the environment. Typically, "specific binding" is at least 2-fold, more typically at least 10-fold, and often at least 100-fold more discriminatory than chance binding in the reaction. Typically, the affinity or avidity of a specific binding reaction, as quantified by a dissociation constant, is about 10 -7 M or stronger (e.g., about 10 -8 M, 10 -9 M or even stronger).
[0230] Stem and Loop As used herein, the term "stem" refers to two or more base pairs, e.g., 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more, formed by inverted repeat sequences connected at their "tips," in which the more 5' or "upstream" strand of the stem bends to allow the more 3' or "downstream" strand to base pair with the upstream strand. The number of base pairs in the stem is the "length" of the stem. The tip of the stem is typically at least 3 nucleotides, but can be 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or more nucleotides.
[0231] Larger tips having more than five nucleotides are also referred to as "loops." Other continuous stems may be interrupted by one or more bulges as defined herein. The number of unpaired nucleotides in the bulge(s) is not included in the length of the stem. The position of the bulge closest to the tip may be represented by the number of base pairs between the bulge and the tip (e.g., the bulge is 4 bp from the tip). The positions of other bulges (if present) further away from the tip may be represented by the number of base pairs in the stem between the bulge of interest and the tip (excluding any unpaired bases in other bulges in between). As used herein, the term "loop" in a polynucleotide refers to a single-stranded extension of one or more nucleotides, e.g., 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides, in which the 5'-most and 3'-most nucleotides of the loop are linked to base-paired nucleotides in the stem, respectively.
[0232] "Stem-loop structure" or "hairpin" refers to a nucleic acid having a secondary structure that includes a region of nucleotides that is known or predicted to form a double strand (the stem portion) connected on one side by a region of primarily single-stranded nucleotides (the loop portion). Such structures are well known in the art, and these terms are used consistently with their known meaning in the art. As known in the art, a stem-loop structure does not require exact base pairing. Thus, the stem may contain one or more base mismatches. Alternatively, the base pairing may be exact, i.e., it may not contain any mismatches.
[0233] As used herein, the terms "operably linked" or "under transcriptional control," when used in conjunction with a promoter, refer to the correct location and orientation of a polynucleotide (e.g., a coding sequence) to control the initiation of transcription and expression of the coding sequence by RNA polymerase.
[0234] Stringent Hybridization Typically, "stringent hybridization" is performed at about 25°C below the thermal melting point (Tm) for a specific DNA hybrid under a particular set of conditions. "Stringent washing" is performed at a temperature about 5°C lower than the Tm for a specific DNA hybrid under a particular set of conditions. The Tm is the temperature at which 50% of the target sequence hybridizes to a perfectly matched probe. See Sambrook et al., Molecular Cloning: A Laboratory Manual, 2nd ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY (1989), page 9.51 (incorporated herein by reference). For purposes herein, "stringent conditions" are defined as solution-phase hybridization (i.e., without formamide) in 6xSSC (20xSSC contains 3.0M NaCl and 0.3M sodium citrate), 1% SDS, at 65°C for 8-12 hours, followed by two washes in 0.2xSSC, 0.1% SDS, at 65°C for 20 minutes. Those skilled in the art will recognize that hybridization at 65°C will occur at different rates depending on numerous factors, including the length and percent identity of the hybridizing sequences. Hybridization does not require that the sequence of a polynucleotide be 100% complementary to that of the target polynucleotide. Hybridization also includes one or more segments (e.g., loop or hairpin structures) such that intervening or adjacent segments are not involved in the hybridization event.
[0235] The nucleic acids (also referred to as polynucleotides) of the present invention can include both sense and antisense strands of RNA, cDNA, genomic DNA, and synthetic forms and mixed polymers of the above. They can be chemically or biochemically modified or contain non-natural or derivatized nucleotide bases, as will be readily understood by those skilled in the art. Such modifications include, for example, labeling, methylation, or substitution of one or more naturally occurring nucleotides with analogs, internucleotide modifications, such as uncharged linkages (e.g., methylphosphonates, phosphotriesters, phosphoramidates, carbamates, etc.), charged linkages (e.g., phosphorothioates, phosphorodithioates, etc.), pendant moieties (e.g., polypeptides), intercalators (e.g., acridine, psoralen, etc.), chelators, alkylators, and modified linkages (e.g., alpha-anomeric nucleic acids, etc.). Synthetic molecules that mimic polynucleotides in their ability to bind to designated sequences via hydrogen bonding and other chemical interactions are also included. Such molecules are known in the art and include, for example, those in which peptide linkages replace phosphate linkages in the backbone of the molecule. Other modifications can include, for example, analogs in which the ribose ring contains a bridging moiety or other structure, e.g., modifications found in "locked" nucleic acids.
[0236] subject As used herein, the term "subject" refers to an individual organism, e.g., an individual mammal. In some embodiments, the subject is a human. In some embodiments, the subject is a non-human mammal. In some embodiments, the subject is a non-human primate. In some embodiments, the subject is a rodent. In some embodiments, the subject is a sheep, goat, cow, cat, or dog. In some embodiments, the subject is a vertebrate, amphibian, reptile, fish, insect, fly, or nematode. In some embodiments, the subject is a research animal. In some embodiments, the subject is genetically engineered, e.g., a genetically engineered non-human subject. The subject can be of either gender and at any stage of development. The terms "individual," "subject," "host," and "patient" are used interchangeably herein.
[0237] synthetic nucleic acid "Synthetic or artificial nucleic acid" refers to a nucleic acid that is a sequence that does not occur in nature. Such a sequence is not derived from or known to exist in any living organism (e.g., based on a sequence search in an existing sequence database).
[0238] Targeted Polynucleotide Sequence As used herein, a "targeted polynucleotide sequence" refers to a DNA polynucleotide that includes a "target site" or "target sequence." The terms "target site," "target sequence," "target protospacer DNA," or "protospacer-like sequence" are used interchangeably herein to refer to a nucleic acid sequence present in a targeted polynucleotide sequence that a DNA-targeting segment of a guide RNA will recognize and / or bind to, provided that sufficient conditions for binding exist. For example, a target site (or target sequence) 5'-GAGCATATC-3' within a targeted polynucleotide sequence is targeted by (or bound by, hybridizes to, or is complementary to) the RNA sequence 5'-GAUAUGCUC-3'. Suitable DNA / RNA binding conditions include physiological conditions normally present in cells.
[0239] Other suitable DNA / RNA binding conditions (e.g., conditions in cell-free systems) are known in the art; see, e.g., Sambrook (supra). The strand of a targeted polynucleotide sequence that is complementary to and hybridizes with a guide RNA is referred to as the "complementary strand," and the strand of a targeted polynucleotide sequence that is complementary to the "complementary strand" (and therefore not complementary to the guide RNA) is referred to as the "non-complementary strand" or "non-complementary strand."
[0240] target site As used herein, a "target site" refers to a polynucleotide (e.g., DNA, e.g., genomic DNA) comprising a site or specific locus ("target site" or "target sequence") targeted by the Cas12a gene editing system disclosed herein. In the context of the Cas12a gene editing system disclosed herein, which comprises an RNA-guided nuclease, a target sequence is a sequence to which a guide sequence of a guide nucleic acid (e.g., a guide RNA) will hybridize. For example, the target site (or target sequence) 5'-GTCAATGGACC-3' within a target nucleic acid is targeted by (or bound by, hybridizes to, or is complementary to) the sequence 5'-GGTCCATTGAC-3'. Suitable hybridization conditions include physiological conditions normally present in a cell. In the case of a double-stranded target nucleic acid, the strand of the target nucleic acid that is complementary to and hybridizes with the guide RNA is referred to as the "complementary strand" or "target strand," while the strand of the target nucleic acid that is complementary to the "target strand" (and therefore not complementary to the guide RNA) is referred to as the "non-target strand" or "non-complementary strand."
[0241] therapeutic agent The term "therapeutic" as used herein means treatment and / or prophylaxis. A therapeutic effect is achieved by suppressing, reducing, ameliorating, or eradicating at least one sign or symptom of a disease or disorder state.
[0242] Therapeutically effective amount The term "therapeutically effective amount" refers to an amount of a compound of interest that elicits the biological or medical response in a tissue, system, or subject sought by a researcher, veterinarian, physician, or other clinician. The term "therapeutically effective amount" includes an amount of a compound that, when administered, is sufficient to prevent or alleviate to some extent one or more of the signs or symptoms of the disorder or disease being treated. The therapeutically effective amount will vary depending on the compound, the disease and its severity, and the age, weight, etc., of the subject being treated.
[0243] Treat or treat "Treating" a disease, as that term is used herein, means reducing the frequency or severity of at least one sign or symptom of the disease or disorder experienced by a subject.
[0244] treatment As used herein, the terms "treatment," "treat," and "treating" refer to a clinical intervention aimed at reversing, alleviating, delaying the onset of, or inhibiting the progression of a disease or disorder described herein, or one or more symptoms thereof. In some embodiments, treatment may be administered after one or more symptoms have developed and / or after a disease has been diagnosed. In other embodiments, treatment may be administered in the absence of symptoms, e.g., to prevent or delay the onset of symptoms or inhibit the onset or progression of a disease. For example, treatment may be administered to a susceptible individual (e.g., in terms of symptom history and / or in terms of genetic or other susceptibility factors) before the onset of symptoms. Treatment may also be continued after symptoms have resolved, e.g., to prevent or delay their recurrence.
[0245] Upstream and downstream As used herein, the terms "upstream" and "downstream" are relative terms that define the linear location of at least two elements located in a nucleic acid molecule (whether single-stranded or double-stranded) oriented in the 5' to 3' direction. A first element is said to be upstream of a second element in a nucleic acid molecule if the first element is located somewhere 5' relative to the second element. Conversely, a first element is downstream of a second element in a nucleic acid molecule if the first element is located somewhere 3' relative to the second element.
[0246] variant As used herein, the term "variant" should be interpreted as meaning a quality display that has a pattern that deviates from that which occurs in nature; for example, a variant retron RT is a retron RT that contains one or more changes in amino acid residues when compared to the wild-type retron RT amino acid sequence. The term "variant" encompasses homologous proteins that have at least 75%, or at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 99% identity with a reference sequence and have the same or substantially the same functional activity(ies) as the reference sequence. The term also encompasses mutants, truncated forms, or domains of the reference sequence, which exhibit the same or substantially the same functional activity(ies) as the reference sequence.
[0247] vector As used herein, the term "vector" allows or facilitates the transfer of a polynucleotide from one environment to another. It is a replicon, such as a plasmid, phage, or cosmid, into which another DNA segment can be inserted so as to bring about replication of the inserted segment. Typically, a vector is capable of replication when associated with appropriate control elements. The term "vector" can include cloning and expression vectors, as well as viral vectors and integrating vectors.
[0248] Wild type As used herein, the term "wild-type" is a term understood by those skilled in the art and means the typical form of an organism, strain, gene, protein, or characteristic as it occurs in nature, as distinguished from mutant or variant forms.
[0249] B. Chemical definition Alkyl "Alkyl" refers to a straight or branched hydrocarbon chain radical consisting solely of carbon and hydrogen atoms, saturated or unsaturated (i.e., containing one or more double and / or triple bonds), having from 1 to 30 or more carbon atoms (e.g., C-C alkyl), from 1 to 12 carbon atoms (C-C alkyl), from 1 to 8 carbon atoms (C-C alkyl), or from 1 to 6 carbon atoms (C-C alkyl), and attached to the rest of the molecule by a single bond, such as methyl, ethyl, n-propyl, 1-methylethyl (isopropyl), n-butyl, n-pentyl, 1,1 dimethylethyl (t-butyl), 3-methylhexyl, 2-methylhexyl, ethenyl, propylenyl, but-l-enyl, pent-l-enyl, penta-l,4-dienyl, ethynyl, propynyl, butynyl, pentynyl, hexynyl, and the like. Alkyl groups containing one or more units of unsaturation (one or more double and / or triple bonds) can be, for example, C2-C24, C2-C12, C2-C8, or C2-C6 groups. Unless otherwise specifically stated, alkyl groups are optionally substituted. The term "alkyl," by itself or as part of another substituent, means, unless otherwise stated, a straight or branched chain hydrocarbon having the specified number of carbon atoms (i.e., C1-6 means 1 to 6 carbon atoms), and includes straight, branched, or cyclic substituents.
[0250] Alkoxy For example, the term "alkoxy," used alone or in combination with other terms, means, unless otherwise stated, an alkyl group having the specified number of carbon atoms as defined above, attached to the remainder of the molecule through an oxygen atom, e.g., methoxy, ethoxy, 1-propoxy, 2-propoxy (isopropoxy), and higher homologs and isomers, etc. (C1-C3)alkoxy, particularly ethoxy and methoxy, are preferred.
[0251] Alkylamino As used herein, the terms "alkoxy," "alkylamino," and "alkylthio" are used in their conventional sense to refer to alkyl groups linked to the molecule via an oxygen atom, an amino group, or a sulfur atom, respectively.
[0252] Alkylene "Alkylene" or "alkylene chain" refers to a straight or branched divalent hydrocarbon chain consisting solely of carbon and hydrogen, saturated or unsaturated (i.e., containing one or more double (alkenylene) and / or triple bonds (alkynylene)), and having, for example, 1 to 30 or more carbon atoms (e.g., C1-C24 alkylene), 1 to 15 carbon atoms (C1-C15 alkylene), 1 to 12 carbon atoms (C1-C12 alkylene), 1 to 8 carbon atoms (C1-C8 alkylene), 1 to 6 carbon atoms (C1-C6 alkylene), 2 to 4 carbon atoms (C2-C4 alkylene), 1 to 2 carbon atoms (C1-C2 alkylene), e.g., methylene, ethylene, propylene, n-butylene, ethenylene, propenylene, n-butenylene, propynylene, n-butynylene, etc. An alkylene group containing one or more units of unsaturation (one or more double and / or triple bonds) can be, for example, a C2-C24, C2-C12, C2-C8, or C2-C6 group. The alkylene chain is attached to the rest of the molecule through a single or double bond and to the radical group through a single or double bond. The points of attachment of the alkylene chain to the rest of the molecule and to the radical group can be through one carbon or any two carbons within the chain. Unless stated otherwise specifically in the specification, an alkylene chain can be optionally substituted.
[0253] Aminoaryl As used herein, the term "aminoaryl" refers to an aryl moiety containing an amino moiety. Such amino moieties may include, but are not limited to, primary amines, secondary amines, tertiary amines, quaternary amines, masked amines, or protected amines. Such tertiary amines, masked amines, or protected amines may be converted to primary amine or secondary amine moieties. Amine moieties may also include amine-like moieties that have similar chemical characteristics to amine moieties, including, but not limited to, chemical reactivity.
[0254] aromatic As used herein, the term "aromatic" refers to a carbocyclic or heterocyclic ring having one or more polyunsaturated rings and having aromatic character, i.e., having (4n+2) delocalized p (π) electrons, where n is an integer.
[0255] Aryl As used herein, the term "aryl," used alone or in combination with other terms, means, unless otherwise stated, a carbocyclic aromatic system containing one or more rings (typically 1, 2, or 3 rings), where such rings may be attached to each other in a pendant manner (e.g., biphenyl) or may be fused (e.g., naphthalene). Examples include phenyl, anthracyl, and naphthyl. Phenyl and naphthyl are preferred, with phenyl being most preferred.
[0256] Cycloalkylene A "cycloalkylene" is a divalent cycloalkyl group. Unless stated otherwise specifically in the specification, a cycloalkylene group may be optionally substituted.
[0257] cycloalkyl "Cycloalkyl" or "carbocyclic ring" refers to a stable non-aromatic monocyclic or polycyclic hydrocarbon radical, consisting solely of carbon and hydrogen atoms, which may include fused or bridged ring systems, having from 3 to 15 carbon atoms, preferably from 3 to 10 carbon atoms, saturated or unsaturated, and attached to the remainder of the molecule by a single bond. Monocyclic radicals include, for example, cyclopropyl, cyclobutyl, cyclopentyl, cyclohexyl, cycloheptyl, and cyclooctyl. Polycyclic radicals include, for example, adamantyl, norvomyl, decalinyl, 7,7 dimethylbicyclo[2.2.1]heptanyl, and the like. Unless otherwise specified, cycloalkyl groups are optionally substituted.
[0258] Halo As used herein, the terms “halo” or “halogen,” by themselves or as part of another substituent, mean, unless otherwise stated, a fluorine, chlorine, bromine, or iodine atom, preferably fluorine, chlorine, or bromine, and more preferably fluorine or chlorine.
[0259] Heteroalkyl As used herein, the term "heteroalkyl," by itself or in combination with another term, means, unless otherwise stated, a stable straight- or branched-chain alkyl group consisting of the stated number of carbon atoms and one or more heteroatoms typically selected from the group consisting of O, N, Si, P, and S, wherein the nitrogen and sulfur atoms may optionally be oxidized, and the nitrogen heteroatom may be primary, secondary, tertiary, or quaternary nitrogen. The heteroatom(s) may be positioned at any position on the heteroalkyl group, including between the remainder of the heteroalkyl group and the fragment to which it is attached, and may also be attached to the most distal carbon atom in the heteroalkyl group. Examples of heteroalkyl groups include -O-CH-CH-CH, -CH-CH-CH-OH, -CH-CH-NH-CH, -CH-S-CH-CH, and -CHCH-S(=O)-CH. Up to two heteroatoms can be consecutive, such as, for example, -CH2-NH-OCH3, or -CH2-CH2-SS-CH3.
[0260] Heteroaryl As used herein, the term "heteroaryl" or "heteroaromatic" refers to an aryl group containing at least one heteroatom, typically selected from N, O, Si, P, and S; the nitrogen and sulfur atoms can be optionally oxidized, and the nitrogen atom(s) can be optionally tertiary or quaternary. Heteroaryl groups can be substituted or unsubstituted. Heteroaryl groups can be attached to the remainder of the molecule through a heteroatom. Polycyclic heteroaryls can include one or more rings that are partially saturated. Examples include tetrahydroquinoline, 2,3-dihydrobenzofuran, 1-pyrrolyl, 2-pyrrolyl, 3-pyrrolyl, 3-pyrazolyl, 2-imidazolyl, 4-imidazolyl, pyrazinyl, 2-oxazolyl, 4-oxazolyl, 2-phenyl-4-oxazolyl, 5-oxazolyl, 3-isoxazolyl, 4-isoxazolyl, 5-isoxazolyl, 2-thiazolyl, 4-thiazolyl, and the like. Examples of quinolyl include thiazolyl, 5-thiazolyl, 2-furyl, 3-furyl, 2-thienyl, 3-thienyl, 2-pyridyl, 3-pyridyl, 4-pyridyl, 2-pyrimidyl, 4-pyrimidyl, 5-benzothiazolyl, purinyl, 2-benzimidazolyl, 5-indolyl, 1-isoquinolyl, 5-isoquinolyl, 2-quinoxalinyl, 5-quinoxalinyl, 3-quinolyl, and 6-quinolyl. Examples of non-aromatic heterocycles include monocyclic groups such as aziridine, oxirane, thiirane, azetidine, oxetane, thietane, pyrrolidine, pyrroline, imidazoline, pyrazolidine, dioxolane, sulfolane, 2,3-dihydrofuran, 2,5-dihydrofuran, tetrahydrofuran, thiophane, piperidine, 1,2,3,6-tetrahydropyridine, 1,4-dihydropyridine, piperazine, morpholine, thiomorpholine, pyran, 2,3-dihydropyran, tetrahydropyran, 1,4-dioxane, 1,3-dioxane, homopiperazine, homopiperidine, 1,3-dioxepane, 4,7-dihydro-1,3-dioxepine, and hexamethylene oxide.Examples of heteroaryl groups include pyridyl, pyrazinyl, pyrimidinyl (especially 2- and 4-pyrimidinyl), pyridazinyl, thienyl, furyl, pyrrolyl (especially 2-pyrrolyl), imidazolyl, thiazolyl, oxazolyl, pyrazolyl (especially 3- and 5-pyrazolyl), isothiazolyl, 1,2,3-triazolyl, 1,2,4-triazolyl, 1,3,4-triazolyl, tetrazolyl, 1,2,3-thiadiazolyl, 1,2,3-oxadiazolyl, 1,3,4-thiadiazolyl and 1,3,4-oxadiazolyl. Examples of polycyclic heterocycles include indolyl (especially 3-, 4-, 5-, 6- and 7-indolyl), indolinyl, quinolyl, tetrahydroquinolyl, isoquinolyl (especially 1- and 5-isoquinolyl), 1,2,3,4-tetrahydroisoquinolyl, cinnolinyl, quinoxalinyl (especially 2- and 5-quinoxalinyl), quinazolinyl, phthalazinyl, 1,8-naphthyridinyl, 1,4-benzodioxanyl, coumarin, dihydrocoumarin, 1,5-naphthyridinyl, benzofuryl (especially 3-, 4-, 5-, 6- and 7-indolyl), ...indolyl), benzofuryl (especially 3-, 4-, 5-indolyl), benzofuryl (especially 3-, 4-, 5-indolyl), benzofuryl (especially 3-, 4-, 5-indolyl
[0023] Examples of heterocyclyl and heteroaryl moieties include benzothienyl (especially 3-, 4-, 5-, 6-, and 7-benzothienyl), benzoxazolyl, benzothiazolyl (especially 2-benzothiazolyl and 5-benzothiazolyl), purinyl, benzimidazolyl (especially 2-benzimidazolyl), benztriazolyl, thioxanthinyl, carbazolyl, carbolinyl, acridinyl, pyrrolidinyl, and quinolidinyl. The foregoing lists of heterocyclyl and heteroaryl moieties are intended to be representative and not limiting.
[0261] Heterocyclyl As used herein, the term "heterocyclyl" or "heterocyclic ring" refers to a stable 3- to 18-membered non-aromatic ring radical, which consists of 2 to 12 carbon atoms and 1 to 6 heteroatoms typically selected from the group consisting of N, O, Si, P, and S. Unless stated otherwise specifically in the specification, the heterocyclyl radical can be a monocyclic, bicyclic, tricyclic, or tetracyclic ring system, which can include fused or bridged ring systems; the nitrogen, carbon, or sulfur atoms in the heterocyclyl radical can be optionally oxidized; the nitrogen atom can be optionally quaternized; and the heterocyclyl radical can be partially or fully saturated. Examples of such heterocyclyl radicals include, but are not limited to, dioxolanyl, thienyl[1,3]dithianyl, decahydroisoquinolyl, imidazolinyl, imidazolidinyl, isothiazolidinyl, isoxazolidinyl, morpholinyl, octahydroindolyl, octahydroisoindolyl, 2-oxopiperazinyl, 2-oxopiperidinyl, 2-oxopyrrolidinyl, oxazolidinyl, piperidinyl, piperazinyl, 4-piperidonyl, pyrrolidinyl, pyrazolidinyl, quinuclidinyl, thiazolidinyl, tetrahydrofuryl, trithianyl, tetrahydropyranyl, thiomorpholinyl, thiamorpholinyl, 1-oxo-thiomorpholinyl, and 1,1-dioxo-thiomorpholinyl. Unless otherwise stated, heterocyclyl groups may be optionally substituted.
[0262] substituent As described herein, compounds of the present disclosure may contain "optionally substituted" moieties. Typically, the term "substituted," whether preceded by the term "optionally," means that one or more hydrogens at a specified site have been replaced with a suitable substituent. Unless otherwise indicated, an "optionally substituted" group may have a suitable substituent at each substitutable position of the group, and when multiple positions in any given structure may be substituted with multiple substituents selected from a specified group, the substituents may be the same or different at all positions. Combinations of substituents envisioned by the present disclosure preferably result in the formation of stable or chemically feasible compounds. The term "stable," as used herein, refers to compounds that are not substantially altered when subjected to conditions that permit their production, detection, and, in certain embodiments, their recovery, purification, and use for one or more of the purposes disclosed herein.
[0263] Suitable monovalent substituents on a substitutable carbon atom of an "optionally substituted" group are independently halogen; -(CH2)o-4R°; -(CH2)o-4OR°; -O(CH2)o-4R°, -O-(CH2)o-4C(O)OR°; -(CH2)o-4CH(OR°)2; -(CH2)o-4SR°; -(CH2)o-4Ph which may be substituted with R°; -(CH2)o-4O(CH2)o-1Ph which may be substituted with R°; -CH=CHPh which may be substituted with R°; -(CH2)o-4O(CH2)o-1Ph which may be substituted with R°; 2)0-1-pyridyl;-NO2;-CN;-N3;-(CH2)0-4N(R゜)2;-(CH2)0-4N(R゜)C (O)R゜;-N(R゜)C(S)R゜;-(CH2)0-4N(R゜)C(O)NR゜2;-N(R゜)C(S)NR゜2 ;-(CH2)0-4N(R゜)C(O)OR゜;-N(R゜)N(R゜)C(O)R゜;-N(R゜)N(R゜)C(O) NR゜2;-N(R゜)N(R゜)C(O)OR゜;-(CH2)0-4C(O)R゜;-C(S)R゜;-(CH2)0- 4C(O)OR゜;-(CH2)0-4C(O)SR゜;-(CH2)0-4C(O)OSiR゜3;-(CH2)0-4OC(O)R゜;-OC(O)(CH2)0-4SR゜, SC(S)SR゜;-(CH2)0-4SC(O)R゜;-(CH 2)0-4C(O)NR゜2;-C(S)NR゜2;-C(S)SR゜;-SC(S)SR゜, -(CH2)0-4OC( O)NR゜2;-C(O)N(OR゜)R゜;-C(O)C(O)R゜;-C(O)CH2C(O)R゜;-C(NOR゜) R゜;-(CH2)0-4SSR゜;-(CH2)0-4S(O)2R゜;-(CH2)0-4S(O)2OR゜;-(C H2)0-4OS(O)2R゜;-S(O)2NR゜2;-(CH2)0-4S(O)R゜;-N(R゜)S(O)2NR゜2;-N(R゜)S(O)2R゜;-N(OR゜)R゜;-C(NH)NR゜2;-P(O)2R゜;-P(O)R゜2;- OP(O)R゜2;-OP(O)(OR゜)2;SiR゜3;-(C1-4 straight chain or branched alkylene)ON(R゜)2;or -(C1-4 straight chain or branched alkylene)C(O)ON(R°), wherein each R° may be substituted as defined below and is independently hydrogen, C1-6 aliphatic, -CH2Ph, -O(CH2)0-1Ph, -CH2- (a 5-6 membered heteroaryl ring), or a 5-6 membered saturated, partially unsaturated, or aryl ring having 0-4 heteroatoms independently selected from nitrogen, oxygen, or sulfur, or, notwithstanding the above definition, two independent occurrences of R° together with their intervening atom(s) form a 3-12 membered saturated, partially unsaturated, or aryl mono- or bicyclic ring having 0-4 heteroatoms independently selected from nitrogen, oxygen, or sulfur, which may be substituted as defined below;
[0264] Suitable monovalent substituents on R° (or the ring formed by two independent occurrences of R° together with their intervening atoms) are independently halogen, —(CH2)0-2R ● ,-(Halo R ● ), -(CH2)0-2OH, -(CH2)0-2OR ● , -(CH2)0-2CH(OR ● )2;-O(HaloR ● ), -CN, -N3, -(CH2)0-2C(O)R ● , -(CH2)0-2C(O)OH, -(CH2)0-2C(O)OR ● , -(CH2)0-2SR ● , -(CH2)0-2SH, -(CH2)0-2NH2, -(CH2)0-2NHR ● , -(CH2)0-2NR ● 2, -NO2, -SiR ● 3. -OSiR ● 3. -C(O)SR ● , -(C1-4 straight or branched alkylene)C(O)OR ● , or -SSR ● (In the formula, each R ●is unsubstituted, or, if preceded by "halo", substituted only with one or more halogens, independently selected from C1-4 aliphatic, CH2Ph, O(CH2)0-1Ph, or a 5-6 membered saturated, partially unsaturated, or aryl ring having 0-4 heteroatoms independently selected from nitrogen, oxygen, or sulfur. Suitable divalent substituents on a saturated carbon atom of R° include =0 and =S.
[0265] Suitable divalent substituents on a saturated carbon atom of an "optionally substituted" group include the following: =0, =S, =NNR*2, =NNHC(O)R*,
[0266] Suitable divalent substituents attached to adjacent substitutable carbons of an "optionally substituted" group include: =NNHC(O)OR*, =NNHS(O)2R*, =NR*, =NOR*, -O(C(R*2))2-3O-, or -S(C(R*2))2-3S-, where each independent occurrence of R* is selected from hydrogen, C1-6 aliphatic which may be substituted as defined below, or an unsubstituted 5-6 membered saturated, partially unsaturated, or aryl ring having 0-4 heteroatoms independently selected from nitrogen, oxygen, or sulfur. Suitable divalent substituents attached to adjacent substitutable carbons of an "optionally substituted" group include -O(CR*2)2-3O-, where each independent occurrence of R* is selected from hydrogen, C1-6 aliphatic which may be substituted as defined below, or an unsubstituted 5-6 membered saturated, partially unsaturated, or aryl ring having 0-4 heteroatoms independently selected from nitrogen, oxygen, or sulfur.
[0267] Suitable substituents for the aliphatic groups of R* include halogen, -R ● ,-(Halo R ● ), -OH, -OR ● , -O(HaloR ● ), -CN, -C(O)OH, -C(O)OR ● , -NH2, -NHR ● , -NR ● 2, or -NO2 (wherein each R ●is unsubstituted, or, if preceded by "halo", is substituted only with one or more halogens, and is independently a C1-4 aliphatic, -CH2Ph, -O(CH2)0-1Ph, or a 5-6 membered saturated, partially unsaturated, or aryl ring having 0-4 heteroatoms independently selected from nitrogen, oxygen, or sulfur.
[0268] Suitable substituents on a substitutable nitrogen of an "optionally substituted" group include -R † , -NR † 2. -C(O)R † , -C(O)OR † , -C(O)C(O)R † , -C(O)CHC(O)R † , -S(O)2R † , -S(O)NR † 2. -C(S)NR † 2. -C(NH)NR † 2, or -N(R † )S(O)2R † (In the formula, each R † are independently hydrogen, C aliphatic, which may be substituted as defined below, unsubstituted -OPh, or an unsubstituted 5-6 membered saturated, partially unsaturated, or aryl ring having 0-4 heteroatoms independently selected from nitrogen, oxygen, or sulfur; or, notwithstanding the above definitions, R † two independent occurrences of, taken together with their intervening atom(s), form an unsubstituted 3-12 membered saturated, partially unsaturated, or aryl mono- or bicyclic ring having 0-4 heteroatoms independently selected from nitrogen, oxygen, or sulfur.
[0269] R † Suitable substituents on the aliphatic group are independently halogen, —R ● ,-(Halo R ● ), -OH, -OR ● , -O(HaloR ● ), -CN, -C(O)OH, -C(O)OR ● , -NH2, -NHR ● , -NR ● 2, or -NO2 (wherein each R ●is unsubstituted, or, if preceded by "halo", substituted only with one or more halogens, and is independently a C1-4 aliphatic, -CH2Ph, -O(CH2)0-1Ph, or a 5-6 membered saturated, partially unsaturated, or aryl ring having 0-4 heteroatoms independently selected from nitrogen, oxygen, or sulfur.
[0270] Heteroatoms such as nitrogen may have hydrogen substituents and / or any permissible substituents of organic compounds described herein that satisfy the valences of the heteroatoms. It is understood that "substituted" or "substituted" includes the implicit proviso that such substitution is in accordance with the permissible valences of the replacing atom and substituents, and that the substitution results in a stable compound, i.e., a compound that does not spontaneously undergo transformation by, for example, rearrangement, cyclization, or elimination.
[0271] In a broad aspect, the permissible substituents include acyclic and cyclic, branched and unbranched, carbocyclic and heterocyclic, aromatic and nonaromatic substituents of organic compounds. Illustrative substituents include, for example, those described herein. The permissible substituents can be one or more and the same or different for appropriate organic compounds. Heteroatoms, such as nitrogen, can have hydrogen substituents and / or any permissible substituents of organic compounds described herein that satisfy the valences of the heteroatoms.
[0272] In various embodiments, the substituents are selected from alkoxy, aryloxy, alkyl, alkenyl, alkynyl, amido, amino, aryl, arylalkyl, carbamate, carboxy, cyano, cycloalkyl, ester, ether, formyl, halogen, haloalkyl, heteroaryl, heterocyclyl, hydroxyl, ketone, nitro, phosphate, sulfide, sulfinyl, sulfonyl, sulfonic acid, sulfonamide, and thioketone, each of which is optionally substituted with one or more suitable substituents. In some embodiments, the substituents are selected from alkoxy, aryloxy, alkyl, alkenyl, alkynyl, amido, amino, aryl, arylalkyl, carbamate, carboxy, cycloalkyl, ester, ether, formyl, haloalkyl, heteroaryl, heterocyclyl, ketone, phosphate, sulfide, sulfinyl, sulfonyl, sulfonic acid, sulfonamide, and thioketone, each of which can be further substituted with one or more suitable substituents.
[0273] Examples of substituents include halogen, azide, alkyl, aralkyl, alkenyl, alkynyl, cycloalkyl, hydroxyl, alkoxyl, amino, nitro, sulfhydryl, imino, amido, phosphonate, phosphinate, carbonyl, carboxyl, silyl, ether, alkylthio, sulfonyl, sulfonamide, ketone, aldehyde, thioketone, ester, heterocyclyl, -CN, aryl, aryloxy, perhaloalkoxy, aralkoxy, heteroaryl, heteroaryloxy, heteroarylalkyl, heteroaralkoxy, azide, aryloxy, aryloxy, aryloxy, aryloxy, heteroarylalkyl, heteroaralkoxy, aryl ... These include, but are not limited to, alkylthio, oxo, acylalkyl, carboxyester, carboxamido, acyloxy, aminoalkyl, alkylaminoaryl, alkylaryl, alkylaminoalkyl, alkoxyaryl, arylamino, aralkylamino, alkylsulfonyl, carboxamidoalkylaryl, carboxamidoaryl, hydroxyalkyl, haloalkyl, alkylaminoalkylcarboxy, aminocarboxamidoalkyl, cyano, alkoxyalkyl, perhaloalkyl, arylalkyloxyalkyl, etc. In some embodiments, the substituents are selected from cyano, halogen, hydroxyl, and nitro.
[0274] Throughout this disclosure, chemical substituents described in Markush configurations are represented by variables. When a variable is given multiple definitions that apply to different Markush formulas in different sections of this disclosure, it should be understood that each definition should apply only to the applicable formula in the appropriate section of this disclosure.
[0275] Abbreviation As used herein, the following abbreviations and initialisms have the indicated meanings: [Table 6]
[0276] The details of one or more embodiments of the present disclosure are set forth in the accompanying description below. Although any materials and methods similar or equivalent to those described herein can be used in the practice or testing of the present disclosure, the preferred materials and methods are now described. Other features, objects, and advantages of the present disclosure will become apparent from the specification. As used herein, the singular forms "a," "an," and "the" include the plural forms unless the context clearly indicates otherwise. Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. In case of conflict, the present specification controls.
[0277] C. Cas12a (or CasV) sequence The present disclosure provides Cas12a (or Cas V-type) polypeptides and nucleic acid molecules encoding same for use in the Cas12a-based gene editing systems described herein for use in a variety of applications, including precise gene editing in cells, tissues, organs, or organisms. In various embodiments, the Cas12a-based gene editing systems include (a) a Cas12a (or Cas V-type) polypeptide (or a nucleic acid molecule encoding a Cas12a (or Cas V-type) polypeptide) and (b) a Cas12a (or Cas V-type) guide RNA that is capable of associating with the Cas12a (or Cas V-type) polypeptide to form a complex, thereby localizing the complex to and binding to a target nucleic acid sequence (e.g., a genomic or plasmid target sequence). In various embodiments, the Cas12a (or Cas V-type) polypeptide has nuclease activity that results in cleavage of both strands of DNA.
[0278] As reviewed in B. Paul, Biomedical Journal, Vol. 43, No. 1, February 2020, pages 8-17, CRISPR-Cas systems are classified into two classes (Class 1 and Class 2), which are further subdivided into six types (Types I to VI). Class 1 (Types I, III, and IV) systems use multiple Cas proteins in their CRISPR ribonucleoprotein effector nucleases, while Class 2 systems (Types II, V, and VI) use a single Cas protein. Class 1 CRISPR-Cas systems are most commonly found in bacteria and archaea and account for approximately 90% of all identified CRISPR-Cas loci. Class 2 CRISPR-Cas systems, comprising the remaining approximately 10%, are widespread in bacteria and assemble a ribonucleoprotein complex consisting of CRISPR RNA (crRNA) and a Cas protein. The crRNA contains the information for targeting specific DNA sequences. These multidomain effector proteins achieve interference through complementarity between the crRNA and the target sequence after recognition of the PAM (protospacer adjacent motif) sequence adjacent to the target DNA. These ribonucleoprotein complexes are designed for precise genome editing by delivering a crRNA with a redesigned guide sequence complementary to the targeted DNA sequence. The most extensively characterized CRISPR-Cas system is the type II subtype II-A found in Streptococcus pyogenes (Sp), which uses the protein SpCas9. Cas9 was the first Cas protein engineered for use in gene editing. Class 2V is further divided into four subtypes (VA, VB, VC, and VU). Currently, VC and VU remain largely uncharacterized, and no structural information is available about these systems. VA encodes the protein Cas12a (also known as Cpf1), and recently, several high-resolution structures of Cas12a have provided insight into its mechanism of action.
[0279] Type II (e.g., Cas9) and type V (e.g., Cas12a) CRISPR-Cas systems possess a characteristic Ruv-C-like nuclease domain, which has been shown to be related to the TnpB protein encoded by IS605 family transposons. Crystallographic and cryo-EM data reveal that Cas12a adopts a bilobed structure formed by the REC and Nuc lobes. The REC lobe contains the REC1 and REC2 domains, while the Nuc lobe contains the RuvC, PAM-interacting (PI), and WED domains, as well as a bridging helix (BH). The RuvC endonuclease domain of this effector protein is composed of three noncontiguous parts (RuvC I-III). The RNase site for its own crRNA is located in the WED-III subdomain, and the DNase site is located at the interface between the RuvC and Nuc domains. These structural studies also indicate that only the 5' repeat region of the crRNA is involved in the assembly of the binary complex. The 19 / 20 nt repeat region forms a pseudoknot structure through intramolecular base pairing. The crRNA is stabilized through interactions with the WED, RuvC, and REC2 domains of the endonuclease and two hydrated Mg2+ ions. This binary interference complex is then responsible for recognizing and degrading foreign DNA.
[0280] PAM recognition is a critical first step in identifying promising DNA molecules for degradation, as it allows the CRISPR-Cas system to distinguish their own genomic DNA from invading nucleic acids. Cas12a employs a multi-step quality control mechanism to ensure accurate and precise recognition of the target spacer sequence. The WED II-III, REC1, and PAM-interacting domains are responsible for PAM recognition and initiation of hybridization between the DNA target and crRNA. After dsDNA recognition by the WED and REC1 domains, a conserved loop-lysine helix-loop (LKL) region in the PI domain, containing three conserved lysines (K667, K671, and K677 in FnCas12a), inserts a helix into the PAM duplex with assistance from two conserved prolines in the LKL region. Structural studies have shown that the helix inserts at a 45° angle relative to the long axis of dsDNA, facilitating the unwinding of the helical DNA. The critical positioning of three conserved lysines on dsDNA initiates the uncoupling of Watson-Crick interactions between dsDNA base pairs after PAM. Target dsDNA dissociation allows hybridization of the crRNA with the PAM-containing strand, i.e., the "target strand (TS)," while the unbound DNA strand, i.e., the non-target strand (NTS), is guided toward the DNase site by the PAM-interacting domain. Cas12a has been shown to efficiently target spacer sequences following 5' T-rich PAM sequences. The PAM for LbCas12a and AsCas12a has the sequence 5'-TTTN-3', while for FnCas12a it has the sequence 5'-TTN-3', located upstream of the 5' end of the non-target strand. In addition to the standard 5'-TTTN-3' PAM, Cas12a has also been shown to exhibit relaxed PAM recognition for suboptimal C-containing PAM sequences by forming modified interactions with the targeted DNA duplex.
[0281] Thus, Cas12a is another RNA-guided nuclease of the Class II CRISPR / Cas system that shares similarities with Cas9 and can be used in a similar manner. Unlike Cas9, Cas12a does not require tracrRNA and relies solely on crRNA for its guide RNA, which provides the advantage that shorter guide RNAs than Cas9 can be used with Cas12a for targeting. Cas12a can cleave either DNA or RNA. In contrast to the G-rich PAM sites recognized by Cas9, the PAM sites recognized by Cas12a have the sequence 5'-YTN-3' (where "Y" is pyrimidine and "N" is any nucleobase) or 5'-TTN-3'. Cas12a cleavage of DNA generates double-strand breaks with sticky ends with 4- or 5-nucleotide overhangs. For further discussion of Cas12a, see, e.g., Ledford et al. (2015) Nature. 526(7571):17-17, Zetsche et al. (2015) Cell. 163(3):759-771, Murovec et al. (2017) Plant Biotechnol. J. 15(8):917-926, Zhang et al. (2017) Front. Plant Sci. 8:177, Fernandes et al. (2016) Postepy Biochem. 62(3):315-326 (incorporated herein by reference).
[0282] Any Cas12a (or Cas V type) polypeptide or variant thereof may be used in the present disclosure, including those described in the tables herein and provided in the accompanying sequence listing.
[0283] In various embodiments, Cas12a (or Cas A Type V) polypeptide is a polypeptide selected from Table S15A (SEQ ID NO:334 (No. ID405), SEQ ID NO:58 (No. ID414), or SEQ ID NO:564 (No. ID418), SEQ ID NO:335 (No. ID406), SEQ ID NO:331 (No. ID411), SEQ ID NO:20 (No. ID415), and SEQ ID NO:445 (No. ID419)), or a polypeptide having at least 70%, 75%, 80%, 85%, 90%, 95%, 99%, or 100% sequence identity to a polypeptide from Table S15A (SEQ ID NO:334 (No. ID405), SEQ ID NO:58 (No. ID414), or SEQ ID NO:564 (No. ID418), SEQ ID NO:335 (No. ID406), SEQ ID NO:331 (No. ID411), SEQ ID NO:20 (No. ID415), and SEQ ID NO:445 (No. ID419)).
[0284] In various embodiments, Cas12a (or Cas The Type V) polypeptide is encoded by a polynucleotide sequence selected from Table S15B (SEQ ID NO:365 (No. ID405), SEQ ID NO:75 (No. ID414), or SEQ ID NO:565 (No. ID418), SEQ ID NO:366 (No. ID406), SEQ ID NO:331 (No. ID411), SEQ ID NO:30 (No. ID415), or SEQ ID NO:445 (No. ID419)), or a polynucleotide having at least 70%, 75%, 80%, 85%, 90%, 95%, 99%, or 100% sequence identity to a polypeptide from Table S15B (SEQ ID NO:365 (No. ID405), SEQ ID NO:75 (No. ID414), or SEQ ID NO:565 (No. ID418), SEQ ID NO:366 (No. ID406), SEQ ID NO:331 (No. ID411), SEQ ID NO:30 (No. ID415), or SEQ ID NO:445 (No. ID419)).
[0285] Any Cas12a (or Cas V) polypeptide can be utilized with the compositions described herein. The Cas12a editing system contemplated herein is not meant to be limiting in any way. The Cas12a editing system disclosed herein can include standard or naturally occurring Cas12a, or any orthologous Cas12a protein, or any variant Cas12a protein known or that can be created or evolved through directed evolution or otherwise mutagenic processes (including any naturally occurring variant, mutant, or otherwise engineered version of Cas12a). In various embodiments, Cas12a or Cas12a variants can only have nickase activity, i.e., strand cleavage of target DNA sequences. In other embodiments, Cas12a or Cas12a variants have inactive nuclease activity, i.e., are "dead" Cas12a proteins. Other variant Cas12a proteins that can be used are those that have a smaller molecular weight than standard Cas12a (e.g., for easier delivery) or have modified amino acid sequences or substitutions.
[0286] In various aspects, the invention provides one or more modifications of a Cas12a (or Cas V) polypeptide, including, for example, mutations and modifications of Cas12a to increase sufficiency and / or efficiency. In some embodiments, one or more domains of Cas12a, such as the RuvC, REC, WED, BH, PI, and NUC domains, are modified. In certain preferred embodiments, the modifications provide editing efficiencies of greater than 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, or 99% compared to SpCas9. Even more preferably, the methods and compositions provide improved transduction efficiency and / or reduced cytotoxicity.
[0287] The Cas12a (or Cas V-type) gene editing systems and therapeutic agents described herein can include one or more nucleic acid components (e.g., guide RNAs or coding RNAs that encode components of the Cas12a system) that can be codon-optimized.
[0288] For example, a nucleotide sequence encoding a nucleic acid base editing system of the present disclosure (e.g., as part of an RNA payload) may be codon-optimized. Codon optimization methods are known in the art. For example, any one or more protein-coding sequences among those provided herein may be codon-optimized. Codon optimization, in some embodiments, can be used to ensure proper folding; bias GC content to increase mRNA stability or reduce secondary structure; minimize tandem repeat codons or base runs that may impair gene assembly or expression; customize transcriptional and translational control regions; insert or remove protein trafficking sequences; remove / add post-translational modification sites (e.g., glycosylation sites) in the encoded protein; add, remove, or shuffle protein domains; insert or delete restriction sites; modify ribosome binding sites and mRNA degradation sites; adjust translation rates to allow various domains of a protein to fold properly; or match codon frequencies in the target and host organisms to reduce or eliminate problematic secondary structures within the polynucleotide. Codon optimization tools, algorithms, and services are known in the art (non-limiting examples include services from GeneArt (Life technologies), DNA2.0 (Menlo Park, Calif.), and / or proprietary methods). In some embodiments, the protein-coding sequence is optimized using an optimization algorithm. In some embodiments, the codon-optimized sequence shares less than 95% sequence identity with a naturally occurring or wild-type sequence (e.g., a naturally occurring or wild-type mRNA sequence encoding a nucleobase editing enzyme). In some embodiments, the codon-optimized sequence shares less than 90% sequence identity with a naturally occurring or wild-type sequence (e.g., a naturally occurring or wild-type mRNA sequence encoding a nucleobase editing enzyme). In some embodiments, the codon-optimized sequence shares less than 85% sequence identity with a naturally occurring or wild-type sequence (e.g., a naturally occurring or wild-type mRNA sequence encoding a nucleobase editing enzyme).In some embodiments, the codon-optimized sequence shares less than 80% sequence identity with a naturally occurring or wild-type sequence (e.g., a naturally occurring or wild-type mRNA sequence encoding a nucleobase editing enzyme). In some embodiments, the codon-optimized sequence shares less than 75% sequence identity with a naturally occurring or wild-type sequence (e.g., a naturally occurring or wild-type mRNA sequence encoding a nucleobase editing enzyme).
[0289] In some embodiments, the codon-optimized sequence shares 65% to 85% (e.g., about 67% to about 85% or about 67% to about 80%) sequence identity with a naturally occurring or wild-type sequence (e.g., a naturally occurring or wild-type mRNA sequence encoding a nucleobase editing enzyme). In some embodiments, the codon-optimized sequence shares 65% to 75% or about 80% sequence identity with a naturally occurring or wild-type sequence (e.g., a naturally occurring or wild-type mRNA sequence encoding a nucleobase editing enzyme).
[0290] When transfected into mammalian cells, the modified mRNA payload has a stability of 12 to 18 hours, or greater than 18 hours, e.g., 24, 36, 48, 60, 72, or greater than 72 hours.
[0291] In some embodiments, codon-optimized RNA may have an increased level of G / C. The G / C content of a nucleic acid molecule (e.g., mRNA) may affect the stability of the RNA. RNA with an increased amount of guanine (G) and / or cytosine (C) residues may be more functionally stable than RNA containing a large amount of adenine (A) and thymine (T) or uracil (U) nucleotides. For example, WO 02 / 098443 discloses a pharmaceutical composition containing mRNA stabilized by sequence modifications in the translated region. Due to the degeneracy of the genetic code, the modifications work by replacing existing codons with ones that promote greater RNA stability without changing the resulting amino acid. The approach is limited to the coding region of the RNA.
[0292] In some embodiments, the disclosure provides engineered Cas12a variants or mutants that have been modified by introducing one or more amino acid substitutions into a baseline sequence (e.g., a wild-type sequence).
[0293] Any available method can be used to obtain or construct variant or mutant Cas12a proteins. The term "mutation," as used herein, refers to the substitution of a residue in a sequence, e.g., a nucleic acid or amino acid sequence, with another residue, or the deletion or insertion of one or more residues in a sequence. Mutations are typically described herein by identifying the position of the original and subsequent residues in the sequence and the identity of the newly substituted residues. Various methods for generating the amino acid substitutions (mutations) provided herein are well known in the art (e.g., site-directed mutagenesis or directed evolution) and are provided, for example, by Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY (2012)). Mutations can include a variety of categories, e.g., single nucleotide polymorphisms, microduplication regions, indels, and inversions, and are not intended to be limiting in any way. Mutations can include "loss-of-function" mutations, which are the normal result of a mutation that reduces or abolishes protein activity. Most loss-of-function mutations are recessive because in heterozygotes, the second chromosome copy carries an unmutated version of the gene that encodes a fully functional protein, the presence of which offsets the effect of the mutation. Mutations also include "gain-of-function" mutations, which confer abnormal activity to proteins or cells that are not normally present. Many gain-of-function mutations are in regulatory sequences rather than coding regions, and therefore can have multiple consequences. For example, a mutation may result in one or more genes being expressed in the wrong tissues, and these tissues will acquire a function that they normally lack. Due to their nature, gain-of-function mutations are usually dominant.
[0294] Mutations can be introduced into a reference Cas12a protein using site-directed mutagenesis. Older methods of site-directed mutagenesis known in the art rely on subcloning the sequence to be mutated into a vector, such as an M13 bacteriophage vector, which allows for the isolation of a single-stranded DNA template. In these methods, a mutagenic primer (i.e., a primer that anneals to the site to be mutated but can possess one or more mismatched nucleotides at the site to be mutated) is annealed to the single-stranded template, and the complement of the template is then polymerized starting from the 3' end of the mutagenic primer. The resulting double strand is then transformed into host bacteria, and plaques are screened for the desired mutation. More recently, site-directed mutagenesis has used PCR methodology, which has the advantage of not requiring a single-stranded template. Methods that do not require subcloning have also been developed. When PCR-based site-directed mutagenesis is performed, several issues must be considered. First, in these methods, it is desirable to reduce the number of PCR cycles to prevent the increase of unwanted mutations introduced by the polymerase. Second, selection must be used to reduce the number of non-mutated parent molecules that persist in the reaction. Third, extended-length PCR methods are preferred because they allow the use of a single PCR primer set. And fourth, due to the non-template-dependent end-extension activity of some thermostable polymerases, it is often necessary to incorporate an end-polishing step into the procedure before blunt-end ligation of the PCR-generated mutant products.
[0295] Mutations can also be introduced by directed evolution processes, such as phage-assisted continuous evolution (PACE) or phage-assisted discontinuous evolution (PANCE). The term "phage-assisted continuous evolution (PACE)" as used herein refers to continuous evolution using phages as viral vectors. The general concept of PACE technology is described, for example, in International PCT Application PCT / US2009 / 056194, filed September 8, 2009, published as WO2010 / 028347 on March 11, 2010; International PCT Application PCT / US2011 / 066747, filed December 22, 2011, published as WO2012 / 088381 on June 28, 2012; and U.S. Patent No. 9,022,029, issued May 5, 2015. 3,594, International PCT Application No. PCT / US2015 / 012022, filed January 20, 2015, published on September 11, 2015 as WO2015 / 134121, and International PCT Application No. PCT / US2016 / 027795, filed April 15, 2016, published on October 20, 2016 as WO2016 / 168631, the entire contents of each of which are incorporated herein by reference. Variant Cas12as may also be obtained by "phage-assisted discontinuous evolution (PANCE)," which, as used herein, refers to discontinuous evolution using phages as viral vectors. PANCE is a simplified technique for in vivo directed evolution that uses serial flask transfer to evolve a "selection phage" (SP) containing the gene of interest to be evolved through a new E. coli host cell, thereby allowing the genes contained in the SP to be continuously evolved while the genes in the host E. coli remain constant. Serial flask transfer has long served as a widely accessible approach for the laboratory evolution of microorganisms, and more recently, a similar approach has been developed for bacteriophage evolution. The PANCE system is characterized by lower stringency than the PACE system.
[0296] The present disclosure contemplates any engineered Cas12a variant or mutant that has been modified by introducing one or more amino acid substitutions into the baseline sequence, including conservative substitutions of one amino acid for another. For example, a mutation of an amino acid with a hydrophobic side chain (e.g., alanine, valine, isoleucine, leucine, methionine, phenylalanine, tyrosine, or tryptophan) can be changed to a second amino acid with a different hydrophobic side chain (e.g., alanine, valine, isoleucine, leucine, methionine, phenylalanine, tyrosine, or tryptophan). For example, a mutation of alanine to threonine (e.g., A262T mutation) can also include a mutation of alanine to an amino acid with similar size and chemical properties to threonine, such as serine. As another example, mutation of an amino acid having a positively charged side chain (e.g., arginine, histidine, or lysine) can include mutation to a second amino acid having a different positively charged side chain (e.g., arginine, histidine, or lysine). As another example, mutation of an amino acid having a polar side chain (e.g., serine, threonine, asparagine, or glutamine) can also include mutation to a second amino acid having a different polar side chain (e.g., serine, threonine, asparagine, or glutamine). Additional similar amino acid pairs include, but are not limited to, the following: phenylalanine and tyrosine; asparagine and glutamine; methionine and cysteine; aspartic acid and glutamic acid; and arginine and lysine. Those skilled in the art will recognize that such conservative amino acid substitutions may have only a minor effect on protein structure and may be well tolerated without impairing function. In some embodiments, any amino acid mutation provided herein from an amino acid to threonine can be an amino acid mutation to serine. In some embodiments, any amino acid mutation provided herein from an amino acid to arginine can be an amino acid mutation to lysine. In some embodiments, any amino acid mutation provided herein from an amino acid to isoleucine can be an amino acid mutation to alanine, valine, methionine, or leucine.In some embodiments, any amino acid mutation provided herein from an amino acid to lysine can be an amino acid mutation to arginine. In some embodiments, any amino acid mutation provided herein from an amino acid to aspartic acid can be an amino acid mutation to glutamic acid or asparagine. In some embodiments, any amino acid mutation provided herein from an amino acid to valine can be an amino acid mutation to alanine, isoleucine, methionine, or leucine. In some embodiments, any amino acid mutation provided herein from an amino acid to glycine can be an amino acid mutation to alanine. However, it should be understood that additional conserved amino acid residues will be recognized by those skilled in the art, and that any amino acid mutation to other conserved amino acid residues is also within the scope of the present disclosure. Amino acid substitutions can also be non-conservative amino acid substitutions.
[0297] In various embodiments, the alanine (A) residue of the Cas12a protein may be substituted with any one of the following amino acids: arginine (R); asparagine (N); aspartic acid (D); cysteine (C); glutamic acid (E); glutamine (N); glycine (G); histidine (H); isoleucine (I); leucine (L); lysine (K); methionine (M); phenylalanine (F); proline (P); serine (S); threonine (T); tryptophan (W); tyrosine (Y); or valine (V).
[0298] In another embodiment, the arginine (R) residue of the Cas12a protein may be substituted with any one of the following amino acids: alanine (A); asparagine (N); aspartic acid (D); cysteine (C); glutamic acid (E); glutamine (N); glycine (G); histidine (H); isoleucine (I); leucine (L); lysine (K); methionine (M); phenylalanine (F); proline (P); serine (S); threonine (T); tryptophan (W); tyrosine (Y); or valine (V).
[0299] In another embodiment, the asparagine (N) residue of the Cas12a protein may be substituted with any one of the following amino acids: alanine (A); arginine (R); aspartic acid (D); cysteine (C); glutamic acid (E); glutamine (N); glycine (G); histidine (H); isoleucine (I); leucine (L); lysine (K); methionine (M); phenylalanine (F); proline (P); serine (S); threonine (T); tryptophan (W); tyrosine (Y); or valine (V).
[0300] In another embodiment, the aspartic acid (D) residue of the Cas12a protein may be substituted with any one of the following amino acids: alanine (A); arginine (R); asparagine (N); cysteine (C); glutamic acid (E); glutamine (N); glycine (G); histidine (H); isoleucine (I); leucine (L); lysine (K); methionine (M); phenylalanine (F); proline (P); serine (S); threonine (T); tryptophan (W); tyrosine (Y); or valine (V).
[0301] In another embodiment, the cysteine (C) residue of the Cas12a protein may be substituted with any one of the following amino acids: alanine (A); arginine (R); asparagine (N); aspartic acid (D); glutamic acid (E); glutamine (N); glycine (G); histidine (H); isoleucine (I); leucine (L); lysine (K); methionine (M); phenylalanine (F); proline (P); serine (S); threonine (T); tryptophan (W); tyrosine (Y); or valine (V).
[0302] In another embodiment, the glutamic acid (E) residue of the Cas12a protein may be substituted with any one of the following amino acids: alanine (A); arginine (R); asparagine (N); aspartic acid (D); cysteine (C); glutamine (N); glycine (G); histidine (H); isoleucine (I); leucine (L); lysine (K); methionine (M); phenylalanine (F); proline (P); serine (S); threonine (T); tryptophan (W); tyrosine (Y); or valine (V).
[0303] In another embodiment, the glutamine (N) residue of the Cas12a protein may be substituted with any one of the following amino acids: alanine (A); arginine (R); asparagine (N); aspartic acid (D); cysteine (C); glutamic acid (E); glycine (G); histidine (H); isoleucine (I); leucine (L); lysine (K); methionine (M); phenylalanine (F); proline (P); serine (S); threonine (T); tryptophan (W); tyrosine (Y); or valine (V).
[0304] In another embodiment, the glycine (G) residue of the Cas12a protein may be substituted with any one of the following amino acids: alanine (A); arginine (R); asparagine (N); aspartic acid (D); cysteine (C); glutamic acid (E); glutamine (N); histidine (H); isoleucine (I); leucine (L); lysine (K); methionine (M); phenylalanine (F); proline (P); serine (S); threonine (T); tryptophan (W); tyrosine (Y); or valine (V).
[0305] In another embodiment, the histidine (H) residue of the Cas12a protein may be substituted with any one of the following amino acids: alanine (A); arginine (R); asparagine (N); aspartic acid (D); cysteine (C); glutamic acid (E); glutamine (N); glycine (G); isoleucine (I); leucine (L); lysine (K); methionine (M); phenylalanine (F); proline (P); serine (S); threonine (T); tryptophan (W); tyrosine (Y); or valine (V).
[0306] In another embodiment, the isoleucine (I) residue of the Cas12a protein may be substituted with any one of the following amino acids: alanine (A); arginine (R); asparagine (N); aspartic acid (D); cysteine (C); glutamic acid (E); glutamine (N); glycine (G); histidine (H); leucine (L); lysine (K); methionine (M); phenylalanine (F); proline (P); serine (S); threonine (T); tryptophan (W); tyrosine (Y); or valine (V).
[0307] In another embodiment, the leucine (L) residue of the Cas12a protein may be substituted with any one of the following amino acids: alanine (A); arginine (R); asparagine (N); aspartic acid (D); cysteine (C); glutamic acid (E); glutamine (N); glycine (G); histidine (H); isoleucine (I); lysine (K); methionine (M); phenylalanine (F); proline (P); serine (S); threonine (T); tryptophan (W); tyrosine (Y); or valine (V).
[0308] In another embodiment, the lysine (K) residue of the Cas12a protein may be substituted with any one of the following amino acids: alanine (A); arginine (R); asparagine (N); aspartic acid (D); cysteine (C); glutamic acid (E); glutamine (N); glycine (G); histidine (H); isoleucine (I); leucine (L); methionine (M); phenylalanine (F); proline (P); serine (S); threonine (T); tryptophan (W); tyrosine (Y); or valine (V).
[0309] In another embodiment, the methionine (M) residue of the Cas12a protein may be substituted with any one of the following amino acids: alanine (A); arginine (R); asparagine (N); aspartic acid (D); cysteine (C); glutamic acid (E); glutamine (N); glycine (G); histidine (H); isoleucine (I); leucine (L); lysine (K); phenylalanine (F); proline (P); serine (S); threonine (T); tryptophan (W); tyrosine (Y); or valine (V).
[0310] In another embodiment, the phenylalanine (F) residue of the Cas12a protein may be substituted with any one of the following amino acids: alanine (A); arginine (R); asparagine (N); aspartic acid (D); cysteine (C); glutamic acid (E); glutamine (N); glycine (G); histidine (H); isoleucine (I); leucine (L); lysine (K); methionine (M); proline (P); serine (S); threonine (T); tryptophan (W); tyrosine (Y); or valine (V).
[0311] In another embodiment, the proline (P) residue of the Cas12a protein may be substituted with any one of the following amino acids: alanine (A); arginine (R); asparagine (N); aspartic acid (D); cysteine (C); glutamic acid (E); glutamine (N); glycine (G); histidine (H); isoleucine (I); leucine (L); lysine (K); methionine (M); phenylalanine (F); serine (S); threonine (T); tryptophan (W); tyrosine (Y); or valine (V).
[0312] In another embodiment, the serine (S) residue of the Cas12a protein may be substituted with any one of the following amino acids: alanine (A); arginine (R); asparagine (N); aspartic acid (D); cysteine (C); glutamic acid (E); glutamine (N); glycine (G); histidine (H); isoleucine (I); leucine (L); lysine (K); methionine (M); phenylalanine (F); proline (P); threonine (T); tryptophan (W); tyrosine (Y); or valine (V).
[0313] In another embodiment, the threonine (T) residue of the Cas12a protein may be substituted with any one of the following amino acids: alanine (A); arginine (R); asparagine (N); aspartic acid (D); cysteine (C); glutamic acid (E); glutamine (N); glycine (G); histidine (H); isoleucine (I); leucine (L); lysine (K); methionine (M); phenylalanine (F); proline (P); serine (S); tryptophan (W); tyrosine (Y); or valine (V).
[0314] In another embodiment, the tryptophan (W) residue of the Cas12a protein may be substituted with any one of the following amino acids: alanine (A); arginine (R); asparagine (N); aspartic acid (D); cysteine (C); glutamic acid (E); glutamine (N); glycine (G); histidine (H); isoleucine (I); leucine (L); lysine (K); methionine (M); phenylalanine (F); proline (P); serine (S); threonine (T); tyrosine (Y); or valine (V).
[0315] In another embodiment, the tyrosine (Y) residue of the Cas12a protein may be substituted with any one of the following amino acids: alanine (A); arginine (R); asparagine (N); aspartic acid (D); cysteine (C); glutamic acid (E); glutamine (N); glycine (G); histidine (H); isoleucine (I); leucine (L); lysine (K); methionine (M); phenylalanine (F); proline (P); serine (S); threonine (T); tryptophan (W); or valine (V).
[0316] In another embodiment, the valine (V) residue of the Cas12a protein may be substituted with any one of the following amino acids: alanine (A); arginine (R); asparagine (N); aspartic acid (D); cysteine (C); glutamic acid (E); glutamine (N); glycine (G); histidine (H); isoleucine (I); leucine (L); lysine (K); methionine (M); phenylalanine (F); proline (P); serine (S); threonine (T); or tryptophan (W).
[0317] Amino acid substitutions may also include any non-naturally occurring amino acid analogue or amino acid derivative known in the art.
[0318] Without intending to be limiting, the following are exemplary embodiments of mutant variants contemplated by the specification and examples and based on Cas12a ID405 (SEQ ID NO: 334), Cas12a ID414 (SEQ ID NO: 58), and Cas12a ID418 (SEQ ID NO: 564). It will be understood that any of the specific substitutions and / or combinations of specific substitutions below can be introduced into the corresponding amino acid residues (as determined by sequence alignment) of any other Type V nuclease enzyme disclosed herein.
[0319] Variant based on ID405 (SEQ ID NO: 334) In various embodiments, the Cas12a may be a Cas12a variant based on ID405 (SEQ ID NO: 334) and may include any of the following substitutions, in any combination (or any amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 99%, or up to 100% sequence identity to SEQ ID NO: 334 with any of the following substitutions): ·D169 substitution; ·C554 replacement; ·N559 replacement; ·Q565 replacement; ·L860 replacement; R950 replacement; and / or ·R954 replacement.
[0320] In various embodiments, the Cas12a may be a Cas12a variant based on ID405 (SEQ ID NO: 334) and may include any of the following substitutions, in any combination (or any amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 99%, or up to 100% sequence identity to SEQ ID NO: 334 with any of the following substitutions): ·D169R replacement; ·C554N substitution; ·C554R substitution; ·N559R replacement; ·Q565R replacement; ·L860Q replacement; R950K replacement; and / or ·R954A replacement.
[0321] In various embodiments, the Cas12a may be a Cas12a variant based on ID405 (SEQ ID NO: 334) and may include any of the following substitutions, in any combination (or any amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 99%, or up to 100% sequence identity to SEQ ID NO: 334 with any of the following substitutions): ·D169 substitution; D169 / R950 / R954 replacement set; D169 / N559 / Q565 replacement set; ·C554 replacement; C554 substitution; and / or ·L860 replacement.
[0322] In various embodiments, the Cas12a may be a Cas12a variant based on ID405 (SEQ ID NO: 334) and may include any of the following substitutions, in any combination (or any amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 99%, or up to 100% sequence identity to SEQ ID NO: 334 with any of the following substitutions): ·D169R replacement; D169R / R950K / R954A replacement set; D169R / N559R / Q565R replacement set; ·C554R substitution; C554N substitution; and / or ·L860Q replacement.
[0323] The full-length amino acid and protein coding sequences of these mutant nucleases are provided in Section K, Subsection P (Cas12a Mutant Type V Nucleases and Related Sequences).
[0324] Variant based on ID414 (SEQ ID NO: 58) In various embodiments, the Cas12a may be a Cas12a variant based on ID414 (SEQ ID NO: 58) and may include any of the following substitutions, in any combination (or any amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 99%, or up to 100% sequence identity to SEQ ID NO: 58 with any of the following substitutions): ·T154 replacement; ·N531 replacement; ·G546 replacement; ·K542 replacement; ·S802 replacement; R887 substitution; and / or ·R891 replacement.
[0325] In various embodiments, the Cas12a may be a Cas12a variant based on ID414 (SEQ ID NO: 58) and may include any of the following substitutions, in any combination (or any amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 99%, or up to 100% sequence identity to SEQ ID NO: 58 with any of the following substitutions): ·T154R substitution; ·N531R replacement; ·G546R replacement; ·K542R substitution; ·S802L replacement; R887K substitution; and / or ·R891A replacement.
[0326] In various embodiments, the Cas12a may be a Cas12a variant based on ID414 (SEQ ID NO: 58) and may include any of the following substitutions, in any combination (or any amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 99%, or up to 100% sequence identity to SEQ ID NO: 58 with any of the following substitutions): ·T154 replacement; ·T154 / R887 / R891 replacement; ·T154 / G536 / K542 replacement; ·N531 / S802 replacement; N531 substitution; and / or ·S802 replacement.
[0327] In various embodiments, the Cas12a may be a Cas12a variant based on ID414 (SEQ ID NO: 58) and may include any of the following substitutions, in any combination (or any amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 99%, or up to 100% sequence identity to SEQ ID NO: 58 with any of the following substitutions): ·T154R substitution; ·T154R / R887K / R891A replacement; ·T154R / G536R / K542R replacement; ·N531R / S802L replacement; N531R substitution; and / or ·S802L replacement.
[0328] The full-length amino acid and protein coding sequences of these mutant nucleases are provided in Section K, Subsection P (Cas12a Mutant Type V Nucleases and Related Sequences).
[0329] Variant based on ID418 (SEQ ID NO: 564) In various embodiments, the Cas12a may be a Cas12a variant based on ID418 (SEQ ID NO: 564) and may include any of the following substitutions, in any combination (or any amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 99%, or up to 100% sequence identity to SEQ ID NO: 564 with any of the following substitutions): ·D161 substitution; ·N527 replacement; ·T532 replacement; ·K538 replacement; ·Q799 substitution; R888 substitution; and / or ·R892 replacement.
[0330] In various embodiments, the Cas12a may be a Cas12a variant based on ID418 (SEQ ID NO: 564) and may include any of the following substitutions, in any combination (or any amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 99%, or up to 100% sequence identity to SEQ ID NO: 564 with any of the following substitutions): ·D161R substitution; ·N527R substitution; ·T532R replacement; ·K538R replacement; ·Q799L replacement; R888K substitution; and / or ·R892A replacement.
[0331] In various embodiments, the Cas12a may be a Cas12a variant based on ID418 (SEQ ID NO: 564) and may include any of the following substitutions, in any combination (or any amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 99%, or up to 100% sequence identity to SEQ ID NO: 564 with any of the following substitutions): ·D161 substitution; ·D161 / R888 / R892 replacement; ·D161 / T532 / K538 replacement; ·N527 / Q799 replacement; N527 substitution; and / or · Q799 replacement.
[0332] In various embodiments, the Cas12a may be a Cas12a variant based on ID418 (SEQ ID NO: 564) and may include any of the following substitutions, in any combination (or any amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 99%, or up to 100% sequence identity to SEQ ID NO: 564 with any of the following substitutions): ·D161R substitution; ·D161R / R888K / R892A replacement; ·D161R / T532R / K538R replacement; ·N527R / Q799L replacement; N527R substitution; and / or ·Q799L replacement.
[0333] The full-length amino acid and protein coding sequences of these mutant nucleases are provided in Section K, Subsection P (Cas12a Mutant Type V Nucleases and Related Sequences).
[0334] Various embodiments of variant Cas12a orthologs are also described in Subsection Q of Section K.
[0335] Additionally, embodiments of Cas12a mutant variants based on ID405, ID414, and ID418 are described in a computational approach to directed mutagenesis described in Example 14.
[0336] The present disclosure provides polypeptides (including in Appendix A and elsewhere herein, including the Examples) that have a percent identity with respect to another amino acid sequence (Reference amino acid sequence), for example, a polypeptide that is at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, ... When referring to a polypeptide that is at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% identical to a reference amino acid sequence, it means that in a polypeptide having a percent identity to a reference amino acid sequence, conserved regions of the reference amino acid sequence (e.g., conserved compared to other Cas12a, e.g., those identified herein, as set forth in the multi-sequence alignment of Figure 31) are conserved, and / or the polypeptide has at least one activity selected from endonuclease activity; endoribonuclease activity, or RNA-guided DNase activity, and / or the polypeptide comprises one or more of: a. one or more α-helix recognition lobes (REC) and nuclease lobes (NUC); b. wedge (WED), α-helix recognition lobes (REC), PAM interaction (PI), RuvC nuclease, bridging helix (BH), and NUC domains; or c.It is noted that it is advantageous to include one or more domains selected from RuvC, REC, WED, BH, PI and NUC domains and / or polypeptides that recognize or bind to crRNA(s) or are bound to crRNA(s), such as the crRNA sequences from Table S15C. Similarly, the present disclosure relates to a nucleic acid sequence or molecule that has a percent identity with respect to another nucleic acid sequence or molecule (reference nucleic acid sequence), for example, a sequence selected from SEQ ID NO: 365 (No. ID405), SEQ ID NO: 74 (No. ID414), or SEQ ID NO: 565 (No. ID418), SEQ ID NO: 366 (No. ID406), SEQ ID NO: 331 (No. ID411), SEQ ID NO: 30 (No. ID415), or SEQ ID NO: 445 (No. ID419)). When referring to a nucleic acid sequence that is 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% identical, it means that in a nucleic acid sequence having a percent identity to a reference nucleic acid sequence, the conserved region of the reference nucleic acid sequence (e.g., conserved compared to other Cas12a, e.g., those identified herein) is conserved, and / or in a polypeptide expressed from a nucleic acid sequence having a percent identity to a reference nucleic acid sequence, the polypeptide contains conserved region(s) (e.g., conserved compared to other Cas12a, e.g., those identified herein), and / or the polypeptide has at least one activity selected from endonuclease activity; endoribonuclease activity, or RNA-guided DNase activity, and / or the polypeptide comprises: a. one or more α-helical recognition lobes (REC) and nuclease lobes (NUC); b.Advantageously, the polypeptide comprises one or more domains selected from the wedge (WED), α-helix recognition lobe (REC), PAM interaction (PI), RuvC nuclease, bridging helix (BH), and NUC domains; or c. RuvC, REC, WED, BH, PI, and NUC domains and / or recognizes or binds to crRNA(s) or is bound to crRNA(s), such as the crRNA sequences from Table S15C.
[0337] D. Cas12a (or Cas V) guide RNA sequence Cas12a (Cas V type) guide sequence The present disclosure further provides guide RNAs for use with the disclosed nucleic acid programmable DNA binding proteins (e.g., Cas12a) for use in editing methods. The present disclosure provides guide RNAs designed to recognize target sequences. Such gRNAs can be designed to have a guide sequence (or "spacer") that is complementary to the target sequence. Such gRNAs can be designed not only to have a guide sequence that is complementary to the target sequence to be edited, but also to have a backbone sequence that specifically interacts with the nucleic acid programmable DNA binding protein.
[0338] In various aspects, one or more guide RNA sequences are provided. In preferred embodiments, the gRNA is cleaved and processed into one or more intermediate crRNAs, which are then processed into one or more mature crRNAs. In some embodiments, the gRNA comprises one or more crRNAs or a precursor CRISPR RNA (pre-crRNA) encoding one or more intermediate or mature crRNAs, each guide RNA comprising, at a minimum, a repeat-spacer in the 5' to 3' direction, the repeat comprising a stem-loop structure, and the spacer comprising a DNA-targeting segment complementary to a target sequence in the targeted polynucleotide sequence. In certain embodiments, the gRNA is cleaved by the RNase activity of the Cas12a polypeptide into one or more mature crRNAs, each comprising at least one repeat and at least one spacer.
[0339] In other embodiments, one or more repeat-spacers direct the Cas12a (or Cas V) polypeptide to two or more distinct sites in the targeted polynucleotide sequence. Preferably, the gRNA is cleaved and processed into one or more intermediate crRNAs, which are then processed into one or more mature crRNAs. More preferably, the pre-crRNA or intermediate crRNA is processed into mature crRNA by the Cas12a (or Cas V) polypeptide, and the mature crRNA becomes available to induce Cas12a (or Cas V) endonuclease activity. In an alternative embodiment, the gRNA is linked to a single- or double-stranded DNA donor template, and the donor template is cleaved from the gRNA by the Cas12a (or Cas V) polypeptide. The donor polynucleotide template remains linked to the gRNA, while the Cas12a (or Cas V) polypeptide cleaves the gRNA to release the intermediate or mature crRNA.
[0340] In an exemplary embodiment, the Cas12a (or Cas V) system: (a) One or more crRNA direct repeat sequences or reverse complements selected from (Group 1) SEQ ID NOs: 7-12; (Group 2) SEQ ID NOs: 24-27; (Group 3) SEQ ID NOs: 36-39; (Group 4) SEQ ID NOs: 49-52; (Group 5) SEQ ID NOs: 63-68; (Group 6) SEQ ID NOs: 84-91; (Group 7) SEQ ID NOs: 106-111; (Group 8) SEQ ID NOs: 122-125; (Group 9) SEQ ID NOs: 211-290; (Group 10) SEQ ID NOs: 343-354; (Group 11) SEQ ID NOs: 374-379; (Group 12) SEQ ID NOs: 390-393; (Group 13) SEQ ID NOs: 411-422; and (Group 14) SEQ ID NOs: 500-541; (b) 20-35 nucleotides from the 3' end of the crRNA direct repeat sequence or reverse complement (a) linked to a targeting guide attached to the 3' end of the direct repeat sequence, which is 16-30 nucleotides in length, or up to the length of the crRNA; (c) (Group 1) SEQ ID NOs: 13 to 15; (Group 2) SEQ ID NOs: 28 to 29; (Group 3) SEQ ID NOs: 40 to 41; (Group 4) SEQ ID NOs: 53 to 54; (Group 5) SEQ ID NOs: 69 to 71; (Group 6) SEQ ID NOs: 92 to 95; (Group 7) SEQ ID NOs: 112 to 114; (Group 8) SEQ ID NOs: 126 to 127; (Group 9) SEQ ID NOs: 291 to 330; (Group 10) SEQ ID NOs: 355 to 360; (Group 11) SEQ ID NOs: 380 to 382; (Group 12) SEQ ID NOs: 394 to 395; (Group 13) SEQ ID NOs: 423 to 428; and (Group 14) SEQ ID NOs: 542 to 563; (d) Nucleic acid sequences that are degenerate variants of (Group 1) SEQ ID NOs: 13 to 15; (Group 2) SEQ ID NOs: 28 to 29; (Group 3) SEQ ID NOs: 40 to 41; (Group 4) SEQ ID NOs: 53 to 54; (Group 5) SEQ ID NOs: 69 to 71; (Group 6) SEQ ID NOs: 92 to 95; (Group 7) SEQ ID NOs: 112 to 114; (Group 8) SEQ ID NOs: 126 to 127; (Group 9) SEQ ID NOs: 291 to 330; (Group 10) SEQ ID NOs: 355 to 360; (Group 11) SEQ ID NOs: 380 to 382; (Group 12) SEQ ID NOs: 394 to 395; (Group 13) SEQ ID NOs: 423 to 428; and (Group 14) SEQ ID NOs: 542 to 563; (e) (Group 1) SEQ ID NOs: 13 to 15; (Group 2) SEQ ID NOs: 28 to 29; (Group 3) SEQ ID NOs: 40 to 41; (Group 4) SEQ ID NOs: 53 to 54; (Group 5) SEQ ID NOs: 69 to 71; (Group 6) SEQ ID NOs: 92 to 95; (Group 7) SEQ ID NOs: 112 to 114; (Group 8) SEQ ID NOs: 126 to 127; (Group 9) SEQ ID NOs: 291 to 330; (Group 10) SEQ ID NOs: 355 to 360; (Group 11) SEQ ID NOs: 380 to 382; (Group 12) SEQ ID NOs: 394 to 395; (Group 13) SEQ ID NOs: Nos. 423-428; and (Group 14) nucleic acid sequences at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or at least 99.9% identical to SEQ ID NOs: 542-563; and (f) Nucleic acid sequences that hybridize under stringent conditions to (Group 1) SEQ ID NOs: 13 to 15; (Group 2) SEQ ID NOs: 28 to 29; (Group 3) SEQ ID NOs: 40 to 41; (Group 4) SEQ ID NOs: 53 to 54; (Group 5) SEQ ID NOs: 69 to 71; (Group 6) SEQ ID NOs: 92 to 95; (Group 7) SEQ ID NOs: 112 to 114; (Group 8) SEQ ID NOs: 126 to 127; (Group 9) SEQ ID NOs: 291 to 330; (Group 10) SEQ ID NOs: 355 to 360; (Group 11) SEQ ID NOs: 380 to 382; (Group 12) SEQ ID NOs: 394 to 395; (Group 13) SEQ ID NOs: 423 to 428; and (Group 14) SEQ ID NOs: 542 to 563. and one or more guide RNAs comprising:
[0341] In preferred embodiments, the Cas12a (or Cas V) protein targets and cleaves a targeted polynucleotide that is complementary to the cognate guide RNA. In certain embodiments, the guide RNA comprises a crRNA that contains a natural CRISPR array. Such variants are derived from the first direct repeat, which is a "leader" sequence and is involved in signal transduction, or the direct repeat retains genetic diversity without affecting functionality. The direct repeat is degenerate and typically located near the 3' end of the repeat array.
[0342] In various embodiments, the crRNA contains a direct repeat sequence comprising approximately 15-40 nucleotides or approximately 20-30 nucleotides. In exemplary embodiments, the direct repeat is selected from: (Group 1) SEQ ID NOS: 7-12; (Group 2) SEQ ID NOS: 24-27; (Group 3) SEQ ID NOS: 36-39; (Group 4) SEQ ID NOS: 49-52; (Group 5) SEQ ID NOS: 63-68; (Group 6) SEQ ID NOS: 84-91; (Group 7) SEQ ID NOS: 106-111; (Group 8) SEQ ID NOS: 122-125; (Group 9) SEQ ID NOS: 211-290; (Group 10) SEQ ID NOS: 343-354; (Group 11) SEQ ID NOS: 374-379; (Group 12) SEQ ID NOS: 390-393; (Group 13) SEQ ID NOS: 411-422; and (Group 14) SEQ ID NOS: 500-541. More preferably, the crRNA contains a guide segment of 16-26 nucleotides or 20-24 nucleotides. Thus, in various embodiments, the crRNA of the Cas12a genome editing system hybridizes to one or more targeted polynucleotide sequences. In certain preferred embodiments, the crRNA is 43 nucleotides. In other embodiments, the crRNA is composed of a 20-nucleotide 5'-handle and a 23-nucleotide leader sequence. In certain embodiments, the leader sequence includes a seed region and a 3' end, both of which are complementary to the target region in the genome. Li, Bin et al. "Engineering CRISPR-Cpf1 crRNAs and mRNAs to maximize genome editing efficiency." Nature biomedical engineering vol. 1, 5 (2017): 0066. doi:10.1038 / s41551-017-0066.
[0343] A single crRNA-guided endonuclease has ribonuclease activity to process pre-crRNA into mature crRNA. Zetsche, Bernd et al. "A Survey of Genome Editing Activity for 16 Cas12a Orthologs." The Keio journal of medicine vol.69,3(2020):59-65.doi:10.2302 / kjm.2019-0009-OA;Fonfara, Ines et al. “The CRISPR-associated DNA-cleaving enzyme Cpf1 also processes precursor CRISPR RNA.”Nature vol.532,7600(2016):517-21.doi:10.1038 / nature17945,which enables multiplex editing in a single crRNA transcript.Campa,Carlo C et al.“Multiplexed genome engineering by Cas12a and CRISPR arrays encoded on single transcripts.”Nature methods vol.16,9(2019):887-893.doi:10.1038 / s41592-019-0508-6;Zetsche,Bernd et al.“Multiplex gene editing by CRISPR-Cpf1 using a single crRNA array.”Nature biotechnology vol.35,1(2017):31-34.doi:10.1038 / nbt.3737.
[0344] Preferably, the crRNA-guided endonuclease provides for the modification of multiple loci in the host cell genome.
[0345] More preferably, Cas12a (or Cas V type) multiplexing is performed using two methods. One method involves expressing many single gRNAs under different small RNA promoters in either the same vector or different vectors. Another method involves multiple single gRNAs fused to tRNA recognition sequences, which are expressed as a single transcript under a single promoter.
[0346] In some embodiments, the guide RNA is 15-100 nucleotides in length and may comprise a sequence of at least 10, at least 15, or at least 20 contiguous nucleotides complementary to the target nucleotide sequence. The guide RNA may comprise a spacer sequence of 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40 contiguous nucleotides complementary to the target nucleotide sequence. In some cases, the guide sequence has a length in the range of 17-30 nucleotides (nt) (e.g., 17-25, 17-22, 17-20, 19-30, 19-25, 19-22, 19-20, 20-30, 20-25, or 20-22 nt). In some cases, the guide sequence has a length ranging from 17 to 25 nucleotides (nt) (e.g., 17-22, 17-20, 19-25, 19-22, 19-20, 20-25, or 20-22 nt). In some cases, the guide sequence has a length of 17 nt or more (e.g., 18 nt or more, 19 nt or more, 20 nt or more, 21 nt or more, or 22 nt or more; 19 nt, 20 nt, 21 nt, 22 nt, 23 nt, 24 nt, 25 nt, etc.). In some cases, the guide sequence has a length of 19 nt or more (e.g., 20 nt or more, 21 nt or more, or 22 nt or more; 19 nt, 20 nt, 21 nt, 22 nt, 23 nt, 24 nt, 25 nt, etc.). In some cases, the guide sequence has a length of 17 nt. In some cases, the guide sequence has a length of 18 nt. In some cases, the guide sequence has a length of 19 nt. In some cases, the guide sequence has a length of 20 nt. In some cases, the guide sequence has a length of 21 nt. In some cases, the guide sequence has a length of 22 nt. In some cases, the guide sequence has a length of 23 nt.
[0347] In some cases, the spacer sequence has a length of 15 to 50 nucleotides (e.g., 15 nucleotides (nt) to 20 nt, 20 nt to 25 nt, 25 nt to 30 nt, 30 nt to 35 nt, 35 nt to 40 nt, 40 nt to 45 nt, or 45 nt to 50 nt).
[0348] The subject guide RNA may interact with a target nucleic acid (e.g., double-stranded DNA (dsDNA), single-stranded DNA (ssDNA), single-stranded RNA (ssRNA), or double-stranded RNA (dsRNA)) in a sequence-specific manner via hybridization (i.e., base pairing).
[0349] The guide RNA can be modified to hybridize to any desired target sequence within a target nucleic acid (e.g., a eukaryotic target nucleic acid, e.g., genomic DNA) (e.g., taking into account PAM when targeting a dsDNA target). In some cases, the complementarity between the spacer sequence of the guide and the target site of the target nucleic acid is 60% or greater (e.g., 65% or greater, 70% or greater, 75% or greater, 80% or greater, 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100%). In some cases, the complementarity between the spacer and the target site of the target nucleic acid is 80% or greater (e.g., 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100%). In some cases, the complementarity between the spacer and the target site of the target nucleic acid is 90% or greater (e.g., 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100%). In some cases, the complementarity between the spacer and the target site of the target nucleic acid is 100%.
[0350] In some cases, the percent complementarity between the spacer sequence and the target site of the target nucleic acid is 100% over a contiguous region of at least 5 nucleotides of the spacer. In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid over a contiguous region of at least 6 nucleotides of the spacer is 60% or greater (e.g., 70% or greater, 75% or greater, 80% or greater, 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100%). In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid over a contiguous region of at least 7 nucleotides of the spacer is 60% or greater (e.g., 70% or greater, 75% or greater, 80% or greater, 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100%). In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid over a contiguous region of at least 8 nucleotides of the spacer is 60% or greater (e.g., 70% or greater, 75% or greater, 80% or greater, 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100%). In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid over a contiguous region of at least 9 nucleotides of the spacer is 60% or greater (e.g., 70% or greater, 75% or greater, 80% or greater, 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100%). In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid over a contiguous region of at least 10 nucleotides of the spacer is 60% or greater (e.g., 70% or greater, 75% or greater, 80% or greater, 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100%). In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid over a contiguous region of at least 11 nucleotides of the spacer is 60% or greater (e.g., 70% or greater, 75% or greater, 80% or greater, 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100%).In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid over a contiguous region of at least 12 nucleotides of the spacer is 60% or greater (e.g., 70% or greater, 75% or greater, 80% or greater, 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100%). In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid over a contiguous region of at least 13 nucleotides of the spacer is 60% or greater (e.g., 70% or greater, 75% or greater, 80% or greater, 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100%). In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid over a contiguous region of at least 14 nucleotides of the spacer is 60% or greater (e.g., 70% or greater, 75% or greater, 80% or greater, 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100%). In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid over a contiguous region of at least 15 nucleotides of the spacer is 60% or greater (e.g., 70% or greater, 75% or greater, 80% or greater, 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100%). In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid over a contiguous region of at least 16 nucleotides of the spacer is 60% or greater (e.g., 70% or greater, 75% or greater, 80% or greater, 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100%). In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid over a contiguous region of at least 17 nucleotides of the spacer is 60% or greater (e.g., 70% or greater, 75% or greater, 80% or greater, 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100%). In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid over a contiguous region of at least 18 nucleotides of the spacer is 60% or greater (e.g., 70% or greater, 75% or greater, 80% or greater, 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100%).In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid over a contiguous region of at least 19 nucleotides of the spacer is 60% or greater (e.g., 70% or greater, 75% or greater, 80% or greater, 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100%). In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid over a contiguous region of at least 20 nucleotides of the spacer is 60% or greater (e.g., 70% or greater, 75% or greater, 80% or greater, 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100%). In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid over a contiguous region of at least 21 nucleotides of the spacer is 60% or greater (e.g., 70% or greater, 75% or greater, 80% or greater, 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100%). In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid over a contiguous region of at least 22 nucleotides of the spacer is 60% or greater (e.g., 70% or greater, 75% or greater, 80% or greater, 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100%).
[0351] In some cases, the percent complementarity between the spacer sequence and the target site of the target nucleic acid is 100% over a contiguous region of at least 5-10 nucleotides of the spacer. In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid over a contiguous region of at least 6-11 nucleotides of the spacer is 60% or greater (e.g., 70% or greater, 75% or greater, 80% or greater, 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100%). In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid over a contiguous region of at least 7-12 nucleotides of the spacer is 60% or greater (e.g., 70% or greater, 75% or greater, 80% or greater, 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100%). In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid over a contiguous region of at least 8-13 nucleotides of the spacer is 60% or greater (e.g., 70% or greater, 75% or greater, 80% or greater, 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100%). In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid over a contiguous region of at least 9-14 nucleotides of the spacer is 60% or greater (e.g., 70% or greater, 75% or greater, 80% or greater, 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100%). In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid over a contiguous region of at least 10-15 nucleotides of the spacer is 60% or greater (e.g., 70% or greater, 75% or greater, 80% or greater, 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100%). In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid over a contiguous region of at least 11-16 nucleotides of the spacer is 60% or greater (e.g., 70% or greater, 75% or greater, 80% or greater, 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100%).In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid over a contiguous region of at least 12-17 nucleotides of the spacer is 60% or greater (e.g., 70% or greater, 75% or greater, 80% or greater, 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100%). In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid over a contiguous region of at least 13-18 nucleotides of the spacer is 60% or greater (e.g., 70% or greater, 75% or greater, 80% or greater, 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100%). In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid over a contiguous region of at least 14-19 nucleotides of the spacer is 60% or greater (e.g., 70% or greater, 75% or greater, 80% or greater, 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100%). In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid over a contiguous region of at least 15-20 nucleotides of the spacer is 60% or greater (e.g., 70% or greater, 75% or greater, 80% or greater, 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100%). In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid over a contiguous region of at least 16-21 nucleotides of the spacer is 60% or greater (e.g., 70% or greater, 75% or greater, 80% or greater, 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100%). In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid over a contiguous region of at least 17-22 nucleotides of the spacer is 60% or greater (e.g., 70% or greater, 75% or greater, 80% or greater, 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100%). In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid over a contiguous region of at least 18-23 nucleotides of the spacer is 60% or greater (e.g., 70% or greater, 75% or greater, 80% or greater, 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100%).In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid over a contiguous region of at least 19-24 nucleotides of the spacer is 60% or greater (e.g., 70% or greater, 75% or greater, 80% or greater, 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100%). In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid over a contiguous region of at least 20-25 nucleotides of the spacer is 60% or greater (e.g., 70% or greater, 75% or greater, 80% or greater, 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100%). In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid over a contiguous region of at least 21-26 nucleotides of the spacer is 60% or greater (e.g., 70% or greater, 75% or greater, 80% or greater, 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100%). In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid over a contiguous region of at least 22-27 nucleotides of the spacer is 60% or greater (e.g., 70% or greater, 75% or greater, 80% or greater, 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100%).
[0352] In various embodiments, the guide RNA can have a backbone or core region that complexes with a cognate nucleic acid programmable DNA binding protein (e.g., CRISPR Cas9 or Cas12a). In some cases, the guide RNA can have two stretches of nucleotides that are complementary to each other and hybridize to form a double-stranded RNA duplex (dsRNA duplex). Thus, in some cases, the protein-binding segment of the guide RNA comprises a dsRNA duplex. In some embodiments, the dsRNA double-stranded region includes a range of 5 to 25 base pairs (bp) (e.g., 5 to 22, 5 to 20, 5 to 18, 5 to 15, 5 to 12, 5 to 10, 5 to 8, 8 to 25, 8 to 22, 8 to 18, 8 to 15, 8 to 12, 12 to 25, 12 to 22, 12 to 18, 12 to 15, 13 to 25, 13 to 22, 13 to 18, 13 to 15, 14 to 25, 14 to 22, 14 to 18, 14 to 15, 15 to 25, 15 to 22, 15 to 18, 17 to 25, 17 to 22, or 17 to 18 bp, e.g., 5 bp, 6 bp, 7 bp, 8 bp, 9 bp, 10 bp, etc.). In some cases, the dsRNA double-stranded region includes a range of 6 to 15 base pairs (bp) (e.g., 6 to 12, 6 to 10, or 6 to 8 bp, e.g., 6 bp, 7 bp, 8 bp, 9 bp, 10 bp, etc.). In some cases, the double-stranded region includes 5 bp or more (e.g., 6 bp or more, 7 bp or more, or 8 bp or more). In some cases, the double-stranded region includes 6 bp or more (e.g., 7 bp or more, or 8 bp or more). In some cases, not all nucleotides in the double-stranded region are paired, and thus the duplex-forming region may include a bulge. The term "bulge" is used herein to mean a stretch of nucleotides (which may be a single nucleotide) that does not contribute to the duplexity of the duplex but is surrounded 5' and 3' by contributing nucleotides, and such a bulge is considered part of the duplex region. In some cases, the dsRNA includes one or more bulges (e.g., two or more, three or more, four or more bulges). In some cases, the dsRNA duplex includes two or more bulges (eg, three or more, four or more bulges).In some cases, the dsRNA duplex includes 1 to 5 bulges (e.g., 1 to 4, 1 to 3, 2 to 5, 2 to 4, or 2 to 3 bulges).
[0353] Thus, in some cases, the stretches of nucleotides that hybridize to each other to form a dsRNA duplex in the guide backbone region have between 70% and 100% complementarity to each other (e.g., 75% to 100%, 80% to 10%, 85% to 100%, 90% to 100%, 95% to 100% complementarity). In some cases, the stretches of nucleotides that hybridize to each other to form a dsRNA duplex have between 70% and 100% complementarity to each other (e.g., 75% to 100%, 80% to 10%, 85% to 100%, 90% to 100%, 95% to 100% complementarity). In some cases, the stretches of nucleotides that hybridize to each other to form a dsRNA duplex have 85% to 100% complementarity (e.g., 90% to 100%, 95% to 100% complementarity) with each other. In some cases, the stretches of nucleotides that hybridize to each other to form a dsRNA duplex have 70% to 95% complementarity (e.g., 75% to 95%, 80% to 95%, 85% to 95%, 90% to 95% complementarity) with each other. That is, in some cases, a dsRNA duplex includes two stretches of nucleotides that have 70% to 100% complementarity (e.g., 75% to 100%, 80% to 10%, 85% to 100%, 90% to 100%, 95% to 100% complementarity) with each other. In some cases, the dsRNA duplex includes two stretches of nucleotides that have 85%-100% complementarity to each other (e.g., 90%-100%, 95%-100% complementarity). In some cases, the dsRNA duplex includes two stretches of nucleotides that have 70%-95% complementarity to each other (e.g., 75%-95%, 80%-95%, 85%-95%, 90%-95% complementarity).
[0354] In various embodiments, the backbone region of the guide RNA may also include one or more (1, 2, 3, 4, 5, etc.) mutations compared to the naturally occurring backbone region. For example, in some cases, base pairs may be maintained, while the nucleotides contributing to the base pair from each segment may be different. In some cases, the double-stranded region of a subject guide RNA includes more paired bases, fewer paired bases, smaller bulges, larger bulges, fewer bulges, more bulges, or any advantageous combination thereof, when compared to the naturally occurring double-stranded region (of a naturally occurring guide RNA).
[0355] Examples of various guide RNAs can be found in the art, and in some cases, variations similar to those introduced into Cas9 guide RNAs can also be introduced into the guide RNAs of the present disclosure (e.g., mutations to the dsRNA double-stranded region, extensions at the 5' or 3' ends to add stability to provide for interaction with another protein, etc.). For example, Jinek et al.,Science.2012 Aug 17;337(6096):816-21;Chylinski et al.,RNA Biol.2013 May;10(5):726-37;Ma et al.,Biomed Res Int.2013;2013:270805;Hou et al.,Proc Natl Acad Sci US A.2013 Sep 24;110(39):15644-9;Jinek et al.,Elife.2013;2:e00471;Pattanayak et al.,Nat Biotechnol.2013 Sep;31(9):839-43;Qi et al,Cell.2013 Feb 28 ;152(5):1173-83 ;Wang et al.,Cell.2013 May 9;153(4):910-8;Auer et al.,Genome Res.2013 Oct 31;Chen et al.,Nucleic Acids Res.2013 Nov 1 ;41(20):el9;Cheng et al.,Cell Res.2013 Oct;23(10):1163-71;Cho et al.,Genetics.2013 Nov;195(3):1177-80;DiCarlo et al.,Nucleic Acids Res.2013 Apr;41(7):4336-43;Dickinson et al.,Nat Methods.2013 Oct;10(10):1028-34;Ebina et al.,Sci Rep.2013;3:2510;Fujii et.al,Nucleic Acids Res.2013 Nov l;41(20):el87;Hu et al.,Cell Res.2013 Nov;23(ll):1322-5;Jiang et al.,Nucleic Acids Res.2013 Nov l;41(20):el88;Larson et al.,Nat Protoc.2013 Nov;8(l l):2180-96;Mali et.at.,Nat Methods.2013 Oct;10(10):957-63;Nakayama et al.,Genesis.2013 Dec;51(12):835-43;Ran et al.,Nat Protoc.2013 Nov;8(l l):2281-308;Ran et al.,Cell.2013 Sep 12;154(6):1380-9;Upadhyay et al.,G3(Bethesda).2013 Dec 9;3(12):2233-8;Walsh et al.,Proc Natl Acad Sci U S A.2013 Sep 24;110(39):15514-5;Xie et al.,Mol Plant.2013 Oct 9;Yang et al.,Cell.2013 Sep 12;154(6):1370-9;Briner et al.,Mol Cell.2014 Oct 23;56(2):333-9; and U.S. patents and patent applications: 8,906,616; 8,895,308; 8,889,418; 8,889,356; 8,871,445; 8,865,406; 8,795,965; 8,771,945; 8,697,359; 20140068797; 20140170753; 20140179006; 20140179770; 20140186843; 2014 0186919;20140186958;20140189896;20140227787;20140234972;20140242664;20140242699;20140242700;20140242702;20140248702;20140256046;20140273037;20140273226;20140273230;20140273231;201402732 32;20140273233;20140273234;20140273235;20140287938;20140295556;20140295557;20140298547;20140304853;20140309487;20140310828;20140310830;20140315985;20140335063;20140335620;20140342456;20 140342457; 20140342458; 20140349400; 20140349405; 20140356867; 20140356956; 20140356958; 20140356959; 20140357523; 20140357530; 20140364333; and 20140377868 (all of which are incorporated by reference herein in their entirety).
[0356] Guide RNA modifications In one embodiment, the guide RNA contemplated herein comprises a non-naturally occurring nucleic acid and / or a non-naturally occurring nucleotide and / or a nucleotide analog, and / or a chemical modification. A non-naturally occurring nucleic acid can include, for example, a mixture of naturally occurring and non-naturally occurring nucleotides. The non-naturally occurring nucleotide and / or nucleotide analog can be modified at the ribose, phosphate, and / or base. In an embodiment of the invention, the guide RNA component nucleic acid comprises a ribonucleotide and a non-ribonucleotide. In one such embodiment, the guide RNA component comprises one or more ribonucleotides and one or more deoxyribonucleotides. In an embodiment of the invention, the guide RNA (including pegylated RNA) component comprises one or more non-naturally occurring nucleotides or nucleotide analogs, such as a nucleotide with a phosphorothioate bond, a locked nucleic acid (LNA) nucleotide comprising a methylene bridge between the 2' and 4' carbons of the ribose ring, or a bridged nucleic acid (BNA).
[0357] Other examples of modified nucleotides include 2'-O-methyl, 2'-deoxy, or 2'-fluoro analogs. Further examples of modified bases include, but are not limited to, 2-aminopurine, 5-bromo-uridine, pseudouridine, inosine, and 7-methylguanosine. Examples of coRNA chemical modifications include, but are not limited to, the incorporation of 2'-O-methyl (M), 2'-O-methyl 3' phosphorothioate (MS), S-constrained ethyl (cEt), or 2'-O-methyl 3' thio PACE (MSP) at one or more terminal nucleotides. Such chemically modified oRNA components may include increased stability and increased activity compared to unmodified oRNA components, although on-target versus off-target specificity cannot be predicted. (Hendel,2015,Nat Biotechnol.33(9):985-9,doi:10.1038 / nbt.3290,published online 29 June 2015 Ragdarm et al.,0215,PNAS,E7110-E7111;Allerson et al. al.,J.Med.Chem.2005,48:901-904;Bramsen et al.,Front.Genet.,2012,3:154;Deng et al.,PNAS,2015,112:11870-11875;Sharma et al.,MedChemComm.,2014,5:1454-1471;Hendel et al. al.,Nat.Biotechnol.(2015)33(9):985-989;Li et al.,Nature Biomedical Engineering, 2017, 1,0066 D01:10.1038 / s41551-017-0066). In one embodiment, the 5' and / or 3' ends of the guide RNA (including pegylated RNA) component are modified with a variety of functional moieties, including fluorescent dyes, polyethylene glycol, cholesterol, proteins, or detection tags. (See Kelly et al., 2016, J. Biotech. 233:74-83).In one embodiment, the guide RNA (including pegRNA) component comprises ribonucleotides in the region that binds to the target sequence and one or more deoxyribonucleotides and / or nucleotide analogs in the region that binds to the nucleic acid programmable DNA-binding protein (e.g., Cas9 nickase).
[0358] In one embodiment, deoxyribonucleotides and / or nucleotide analogs are incorporated into the engineered guide RNA component structure. In one embodiment, 3 to 5 nucleotides at either the 3' or 5' end of the guide RNA component are chemically modified. In one embodiment, only minor modifications, such as 2'-F modifications, are introduced in the seed region. In one embodiment, 2'-F modifications are introduced at the 3' end of the guide RNA component. In one embodiment, 3 to 5 nucleotides at the 5' and / or 3' end of the reRNA component are chemically modified with 2'-O-methyl (M), 2'-O-methyl 3' phosphorothioate (MS), S-constrained ethyl (cEt), or 2'-O-methyl 3' thio PACE (MSP). Such modifications can improve genome editing efficiency (see Hendel et al., Nat. Biotechnol. (2015) 33(9):985-989). In one embodiment, all of the phosphodiester bonds of a guide RNA (including pegylated RNA) component are replaced with phosphorothioate (PS) to enhance the level of gene disruption. In one embodiment, more than five nucleotides at the 5' and / or 3' end of a guide RNA (including pegylated RNA) component are chemically modified with 2'-0-Me, 2'-F, or S-constrained ethyl (cEt). Such chemically modified guide RNA (including pegylated RNA) components can mediate enhanced levels of gene disruption (see Ragdarm et al., 2015, PNAS, E7110-E7111). In an embodiment of the present invention, a guide RNA (including pegylated RNA) component is modified at its 3' and / or 5' end to include chemical moieties. Such moieties include, but are not limited to, amine, azide, alkyne, thio, dibenzocyclooctyne (DBCO), or rhodamine. In certain embodiments, the chemical moiety is conjugated to the guide RNA (including pegylated RNA) component by a linker, e.g., an alkyl chain. In one embodiment, the chemical moiety of the modified nucleic acid component can be used to attach the guide RNA (including pegylated RNA) component to another molecule, e.g., DNA, RNA, a protein, or a nanoparticle.Such chemically modified guide RNA (including pegRNA) components can be used to identify or enrich for cells that have been genetically edited by the gene editing systems described herein.
[0359] Other guide RNA modifications are described in Kim, DY, Lee, JM, Moon, SB et al. Efficient CRISPR editing with a hypercompact Cas12f1 and engineered guide RNAs delivered by adeno-associated virus. Nat Biotechnol 40, 94-102 (2022).
[0360] Thus, in various embodiments of the invention, the guide RNA is modified at one or more positions within the molecule: MS1, an internal penta(uridinylate) (UUUUU) sequence in the tracrRNA; MS2, the 3' end of the crRNA; MS3, the "stem 1" region of the tracrRNA; MS4, the tracrRNA-crRNA complementary region; and MS5, the "stem 2" region of the tracrRNA.
[0361] Various aspects of the present invention provide methods and compositions for improved guide RNA stability through chemical modification.Braasch,D.A.,Jensen,S.,Liu,Y.,Kaur,K.,Arar,K.,White,M.A.,et al.(2003).RNA interference in mammalian cells by chemically-modified RNA.Biochemistry 42,7967-7975.doi:10.1021 / bi0343774.Chiu,Y.L.,and Rana,T.M.(2003).siRNA function in RNAi:a chemical modification analysis.RNA 9,1034-1048.doi:10.1261 / rna.5103703.Behlke,M.A.(2008).Chemical modification of siRNAs for in vivo use.Oligonucleotides18,305-319.doi:10.1089 / oli.2008.0164.Bennett,C.F.,and Swayze,E.E.(2010).RNA targeting therapeutics:molecular mechanisms of antisense oligonucleotides as a therapeutic platform.Annu.Rev.Pharmacol.Toxicol.50,259-293.doi:10.1146 / annurev.pharmtox.010909.105654.Deleavey,G.F.,and Damha,M.J.(2012).Designing chemically modified oligonucleotides for targeted gene silencing.Chem.Biol.19,937-954.doi:10.1016 / j.chembiol.2012.07.011.Lennox,K.A.,and Behlke,M.A.(2020).Chemical modifications in RNA interference and CRISPR / Cas genome editing reagents.Methods Mol.Biol.2115,23-55.doi:10.1007 / 978-1-0716-0290-4_2.。
[0362] For example, Hendel et al. improved guide RNA stability by chemically modifying the gRNA ends to reduce degradation by exonucleases and RNA nucleases. Hendel, A., Bak, RO, Clark, JT, Kennedy, AB, Ryan, DE, Roy, S., et al. (2015a). Chemically modified guide RNAs improve CRISPR-Cas genome editing in human primary cells. Nat. Biotechnol. 33, 985-989. doi:10.1038 / nbt.3290. Chemical modification of gRNAs may enable more efficient and safer gene editing in primary cells, suitable for clinical applications.
[0363] A description of the types of chemical modifications is provided in the following table: Allen, Daniel et al. "Using Synthetically Engineered Guide RNAs to Enhance CRISPR Genome Editing Systems in Mammalian Cells." Frontiers in Genome Editing vol. 2 617910. 28 Jan. 2021, doi:10.3389 / fgeed.2020.617910. [Table 7]
[0364] Thus, in various embodiments of the present invention, the genome editing system comprises a guide RNA and further comprises one or more chemical modifications selected from, but not limited to, the modifications in the table above.
[0365] In exemplary embodiments, chemical modifications to guide RNAs (including PEG RNAs) include modifications on the ribose ring and phosphate backbone of guide RNAs (including PEG RNAs), with modifications at the 2'OH including 2'-O-Me, 2'-F, and 2'F-ANA. A more extensive ribose modification includes 2'F-4'-Cα-OMe, with 2',4'-di-Cα-OMe combining modifications at both the 2' and 4' carbons. Phosphodiester modifications include sulfide-based phosphorothioate (PS) or acetate-based phosphonoacetate modifications. Combinations of ribose and phosphodiester modifications have resulted in formulations such as 2'-O-methyl 3' phosphorothioate (MS), or 2'-O-methyl-3'-thioPACE (MSP), and 2'-O-methyl-3'-phosphonoacetate (MP) RNAs. Locked and unlocked nucleotides, such as locked nucleic acids (LNAs), bridged nucleic acids (BNAs), S-constrained ethyl (cEt), and unlocked nucleic acids (UNAs), are examples of sterically hindered nucleotide modifications. Modifications to the phosphodiester bond between the 2' and 5' carbons of adjacent RNAs (2',5'-RNA) and the butane 4-carbon linkage between adjacent RNAs have been described.
[0366] In certain embodiments, the guide RNA comprises one or more hairpins as shown in the accompanying figures. Preferably, the guide RNA comprises 0 to 10 hairpins. In some embodiments, the guide RNA comprises 1 to 3 hairpins. In some embodiments, the guide RNA comprises 2 hairpins. More preferably, the hairpin comprises 6 to 20 ribonucleotides.
[0367] Modifying sgRNA is also an effective way to improve the efficiency of the CRISPR-Cas system. Kim, Daesik et al. "Evaluating and Enhancing Target Specificity of Gene-Editing Nucleases and Deaminases." Annual review of biochemistry vol. 88(2019):191-220. doi:10.1146 / annurev-biochem-013118-111730. For example, adding a "U4AU6" motif to the end of crRNA has been demonstrated (Bin Moon, Su et al., "Highly efficient genome editing by CRISPR-Cpf1 using CRISPR RNA with a uridinylate-rich 3'-overhang." Nature communications vol. 9, 1, 3651. 7 Sep. 2018, doi:10.1038 / s41467-018-06129-w) or using Pol II-induced truncated pre-tRNA (Zhang, Xuhua et al., "Genetic editing and interrogation with Cpf1 and caged truncated pre-tRNA-like crRNA in mammalian cells." Cell discovery vol. 4, 36, 10 Jul. 2018, doi:10.1038 / s41421-018-0035-0).
[0368] Thus, various embodiments provide modifications to the sgRNA to improve the efficiency of the CRISPR-Cas12a system and modifications to express the crRNA to improve the activity of the CRISPR-Cas12a system.
[0369] Additional embodiments provide guide RNA modifications including, but not limited to, one or more chemical modifications selected from 2'-O-Me, 2'-F, and 2'F-ANA at the 2'OH; 2'F-4'-Cα-OMe and 2',4'-di-Cα-OMe at the 2' and 4' carbons; phosphodiester modifications including sulfide-based phosphorothioate (PS) or acetate-based phosphonoacetate modifications; combinations of ribose and phosphodiester modifications; locked nucleic acids (LNA), bridged nucleic acids (BNA), S-constrained ethyl (cEt), and unlocked nucleic acids (UNA); modifications to generate phosphodiester bonds between the 2' and 5' carbons of adjacent RNAs (2',5'-RNA); and butane 4-carbon chain linkages between adjacent RNAs.
[0370] In yet other embodiments, the guide RNAs disclosed herein may be modified by introducing additional RNA motifs into the guide RNA, for example, at the 5' and 3' ends of the guide RNA. Such structures may include, but are not limited to, RNA hairpins, RNA step-loops, RNA quadruplex structures, cap structures, and poly(A) tails, or ribozyme functions. Additionally, the guide RNAs may also be modified to include one or more nuclear localization sequences.
[0371] Additional RNA motifs can also improve the function or stability of guide RNAs. Addition of dimerization motifs (e.g., kissing loops or GNRA tetraloop / tetraloop acceptor pairs) at the 5' and 3' ends of guide RNAs can also result in efficient circularization of guide RNAs and improve stability. It is also envisioned that the addition of these motifs can enable physical separation of guide RNA components, for example, separating the Cas12a binding region from the spacer sequence. Short 5' or 3' extensions of guide RNAs that form small toehold hairpins at either or both ends of the guide RNA can also favorably compete for annealing of complementary regions along the length of the guide RNA. Finally, kissing loops can also be used to recruit other RNAs or proteins to the genomic site targeted by the guide RNA.
[0372] Guide RNAs can be further improved through directed evolution in a manner similar to the way protein function can be improved: directed evolution can improve guide RNA function, and / or reduce off-site targeting and / or indels, and / or improve precision editing efficiency.
[0373] The present disclosure contemplates any such methods to further improve the stability and / or functionality of the guide RNAs disclosed herein.
[0374] In some embodiments, the RNA (including guide RNA) used in the compositions of the present disclosure undergoes chemical or biological modification to make them more stable.Exemplary modifications to RNA include base depletion (for example, by deletion or by substituting one nucleotide for another) or base modification, for example, chemical modification of base.The term "chemical modification" as used herein includes modification that introduces chemistry different from that found in naturally occurring RNA, for example, covalent modification, for example, introduction of modified nucleotide, (for example, nucleotide analogue, or the inclusion of a pendant group that is not naturally found in such mRNA molecules).
[0375] Other suitable polynucleotide modifications that can be incorporated into RNA used in the compositions of the present disclosure include 4'-thio modified bases: 4'-thio-adenosine, 4'-thio-guanosine, 4'-thio-cytidine, 4'-thio-uridine, 4'-thio-5-methyl-cytidine, 4'-thio-pseudouridine, and 4'-thio-2-thiouridine, pyridin-4-one ribonucleosides, 5-aza-uridine, 2-thio-5-aza-uridine, 2-thiouridine, 4-thio-pseudouridine, 2-thio-pseudouridine, 5-hydroxyuridine, 3-methyluridine, 4'-thio-5-methyl-cytidine, 4'-thio-pseudouridine, and 4'-thio-2-thiouridine. uridine, 5-carboxymethyl-uridine, 1-carboxymethyl-pseudouridine, 5-propynyl-uridine, 1-propynyl-pseudouridine, 5-taurinomethyluridine, 1-taurinomethyl-pseudouridine, 5-taurinomethyl-2-thio-uridine, 1-taurinomethyl-4-thio-uridine, 5-methyl-uridine, 1-methyl-pseudouridine, 4-thio-1-methyl-pseudouridine, 2-thio-1-methyl-pseudouridine, 1-methyl-1-deaza-pseudouridine, 2-thio-1-methyl-1-deaza-pseudouridine Uridine, Dihydrouridine, Dihydropseudouridine, 2-thio-dihydrouridine, 2-thio-dihydropseudouridine, 2-methoxyuridine, 2-methoxy-4-thio-uridine, 4-methoxy-pseudouridine, 4-methoxy-2-thio-pseudouridine, 5-aza-cytidine, Pseudoisocytidine, 3-methyl-cytidine, N4-acetylcytidine, 5-formylcytidine, N4-methylcytidine, 5-hydroxymethylcytidine, 1-methyl-pseudoisocytidine, Pyrrolo-cytidine, Pyrrolo-pseudoisocytidine, 2-thio -cytidine, 2-thio-5-methyl-cytidine, 4-thio-pseudoisocytidine, 4-thio-1-methyl-pseudoisocytidine, 4-thio-1-methyl-1-deaza-pseudoisocytidine, 1-methyl-1-deaza-pseudoisocytidine, zebularine, 5-aza-zebularine, 5-methyl-zebularine, 5-aza-2-thio-zebularine, 2-thio-zebularine, 2-methoxy-cytidine, 2-methoxy-5-methyl-cytidine, 4-methoxy-pseudoisocytidine, 4-methoxy-1-methyl-pseudoisocytidine, 2-aminopurine, 2,6-diaminopurine, 7-deaza-adenine, 7-deaza-8-aza-adenine, 7-deaza-2-aminopurine, 7-deaza-8-aza-2-aminopurine, 7-deaza-2,6-diaminopurine, 7-deaza-8-aza-2,6-diaminopurine, 1-methyladenosine, N6-methyladenosine, N6-isopentenyladenosine, N6-(cis-hydroxyisopentenyl)adenosine, 2-methylthio-N6-(cis-hydroxyisopentenyl)adenosine, N6-glycinylcarbamoyladenosine, N6-threonylcarbamoyladenosine, 2-methylthio-N6-threonylcarbamoyladenosine, N6,N6-dimethyladenosine, 7-methyladenine, 2-methylthio-adenine, and 2- These include, but are not limited to, methoxy-adenine, inosine, 1-methyl-inosine, wyosine, wybutosine, 7-deaza-guanosine, 7-deaza-8-aza-guanosine, 6-thio-guanosine, 6-thio-7-deaza-guanosine, 6-thio-7-deaza-8-aza-guanosine, 7-methyl-guanosine, 6-thio-7-methyl-guanosine, 7-methylinosine, 6-methoxy-guanosine, 1-methylguanosine, N2-methylguanosine, N2,N2-dimethylguanosine, 8-oxo-guanosine, 7-methyl-8-oxo-guanosine, 1-methyl-6-thio-guanosine, N2-methyl-6-thio-guanosine, and N2,N2-dimethyl-6-thio-guanosine, and combinations thereof. The term modification also includes, for example, the incorporation of non-nucleotide linkages or modified nucleotides into the mRNA sequences of the invention (e.g., modifications to one or both of the 3' and 5' ends of an mRNA molecule encoding a functional protein or enzyme). Such modifications include the addition of bases to the mRNA sequence (e.g., inclusion of a polyA tail or a longer polyA tail), alteration of the 3' UTR or 5' UTR, complexing the mRNA with an agent (e.g., a protein or a complementary nucleic acid molecule), and the inclusion of elements that alter the structure of the RNA molecule (e.g., form secondary structures).
[0376] In some embodiments, the RNA (e.g., guide RNA) includes a 5' cap structure. The 5' cap is typically added as follows: first, an RNA terminal phosphatase removes one of the terminal phosphate groups from the 5' nucleotide, leaving two terminal phosphates; then, guanosine triphosphate (GTP) is added to the terminal phosphate via a guanylyltransferase to generate a 5'5'5 triphosphate bond; and then, the 7-nitrogen of guanine is methylated by a methyltransferase. Examples of cap structures include, but are not limited to, mG(5')ppp(5')(A), G(5')ppp(5')A, and G(5')ppp(5')G. The naturally occurring cap structure is 7-methylguanosine, which is linked via a triphosphate bridge to the 5'-end of the first transcribed nucleotide, resulting in a dinucleotide cap of mG(5')ppp(5')N (where N is any nucleoside). In vivo, the cap is added enzymatically. The cap is added in the nucleus and is catalyzed by the enzyme guanylyltransferase. Addition of the cap to the 5' end of RNA occurs immediately after the initiation of transcription. The terminal nucleoside is typically guanosine, in the reverse orientation relative to all other nucleotides, i.e., G(5')ppp(5')GpNpNp.
[0377] Additional cap analogs include, but are not limited to, chemical structures selected from the group consisting of m7GpppG, m7GpppA, m7GpppC; unmethylated cap analogs (e.g., GpppG); dimethylated cap analogs (e.g., m2,7GpppG), trimethylated cap analogs (e.g., m2,2,7GpppG), dimethylated symmetric cap analogs (e.g., m7Gpppm7G), or anti-reverse cap analogs (e.g., ARCA; m7,2'OmeGpppG, m72'dGpppG, m7,3'OmeGpppG, m7,3'dGpppG and their tetraphosphate derivatives) (see, e.g., Jemielity, J. et al., "Novel 'anti-reverse' cap analogs with superior translational properties", RNA, 9:1108-1122 (2003)).
[0378] Typically, the presence of a "tail" serves to protect RNA (e.g., guide RNA) from exonuclease degradation. PolyA or polyU tails are thought to stabilize natural messenger and synthetic sense RNA. Therefore, in certain embodiments, long polyA or polyU tails can be added to RNA molecules, thereby making the RNA more stable. PolyA or polyU tails can be added using a variety of art-recognized techniques. For example, long polyA tails can be added to synthetic or in vitro transcribed RNA using polyA polymerase (Yokoe, et al. Nature Biotechnology. 1996;14:1252-1256). Transcription vectors can also encode long polyA tails. PolyA tails can also be added by transcription directly from PCR products. Poly A can also be ligated to the 3' end of the sense RNA with RNA ligase (see, e.g., Molecular Cloning A Laboratory Manual, 2nd Ed., ed. by Sambrook, Fritsch and Maniatis (Cold Spring Harbor Laboratory Press: 1991 edition)).
[0379] Typically, the length of the poly-A or poly-U tail can be at least about 10, 50, 100, 200, 300, 400, or at least 500 nucleotides. In some embodiments, the poly-A tail at the 3' end of the mRNA typically contains about 10-300 adenosine nucleotides (e.g., about 10-200 adenosine nucleotides, about 10-150 adenosine nucleotides, about 10-100 adenosine nucleotides, about 20-70 adenosine nucleotides, or about 20-60 adenosine nucleotides). In some embodiments, the mRNA includes a 3' poly(C) tail structure. A suitable poly-C tail at the 3' end of an mRNA typically contains about 10 to 200 cytosine nucleotides (e.g., about 10 to 150 cytosine nucleotides, about 10 to 100 cytosine nucleotides, about 20 to 70 cytosine nucleotides, about 20 to 60 cytosine nucleotides, or about 10 to 40 cytosine nucleotides). The poly-C tail can be added to or replace a poly-A or poly-U tail.
[0380] RNAs (e.g., Cas12a guide RNAs) according to the present disclosure can be synthesized according to any of a variety of known methods. For example, RNAs according to the present disclosure can be synthesized via in vitro transcription (IVT). Briefly, IVT is typically performed using a linear or circular DNA template containing a promoter, a pool of ribonucleotide triphosphates, a buffer system that may include DTT and magnesium ions, and an appropriate RNA polymerase (e.g., T3, T7, or SP6 RNA polymerase), DNAse I, pyrophosphatase, and / or an RNAse inhibitor. The exact conditions will vary depending on the particular application.
[0381] In certain embodiments, the guide RNA may contain MS2 modifications as specific RNA hairpin structures that are naturally recognized by certain MS2 binding proteins. This domain may help stabilize the guide RNA and improve editing efficiency. The present disclosure contemplates other similar modifications. Reviews of other such MS2-like domains are provided in the art, e.g., Johansson et al., "RNA recognition by the MS2 phage coat protein," Sem. Virol., 1997, Vol. 8(3):176-185; Delebecque et al., "Organization of intracellular reactions with rationally designed RNA assemblies," Science, 2011, Vol. 333:470-474; Mali et al., "Cas9 transcriptional activators for target specificity screening and paired nickases for cooperative genome engineering," Nat. Biotechnol., 2013, Vol. 31:833-838; and Zalatan et al., "Engineering complex synthetic transcriptional programs with CRISPR RNA scaffolds," Cell, 2015, Vol. 160:339-350 (each of which is incorporated by reference herein in its entirety). Other systems include the PP7 hairpin, which specifically recruits the PCP protein, and the "com" hairpin, which specifically recruits the Com protein. See Zalatan et al. The nucleotide sequence of the MS2 hairpin (or equivalently referred to as the "MS2 aptamer") is GCCAACATGAGGATCACCCATGTCTGCAGGGCC (SEQ ID NO: 1391).
[0382] E. Cas12a (or Cas V) editing system The present disclosure relates to a novel genome editing system. In an exemplary embodiment, the editing system comprises: (a) one or more polypeptide sequences comprising at least 45%, 46%, 47%, 48%, 49%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, or 65% sequence identity to any one of the sequences selected from SEQ ID NO: 334 (No. ID405), SEQ ID NO: 58 (No. ID414), or SEQ ID NO: 564 (No. ID418), SEQ ID NO: 335 (No. ID406), SEQ ID NO: 331 (No. ID411), SEQ ID NO: 20 (No. ID415), and SEQ ID NO: 445 (No. ID419); and (b) one or more polynucleotide sequences comprising a guide RNA, wherein the guide RNA comprises a sequence complementary to that of the targeted polynucleotide sequence; Includes.
[0383] In other aspects, the Cas12a-based gene editing system can include one or more additional accessory proteins with genome modification functions, including recombinases, invertases, nucleases, polymerases, ligases, deaminases, reverse transcriptases, or epigenetic modification functions. In various embodiments, the accessory proteins can be provided separately. In other embodiments, the accessory proteins can be fused to Cas12a, optionally using a linker.
[0384] In various embodiments, the genome editing system may include a guide RNA that hybridizes to one or more targeted polynucleotide sequences. In a preferred embodiment, the guide RNA of the genome editing system includes 12 to 40 nucleotides.
[0385] In various embodiments, the genome editing system includes a targeted polynucleotide sequence comprising one or more protospacer adjacent motif (PAM) recognition domains selected from 5'-TTTN-3', 5'-TTN-3', 5'-TNN-3', 5'-TTV-3', or 5'-TTTV-3' (where N = A, T, C, or G and V = A, C, or G). In additional embodiments, the targeted polynucleotide sequence comprises one or more relaxed PAM recognition domains. Jacobsen, Thomas et al. "Characterization of Cas12a nucleases reveals diverse PAM profiles between closely-related orthologs." Nucleic acids research vol. 48, 10 (2020): 5624-5638. doi:10.1093 / nar / gkaa272. Previous studies have demonstrated that non-canonical PAMs (e.g., ATTA, CTTA, GTTA, and TCTA) address the limitation of the requirement for an extended TTTV protospacer adjacent motif (PAM) by expanding targeting range. Kleinstiver, Benjamin P et al. "Engineered CRISPR-Cas12a variants with increased activities and improved targeting ranges for gene, epigenetic, and base editing." Nature biotechnology vol. 37, 3 (2019): 276-282. doi:10.1038 / s41587-018-0011-0. Most Cpf1 nucleases require a thymine-rich PAM. Different studies have demonstrated increased Cpf1 targeting range using in vitro and in vivo (E. coli) PAM identification assays. Zhang, Xiaochun, et al. “Multiplex gene regulation by CRISPR-ddCpf1.”Cell discovery 3.1(2017):1-9.Two Cpf1 endonucleases, AsCpf1 and LbCpf1, require TTTV as the PAM sequence, where V can be an A, C, or G nucleotide. Mutations at positions S542R / K607R and S542R / K548V / N552R generate AsCpf1 variants that can recognize TYCV and TATV PAMs (where Y can be C or T), respectively. Gao, Linyi, et al. "Engineered Cpf1 variants with altered PAM specificities." Nature biotechnology 35.8(2017):789-792. AsCpf1 showed increased activity with TTTV PAM and decreased activity with TTTT PAM. Kim, Hui K., et al. “In vivo high-throughput profiling of CRISPR-Cpf1 activity.” Nature methods 14.2 (2017): 153-159.
[0386] Therefore, it is within the scope of the present disclosure to devise a Cas12a editing system that recognizes modified PAM recognition domains for genome editing. In a preferred embodiment, the Cas12a polypeptide recognizes one or more non-canonical PAM sequences in the targeted polynucleotide sequence, and the PAM is located upstream of the crRNA complementary DNA sequence on the non-target strand. In a related embodiment, the gRNA has an 8-nucleotide seed sequence located at the 5' end of the spacer, close to the PAM sequence on the targeted polynucleotide sequence. Preferably, the Cas12a polypeptide cleaves the targeted polynucleotide sequence about 20 nucleotides upstream of the PAM sequence.
[0387] In further embodiments, the one or more polypeptide sequences and the one or more polynucleotide sequences comprising the cognate guide RNA of the genome editing system form a ribonucleoprotein complex.
[0388] In various embodiments, the one or more polypeptide sequences of the genome editing system include: one or more α-helical recognition lobes (REC) and nuclease lobes (NUC); Wedge (WED), α-helix recognition lobe (REC), PAM interaction (PI), RuvC nuclease, bridge helix (BH) and NUC domain; or One or more domains selected from the RuvC, REC, WED, BH, PI, and NUC domains Includes.
[0389] Preferably, the REC lobe comprises the REC1 and REC2 domains. More preferably, the NUC lobe comprises the RuvC, PI, WED, and bridge helix (BH) domains. The RuvC domain also comprises the subdomains RuvCI, RuvCII, and RuvCIII. In a preferred embodiment, the RuvCIII domain is located at the C-terminus.
[0390] In various embodiments, one or more polypeptide sequences of the genome editing system lack an HNH endonuclease domain.
[0391] Without being bound by theory, the Cas12a genome editing system is characterized as a class 2, type V Cas endonuclease.
[0392] In various embodiments, the molecular weight of the Cas12a nuclease is characterized by its molecular weight to be between about 50 kDa and 100 kDa, between 100 kDa and 200 kDa, between 200 kDa and 500 kDa.
[0393] In additional embodiments, the polypeptide sequence comprises at least one activity selected from an endonuclease activity, an endoribonuclease activity, or an RNA-guided DNase activity. In such embodiments, the cognate guide RNA and Cas12a protein modify a targeted polynucleotide sequence in the host cell genome. In certain instances, the targeted polynucleotide sequence is modified by an insertion, deletion, or alteration of one or more base pairs in the targeted polynucleotide sequence in the host cell genome.
[0394] In related embodiments, the genome editing system is characterized by improved efficiency and accuracy of site-specific integration. Preferably, the efficiency and accuracy of site-specific integration enabled by the genome editing system is improved by cohesive overhangs on the donor nucleic acid sequence. In certain embodiments, the targeted polynucleotide sequence is double-stranded and contains a 5' overhang, the overhang preferably comprising 5 nucleotides.
[0395] In various embodiments, the cut or nicking in the targeted polynucleotide sequence is preferably repaired by the endogenous DNA polymerase repair mechanism present in cells.In some embodiments, the method provides introducing a donor DNA sequence under conditions that allow the targeted polynucleotide sequence to be edited by homologous recombination repair.Preferably, the Cas12a genome editing system is characterized as exhibiting reduced specificity, such as off-target effects, compared to Cas9.More preferably, the Cas12a system comprises improved activity of at least 1, 1.5, 2, 2.5, 3, 3.5, 4, 4.5, 5, 5.5, 6, 6.5, 7, 7.5, 8, 8.5, 9, 9.5, 10 times or more.
[0396] In various embodiments, the RuvC domain of the Cas12a polypeptide of the Cas12a genome editing system, comprising RuvC subdomains I, II, and II, cleaves the targeted polynucleotide sequence and / or the non-target DNA strand. Preferably, the genome editing system expresses multiple copies of the guide RNA in the host cell of interest.
[0397] In various other embodiments, the polypeptide of the genome editing system comprises one or more mutations. Preferably, the mutations are selected from one or more domains selected from the RuvC, REC, WED, BH, PI, and NUC domains. More preferably, the mutations encode a nuclease-deficient polypeptide. In various embodiments, the genome editing system comprises a fusion of one or more deaminases to the nuclease-deficient polypeptide. Preferably, the one or more deaminases of the genome editing system are selected from adenine deaminase or cytosine deaminase. The use of cytidine deaminase and adenosine deaminase base editing is disclosed in U.S. Patent No. 9,840,699. One approach is to generate a Cas12a fusion protein (preferably an inactive or nickase variant) and a base-editing enzyme or an active domain of a base-editing enzyme. Cytidine deaminase and adenosine deaminase base editing are disclosed in U.S. Patent No. 9,840,699. In various embodiments, the composition comprises contacting a targeted polynucleotide sequence with a fusion protein comprising Cas12a and one or more base-editing polypeptides, e.g., deaminases; and a gRNA that targets the fusion protein to the targeted polynucleotide sequence of a DNA strand. Thus, fusing one or more deaminases to the nuclease-deficient polypeptide of the Cas12a genome editing system enables base editing on DNA and / or RNA. In select embodiments, the system modifies one or more nucleobases on DNA and RNA. In related embodiments, the system enables multiplexed gene editing. Preferably, the genome editing system includes a single crRNA. More preferably, the system enables simultaneous targeting of multiple genes.
[0398] Recent studies have demonstrated near-100% on-target gene editing efficiency of Cas12a through modifications to the NLS framework. Luk et al., GEN Biotechnology. Jun 2022, pp. 271–284. http: / / doi.org / 10.1089 / genbio.2022.0003. Previous studies have also demonstrated that NLS-optimized SpCas9-based prime editors improve genome editing efficiency. Liu, Pengpeng et al., “Improved prime editors enable pathogenic allele correction and cancer modeling in adult mice.” Nature Communications, vol. 12, 1, 2121, 9 Apr 2021, doi:10.1038 / s41467-021-22295-w.
[0399] In yet other embodiments, the Cas12a polypeptide is operably linked to a nuclear localization signal (NLS). Preferably, the Cas12a polypeptide comprises an NLS at the N-terminus, C-terminus, or both, or multiple NLSs on the Cas12a polypeptide. In some embodiments, the polypeptide linked to the NLS further comprises a crRNA to form a ribonucleoprotein complex. In some embodiments, the polypeptide comprises one or more NLS repeats at either the N- or C-terminus of the polypeptide.
[0400] In selected embodiments, one or more polypeptide sequences of the genome editing system comprise a modification, the modification comprising a nuclease-deficient polypeptide (dCas). In a related embodiment, the guide RNA of the genome editing system comprises a prime editing guide RNA (pegRNA). Preferably, the pegRNA of the genome editing system hybridizes to the targeted polynucleotide sequence and acts as a primer for one or more reverse transcriptases. More preferably, the pegRNA of the genome editing system binds to the nicked strand for initiation of reverse transcriptase-mediated repair using a repair template.
[0401] In various additional embodiments, the nuclease-deficient polypeptide of the genome editing system comprises nickase activity. Preferably, the genome editing system comprises a fusion of one or more reverse transcriptases to a nuclease-deficient Cas (dCas). In certain examples, the fusion of one or more reverse transcriptases comprises:
[0402] Moloney murine leukemia virus (M-MLV). In some embodiments, the guide RNA or pegRNA comprises or consists of an extended single guide RNA containing a primer binding site (PBS) and a reverse transcriptase (RT) template sequence.
[0403] The Cas12a genome editing system includes improved genome editing features selected from efficiency, specificity, accuracy, intended / unintended editing, and indels compared to Cas9. Accordingly, an object of the present invention is to reduce off-target effects in host cells compared to SpCas9 with comparable endonuclease activity in host cells.
[0404] Optional components / modifications Donor Template In one embodiment, the compositions and systems herein may further comprise one or more donor templates for use in editing. In some cases, the donor template may comprise one or more polynucleotides. In certain cases, the donor template may comprise a coding sequence for one or more polynucleotides. The donor template may be a DNA template. It may be single-stranded or double-stranded. It may also be circular, single-stranded, or double-stranded. It may also be linear, single-stranded, or double-stranded. Without being bound by theory, the donor template may be integrated into the genome after targeted cleavage by the Cas12a gene editing system described herein via cellular repair mechanisms, including HDR and NHEJ.
[0405] A donor template can be used to edit a target polynucleotide. In some cases, the donor polynucleotide contains one or more mutations introduced into the target polynucleotide. Examples of such mutations include substitutions, deletions, insertions, or a combination thereof. The mutations can cause a shift in the open reading frame on the target polynucleotide. In some cases, the donor template modifies a stop codon in the target polynucleotide. For example, the donor template can correct a premature stop codon. Correction can be achieved by deleting the stop codon or introducing one or more mutations into the stop codon. In other exemplary embodiments, the donor polynucleotide inserts or repairs a functional copy of a gene, or a functional fragment thereof, or a functional regulatory sequence or functional fragment of a regulatory sequence, to address, for example, loss-of-function mutations, deletions, or translocations that may occur in a given disease environment. A functional fragment refers to less than an entire copy of a gene by providing sufficient nucleotide sequence to restore functionality of a wild-type gene or non-coding regulatory sequence (e.g., a sequence encoding a long non-coding RNA). In certain exemplary embodiments, the systems disclosed herein can be used to replace a single allele of a defective gene or defective fragment thereof. In another exemplary embodiment, the systems disclosed herein can be used to replace both alleles of a defective gene or defective gene fragment. A "defective gene" or "defective gene fragment" is a gene or portion of a gene that, when expressed, fails to produce a functional protein or non-coding RNA with the functionality of the corresponding wild-type gene. In certain exemplary embodiments, these defective genes can be associated with one or more disease phenotypes. In certain exemplary embodiments, the defective gene or gene fragment is not replaced, but the systems described herein are used to insert a donor template encoding a gene or gene fragment that complements or overrides the defective gene expression, such that cell types associated with the defective gene expression are eliminated or transformed into a different or desired cell phenotype.
[0406] In embodiments of the present invention, the donor template may include, but is not limited to, a gene or gene fragment encoding a protein or RNA transcript to be expressed, a regulatory element, a repair template, etc. In accordance with the present invention, the donor template may include left and right terminal sequence elements that function with the transposition element to mediate insertion.
[0407] In certain cases, the donor template manipulates a splice site on the target polynucleotide. In some instances, the donor template disrupts the splice site. Disruption can be achieved by inserting a polynucleotide into the splice site and / or introducing one or more mutations into the splice site. In certain instances, the donor template can repair the splice site. For example, the polynucleotide can include a splice site sequence.
[0408] The donor template to be inserted can be from 10 base pairs or nucleotides to 50 kb in length, e.g., from 50 to 40 kb, 100 and 30 kb, 100 to 10,000, 100 to 300, 200 to 400, 300 to 500, 400 to 600, 500 to 700, 600 to 800, 700 to 900, 800 to 1,000, 900 to 1,100, 1,000 to 1,200, 1,100 to 1,300, 1,200 to 1,400, 1,300 to 1,500, 1,400 to 1,600, 1,500 to 1,700, 600 to 1,800, 1,700 to 1,900, or 1,800 to 2,000 base pairs (bp) or nucleotides in size.
[0409] In some embodiments, the heterologous nucleic acid sequence is a donor DNA template that can be integrated into the host genome via HDR. In other embodiments, the heterologous nucleic acid sequence is a donor DNA template that can be integrated into the host genome via NHEJ.
[0410] In certain embodiments, the heterologous nucleic acid comprises or encodes a donor / template sequence, and the donor / template corrects / repairs / removes a mutation at a target genomic site. For example, the mutation can be a mutated exon in a disease gene.
[0411] In certain embodiments, the donor / template may encode or contain functional DNA elements, such as promoters, enhancers, protein binding sequences, methylation sites, or homology regions to assist in gene editing.
[0412] "Donor DNA" or "donor DNA template" refers to a DNA segment (which may be single-stranded or double-stranded DNA) to be inserted at the site cleaved by a gene editing nuclease (e.g., Cas12a nuclease) (e.g., after dsDNA cleavage, after target DNA nicking, after target DNA double-nicking, etc.). The donor DNA template contains sufficient homology to the genomic sequence at the target site, e.g., 70%, 80%, 85%, 90%, 95%, or 100% homology to the nucleotide sequences flanking the target site (e.g., within about 50 bases or less, e.g., within about 30 bases, within about 15 bases, within about 10 bases, within about 5 bases, or within about 5 bases of the target site), or immediately flanking the target site, to support homologous recombination repair between it and the genomic sequence to which it is homologous. In the case of repair by NHEJ, the donor DNA template does not need to be homologous to the site targeted for editing.
[0413] Approximately 25, 50, 100, or 200 nucleotides, or greater than 200 nucleotides, of sequence homology between the donor DNA template and the genomic sequence (or any integer value between 10 and 200 nucleotides, or greater) can support homology-directed repair. The donor DNA template can be of any length, e.g., 50 nucleotides or more, 100 nucleotides or more, 250 nucleotides or more, 500 nucleotides or more, 1000 nucleotides or more, 5000 nucleotides or more, etc. Suitable donor DNA templates can be 50 nucleotides to 100 nucleotides, 100 nucleotides to 500 nucleotides, 500 nucleotides to 1000 nucleotides, 1000 nucleotides to 5000 nucleotides, or 5000 nucleotides to 10,000 nucleotides, or greater than 10,000 nucleotides in length.
[0414] As described above, in some embodiments, the donor DNA template comprises a first homology arm and a second homology arm. The first homology arm is located at or near the 5' end of the donor DNA and comprises a nucleotide sequence that is at least partially complementary to a first nucleotide sequence in the target nucleic acid. The second homology arm is located at or near the 3' end of the donor DNA and comprises a nucleotide sequence that is at least partially complementary to a second nucleotide sequence in the target nucleic acid. The first and second homology arms can each independently have a length of about 10 nucleotides to 400 nucleotides; for example, 10 nucleotides (nt) to 15 nt, 15 nt to 20 nt, 20 nt to 25 nt, 25 nt to 30 nt, 30 nt to 35 nt, 35 nt to 40 nt, 40 nt to 45 nt, 45 nt to 50 nt, 50 nt to 75 nt, 75 nt to 100 nt, 100 nt to 125 nt, 125 nt to 150 nt, 150 nt to 175 nt, 175 nt to 200 nt, 200 nt to 225 nt, 225 nt to 250 nt, 250 nt to 275 nt, 275 nt to 300 nt, 325 nt to 350 nt, 350 nt to 375 nt, or 375 nt to 400 nt.
[0415] In certain embodiments, a donor DNA template is used to edit a target nucleotide sequence. In certain embodiments, the donor DNA template includes one or more mutations to be introduced into the target polynucleotide. Examples of such mutations include substitutions, deletions, insertions, or a combination thereof. In certain embodiments, the mutations cause a shift in the open reading frame on the target polynucleotide. In certain embodiments, the donor polynucleotide modifies a stop codon in the target polynucleotide. In certain embodiments, the donor polynucleotide corrects a premature stop codon. Correction can be achieved by deleting the stop codon or by introducing one or more sequence changes to change the stop codon to a codon. In certain embodiments, the donor polynucleotide inserts or repairs a functional copy of a gene, or a functional fragment thereof, or a functional regulatory sequence or functional fragment of a regulatory sequence, to address loss-of-function mutations, deletions, or translocations that may occur in certain disease environments, for example. A functional fragment comprises less than an entire copy of a gene, but provides sufficient nucleotide sequence to r...
Claims
1. A polypeptide or an isolated polypeptide, (a) a polypeptide having the amino acid sequence of one of SEQ ID NO:334 (No. ID405), SEQ ID NO:58 (No. ID414), or SEQ ID NO:564 (No. ID418), SEQ ID NO:335 (ID406), SEQ ID NO:331 (No. ID411), SEQ ID NO:20 (No. ID415), or SEQ ID NO:445 (No. ID419); (b) a polypeptide at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8% or at least 99.9% identical to one of SEQ ID NO:334 (No. ID405), SEQ ID NO:58 (No. ID414), or SEQ ID NO:564 (No. ID418), SEQ ID NO:335 (No. ID406), SEQ ID NO:331 (No. ID411), SEQ ID NO:20 (No. ID415), or SEQ ID NO:445 (No. ID419); (c) has at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 100%, at least 101%, at least 102%, at least 103%, at least 104%, at least 105%, at least 106%, at least 107%, at least 108%, at least 109%, at least 110%, at least 111%, at least 112%, at least 113%, at least 114%, at least 115%, at least 116%, at least 117%, at least 118%, at least 119 ... a polypeptide which is at least 99%%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8% or at least 99.9% identical to SEQ ID NO:334 (No. ID405), SEQ ID NO:58 (No. ID414), or SEQ ID NO:564 (No. ID418), SEQ ID NO:335 (No. ID406), SEQ ID NO:331 (No. ID411), SEQ ID NO:20 (No. ID415), or SEQ ID NO:445 (No. ID419); (d) a sequence identical to or different from one of SEQ ID NO:334 (No. ID405), SEQ ID NO:58 (No. ID414), or SEQ ID NO:564 (No. ID418), SEQ ID NO:335 (No. ID406), SEQ ID NO:331 (No. ID411), SEQ ID NO:20 (No. ID415), or SEQ ID NO:445 (No. ID419), by at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 100%, at least 101%, at least 102%, at least 103%, at least 104%, at least 105%, at least 106%, at least 107%, at least 108%, at least 109%, at least 110%, at least 111%, at least 112%, at least 113%, at least 114%, at least 115%, at least 116%, at least 117%, at least 118%, at least 119%, at least 119%, at least 119%, at least 119%, at least 120%, at least 121%, at least 122%, at least 123%, at least 124%, at least 125%, at least 126%, at least 127%, at least 128%, at least 129%, at least 129%. 7%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8% or at least 99.9% identical polypeptide compared to SEQ ID NO:334 (No. ID405), SEQ ID NO:58 (No. ID414), or SEQ ID NO:564 (No. ID418), SEQ ID NO:335 (No. ID406), SEQ ID NO:331 (No. ID411), SEQ ID NO:20 (No. ID415), or SEQ ID NO:445 (No. ID419), i) an alanine (A) residue of SEQ ID NO:334 (No. ID405), SEQ ID NO:58 (No. ID414), or SEQ ID NO:564 (No. ID418), SEQ ID NO:335 (No. ID406), SEQ ID NO:331 (No. ID411), SEQ ID NO:20 (No. ID415), or SEQ ID NO:445 (No. ID419), substituted with any one of: arginine (R); asparagine (N); aspartic acid (D); cysteine (C); glutamic acid (E); glutamine (N); glycine (G); histidine (H); isoleucine (I); leucine (L); lysine (K); methionine (M); phenylalanine (F); proline (P); serine (S); threonine (T); tryptophan (W); tyrosine (Y); or valine (V); and / or ii) an arginine (R) residue of SEQ ID NO:334 (No. ID405), SEQ ID NO:58 (No. ID414), or SEQ ID NO:564 (No. ID418), SEQ ID NO:335 (No. ID406), SEQ ID NO:331 (No. ID411), SEQ ID NO:20 (No. ID415), or SEQ ID NO:445 (No. ID419), substituted with any one of alanine (A); asparagine (N); aspartic acid (D); cysteine (C); glutamic acid (E); glutamine (N); glycine (G); histidine (H); isoleucine (I); leucine (L); lysine (K); methionine (M); phenylalanine (F); proline (P); serine (S); threonine (T); tryptophan (W); tyrosine (Y); or valine (V); and / or iii) an asparagine (N) residue of SEQ ID NO:334 (No. ID405), SEQ ID NO:58 (No. ID414), or SEQ ID NO:564 (No. ID418), SEQ ID NO:335 (No. ID406), SEQ ID NO:331 (No. ID411), SEQ ID NO:20 (No. ID415), or SEQ ID NO:445 (No. ID419), substituted with any one of alanine (A); arginine (R); aspartic acid (D); cysteine (C); glutamic acid (E); glutamine (N); glycine (G); histidine (H); isoleucine (I); leucine (L); lysine (K); methionine (M); phenylalanine (F); proline (P); serine (S); threonine (T); tryptophan (W); tyrosine (Y); or valine (V); and / or iv) an aspartic acid (D) residue of SEQ ID NO:334 (No. ID405), SEQ ID NO:58 (No. ID414), or SEQ ID NO:564 (No. ID418), SEQ ID NO:335 (No. ID406), SEQ ID NO:331 (No. ID411), SEQ ID NO:20 (No. ID415), or SEQ ID NO:445 (No. ID419), substituted with any one of alanine (A); arginine (R); asparagine (N); cysteine (C); glutamic acid (E); glutamine (N); glycine (G); histidine (H); isoleucine (I); leucine (L); lysine (K); methionine (M); phenylalanine (F); proline (P); serine (S); threonine (T); tryptophan (W); tyrosine (Y); or valine (V); and / or v) a cysteine (C) residue of SEQ ID NO:334 (No. ID405), SEQ ID NO:58 (No. ID414), or SEQ ID NO:564 (No. ID418), SEQ ID NO:335 (No. ID406), SEQ ID NO:331 (No. ID411), SEQ ID NO:20 (No. ID415), or SEQ ID NO:445 (No. ID419), substituted with any one of alanine (A); arginine (R); asparagine (N); aspartic acid (D); glutamic acid (E); glutamine (N); glycine (G); histidine (H); isoleucine (I); leucine (L); lysine (K); methionine (M); phenylalanine (F); proline (P); serine (S); threonine (T); tryptophan (W); tyrosine (Y); or valine (V); and / or vi) a glutamic acid (E) residue of SEQ ID NO:334 (No. ID405), SEQ ID NO:58 (No. ID414), or SEQ ID NO:564 (No. ID418), SEQ ID NO:335 (No. ID406), SEQ ID NO:331 (No. ID411), SEQ ID NO:20 (No. ID415), or SEQ ID NO:445 (No. ID419), substituted with any one of alanine (A); arginine (R); asparagine (N); aspartic acid (D); cysteine (C); glutamine (N); glycine (G); histidine (H); isoleucine (I); leucine (L); lysine (K); methionine (M); phenylalanine (F); proline (P); serine (S); threonine (T); tryptophan (W); tyrosine (Y); or valine (V); and / or vii) a glutamine (N) residue of SEQ ID NO:334 (No. ID405), SEQ ID NO:58 (No. ID414), or SEQ ID NO:564 (No. ID418), SEQ ID NO:335 (No. ID406), SEQ ID NO:331 (No. ID411), SEQ ID NO:20 (No. ID415), or SEQ ID NO:445 (No. ID419), substituted with any one of alanine (A); arginine (R); asparagine (N); aspartic acid (D); cysteine (C); glutamic acid (E); glycine (G); histidine (H); isoleucine (I); leucine (L); lysine (K); methionine (M); phenylalanine (F); proline (P); serine (S); threonine (T); tryptophan (W); tyrosine (Y); or valine (V); and / or viii) a glycine (G) residue of SEQ ID NO:334 (No. ID405), SEQ ID NO:58 (No. ID414), or SEQ ID NO:564 (No. ID418), SEQ ID NO:335 (No. ID406), SEQ ID NO:331 (No. ID411), SEQ ID NO:20 (No. ID415), or SEQ ID NO:445 (No. ID419), substituted with any one of alanine (A); arginine (R); asparagine (N); aspartic acid (D); cysteine (C); glutamic acid (E); glutamine (N); histidine (H); isoleucine (I); leucine (L); lysine (K); methionine (M); phenylalanine (F); proline (P); serine (S); threonine (T); tryptophan (W); tyrosine (Y); or valine (V); and / or ix) a histidine (H) residue of SEQ ID NO:334 (No. ID405), SEQ ID NO:58 (No. ID414), or SEQ ID NO:564 (No. ID418), SEQ ID NO:335 (No. ID406), SEQ ID NO:331 (No. ID411), SEQ ID NO:20 (No. ID415), or SEQ ID NO:445 (No. ID419), substituted with any one of alanine (A); arginine (R); asparagine (N); aspartic acid (D); cysteine (C); glutamic acid (E); glutamine (N); glycine (G); isoleucine (I); leucine (L); lysine (K); methionine (M); phenylalanine (F); proline (P); serine (S); threonine (T); tryptophan (W); tyrosine (Y); or valine (V); and / or x) an isoleucine (I) residue of SEQ ID NO: 334 (No. ID405), SEQ ID NO: 58 (No. ID414), or SEQ ID NO: 564 (No. ID418), SEQ ID NO: 335 (No. ID406), SEQ ID NO: 331 (No. ID411), SEQ ID NO: 20 (No. ID415), or SEQ ID NO: 445 (No. ID419), substituted with any one of alanine (A); arginine (R); asparagine (N); aspartic acid (D); cysteine (C); glutamic acid (E); glutamine (N); glycine (G); histidine (H); leucine (L); lysine (K); methionine (M); phenylalanine (F); proline (P); serine (S); threonine (T); tryptophan (W); tyrosine (Y); or valine (V); and / or xi) a leucine (L) residue of SEQ ID NO:334 (No. ID405), SEQ ID NO:58 (No. ID414), or SEQ ID NO:564 (No. ID418), SEQ ID NO:335 (No. ID406), SEQ ID NO:331 (No. ID411), SEQ ID NO:20 (No. ID415), or SEQ ID NO:445 (No. ID419), substituted with any one of alanine (A); arginine (R); asparagine (N); aspartic acid (D); cysteine (C); glutamic acid (E); glutamine (N); glycine (G); histidine (H); isoleucine (I); lysine (K); methionine (M); phenylalanine (F); proline (P); serine (S); threonine (T); tryptophan (W); tyrosine (Y); or valine (V); and / or xii) a lysine (K) residue of SEQ ID NO:334 (No. ID405), SEQ ID NO:58 (No. ID414), or SEQ ID NO:564 (No. ID418), SEQ ID NO:335 (No. ID406), SEQ ID NO:331 (No. ID411), SEQ ID NO:20 (No. ID415), or SEQ ID NO:445 (No. ID419), substituted with any one of alanine (A); arginine (R); asparagine (N); aspartic acid (D); cysteine (C); glutamic acid (E); glutamine (N); glycine (G); histidine (H); isoleucine (I); leucine (L); methionine (M); phenylalanine (F); proline (P); serine (S); threonine (T); tryptophan (W); tyrosine (Y); or valine (V); and / or xiii) a methionine (M) residue of SEQ ID NO:334 (No. ID405), SEQ ID NO:58 (No. ID414), or SEQ ID NO:564 (No. ID418), SEQ ID NO:335 (No. ID406), SEQ ID NO:331 (No. ID411), SEQ ID NO:20 (No. ID415), or SEQ ID NO:445 (No. ID419), substituted with any one of alanine (A); arginine (R); asparagine (N); aspartic acid (D); cysteine (C); glutamic acid (E); glutamine (N); glycine (G); histidine (H); isoleucine (I); leucine (L); lysine (K); phenylalanine (F); proline (P); serine (S); threonine (T); tryptophan (W); tyrosine (Y); or valine (V); and / or xiv) a phenylalanine (F) residue of SEQ ID NO: 334 (No. ID405), SEQ ID NO: 58 (No. ID414), or SEQ ID NO: 564 (No. ID418), SEQ ID NO: 335 (No. ID406), SEQ ID NO: 331 (No. ID411), SEQ ID NO: 20 (No. ID415), or SEQ ID NO: 445 (No. ID419), substituted with any one of alanine (A); arginine (R); asparagine (N); aspartic acid (D); cysteine (C); glutamic acid (E); glutamine (N); glycine (G); histidine (H); isoleucine (I); leucine (L); lysine (K); methionine (M); proline (P); serine (S); threonine (T); tryptophan (W); tyrosine (Y); or valine (V); and / or xv) a proline (P) residue of SEQ ID NO: 334 (No. ID405), SEQ ID NO: 58 (No. ID414), or SEQ ID NO: 564 (No. ID418), SEQ ID NO: 335 (No. ID406), SEQ ID NO: 331 (No. ID411), SEQ ID NO: 20 (No. ID415), or SEQ ID NO: 445 (No. ID419), substituted with any one of alanine (A); arginine (R); asparagine (N); aspartic acid (D); cysteine (C); glutamic acid (E); glutamine (N); glycine (G); histidine (H); isoleucine (I); leucine (L); lysine (K); methionine (M); phenylalanine (F); serine (S); threonine (T); tryptophan (W); tyrosine (Y); or valine (V); and / or xvi) a serine (S) residue of SEQ ID NO:334 (No. ID405), SEQ ID NO:58 (No. ID414), or SEQ ID NO:564 (No. ID418), SEQ ID NO:335 (No. ID406), SEQ ID NO:331 (No. ID411), SEQ ID NO:20 (No. ID415), or SEQ ID NO:445 (No. ID419), substituted with any one of alanine (A); arginine (R); asparagine (N); aspartic acid (D); cysteine (C); glutamic acid (E); glutamine (N); glycine (G); histidine (H); isoleucine (I); leucine (L); lysine (K); methionine (M); phenylalanine (F); proline (P); threonine (T); tryptophan (W); tyrosine (Y); or valine (V); and / or xvii) a threonine (T) residue of SEQ ID NO:334 (No. ID405), SEQ ID NO:58 (No. ID414), or SEQ ID NO:564 (No. ID418), SEQ ID NO:335 (No. ID406), SEQ ID NO:331 (No. ID411), SEQ ID NO:20 (No. ID415), or SEQ ID NO:445 (No. ID419), substituted with any one of alanine (A); arginine (R); asparagine (N); aspartic acid (D); cysteine (C); glutamic acid (E); glutamine (N); glycine (G); histidine (H); isoleucine (I); leucine (L); lysine (K); methionine (M); phenylalanine (F); proline (P); serine (S); tryptophan (W); tyrosine (Y); or valine (V); and / or xviii) a tryptophan (W) residue of SEQ ID NO:334 (No. ID405), SEQ ID NO:58 (No. ID414), or SEQ ID NO:564 (No. ID418), SEQ ID NO:335 (No. ID406), SEQ ID NO:331 (No. ID411), SEQ ID NO:20 (No. ID415), or SEQ ID NO:445 (No. ID419), substituted with any one of alanine (A); arginine (R); asparagine (N); aspartic acid (D); cysteine (C); glutamic acid (E); glutamine (N); glycine (G); histidine (H); isoleucine (I); leucine (L); lysine (K); methionine (M); phenylalanine (F); proline (P); serine (S); threonine (T); tyrosine (Y); or valine (V); and / or xix) a tyrosine (Y) residue of SEQ ID NO:334 (No. ID405), SEQ ID NO:58 (No. ID414), or SEQ ID NO:564 (No. ID418), SEQ ID NO:335 (No. ID406), SEQ ID NO:331 (No. ID411), SEQ ID NO:20 (No. ID415), or SEQ ID NO:445 (No. ID419), substituted with any one of alanine (A); arginine (R); asparagine (N); aspartic acid (D); cysteine (C); glutamic acid (E); glutamine (N); glycine (G); histidine (H); isoleucine (I); leucine (L); lysine (K); methionine (M); phenylalanine (F); proline (P); serine (S); threonine (T); tryptophan (W); or valine (V); and / or xx) a valine (V) residue of SEQ ID NO:334 (No. ID405), SEQ ID NO:58 (No. ID414), or SEQ ID NO:564 (No. ID418), SEQ ID NO:335 (No. ID406), SEQ ID NO:331 (No. ID411), SEQ ID NO:20 (No. ID415), or SEQ ID NO:445 (No. ID419), substituted with any one of alanine (A); arginine (R); asparagine (N); aspartic acid (D); cysteine (C); glutamic acid (E); glutamine (N); glycine (G); histidine (H); isoleucine (I); leucine (L); lysine (K); methionine (M); phenylalanine (F); proline (P); serine (S); threonine (T); or tryptophan (W). or (e) is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% identical to SEQ ID NO: 334 (No. ID405); and i) a D169 substitution; a C554 substitution; an N559 substitution; a Q565 substitution; an L860 substitution; an R950 substitution; and / or an R954 substitution; or ii) a D169R substitution; a C554N substitution; a C554R substitution; an N559R substitution; a Q565R substitution; an L860Q substitution; an R950K substitution; and / or an R954A substitution; or iii) a D169 substitution; a D169 / R950 / R954 substitution set; a D169 / N559 / Q565 substitution set; a C554 substitution; a C554 substitution; and / or an L860 substitution; or iv) D169R substitution; D169R / R950K / R954A substitution set; D169R / N559R / Q565R substitution set; C554R substitution; C554N substitution; and / or L860Q substitution or a polypeptide comprising (f) is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8% or at least 99.9% identical to SEQ ID NO: 58 (ID414); and i) a T154 substitution; an N531 substitution; a G546 substitution; a K542 substitution; an S802 substitution; an R887 substitution; and / or an R891 substitution; or ii) a T154R substitution; an N531R substitution; a G546R substitution; a K542R substitution; an S802L substitution; an R887K substitution; and / or an R891A substitution; or iii) a T154 substitution; a T154 / R887 / R891 substitution set; a T154 / G536 / K542 substitution set; a N531 / S802 substitution set; an N531 substitution; and / or an S802 substitution; or iv) T154R substitution; T154R / R887K / R891A substitution set; T154R / G536R / K542R substitution set; N531R / S802L substitution set; N531R substitution; and / or S802L substitution or a polypeptide comprising (g) is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8% or at least 99.9% identical to SEQ ID NO: 564 (ID418); and i) a D161 substitution; an N527 substitution; a T532 substitution; a K538 substitution; a Q799 substitution; an R888 substitution; and / or an R892 substitution; or ii) a D161R substitution; an N527R substitution; a T532R substitution; a K538R substitution; a Q799L substitution; an R888K substitution; and / or an R892A substitution; or iii) a D161 substitution; a D161 / R888 / R892 substitution; a D161 / T532 / K538 substitution; a N527 / Q799 substitution; a N527 substitution; and / or a Q799 substitution; or iv) D161R substitution; D161R / R888K / R892A substitution; D161R / T532R / K538R substitution set; N527R / Q799L substitution; N527R substitution; and / or Q799L substitution or a polypeptide comprising (h) any one of (b) to (g) above, wherein the polypeptide comprises at least one non-naturally occurring amino acid analog or amino acid derivative; or (i) a sequence selected from SEQ ID NO: 365 (Nr. ID405), SEQ ID NO: 74 (Nr. ID414), or SEQ ID NO: 565 (Nr. ID418), SEQ ID NO: 366 (Nr. ID406), SEQ ID NO: 331 (Nr. ID411), SEQ ID NO: 30 (Nr. ID415), or SEQ ID NO: 445 (Nr. ID419), which has at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least a polypeptide encoded by or expressed by a nucleic acid sequence that is at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9% identical, or 100% identical The polypeptide or isolated polypeptide comprising:
2. The polypeptide of claim 1, having at least one activity selected from an endonuclease activity; an endoribonuclease activity; or an RNA-guided DNase activity.
3. a. one or more α-helical recognition lobes (REC) and nuclease recognition lobes (NUC); b. Wedge (WED), α-Helix Recognition Lobe (REC), PAM Interaction (PI), RuvC Nuclease, Bridge Helix (BH), and NUC Domains; or c. one or more domains selected from the RuvC, REC, WED, BH, PI, and NUC domains The polypeptide of claim 1 or 2, comprising:
4. The polypeptide of any one of claims 1 to 3, wherein the polypeptide recognizes or binds to crRNA(s) or is bound to crRNA(s).
5. The polypeptide of claim 4, wherein the crRNA comprises a crRNA sequence from Table S15C.
6. The polypeptide of any one of claims 1 to 5, wherein the polypeptide comprises one or more mutations.
7. 7. The polypeptide of claim 6, wherein the polypeptide comprises one or more mutations in one or more of the RuvC, REC, WED, BH, PI, and NUC domains.
8. The polypeptide according to any one of claims 1 to 7, wherein the polypeptide is a nickase.
9. The polypeptide of any one of claims 1 to 7, wherein the polypeptide is a nuclease-deficient polypeptide.
10. The polypeptide of any one of claims 1 to 9, wherein the polypeptide is operably fused to one or more other polypeptides each having enzymatic activity.
11. The polypeptide of any one of claims 1 to 9, wherein the polypeptide is operably fused to one or more deaminases.
12. The polypeptide of any one of claims 1 to 9, wherein the polypeptide is operably fused to one or more reverse transcriptases.
13. The polypeptide of claim 12, wherein the one or more reverse transcriptases comprise Moloney murine leukemia virus (M-MLV) reverse transcriptase.
14. The polypeptide of claim 12 or 13, wherein the polypeptide is a nickase.
15. The polypeptide of claim 12 or 13, wherein the polypeptide is a nuclease-deficient polypeptide.
16. The polypeptide of any one of claims 1 to 15, wherein the polypeptide is operably linked to one or more nuclear localization signals.
17. The polypeptide of any one of claims 1 to 16, wherein the polypeptide is expressed from a nucleic acid sequence operably linked to one or more expression control sequences.
18. An isolated or recombinant nucleic acid sequence encoding a polypeptide according to any one of claims 1 to 17.
19. A vector expressing or containing a nucleic acid sequence according to claim 18, or a nucleic acid molecule encoding a polypeptide according to any one of claims 1 to 17.
20. 20. The vector of claim 19, wherein the vector comprises a viral vector, a retroviral vector, a lentiviral vector, an adenoviral vector, an adeno-associated viral vector, a vaccinia viral vector, a pox viral vector, a herpes simplex viral vector, a liposome, a lipid nanoparticle (LNP), a cationic polymer, a vesicle, or a gold nanoparticle.
21. A host cell comprising a vector according to claim 19 or 20, or a nucleic acid sequence according to claim 18, or a nucleic acid molecule encoding a polypeptide according to any one of claims 1 to 17, or a polypeptide according to any one of claims 1 to 17.
22. 22. The host cell of claim 21, wherein the host cell comprises a prokaryotic cell, a mammalian cell, a human cell, or a synthetic cell.
23. (a) a nucleic acid sequence encoding a polypeptide having the amino acid sequence of one of SEQ ID NO:334 (No. ID405), SEQ ID NO:58 (No. ID414), or SEQ ID NO:564 (No. ID418), SEQ ID NO:335 (ID406), SEQ ID NO:331 (No. ID411), SEQ ID NO:20 (No. ID415), or SEQ ID NO:445 (No. ID419); (b) has at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity with one of SEQ ID NO:334 (No. ID405), SEQ ID NO:58 (No. ID414), or SEQ ID NO:564 (No. ID418), SEQ ID NO:335 (No. ID406), SEQ ID NO:331 (No. ID411), SEQ ID NO:20 (No. ID415), or SEQ ID NO:445 (No. ID419). %, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9% identical, or 100% identical; a nucleic acid sequence encoding a polypeptide; (c) at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% of one of SEQ ID NO:334 (Nr. ID405), SEQ ID NO:58 (Nr. ID414), or SEQ ID NO:564 (Nr. ID418), SEQ ID NO:335 (Nr. ID406), SEQ ID NO:331 (Nr. ID411), SEQ ID NO:20 (Nr. ID415), or SEQ ID NO:445 (Nr. ID419). 9.5%, at least 99.6%, at least 99.7%, at least 99.8% or at least 99.9% identical to a nucleic acid sequence encoding a polypeptide which is at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 9 ... comprises at least one conservative substitution of an amino acid of at least SEQ ID NO:334 (No. ID405), SEQ ID NO:58 (No. ID414), or SEQ ID NO:564 (No. ID418), SEQ ID NO:335 (No. ID406), SEQ ID NO:331 (No. ID411), SEQ ID NO:20 (No. ID415), or SEQ ID NO:445 (No. ID419); (d) has at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 100%, at least 101%, at least 102%, at least 103%, at least 104%, at least 105%, at least 106%, at least 107%, at least 108%, at least 109%, at least 110%, at least 111%, at least 112%, at least 113%, at least 114%, at least 115%, at least 116%, at least 117%, at least 118%, at least 119 ...
1. A nucleic acid sequence encoding a polypeptide that is at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8% or at least 99.9% identical to SEQ ID NO:334 (No. ID405), SEQ ID NO:58 (No. ID414), or SEQ ID NO:564 (No. ID418), SEQ ID NO:335 (No. ID406), SEQ ID NO:331 (No. ID411), SEQ ID NO:20 (No. ID415), or SEQ ID NO:445 (No. ID419), i) an alanine (A) residue of SEQ ID NO:334 (No. ID405), SEQ ID NO:58 (No. ID414), or SEQ ID NO:564 (No. ID418), SEQ ID NO:335 (No. ID406), SEQ ID NO:331 (No. ID411), SEQ ID NO:20 (No. ID415), or SEQ ID NO:445 (No. ID419), substituted with any one of: arginine (R); asparagine (N); aspartic acid (D); cysteine (C); glutamic acid (E); glutamine (N); glycine (G); histidine (H); isoleucine (I); leucine (L); lysine (K); methionine (M); phenylalanine (F); proline (P); serine (S); threonine (T); tryptophan (W); tyrosine (Y); or valine (V); and / or ii) an arginine (R) residue of SEQ ID NO:334 (No. ID405), SEQ ID NO:58 (No. ID414), or SEQ ID NO:564 (No. ID418), SEQ ID NO:335 (No. ID406), SEQ ID NO:331 (No. ID411), SEQ ID NO:20 (No. ID415), or SEQ ID NO:445 (No. ID419), substituted with any one of alanine (A); asparagine (N); aspartic acid (D); cysteine (C); glutamic acid (E); glutamine (N); glycine (G); histidine (H); isoleucine (I); leucine (L); lysine (K); methionine (M); phenylalanine (F); proline (P); serine (S); threonine (T); tryptophan (W); tyrosine (Y); or valine (V); and / or iii) an asparagine (N) residue of SEQ ID NO:334 (No. ID405), SEQ ID NO:58 (No. ID414), or SEQ ID NO:564 (No. ID418), SEQ ID NO:335 (No. ID406), SEQ ID NO:331 (No. ID411), SEQ ID NO:20 (No. ID415), or SEQ ID NO:445 (No. ID419), substituted with any one of alanine (A); arginine (R); aspartic acid (D); cysteine (C); glutamic acid (E); glutamine (N); glycine (G); histidine (H); isoleucine (I); leucine (L); lysine (K); methionine (M); phenylalanine (F); proline (P); serine (S); threonine (T); tryptophan (W); tyrosine (Y); or valine (V); and / or iv) an aspartic acid (D) residue of SEQ ID NO:334 (No. ID405), SEQ ID NO:58 (No. ID414), or SEQ ID NO:564 (No. ID418), SEQ ID NO:335 (No. ID406), SEQ ID NO:331 (No. ID411), SEQ ID NO:20 (No. ID415), or SEQ ID NO:445 (No. ID419), substituted with any one of alanine (A); arginine (R); asparagine (N); cysteine (C); glutamic acid (E); glutamine (N); glycine (G); histidine (H); isoleucine (I); leucine (L); lysine (K); methionine (M); phenylalanine (F); proline (P); serine (S); threonine (T); tryptophan (W); tyrosine (Y); or valine (V); and / or v) a cysteine (C) residue of SEQ ID NO:334 (No. ID405), SEQ ID NO:58 (No. ID414), or SEQ ID NO:564 (No. ID418), SEQ ID NO:335 (No. ID406), SEQ ID NO:331 (No. ID411), SEQ ID NO:20 (No. ID415), or SEQ ID NO:445 (No. ID419), substituted with any one of alanine (A); arginine (R); asparagine (N); aspartic acid (D); glutamic acid (E); glutamine (N); glycine (G); histidine (H); isoleucine (I); leucine (L); lysine (K); methionine (M); phenylalanine (F); proline (P); serine (S); threonine (T); tryptophan (W); tyrosine (Y); or valine (V); and / or vi) a glutamic acid (E) residue of SEQ ID NO:334 (No. ID405), SEQ ID NO:58 (No. ID414), or SEQ ID NO:564 (No. ID418), SEQ ID NO:335 (No. ID406), SEQ ID NO:331 (No. ID411), SEQ ID NO:20 (No. ID415), or SEQ ID NO:445 (No. ID419), substituted with any one of alanine (A); arginine (R); asparagine (N); aspartic acid (D); cysteine (C); glutamine (N); glycine (G); histidine (H); isoleucine (I); leucine (L); lysine (K); methionine (M); phenylalanine (F); proline (P); serine (S); threonine (T); tryptophan (W); tyrosine (Y); or valine (V); and / or vii) a glutamine (N) residue of SEQ ID NO:334 (No. ID405), SEQ ID NO:58 (No. ID414), or SEQ ID NO:564 (No. ID418), SEQ ID NO:335 (No. ID406), SEQ ID NO:331 (No. ID411), SEQ ID NO:20 (No. ID415), or SEQ ID NO:445 (No. ID419), substituted with any one of alanine (A); arginine (R); asparagine (N); aspartic acid (D); cysteine (C); glutamic acid (E); glycine (G); histidine (H); isoleucine (I); leucine (L); lysine (K); methionine (M); phenylalanine (F); proline (P); serine (S); threonine (T); tryptophan (W); tyrosine (Y); or valine (V); and / or viii) a glycine (G) residue of SEQ ID NO:334 (No. ID405), SEQ ID NO:58 (No. ID414), or SEQ ID NO:564 (No. ID418), SEQ ID NO:335 (No. ID406), SEQ ID NO:331 (No. ID411), SEQ ID NO:20 (No. ID415), or SEQ ID NO:445 (No. ID419), substituted with any one of alanine (A); arginine (R); asparagine (N); aspartic acid (D); cysteine (C); glutamic acid (E); glutamine (N); histidine (H); isoleucine (I); leucine (L); lysine (K); methionine (M); phenylalanine (F); proline (P); serine (S); threonine (T); tryptophan (W); tyrosine (Y); or valine (V); and / or ix) a histidine (H) residue of SEQ ID NO:334 (No. ID405), SEQ ID NO:58 (No. ID414), or SEQ ID NO:564 (No. ID418), SEQ ID NO:335 (No. ID406), SEQ ID NO:331 (No. ID411), SEQ ID NO:20 (No. ID415), or SEQ ID NO:445 (No. ID419), substituted with any one of alanine (A); arginine (R); asparagine (N); aspartic acid (D); cysteine (C); glutamic acid (E); glutamine (N); glycine (G); isoleucine (I); leucine (L); lysine (K); methionine (M); phenylalanine (F); proline (P); serine (S); threonine (T); tryptophan (W); tyrosine (Y); or valine (V); and / or x) an isoleucine (I) residue of SEQ ID NO: 334 (No. ID405), SEQ ID NO: 58 (No. ID414), or SEQ ID NO: 564 (No. ID418), SEQ ID NO: 335 (No. ID406), SEQ ID NO: 331 (No. ID411), SEQ ID NO: 20 (No. ID415), or SEQ ID NO: 445 (No. ID419), substituted with any one of alanine (A); arginine (R); asparagine (N); aspartic acid (D); cysteine (C); glutamic acid (E); glutamine (N); glycine (G); histidine (H); leucine (L); lysine (K); methionine (M); phenylalanine (F); proline (P); serine (S); threonine (T); tryptophan (W); tyrosine (Y); or valine (V); and / or xi) a leucine (L) residue of SEQ ID NO:334 (No. ID405), SEQ ID NO:58 (No. ID414), or SEQ ID NO:564 (No. ID418), SEQ ID NO:335 (No. ID406), SEQ ID NO:331 (No. ID411), SEQ ID NO:20 (No. ID415), or SEQ ID NO:445 (No. ID419), substituted with any one of alanine (A); arginine (R); asparagine (N); aspartic acid (D); cysteine (C); glutamic acid (E); glutamine (N); glycine (G); histidine (H); isoleucine (I); lysine (K); methionine (M); phenylalanine (F); proline (P); serine (S); threonine (T); tryptophan (W); tyrosine (Y); or valine (V); and / or xii) a lysine (K) residue of SEQ ID NO:334 (No. ID405), SEQ ID NO:58 (No. ID414), or SEQ ID NO:564 (No. ID418), SEQ ID NO:335 (No. ID406), SEQ ID NO:331 (No. ID411), SEQ ID NO:20 (No. ID415), or SEQ ID NO:445 (No. ID419), substituted with any one of alanine (A); arginine (R); asparagine (N); aspartic acid (D); cysteine (C); glutamic acid (E); glutamine (N); glycine (G); histidine (H); isoleucine (I); leucine (L); methionine (M); phenylalanine (F); proline (P); serine (S); threonine (T); tryptophan (W); tyrosine (Y); or valine (V); and / or xiii) a methionine (M) residue of SEQ ID NO:334 (No. ID405), SEQ ID NO:58 (No. ID414), or SEQ ID NO:564 (No. ID418), SEQ ID NO:335 (No. ID406), SEQ ID NO:331 (No. ID411), SEQ ID NO:20 (No. ID415), or SEQ ID NO:445 (No. ID419), substituted with any one of alanine (A); arginine (R); asparagine (N); aspartic acid (D); cysteine (C); glutamic acid (E); glutamine (N); glycine (G); histidine (H); isoleucine (I); leucine (L); lysine (K); phenylalanine (F); proline (P); serine (S); threonine (T); tryptophan (W); tyrosine (Y); or valine (V); and / or xiv) a phenylalanine (F) residue of SEQ ID NO: 334 (No. ID405), SEQ ID NO: 58 (No. ID414), or SEQ ID NO: 564 (No. ID418), SEQ ID NO: 335 (No. ID406), SEQ ID NO: 331 (No. ID411), SEQ ID NO: 20 (No. ID415), or SEQ ID NO: 445 (No. ID419), substituted with any one of alanine (A); arginine (R); asparagine (N); aspartic acid (D); cysteine (C); glutamic acid (E); glutamine (N); glycine (G); histidine (H); isoleucine (I); leucine (L); lysine (K); methionine (M); proline (P); serine (S); threonine (T); tryptophan (W); tyrosine (Y); or valine (V); and / or xv) a proline (P) residue of SEQ ID NO: 334 (No. ID405), SEQ ID NO: 58 (No. ID414), or SEQ ID NO: 564 (No. ID418), SEQ ID NO: 335 (No. ID406), SEQ ID NO: 331 (No. ID411), SEQ ID NO: 20 (No. ID415), or SEQ ID NO: 445 (No. ID419), substituted with any one of alanine (A); arginine (R); asparagine (N); aspartic acid (D); cysteine (C); glutamic acid (E); glutamine (N); glycine (G); histidine (H); isoleucine (I); leucine (L); lysine (K); methionine (M); phenylalanine (F); serine (S); threonine (T); tryptophan (W); tyrosine (Y); or valine (V); and / or xvi) a serine (S) residue of SEQ ID NO:334 (No. ID405), SEQ ID NO:58 (No. ID414), or SEQ ID NO:564 (No. ID418), SEQ ID NO:335 (No. ID406), SEQ ID NO:331 (No. ID411), SEQ ID NO:20 (No. ID415), or SEQ ID NO:445 (No. ID419), substituted with any one of alanine (A); arginine (R); asparagine (N); aspartic acid (D); cysteine (C); glutamic acid (E); glutamine (N); glycine (G); histidine (H); isoleucine (I); leucine (L); lysine (K); methionine (M); phenylalanine (F); proline (P); threonine (T); tryptophan (W); tyrosine (Y); or valine (V); and / or xvii) a threonine (T) residue of SEQ ID NO:334 (No. ID405), SEQ ID NO:58 (No. ID414), or SEQ ID NO:564 (No. ID418), SEQ ID NO:335 (No. ID406), SEQ ID NO:331 (No. ID411), SEQ ID NO:20 (No. ID415), or SEQ ID NO:445 (No. ID419), substituted with any one of alanine (A); arginine (R); asparagine (N); aspartic acid (D); cysteine (C); glutamic acid (E); glutamine (N); glycine (G); histidine (H); isoleucine (I); leucine (L); lysine (K); methionine (M); phenylalanine (F); proline (P); serine (S); tryptophan (W); tyrosine (Y); or valine (V); and / or xviii) a tryptophan (W) residue of SEQ ID NO:334 (No. ID405), SEQ ID NO:58 (No. ID414), or SEQ ID NO:564 (No. ID418), SEQ ID NO:335 (No. ID406), SEQ ID NO:331 (No. ID411), SEQ ID NO:20 (No. ID415), or SEQ ID NO:445 (No. ID419), substituted with any one of alanine (A); arginine (R); asparagine (N); aspartic acid (D); cysteine (C); glutamic acid (E); glutamine (N); glycine (G); histidine (H); isoleucine (I); leucine (L); lysine (K); methionine (M); phenylalanine (F); proline (P); serine (S); threonine (T); tyrosine (Y); or valine (V); and / or xix) a tyrosine (Y) residue of SEQ ID NO:334 (No. ID405), SEQ ID NO:58 (No. ID414), or SEQ ID NO:564 (No. ID418), SEQ ID NO:335 (No. ID406), SEQ ID NO:331 (No. ID411), SEQ ID NO:20 (No. ID415), or SEQ ID NO:445 (No. ID419), substituted with any one of alanine (A); arginine (R); asparagine (N); aspartic acid (D); cysteine (C); glutamic acid (E); glutamine (N); glycine (G); histidine (H); isoleucine (I); leucine (L); lysine (K); methionine (M); phenylalanine (F); proline (P); serine (S); threonine (T); tryptophan (W); or valine (V); and / or xx) a valine (V) residue of SEQ ID NO:334 (No. ID405), SEQ ID NO:58 (No. ID414), or SEQ ID NO:564 (No. ID418), SEQ ID NO:335 (No. ID406), SEQ ID NO:331 (No. ID411), SEQ ID NO:20 (No. ID415), or SEQ ID NO:445 (No. ID419), substituted with any one of alanine (A); arginine (R); asparagine (N); aspartic acid (D); cysteine (C); glutamic acid (E); glutamine (N); glycine (G); histidine (H); isoleucine (I); leucine (L); lysine (K); methionine (M); phenylalanine (F); proline (P); serine (S); threonine (T); or tryptophan (W). or the nucleic acid sequence comprising: (e) is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% identical to SEQ ID NO: 334 (No. ID405); and i) a D169 substitution; a C554 substitution; an N559 substitution; a Q565 substitution; an L860 substitution; an R950 substitution; and / or an R954 substitution; or ii) a D169R substitution; a C554N substitution; a C554R substitution; an N559R substitution; a Q565R substitution; an L860Q substitution; an R950K substitution; and / or an R954A substitution; or iii) a D169 substitution; a D169 / R950 / R954 substitution set; a D169 / N559 / Q565 substitution set; a C554 substitution; a C554 substitution; and / or an L860 substitution; or iv) D169R substitution; D169R / R950K / R954A substitution set; D169R / N559R / Q565R substitution set; C554R substitution; C554N substitution; and / or L860Q substitution a nucleic acid sequence encoding a polypeptide comprising: (f) is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8% or at least 99.9% identical to SEQ ID NO: 58 (ID414); and i) a T154 substitution; an N531 substitution; a G546 substitution; a K542 substitution; an S802 substitution; an R887 substitution; and / or an R891 substitution; or ii) a T154R substitution; an N531R substitution; a G546R substitution; a K542R substitution; an S802L substitution; an R887K substitution; and / or an R891A substitution; or iii) a T154 substitution; a T154 / R887 / R891 substitution set; a T154 / G536 / K542 substitution set; a N531 / S802 substitution set; an N531 substitution; and / or an S802 substitution; or iv) T154R substitution; T154R / R887K / R891A substitution set; T154R / G536R / K542R substitution set; N531R / S802L substitution set; N531R substitution; and / or S802L substitution a nucleic acid sequence encoding a polypeptide comprising: (g) is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8% or at least 99.9% identical to SEQ ID NO: 564 (ID418); and i) a D161 substitution; an N527 substitution; a T532 substitution; a K538 substitution; a Q799 substitution; an R888 substitution; and / or an R892 substitution; or ii) a D161R substitution; an N527R substitution; a T532R substitution; a K538R substitution; a Q799L substitution; an R888K substitution; and / or an R892A substitution; or iii) a D161 substitution; a D161 / R888 / R892 substitution; a D161 / T532 / K538 substitution; a N527 / Q799 substitution; a N527 substitution; and / or a Q799 substitution; or iv) D161R substitution; D161R / R888K / R892A substitution; D161R / T532R / K538R substitution set; N527R / Q799L substitution; N527R substitution; and / or Q799L substitution a nucleic acid sequence encoding a polypeptide comprising: (h) the nucleic acid sequence of any one of (b) to (g) above, wherein the polypeptide comprises at least one non-naturally occurring amino acid analog or amino acid derivative; or (i) a nucleic acid sequence at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8% or at least 99.9% identical to a sequence selected from SEQ ID NO: 365 (No. ID405), SEQ ID NO: 74 (No. ID414), or SEQ ID NO: 565 (No. ID418), SEQ ID NO: 366 (No. ID406), SEQ ID NO: 331 (No. ID411), SEQ ID NO: 30 (No. ID415), or SEQ ID NO: 445 (No. ID419).
1. An isolated or recombinant polynucleotide comprising a nucleic acid sequence comprising:
24. 24. The isolated or recombinant nucleic acid sequence of claim 18 or 23, wherein the nucleic acid sequence encodes a polypeptide having at least one activity selected from an endonuclease activity; an endoribonuclease activity, or an RNA-guided DNase activity.
25. The nucleic acid sequence a. one or more α-helical recognition lobes (REC) and nuclease recognition lobes (NUC); b. Wedge (WED), α-Helix Recognition Lobe (REC), PAM Interaction (PI), RuvC Nuclease, Bridge Helix (BH), and NUC Domains; or c. one or more domains selected from the RuvC, REC, WED, BH, PI, and NUC domains 25. The isolated or recombinant nucleic acid sequence of claim 18, 23 or 24, encoding a polypeptide comprising:
26. 26. The isolated or recombinant nucleic acid sequence of any one of claims 18 or 23 to 25, wherein the nucleic acid sequence encodes a polypeptide that recognizes or binds to the crRNA(s).
27. 27. The isolated or recombinant nucleic acid sequence of claim 26, wherein the crRNA comprises a crRNA sequence from Table S15C.
28. 28. The isolated or recombinant nucleic acid sequence of any one of claims 18 or 23 to 27, wherein the nucleic acid sequence encodes a polypeptide comprising one or more mutations.
29. 29. The isolated or recombinant nucleic acid sequence of claim 28, wherein said polypeptide comprises one or more mutations in one or more of the RuvC, REC, WED, BH, PI and NUC domains.
30. 30. The isolated or recombinant nucleic acid sequence of any one of claims 18 or 23 to 29, wherein the nucleic acid sequence encodes a polypeptide comprising nickase activity, or wherein the polypeptide is a nickase.
31. 30. The isolated or recombinant nucleic acid sequence of any one of claims 18 or 23 to 29, wherein the nucleic acid sequence encodes a nuclease-deficient polypeptide.
32. 31. The isolated or recombinant nucleic acid sequence of any one of claims 18 or 23 to 30, wherein the nucleic acid sequence is operably fused to nucleic acids encoding one or more other polypeptides, each having enzymatic activity.
33. 31. The isolated or recombinant nucleic acid sequence of any one of claims 18 or 23 to 30, wherein the nucleic acid sequence is operably fused to a nucleic acid encoding one or more deaminases.
34. 31. The isolated or recombinant nucleic acid sequence of any one of claims 18 or 23 to 30, wherein the nucleic acid sequence is operably fused to one or more nucleic acid sequences encoding reverse transcriptases.
35. 35. The isolated or recombinant nucleic acid sequence of claim 34, wherein the one or more reverse transcriptases comprises Moloney murine leukemia virus (M-MLV) reverse transcriptase.
36. 36. The isolated or recombinant nucleic acid sequence of claim 34 or 35, wherein the nucleic acid sequence encodes a nickase.
37. 36. The isolated or recombinant nucleic acid sequence of claim 34 or 35, wherein the nucleic acid sequence encodes a nuclease-deficient polypeptide.
38. 38. The isolated or recombinant nucleic acid sequence of any one of claims 18 or 23 to 37, wherein the nucleic acid sequence is operably linked to a nucleic acid sequence encoding one or more nuclear localization signals.
39. 39. The isolated or recombinant nucleic acid sequence of any one of claims 18 or 23 to 38, wherein the nucleic acid sequence is operably linked to one or more expression control sequences.
40. A vector comprising the isolated or recombinant nucleic acid sequence of any one of claims 18 or 23 to 39.
41. 41. The vector of claim 40, wherein the vector comprises a viral vector, a retroviral vector, a lentiviral vector, an adenoviral vector, an adeno-associated viral vector, a vaccinia viral vector, a pox viral vector, a herpes simplex viral vector, a liposome, a lipid nanoparticle (LNP), a cationic polymer, a vesicle, or a gold nanoparticle.
42. A host cell comprising a vector according to claim 40 or 41 or an isolated or recombinant nucleic acid sequence according to any one of claims 18, or 23 to 39.
43. 43. The host cell of claim 42, wherein the host cell comprises a prokaryotic cell, a mammalian cell, a human cell, or a synthetic cell.
44. 40. A polypeptide or isolated polypeptide encoded by or expressed from an isolated or recombinant nucleic acid sequence according to any one of claims 23 to 39.
45. A genetic engineering or gene editing or nucleic acid molecule editing or manipulation system comprising an isolated or recombinant nucleic acid sequence according to any one of claims 18, or 23 to 39, and / or a polypeptide according to any one of claims 1 to 17, and / or a vector according to any one of claims 19, 20, 40 or 41, and / or resulting in or resulting in a host cell according to any one of claims 21, 22, 42 or 43.
46. 42. A method, or a method of genetic manipulation, or a method of gene editing, or a method of editing or manipulating a nucleic acid molecule, comprising contacting said gene or nucleic acid molecule with a polypeptide according to any one of claims 1 to 17, and / or introducing a polypeptide according to any one of claims 1 to 17 and / or a vector according to claims 19, 20, 40 or 41 and / or an isolated or recombinant nucleic acid sequence according to any one of claims 18, or 23 to 39 into a cell or medium or solution containing said gene or nucleic acid molecule, and / or wherein said method obtains or results in a host cell according to any one of claims 21, 22, 42 or 43.
47. 42. Use of a polypeptide according to any one of claims 1 to 17, and / or a vector according to claim 19, 20, 40 or 41, and / or an isolated or recombinant nucleic acid sequence according to any one of claims 18, or 23 to 39, in a method for genetic engineering, or a method for gene editing, or a method for nucleic acid molecule editing or manipulation and / or in a genetic engineering or gene editing or nucleic acid molecule editing or manipulation system and / or in a method or system for obtaining or resulting in a host cell according to any one of claims 21, 22, 42 or 43.
48. 48. The method of claim 46 or the use of claim 47, wherein the method or use is ex vivo or in vitro.
49. 48. The method according to claim 46 or the use according to claim 47, excluding methods for the treatment of the human or animal body by surgery or therapy and diagnostic methods performed on the human or animal body.