Novel class 2 type ii and type v crisper-cas rna-guided endonucleases
By developing new Cas9 and Cas12 variants and their engineered systems, the shortcomings of existing CRISPR-Cas RNA-guided enzymes in terms of targeting and cleavage efficiency have been overcome, enabling efficient recognition and cleavage of viruses and other DNA, suitable for virus detection and gene therapy.
Patent Information
- Application Number
- CN202080077872.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-07-29
- Filing Date
- 2020-09-10
- Publication Date
- 2026-02-13
- Estimated Expiration
- 2040-09-10
AI Technical Summary
There is room for improvement in the existing type II and type V CRISPR-Cas RNA-guided endonucleases, particularly in terms of targeting and cleavage efficiency, especially in applications such as virus detection and gene therapy.
New engineered systems for Cas9 variants, Cas12a variants, and Cas12 isotypes are provided, including Cas9.1, Cas9.2, Cas9.3, Cas9.4 proteins and Cas12a.1, Cas12p, or Cas12q proteins, as well as corresponding engineered single-molecule gRNAs that can hybridize with target DNA to form complexes with improved targeting and cleavage capabilities, particularly in the absence of tracrRNA, enabling paracleavage of RNA and single-stranded polynucleotides.
It achieves efficient identification and cleavage of targets such as viral DNA, plant DNA, and fungal DNA, and improves targeting and cleavage efficiency, especially in virus detection and gene therapy, making it suitable for diagnostic and therapeutic applications.
Smart Images

Figure CN114729343B_ABST
Abstract
Description
[0001] Cross Reference to Related Applications
[0002] This application claims priority to U.S. Provisional Patent Application Serial No. 63 / 058,448, filed July 29, 2020, and U.S. Provisional Patent Application Serial No. 62 / 898,340, filed September 10, 2019, each of which is incorporated herein by reference in its entirety.
[0003] Description of Text Files Submitted Electronically
[0004] The Sequence Listing associated with this application is provided in text format in lieu of a paper copy, and is hereby incorporated by reference into this specification. The name of the text file containing this Sequence Listing is “CABI_002_02WO_SeqList_ST25.txt”. The text file is 456 kilobytes, was created on September 10, 2020, and is being submitted electronically via EFS-Web. BACKGROUND
[0006] Bacterial adaptive immune systems appropriately have CRISPR (Clustered Regularly Interspaced Short Palindromic Repeats) and CRISPR-associated (Cas) proteins to perform RNA-guided nucleic acid cleavage. CRISPR-Cas systems endow bacteria and archaea with adaptive immunity through RNA-guided nucleic acid interference. To provide immunity against invaders, processed CRISPR array transcripts (crRNAs) are assembled with Cas protein-containing surveillance complexes that recognize nucleic acids with sequences (called spacers) complementary to a segment of the invading crRNA.
[0007] Class 2 CRISPR-Cas systems are streamlined versions in which a single Cas protein (effector endonuclease protein) bound to RNA is responsible for both binding and cleaving the targeted sequence. The programmable nature of these minimal systems has facilitated their use as a versatile technology that continues to revolutionize the field of genome manipulation.
[0008] However, improved Class 2 Type II and Type V CRISPR-Cas RNA-guided endonuclease variants are needed. Provided herein are such variants and methods of making, testing, and using them. SUMMARY
[0009] Provided herein are new Class 2 Type II and new Class 2 Type V CRISPR-Cas RNA-guided systems, methods of making, and methods of using. More specifically, new Cas9 variants, new Cas12a variants, and new Cas12 subtypes are provided.
[0010] In one aspect, provided herein is an engineered system comprising: (a) a Cas9.1, Cas9.2, Cas9.3, or Cas9.4 protein or a nucleic acid encoding a Cas9.1, Cas9.2, Cas9.3, or Cas9.4 protein; (b) a Cas9.1, Cas9.2, Cas9.3, or Cas9.4 guide RNA (gRNA) or a nucleic acid encoding a Cas9.1, Cas9.2, Cas9.3, or Cas9.4 gRNA, wherein the gRNA and the Cas9.1, Cas9.2, Cas9.3, or Cas9.4 protein do not naturally occur together, wherein the gRNA is capable of hybridizing to a target sequence in a target DNA, and the gRNA is capable of forming a complex with the Cas9.1, Cas9.2, Cas9.3, or Cas9.4 protein.
[0011] In another aspect, provided herein is an engineered single-molecule gRNA comprising: (a) a target-RNA comprising a spacer sequence capable of hybridizing to a target sequence in a target DNA; and (b) an activator-RNA capable of hybridizing to the target-RNA to form a double-stranded RNA duplex, the activator-RNA comprising an activator-RNA, wherein the target-RNA and the activator-RNA are covalently linked to each other, wherein the single-molecule gRNA is capable of forming a complex with a Cas9.1, Cas9.2, Cas9.3, or Cas9.4 protein, and wherein hybridization of the spacer sequence to the target sequence is capable of targeting the Cas9.1, Cas9.2, Cas9.3, or Cas9.4 protein to the target DNA.
[0012] In another aspect, provided herein is an engineered system comprising: a Class 2 Type V CRISPR-Cas RNA-guided endonuclease protein and a single guide RNA, wherein the gRNA and the Class 2 Type V CRISPR-Cas RNA-guided endonuclease protein do not naturally occur together, wherein the gRNA is capable of hybridizing to a target sequence in a target DNA, wherein the gRNA is capable of forming a complex with the Class 2 Type V CRISPR-Cas RNA-guided endonuclease protein, and wherein the Class 2 Type V CRISPR-Cas RNA-guided endonuclease protein has collateral activity and is capable of collateral cleaving a single-stranded polynucleotide comprising RNA without the use of a tracrRNA. In some embodiments, the Class 2 Type V CRISPR-Cas RNA-guided endonuclease protein comprises the amino acid sequence of SEQ ID NO: 4 or an amino acid sequence having at least 70% sequence identity to SEQ ID NO: 4. In some embodiments, the Class 2 Type V CRISPR-Cas RNA-guided endonuclease protein is capable of collateral cleaving single-stranded RNA. In some embodiments, the Class 2 Type V CRISPR-Cas RNA-guided endonuclease protein is capable of collateral cleaving single-stranded DNA / RNA hybrids.
[0013] In another aspect, provided herein is an engineered system comprising: (a) a Cas12a.1, Cas12p, or Cas12q protein or a nucleic acid encoding the Cas12a.1, Cas12p, or Cas12q protein; and (b) a Cas12a.1, Cas12p, or Cas12q gRNA or a nucleic acid encoding a Cas12a.1, Cas12p, or Cas12q gRNA, wherein the gRNA and the Cas12a.1, Cas12p, or Cas12q protein do not naturally occur together, wherein the gRNA is capable of hybridizing to a target sequence in a target DNA, and the gRNA is capable of forming a complex with the Cas12a.1, Cas12p, or Cas12q protein.
[0014] In another aspect, provided herein is an engineered single-molecule gRNA comprising a scaffold sequence of SEQ ID NO: 116 or SEQ ID NO: 117 and a spacer sequence capable of hybridizing to a target sequence in a target DNA. In some embodiments, the target DNA is viral DNA, plant DNA, fungal DNA, or bacterial DNA. In some embodiments, the target sequence is the sequence of a target provided in any one of Tables 6a-6f. In some embodiments, the target is a coronavirus. In some embodiments, the target is a SARS-CoV-2 virus. In some embodiments, the target DNA is cDNA and has been obtained by reverse transcription.
[0015] In another aspect, provided herein is a method of detecting a target DNA in a sample, the method comprising: (a) contacting the sample with: (i) a Casl2a.l, Casl2p, or Casl2q protein; (ii) a Casl2a.l, Casl2p, or Casl2q gRNA comprising a spacer sequence capable of hybridizing to a target sequence in a target DNA; and (iii) a labeled reporter oligonucleotide that does not hybridize to the spacer sequence of the gRNA; and (b) measuring a detectable signal resulting from cleavage of the labeled reporter by the Casl2a.l, Casl2p, or Casl2q protein, thereby detecting the target DNA. Such a method can be used for diagnostics, e.g., detecting viral or bacterial pathogens in a sample.
[0016] In another aspect, provided herein is a method of modifying a target DNA, the method comprising (a) contacting the target DNA with: (i) a Cas9.1, Cas9.2, Cas9.3, Cas9.4, Casl2a.l, Casl2p, or Casl2q protein or nucleotides encoding the same; and (ii) a Cas9.1, Cas9.2, Cas9.3, Cas9.4, Casl2a.l, Casl2p, or Casl2q gRNA comprising a spacer sequence capable of hybridizing to a target sequence in a target DNA. This method can be used for gene therapy applications, as well as to generate cells for therapeutic delivery purposes and to make cell lines.
[0017] In various embodiments, provided herein are compositions, pharmaceutical compositions, vectors, host cells, and kits comprising any of the proteins or polynucleotides of the engineered systems described herein. BRIEF DESCRIPTION OF DRAWINGS
[0018] FIG. 1A to FIG. 1B Expression vector maps for Cas9.1 and Cas9.2 are shown.
[0019] FIG. 2A to FIG. 2C Expression vector maps for Cas12a.1, Cas12p, and Cas12q are shown.
[0020] FIG. 3A is a schematic representation of the CRISPR Cas cluster around the new Cas9.1 gene. FIG. 3B The secondary structure of the direct repeat of the Cas9.1 precursor crRNA is shown. FIG. 3C is a schematic representation of the CRISPR Cas cluster around the new Cas9.2 gene. FIG. 3D is a schematic representation of the CRISPR Cas cluster around the new Cas9.3 gene. FIG. 3E The secondary structure of the direct repeat of the Cas9.3 precursor crRNA is shown. FIG. 3F is a schematic representation of the CRISPR Cas cluster around the new Cas9.4 gene. FIG. 3G The secondary structure of the direct repeat of the Cas9.4 precursor crRNA is shown.
[0021] FIG. 4A Key catalytic amino acids of Cas9 proteins (SEQ ID NOs: 137-168) are shown, as well as an alignment of conserved motifs in selected representatives of the Cas9 protein family. FIG. 4B An alignment of the RuvC1, Bridge Helix, RuvCII, and RuvCIII domains of Cas9.1 (SEQ ID NO: 1) and other selected representatives of the Cas9 protein family (SEQ ID NOs: 169-176) is shown. FIG. 4C An alignment of the RuvC1, Bridge Helix, RuvCII, and RuvCIII domains of Cas9.2 (SEQ ID NO: 2) and other selected representatives of the Cas9 protein family (SEQ ID NOs: 170-174 and 169) is shown. FIG. 4D An alignment of the RuvC1, Bridge Helix, RuvCII, and RuvCIII domains of Cas9.3 (SEQ ID NO: 10) and other selected representatives of the Cas9 protein family (SEQ ID NOs: 169-176) is shown. FIG. 4E An alignment of the RuvC1, Bridge Helix, RuvCII, and RuvCIII domains of Cas9.4 (SEQ ID NO: 11) and other selected representatives of the Cas9 protein family (SEQ ID NOs: 169-176) is shown.
[0022] FIG. 5A is a schematic representation of the CRISPR Cas cluster around the new Cas12a.1 gene. FIG. 5BSecondary structure of the direct repeat of the Cas12a.1 precursor crRNA (SEQ ID NO: 177) is shown. FIG. 5C is a schematic representation of the CRISPR Cas cluster around the new Cas12p gene. FIG. 5D Secondary structure of the direct repeat of the first Cas12p precursor crRNA (SEQ ID NO: 178) and the second Cas12p precursor crRNA (SEQ ID NO: 179) is shown. FIG. 5E is a schematic representation of the CRISPR Cas cluster around the new Cas12q gene. FIG. 5F Secondary structure of the direct repeat of the Cas12q precursor crRNA (SEQ ID NO: 180 and 181) is shown.
[0023] FIG. 6A Key catalytic amino acids of Cas12 proteins (SEQ ID NO: 182-217) are shown, as well as alignment of conserved motifs in select representatives of the Cas12a protein family.
[0024] FIG. 6B Alignment of Cas12a.1 (SEQ ID NO: 3) with SEQ ID NO: 81 of US 20160208243 (SEQ ID NO: 218) is shown, and has 46.8% sequence identity; and FIG. 6C Alignment of Cas12a.1 (SEQ ID NO: 3) with SEQ ID NO: 3 of US 10,253,365 (SEQ ID NO: 219) is shown, and has 46.5% sequence identity.
[0025] FIG. 6D Amino acid sequence of Cas12p (SEQ ID NO: 4) is shown, with the RuvC motif underlined. FnCas12a sequence mentioned by Shmakov et al., 2015 was used as a reference to identify the Ruv motif.
[0026] FIG. 6E Alignment of Cas12p (SEQ ID NO: 4) with Cas12gl (SEQ ID NO: 220) is shown. This figure shows the alignment of Cas12p with Cas12gl.
[0027] In the following figure, the structure of the Cas12p protein was modeled based on the Fn Cas12a structure using the Swiss Model server. FIG. 6F Structure analysis of Cas12p using the Swiss Model server is shown. FIG. 6G Spatial prediction of non-conserved amino acid residues in Cas12p is shown.FIG. 6H Approximation of charge distribution on the surface of Cas12p is shown. FIG. 6I Predicted structural differences between Cas12p (SEQ ID NO: 4) and FnCas12a (SEQ ID NO: 221) based on protein sequence are shown. FIG. 6J RuvCIII domain structure analysis of Cas12p (SEQ ID NO: 4) and Cas12a proteins (AsCas12a (SEQ ID NO: 223), LbCas12a (SEQ ID NO: 224), and FnCas12a (SEQ ID NO: 221)) based on Swiss Model server structure analysis is shown.
[0028] FIG. 6K Amino acid sequence of Cas12q (SEQ ID NO: 5) is shown, with the RuvC motif underlined.
[0029] FIG. 7A FIG. 7B FIG. 7C Predicted RNA secondary structure of non-naturally occurring direct repeats (artificial variants; SEQ ID NOs: 225-239) are shown, which were generated to improve stem loop stability of the guides of the disclosure.
[0030] FIG. 8 Bar graphs showing PAM sequence preference of Cas12a.1 and Cas12p for ten PAM motifs, using a fluorometric assay to measure performance of Cas12a.1 and Cas12p.
[0031] FIG. 9A Specific cleavage activity of Cas12a.1 (designated as Cas12.1 in the figure) and Cas12p proteins of the disclosure with an exemplary hanta virus target is shown. FIG. 9B Both Cas12a.1 and Cas12p exhibit collateral cleavage activity and can cleave non-targets containing ssDNA are shown. FIG. 9C Cas12p exhibits ssDNA and RNA reporter collateral cleavage, with SARS-CoV-2 inactivated virus as sample as target is shown.
[0032] FIG. 10 Activity of new cas12 proteins at 25 °C is shown.
[0033] FIG. 11 Activity of new Cas12 proteins at various salt concentrations is shown.
[0034] FIG. 12 Performance of Cas12a.1 and Cas12p of the disclosure in three different commercial buffers is shown.
[0035] FIG. 13 RPA-free sensitivity curves for Cas12a.1 and Cas12p of the disclosure are shown, measuring each target concentration for 30 minutes.
[0036] FIG. 14 Fluorescence detection amounts obtained by Cas12a.1 and Cas12p are shown to be equal for target DNA reverse transcribed from SARS-CoV-2 RNA at 37°C and 25°C, indicating thermal stability and functionality at room temperature.
[0037] FIG. 15 Differential performance of Cas12p versus LbCas12a at 25°C is shown.
[0038] FIG. 16 Differential performance of Cas12p versus LbCas12a at 25°C using SARS-CoV-2 as a target is shown, described in Example 10.
[0039] FIG. 17 Ability of Cas12p to cleave ssDNA and RNA reporters is shown.
[0040] FIG. 18 A schematic workflow for detecting SARS-CoV-2 described herein is shown.
[0041] FIG. 19 A schematic workflow for detecting SARS-CoV-2 described herein is shown.
[0042] FIG. 20 Cas12p is shown to have minimal background signal after 30-60 minute cleavage activity. This provides an advantage for low viral concentrations and indicates stability of the lyophilized format.
[0043] FIG. 21 A diagnostic assay using Cas12p at room temperature is shown to be able to be read out in a paper format.
[0044] FIG. 22 A diagnostic assay using Cas12p at room temperature is shown to be able to be read out in a well plate with a fluorescence detector.
[0045] FIG. 23 Exemplary lyophilized beads of the disclosure are shown.
[0046] FIG. 24Results of SARS-CoV-2 detection using Cas12p / guide, using an RNA reporter, on lyophilized versions of patient samples and negative control samples are shown.
[0047] FIG. 25 Specific dsDNA cleavage time course of the Cas12a.1 and Cas12p proteins of the disclosure complexed with sgRNA to an exemplary Hantavirus target is shown. Time points: 0, 30, 60, and 90 minutes.
[0048] FIG. 26 Specific ssDNA cleavage time course of the Cas12a.1 and Cas12p proteins of the disclosure complexed with sgRNA to an exemplary Hantavirus target is shown. (S): 3’FAM-ssDNA target substrate. (P): 3’FAM-ssDNA target product. (NTC): AS ssDNA non-target control. Time points: 0, 0.5, 1, and 5 minutes.
[0049] FIG. 27 Specific ssRNA cleavage time course of the Cas12a.1 and Cas12p proteins of the disclosure complexed with sgRNA to an exemplary Hantavirus target is shown. (S): ssRNA target substrate. (TC): ssDNA target control. (NTC): ssRNA non-target control. Time points: 0, 1, and 3 hours.
[0050] FIG. 28 Mass spectrometry data for Cas12p reactions using DNA oligos as reporters is shown.
[0051] FIG. 29 Mass spectrometry data for Cas12p reactions using DNA oligos as reporters is shown.
[0052] FIG. 30 Mass spectrometry data for Cas12p reactions using RNA oligos as reporters is shown.
[0053] FIG. 31 Mass spectrometry data for Cas12p reactions using RNA oligos as reporters is shown.
[0054] FIG. 32 DNA-RNA chimeric guides are able to achieve efficient collateral activity when used with Cas12p is shown.
[0055] FIG. 33 Agarose gels demonstrating collateral activity of Cas12a.1 and Cas12p on ssDNA but not dsDNA is shown.
[0056] FIG. 34The differential cleavage efficiency of homopolymer reporter molecules at 25 °C and 37 °C was shown. The results indicate that Cas12p cleaves poly(T), poly(A), and poly(C), while Cas12a.1 shows a preference for cleavage of poly(C).
[0057] FIG. 35 This demonstrates the ability of Cas12p, rather than Cas12a.1, to perform paracleavage (also referred to as trans-cleavage in this paper) of RNA reporter cleavage.
[0058] FIG. 36 The kinetics of paracleavage activity of Cas12p and Cas12a.1 using DNA and RNA as reporters are shown.
[0059] FIG. 37 Paracuts of Cas12p and Cas12a.1 using FAMQ DNA-RNA chimeric reporter are shown.
[0060] FIG. 38 The sequence and secondary structure of the mature guide scaffolds of Cas12a.1 (SEQ ID NO:116) and Cas12p (SEQ ID NO:117) are shown.
[0061] FIG. 39 The validation shown demonstrates that, using Cas12a.1 and Cas12p, when used in conjunction with spacers targeting the N gene of SARS-CoV-2, a maturation guide scaffold can be used to detect SARS-CoV-2. Detailed Implementation
[0062] This article provides new type II and new type V CRISPR-Cas RNA-guided systems, preparation methods, and usage methods.
[0063] definition
[0064] In this document, the terms “polynucleotide” or “nucleic acid” are used interchangeably to refer to a polymer of nucleotides (ribonucleotides or deoxyribonucleotides) of any length. Therefore, the terms “polynucleotide” and “nucleic acid” encompass single-stranded DNA; double-stranded DNA; multi-stranded DNA; single-stranded RNA; double-stranded RNA; multi-stranded RNA; genomic DNA; cDNA; DNA-RNA hybrids; and polymers containing purine and pyrimidine bases or other natural, chemically or biochemically modified, non-natural, or derivatized nucleotide bases.
[0065] “Hybridizable” or “complementary” or “substantially complementary” refers to nucleic acids (e.g., RNA, DNA) comprising nucleotide sequences that are capable of non-covalent binding (i.e., forming Watson-Crick base pairs and / or G / U base pairs) in a sequence-specific (anti-parallel) manner, “annealing” or “hybridizing” to another nucleotide under conditions of appropriate temperature and solution ionic strength found in suitable in vitro and / or in vivo temperatures and solution ionic strengths, to another nucleotide (i.e., a nucleic acid specifically binds to a complementary nucleic acid).
[0066] It is understood that the sequence of a polynucleotide need not be 100% complementary to that of its target nucleic acid to be capable of specifically hybridizing. Moreover, a polynucleotide can hybridize over one or more segments such that intervening or adjacent segments do not participate in the hybridization event (e.g., a loop structure or hairpin structure, ‘bulge’, etc.).
[0067] The percent complementarity of a particular nucleic acid sequence segment within a nucleic acid can be determined using any convenient method. Exemplary methods include the BLAST program (Basic Local Alignment Search Tool) and the PowerBLAST program (Altschul et al., J. Mol. Biol., 1990, 215, 403-410; Zhang and Madden, Genome Res., 1997, 7, 649-656), or using the Gap program (Wisconsin Sequence Analysis Package, Version 8 for Unix, Genetics Computer Group, University Research Park, Madison Wis.), for example, using default settings, which uses the algorithm of Smith and Waterman (Adv. Appl. Math., 1981, 2, 482-489).
[0068] The terms “peptide,” “polypeptide,” and “protein” are used interchangeably herein and refer to polymeric forms of amino acids of any length, which can include coded and non-coded amino acids, chemically or biochemically modified or derivatized amino acids, and polypeptides having modified peptide backbones.
[0069] A “vector” or “expression vector” is a replicon (e.g., a plasmid, phage, virus, or cosmid) to which another DNA segment (i.e., an “insert”) can be attached so as to bring the attached segment into association with the vector’s DNA-dependent DNA polymerase and other proteins necessary for replication.
[0070] General methods in molecular and cellular biochemistry can be found in the following standard texts: Molecular Cloning: A Laboratory Manual, 3rded. (Sambrook et al., HaRBor Laboratory Press 2001); Short Protocols in Molecular Biology, 4thed. (Ausubel et al., eds., John Wiley & Sons 1999); Protein Methods (Bollag et al., John Wiley & Sons 1996); Nonviral Vectors for Gene Therapy (Wagner et al., eds., Academic Press 1999); Viral Vectors (Kaplift & Loewy, eds., Academic Press 1995); Immunology Methods Manual (I. Lefkovits, ed., Academic Press 1997); and Cell and Tissue Culture: Laboratory Procedures in Biotechnology (Doyle & Griffiths, John Wiley & Sons 1998), the disclosures of which are incorporated herein by reference.
[0071] Where a series of values is provided, it is understood that each intervening value, to the tenth of the unit of the lower magnitude (unless the context clearly indicates otherwise), between the upper and lower limits of that range and any other stated or intervening value in that stated range, is encompassed within the application. The upper and lower limits of these smaller ranges can independently be included in the smaller ranges, and are also encompassed within the application, subject to any explicitly excluded limit in the stated range. Where the stated range includes one or both of the limits, ranges excluding either or both of the included limits are also included in the application.
[0072] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. All publications mentioned herein are incorporated by reference to disclose and describe the methods and / or materials in connection with which the publications are cited.
[0073] It must be noted that, as used herein and in the appended claims, the singular form "a", "an", and "the" include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to "a Casl2a.1 protein" includes a plurality of such Casl2a.1 proteins, and reference to "a gRNA" or "a guide RNA" includes reference to one or more gRNAs and equivalents thereof known to those skilled in the art, and so forth. It should also be noted that the claims can be drafted to exclude any optional element. Thus, this statement is intended to serve as antecedent basis for using such exclusive terminology to disclaim any optional element in each claim in this specification.
[0074] I. Class 2 Type II CRISPR-Cas RNA-guided systems
[0075] Provided herein are novel Class 2 Type II CRISPR-Cas RNA-guided proteins and their guide RNAs ("guide RNAs" are referred to herein interchangeably as "gRNAs"), which make up the Class 2 Type II CRISPR-Cas RNA-guided systems of the disclosure. As used herein, a gRNA can comprise only RNA nucleotides, can comprise RNA and DNA nucleotides, or can comprise only DNA nucleotides, and thus when referred to as a gRNA, can comprise non-RNA nucleotides.
[0076] Thus, provided herein are systems comprising: (a) a Cas9.1, Cas9.2, Cas9.3, or Cas9.4 protein or a nucleic acid encoding the Cas9.1, Cas9.2, Cas9.3, or Cas9.4 protein; (b) a Cas9.1, Cas9.2, Cas9.3, or Cas9.4 gRNA or a nucleic acid encoding a Cas9.1, Cas9.2, Cas9.3, or Cas9.4 molecular RNA, wherein the gRNA and the Cas9.1, Cas9.2, Cas9.3, or Cas9.4 protein do not naturally occur together, wherein the gRNA is capable of hybridizing to a target sequence in a target DNA, and the gRNA is capable of forming a complex with the Cas9.1, Cas9.2, Cas9.3, or Cas9.4 protein. It should be understood that "Cas9.1-Cas9.4" as used herein refers to the following: Cas9.1, Cas9.2, Cas9.3, Cas9.4.
[0077] These components are in turn described below.
[0078] a. Class 2 Type II CRISPR-Cas RNA-guided proteins
[0079] Provided herein are novel Class 2 Type II and Type V CRISPR-Cas RNA-guided endonucleases, e.g., novel Cas9 proteins (Cas9 variants) and novel Cas12a proteins (Cas12a variants) and novel Cas12 subtypes.
[0080] Table 1 shows the protein sequences of the novel Cas9 proteins of the disclosure. In some embodiments, the novel Cas9 proteins of the disclosure from metagenomic samples have been deduced using bioinformatics methods.
[0081] SEQ ID NO: 1 represents a novel Cas9 variant of the disclosure, Cas9.1 (length of 1038 amino acids). FIG. 3A is a schematic representation of the CRISPR Cas cluster surrounding the novel Cas9.1 gene. FIG. 4A The key catalytic amino acids of the Cas9 protein are shown, as well as an alignment of conserved motifs in selected representatives of the Cas9 protein family. FIG. 4B An alignment of the RuvCl, bridge helix, RuvCII, and RuvCIII domains of Cas9.1, as well as other selected representatives of the Cas9 protein family, is shown.
[0082] SEQ ID NO: 2 represents a novel Cas9 variant of the disclosure, Cas9.2 (length of 1375 amino acids). FIG. 3C is a schematic representation of the CRISPR Cas cluster surrounding the novel Cas9.2 gene. FIG. 4C An alignment of the RuvCl, bridge helix, RuvCII, and RuvCIII domains of Cas9.2, as well as other selected representatives of the Cas9 protein family, is shown.
[0083] SEQ ID NO: 10 represents a novel Cas9 variant of the disclosure, Cas9.3 (length of 1031 amino acids). FIG. 3D is a schematic representation of the CRISPR Cas cluster surrounding the novel Cas9.3 gene. FIG. 4D An alignment of the RuvCl, bridge helix, RuvCII, and RuvCIII domains of Cas9.3, as well as other selected representatives of the Cas9 protein family, is shown.
[0084] SEQ ID NO: 11 represents a novel Cas9 variant of the disclosure, Cas9.4 (length of 1329 amino acids). FIG. 3F is a schematic representation of the CRISPR Cas cluster surrounding the novel Cas9.4 gene. FIG. 4E An alignment of the RuvCl, bridge helix, RuvCII, and RuvCIII domains of Cas9.4, as well as other selected representatives of the Cas9 protein family, is shown.
[0085] Table 1
[0086]
[0087]
[0088]
[0089] As used herein, Cas9.1 includes SEQ ID NO: 1 and proteins having at least 70%-99.5% sequence identity to SEQ ID NO: 1. Accordingly, provided herein are proteins comprising the amino acid sequence of SEQ ID NO: 1 and proteins having at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% sequence identity thereto. Also provided herein are nucleic acids encoding proteins comprising the amino acid sequence of SEQ ID NO: 1 and proteins having at least 70%-99.5% sequence identity thereto.
[0090] As used herein, Cas9.2 includes SEQ ID NO: 2 and proteins having at least 70%-99.5% sequence identity to SEQ ID NO: 2. Accordingly, provided herein are proteins comprising the amino acid sequence of SEQ ID NO: 2 and proteins having at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% sequence identity thereto. Also provided herein are nucleic acids encoding proteins comprising the amino acid sequence of SEQ ID NO: 2 and proteins having at least 70%-99.5% sequence identity thereto.
[0091] As used herein, Cas9.3 includes SEQ ID NO: 10 and proteins having at least 70%-99.5% sequence identity to SEQ ID NO: 10. Accordingly, provided herein are proteins comprising the amino acid sequence of SEQ ID NO: 10 and proteins having at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% sequence identity thereto. Also provided herein are nucleic acids encoding proteins comprising the amino acid sequence of SEQ ID NO: 10 and proteins having at least 70%-99.5% sequence identity thereto.
[0092] As used herein, Cas9.4 includes SEQ ID NO: 11 and proteins having at least 70%-99.5% sequence identity to SEQ ID NO: 11. Accordingly, provided herein are proteins comprising the amino acid sequence of SEQ ID NO: 11 and proteins having at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% sequence identity thereto. Also provided herein are nucleic acids encoding proteins comprising the amino acid sequence of SEQ ID NO: 11 and proteins having at least 70%-99.5% sequence identity thereto.
[0093] In some embodiments, the Cas9 protein of the present disclosure is a catalytically active Cas9 protein, e.g., a catalytically active Cas9.1, Cas9.2, Cas9.3, or Cas9.4 protein.
[0094] In some embodiments, the Cas9 protein of the present disclosure cleaves at a site distal to the target sequence, e.g., a Cas9.1, Cas9.2, Cas9.3, or Cas9.4.4 protein cleaves at a site distal to the target sequence.
[0095] In some embodiments, the Cas9 protein of the present disclosure is a catalytically inactive Cas9 protein, e.g., a Cas9.1, Cas9.2, Cas9.3, or Cas9.4 protein is catalytically inactive (a dCas9.1, dCas9.2, dCas9.3, or dCas9.4 protein).
[0096] In some embodiments, the Cas9 protein of the present disclosure is a nickase Cas9 protein, e.g., a Cas9.1 nickase, Cas9.2 nickase, Cas9.3 nickase, or Cas9.4 nickase protein.
[0097] A Cas9 protein of the present disclosure can be modified to comprise an aptamer.
[0098] A Cas9 protein of the present disclosure can be further fused to a domain (e.g., a catalytic domain) to produce a dual-acting Cas protein. In some embodiments, a Cas9 protein is further fused to a base editor.
[0099] b. gRNAs for Class 2 Type II CRISPR-Cas RNA-guided proteins
[0100] The present disclosure provides a DNA-targeting RNA that directs the activity of the novel Cas9 proteins of the present disclosure to a specific target sequence within a target DNA. As provided herein, these DNA-targeting RNAs are referred to herein as "gRNAs" or "gRNAs." In general, as provided herein, a Cas9 variant gRNA comprises a first segment (also referred to herein as a "target-RNA," a "DNA-targeting segment," or a "DNA-targeting sequence") and a second segment (also referred to herein as an "activator-RNA," an "activator-RNA," or a "protein-binding sequence"). Nucleotide sequences encoding the Cas9 gRNAs of the present disclosure are also provided herein.
[0101] i. target-RNA
[0102] The target-RNA of a Cas9 variant gRNA of the present disclosure comprises a nucleotide sequence that is complementary to a sequence in a target DNA (the target sequence of the gRNA; the DNA-targeting sequence; the spacer sequence). The target-RNA is interchangeably referred to as a crRNA. The target-RNA of a gRNA interacts with a target DNA in a sequence-specific manner through hybridization (i.e., base pairing). Thus, the nucleotide sequence of a target-RNA can vary and determines the location within the target DNA at which the gRNA and the target DNA interact. The target-RNA of a subject gRNA can be modified (e.g., by genetic engineering) to hybridize to any desired sequence within a target DNA.
[0103] The target-RNA can be about 12 nucleotides to about 100 nucleotides in length. For example, the target-RNA can be about 12 nucleotides (nt) to about 80 nt, about 12 nt to about 50 nt, about 12 nt to about 40 nt, about 12 nt to about 30 nt, about 12 nt to about 25 nt, about 12 nt to about 20 nt, or about 12 nt to about 19 nt in length. For example, the target-RNA can be about 19 nt to 20 nt, about 19 nt to 25 nt, about 19 nt to 30 nt, about 19 nt to 35 nt, about 19 nt to 40 nt, about 19 nt to 45 nt, about 19 nt to 50 nt, about 19 nt to 60 nt, about 19 nt to 70 nt, about 19 nt to 80 nt, about 19 nt to 90 nt, about 19 nt to 100 nt, about 20 nt to 25 nt, about 20 nt to 30 nt, about 20 nt to 35 nt, about 20 nt to 40 nt, about 20 nt to 45 nt, about 20 nt to 50 nt, about 20 nt to 60 nt, about 20 nt to 70 nt, about 20 nt to 80 nt, about 20 nt to 90 nt, or about 20 nt to about 100 nt in length.
[0104] Generally, the native unprocessed precursor crRNA of Cas9 comprises a forward repeat and an adjacent spacer (the portion of the crRNA that allows for targeting of a DNA molecule). In some embodiments, mutating the forward repeat from the unprocessed precursor crRNA and including the forward repeat in the mature gRNA can improve gRNA stability.
[0105] Table 2 shows the naturally occurring forward repeat sequences of the naturally occurring crRNAs of the Cas9 variants of the disclosure.
[0106] Table 2: Forward Repeat Sequences
[0107]
[0108] In some embodiments, the gRNAs of the disclosure comprise a non-naturally occurring engineered forward repeat sequence, which can be incorporated into the engineered gRNAs of the disclosure.
[0109] ii. Spacer Sequences
[0110] The gRNAs of the present disclosure comprise a spacer sequence that is complementary to a target DNA. More specifically, the length of the nucleotide sequence of the target-RNA that is complementary to the target nucleotide sequence of the target DNA (the DNA targeting sequence or spacer sequence) can be at least about 12 nt. For example, the length of the DNA targeting sequence of the target-RNA that is complementary to the target sequence of the target DNA can be at least about 12 nt, at least about 15 nt, at least about 18 nt, at least about 19 nt, at least about 20 nt, at least about 25 nt, at least about 30 nt, at least about 35 nt, or at least about 40 nt. For example, the length of the DNA targeting sequence of the target-RNA that is complementary to the target sequence of the target DNA can be from about 12 nucleotides (nt) to about 80 nt, from about 12 nt to 50 nt, from about 12 nt to 45 nt, from about 12 nt to 40 nt, from about 12 nt to 35 nt, from about 12 nt to 30 nt, from about 12 nt to 25 nt, from about 12 nt to 20 nt, from about 12 nt to 19 nt, from about 19 nt to 20 nt, from about 19 nt to 25 nt, from about 19 nt to 30 nt, from about 19 nt to 35 nt, from about 19 nt to 40 nt, from about 19 nt to 45 nt, from about 19 nt to 50 nt, from about 19 nt to 60 nt, from about 20 nt to 25 nt, from about 20 nt to 30 nt, from about 20 nt to 35 nt, from about 20 nt to 40 nt, from about 20 nt to 45 nt, from about 20 nt to 50 nt, or from about 20 nt to about 60 nt. The length of the nucleotide sequence of the target-RNA that is complementary to the nucleotide sequence of the target DNA (the target sequence) (the DNA targeting sequence) can be at least about 12 nt. In some embodiments, the length of the DNA targeting sequence of the target-RNA that is complementary to the target sequence of the target DNA is 20 nucleotides. In some embodiments, the length of the DNA targeting sequence of the target-RNA that is complementary to the target sequence of the target DNA is 19 nucleotides.
[0111] The percent complementarity between the spacer sequence of the target-RNA and the target sequence of the target DNA can be at least 60% (e.g., at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99%, or 100%). In some embodiments, the percent complementarity between the DNA targeting sequence of the target-RNA and the target sequence of the target DNA is 100% over the 1-25 contiguous most 5' nucleotides of the target sequence of the complementary strand of the target DNA. In some embodiments, the percent complementarity between the DNA targeting sequence of the target-RNA and the target sequence of the target DNA is at least 60% over about 1-25 contiguous nucleotides. In some embodiments, the percent complementarity between the DNA targeting sequence of the target-RNA and the target sequence of the target DNA is 100% over the 1-25 contiguous most 5' nucleotides of the target sequence of the complementary strand of the target DNA and 0% over the remainder. In this case, the length of the DNA targeting sequence can be considered to be 1-25 nucleotides.
[0112] In some embodiments, the spacer sequence of a Cas9 gRNA of the disclosure is directed to a target sequence in a mammalian organism. In some embodiments, the spacer sequence is directed to a target sequence in a non-mammalian organism.
[0113] In some embodiments, the spacer sequence of a Cas9 gRNA of the disclosure is directed to a target sequence that is a sequence of a human. In some embodiments, the target sequence is a sequence of a non-human primate.
[0114] In some embodiments, the spacer sequence of a Cas9 gRNA of the disclosure is directed to a selected target sequence of a therapeutic target.
[0115] In some embodiments, the spacer sequence of a Cas9 gRNA of the disclosure is directed to a selected target sequence of a diagnostic target - e.g., in such embodiments, a labeled dCas9 of the disclosure and a gRNA directed to a diagnostic target DNA are contacted with the target DNA, or a cell comprising the target DNA, or a sample comprising the target DNA.
[0116] iii. Activator-RNA
[0117] The activator-RNA of a Cas9 variant gRNA of the disclosure binds to its cognate Cas9 variant of the disclosure. The activator-RNA is interchangeably referred to as a tracrRNA. The gRNA directs the bound Cas9 protein to a specific nucleotide sequence within a target DNA by the target-RNA described above. The activator-RNA of a Cas9 variant gRNA comprises two stretches of nucleotides that are complementary to each other.
[0118] iv. Bis-molecular Cas9 gRNA
[0119] In some embodiments, provided herein are bis-molecular (two-molecular) Cas9 gRNAs of the novel Cas9 proteins of the disclosure. Such gRNAs comprise two separate RNA molecules (an activator RNA - tracRNA; and a targeting RNA - crRNA). Each of the two RNA molecules of the subject bis-molecular gRNA comprises two stretches of nucleotides that are complementary to each other, such that the complementary nucleotides of the two RNA molecules hybridize to form a double-stranded RNA duplex of the gRNA.
[0120] The bis-molecular gRNAs can be designed to allow for controlled binding (i.e., conditional binding) of the target-RNA to the activator-RNA. Because the bis-molecular gRNAs are non-functional unless both the activator-RNA and the target-RNA are bound in a functional complex with the Cas9 variants of the disclosure, the bis-molecular gRNAs can be inducible (e.g., drug-inducible) by making the binding between the activator-RNA and the target-RNA inducible. As a non-limiting example, RNA aptamers can be used to modulate (i.e., control) the binding of the activator-RNA to the target-RNA. Thus, the activator-RNA and / or the target-RNA can comprise an RNA aptamer sequence.
[0121] Bis-molecular guides can be modified to comprise aptamers
[0122] v. Single-molecular Cas9 variant gRNA
[0123] In some embodiments, provided herein are Cas9 gRNAs of the novel Cas9 proteins of the disclosure, the Cas9 gRNAs comprising a single-molecular gRNA (interchangeably referred to herein as sgRNA).
[0124] Accordingly, provided herein is an engineered single-molecular gRNA, the engineered single-molecular gRNA comprising:
[0125] a. a target-RNA capable of hybridizing to a target sequence in a target DNA; and
[0126] b. an activator-RNA capable of hybridizing to the target-RNA to form a double-stranded RNA duplex, the activator-RNA comprising an activator-RNA,
[0127] wherein the target-RNA and the activator-RNA are covalently linked to each other, wherein the single-molecular gRNA is capable of forming a complex with the novel Cas9 proteins of the disclosure, and wherein the hybridization of the target-RNA to the target sequence is capable of targeting the Cas9 proteins of the disclosure to the target DNA.
[0128] The subject single-molecule gRNA comprises two segments of nucleotides that are complementary to each other (a target-RNA and an activator-RNA) that can be covalently linked and hybridized to form a double-stranded RNA duplex (a dsRNA duplex) of the activator-RNA, creating a stem-loop structure via intervening nucleotides ("linker" or "linker nucleotides"). In some embodiments, the target-RNA and the activator-RNA are covalently linked via the 3' end of the target-RNA and the 5' end of the activator-RNA. In other embodiments, the activator-RNA is covalently linked via the 5' end of the target-RNA and the 3' end of the activator-RNA.
[0129] In some embodiments, the target-RNA and the activator-RNA are arranged in a 5' to 3' orientation.
[0130] In some embodiments, the activator-RNA and the target-RNA are arranged in a 5' to 3' orientation.
[0131] In some embodiments, the single-molecule gRNA comprises one or more sequence modifications compared to the sequence of a corresponding wild-type tracrRNA and / or crRNA.
[0132] In some embodiments, the target-RNA and the activator-RNA are covalently linked to each other via a linker.
[0133] When present, the linker of the single-molecule gRNA can be from about 3 nucleotides to about 30 nucleotides in length. In exemplary embodiments, the linker of the single-molecule gRNA is 4, 5, 6, or 7 nt.
[0134] Exemplary single-molecule gRNAs comprise two segments of complementary nucleotides that are hybridized to form a dsRNA duplex. In some embodiments, one of the two segments of complementary nucleotides of the single-molecule gRNA (or DNA encoding the segment) is at least about 60% identical to one of the activator-RNAs. For example, one of the two segments of complementary nucleotides of the single-molecule gRNA (or DNA encoding the segment) is at least about 65% identical, at least about 70% identical, at least about 75% identical, at least about 80% identical, at least about 85% identical, at least about 90% identical, at least about 95% identical, at least about 98% identical, at least about 99% identical, or 100% identical to one of the activator-RNAs.
[0135] The activator-RNA and target-RNA segments can be engineered while ensuring that the structure of the protein-binding domain of the gRNA is conserved. Thus, the RNA folding structure of the naturally occurring protein-binding domain of the RNA targeting DNA can be considered to design an artificial protein-binding domain (either a dual-molecule or single-molecule version).
[0136] The activator-RNA in a single-molecule gRNA can be from about 10 nucleotides to about 100 nucleotides in length. For example, the activator-RNA can be from about 15 nucleotides (nt) to about 80 nt, from about 15 nt to about 50 nt, from about 15 nt to about 40 nt, from about 15 nt to about 30 nt, or from about 15 nt to about 25 nt in length.
[0137] Likewise with respect to the single-molecule and double-molecule gRNAs of the present disclosure, the dsRNA duplex of the activator-RNA can be from about 6 nucleotides (nt) to about 50 bp in length. For example, the dsRNA duplex of the activator-RNA can be from about 6 nt to about 40 nt, from about 6 nt to about 30 bp, from about 6 nt to about 25 nt, from about 6 nt to about 20 nt, from about 6 nt to about 15 nt, from about 8 nt to about 40 nt, from about 8 nt to about 30 bp, from about 8 nt to about 25 nt, from about 8 nt to about 20 nt, or from about 8 nt to about 15 nt in length. For example, the dsRNA duplex of the activator-RNA can be from about 8 nt to about 10 nt, from about 10 nt to about 15 nt, from about 15 nt to about 18 nt, from about 18 nt to about 20 nt, from about 20 nt to about 25 nt, from about 25 nt to about 30 nt, from about 30 nt to about 35 nt, from about 35 nt to about 40 nt, or from about 40 nt to about 50 nt in length. In some embodiments, the dsRNA duplex of the activator-RNA is 8-15 base pairs in length. The percent complementarity between the nucleotide sequences hybridized to form the dsRNA duplex of the activator-RNA can be at least about 60%. For example, the percent complementarity between the nucleotide sequences hybridized to form the dsRNA duplex of the activator-RNA can be at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%. In some embodiments, the percent complementarity between the nucleotide sequences hybridized to form the dsRNA duplex of the activator-RNA is 100%.
[0138] In some embodiments, the spacer sequence of a Cas9 gRNA of the present disclosure, whether a single-molecule gRNA or a double-molecule gRNA, is directed to a target sequence in a mammalian organism (e.g., a human or non-human primate). In some embodiments, the spacer sequence of a Cas9 gRNA of the present disclosure is directed to a target sequence in a bacterium.
[0139] In some embodiments, the spacer sequence of a Cas9 gRNA of the present disclosure is directed to a target sequence in a virus. In some embodiments, the spacer sequence of a Cas9 gRNA of the present disclosure is directed to a target sequence in a plant.
[0140] In some embodiments, a single molecule Cas9 gRNA of the disclosure can be modified to comprise an aptamer.
[0141] vi. gRNA arrays
[0142] A Cas9 gRNA of the disclosure can be provided as a gRNA array.
[0143] A gRNA array comprises more than one gRNA arranged in tandem, and can be processed into two or more individual gRNAs. Thus, in some embodiments, a precursor Cas9 gRNA array comprises two or more (e.g., 3 or more, 4 or more, 5 or more, 2, 3, 4, or 5) gRNAs (e.g., arranged in tandem as a precursor molecule). In some embodiments, two or more gRNAs can be present on an array (a precursor gRNA array). A Cas9 protein of the disclosure can cleave a precursor gRNA array into individual gRNAs.
[0144] In some embodiments, a Cas9 gRNA array comprises 2 or more gRNAs (e.g., 3 or more, 4 or more, 5 or more, 6 or more, or 7 or more gRNAs). The gRNAs of a given array can target different target sites (i.e., can comprise guide sequences that are hybridized thereto) of the same target DNA. In some embodiments, two or more gRNAs of a precursor gRNA array have the same guide sequence. In some embodiments, a precursor gRNA array comprises two or more gRNAs that target different target sites within the same target DNA. In some embodiments, a precursor gRNA array comprises two or more gRNAs that target different target DNAs.
[0145] II. Class 2 Type V CRISPR-Cas RNA-guided systems
[0146] Provided herein are novel Class 2 Type V CRISPR-Cas RNA-guided proteins and their gRNAs, which make up the novel Class 2 Type V CRISPR-Cas RNA-guided systems of the disclosure.
[0147] Provided herein are engineered systems comprising: a Class 2 Type V CRISPR-Cas RNA-guided endonuclease protein and a single guide RNA, wherein the gRNA and the Class 2 Type V CRISPR-Cas RNA-guided endonuclease protein do not naturally occur together, wherein the gRNA is capable of hybridizing to a target sequence in a target DNA, wherein the gRNA is capable of forming a complex with the Class 2 Type V CRISPR-Cas RNA-guided endonuclease protein, and wherein the Class 2 Type V CRISPR-Cas RNA-guided endonuclease protein has collateral activity and is capable of collateral cleaving a single-stranded polynucleotide comprising RNA without the use of a tracrRNA. In some embodiments, the Class 2 Type V CRISPR-Cas RNA-guided endonuclease protein comprises the amino acid sequence of SEQ ID NO: 4 or an amino acid sequence having at least 70% sequence identity to SEQ ID NO: 4. In some embodiments, the Class 2 Type V CRISPR-Cas RNA-guided endonuclease protein is capable of collateral cleaving single-stranded RNA. In some embodiments, the Class 2 Type V CRISPR-Cas RNA-guided endonuclease protein is capable of collateral cleaving single-stranded DNA / RNA hybrids.
[0148] Also provided herein are engineered systems comprising: (a) a Cas12a.1, Cas12p, or Cas12q protein or a nucleic acid encoding the Cas12a.1, Cas12p, or Cas12q protein; and (b) a Cas12a.1, Cas12p, or Cas12q gRNA or a nucleic acid encoding a Cas12a.1, Cas12p, or Cas12q gRNA, wherein the gRNA and the Cas12a.1, Cas12p, or Cas12q protein do not naturally occur together, wherein the gRNA is capable of hybridizing to a target sequence in a target DNA, and the gRNA is capable of forming a complex with the Cas12a.1, Cas12p, or Cas12q protein.
[0149] These components are in turn described below.
[0150] a. Class 2 Type V CRISPR-Cas RNA-guided proteins
[0151] Provided herein are new Class 2 Type V CRISPR-Cas RNA-guided endonucleases, e.g., new Cas12 proteins of the disclosure, including new Cas12a variants and new Cas12 subtypes. In some embodiments, the new Cas12 proteins of the disclosure have been deduced using bioinformatic methods.
[0152] Table 3a shows protein sequences of the new Cas12 proteins of the disclosure. Table 2b shows nucleotide sequences encoding the new Cas12a proteins of the disclosure.
[0153] SEQ ID NO:3 represents a new Cas12a variant of the disclosure, Cas12a.1 (length of 1254 amino acids). Cas12a.1 was isolated from a metagenomics sample and was deduced to be from the Candidatus Micrarcheota archaea. Based on sequence, functional, and structural features, Cas12a.1 is believed to be a Cas12a subtype. FIG. 5A is a schematic representation of the CRISPR Cas cluster surrounding the new Cas12a.1 gene. FIG. 6A shows key catalytic amino acids of Cas12a proteins, as well as alignment of conserved motifs in selected representatives of the Cas12a protein family. FIG. 6B shows alignment of RuvC1, bridge helix, RuvCII, and RuvCIII domains of Cas12a.1, as well as other selected representatives of the Cas12a protein family. SEQ ID NO: 13 shows a nucleotide sequence encoding Cas12a.1 of the disclosure.
[0154] SEQ ID NO:4 represents a new Cas12 subtype of the disclosure, Cas12p (length of 1281 amino acids). Cas12a.1 was isolated from a metagenomics sample and was deduced to be from the Candidatus Peregrinibacteria bacteria. Based on sequence, functional, and structural features described herein, Cas12p is distinct from other members of the Cas12 family identified to date, and thus is a new Cas12 enzyme. This new Cas12 subtype has unique properties not seen in other Cas12 proteins, for example, the ability to nick sequences containing RNA or DNA (e.g., single-stranded DNA, single-stranded RNA, and single-stranded chimeric RNA / DNA) without the use of a tracrRNA. It should be noted that SEQ ID NO:222 in Table 3a is also an N-terminal truncation of Cas12p of SEQ ID NO:4.
[0155] SEQ ID NO: 14 provides a nucleotide sequence encoding Cas12p of the disclosure. FIG. 5C is a schematic representation of the CRISPR Cas cluster surrounding the new Cas12p gene. FIG. 6B .1 shows alignment of Cas12a.1 with SEQ ID NO:81 of US 20160208243, and has 46.8% sequence identity; and FIG. 6CAn alignment of Cas12a.1 with SEQ ID NO: 3 of US 10,253,365 is shown and has 46.5% sequence identity.
[0156] FIG. 6D An amino acid sequence of Cas12p is shown with the RuvC motif underlined (SEQ ID NO: 4). The FnCas12a sequence referenced by Shmakov et al., 2015 was used as a reference to identify the RuvC motif. FIG. 6E An alignment of Cas12p with Cas12gl (another Cas12 enzyme) is shown. This figure shows an alignment of Cas12p with Cas12gl. While Cas12gl has been reported to have the ability to make a collateral RNA (trans-cleavage), there is less than 8.9% sequence identity (as retrieved by the program Clustal Omega). The very low homology between these enzymes and the lack of conserved domains indicate that they are members of different enzyme families. In addition, Cas12gl requires the presence of a tracr sequence, while Cas12p does not, which provides another functional distinction.
[0157] In the following figure, the structure of the Cas12p protein was modeled using the Swiss Model server based on the Fn Cas12a structure. The sequence identity between these proteins is 38.34%. This model covers the entire sequence of the Cas12p protein. FIG. 6F A structural analysis of Cas12p using the Swiss Model server is shown. FIG. 6G A spatial prediction of non-conserved amino acid residues in Cas12p is shown. It can be seen that the non-conserved residues are located on the exposed surface of the protein. These differences can reflect changes in first substrate contact and solvent interactions. FIG. 6H An approximation of the charge distribution on the surface of Cas12p is shown. Using the model shown in FIG. 6F The vacuum electrostatics generated by Pymol software allowed an approximation of the charge distribution on the surface of the protein to be modeled using the model shown in FIG. 6IPredicted structural differences between Cas12p and FnCas12a based on protein sequence are shown. On FnCas12a, the 696-706 region on the PAM-interacting domain is involved in binding and cleavage of the target DNA, and the 842-852 region of the Wedge III region is involved in precursor cRNA processing (Swarts et al., 2017). In comparison to Cas12p, the enzyme presents low homology on those regions given the absence of the sequence KNGNPQKGY (SEQ ID NO: 113) on position 699 and PAKE (SEQ ID NO: 114) on position 844. As these regions are catalytically relevant, the sequence changes can be related to the changes seen on the catalysis. The predicted absence has an impact on the secondary structure of Cas12p. These figures show the superimposition of the model of Cas12p (light grey) and the structure of FnCas12a (dark grey), the absent sequence is shown in black. The absence of the sequence KNGNPQY (SEQ ID NO: 115) is reflected on the loop shortening. The absence of the PAKE sequence (SEQ ID NO: 114; plus other changes on the loop) reduces the loop length and decreases the negative charge on this position for Cas12p. FIG. 6J RuvCIII domain structure analysis of Cas12p based on Swiss Model server structure analysis is shown. FnCas12a sequence mentioned by Shmakov et al., 2015 was used as a reference to identify the Ruv motif. Although the RuvCIII region is conserved on Cas12p and the prototype Cas12a protein, Cas12p has several differences on the sequence surrounding the domain. The presence of these changes has an impact on the secondary structure of Cas12p (shown in black) and can explain the differential RNA cleavage activity of the enzyme. In the structural model depicted in the figure, the superimposition of the structure of the RuvCIII region of the Cas12a enzyme studied and the model of Cas12p. The changes in the secondary structure of Cas12p are circled and shown in black. FIG. 9B 、 FIG. 9C and FIG. 17 The unique collateral cleavage activity of the new Cas12p enzyme is shown.
[0158] SEQ ID NO: 5 represents the new Cas12 of the present disclosure, Cas12q (length of 1137 amino acids). FIG. 5E is a schematic representation of the CRISPR Cas cluster surrounding the new Cas12q gene. FIG. 6KThe Cas12q sequence of the new Cas12 protein of the disclosure, Cas12q, is shown, with the RuvC motif underlined. The FnCas12a sequence referenced by Shmakov et al., 2015 was used as a reference to identify the Ruv motif. SEQ ID NO: 15 shows the nucleotide sequence encoding the Cas12q of the disclosure.
[0159] Table 3a
[0160]
[0161]
[0162]
[0163] As used herein, Cas12a.1 includes SEQ ID NO: 3 and proteins having at least 70-99.5% sequence identity to SEQ ID NO: 3. Accordingly, provided herein are proteins comprising the amino acid sequence of SEQ ID NO: 3 and proteins having at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% sequence identity thereto. Also provided herein are nucleic acids encoding proteins comprising the amino acid sequence of SEQ ID NO: 3 and proteins having at least 70-99.5% sequence identity thereto.
[0164] As used herein, Cas12p includes SEQ ID NO: 4 and proteins having at least 70-99.5% sequence identity to SEQ ID NO: 4. Accordingly, provided herein are proteins comprising the amino acid sequence of SEQ ID NO: 4 and proteins having at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% sequence identity thereto. Also provided herein are nucleic acids encoding proteins comprising the amino acid sequence of SEQ ID NO: 4 and proteins having at least 70-99.5% sequence identity thereto.
[0165] Also provided herein are proteins comprising the amino acid sequence of SEQ ID NO: 222 and proteins having at least 70-99.5% sequence identity thereto. Also provided herein are nucleic acids encoding proteins comprising the amino acid sequence of SEQ ID NO: 222 and proteins having at least 70-99.5% sequence identity thereto.
[0166] As used herein, Cas12q includes SEQ ID NO: 5 and proteins having at least 70-99.5% sequence identity to SEQ ID NO: 5. Accordingly, provided herein are proteins comprising the amino acid sequence of SEQ ID NO: 5 and proteins having at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% sequence identity thereto. Also provided herein are nucleic acids encoding proteins comprising the amino acid sequence of SEQ ID NO: 5 and proteins having at least 70-99.5% sequence identity thereto.
[0167] Table 3b shows exemplary nucleotide sequences and exemplary codon-optimized nucleic acid sequences of the new Cas12 proteins of the disclosure.
[0168] Table 3b
[0169]
[0170]
[0171]
[0172]
[0173]
[0174]
[0175]
[0176]
[0177]
[0178] Table 4a shows the structural and functional features of the new Cas12 proteins of the present disclosure as exemplified herein. Table 4b shows the number and sequence of the natural spacers of the corresponding CRISPR array. Blank cells in the table do not indicate the absence of a value / property, but that it has not been exemplified herein.
[0179] Table 4a
[0180]
[0181]
[0182] Table 4b
[0183]
[0184]
[0185] In some embodiments, the Cas12 protein of the present disclosure is a catalytically active Cas12 protein, e.g., a catalytically active Cas12a.1, Cas12p, or Cas12q protein.
[0186] In some embodiments, the Cas12 protein of the present disclosure cleaves at a site distal to the target sequence, e.g., a Cas12a.1, Cas12p, or Cas12q protein cleaves at a site distal to the target sequence.
[0187] In some embodiments, the Cas12 protein of the present disclosure is a catalytically inactive Cas12 protein, e.g., a Cas12a.1, Cas12p, or Cas12q protein is catalytically inactive (dCas12a.1, dCas12p, or dCas12q protein).
[0188] In some embodiments, the Cas12 proteins of the present disclosure are nickase Cas12 proteins, e.g., Cas12a.1 nickase, Cas12p nickase, or Cas12q nickase proteins.
[0189] In some embodiments, the Cas12 proteins of the present disclosure can be modified to comprise aptamers.
[0190] In some embodiments, the Cas12 proteins of the present disclosure can be further fused to a domain (e.g., a catalytic domain) to generate dual acting Cas proteins. In some embodiments, the Cas12a proteins are further fused to a base editor.
[0191] b.2 class V type CRISPR-Cas RNA-guided protein collateral activity
[0192] In addition to the ability to cleave target sequences in targeted DNA, the Cas12 proteins of the present disclosure have collateral ability (trans-cleavage activity), i.e., the ability to promiscuously cleave non-targeted single-stranded DNA (ssDNA) or RNA once activated upon detection of a target DNA. Without being bound by any theory or mechanism, in general, once the Cas12 proteins of the present disclosure are activated by a gRNA when a sample contains a target sequence hybridized to the gRNA (i.e., the sample contains targeted DNA), Cas12 can become a nuclease that randomly cleaves oligonucleotides (e.g., ssDNA, RNA, chimeric RNA / DNA) that do not contain a target sequence of the gRNA (non-targeted oligonucleotides, to which the guide sequence of the gRNA is not hybridized). Thus, when targeted DNA (double-stranded or single-stranded) is present in a sample (e.g., in some embodiments, above a threshold amount), the result can be that single-stranded oligonucleotides (e.g., ssDNA, ssRNA, single-stranded chimeric RNA / DNA) in the sample are cleaved, which can be detected using any convenient detection method (e.g., using a labeled detector DNA, RNA, or DNA / RNA chimera).
[0193] Accordingly, provided herein are methods and compositions for detecting targeted DNA (dsDNA or ssDNA) in a sample. Also provided are methods and compositions for cleaving non-targeted oligonucleotides, which can utilize a detector. These embodiments are described in further detail below.
[0194] c. gRNAs for 2 class V type CRISPR-Cas RNA-guided proteins
[0195] The present disclosure provides DNA-targeting RNAs that direct the activity of the novel Cas12 proteins of the present disclosure to a specific target sequence within a target DNA. As described above for the novel Cas9 proteins of the present disclosure, these DNA-targeting RNAs are referred to herein as “gRNAs” or “gRNAs.” In general, as provided herein, gRNAs for Cas12 comprise a single segment containing both a spacer (DNA-targeting sequence) and a Cas12a “protein-binding sequence” (together referred to as a crRNA). Nucleotide sequences encoding the Cas12a gRNAs of the present disclosure are also provided herein.
[0196] i. spacer sequence
[0197] The Cas12 proteins of the present disclosure are single-crRNA guided endonucleases (single guide RNA, sgRNA), whereas the Cas9 proteins of the present disclosure are guided by a dual RNA system consisting of a crRNA and a trans-activating crRNA (tracrRNA). The crRNA of the Cas12 guides of the present disclosure comprises a nucleotide sequence that is complementary to a sequence in the target DNA (DNA-targeting sequence or spacer).
[0198] The crRNA portion of the Cas12 gRNAs of the present disclosure can be about 25-50 nt in length. In some embodiments, the length can be about 40-43 nt.
[0199] The mature guide scaffold for Cas12a.1 and Cas12p were deduced on the computer from the corresponding CRISPR locus. FIG. 38 The secondary structure of the scaffold for Cas12a.1 (5’ aaauuucuacuguaguagau 3’) (SEQ ID NO: 116; panel A) and Cas12p (5’ agauuucuacuuuuguagau 3’) (SEQ ID NO: 117; panel B) is shown. These mature scaffolds can then be linked to a variable targeting spacer sequence, resulting in an sgRNA. Thus, in some embodiments, provided herein is an engineered single-molecule gRNA comprising the scaffold sequence of SEQ ID NO: 116 or SEQ ID NO: 117 and a spacer sequence capable of hybridizing to a target sequence in a target DNA. In some embodiments, the target DNA is viral DNA, plant DNA, fungal DNA, or bacterial DNA. In some embodiments, the target sequence is the sequence of a target provided in any one of Tables 6a-6f. In some embodiments, the target is a coronavirus. In some embodiments, the target is a SARS-CoV-2 virus. In some embodiments, the target DNA is cDNA and has been obtained by reverse transcription.
[0200] The DNA targeting spacer sequence of a Cas12 gRNA typically interacts with a target DNA in a sequence-specific manner through hybridization (i.e., base pairing). Thus, the nucleotide sequence of the DNA targeting sequence can vary and determines the location within the target DNA where the gRNA and the target DNA interact. The DNA targeting sequence of a subject Cas12 gRNA can be modified (e.g., by genetic engineering) to hybridize with a desired sequence within a target DNA.
[0201] The DNA targeting sequence of a subject Cas12 gRNA can be about 8 nucleotides to about 30 nucleotides in length. For example, the length can be 23 nucleotides.
[0202] The percent complementarity between the DNA targeting spacer sequence of a crRNA and a target sequence of a target DNA can be at least 60% (e.g., at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99%, or 100%). In some embodiments, the percent complementarity between the DNA targeting sequence of a crRNA and a target sequence of a target DNA is 100% over the 1-23 contiguous most 5' nucleotides of the target sequence of the complementary strand of the target DNA. In some embodiments, the percent complementarity between the DNA targeting sequence of a crRNA and a target sequence of a target DNA is at least 60% over about 1-23 contiguous nucleotides. In some embodiments, the percent complementarity between the DNA targeting sequence of a crRNA and a target sequence of a target DNA is 100% over the 1-23 contiguous most 5' nucleotides of the target sequence of the complementary strand of the target DNA and 0% over the remainder. In this case, the length of the DNA targeting sequence can be considered to be 1-23 nucleotides.
[0203] Generally, the native unprocessed precursor crRNA of a Cas12 comprises a forward repeat and an adjacent spacer (the portion of the crRNA that allows for targeting of a DNA molecule). In some embodiments, the forward repeat from the unprocessed precursor crRNA and the forward repeat mutation are included in a Cas12 gRNA of the disclosure to improve gRNA stability.
[0204] Table 5a shows the predicted (putative) naturally occurring forward repeat sequences in the CRISPR locus (as found in bacterial DNA) of Cas12 proteins of the disclosure. These are the predicted natural sequences in the CRISPR locus contig (as found in bacterial DNA). A gRNA of the disclosure has a portion of the forward repeat linked to the spacer.
[0205] Table 5a: Forward repeat sequences
[0206]
[0207] In some embodiments, the crRNA comprises a non-naturally occurring engineered forward repeat sequence. Table 5b shows non-naturally occurring engineered forward repeat sequences that can be incorporated into the engineered gRNAs of the present disclosure.
[0208] Predicted RNA secondary structures of predicted non-naturally occurring engineered forward repeat sequences are shown in FIG. 7A to FIG. 7C .
[0209] Table 5b
[0210]
[0211]
[0212] In some embodiments, the spacer sequence of a Cas12 gRNA of the present disclosure is directed to a target sequence in a mammalian organism. In some embodiments, the spacer sequence is directed to a target sequence in a non-mammalian organism.
[0213] In some embodiments, the spacer sequence of a Cas12 gRNA of the present disclosure is directed to a target sequence that is a sequence of a human. In some embodiments, the target sequence is a sequence of a non-human primate.
[0214] In some embodiments, the spacer sequence of a Cas12 gRNA of the present disclosure is directed to a target sequence in a mammalian organism (e.g., a human or a non-human primate).
[0215] In some embodiments, the spacer sequence of a Cas12 gRNA of the present disclosure is directed to a target sequence in a bacterium.
[0216] In some embodiments, the spacer sequence of a Cas12 gRNA of the present disclosure is directed to a target sequence in a virus.
[0217] In some embodiments, the spacer sequence of a Cas12 gRNA of the present disclosure is directed to a target sequence in a plant.
[0218] A Cas12 gRNA of the present disclosure can be modified to comprise an aptamer.
[0219] ii. PAM specificity
[0220] TCTN and TGTN were identified as effective PAM sequences for Cas12a.1 and Cas12p, respectively.
[0221] iii. gRNA arrays
[0222] In some embodiments, the Cas12 gRNAs of the disclosure can be provided as gRNA arrays.
[0223] Such gRNA arrays of the disclosure comprise more than one gRNA arranged in tandem, and can be processed into two or more individual gRNAs. Thus, in some embodiments, a precursor Cas12 gRNA array comprises two or more (e.g., 3 or more, 4 or more, 5 or more, 2, 3, 4, or 5) gRNAs (e.g., arranged in tandem as a precursor molecule). In some embodiments, two or more gRNAs can be present on an array (a precursor gRNA array). Cas12 proteins of the disclosure can cleave a precursor gRNA array into individual gRNAs.
[0224] In some embodiments, a Cas12 gRNA array comprises 2 or more gRNAs (e.g., 3 or more, 4 or more, 5 or more, 6 or more, or 7 or more gRNAs). The gRNAs of a given array can target different target sites (i.e., can comprise a guide sequence that is hybridized thereto) of the same target DNA. In some embodiments, two or more gRNAs of a precursor gRNA array have the same guide sequence. In some embodiments, a precursor gRNA array comprises two or more gRNAs that target different target sites within the same target DNA. In some embodiments, a precursor gRNA array comprises two or more gRNAs that target different target DNAs.
[0225] III. Methods of use - modification and therapeutic agents
[0226] a. Modification of target DNA
[0227] Provided herein are uses of the novel Cas9 and Cas12 proteins of the disclosure. Thus, provided herein is a method of modifying a target DNA, the method comprising contacting the target DNA with any of the Cas9 systems or Cas12 systems described herein. These methods can be used in therapeutic applications
[0228] In some embodiments, the target DNA is part of a chromosome in vitro. In some embodiments, the target DNA is part of a chromosome in vivo.
[0229] In some embodiments, the target DNA is part of a chromosome in a cell.
[0230] In some embodiments, the target DNA is extrachromosomal DNA.
[0231] In some embodiments, the target DNA is in a cell, wherein the cell is selected from the group consisting of an archaeal cell, a bacterial cell, a eukaryotic cell, a eukaryotic single-celled organism, a somatic cell, a germ cell, a stem cell, a plant cell, an algal cell, an animal cell, an invertebrate animal cell, a vertebrate animal cell, a fish cell, a frog cell, a bird cell, a mammalian cell, a pig cell, a cow cell, a goat cell, a sheep cell, a rodent cell, a rat cell, a mouse cell, a non-human primate cell, and a human cell.
[0232] In some embodiments, the target DNA is DNA of a parasite.
[0233] In some embodiments, the target DNA is viral DNA.
[0234] In some embodiments, the target DNA is bacterial DNA.
[0235] In some embodiments, the modification comprises introducing a double-strand break in the target DNA.
[0236] In some embodiments, the contacting occurs under conditions that allow for non-homologous end joining or homology directed repair.
[0237] In some embodiments, the method comprises contacting the target DNA with a donor polynucleotide, wherein the donor polynucleotide, a portion of the donor polynucleotide, a copy of the donor polynucleotide, or a portion of a copy of the donor polynucleotide is integrated into the target DNA.
[0238] In some embodiments, the method does not comprise contacting the cell with a donor polynucleotide, wherein the target DNA is modified such that nucleotides within the target DNA are deleted.
[0239] b. Therapeutic applications
[0240] The present disclosure provides novel Cas9 proteins, novel Cas12a proteins and novel Cas12 protein subtypes, engineered systems, one or more polynucleotides encoding components of the systems, and vectors or delivery systems comprising one or more polynucleotides encoding components of the systems, for use in therapeutic methods. Therapeutic methods can include gene or genome editing, or gene therapy. Therapeutic methods include the use and delivery of the novel Cas9 and Cas12 proteins of the present disclosure. Accordingly, in some embodiments, provided herein is a method of modifying a target DNA, the method comprising contacting a target DNA, a cell comprising a target DNA, or a subject having a cell comprising a target DNA, with any one of the Cas9 systems or Cas12 systems described herein.
[0241] In some embodiments, the target DNA is part of a chromosome in vitro. In some embodiments, the target DNA is part of a chromosome in vivo.
[0242] In some embodiments, the target DNA is part of a chromosome in a cell.
[0243] In some embodiments, the target DNA is extrachromosomal DNA.
[0244] In some embodiments, the target DNA is in a cell, wherein the cell is selected from the group consisting of an archaeal cell, a bacterial cell, a eukaryotic cell, a eukaryotic unicellular organism, a somatic cell, a germ cell, a stem cell, a plant cell, an algal cell, an animal cell, an invertebrate animal cell, a vertebrate animal cell, a fish cell, a frog cell, a bird cell, a mammalian cell, a pig cell, a cow cell, a goat cell, a sheep cell, a rodent cell, a rat cell, a mouse cell, a non-human primate cell, and a human cell.
[0245] In some embodiments, the target DNA is extracellular.
[0246] In some embodiments, the target DNA is in a cell in vitro.
[0247] In some embodiments, the target DNA is in a cell in vivo.
[0248] In some embodiments, the modification comprises introducing a double-strand break in the target DNA.
[0249] In some embodiments, the contacting occurs under conditions that allow for non-homologous end joining or homology directed repair.
[0250] In some embodiments, the method comprises contacting the target DNA with a donor polynucleotide, wherein the donor polynucleotide, a portion of the donor polynucleotide, a copy of the donor polynucleotide, or a portion of a copy of the donor polynucleotide is integrated into the target DNA.
[0251] In some embodiments, the method does not comprise contacting the cell with a donor polynucleotide, wherein the target DNA is modified such that nucleotides within the target DNA are deleted.
[0252] In some embodiments, the method of treatment comprises modifying a target DNA comprising a target sequence of a gene of interest and / or a regulatory region of a gene of interest, the method comprising delivering to a cell, including the target DNA, a Cas9 protein of the disclosure and one or more Cas9 gRNAs, a Cas12 protein of the disclosure and one or more Cas12 gRNAs, one or more nucleotides encoding a Cas9 protein of the disclosure and one or more Cas9 gRNAs, or one or more nucleotides encoding a Cas12 protein of the disclosure and one or more Cas12 gRNAs.
[0253] In some embodiments, the gene of interest is within a eukaryotic cell, e.g., a human or non-human primate cell.
[0254] In some embodiments, the gene of interest is within a plant cell.
[0255] In some embodiments, the delivering comprises delivering to the cell a Cas9 protein of the disclosure (or one or more nucleotides encoding the protein) and one or more Cas9 gRNAs.
[0256] In some embodiments, the delivering comprises delivering to the cell a Cas12 protein of the disclosure (or one or more nucleotides encoding the protein) and one or more Cas12 gRNAs.
[0257] In some embodiments, the delivering comprises delivering to the cell one or more nucleotides encoding a Cas9 protein of the disclosure and one or more Cas9 gRNAs.
[0258] In some embodiments, the delivering comprises delivering to the cell one or more nucleotides encoding a Cas12 protein of the disclosure and one or more Cas12 gRNAs.
[0259] Delivery of Cas9 or Cas12 components to a cell can be achieved by any of a variety of delivery methods known to those of skill in the art. As a non-limiting example, the components can be combined with lipids. As another non-limiting example, the components are combined with or formulated into particles, e.g., nanoparticles.
[0260] Methods of introducing nucleic acids and / or proteins into host cells are known in the art, and any convenient method can be used to introduce the subject nucleic acids (e.g., expression constructs / vectors) into target cells (e.g., prokaryotic cells, eukaryotic cells, plant cells, animal cells, mammalian cells, human cells, etc.). Suitable methods include, for example, viral infection, transfection, conjugation, protoplast fusion, lipofection, electroporation, calcium phosphate precipitation, polyethylenimine (PEI)-mediated transfection, DEAE-dextran-mediated transfection, liposome-mediated transfection, particle gun technology, calcium phosphate precipitation, direct microinjection, nanoparticle-mediated nucleic acid delivery, and the like.
[0261] The gRNA can be introduced, e.g., as a DNA molecule encoding the gRNA, or can be provided directly as an RNA molecule (or, where applicable, a chimeric / hybrid molecule).
[0262] In some embodiments, the Cas9 or Cas12 protein is provided as a nucleic acid encoding the protein (e.g., mRNA, DNA, plasmid, expression vector, viral vector, etc.).
[0263] In some embodiments, the Cas9 or Cas12 protein is provided directly as a protein (e.g., without an associated gRNA or with an associated gRNA, i.e., a ribonucleoprotein complex - RNP). As with the gRNA, the Cas9 or Cas12 protein of the disclosure can be introduced into a cell (provided to a cell) by any convenient method; such methods are known to those of ordinary skill in the art. As an illustrative example, the Cas9 or Cas12 protein of the disclosure can be injected directly into a cell (e.g., with or without a gRNA or nucleic acid encoding a gRNA). As another example, a preformed complex of the Cas9 or Cas12 protein and the gRNA can be introduced into a cell (e.g., a eukaryotic cell) (e.g., by injection, by nucleofection; by a protein transduction domain (PTD) conjugated to one or more components, e.g., conjugated to the Cas9 or Cas12 protein of the disclosure, conjugated to the gRNA; etc.).
[0264] In some embodiments, a nucleic acid (e.g., a gRNA; a nucleic acid comprising a nucleotide sequence encoding a Cas9 or Cas12 protein of the disclosure; etc.) and / or a polypeptide (e.g., a Cas9 or Cas12 protein of the disclosure) is delivered to a cell (e.g., a target host cell) in a particle, or is associated with a particle. In some embodiments, the particle is a nanoparticle.
[0265] The Cas9 or Cas12 protein (or mRNA comprising a nucleotide sequence encoding the protein) and / or the gRNA (or nucleic acid encoding the gRNA, such as one or more expression vectors) of the disclosure can be delivered simultaneously using a particle or a lipid envelope.
[0266] i. Target cell of interest
[0267] Suitable target cells, which can comprise target DNA, such as genomic DNA, include, but are not limited to: bacterial cells; archaeal cells; cells of single-celled eukaryotes; plant cells; algal cells, e.g., Botryococcus Braunii, Chlamydomonas Reinhardtii, Nannchloropsis gaditana, Chlorella pyrenoidosa, Sargassum patens, C. agardh, and the like; fungal cells (e.g., yeast cells); animal cells; cells of invertebrates (e.g., fruit flies, cnidarians, echinoderms, nematodes, and the like); cells of insects (e.g., mosquitoes; bees; agricultural pests; and the like); cells of arachnids (e.g., spiders; ticks; and the like); cells of vertebrates (e.g., fish, amphibians, reptiles, birds, mammals); cells of mammals (e.g., cells from rodents; human cells; cells of non-human mammals; cells of rodents (e.g., mice, rats); cells of lagomorphs (e.g., rabbits); cells of ungulates (e.g., cows, horses, camels, llamas, tapirs, sheep, goats, and the like); cells of marine mammals (e.g., whales, seals, elephant seals, dolphins, sea lions, and the like); and the like.
[0268] Any type of cell can be a cell of interest (e.g., stem cells, such as embryonic stem (ES) cells, induced pluripotent stem cells (iPSCs), germ cells (e.g., oocytes, sperm, oogonia, spermatogonia, and the like), somatic stem cells, somatic cells (e.g., fibroblasts), hematopoietic cells, neurons, muscle cells, bone cells, liver cells, pancreatic cells; in vitro or in vivo embryonic cells of an embryo at any stage, such as 1-cell, 2-cell, 4-cell, 8-cell, and the like stage zebrafish embryos; and the like).
[0269] Cells can be from a cell line or primary cells. Target cells can be unicellular organisms and / or can be grown in culture. If the cells are primary cells, they can be harvested from an individual by any convenient method. For example, white blood cells can be conveniently harvested by apheresis, leukapheresis, density gradient separation, and the like, while cells from tissues of the skin, muscle, bone marrow, spleen, liver, pancreas, lung, intestine, stomach, and the like can be conveniently harvested by biopsy.
[0270] Because the gRNA provides specificity by hybridizing to a target nucleic acid, the mitotic cell and / or post-mitotic cell of interest in the disclosed methods can include cells of any organism (e.g., bacterial cells, archaeal cells, cells of single-celled eukaryotes, plant cells, algal cells (e.g., Botrydiopsis arbuscula, Chlamydomonas reinhardtii, Nannochloropsis gaditana, Chlorella variabilis, Sargassum patens, Agardhiella sp., etc.), fungal cells (e.g., yeast cells), animal cells, cells of invertebrates (e.g., fruit flies, cnidarians, echinoderms, nematodes, etc.), cells of vertebrates (e.g., fish, amphibians, reptiles, birds, mammals), cells of mammals, cells of rodents, cells of humans, etc.).
[0271] Plant cells include cells of monocotyledonous plants and cells of dicotyledonous plants. Cells can be root cells, leaf cells, xylem cells, phloem cells, cambium cells, apical meristem cells, parenchymal cells, parenchymal cells, sclerenchymal cells, sclerotic cells, etc. Plant cells include cells of agricultural crops such as wheat, corn, rice, sorghum, millet, soybean, etc. Plant cells include cells of agricultural fruit and nut plants such as plants that produce apricots, oranges, lemons, apples, plums, pears, almonds, etc.
[0272] Non-limiting examples of cells (target cells) include: prokaryotic cells, eukaryotic cells, bacterial cells, archaeal cells, cells of single-celled eukaryotes, cells of plants (e.g., cells of plant crops, fruits, vegetables, grains, soybeans, corn, maize, wheat, seeds, tomatoes, rice, cassava, sugar cane, pumpkin, hay, potatoes, cotton, hemp, tobacco, flowering plants, gymnosperms, angiosperms, ferns, clubmosses, hornworts, liverworts, mosses, dicotyledons, monocotyledons, etc.), algal cells (e.g., Botrydiopsis arbuscula, Chlamydomonas reinhardtii, Nannochloropsis gaditana, Chlorella variabilis, Sargassum patens, Agardhiella sp., etc.), seaweed (e.g., kelp), fungal cells (e.g., yeast cells, mushroom cells), animal cells, cells of invertebrates (e.g., fruit flies, cnidarians, echinoderms, nematodes, etc.), cells of vertebrates (e.g., fish, amphibians, reptiles, birds, mammals), cells of mammals (e.g., cells of ungulates (e.g., pigs, cows, goats, sheep); rodents (e.g., rats, mice); non-human primates; humans; felines (e.g., cats); canines (e.g., dogs) etc.), etc. In some embodiments, the cells are cells that are not derived from a natural organism (e.g., the cells can be synthetically manufactured cells; also referred to as artificial cells).
[0273] The cell can be an in vitro cell (e.g., an established cell line). The cell can be an ex vivo cell (a cultured cell from an individual). The cell can be an in vivo cell (e.g., a cell in an individual). The cell can be an isolated cell. The cell can be a cell within an organism. The cell can be an organism.
[0274] Suitable cells include human embryonic stem cells, fetal cardiomyocytes, myofibroblasts, mesenchymal stem cells, autograft expanded cardiomyocytes, adipocytes, totipotent cells, pluripotent cells, blood stem cells, myoblasts, adult stem cells, bone marrow cells, mesenchymal cells, embryonic stem cells, parenchymal cells, epithelial cells, endothelial cells, mesothelial cells, fibroblasts, osteoblasts, chondrocytes, xenogeneic cells, homologous cells, and postnatal stem cells.
[0275] In some embodiments, the cell is an immune cell, a neuron, an epithelial cell, and an endothelial cell, or a stem cell. In some embodiments, the immune cell is a T cell, a B cell, a monocyte, a natural killer cell, a dendritic cell, or a macrophage. In some embodiments, the immune cell is a cytotoxic T cell. In some embodiments, the immune cell is a helper T cell. In some embodiments, the immune cell is a regulatory T cell (Treg).
[0276] In some embodiments, the cell is a stem cell. Stem cells include adult stem cells. Adult stem cells are also known as somatic stem cells.
[0277] Adult stem cells reside in differentiated tissues, but retain the ability to self-renew and generate multiple cell types, typically the cell types typical of the tissue in which the stem cell is found. Many examples of somatic stem cells are known to those of skill in the art, including muscle stem cells; hematopoietic stem cells; epithelial stem cells; neural stem cells; mesenchymal stem cells; mammary stem cells; intestinal stem cells; mesodermal stem cells; endothelial stem cells; olfactory stem cells; neural crest stem cells; and the like.
[0278] Stem cells of interest include mammalian stem cells, where the term "mammal" refers to any animal classified as a mammal, including humans; non-human primates; domestic and farm animals; and zoo, laboratory, sport, or pet animals, such as dogs, horses, cats, cows, mice, rats, rabbits, etc. In some embodiments, the stem cell is a human stem cell. In some embodiments, the stem cell is a rodent (e.g., mouse; rat) stem cell. In some embodiments, the stem cell is a non-human primate stem cell.
[0279] ii. Target
[0280] Any gene of interest can be a target for modification.
[0281] In particular embodiments, the target is a cancer implicated gene. In particular embodiments, the target is an immune disease (e.g., an autoimmune disease) implicated gene. In particular embodiments, the target is a neurodegenerative disease implicated gene. In particular embodiments, the target is a neuropsychiatric disease implicated gene. In particular embodiments, the target is a muscle disease implicated gene. In particular embodiments, the target is a cardiac disease implicated gene. In particular embodiments, the target is a diabetes implicated gene. In particular embodiments, the target is a kidney disease implicated gene.
[0282] iii. Precursor gRNA array
[0283] The therapeutic methods provided herein can include delivery of a precursor gRNA array. The Cas9 or Cas12 proteins of the disclosure can cleave a precursor gRNA into a mature gRNA, for example by endoribonuclease cleavage of the precursor. The Cas9 or Cas12 proteins of the disclosure can cleave a precursor gRNA array (including a plurality of gRNAs arranged in tandem) into two or more individual gRNAs.
[0284] IV. Methods of use - detection and diagnostic applications
[0285] In addition to the ability to cleave a target sequence in a targeted DNA, the Cas12 proteins of the disclosure have collateral activity (trans-cleavage activity), i.e., the ability to promiscuously cleave non-targeted oligonucleotides (ssDNA, RNA, DNA / RNA hybrids) once activated upon detection of a target DNA. Without being bound by any theory or mechanism, generally speaking, once the Cas12 proteins of the disclosure are activated by a gRNA when a sample contains a target sequence hybridized to the gRNA (i.e., the sample contains a targeted DNA), Cas12 becomes a nuclease that randomly cleaves single-stranded oligonucleotides (i.e., non-target single-stranded oligonucleotides, i.e., single-stranded oligonucleotides that are not hybridized to the guide sequence of the gRNA). Thus, when targeted DNA (double-stranded or single-stranded) is present in a sample (e.g., in some embodiments, above a threshold amount), the result can be cleavage (collateral cleavage) of oligonucleotides in the sample, which can be detected using any convenient detection method (e.g., using a labeled single-stranded detector DNA, a labeled detector RNA, or a labeled detector DNA / RNA chimeric oligonucleotide).
[0286] Accordingly, provided herein are methods and compositions for detecting a target DNA (dsDNA or ssDNA) in a sample. Also provided are methods and compositions for cleaving a non-target oligonucleotide (e.g., as a detector).
[0287] As used herein, a “detector” includes a single-stranded or double-stranded oligonucleotide of any nature and is not hybridized to the guide sequence of a gRNA (i.e., the detector oligonucleotide is not a target). Exemplary detectors include, but are not limited to, ssDNA, dsDNA, ssRNA, ssDNA / RNA chimera, dsRNA, RNA comprising ss or ds regions, and RNA and DNA nucleotides containing ss or ds oligonucleotides (as used herein, ss = single-stranded; and ds = double-stranded).
[0288] Methods of detection based on the collateral cleavage activity of a Cas12 protein of the disclosure can include:
[0289] (a) contacting the sample with: (i) a Cas12 protein of the disclosure; (ii) a gRNA comprising a region that binds the Cas12 protein and a guide sequence that is hybridized to the target DNA; and (iii) a detector that is not hybridized to the guide sequence of the gRNA; and
[0290] (b) measuring a detectable signal resulting from cleavage of the detector by the Cas12 protein, thereby detecting the target DNA.
[0291] Once the subject Cas12 protein is activated by the gRNA activation event when the sample comprises a target DNA that is hybridized to the gRNA (i.e., the sample comprises the targeted sequence in the target DNA), the Cas12 can be activated to function as an endoribonuclease that non-specifically cleaves detector oligonucleotides present in the sample (including non-target ss oligonucleotides). Thus, when the target DNA is present in the sample, the result is cleavage of the detector oligonucleotides in the sample, which can be detected using any convenient method of detection (e.g., using a labeled detector oligonucleotide).
[0292] Also provided are methods and compositions for cleaving a detector oligonucleotide (e.g., ssDNA, ssRNA, ssDNA / RNA chimera, or a detector comprising ss and ds regions). These methods can include contacting a population of nucleic acids (wherein the population comprises a target DNA and a plurality of non-target ss oligonucleotides) with: (i) a Cas12 protein of the disclosure; and (ii) a gRNA comprising a region that binds the Cas12 effector protein and a guide sequence that is hybridized to the target DNA, wherein the Cas12 protein cleaves non-target ss oligonucleotides
[0293] Accordingly, provided herein is a method of detecting a target DNA in a sample, the method comprising:
[0294] (a) contacting the sample with:
[0295] (i) a Cas12 protein of the disclosure (e.g., a Cas12a.1, Cas12p, or Cas12q protein);
[0296] (ii) a gRNA comprising a spacer sequence capable of hybridizing to a target sequence in the target DNA; and
[0297] (iii) a labeled detector oligonucleotide that does not hybridize to the spacer sequence of the gRNA; and
[0298] (b) measuring a detectable signal resulting from cleavage of the labeled detector oligonucleotide by the Cas12 protein, thereby detecting the target oligonucleotide.
[0299] In some embodiments, the method further comprises the above described in connection with detecting a positive control target DNA in a positive control sample, the detecting comprising the additional step of:
[0300] (c) contacting the positive control sample with:
[0301] (i) a Cas12 protein of the disclosure (e.g., a Cas12a.1, Cas12p, or Cas12q protein);
[0302] (ii) a positive control gRNA comprising a region that binds the Cas12a.1, Cas12p, or Cas12q protein and a positive control spacer sequence that hybridizes to the positive control target DNA; and
[0303] (iii) a labeled detector oligonucleotide that does not hybridize to the positive control spacer sequence of the positive control gRNA; and
[0304] (d) measuring a detectable signal resulting from cleavage of the labeled detector by the Cas12 protein, thereby detecting the positive control target DNA.
[0305] In some embodiments, the contacting step can be performed in a non-cellular environment, e.g., extracellularly. In other embodiments, the contacting step can be performed intracellularly. The contacting step can be performed in a cell in vitro. The contacting step can be performed in a cell in vivo. The contacting step of the detection method can be performed in a composition comprising a divalent metal ion.
[0306] The gRNA can be provided as RNA, or as a nucleic acid (e.g., DNA, such as a recombinant expression vector) that encodes a gRNA as described herein.
[0307] The contacting prior to the measuring step can be for any period of time prior to the measuring step, such as 5 seconds to 2 hours or more. In some embodiments, the sample is contacted for 45 minutes or less prior to the measuring step. In some embodiments, the sample is contacted for 30 minutes or less prior to the measuring step. In some embodiments, the sample is contacted for 10 minutes or less prior to the measuring step. In some embodiments, the sample is contacted for 5 minutes or less prior to the measuring step. In some embodiments, the sample is contacted for 1 minute or less prior to the measuring step. In some embodiments, the sample is contacted for 50 seconds to 60 seconds prior to the measuring step. In some embodiments, the sample is contacted for 40 seconds to 50 seconds prior to the measuring step. In some embodiments, the sample is contacted for 30 seconds to 40 seconds prior to the measuring step. In some embodiments, the sample is contacted for 20 seconds to 30 seconds prior to the measuring step. In some embodiments, the sample is contacted for 10 seconds to 20 seconds prior to the measuring step.
[0308] The detection methods provided herein can detect target DNA with high sensitivity. Thus, in some embodiments, the detection methods of the disclosure can be used to detect target DNA present in a sample comprising a plurality of DNA, including target DNA and a plurality of non-target DNA, wherein one or more copies of the target DNA are present per 5 to 10^9 copies of the non-target DNA
[0309] In some embodiments, the detection methods detect target DNA in a sample with a detection threshold of 10 nM or less. The term "detection threshold" is used herein to describe the minimum amount of target DNA that must be present in a sample for detection to occur. Thus, as an illustrative example, when the detection threshold is 10 nM, then a signal can be detected when the target DNA is present in the sample at a concentration of 10 nM or greater. In some embodiments, the subject compositions or methods exhibit detection sensitivity in the attomolar (aM) range. In some embodiments, the subject compositions or methods exhibit detection sensitivity in the femtomolar (fM) range. In some embodiments, the subject compositions or methods exhibit detection sensitivity in the picomolar (pM) range. In some embodiments, the subject compositions or methods exhibit detection sensitivity in the nanomolar (nM) range.
[0310] a. target DNA
[0311] The target DNA can be single-stranded (ssDNA) or double-stranded (dsDNA). There is no preference or requirement for PAM sequence in single-stranded target DNA.
[0312] The source of the target DNA can be any source. In some embodiments, the target DNA is viral or bacterial DNA (e.g., genomic DNA of a DNA virus or bacterium). Thus, the detection method can be used to detect the presence of viral or bacterial DNA in a population of nucleic acids (e.g., in a sample). In the case of an RNA-carrying organism (e.g., an RNA virus (e.g., a coronavirus)), it will be appreciated that a step such as reverse transcription can be performed on a sample comprising the RNA-carrying organism to generate cDNA, and for the purposes of the present disclosure, the cDNA is the target DNA.
[0313] Exemplary non-limiting sources of target DNA are provided in Tables 6a-6f.
[0314] Table 6a
[0315] Bacterial resistance gene targets KPC: Class A beta-lactamase that hydrolyzes carbapenems NDM: Metallo-beta-lactamase OXA: Class D beta-lactamase that hydrolyzes oxacillin MecA: PBP2a family beta-lactam resistance peptidoglycan transpeptidase vanA / B: Vancomycin resistance
[0316] Table 6b
[0317] Viral genome targets Dengue (DENV) fever virus (subtypes 1, 2, 3, and 4) Zika Virus Chikungunya virus Coronavirus
[0318] Respiratory system targets
[0319] DNA obtained from viruses and bacteria associated with respiratory infections can also be targeted. A list of targets of interest can include the examples shown in Table 6c.
[0320] Table 6c
[0321] Respiratory targets Adenovirus Coronavirus SARS-CoV SARS-CoV-2 MERS-CoV Coronavirus HKU1 Coronavirus NL63 Coronavirus 229E Coronavirus OC43 Coronavirus HKU1 Human metapneumovirus Human rhinovirus / enterovirus Influenza A Influenza A / H1 Influenza A / H3 Influenza A / H1-2009 Influenza B Parainfluenza virus 1 Parainfluenza virus 2 Parainfluenza virus 3 Parainfluenza virus 4 Respiratory syncytial virus Bacteria: Bordetella parapertussis Bordetella pertussis Chlamydophila pneumoniae Mycoplasma pneumoniae
[0322] Sexually transmitted disease targets
[0323] DNA obtained from viruses and bacteria associated with sexually transmitted diseases can also be targeted. A list of targets of interest can include the examples shown in Table 6d.
[0324] Table 6d
[0325] Sexually transmitted disease targets HIV (types 1 and 2) Herpes simplex virus 1 (HSV-1) Herpes simplex virus 2 (HSV-2) Hepatitis A Hepatitis B Hepatitis C Bacteria: Treponema pallidum Chlamydia Neisseria gonorrhoeae
[0326] Other targets
[0327] Other DNA can also be targeted. As another example, male genes used to determine the sex of a fetus of a pregnant woman / animal, as well as male genes used to determine the sex of plants and seeds, can also be targeted. Examples of additional targets of interest can include the following shown in Table 6e.
[0328] Table 6e
[0329]
[0330]
[0331] Other miscellaneous targets of interest that provide sources of DNA targets are shown in Table 6f.
[0332] Table 6f
[0333] Sex determination targets SRY gene in mammals and non-mammals Other miscellaneous targets of interest hHPRT1 (hypoxanthine phosphoribosyltransferase 1) 16S E. coli
[0334] A list of non-limiting exemplary target sequences is provided in Table 6g.
[0335] Table 6g
[0336]
[0337]
[0338] b. Sample
[0339] The term“sample” is used herein to refer to any sample comprising DNA (e.g., to determine whether target DNA is present in a population of DNA). As described above, the DNA can be single-stranded DNA, double-stranded DNA, complementary DNA, etc.
[0340] The sample intended to be tested comprises a plurality of nucleic acids. Thus, in some embodiments, the sample comprises two or more (e.g., 3 or more, 5 or more, 10 or more, 20 or more, 50 or more, 100 or more, 500 or more, 1,000 or more, or 5,000 or more) nucleic acids (e.g., DNA). The detection method can be used as a way to very sensitively detect target DNA present in a sample (e.g., in a complex mixture of nucleic acids, such as DNA).
[0341] In some embodiments, the sample comprises 5 or more DNA that differ from each other in sequence (e.g., 10 or more, 20 or more, 50 or more, 100 or more, 500 or more, 1,000 or more, or 5,000 or more DNA). In some embodiments, the sample comprises 10 or more, 20 or more, 50 or more, 100 or more, 500 or more, 10^3 or more, 5x10^3 or more, 10^4 or more, 5x10^4 or more, 10^5 or more, 5x10^5 or more, 10^6 or more, 5x10^6 or more, or 10^7 or more DNA. In some embodiments, the sample comprises 10 to 20, 20 to 50, 50 to 100, 100 to 500, 500 to 10^3, 10^3 to 5x10^3, 5x10^3 to 10^4, 10^4 to 5x10^4, 5x10^4 to 10^5, 10^5 to 5x10^5, 5x10^5 to 10^6, 10^6 to 5x10^6, or 5x10^6 to 10^7, or more than 10^7 DNA. In some embodiments, the sample comprises 5 to 10^7 DNA (e.g., that differ in sequence from each other) (e.g., 5 to 10^6, 5 to 10^5, 5 to 50,000, 5 to 30,000, 10 to 10^6, 10 to 10^5, 10 to 50,000, 10 to 30,000, 20 to 10^6, 20 to 10^5, 20 to 50,000, or 20 to 30,000 DNA).
[0342] In some embodiments, the sample comprises 20 or more DNA that differ from each other in sequence. In some embodiments, the sample comprises DNA from a cell lysate (e.g., a eukaryotic cell lysate, a mammalian cell lysate, a human cell lysate, a prokaryotic cell lysate, a plant cell lysate, etc.). For example, in some embodiments, the sample comprises DNA from a cell (e.g., a eukaryotic cell, e.g., a mammalian cell, e.g., a human cell).
[0343] The sample can be derived from any source, e.g., the sample can be a synthetic combination of purified DNA; the sample can be a cell lysate, a DNA-rich cell lysate, or DNA isolated and / or purified from a cell lysate. The sample can be from a patient (e.g., for diagnostic purposes). The sample can be from a permeabilized cell. The sample can be from a cross-linked cell. The sample can be in a tissue section.
[0344] A sample can comprise target DNA and a plurality of non-target DNA. In some embodiments, there is one or more copies of target DNA per 5 to 10^9 copies of non-target DNA in the sample.
[0345] Suitable samples include, but are not limited to, a urine sample, a blood sample, a serum sample, a plasma sample, a lymphatic fluid sample, a cerebrospinal fluid sample, a saliva sample, a nasopharyngeal sample, an oropharyngeal sample, a nasopharyngeal / oropharyngeal sample, an aspirate sample, or a biopsy sample. Thus, the term“sample” with respect to a patient encompasses a blood and other liquid samples of biological origin, a solid tissue sample such as a biopsy specimen, or a culture of cells or descendants of such cells derived from a tissue culture. The sample can also be one that has been manipulated in any way after their procurement, such as by treatment with a reagent; washing; or enrichment for certain cell populations, such as cancer cells. The sample can be obtained using a swab, for example, a nasopharyngeal swab, an oropharyngeal swab, or a nasopharyngeal / oropharyngeal swab. The sample can also be one that has been enriched for a particular type of molecule (e.g., DNA). Samples encompass biological samples, such as clinical samples, such as blood, plasma, serum, aspirates, cerebrospinal fluid (CSF), and also include tissue obtained by surgical resection, tissue obtained by biopsy, cultured cells, cell supernatants, cell lysates, tissue samples, organs, bone marrow, and the like. A“biological sample” includes biological fluids derived therefrom (e.g., cancer cells, infected cells, etc.), for example, a sample comprising DNA obtained from such cells (e.g., a cell lysate or other cell extract comprising DNA).
[0346] A sample can comprise or can be obtained from any of a plurality of cells, tissues, organs, or acellular fluids. Suitable sample sources include eukaryotic cells, bacterial cells, and archaeal cells. Suitable sample sources include unicellular organisms and multicellular organisms. Suitable sample sources include unicellular eukaryotes; plant or plant cells; algal cells; fungal cells; animal cells, tissues, or organs; cells, tissues, or organs of invertebrates; cells, tissues, fluids, or organs of vertebrates; cells, tissues, fluids, or organs of mammals (e.g., humans; non-human primates; ungulates; felines; bovines; ovines; caprines; etc.). Suitable sample sources include nematodes, protozoa, and the like. Suitable sample sources include parasites, such as helminths, malarial parasites, and the like.
[0347] Suitable sample sources include cells, tissues, or organisms of any of the six kingdoms.
[0348] Suitable sample sources include cells, fluids, tissues, or organs collected from an organism; a particular cell or set of cells isolated from an organism; and the like. For example, where the organism is a plant, suitable sources include xylem, phloem, cambium, leaves, roots, and the like. Where the organism is an animal, suitable sources include particular tissues (e.g., lung, liver, heart, kidney, brain, spleen, skin, fetal tissue, and the like), or particular cell types (e.g., neuronal cells, epithelial cells, endothelial cells, astrocytes, macrophages, glial cells, islet cells, T lymphocytes, B lymphocytes, and the like).
[0349] In some embodiments, the source of the sample is (or is suspected of being) a diseased cell, fluid, tissue, or organ.
[0350] In some embodiments, the source of the sample is a normal (non-diseased) cell, fluid, tissue, or organ.
[0351] In some embodiments, the source of the sample is (or is suspected of being) a cell, tissue, or organ infected by a pathogen. For example, the source of the sample can be an individual who can or can not be infected, and the sample can be any biological sample (e.g., blood, saliva, biopsy sample, plasma, serum, bronchoalveolar lavage sample, sputum, fecal sample, cerebrospinal fluid, fine needle aspirate, swab sample (e.g., buccal swab, cervical swab, nasal swab), interstitial fluid, synovial fluid, nasal discharge, tear fluid, buffy coat, mucosal sample, epithelial cell sample (e.g., epithelial cell scrape), and the like) collected from the individual. In some embodiments, the sample is a cell-free fluid sample.
[0352] In some embodiments, the sample is a liquid sample that can include cells (urine, blood, serum, plasma, lymph, cerebrospinal fluid, saliva, nasopharyngeal samples, oropharyngeal samples, nasopharyngeal / oropharyngeal samples, aspirates, and biopsy samples). Pathogens include viruses, fungi, helminths, protozoa, Plasmodium parasites, Toxoplasma parasites, Schistosoma parasites, and the like. "Helminths" include roundworms, heartworms, and plant feeding nematodes (Nematoda), flukes (Tematoda), Acanthocephala, and tapeworms (Cestoda). Protozoal infections include infections from Giardia spp., Trichomonas spp., African trypanosomiasis, amoebic dysentery, babesiosis, balantidial dysentery, Chagas disease, coccidiosis, malaria, and toxoplasmosis. Examples of pathogens such as parasitic / protozoal pathogens include, but are not limited to: Plasmodium falciparum, Plasmodium vivax, Trypanosoma cruzi, and Toxoplasma gondii. Fungal pathogens include, but are not limited to: Cryptococcus neoformans, Histoplasma capsulatum, Coccidioides immitis, Blastomyces dermatitidis, Chlamydia trachomatis, and Candida albicans.pathogenic viruses include RNA or DNA viruses, such as coronaviruses (e.g., SARS-CoV, SARS-CoV-2, MERS-CoV); immunodeficiency viruses (e.g., HIV); influenza viruses; dengue; West Nile virus; herpes viruses; yellow fever virus; hepatitis C virus; hepatitis A virus; hepatitis B virus; papillomavirus; and the like. Pathogenic viruses can include DNA viruses such as: papovaviruses (e.g., human papillomavirus (HPV), polyomavirus); hepadnaviruses (e.g., hepatitis B virus (HBV)); herpesviruses (e.g., herpes simplex virus (HSV), varicella-zoster virus (VZV), Epstein-Barr virus (EBV), cytomegalovirus (CMV), herpesvirus lymphotrophic, roseola, Kaposi's sarcoma-associated herpesvirus); adenoviruses (e.g., atadenovirus, aviadenovirus, ichtadenovirus, mastadenovirus, siaenovirus); poxviruses (e.g., smallpox, vaccinia virus, cowpox virus, monkeypox virus, goatpox virus, pseudocowpox virus, bovine papular dermatitis virus; tanapoxvirus, yaba monkey tumor virus; molluscum contagiosum virus (MCV)); parvoviruses (e.g., adeno-associated virus (AAV), parvovirus B19, human bocavirus, bufavirus, human parv4G1); gyroviridae; asfarviridae; algal deoxyriboviridae; and the like. Pathogens can include, for example, DNA viruses [e.g.: papovaviruses (e.g., human papillomavirus (HPV), polyomavirus); hepadnaviruses (e.g., hepatitis B virus (HBV)); herpesviruses (e.g., herpes simplex virus (HSV), varicella-zoster virus (VZV), Epstein-Barr virus (EBV), cytomegalovirus (CMV), herpesvirus lymphotrophic, roseola, Kaposi's sarcoma-associated herpesvirus); adenoviruses (e.g., atadenovirus, aviadenovirus, ichtadenovirus, mastadenovirus, siaenovirus); poxviruses (e.g., smallpox, vaccinia virus, cowpox virus, monkeypox virus, goatpox virus, pseudocowpox virus, bovine papular dermatitis virus; tanapoxvirus, yaba monkey tumor virus; molluscum contagiosum virus (MCV)); parvoviruses (e.g., adeno-associated virus (AAV), parvovirus B19, human bocavirus, bufavirus, human parv4G1); gyroviridae; asfarviridae; algal deoxyriboviridae; and the like], Mycobacterium tuberculosis, Streptococcus agalactiae,agalactiae, methicillin-resistant Staphylococcus aureus, Legionella pneumophila, Streptococcus pyogenes, Escherichia coli, Neisseria gonorrhoeae, Neisseria meningitidis, Pneumococcus, Cryptococcus neoformans, Histoplasma capsulatum, Haemophilus influenzae type b, Treponema pallidum, Lyme disease spirochetes, Pseudomonas aeruginosa, Mycobacterium leprae, Brucella abortus Herpes simplex virus (HSV), rabies virus, influenza virus, cytomegalovirus, herpes simplex virus I, herpes simplex virus II, human serum parvovirus, respiratory syncytial virus, varicella-zoster virus, hepatitis B virus, hepatitis C virus, measles virus, adenovirus, human T-cell leukemia virus, Epstein-Barr virus, murine leukemia virus, mumps virus, vesicular stomatitis virus, Sindbis virus, lymphocytic choriomeningitis virus, verrucavirus, bluetongue virus, Sendai virus, feline leukemia virus, reovirus, poliovirus, simian virus 40, mouse mammary tumor virus, dengue virus, rubella virus, West Nile virus, Plasmodium falciparum, Plasmodium vivax, Toxoplasma gondii, Trypanosoma rangeli, Trypanosoma krusei, Trypanosoma rhodesiense, Trypanosoma brucei brucei), Schistosoma mansoni, Schistosoma japonicum, Babesia bovis, Eimeria tenella, Onchocerca volvulus, Leishmania tropica, Mycobacterium tuberculosis, Trichinella spiralis, Theileria parva, Taenia hydatigena, Taenia sheep.Mycoplasma ovis, Taenia saginata, Echinococcus granulosus, Mesocestoides corti, Mycoplasma arthritidis, M. hyorhinis, M. orale, M. arginini, Acholeplasma laidlawii, M. salivarium, and M. pneumoniae.
[0353] c. Measuring a detectable signal
[0354] The detection method generally includes a step of measuring a detectable signal produced by the Cas12 of the present disclosure. The detectable signal can be any signal produced when the ss oligonucleotide is cleaved. The detection step can involve fluorescence-based detection. The readout of such a detection method can be any convenient readout. Examples of possible readouts include, but are not limited to: the amount of detectable fluorescent signal measured; visual analysis of bands on a gel (e.g., bands representing cleavage products versus uncleaved substrate), visual or sensor-based detection of the presence or absence of a color (i.e., colorimetric detection methods), the presence or absence (or a particular amount) of a magnetic signal, and the presence or absence (or a particular amount) of an electrical signal.
[0355] In some embodiments, the measurement can be quantitative, e.g., in the sense that the amount of signal detected can be used to determine the amount of target DNA present in the sample. In some embodiments, the measurement can be qualitative, e.g., in the sense that the presence or absence of a detectable signal can be indicative of the presence or absence of a targeted DNA (e.g., a virus, a SNP, etc.). In some embodiments, there will be no detectable signal (e.g., above a given threshold level) unless the one or more targeted DNA (e.g., a virus, a SNP, etc.) is present at a concentration that exceeds a particular threshold. In some embodiments, the detection threshold can be titrated by modifying the amount of Cas12 protein provided.
[0356] The compositions and methods of the present disclosure can be used to detect any DNA target.
[0357] In some embodiments, the detection methods of the present disclosure can be used to determine the amount of target DNA in a sample (e.g., a sample comprising target DNA and a plurality of non-target DNA). Determining the amount of target DNA in a sample can comprise comparing the amount of detectable signal generated from the test sample to the amount of detectable signal generated from a reference sample. Determining the amount of target DNA in a sample can comprise: measuring the detectable signal to generate a test measurement; measuring the detectable signal generated by the reference sample to generate a reference measurement; and comparing the test measurement to the reference measurement to determine the amount of target DNA present in the sample.
[0358] In some embodiments, the detectable signal is detectable within less than 1, 2, 3, 4, 5, 10, 15, 20, 30, 60, 90, 120, 150, 180, 210, or 240 minutes.
[0359] In some embodiments, the sensitivity of the subject compositions and / or methods can be increased by coupling detection with nucleic acid amplification (e.g., for detecting the presence of target DNA, such as viral DNA or SNPs in cellular genomic DNA).
[0360] In some embodiments, the nucleic acid in the sample is amplified prior to contact with the Cas12; in particular embodiments, the Cas12 remains in an inactive state until amplification has ended. In some embodiments, the nucleic acid in the sample is amplified at the same time as contact with the Cas12. Amplification can occur for 5 seconds or more, up to 240 minutes or more, because of the entire processing time involved in the detection method.
[0361] Various amplification methods and components will be known to one of ordinary skill in the art, and any convenient method can be used.
[0362] Nucleic acid amplification can include polymerase chain reaction (PCR), reverse transcription PCR (RT-PCR), quantitative PCR (qPCR), reverse transcription qPCR (RT-qPCR), isothermal PCR, nested PCR, multiplex PCR, asymmetric PCR, touchdown PCR, random primer PCR, heminested PCR, polymerase cycle assembly (PCA), colony PCR, ligase chain reaction (LCR), digital PCR, methylation-specific PCR (MSP), coamplification at lower denaturation temperature-PCR (COLD-PCR), allele-specific PCR, intersequence-specific PCR (ISS-PCR), whole genome amplification (WGA), inverse PCR, and thermal asymmetric interlaced PCR (TAIL-PCR).
[0363] In some embodiments, amplification is isothermal amplification. Thus, isothermal nucleic acid amplification methods can be performed inside or outside of a laboratory setting. Examples of isothermal amplification methods include, but are not limited to: loop-mediated isothermal amplification (LAMP), helicase-dependent amplification (HDA), recombinase polymerase amplification (RPA), strand displacement amplification (SDA), nucleic acid sequence-based amplification (NASBA), transcription-mediated amplification (TMA), nicking enzyme amplification reaction (NEAR), rolling circle amplification (RCA), multiple displacement amplification (MDA), ramification (RAM), circular helicase-dependent amplification (cHDA), single-primer isothermal amplification (SPIA), signal-mediated amplification of RNA technology (SMART), self-sustained sequence replication (3SR), genome exponential amplification reaction (GEAR), and isothermal multiple displacement amplification (IMDA).
[0364] d. detection oligonucleotide
[0365] The novel Cas12 proteins of the present disclosure have collateral (trans-cleavage) activity. As in the case of Cas12a.1, upon binding to DNA targeted by a guide, the proteins have the ability to collateral cleave ssDNA. In the case of Cas12p, the proteins have the dual ability to collateral cleave all types of oligonucleotides including ssDNA, ssRNA, chimeric ssDNA / RNA, and other RNA-containing oligonucleotides. These features are taken into account when designing detection oligonucleotides for use in the assay.
[0366] In some embodiments, the detection method comprises contacting a sample (e.g., a sample comprising target DNA and a plurality of non-target ssDNA) with: i) a Cas12 protein of the present disclosure; ii) a gRNA (or an array of precursor gRNAs); and iii) a detection oligonucleotide that is not hybridized to the guide sequence of the gRNA. For example, in some embodiments, the detection method comprises contacting a sample with a labeled detection oligonucleotide (a detection ssDNA in the case of Cas12a.1, or a detection oligonucleotide comprising RNA, DNA, and combinations thereof in the case of Cas12p) comprising a fluorescently emitting dye pair; the Cas12 protein of the present disclosure has the ability to cleave the labeled detection oligonucleotide upon activation (by a gRNA hybridized to the target DNA); and the detectable signal measured is produced by the fluorescently emitting dye pair. For example, in some embodiments, the detection method comprises contacting a sample with a labeled detection oligonucleotide comprising a fluorescent resonance energy transfer (FRET) pair or a quencher / fluor pair or both. In some embodiments, the detection method comprises contacting a sample with a labeled detection oligonucleotide comprising a FRET pair. In some embodiments, the detection method comprises contacting a sample with a labeled detection oligonucleotide comprising a fluor / quencher pair.
[0367] Fluorescent emission dyes pairs include FRET pairs or quencher / fluorophore pairs. In both embodiments of FRET pairs and quencher / fluorophore pairs, the emission spectrum of one dye overlaps the region of the absorption spectrum of the other dye in the pair. As used herein, the term “fluorescent emission dye pair” is a generic term used to encompass both “fluorescent resonance energy transfer (FRET) pairs” and “quencher / fluorophore pairs.” The term “fluorescent emission dye pair” is used interchangeably with the phrase “FRET pair and / or quencher / fluorophore pair.”
[0368] In some embodiments (e.g., when the detection subunit comprises a FRET pair), the labeled detection subunit produces an amount of detectable signal prior to being cleaved, and the amount of detectable signal measured decreases when the labeled detection subunit is cleaved. In some embodiments, the labeled detection subunit produces a first detectable signal prior to being cleaved (e.g., by a FRET pair), and a second detectable signal when the labeled detection subunit is cleaved (e.g., by a quencher / fluorophore pair). Thus, in some embodiments, the labeled detection subunit comprises both a FRET pair and a quencher / fluorophore pair.
[0369] In some embodiments, the labeled detection subunit comprises a FRET pair.
[0370] FRET donor and acceptor moieties (FRET pairs) will be known to those of ordinary skill in the art, and any convenient FRET pair (e.g., any convenient pair of donor and acceptor moieties) can be used. Examples of suitable FRET pairs include, but are not limited to, those presented in Table 7. The FRET pairs provided in US 10,253,365 are incorporated by reference herein in their entirety. In some embodiments, the FRET pair is 5' 6-FAM and 3IABkFQ (Iowa Black (registered trademark)-FQ).
[0371] Table 7
[0372] Examples of FRET pairs (donor and acceptor pairs)
[0373]
[0374]
[0375] In some embodiments, the labeled detection subunit produces a detectable signal when it is cleaved (e.g., in some embodiments, the labeled detection subunit comprises a quencher / fluorophore pair).
[0376] Any fluorescent label can be used. Examples of fluorescent labels include, but are not limited to: Alexa dyes, ATTO dyes (e.g., ATTO 390, ATTO 425, ATTO 465, ATTO 488, ATTO 495, ATTO 514, ATTO 520, ATTO 532, ATTO Rho6G, ATTO 542, ATTO 550, ATTO 565, ATTO Rho3B, ATTO Rho11, ATTO Rho12, ATTO Thio12, ATTO Rho101, ATTO 590, ATTO 594, ATTO Rho13, ATTO 610, ATTO 620, ATTO Rho14, ATTO 633, ATTO 647, ATTO 647N, ATTO 655, ATTO Oxa12, ATTO 665, ATTO 680, ATTO 700, ATTO 725, ATTO 740), DyLight dyes, Cyanine dyes (e.g., Cy2, Cy3, Cy3.5, Cy3b, Cy5, Cy5.5, Cy7, Cy7.5), FluoProbes dyes, Sulfo Cy dyes, Seta dyes, IRIS dyes, SeTau dyes, SRfluor dyes, Square dyes, fluorescein isothiocyanate (FITC), fluorescein amidite (FAM), tetramethylrhodamine (TRITC), Texas Red, Oregon Green, Pacific Blue, Pacific Green, Pacific Orange, quantum dots, and tethered fluorescent proteins.
[0377] Examples of quencher moieties include, but are not limited to: dark quenchers, Black Hole (e.g., BHQ-0, BHQ-1, BHQ-2, BHQ-3), Qxl quenchers, ATTO quenchers (e.g., ATTO 540Q, ATTO 580Q, and ATTO 612Q), dimethylaminoazobenzenesulfonic acid (Dabsyl), Iowa Black RQ, Iowa Black FQ, IRDye QC-1, QSY dyes (e.g., QSY 7, QSY 9, QSY 21), Absolute Quencher, Eclipse, and metal clusters (such as gold nanoparticles), and the like.
[0378] In some embodiments, the quencher moiety is selected from: dark quenchers, Black Hole (e.g., BHQ-0, BHQ-1, BHQ-2, BHQ-3), Qxl quenchers, ATTO quenchers (e.g., ATTO 540Q, ATTO 580Q, and ATTO 612Q), dimethylaminoazobenzene sulfonic acid (Dabsyl), Iowa Black RQ, Iowa Black FQ, IRDye QC-1, QSY dyes (e.g., QSY 7, QSY 9, QSY 21), Absolute Quencher, Eclipse, and metal clusters.
[0379] In some embodiments, cleavage of a labeled detector can be detected by a colorimetric readout. For example, release of a fluorophore (e.g., from a FRET pair, from a quencher / fluorophore pair) can result in a shift in the wavelength of the detectable signal (and thus a color shift). Thus, in some embodiments, cleavage of a subject labeled detector can be detected by a color shift. Such a shift can be represented as a loss in the amount of signal of one color (wavelength), an increase in the amount of another color, a change in the ratio of one color to another, etc.
[0380] As provided herein, a labeled detector can be a nucleic acid mimic. Polynucleotide mimics include PNAs, LNAs, CeNAs, and morpholino nucleic acids.
[0381] A labeled detector can also comprise one or more substituted sugar moieties.
[0382] A labeled detector can also comprise modified nucleotides.
[0383] e. Positive control
[0384] The detection methods provided herein can also include a positive control target DNA. In some embodiments, the methods include using a positive control gRNA comprising a nucleotide sequence that is hybridized to a control target DNA. In some embodiments, a positive control target DNA is provided in various amounts. In some embodiments, a positive control target DNA is provided in various known concentrations, along with a control non-target DNA.
[0385] f. gRNA array
[0386] In some embodiments, the methods include contacting a sample with a precursor gRNA array, wherein a new Cas12 protein of the disclosure cleaves the precursor gRNA array to produce the gRNAs.
[0387] In some embodiments, such an array of gRNAs comprises 2 or more gRNAs (e.g., 3 or more, 4 or more, 5 or more, 6 or more, or 7 or more gRNAs). The gRNAs of a given array can target different target sites of the same target DNA (i.e., can comprise guide sequences hybridized to different target sites of the same target DNA) (e.g., which can increase detection sensitivity) and / or can target different target DNAs (e.g., single nucleotide polymorphisms (SNPs), different strains of a particular virus, etc.), and this can be used, for example, to detect multiple strains of a virus. In some embodiments, each gRNA of a precursor array of gRNAs has a different guide sequence.
[0388] In some embodiments, a precursor array of gRNAs comprises two or more gRNAs that target different target sites within the same target DNA. For example, in some embodiments, such a scenario can increase the sensitivity of detection (when either of them is hybridized to the target DNA) by activating a Cas9 or Cas12 protein of the disclosure. Thus, in some embodiments, a subject composition (e.g., kit) or method comprises two or more gRNAs (either in the context of a precursor array of gRNAs, or not in the context of a precursor array of gRNAs, e.g., the gRNAs can be mature gRNAs).
[0389] In some embodiments, a precursor array of gRNAs comprises two or more gRNAs that target different target DNAs. For example, such a scenario can result in a positive signal when any one of a family of potential target DNAs is present. Such an array can be used to target a family of transcripts, e.g., based on a variation such as a single nucleotide polymorphism (SNP) (e.g., for diagnostic purposes). This can also be useful for detecting whether any of a number of different strains of a virus is present. This can also be useful for detecting whether any of a number of different species, strains, isolates, or variants of a bacterium or virus is present. Thus, in some embodiments, a subject composition (e.g., kit) or method comprises two or more gRNAs (either in the context of a precursor array of gRNAs, or not in the context of a precursor array of gRNAs, e.g., the gRNAs can be mature gRNAs).
[0390] V. Compositions of Matter
[0391] Provided herein are compositions and pharmaceutical compositions comprising a Cas9 protein and / or Cas9 gRNA of the disclosure, which can optionally comprise a pharmaceutically acceptable carrier and / or a protein stabilizing buffer and / or a nucleic acid stabilizing buffer. In some embodiments, the Cas9 protein and / or Cas9 gRNA is provided in lyophilized form.
[0392] Provided herein are compositions and pharmaceutical compositions comprising a Cas12 protein and / or a Cas12 gRNA of the disclosure, which can optionally comprise a pharmaceutically acceptable carrier and / or a protein stabilizing buffer and / or a nucleic acid stabilizing buffer. In some embodiments, the Cas12 protein and / or Cas12 gRNA is provided in lyophilized form.
[0393] Provided herein are compositions comprising a gRNA and / or a gRNA array of the disclosure (compatible for use with a Cas9 protein of the disclosure and / or a Cas12 protein of the disclosure) and optionally a protein stabilizing buffer.
[0394] Provided herein are proteins comprising an amino acid sequence that is 70-99.5% homologous to SEQ ID NO: 1, 2, 3, 4, 222, 5, 10, 11, or 12. Provided herein are compositions comprising these proteins and optionally a pharmaceutically acceptable carrier. Provided herein are these proteins and optionally a protein stabilizing buffer.
[0395] Provided herein are DNA polynucleotides encoding a sequence that encodes any of the Cas9 or Cas12 proteins of the disclosure. Also provided are recombinant expression vectors comprising such DNA polynucleotides. In some embodiments, the nucleotide sequence encoding the Cas9 or Cas12 of the disclosure is operably linked to a promoter. In some embodiments, the nucleic acid encoding the Cas9 or Cas12 further comprises a nuclear localization signal (NLS), which can be used for expression in eukaryotic systems.
[0396] Provided herein are DNA polynucleotides or RNAs comprising a sequence that encodes any of the gRNAs of the disclosure. Also provided are recombinant expression vectors comprising such DNA polynucleotides. In some embodiments, the nucleotide sequence encoding the gRNA of the disclosure is operably linked to a promoter.
[0397] Also provided herein are host cells comprising any of the recombinant vectors provided herein.
[0398] VI. Kits
[0399] Provided herein are kits comprising one or more components of the Cas9 and Cas12 engineering systems described herein, which can be used for a variety of applications, including but not limited to therapeutic and diagnostic applications.
[0400] In some embodiments, provided herein is a kit comprising: (a) a Cas9.1, Cas9.2, Cas9.3, or Cas9.4 protein or a nucleic acid encoding a Cas9.1, Cas9.2, Cas9.3, or Cas9.4 protein; (b) a Cas9.1, Cas9.2, Cas9.3, or Cas9.4 gRNA or a nucleic acid encoding a Cas9.1, Cas9.2, Cas9.3, or Cas9.4 gRNA, wherein the gRNA and the Cas9.1, Cas9.2, Cas9.3, or Cas9.4 protein do not naturally occur together, wherein the gRNA is capable of hybridizing to a target sequence in a target DNA, and the gRNA is capable of forming a complex with the Cas9.1, Cas9.2, Cas9.3, or Cas9.4 protein.
[0401] In some embodiments, provided herein is a kit comprising: (a) a Cas12a.1, Cas12p, or Cas12q protein or a nucleic acid encoding the Cas12a.1, Cas12p, or Cas12q protein; and (b) a Cas12a.1, Cas12p, or Cas12q gRNA or a nucleic acid encoding a Cas12a.1, Cas12p, or Cas12q gRNA, wherein the gRNA and the Cas12a.1, Cas12p, or Cas12q protein do not naturally occur together, wherein the gRNA is capable of hybridizing to a target sequence in a target DNA, and the gRNA is capable of forming a complex with the Cas12a.1, Cas12p, or Cas12q protein.
[0402] In exemplary embodiments, diagnostic kits are provided herein. In exemplary embodiments, reagent components are provided in lyophilized form. In some embodiments, reagent components are provided separately (lyophilized or non-lyophilized), in other embodiments, reagent components are provided in pre-mixed form (lyophilized or non-lyophilized).
[0403] The following are exemplary kit reagent components for detecting SARS-CoV-2 (an RNA virus) using one of the new Cas12 proteins of the disclosure (Cas12a.1, Cas12p, and Cas12q) as exemplified in Example 10.
[0404] (1) Reagents containing lyophilized reaction mix, SARS-CoV-2 primer set, and enzymes for reverse transcription and loop-mediated isothermal amplification (RT-LAMP) of the SARS-CoV-2 genome gene for the disease.
[0405] (2) reagents containing lyophilized reaction mix, control RNAse P primer set, and enzymes for reverse transcription and RT-LAMP amplification of human housekeeping gene RNAse P.
[0406] (3) reagents containing lyophilized reaction mix and Cas12p-gRNA RNP complex for detection of SARS-CoV-2 amplification products. This mix can also contain a labeled reporter, such as a 5’FAM-3’Quencher ssRNA based oligonucleotide reporter or a 5’FAM-3’Quencher single stranded DNA / RNA chimera based oligonucleotide reporter.
[0407] (4) reagents containing lyophilized reaction mix and Cas12p-gRNA RNP complex for detection of RNAse P amplification products. This mix can also contain a labeled reporter, such as a 5’FAM-3’Quencher RNA based oligonucleotide reporter.
[0408] FIG. 23 Exemplary strips showing lyophilized beads of the disclosure contained in an exemplary kit are shown. Each bead can be resuspended with water and used in a detection assay. Exemplary beads each contain a CRISPR protein (e.g., Cas12p), a gRNA for a desired target (e.g., a gRNA for SARS-CoV-2), a labeled reporter, a buffer, and nuclease-free water.
[0409] VII. Enumerated Embodiments
[0410] Provided herein are illustrative, non-limiting, enumerated embodiments of the disclosure.
[0411] Embodiment 1. An engineered system comprising:
[0412] a. a Cas9.1, Cas9.2, Cas9.3, or Cas9.4 protein or a nucleic acid encoding the Cas9.1, Cas9.2, Cas9.3, or Cas9.4 protein; and
[0413] b. a Cas9.1, Cas9.2, Cas9.3, or Cas9.4 guide RNA (gRNA) or a nucleic acid encoding a Cas9.1, Cas9.2, Cas9.3, or Cas9.4 gRNA, wherein the gRNA and the Cas9.1, Cas9.2, Cas9.3, or Cas9.4 protein do not naturally occur together, wherein the gRNA is capable of hybridizing to a target sequence in a target DNA, and the gRNA is capable of forming a complex with the Cas9.1, Cas9.2, Cas9.3, or Cas9.45 protein.
[0414] Embodiment 2. The system of embodiment 1, comprising:
[0415] a. a Cas9.1, Cas9.2, Cas9.3, Cas9.4 protein; and
[0416] b. a Cas9.1, Cas9.2, Cas9.3, or Cas9.4 gRNA.
[0417] Embodiment 3. The system of embodiment 1, comprising:
[0418] a. a nucleic acid encoding the Cas9.1, Cas9.2, Cas9.3, or Cas9.4 protein; and
[0419] b. a nucleic acid encoding the Cas9.1, Cas9.2, Cas9.3, or Cas9.4 gRNA.
[0420] Embodiment 4. The system of any one of embodiments 1-3, wherein the gRNA is a single molecule gRNA.
[0421] Embodiment 5. The system of any one of embodiments 1-3, wherein the gRNA is a dual molecule gRNA.
[0422] Embodiment 6. The system of any one of embodiments 1-5, wherein the Cas9.1 protein comprises the amino acid sequence of SEQ ID NO: 1 or an amino acid sequence having at least 70% sequence identity to SEQ ID NO: 1.
[0423] Embodiment 7. The system of any one of embodiments 1-5, wherein the Cas9.2 protein comprises the amino acid sequence of SEQ ID NO: 2 or an amino acid sequence having at least 70% sequence identity to SEQ ID NO: 2.
[0424] Embodiment 8. The system of any one of embodiments 1-5, wherein the Cas9.3 protein comprises the amino acid sequence of SEQ ID NO: 10 or an amino acid sequence having at least 70% sequence identity to SEQ ID NO: 10.
[0425] Embodiment 9. The system of any one of embodiments 1-5, wherein the Cas9.4 protein comprises the amino acid sequence of SEQ ID NO: 11 or an amino acid sequence having at least 70% sequence identity to SEQ ID NO: 11.
[0426] Embodiment 10. The system of any one of embodiments 1 to 7, wherein the target sequence is the sequence of a target provided in any one of Table 6a to Table 6f.
[0427] Embodiment 11. The system of any one of embodiments 1 to 7, wherein the target sequence is a human sequence.
[0428] Embodiment 12. The system of any one of embodiments 1 to 7, wherein the target sequence is a non-human primate sequence.
[0429] Embodiment 13. The system of any one of embodiments 1 to 12, wherein the Cas9.1, Cas9.2, Cas9.3, or Cas9.4 protein is a catalytically active protein.
[0430] Embodiment 14. The system of embodiment 13, wherein the Cas9.1, Cas9.2, Cas9.3, or Cas9.4 protein cleaves at a site distal to the target sequence.
[0431] Embodiment 15. The system of any one of embodiments 1 to 12, wherein the Cas9.1, Cas9.2, Cas9.3, or Cas9.4 protein is a catalytically inactive protein.
[0432] Embodiment 16. The system of any one of embodiments 1 to 12, wherein the Cas9.1, Cas9.2, Cas9.3, or Cas9.4 protein comprises a nickase activity.
[0433] Embodiment 17. An engineered system, comprising:
[0434] a. a Class 2 Type V CRISPR-Cas RNA-guided endonuclease protein; and
[0435] b. a single guide RNA (gRNA),
[0436] wherein the gRNA and the Class 2 Type V CRISPR-Cas RNA-guided endonuclease protein do not naturally occur together, wherein the gRNA is capable of hybridizing to a target sequence in a target DNA, wherein the gRNA is capable of forming a complex with the Class 2 Type V CRISPR-Cas RNA-guided endonuclease protein, and wherein the Class 2 Type V CRISPR-Cas RNA-guided endonuclease protein has collateral cleavage activity and is capable of collateral cleaving a single-stranded polynucleotide comprising RNA in the absence of a tracrRNA.
[0437] Embodiment 18. The system of embodiment 17, wherein the Class 2, Type V CRISPR-Cas RNA-guided endonuclease protein comprises the amino acid sequence of SEQ ID NO: 4 or an amino acid sequence having at least 70% sequence identity to SEQ ID NO: 4.
[0438] Embodiment 19. The system of any one of embodiments 17-18, wherein the target sequence is the sequence of a target provided in any one of Tables 6a-6f.
[0439] Embodiment 20. The system of any one of embodiments 17-18, wherein the target sequence is a human sequence.
[0440] Embodiment 21. The system of any one of embodiments 17-18, wherein the target sequence is a non-human primate sequence.
[0441] Embodiment 22. The system of any one of embodiments 17-18, wherein the target sequence is a bacterial or viral sequence.
[0442] Embodiment 23. The system of any one of embodiments 17-22, wherein the Class 2, Type V CRISPR-Cas RNA-guided endonuclease protein is capable of flanking cutting single stranded RNA.
[0443] Embodiment 24. The system of any one of embodiments 17-22, wherein the Class 2, Type V CRISPR-Cas RNA-guided endonuclease protein is capable of flanking cutting single stranded DNA / RNA hybrids.
[0444] Embodiment 25. An engineered system, comprising:
[0445] a. a Cas12a.1, Cas12p, or Cas12q protein or a nucleic acid encoding the Cas12a.1, Cas12p, or Cas12q protein; and
[0446] b. a Cas12a.1, Cas12p, or Cas12q gRNA or a nucleic acid encoding a Cas12a.1, Cas12p, or Cas12q gRNA,
[0447] wherein the gRNA and the Cas12a.1, Cas12p, or Cas12q protein do not naturally occur together, wherein the gRNA is capable of hybridizing to a target sequence in a target DNA, and the gRNA is capable of forming a complex with the Cas12a.1, Cas12p, or Cas12q protein.
[0448] Embodiment 26. The system of embodiment 25, comprising:
[0449] a. a Cas12a.1, Cas12p, or Cas12q protein; and
[0450] b. a Cas12a.1, Cas12p, or Cas12q gRNA.
[0451] Embodiment 27. The system of embodiment 25, comprising:
[0452] a. a nucleic acid encoding the Cas12a.1, Cas12p, or Cas12q protein; and
[0453] b. a nucleic acid encoding a Cas12a.1, Cas12p, or Cas12q gRNA.
[0454] Embodiment 28. The system of any one of embodiments 25-27, wherein the Cas12a.1 protein comprises the amino acid sequence of SEQ ID NO: 3 or an amino acid sequence having at least 70% sequence identity to SEQ ID NO: 3.
[0455] Embodiment 29. The system of any one of embodiments 25-27, wherein the Cas12p protein comprises the amino acid sequence of SEQ ID NO: 4 or an amino acid sequence having at least 70% sequence identity to SEQ ID NO: 4.
[0456] Embodiment 30. The system of any one of embodiments 25-27, wherein the Cas12q protein comprises the amino acid sequence of SEQ ID NO: 222 or an amino acid sequence having at least 70% sequence identity to SEQ ID NO: 222.
[0457] Embodiment 31. The system of any one of embodiments 25-27, wherein the Cas12q protein comprises the amino acid sequence of SEQ ID NO: 5 or an amino acid sequence having at least 70% sequence identity to SEQ ID NO: 5.
[0458] Embodiment 32. The system of any one of embodiments 25-31, wherein the target sequence is the sequence of a target provided in any one of Table 6a-Table 6f.
[0459] Embodiment 33. The system of any one of embodiments 25-31, wherein the target sequence is a sequence of a human.
[0460] Embodiment 34. The system of any one of embodiments 25-31, wherein the target sequence is a sequence of a non-human primate.
[0461] Embodiment 35. The system of any one of embodiments 25-31, wherein the target sequence is a bacterial or viral sequence.
[0462] Embodiment 36. The system of any one of embodiments 25-34, wherein the Casl2a.l, Casl2p, or Casl2q protein is a catalytically active Casl2a.l, Casl2p, or Casl2q protein.
[0463] Embodiment 37. The system of embodiment 36, wherein the Casl2a.l, Casl2p, or Casl2q protein cleaves at a site distal to the target sequence.
[0464] Embodiment 38. The system of any one of embodiments 25-34, wherein the Casl2a.l, Casl2p, or Casl2q protein is a catalytically inactive Casl2a.l, Casl2p, or Casl2q protein.
[0465] Embodiment 39. The system of any one of embodiments 25-34, wherein the Casl2a.l, Casl2p, or Casl2q protein comprises nickase activity.
[0466] Embodiment 40. An engineered single molecule gRNA, comprising:
[0467] a. a target-RNA comprising a spacer sequence capable of hybridizing to a target sequence in a target DNA; and
[0468] b. an activator-RNA capable of hybridizing to the target-RNA to form a double-stranded RNA duplex, the activator-RNA comprising an activator-RNA,
[0469] wherein the target-RNA and the activator-RNA are covalently linked to one another, wherein the single molecule gRNA is capable of forming a complex with a Cas9.1, Cas9.2, Cas9.3, or Cas9.4 protein, and wherein hybridization of the spacer sequence to the target sequence is capable of targeting the Cas9.1, Cas9.2, Cas9.3, or Cas9.4 protein to the target DNA.
[0470] Embodiment 41. The gRNA of embodiment 40, wherein the target-RNA and the activator-RNA are arranged in a 5' to 3' orientation.
[0471] Embodiment 42. The gRNA of embodiment 40, wherein the activator-RNA and the target-RNA are arranged in a 5' to 3' orientation.
[0472] Embodiment 43. The gRNA of any one of embodiments 40 to 42, wherein the target-RNA and the activator-RNA are covalently linked to each other by a linker.
[0473] Embodiment 44. The gRNA of any one of embodiments 40 to 43, wherein the single molecule gRNA comprises one or more sequence modifications compared to the sequence of a corresponding wild type tracrRNA and / or crRNA.
[0474] Embodiment 45. The gRNA of any one of embodiments 40 to 44, wherein the target-RNA comprises a spacer sequence of about 10-50 nucleotides having 100% complementarity to a sequence in the target DNA.
[0475] Embodiment 46. The gRNA of any one of embodiments 40 to 44, wherein the target-RNA comprises a spacer sequence of about 10-50 nucleotides having less than 100% complementarity to a sequence in the target DNA.
[0476] Embodiment 47. The gRNA of any one of embodiments 40 to 46, wherein the target sequence is the sequence of a target provided in any one of Tables 6a to 6f.
[0477] Embodiment 48. The gRNA of any one of embodiments 40 to 47, wherein the Cas9.1 protein comprises the sequence of SEQ ID NO: 1 or a sequence having at least 70% sequence identity to SEQ ID NO: 1.
[0478] Embodiment 49. The gRNA of any one of embodiments 40 to 47, wherein the Cas9.2 protein comprises the sequence of SEQ ID NO: 2 or a sequence having at least 70% sequence identity to SEQ ID NO: 2.
[0479] Embodiment 50. The gRNA of any one of embodiments 40 to 47, wherein the Cas9.3 protein comprises the sequence of SEQ ID NO: 10 or a sequence having at least 70% sequence identity to SEQ ID NO: 10.
[0480] Embodiment 51. The gRNA of any one of embodiments 40 to 47, wherein the Cas9.4 protein comprises the sequence of SEQ ID NO: 11 or a sequence having at least 70% sequence identity to SEQ ID NO: 11.
[0481] Embodiment 52. An engineered single molecule gRNA comprising a scaffold sequence of SEQ ID NO: 116 or SEQ ID NO: 117 and a spacer sequence capable of hybridizing to a target sequence in a target DNA.
[0482] Embodiment 53. The gRNA of embodiment 52, wherein the target DNA comprises viral DNA, plant DNA, fungal DNA, or bacterial DNA.
[0483] Embodiment 54. The gRNA of embodiment 52, wherein the target sequence is the sequence of a target provided in any one of Tables 6a-6f.
[0484] Embodiment 55. The gRNA of embodiment 52, wherein the target is a coronavirus.
[0485] Embodiment 56. The gRNA of embodiment 52, wherein the target is a SARS-CoV-2 virus.
[0486] Embodiment 57. The gRNA of embodiment 52, wherein the target DNA is cDNA and has been obtained by reverse transcription.
[0487] Embodiment 58. A method of modifying a target DNA, the method comprising contacting the target DNA with any one of the systems of embodiments 1-39, wherein the gRNA hybridizes to the target sequence, whereby modification of the target DNA occurs.
[0488] Embodiment 59. The method of embodiment 58, wherein the target DNA is extrachromosomal DNA.
[0489] Embodiment 60. The method of embodiment 58, wherein the target DNA is part of a chromosome.
[0490] Embodiment 61. The method of embodiment 58, wherein the target DNA is part of an in vitro chromosome.
[0491] Embodiment 62. The method of embodiment 58, wherein the target DNA is part of an in vivo chromosome.
[0492] Embodiment 63. The method of embodiment 58, wherein the target DNA is extracellular.
[0493] Embodiment 64. The method of embodiment 58, wherein the target DNA is intracellular.
[0494] Embodiment 65. The method of embodiment 64, wherein the target DNA comprises a gene and / or a regulatory region thereof.
[0495] Embodiment 66. The method of embodiment 64 or 65, wherein the cell is selected from the group consisting of an archaeal cell, a bacterial cell, a eukaryotic cell, a eukaryotic unicellular organism, a somatic cell, a germ cell, a stem cell, a plant cell, an algal cell, an animal cell, an invertebrate animal cell, a vertebrate animal cell, a fish cell, a frog cell, a bird cell, a mammalian cell, a pig cell, a cow cell, a goat cell, a sheep cell, a rodent cell, a rat cell, a mouse cell, a non-human primate cell, and a human cell.
[0496] Embodiment 67. The method of any one of embodiments 58-66, wherein the modification comprises introducing a double-strand break in the target DNA.
[0497] Embodiment 68. The method of any one of embodiments 58-67, wherein the contacting occurs under conditions that allow for non-homologous end joining or homology directed repair.
[0498] Embodiment 69. The method of any one of embodiments 58-67, wherein the target DNA is contacted with a donor polynucleotide, wherein the donor polynucleotide, a portion of the donor polynucleotide, a copy of the donor polynucleotide, or a portion copy of the donor polynucleotide is integrated into the target DNA.
[0499] Embodiment 70. The method of any one of embodiments 58-67, wherein the method does not comprise contacting the cell with a donor polynucleotide, or wherein the target DNA is modified such that nucleotides within the target DNA are deleted.
[0500] Embodiment 71. A method of detecting a target DNA in a sample, the method comprising:
[0501] a. contacting the sample with:
[0502] i. a Casl2a.l, Casl2p, or Casl2q protein;
[0503] ii. a Casl2a.l, Casl2p, or Casl2q gRNA comprising a spacer sequence capable of hybridizing to a target sequence in a target DNA; and
[0504] iii. a labeled detector that does not hybridize to the spacer sequence of the gRNA; and
[0505] b. measuring a detectable signal produced by the Cas12a.1, Cas12p, or Cas12q protein cleaving the labeled detector, thereby detecting the target DNA.
[0506] Embodiment 72. The method of embodiment 71, wherein the labeled detector comprises a labeled single-stranded DNA.
[0507] Embodiment 73. The method of embodiment 71, wherein the labeled detector comprises a labeled RNA.
[0508] Embodiment 74. The method of embodiment 72, wherein the labeled RNA is a single-stranded RNA.
[0509] Embodiment 75. The method of embodiment 71, wherein the labeled detector comprises a labeled single-stranded DNA / RNA chimera.
[0510] Embodiment 76. The method of any one of embodiments 71-75, wherein the labeled detector comprises one or more modified nucleotides.
[0511] Embodiment 77. The method of any one of embodiments 71-76, comprising contacting the sample with an array of precursor gRNAs, wherein the Cas12a.1, Cas12p, or Cas12q protein cleaves the array of precursor gRNAs to produce the gRNAs.
[0512] Embodiment 78. The method of any one of embodiments 71-77, wherein the target DNA is single-stranded.
[0513] Embodiment 79. The method of any one of embodiments 71-78, wherein the target DNA is double-stranded.
[0514] Embodiment 80. The method of any one of embodiments 71-79, wherein the target DNA is viral DNA, plant DNA, fungal DNA, or bacterial DNA.
[0515] Embodiment 81. The method of embodiment 80, wherein the target sequence is the sequence of a target provided in any one of Tables 6a-6f.
[0516] Embodiment 82. The method of embodiment 81, wherein the target is a coronavirus.
[0517] Embodiment 83. The method of embodiment 82, wherein the target is a SARS-CoV-2 virus.
[0518] Embodiment 84. The method of any one of embodiments 71-83, wherein the target DNA is cDNA and has been obtained by reverse transcription.
[0519] Embodiment 85. The method of any one of embodiments 71-79, wherein the target DNA is from a human cell.
[0520] Embodiment 86. The method of embodiment 85, wherein the target DNA is human fetal or cancer cell DNA.
[0521] Embodiment 87. The method of any one of embodiments 71-86, wherein the protein is a Cas12a.1 comprising the amino acid sequence of SEQ ID NO: 3 or an amino acid sequence having at least 70% sequence identity to SEQ ID NO: 3.
[0522] Embodiment 88. The method of any one of embodiments 71-86, wherein the protein is a Cas12p comprising the amino acid sequence of SEQ ID NO: 4 or an amino acid sequence having at least 70% sequence identity to SEQ ID NO: 4.
[0523] Embodiment 89. The method of any one of embodiments 71-86, wherein the protein is a Cas12p comprising the amino acid sequence of SEQ ID NO: 222 or an amino acid sequence having at least 70% sequence identity to SEQ ID NO: 222.
[0524] Embodiment 90. The method of any one of embodiments 71-86, wherein the protein is a Cas12q comprising the amino acid sequence of SEQ ID NO: 5 or an amino acid sequence having at least 70% sequence identity to SEQ ID NO: 5.
[0525] Embodiment 91. The method of any one of embodiments 71-87, wherein the sample comprises DNA from a cell lysate.
[0526] Embodiment 92. The method of any one of embodiments 71-87, wherein the sample comprises a cell.
[0527] Embodiment 93. The method of any one of embodiments 71-87, wherein the sample is a urine sample, a blood sample, a serum sample, a plasma sample, a lymphatic fluid sample, a cerebrospinal fluid sample, a saliva sample, a nasopharyngeal sample, an oropharyngeal sample, a nasopharyngeal / oropharyngeal sample, an aspirate sample, or a biopsy sample.
[0528] Embodiment 94. The method of any one of embodiments 71-93, comprising determining the amount of the target DNA present in the sample.
[0529] Embodiment 95. The method of embodiment 94, wherein the measuring a detectable signal comprises one or more of: visual-based detection, sensor-based detection, color detection, gold nanoparticle-based detection, fluorescence polarization, colloid phase change / dispersion, electrochemical detection, and semiconductor-based sensing.
[0530] Embodiment 96. The method of any one of embodiments 71-95, wherein the labeled detector comprises a modified nucleobase, a modified sugar moiety, and / or a modified nucleic acid linkage.
[0531] Embodiment 97. The method of any one of embodiments 71-96, further comprising detecting a positive control target DNA in a positive control sample, the detecting comprising:
[0532] a. contacting the positive control sample with:
[0533] i. a Cas12a.1, Cas12p, or Cas12q protein;
[0534] ii. a positive control gRNA comprising a region that binds the Cas12a.1, Cas12p, or Cas12q protein and a positive control spacer sequence that is hybridized to the positive control target DNA; and
[0535] iii. a labeled detector that is not hybridized to the positive control spacer sequence of the positive control gRNA; and
[0536] b. measuring a detectable signal resulting from cleavage of the labeled detector by the Cas12a.1, Cas12p, or Cas12q protein, thereby detecting the positive control target DNA.
[0537] Embodiment 98. The method of any one of embodiments 71-97, wherein the detectable signal is detectable in less than 15, 30, 45, 60, 90, 120, 150, 180, 210, or 240 minutes.
[0538] Embodiment 99. The method of any one of embodiments 71-98, further comprising amplifying the target DNA in the sample by loop-mediated isothermal amplification (LAMP), helicase-dependent amplification (HDA), recombinase polymerase amplification (RPA), strand displacement amplification (SDA), nucleic acid sequence-based amplification (NASBA), transcription-mediated amplification (TMA), nicking enzyme amplification reaction (NEAR), rolling circle amplification (RCA), multiple displacement amplification (MDA), ramification (RAM), circular helicase-dependent amplification (cHDA), single-primer isothermal amplification (SPIA), signal-mediated RNA amplification technology (SMART), self-sustained sequence replication (3SR), genome exponential amplification reaction (GEAR), or isothermal multiple displacement amplification (IMDA).
[0539] Embodiment 100. The method of any one of embodiments 71-99, wherein the target DNA in the sample is present at a concentration of less than 100 uM.
[0540] Embodiment 101. A protein comprising an amino acid sequence that is 70-99.5% homologous to SEQ ID NO: 1, 2, 3, 4, 5, 10, 11, or 222.
[0541] Embodiment 102. The protein of embodiment 101, wherein the sequence of the protein has been deduced bioinformatically.
[0542] Embodiment 103. A composition comprising any one of the proteins of embodiment 101 and optionally a pharmaceutically acceptable carrier.
[0543] Embodiment 104. A composition comprising any one of the proteins of embodiment 101, optionally comprising a pharmaceutically acceptable carrier, a nucleic acid stabilizing buffer, and / or a protein stabilizing buffer.
[0544] Embodiment 105. A composition comprising any one of the proteins of embodiment 101, wherein the protein is lyophilized, and optionally further comprising any one or more of: a labeled detector, a reverse transcriptase, and reagents for loop-mediated isothermal amplification.
[0545] Embodiment 106. A DNA polynucleotide comprising a nucleotide sequence encoding any one of the proteins of embodiment 101.
[0546] Embodiment 107. A recombinant expression vector comprising the DNA polynucleotide of embodiment 106.
[0547] Embodiment 108. The recombinant expression vector of embodiment 107, wherein the nucleotide sequence encoding a single protein is operably linked to a promoter.
[0548] Embodiment 109. A host cell comprising the DNA polynucleotide of any one of embodiments 106-108.
[0549] Embodiment 110. A pharmaceutical composition comprising any one of the engineered systems of embodiments 1-39, and optionally a pharmaceutically acceptable carrier.
[0550] Embodiment 111. A composition comprising any one of the engineered systems of embodiments 1-39, and optionally comprising a nucleic acid stabilizing buffer and / or a protein stabilizing buffer.
[0551] Embodiment 112. A pharmaceutical composition comprising any one of the single molecule gRNAs of embodiments 40-57, and optionally a pharmaceutically acceptable carrier.
[0552] Embodiment 113. A composition comprising any one of the single molecule gRNAs of embodiments 40-51, and optionally a nucleic acid stabilizing buffer and / or a protein stabilizing buffer.
[0553] Embodiment 114. A DNA polynucleotide comprising a nucleotide sequence encoding any one of: any nucleic acid of embodiments 3, 27, or a gRNA of embodiments 40-51.
[0554] Embodiment 115. A recombinant expression vector comprising the DNA polynucleotide of embodiment 114.
[0555] Embodiment 116. The recombinant expression vector of embodiment 115, wherein the nucleotide sequence encoding a single gRNA is operably linked to a promoter.
[0556] Embodiment 117. A host cell comprising the DNA polynucleotide of any one of embodiments 114-116.
[0557] Embodiment 118. A kit comprising one or more components of any one of the engineered systems of embodiments 1-39.
[0558] Embodiment 119. The kit of embodiment 118, wherein one or more components are lyophilized.
[0559] Embodiment 120. The kit of any one of embodiments 118-119, wherein the one or more components comprise a Cas12p, a labeled RNA reporter, and a gRNA directed to SARS-CoV-2.
[0560] Embodiment 121. A method of isolating a Class 2 Type II or Class 2 Type V CRISPR-Cas protein from a metagenomic sample, the method comprising using a bioinformatics-based method.
[0561] Embodiment 122. The method of embodiment 121, wherein the Class 2 Type II or Class 2 Type V CRISPR-Cas protein is selected from the group consisting of SEQ ID NOs: 1, 2, 3, 4, 5, 10, 11, and 222.
[0562] Example
[0563] The following examples are included to demonstrate the present application and are not intended to limit the scope of the application.
[0564] Example 1: Identification and validation of novel Class II Type II and Type V endonucleases
[0565] 2 Class II and V Type CRISPR-Cas Locus Identification
[0566] Metagenomic sequences were obtained from NCBI and curated to build a database of putative CRISPR-Cas loci. CRISPR arrays were identified using the CrisprCasFinder software. Filtering criteria was for putative Class II Type II and Type V effectors >500 aa adjacent to Cas genes and CRISPR arrays. Sequence alignments were performed using Clustal Omega using HMM profiles. Novel Cas9.1, Cas9.2, Cas9.3, Cas9.4, Cas12a.1, Cas12p, and Cas12q proteins were identified as described herein.
[0567] Expression Plasmid and Non-coding Element Production
[0568] The minimum conditions for validating the Cas protein were established as a cloning strategy. A minimal CRISPR locus was designed by removing the acquisition protein and generating a minimal array with a single spacer (Sp1). The native Sp1 sequence was replaced with a known specific target sequence of the naturally occurring sequence length (GTGGCAGCTCAAAAATTGGCTACAAAACCAGTT; SEQ ID NO: 118) for target detection and PAM screening assays. The codon-optimized protein sequence of the CRISPR effector and / or accessory proteins was placed in a pET-based expression vector (EMD-Millipore) under the transcriptional control of the lac and IPTG-inducible T7 promoters.
[0569] artificial synthesis
[0570] For Cas12a.1, Cas12p, Cas9.1, and Cas9.2, the expression vectors were synthetically produced. Effector plasmid codon optimization, synthesis, and cloning were performed by the provider (GeneScript). Flanking restriction sites were added to the CRISPR array to clone DNA fragments (IDTs), taking into account two putative transcriptional directions. This was done using the same elements in the opposite direction to produce a second construct variant. FIG. 1A to FIG. 1B The expression vector maps of Cas9.1 and Cas9.2 are shown. FIG. 2A to FIG. 2C The expression vector maps of Cas12a.1, Cas12p, and Cas12q are shown. The vector sequences are provided in Table 8.
[0571] Table 8. Expression vector sequences
[0572]
[0573]
[0574]
[0575]
[0576]
[0577]
[0578]
[0579]
[0580]
[0581]
[0582]
[0583]
[0584]
[0585]
[0586]
[0587]
[0588]
[0589]
[0590] Protein expression and purification
[0591] Cas12 coding sequences were codon-optimized and synthesized by GeneScript and then cloned into pET28a (Novagen) with N-terminal 6xHis tagging. Cas12 expression plasmids were transformed into E. coli NiCo21(DE3) (NEB). For protein expression, individual clones were first cultured overnight in 5-mL liquid LB tubes and then inoculated into 400 ml fresh liquid LB (OD 600 0.1). Cells were grown at 200 rpm and 37°C until OD 600 reached 0.8, then IPTG was added to a final concentration of 0.1 mM, then cells were further cultured at 37°C for about 2 hours before cell harvesting. Cells were resuspended in 20 mL buffer A (50 mM Tris-HCl pH 8.0, 0.5 M NaCl, 1 mM DTT and 5% glycerol) with protease inhibitor cocktail (Promega) and 5 mg / ml lysozyme. After 15 minutes incubation at 37°C, cells were lysed by sonication with a 10 minutes 10 seconds run and 10 seconds stop cycle. Cell debris and insoluble particles were removed by centrifugation (15,000 rpm, 30 minutes). After centrifugation, the supernatant was loaded onto a 5 mL Crude HisTrap column (GE Healthcare) equilibrated with 20 mM imidazole buffer A on an AKTA Pure 25L device (GE Healthcare Life Sciences). Elution was performed by a step gradient of buffer B (buffer A plus 0.5 M imidazole). Eluate was dialyzed against dialysis buffer (50 mM Tris-HCl pH 8.0, 200 mM NaCl, 1 mM DTT and 5% glycerol).
[0592] Guide RNA (gRNA) and variants
[0593] Including positive repeat mutations can improve gRNA stability. The positive repeats from the three CRISPR Cas12 systems presented herein contain two A:U base pairs within the stem-loop region. Increasing the thermal stability of the stem-loop is expected to increase the portion of properly folded crRNA available for loading into its homologous Cas12, thereby increasing nuclease activity (Pengpeng et al., 2019). In the positive repeats of the CRISPR systems disclosed herein, those A:U base pairs are replaced with C:G to create new, more stable, non-naturally occurring variants based on minimum free energy predictions of RNA folding.
[0594] The predicted (proposed) naturally occurring normal repeat sequences of the Cas protein found in bacterial DNA at CRISPR loci are shown in Tables 2 and 5a above (shown as DNA sequences). Novel variants are shown in Table 5b above (shown as DNA sequences). The predicted secondary structures are as follows... FIG. 7A to FIG. 7C As shown in the diagram. It is anticipated that the entire or partial orthogonal repeat sequence will form a functional, non-naturally occurring gRNA and bind to the Cas protein of this disclosure. The RNA forming the orthogonal repeat variant and spacer used in this embodiment was synthesized by Synthego.
[0595] FIG. 3B , FIG. 3E , FIG. 3G , FIG. 5B , FIG. 5D and FIG. 5F The predicted secondary structures (folds) of the repetitive sequences of Cas9.1, Cas9.3, Cas9.4, Cas12a.1, Cas12p, and Cas12q pre-crRNAs are shown. The publicly available RNAfold webserver tool was used to assemble these predictions.
[0596] In vitro transcription (IVT)
[0597] Use MEGAscript according to the manufacturer's instructions. TM In vitro transcription was performed using the T7 transcription kit (Ambion, Invitrogen), and the transcription was performed according to the manufacturer's instructions. RNA cleaning was performed using an RNA cleaning kit (New England Biolab). RNA was visualized on a 2% agarose gel using gel loading buffer II (Ambion, Invitrogen).
[0598] In vitro target cutting assay
[0599] The template sequences used for in vitro target cutting assays are shown in Table 9.
[0600] Table 9
[0601]
[0602]
[0603] gBlocks (Table 9) are double stranded DNA templates of approximately 100-500 nt synthesized by IDT with sequences including the target of interest. Specific cleavage assays containing 1 ug of gBlock target sequence were performed in buffer NEB 3 with 30 nM Cas (Cas9.1, Cas9.2, Cas9.3, Cas9.4 Cas12a.1, Cas12p, Cas12q), 30 nM crRNA directed against the specific sequence for 2 hours at 37 °C. Reactions were stopped at 70 °C for 10 minutes. Products were cleaned using PCR purification columns (QIAGEN) and visualized in a 1% agarose gel pre-stained with SYBER Gold (Invitrogen). To identify the type of cleavage (staggered / flush), an aliquot of the digestion products was run in a 1% agarose gel and bands corresponding to the cleaved target were gel extracted using the DNA clean and concentrate kit (Zymo Research). Purified products were sequenced using specific primers and analyzed by DNA STAR. For the nicking activity assays, buffer NEB 3 with 30 nM Cas (Cas9.1, Cas9.2, Cas9.3, Cas9.4, Cas12a.1, Cas12p, Cas12q), 30 nM crRNA and 1 nM ssDNA activator containing the target sequence were used for 10, 20, 40 and 60 minutes at 37 °C. Reactions were initiated by the addition of 250 nM M13 ssDNA or M13 dsDNA plasmid (NEB). Reactions were stopped at 70 °C for 10 minutes. Product separation was performed in a 2% agarose gel pre-stained with SYBER Gold (Invitrogen)
[0604] Fluorescent detection of nicking activity
[0605] Fluorescent detection can be performed to determine collateral cleavage activity. With a 40 pl reaction final volume, 30 nM Cas12 was complexed with 30 nM crRNA and 50 nM DNaseAlert™ substrate (IDT) in buffer NEB 2.1 at 37 °C. Reactions can be monitored in a fluorescent plate reader at 37 °C for up to 30 minutes with a fluorescence measurement every 2 minutes in the HEX channel (l ex : 536 nm; l em : 556 nm). The resulting data can be background corrected using readings obtained in the absence of target. For FQ detection of collateral cleavage of dsDNA / ssDNA and dsRNA / ssRNA, DNaseAlert™ (IDT) and
[0606] Cis and trans cleavage rates
[0607] Initial velocity (V0) can be calculated by fitting a linear regression and plotted against substrate concentration according to the following equation to determine Michaelis-Menten constants (GraphPad software): Y = (Vmax x X) / (Km+X), where X is substrate concentration, y is enzyme velocity. Turnover number (kcat) is determined from the following equation: kcat = Vmax / Et, where Et = 0.1 nM.
[0608] Example 2: Determining endonuclease activity
[0609] It was investigated whether providing only crRNA with the novel Cas12a.1 and Cas12p of the present disclosure can cleave target DNA in vitro. Cas12a.1 and Cas12p were designed, overexpressed, purified in vitro and used to form complexes with crRNA against specific targets. It was found that the presence of Cas12 proteins and cRNA is sufficient to form active complexes for mediating DNA cleavage.
[0610] Example 3: Determining PAM sequence specificity
[0611] To demonstrate the PAM sequence cleavage-dependent action of Cas12a.1 and Cas12p of the present disclosure, ten different PAM motifs were designed following specific target sequences. Using these, TCTN and TGTN were identified as effective PAM sequences for Cas12a.1 and Cas12p, respectively, among the ten motifs tested. FIG. 8 Bar graphs showing PAM sequence preference of Cas12a.1 and Cas12p for ten PAM motifs are shown, using a fluorescence assay to measure the performance of Cas12a.1 and Cas12p. The resulting fluorescence data is background subtracted.
[0612] Example 4: Demonstration of collateral cleavage activity of Cas12a.1 and Cas12p, and their ability to cleave ssDNA and RNA reporters FIG. 9B
[0613] The Cas12a.1 and Cas12p proteins of the disclosure were investigated for their ability to cleave dsDNA or RNA. Cas12a.1-gRNA or Casp-gRNA complexes were mixed with samples (positive and negative) and reporters to react in the presence of the target. In these examples, a custom ssDNA fluorescently labeled reporter (5’FAM-TTATTATT-3’IABkFQ 3’-IDT) (SEQ ID NO: 121) and a commercial fluorescently labeled reporter RNA reporter (Cat N 11-04-03-03-IDT) were used.
[0614] FIG. 9C The Cas12a.1 and Cas12p proteins of the disclosure were investigated for their ability to cleave dsDNA or RNA. Cas12a.1-gRNA or Casp-gRNA complexes were mixed with samples (positive and negative) and reporters to react in the presence of the target. In these examples, a custom ssDNA fluorescently labeled reporter (5’FAM-TTATTATT-3’IABkFQ 3’-IDT) (SEQ ID NO: 121) and a commercial fluorescently labeled reporter RNA reporter (Cat N 11-04-03-03-IDT) were used. -1 substrate (25 single use tubes. Catalog number. 11-04-03-03-IDT). The exemplary ssDNA reporter used in this and other examples provided herein is (5’FAM-TTATTATT-3’IABkFQ 3’-IDT) (SEQ ID NO: 121).
[0615] Example 5: Thermostability testing The Cas12p was shown to exhibit ssDNA and RNA reporter cleavage, using SARS-CoV-2 inactivated virus as sample as target.
[0616] FIG. 10
[0617] The activity of Cas12a.1 and Cas12p was tested at different temperatures.
[0618] FIG. 10 The Cas12a.1 and Cas12p proteins were shown to be active at 25°C, using 1 uM complex, 300 nM reporter SARS-CoV-2 (Spn2 target) at 1 minute and 5 minutes as end point for readout.
[0619] FIG. 14 and FIG. 15 shows that Cas12p performs as well at 25°C as at 37°C.
[0620] FIG. 16 shows the differential performance of Cas12p from LbCas12a at 25°C in generating a fluorescent signal by reporter cleavage. LbCas12a and Cas12p were incubated with their respective gRNAs to target the N gene of SARS-CoV-2 to form 1 uM complexes. The targets were the same for both and were provided at a concentration of 10 nM. 600 nM ssDNA reporter was added to the reaction mixture (50 mM NaCl, 10 mM Tris-HCl, 10 mM MgCl2, and 100 pg / ml BSA). The collateral cleavage was measured by fluorescence and read out in real time. Example 6: Testing multiple salt concentrations shows the differential performance of Cas12p from LbCas12a at 25°C using SARS-CoV-2 as a target, described in Example 10.
[0621] FIG. 11
[0622] Activity of Cas12a.1 and Cas12p was tested at various concentrations of NaCl; Cas12a.1 and Cas12p showed that they remained functional. shows the activity of the two proteins at various concentrations of NaCl. The resulting fluorescence data is background subtracted.
[0623] Example 7: Testing of various commercial buffers
[0624] Cas12a.1 and Cas12p of the disclosure show different performance in various commercial buffers. Figure 12 shows the performance of Cas12a.1 and Cas12p of the disclosure in three different commercial buffers. The resulting fluorescence data is background subtracted.
[0625] Example 8: Detection of Hantavirus using Cas12a.1 and Cas12p
[0626] Hantaviruses are a family of viruses that are primarily transmitted by rodents and can cause a variety of disease symptoms in humans worldwide. Hantavirus disease can be produced in humans by infection with any of the Hantaviruses. The use of the novel Cas12a.1 and Cas12p proteins of the disclosure to detect Hantaviruses is described below.
[0627] Primer design and CRISPR RNA guide selection
[0628] The complete sequence of the Hantavirus genome Andes virus segment S is provided below
[0629]
[0630]
[0631] The following exemplary sequence was chosen as a target for the spacer gRNA for Hantavirus detection: GTGGCAGCTCAAAAATTGGCTAC (SEQ ID NO: 70) (underlined above). Other sequences can be chosen for targeting.
[0632] CRISPR guide design and synthesis
[0633] A gRNA was designed using a spacer specific to the Hantavirus target sequence. The guide (including the forward repeat (single underline) + target-complementary sequence (double underline)) is shown below: AAATTTCTACTGTAGTAGAT GTGGCAGCTCAAAAATTGGCTAC (SEQ ID NO: 249)
[0634] For native expression and processing of the gRNA, a minimal array with forward repeats and target-complementary sequences from Cas12a.1 and Cas12p was cloned in a Cas expression vector. In vivo CRISPR complex formation was performed in bacterial NiCo21(DE3) competent E. coli and purified from bacterial extracts. In other variations, the guide can be synthesized in vitro and complexed with the Cas protein.
[0635] The complex was added to a mixture containing a molecular reporter and a fluorescent dye. The sample to be tested was added to the mixture. The sample to be tested can be: a sample obtained directly from a subject; a sample obtained from a subject, then diluted and / or treated; DNA (which can be amplified) or RNA in a sample taken from a subject; or the sample to be tested can be cDNA made from RNA of the sample. The sample can be further amplified, for example using RPA (Recombinase Polymerase Amplification, for example, using RPA TwistAmp Basic (TABAS03)).
[0636] The components for forming the CRISPR complex were mixed in order as shown in Table 10. The complex was prepared and allowed to incubate at room temperature for 10 minutes.
[0637] Table 10
[0638] Components [Master Mix] [Final] Volume (1X) Nuclease-free water 15.05 Buffer NEB 2.1 10X 1X 2.0 RNA guide working solution 300 nM 30 nM 2.2 Cas12a.1 working solution 1 uM 30 nM 0.8 Total 20.0
[0639] The components for forming the CRISPR mixture were mixed in order as shown in Table 11.
[0640] Table 11
[0641]
[0642]
[0643] Result readout
[0644] Reactions were monitored in a fluorescent plate reader at 37°C for up to 30 minutes, with a fluorescence measurement every 2 minutes in the HEX channel (λex: 536 nm; λem: 556 nm) or at the final endpoint. The resulting data was background corrected using the readout obtained in the absence of target.
[0645] Figure 9A Specific cleavage activity of Cas12a.1 and Cas12p proteins of the disclosure and Hanta target is shown. The pGEM plasmid was cloned with the Hanta target (pGEM-Hanta) and used to demonstrate specific cleavage activity of Cas12a.1 and Cas12p. Cas12a.1 and Cas12p were incubated with their respective gRNA at 37°C for 2 hours to target the Hanta target and exposed to either the gGEM-Hanta plasmid or the gGEM plasmid without target. The arrows show cleavage of the pGEM-Hanta plasmid but not the pGEM, demonstrating that the cleavage is specific to the Hanta target.
[0646] Using the nicking activity, it is possible to detect Hanta virus RNA in less than 1 hour at picomolar concentrations, as shown in Figure 13 Figure 13 RPA-free sensitivity curves of Cas12a.1 and Cas12p of the disclosure are shown, measuring each target concentration for 30 minutes.
[0647] Example 9: Cas12p characterization
[0648] Cas12p was further characterized and compared to LbCas12a (SEQ ID NO: 122 (SEQ ID NO: 242, from US 9790490)) to support the characterization of this new subtype of Cas12.
[0649] Figure 14 Fluorescence detection amounts by Cas12p are shown to be equivalent at 37°C and 25°C for a target DNA reverse transcribed from SARS-CoV-2 RNA, indicating thermal stability and functionality at room temperature.
[0650] Figure 15 and below show the kinetic performance of Cas12p with LbCas12a at room temperature.
[0651]
[0652] Figure 16 Further showing the differential performance of Cas12p from LbCas12a at room temperature.
[0653] As described above, Figure 9A The Cas12a.1 and Cas12p proteins of the disclosure are shown to have specific cleavage activity with exemplary Hantavirus targets, as described in the previous example. Figure 9B The Cas12a.1 and Cas12p proteins of the disclosure are shown to have collateral cleavage activity, using Hantavirus as an exemplary target, as described in the previous example. Figure 9C Collateral cleavage activity of novel Cas12p proteins against SARS-CoV-2 targets is shown, as described in Example 10.
[0654] Figure 17 The ability of Cas12p to cleave ssDNA and RNA reporters is shown, as tested in various targets (Hantavirus, SARS-CoV-2) as examples. Cas12p was incubated with gRNA against either Hantavirus or SARS-CoV-2 virus to form a 1 uM complex and exposed to DNA targets at 10 nM concentration, with ssDNA or RNA fluorescently labeled reporters added to the mixture at concentrations between 1 and 0,5 uM. Controls contained no specific DNA target. Collateral cleavage activity was only visible in the presence of targets for ssDNA and RNA.
[0655] Example 10: Detection of SARS-CoV-2 using Cas12a.1
[0656] An example of using Cas12p to detect SARS-CoV-2 in upper respiratory samples during acute infection is provided herein. A positive result indicates the presence of SARS-CoV-2 RNA. Further clinical correlation with patient history and other diagnostic information can be utilized to determine the patient’s infection status.
[0657] Assay
[0658] RNA was purified from 140 μΐ of nasopharyngeal / oropharyngeal sample using the QIAmp Viral RNA Mini Kit (QIAGEN) following the instructions in the user guide and eluted in 60 μΐ. If the RNA is not tested immediately, the RNA was stored at -70 °C.
[0659] Following RNA purification, detection of SARS-CoV-2 genomic RNA using the CASPR Lyo-CRISPR SARS-CoV-2 kit was performed using Figure 18 The two-step procedure summarized in Example 1 and outlined below was performed.
[0660] Step 1: Purified RNA is subjected to reverse transcription and amplification. Reverse transcription loop-mediated isothermal amplification (RT-LAMP) is used to reverse transcribe and amplify 5 pl of purified RNA, where the primer set is specifically designed to target the highly conserved N gene of the SARS-CoV-2 viral genome.
[0661] The RT-LAMP reaction is based on a total of three (3) pairs of primers that amplify specific sequences in the N gene of the SARS-CoV-2 RNA.
[0662] The RT-LAMP reaction is performed by incubating at 62°C for 30 minutes.
[0663] Step 2: Following the RT-LAMP reaction, detection of the amplified viral target is performed using a Cas12a.1 ribonucleoprotein complex (RNP complex) that contains Cas12a.1 + a gRNA (single molecule guide) that targets the amplified viral N gene sequence of Step 1. The sequence targeted by the gRNA in the cDNA made from the viral RNA is as follows: GATCGCGCCCCACTGCGTTCTCC (SEQ ID NO: 119)
[0664] If SARS-CoV-2 genomic RNA is present in the sample and is amplified during the RT-LAMP reaction, the gRNA from the RNP complex can bind to the DNA target and trigger the collateral activity of Cas12a.1, degrading the 5’FAM-3’Quencher single-stranded DNA (ss-DNA) reporter molecule to cause fluorescence emission. Fluorescence measurement can be performed in a standard plate reader with fluorescence capability.
[0665] From start to finish - from obtaining the sample to reading out the result, the assay is completed in less than 60 minutes. Figure 18 An illustrative workflow for detecting SARS-CoV-2 described in this example is shown.
[0666] Additional negative control, positive control, and extraction control are included.
[0667] Negative control: Nuclease-free water is used to identify any potential contamination of the assay run.
[0668] Positive control: A synthetic sequence identical to the target sequence is provided in a separate vial at a concentration of 2000 cp / ml. The positive control confirms that the assay is performing as expected.
[0669] Extraction control: A primer set targeting the human housekeeping gene RNAse P (e.g.) is included in the RT-LAMP reaction mixture to ensure proper execution of the extraction process.
[0670] The reagents used are provided in lyophilized form, reducing the human source of operator error.
[0671] Results
[0672] For the negative control (NTC), the ratio between the fluorescence measured at the end point (t=20 minutes) and the fluorescence at the start of the run (t=0 minutes) was calculated
[0673]
[0674] For the positive control and the clinical samples, the ratio between the sample reaction fluorescence measured at the end point (t=20 minutes) and the corresponding effective negative template control reaction fluorescence measurement at 20 minutes was calculated.
[0675] - for the positive control (PC)
[0676]
[0677] - for the clinical samples
[0678]
[0679] Once the ratios for the controls and samples were calculated, the results were calculated according to the following control assay criteria:
[0680]
[0681] In this example, for the unknown clinical samples: the ratio for the positive samples should be >3 (at least a 3-fold increase in fluorescence emission between the sample reaction and the negative control reaction at t=20 minutes).
[0682] In this example, the ratio for the negative samples should be <3 (less than a 3-fold increase in fluorescence emission between the sample reaction and the negative control reaction at t=20 minutes). To confirm the negative result, the RNAse P should have a value of >3 (an increase in fluorescence emission between the sample reaction and the negative control reaction at t=20 minutes).
[0683] Performance evaluation - Analytical sensitivity, limit of detection
[0684] The limit of detection (LoD) study established the lowest SARS-CoV-2 concentration (genome copies (cp) per pL input) that can be detected at least 95% of the time.
[0685] To determine the LoD, serial dilutions of whole inactivated SARS-CoV-2 were spiked into negative nasopharyngeal samples and processed according to the procedure described above.
[0686] LoD was determined by testing three (3) different dilutions (10 copies / μΐ, 5 copies / μΐ, 2.5 copies / μΐ) in triplicate (3) and corresponds to the lowest concentration (5 copies / μΐ) at which 3 / 3 replicates were tested positive. This preliminary LoD (5 copies / μΐ) was confirmed by testing in twenty (20) replicates at each concentration with a preliminary LoD of 0.5X-1X - 1.5X - 2X. LoD is the lowest concentration at which at least 19 / 20 replicates were tested positive for the target.
[0687] LoD was confirmed as 7.5 copies / μL with a detection rate of 95% (19 / 20). Results are summarized in the table below:
[0688] Table 12
[0689] Replicate Ratio Result 1 >3 Positive 2 >3 Positive 3 >3 Positive 4 >3 Positive 5 >3 Positive 6 >3 Positive 7 >3 Positive 8 >3 Positive 9 >3 Positive 10 >3 Positive 11 >3 Positive 12 <3 Negative 13 >3 Positive 14 >3 Positive 15 >3 Positive 16 >3 Positive 17 >3 Positive 18 >3 Positive 19 >3 Positive 20 >3 Positive
[0690] Performance evaluation - Analytical sensitivity, inclusivity
[0691] Inclusivity was demonstrated by comparing the SARS-CoV-2 assay primers and gRNAs to the alignment of 4703 SARS-CoV-2 sequences available in GISAID as of May 16, 2020. The dataset was further refined by considering only full genome sequences (>29000bp) and by eliminating low quality sequences with ambiguous sequencing data (N) and of animal origin. This computer analysis showed that the primers and gRNA sequences used have 99.9% homology with all available circulating SARS-CoV-2 sequences.
[0692] Performance evaluation - Analytical specificity
[0693] Assay 2 is based on a set of primers and unique gRNAs designed for the specific detection of SARS-CoV-2.
[0694] To evaluate the analytical specificity, a computer analysis was first performed using the NCBI Blast tool to confirm that there is no potential cross-reactivity between any of the primers / gRNA sequences and normal and pathogenic organisms of the respiratory tract.
[0695] Results are summarized in Table 13:
[0696] Table 13
[0697]
[0698]
[0699] These results indicate that only a few microorganisms have >80% homology between their genomic sequence and at least one of the SARS-CoV-2 primers or gRNAs included in the assay.
[0700] To confirm the computer evaluation, the same pathogens were examined in vitro to check for potential cross-reactivity and interference.
[0701] A total of 22 pathogens were analyzed during the lysis step of the extraction procedure by spiking genomic DNA / RNA or inactivated strains into SARS-CoV-2 negative nasopharyngeal samples at the concentrations indicated in Table 15 and tested using the assay described herein. Each pathogen was tested in triplicate. To discard any false negative results, a RNAseP assay was run in parallel for each sample,
[0702] Interference analysis was also performed for microorganisms showing >80% homology with the SARS-CoV-2 primers or gRNAs included in the kit. To detect any potential interference, the analysis was performed following the same protocol used for the cross-reactivity test in the presence of 3X LoD SARS-CoV-2 (22.5 cp / µl).
[0703] All negative results for the tested pathogens were confirmed by positive results in the RNAseP assay.
[0704] Table 14
[0705]
[0706]
[0707] In summary, based on the computer and in vitro analysis, no cross-reactivity or interference between the primers / gRNAs included in the assay and the most common pathogens in the respiratory tract was expected.
[0708] Clinical evaluation
[0709] The assay was clinically evaluated using nasopharyngeal swabs as clinical samples from male and female adult patients with signs and symptoms of upper respiratory tract infection.
[0710] A total of 30 positive and 30 negative samples were collected to evaluate the performance and tested using RNA extraction with the QIAmp viral RNA mini kit following the described procedure (indicated in Table 15 as “Cas12a.1-based assay”). All samples were also tested using the RT-PCR test as a comparative method to obtain the positive and negative percent agreement values. The results are presented in Table 15 and show 100% positive percent agreement (PPA) and 100% negative percent agreement (NPA) with the comparator method.
[0711] Table 15
[0712]
[0713] Example 11: Detection of SARS-CoV-2 using Cas12p
[0714] Provided herein are examples of using Cas12p to detect SARS-CoV-2 in upper respiratory samples during acute infection. A positive result indicates the presence of SARS-CoV-2 RNA. Further clinical correlation with patient history and other diagnostic information can be utilized to determine the patient’s infection status.
[0715] Assay
[0716] Nasopharyngeal / nasal swabs were inserted into 500 uL lysis buffer, vortexed for 2 minutes, and 100 uL of lysed sample was transferred to a 1,5 mL capacity tube and heated at 95 °C for 5 minutes.
[0717] Following sample processing, detection of SARS-CoV-2 genomic RNA using the CASPR Direct Lyo-CRISPR SARS-CoV-2 kit was performed using the two-step procedure summarized in Table 15 and outlined below. Figure 19
[0718] Step 1: Lysed samples were subjected to reverse transcription and amplification. Reverse transcription and amplification was performed on 10 μΐ of lysed sample using reverse transcription loop-mediated isothermal amplification (RT-LAMP) with a primer set specifically designed to target two highly conserved N genes and one highly conserved ORF1ab gene of the SARS-CoV-2 viral genome.
[0719] The RT-LAMP reaction was based on a total of three (9) pairs of primers that amplify two specific sequences in the N gene and one specific sequence in the ORF1ab gene of the SARS-CoV-2 RNA.
[0720] The RT-LAMP reaction was performed by incubation at 62 °C for 60 minutes.
[0721] Step 2: Following the RT-LAMP reaction, detection of the amplified viral target is performed using a Cas12p ribonucleoprotein complex (RNP complex) comprising Cas12p + three gRNAs (single molecule guides) targeting the amplified viral N and ORF1ab gene sequences of Step 1. The sequences targeted by the gRNAs in the cDNA made from the viral RNA are as follows: GATCGCGCCCCACTGCGTTCTCC (SEQ ID NO: 119), AUGGCACCUGUGUAGGUCAACCA (SEQ ID NO: 120), and UGUGCUGACUCUAUCAUUAUUGG (SEQ ID NO: 123).
[0722] If SARS-CoV-2 genomic RNA is present in the sample and is amplified during the RT-LAMP reaction, the gRNAs from the RNP complex can bind to the DNA target and trigger the collateral activity of Cas12p, degrading the 5’FAM-3’Quencher single-stranded reporter molecule to cause fluorescence emission. Fluorescence measurement can be performed in a standard plate reader with fluorescence capability.
[0723] From start to finish - from obtaining the sample to reading the results, the assay is completed in less than 75 minutes. Figure 18 Figure 19 An illustrative workflow for detecting SARS-CoV-2 is shown.
[0724] Additional negative control, positive control, and extraction control are included.
[0725] Negative control: Nuclease-free water is used to identify any potential contamination of the assay run.
[0726] Positive control: A synthetic sequence identical to the target sequence is provided in a separate vial at a concentration of 2000 cp / ml. The positive control confirms that the assay is performing as expected.
[0727] Extraction control: A primer set targeting the human housekeeping gene RNAse P (e.g.) is included in the RT-LAMP reaction mix to ensure proper execution of the extraction process.
[0728] The reagents used are provided in lyophilized form, reducing the human source of operator error.
[0729] Results
[0730] For the negative control (NTC), the ratio between the fluorescence measured at the end point (t = 5 minutes) and the fluorescence at the start of the run (t = 0 minutes) is calculated.
[0731]
[0732] For positive controls and clinical samples, the ratio between the sample reaction fluorescence (IF) measured at endpoint (t=5 minutes) and the corresponding effective negative template control reaction fluorescence measurement at 5 minutes was calculated.
[0733]
[0734] Once the ratio of controls and samples was calculated, the results were calculated according to the following control assay criteria:
[0735] Table 16
[0736]
[0737] In this example, for unknown clinical samples: the ratio of positive samples should be > 2.5 (at least 2.5-fold increase in fluorescence emission between sample reaction and negative control reaction at t=5 minutes).
[0738] In this example, the ratio of negative samples should be < 2.5 (less than 2.5-fold increase in fluorescence emission between sample reaction and negative control reaction at t=5 minutes). To confirm negative results, RNAse P should have a value of > 2.5 (increase in fluorescence emission between sample reaction and negative control reaction at t=5 minutes).
[0739] Performance evaluation - Analytical sensitivity, limit of detection
[0740] The limit of detection (LoD) study established the lowest SARS-CoV-2 concentration (genome copies (cp) per pL input) that can be detected at least 95% of the time.
[0741] To determine the LoD, serial dilutions of whole inactivated SARS-CoV-2 were spiked into lysis buffer with negative nasal matrix and processed according to the procedure described above.
[0742] The LoD was determined by testing three (3) different dilutions (25 copies / pL, 12.5 copies / pL, 6.125 copies / pL) in three (5) replicates and corresponds to the lowest concentration (25 copies / pL) at which 3 / 3 replicates were tested positive. This preliminary LoD (25 copies / pL) was confirmed in twenty (20) replicates. The LoD is the lowest concentration at which at least 20 / 20 replicates were tested positive for the target.
[0743] The LoD was confirmed to be 25 copies / pL with a detection rate of 100% (20 / 20). The results are summarized in the table below:
[0744] Table 17
[0745]
[0746]
[0747] Performance evaluation - analysis sensitivity, inclusivity
[0748] Inclusivity was demonstrated by comparing the alignment of SARS-CoV-2 assay primers and gRNAs to 4703 SARS-CoV-2 sequences available in GISAID as of May 16, 2020. The dataset was further refined by only considering full genome sequences (>29000bp) and by eliminating low quality sequences with ambiguous sequencing data (N) and of animal origin. This computer analysis showed that all the primer and gRNA sequences used were 100% homologous to all the available circulating SARS-CoV-2 sequences
[0749] with 100% homology.
[0750] Performance evaluation - analysis specificity
[0751] Assay 2 is based on a set of primers and gRNAs designed for the specific detection of SARS-CoV-2.
[0752] To evaluate the analysis specificity, a computer analysis was first performed using the NCBI Blast tool to confirm that there was no potential cross-reactivity between any of the primer / gRNA sequences and normal and pathogenic organisms of the respiratory tract.
[0753] The results are summarized in Table 18:
[0754] Table 18
[0755]
[0756]
[0757] These results show that only a few microorganisms have a genomic sequence with >80% homology with at least one of the SARS-CoV-2 primers or gRNAs included in the assays.
[0758] To confirm the computer evaluation, the same pathogens were examined in vitro to check for potential cross-reactivity and interference.
[0759] A total of 22 pathogens were analyzed by spiking genomic DNA / RNA or inactivated strains into SARS-CoV-2 negative lysed samples at the concentrations shown in Table 19 and tested using the assays described herein. Each pathogen was tested in triplicate. To discard any false negative results, a RNAseP assay was run in parallel for each sample,
[0760] Microorganisms that showed >80% homology to the SARS-CoV-2 primers or gRNAs included in the kit were also subjected to interference analysis. To detect any potential interference, the analysis was performed in the presence of 3X LoD SARS-CoV-2 (75 cp / µl) following the same protocol used for cross-reactivity testing.
[0761] All negative results for the pathogens tested were confirmed by positive results in the RNAseP assay.
[0762] Table 19
[0763]
[0764]
[0765] In summary, based on computer and in vitro analysis, no cross-reactivity or interference between the primers / gRNAs included in the assay and the most common pathogens in the respiratory tract was expected.
[0766] Clinical evaluation
[0767] The assay was clinically evaluated using nasopharyngeal swabs as clinical samples from male and female adult patients with signs and symptoms of upper respiratory tract infection.
[0768] A total of 47 positive samples and 43 negative samples were collected to evaluate the performance. All samples were also tested using the RT-PCR test as a comparative method to obtain the positive and negative percent agreement values. The results are presented in Table 20 and show a 97.9% positive percent agreement (PPA) and 100% negative percent agreement (NPA) with the comparator method.
[0769] Table 20
[0770]
[0771]
[0772] Figure 20 It was shown that Cas12p has minimal background signal after 30-60 minutes of cleavage activity. This provides an advantage for low viral concentrations and indicates stability of the lyophilized format. Figure 21 It was shown that diagnostic assays using Cas12p at room temperature can be read out in a paper format. Figure 22 It was shown that diagnostic assays using Cas12p at room temperature can be read out in a well plate with a fluorescence detector.
[0773] Example 12: Detection of SARS-CoV-2 using Cas12p and RNA guide
[0774] SARS-CoV-2 RNA in patient and control samples was detected using lyophilized beads with RNA-based reporters. For this example, a subset of samples described in Example 11 was used. Cas12p were pre-incubated with their respective sgRNA and the labeled RNA reporter was added prior to the lyophilization process. Pre-amplified RT-LAMP products were used as input. The input for the RT-LAMP reaction was lysed samples from patient and negative control nasopharyngeal swabs. Figure 19 A workflow for the detection of SARS-CoV-2 from samples using Cas12p / guide complex using RNA reporters is shown. Figure 24 Results of SARS-CoV-2 detection using Cas12p and RNA reporters at 37°C for 30 minutes on lyophilized formats of patient samples and negative control samples (n=16) are shown.
[0775] Example 13: Detection of specific cleavage activity
[0776] Figure 25 : It was investigated whether Cas12a.1 and Cas12p of the present disclosure, when complexed with their guide, are able to cleave dsDNA. In these examples, the target was a Hantavirus dsDNA sequence (100 pb) cloned into the commercial pGEM®-T Easy vector from Promega (cat. A1360). Negative controls included empty pGEM®-T Easy vector. Hantavirus dsDNA sequence (100 pb) in pGEM®-T Easy vector. Negative controls included empty pGEM®-T Easy vector. Hantavirus dsDNA sequence (100 pb) in pGEM®-T Easy vector. Negative controls included empty pGEM®-T Easy vector. Hantavirus dsDNA sequence (100 pb) in pGEM®-T Easy vector. Negative controls included empty pGEM®-T Easy vector. TM 2.1 (cat. B7202S) for 15 minutes at room temperature, 100 nM of Cas12a.1 or Cas12p were complexed with 100 nM of sgRNA to target the Hantavirus sequence. Controls with Cas enzymes that were not complexed with their guide were included. Then, 5 ng / uL of target was added in a final reaction volume of 20 uL. Reactions were incubated at 37°C or 25°C for 0, 30, 60 or 90 minutes and ended by the addition of 50 mM EDTA. Then, samples were centrifuged at 12000 g for 10 minutes and mixed with 6X gel loading dye from NEB (cat. B7024S). Samples were analyzed in 0.8% TBE agarose gels. The molecular weight of the material was assessed using the Fast DNA Ladder from NEB (cat. N3238S). After electrophoresis, gels were stained for 30 minutes with a fresh solution of SYBR® Gold nucleic acid gel stain from Invitrogen (cat. S11494) and imaged in a VersaDoc 3000 imager from Bio-Rad. TM Gold nucleic acid gel stain (cat. S11494) and imaged in a VersaDoc 3000 imager from Bio-Rad.TM Imaged on a ChemiDoc™ Touch Imaging System (Bio-rad). Figure 1 shows the results of the assay. Cas12a.1 can linearize all plasmids after 90 minutes at 37°C, while Cas12p achieves comparable results in only 60 minutes.
[0777] Figure 26 : It was investigated whether Cas12a.1 and Cas12p of the present disclosure, when complexed with their guide, can cleave ssDNA. In these examples, the target includes a custom ssDNA fluorescently labeled sequence (3’FAM-ssDNA) from IDT that is 70 nucleotides in length that targets Hantavirus (5’-TCA TTT AGA AAG TAG ATA TTG ATT GAT TTT AGC GAA AGC CAA TTT TTG AGC TGC CAC TGA TGT AAA AGT T-3’-6-FAM; SEQ ID NO: 124). The negative control includes a custom antisense ssDNA sequence (ASssDNA) from IDT that is 120 nucleotides in length that also targets Hantavirus (5’-GCT ATC TTA ATC CTT AAT CTA TCC TCA AAC GTT CTA TTA ATG GCC GTG TCA ATC AAT ATC TAC TTT CTA AAT GAA ACT TTT ACA TCA GTG GCA GCT CAA AAA TTG GCT TTC GCT AAA ATC-3’; SEQ ID NO: 125). The procedure was as follows: 10 pmol of Cas12a.1 or Cas12p was complexed with 10 pmol of sgRNA to target the Hantavirus sequence in NEBuffer 2.1 (Cat# B7202S) for 15 minutes at room temperature. Controls included Cas enzymes that were not complexed with their guide. Then, 10 pmol of 3’FAM-ssDNA or optional ASssDNA was added in a final reaction volume of 10 uL. The reactions were incubated at 37°C for 0, 0.5, 1, or 5 minutes and the reactions were stopped by adding 2X Novex TBE-Urea sample buffer (Cat# LC6876) from Invitrogen followed by heating at 95°C for 3 minutes. The samples were centrifuged at 12000g for 10 minutes and analyzed on a 15% Mini-PROTEAN® TGX™ Precast TM 2.1 (Cat# B7202S) for 15 minutes at room temperature. Controls included Cas enzymes that were not complexed with their guide. Then, 10 pmol of 3’FAM-ssDNA or optional ASssDNA was added in a final reaction volume of 10 uL. The reactions were incubated at 37°C for 0, 0.5, 1, or 5 minutes and the reactions were stopped by adding 2X Novex TM TBE-Urea sample buffer (Cat# LC6876) from Invitrogen followed by heating at 95°C for 3 minutes. The samples were centrifuged at 12000g for 10 minutes and analyzed on a 15% Mini-PROTEAN® TGX™ Precast TBE-Urea gel (Cat# 4566056) from Bio-Rad. The molecular weight of the material was assessed using oligonucleotide length standards (Cat# 51-05-15-02) from IDT. The samples were imaged on a ChemiDoc™ Touch Imaging System (Bio-rad). Figure 1 shows the results of the assay. Cas12a.1 can linearize all plasmids after 90 minutes at 37°C, while Cas12p achieves comparable results in only 60 minutes. TMgels were first imaged on a ChemiDoc™ Touch Imaging System (Bio-Rad) and then stained with a fresh solution of SYBR® Gold nucleic acid gel stain (Cat# S11494) from Invitrogen for 30 minutes to visualize the non-fluorescently labeled sequence of the AS ssDNA and the non-fluorescently labeled ladder. Figure 2 shows the results of the assay. Cas12a.1 and Cas12p show specific ssDNA cleavage of the 3’FAM-ssDNA substrate (S) to produce a ~40 nucleotide length product (P). Neither of these two Cas enzymes can cleave the AS ssDNA sequence (NTC). Reactions occur in the time range of seconds to minutes. TM Gold nucleic acid gel stain (Cat# S11494) from Invitrogen for 30 minutes to visualize the non-fluorescently labeled sequence of the AS ssDNA and the non-fluorescently labeled ladder. Figure 2 shows the results of the assay. Cas12a.1 and Cas12p show specific ssDNA cleavage of the 3’FAM-ssDNA substrate (S) to produce a ~40 nucleotide length product (P). Neither of these two Cas enzymes can cleave the AS ssDNA sequence (NTC). Reactions occur in the time range of seconds to minutes.
[0778] Figure 27 : It was investigated whether Cas12a.1 and Cas12p of the present disclosure, when complexed with their guide, are able to cleave ssRNA. In these examples, the targets include ssRNA sequences targeting Hantavirus obtained by in vitro transcription (IVT). The negative control includes a 65 nucleotide length custom non-target ssRNA sequence from IDT (5’-TAA GCG CCC TTG C GC TTT CCC CAG CCT TCG GGT TGG TTG CCT TTT AGT GCA AGG GCG CGA TTA TT-3’; SEQ ID NO: 126). The positive control includes a 120 nucleotide length custom ssDNA sequence from IDT targeting Hantavirus (5’-GAT TTT AGC GAA AGC CAA TTT TTG AGC TGC CAC TGA TGT AAA AGT TTC ATT TAG AAA GTA GAT ATT GAT TG ACAC GGC CAT TAA TAG AAC GTT TGA GGA TAG ATT AAG GAT TAA GAT AGC-3’; SEQ ID NO: 127). The procedure was as follows: 150 nM of Cas12a.1 or Cas12p was complexed with 150 nM of sgRNA to target the Hantavirus sequence in commercial NEBuffer TM 2.1 (Cat# B7202S) for 15 minutes at room temperature. Controls with Cas enzymes that were not complexed with their guide were included. Then, 5 ng / uL of ssRNA or optionally non-target ssRNA or ssDNA was added in a final reaction volume of 10 uL. Reactions were incubated at 37°C for 0, 1, or 3 hours and visualized by adding 2X Novex® Pre-Cast TGX® Stain Reagent from Invitrogen. TMTBE-Urea sample buffer (catalog number LC6876) was then added, followed by heating at 65°C for 3 minutes to terminate the reaction. The sample was centrifuged at 12000g for 10 minutes and distilled into a 15% buffer solution from Bio-Rad. Analysis was performed on a TBE-Urea gel (catalog number 4566056). Molecular weights were assessed using a low-range ssRNA ladder from NEB (catalog number N0364S). The gel was analyzed using SYBR from Invitrogen. TM Stain with fresh solution of gold nucleic acid gel dye (catalog number S11494) for 30 minutes and then in VersaDoc. TM Imaging was performed on a Bio-rad imaging system. Figure 3 shows the results of the assay. Neither Cas12a.1 nor Cas12p demonstrated specific ssRNA cleavage activity.
[0779] Example 14: Characterization of Cas12p non-specific nuclease activity
[0780] MALDI-TOF MS Experiment Description: Matrix-assisted laser desorption / ionization time-of-flight mass spectrometry (MALDI-TOF MS) was used to monitor the products generated by the non-specific nuclease activity of the Cas12p enzyme. Protected DNA (C*C*C*C*C*C*C*C*C*C*C*C*C*C*C*C*C*C*TTATT; SEQ ID NO:128) and RNA (rC*rC*rC*rC*rC*rC*rC*rC*rC*rC*rC*rC*rC*rC*rC*rUrUrArUrU; SEQ ID NO:129) reporters were used to ensure minimum length and minimize the number of possible hydrolysis products. The symbol (*) on the C and rC bases indicates the presence of a phosphate thioester bond resistant to nuclease degradation. When performing CRISPR reactions using the corresponding reporter, the final concentration of the complex is 75 nM Cas12p:75 nM sgRNA:20 nM activator:2.5 μM DNA reporter or 75 nM Cas12p:75 nM sgRNA:10 nM activator:1.25 μM RNA reporter in a solution containing 1X binding buffer (50 mM NaCl, 10 mM Tris-HCl, 10 mM MgCl2, 1 mM DTT, 100 g / ml BSA, pH 7.9). The reaction is incubated at 25°C for 1 hour for DNA reporters and at 37°C for 6 hours for RNA reporters (T1 of reaction). Figure 28 and Figure 30 By heating the CRISPR reaction before adding the reporter, the reaction time was reduced to zero (T0, T0). Figure 29 and Figure 31) as a negative control. Reactions were purified and analyzed on a PerSeptive Biosystems (ABI)- Voyager-DE RP-MALDI-TOF mass spectrometer at Stanford University. For each reaction, a list was generated containing the predicted m / z (mass to charge ratio) of all possible DNA / RNA cleavage products and all expected overhangs as outlined by Joyner et al. 2012. The observed m / z was correlated to the above list by using a Perl script. The relative intensity of the peaks was calculated relative to the main signal of the reaction. DNA reporter hydrolysis yields unique cleavage products Figure 28 ), while RNA reporter hydrolysis yields multiple fragments including 3’ hydroxyl and phosphate ends Figure 30 ). In all cases, the main hydrolysis species is the hydrolysis species containing two nucleotides after the protected sequence and the 3’ hydroxyl end.
[0781] Figure 28 to Figure 29 Mass spectrometry data of Cas12p reactions using DNA oligonucleotides as reporters are shown. Figure 30 to Figure 31 Mass spectrometry data of Cas12p reactions using RNA oligonucleotides as reporters are shown.
[0782] Example 15: Characterization of chimeric (hybrid) guide use
[0783] Guide sequences composed partially of DNA and RNA nucleotides (hybrid guides, chimeric guide) were tested and determined that they can support efficient collateral Cas12p activity. Replacing the 3` end portion of the sgRNA with DNA nucleotides (hybrid 4 DNA; 5’ AGAUUUCUACUUUUGUAGAUGUGGCAGCUCAAAAAU(TGGC) 3’; SEQ ID NO: 130) or replacing both the 5` and 3` ends with DNA nucleotides (hybrid 3 / 4 DNA; 5’ (AGA)UUUCUACUUUUGUAGAUGUGGCAGCUCAAAAAU(TGGC) 3’; SEQ ID NO: 131) retained their activity compared to the unmodified guide sequence (sgRNA; 5’ AGAUUUCUACUUUUGUAGAUGUGGCAGCUCAAAAAUUGGC 3’; SEQ ID NO: 132). Partial replacement of 8 DNA nucleotides at the 3` end resulted in complete loss of Cas12p collateral activity (hybrid 8 DNA; 5’ AGAUUUCUACUUUUGUAGAUGUGGCAGCUCAA(AAATTGGC) 3’; SEQ ID NO: 133).
[0784] Cas12p were pre-incubated with their respective sgRNA or hybrid guide (1 uM complex). Reactions were initiated by dilution of Cas12p complex to a final concentration of 37.5 nM Cas12p: 37.5 nM sgRNA: 10 nM activator in a solution containing 1 x binding buffer (50 mM NaCl, 10 mM Tris-HCl, 10 mM MgCl2, 1 mM DTT, 100 g / ml BSA, pH 7.9) and 600 nM TTATTATT ssDNA FQ reporter (SEQ ID NO: 121) substrate in a 40 μΐ reaction. Reactions (40 μΐ, 384-well microplate format) were incubated at 25 °C for 40 min with fluorescence measurements every 1 min (ssDNA FQ substrate = λex: 485 nm; λem: 538 nm) in a fluorescence plate reader (Synergy HTX, BioTek). Results show quantification of maximum fluorescence signal produced after 30 min. Non-template negative control (NTC) fluorescence values were calculated from reactions performed in the absence of target plasmid. Error bars represent mean ± s.d. with n = 3 replicates. M2). Non-template negative control (NTC) fluorescence values were calculated from reactions performed in the absence of target plasmid. Error bars represent mean ± s.d. with n = 3 replicates.
[0785] Figure 32 It is shown that effective collateral cleavage Cas12p activity can be achieved using DNA-RNA chimeric guides.
[0786] Example 16: Characterization of the collateral cleavage activity
[0787] Figure 33 It is shown that the following substrates were used to show agarose gels of Cas12a.1 and Cas12p protein / guide complex activity: (A) M13mp18 single-stranded DNA (Cat. No. N4040S, NEB); and (B) M13mp18 RF I double-stranded DNA (Cat. No. N4018S, NEB). Cas12a.1 and Cas12p exhibit collateral cleavage activity and cleave ssDNA circular DNA Figure 33 , panel A), but not dsDNA circular DNA Figure 33, panel B). Reactions were initiated at 25 °C by diluting Cas12p / guide or Cas12a.1 / guide complexes to a final concentration of 37.5 nM Cas12p: 37.5 nM sgRNA: 10 nM activator or 75 nM Cas12a.1: 75 nM sgRNA: 10 nM activator in a solution containing 1x binding buffer (50 mM NaCl, 10 mM Tris-HCl, 10 mM MgCl2, 1 mM DTT, 100 g / ml BSA, pH 7.9) and 1 uL of M13mp18 single-stranded DNA (cat# N4040S, NEB) and M13mp18 RF I double-stranded DNA (cat# N4018S, NEB). Controls containing no Cas enzyme, guide, or activator were included, and non-flanking was observed.
[0788] Cutting efficiency Cas12p showed similar cutting efficiency for homopolymeric reporters of T, A, or C (7 nt in length), while Cas12a.1 showed higher cutting efficiency for poly-C, but also cut poly-A and poly-T sequences. Cas12p showed cleavage for homopolymeric reporters of T, A, or C at 25 °C, evidenced by an increase in fluorescence, while Cas12a.1 showed cleavage reactions for the 5’ 6-FAM-TTATTATT-3’ IABkFQ 3’ reporter sequence (SEQ ID NO: 121) only at 37 °C.
[0789] Reactions were initiated at 25 °C by diluting Cas12p or Cas12a.1 complexes to a final concentration of 37.5 nM Cas12p: 37.5 nM sgRNA: 10 nM activator or 75 nM Cas12a.1: 75 nM sgRNA: 10 nM activator in a solution containing 1x binding buffer (50 mM NaCl, 10 mM Tris-HCl, 10 mM MgCl2, 1 mM DTT, 100 g / ml BSA, pH 7.9) and 600 nM ssDNA FQ reporter substrate (5’ 6-FAM-TTATTATT-3’ IABkFQ 3’ (SEQ ID NO: 121), 5’ 6-FAM-AAAAAAA-3’ IABkFQ 3’, 5’ 6-FAM-TTTTTTT-3’ IABkFQ 3’, 5’ 6-FAM-CCCCCCC-3’ IABkFQ 3’, or 5’ 6-FAM-C*GGGC*GGG-3’ IABkFQ 3’, from IDT (Integrated DNA Technologies, Inc)). Reactions (40 μΐ, 384-well microplate format) were read on a fluorescent plate reader (EnVision Multilabel 2103, PerkinElmer) at 25 °C for 1 hour. M2) incubated at 25°C or 37°C with fluorescence measurements every 1 min (ssDNA FQ substrate = λex: 485 nm; λem: 538 nm). Background corrected fluorescence values were calculated by subtracting the fluorescence values obtained from reactions performed in the absence of target plasmid. Error bars represent mean ± s.d. with n = 3 replicates. Figure 34 The differential efficiency of homopolymeric reporter cleavage at 25°C and 37°C is shown. The results show that Cas12p cleaves polyT, polyA and polyC, while Cas12a.1 shows a preference for polyC cleavage.
[0790] The specificity of the trans-cleavage activity (side-cutting activity) was tested using a custom ssRNA 5’6-FAM rArUrArUrArUrA-3 IABkFQ 3’ from IDT (Integrated DNA Technologies, Inc) and RNaseAlert™ (commercial RNA reporter). The results show that Cas12p is able to cleave the RNA reporter used, but Cas12a.1 is not. The assay was performed in 40 μΐ reactions using Cas12p or Cas12a.1 complexes at 37°C, with a final concentration of 37.5 nM Cas12p: 37.5 nM sgRNA: 10 nM activator or 75 nM Cas12a.1 : 75 nM sgRNA: 10 nM activator in a solution containing 1 x binding buffer (50 mM NaCl, 10 mM Tris-HCl, 10 mM MgCl2, 1 mM DTT, 100 g / ml BSA, pH 7.9) and 600 nM of the RNA FAM Q reporter substrate (ssRNA 5’6-FAM rArUrArUrArUrA-3 IABkFQ 3 and RNaseAlert (Cat N 11-04-03-03-IDT)). Reactions were read in a fluorescence plate reader (Tecan Spark® 10M) every 1 min for 60 min at 37°C. The background corrected fluorescence values were calculated by subtracting the fluorescence values obtained from reactions performed in the absence of target plasmid. Error bars represent mean ± s.d. with n = 3 replicates. M2) incubated at 25°C or 37°C with fluorescence measurements every 1 min (ssDNA FQ substrate = λex: 485 nm; λem: 538 nm). Background corrected fluorescence values were calculated by subtracting the fluorescence values obtained from reactions performed in the absence of target plasmid. Error bars represent mean ± s.d. with n = 3 replicates. Figure 35 The results of these data are shown and demonstrate the side-cutting ability of Cas12p but not Cas12a.1 to cleave the RNA reporter.
[0791] The kinetics of the collateral (trans-cleavage) activity of Cas12p were assessed using both DNA and RNA reporters. Experiments with RNA substrates showed that the rate of cleavage of ssRNA was only 3-fold slower than ssDNA reporters. Cas12a.1 cleaved ssRNA substrates at least 104-fold slower than ssDNA, confirming that ssDNA is the substrate of choice for Cas12a.1 collateral cleavage. Assay detection was performed in 40 μΐ reactions using Cas12p or Cas12a.1 complexes at a final concentration of 37.5 nM Cas12p: 37.5 nM sgRNA: 10 nM activator or 75 nM Cas12a.1 : 75 nM sgRNA: 10 nM activator in a solution containing 1 x binding buffer (50 mM NaCl, 10 mM Tris-HCl, 10 mM MgCl2, 1 mM DTT, 100 g / ml BSA, pH 7.9) and 600 nM of either ssDNA FAM Q reporter substrate (ssDNA 5’ 6-FAM TTATTATT-3’ IABkFQ 3 (SEQ ID NO: 121)) or RNaseAlert (Cat N 11-04-03-03-IDT)). Reactions were incubated in a fluorescence plate reader (Tecan Spark®M1000) at 37 °C for up to 40 minutes, with a fluorescence measurement taken every 1 minute (λex: 535 nm; λem: 595 nm). Background corrected fluorescence values were calculated by subtracting the fluorescence values obtained from reactions performed without target plasmid. The resulting data were fitted to a single exponential decay curve (GraphPad software) according to the following equation: fraction cleaved = A x (1 - exp(-k x t)), where A is the amplitude of the curve, k is the first order rate constant, and t is time. Error bars represent mean ± s.d. with n = 3 replicates. M2) at 37 °C for up to 40 minutes, with a fluorescence measurement taken every 1 minute (λex: 535 nm; λem: 595 nm). Background corrected fluorescence values were calculated by subtracting the fluorescence values obtained from reactions performed without target plasmid. The resulting data were fitted to a single exponential decay curve (GraphPad software) according to the following equation: fraction cleaved = A x (1 - exp(-k x t)), where A is the amplitude of the curve, k is the first order rate constant, and t is time. Error bars represent mean ± s.d. with n = 3 replicates. Figure 36 The results of these data are shown and demonstrate the kinetics of collateral cleavage activity of Cas12p and Cas12a.1 using DNA and RNA as reporters.
[0792] The reporter composed of DNA and RNA nucleotides resulted in efficient collateral Cas12p and Cas12a.1 activity. In comparison to ssDNA or RNA reporters (ssDNA FAMQ reporter substrate (ssDNA 5’6-FAM TTATTATT-3IABkFQ 3 (SEQ ID NO: 121)) or RNaseAlert (Cat N 11-04-03-03-IDT)), the FQ hybrid / 56-FAM / TT rArUrU ATT / 3IABkFQ / or / 56-FAM / TT ATrU rArUrU / 3IABkFQ / resulted in maintained Cas12p collateral activity. While Cas12a.1 showed a slight decrease in trans-cleavage efficiency of the chimeric reporter compared to ssDNA. In 40 μΐ reactions, the reaction was initiated by diluting Cas12p or Cas12a.1 complexes to a final concentration of 37.5 nM Cas12p:37.5 nM sgRNA:10 nM activator or 75 nM Cas12a.1 :75 nM sgRNA:10 nM activator in a solution containing 1 x binding buffer (50 mM NaCl, 10 mM Tris-HCl, 10 mM MgCl2, 1 mM DTT, 100 g / ml BSA, pH 7.9) and 600 nM of ssDNA FAMQ reporter substrate (ssDNA 5’6-FAM TTATTATT-3IABkFQ 3 (SEQ ID NO: 121), DNA-RNA chimeric reporter ( / 56-FAM / TT rArUrU ATT / 3IABkFQ / , / 56-FAM / TT ATrU rArUrU / 3IABkFQ / or RNaseAlert (Cat N 11-04-03-03-IDT)). Reactions were incubated in a fluorescent plate reader (Tecan Spark® 10M) for up to 40 minutes with a fluorescence measurement every 1 minute (λex: 535 nm; λem: 595 nm). Background corrected fluorescence values were calculated by subtracting the fluorescence values obtained from reactions performed without target plasmid. Error bars represent mean ± s.d. with n = 3 replicates. M2) for up to 40 minutes with a fluorescence measurement every 1 minute (λex: 535 nm; λem: 595 nm). Background corrected fluorescence values were calculated by subtracting the fluorescence values obtained from reactions performed without target plasmid. Error bars represent mean ± s.d. with n = 3 replicates. Figure 37 Results showing these data are presented.
[0793] While the application has been described in connection with specific embodiments thereof, the foregoing description is intended to illustrate and not limit the scope of the application as defined by the appended claims. Other aspects, advantages, and modifications are within the scope of the following claims.
[0794] Example 17: Cas12a.1 and Cas12p mature guide characterization and validation
[0795] The scaffold sequence of the mature guide was computationally inferred from the corresponding CRISPR locus.Figure 38 The secondary structures of the mature guide scaffolds of Cas12a.1 (5'aaauuucuacuguaguagau 3') (SEQ ID NO:116; Sub-figure A) and Cas12p (5'agauuucuacuuuuguagau 3') (SEQ ID NO:117; Sub-figure B) are shown. These are verified below.
[0796] Mature guide scaffolds for Cas12a.1 and Cas12p were evaluated in vitro. These mature scaffold sequences, along with spacers targeting the N gene from the SARS-CoV-2 virus, were used in this embodiment. In a 40 μl reaction, the reaction was initiated by diluting the Cas12p or Cas12a.1 complex in a solution containing 1× binding buffer (50 mM NaCl, 10 mM Tris-HCl, 10 mM MgCl2, 1 mM MTT, 100 g / ml BSA, pH 7.9) and 600 nM of ssDNA FAMQ reporter substrate (ssDNA 5'6-FAM TTATTATT-3IABkFQ3 (SEQ ID NO: 121)) to a final concentration of 37.5 nM Cas12p:37.5 nM sgRNA:10 nM activator or 75 nM Cas12a.1:75 nM sgRNA:10 nM activator. The reaction was then run on a fluorescent plate reader (…). Incubation in M2 for up to 20 minutes, with fluorescence measurements taken every minute (λex: 535 nm; λem: 595 nm). Background-corrected fluorescence values were calculated by subtracting the fluorescence values obtained from reactions performed without the target plasmid. Error bars represent the mean ± sd, where n = 3 replicates ( Figure 39 The data in this figure indicate that these mature scaffold sequences provide CRISPR-mediated detection of SARS-CoV-2. sequence list <110> CASPR Biotechnology Corporation National Science and Technology Research Council (CONSEJO NACIONAL DE INVESTIGACIONESCIENTIFICAS Y TECNICAS) (CONICET) <120> Two new classes of CRISPR-CAS RNA-guided endonucleases <130> CABI-002 / 02WO 337081-2006 <150> US 63 / 058,448 <151> 29-JUL-20 <150> US 62 / 898,340 <151> 10-SEP-19 <160> 249 <170> PatentIn version 3.5 <210> 1 <211> 1038 <212> PRT <213> Unknown <220> <223> Cas9.1 <400> 1 Met Gin Arg lie Phe Gly Leu Asp lie Gly Thr Thr Ser lie Gly Phe 1 5 10 15 Ala Val lie Asp His Asp Arg Asp Gin Gly Val Gly Arg lie His Arg 20 25 30 Leu Gly Ala Arg lie Phe Pro Glu Ala Arg Asp Glu Lys Gly Thr Pro 35 40 45 Leu Asn Gin His Arg Arg Gin Lys Arg Leu Ala Arg Arg Gin Leu Arg 50 55 60 Arg Arg Arg Leu Arg Arg Lys Ala Leu Asn Glu Leu Leu Ser Ala Arg 65 70 75 80 Gly Met Leu Pro Arg Phe Gly Thr Ser Ala Trp His Asp Ala Met Ala 85 90 95 Leu Asp Pro Tyr Ala Leu Arg Ala Arg Gly Thr Glu Glu Ala Leu Gin 100 105 110 Pro Val Glu Val Gly Arg Ala Leu Tyr His Leu Ala Gln Arg Arg His 115 120 125 Phe Lys Pro Arg Asp Glu Ala Ala Glu Ala Asp Glu Gln Glu Val Gly 130 135 140 Asp Gln Glu Ala Glu Thr Lys Arg Glu Lys Leu Leu Gln Ala Leu Arg 145 150 155 160 Arg Ser Gly Arg Thr Leu Gly Gln Glu Leu Ala Ala Arg Gly Pro His 165 170 175 Glu Arg Lys Arg His Glu His Ala Leu Arg Ser Thr Val Glu Thr Glu 180 185 190 Phe Glu Arg Leu Leu Thr Ala Gln Ala Arg His His Glu Ile Leu Arg 195 200 205 Asp Pro Glu Phe Val Glu Glu Leu Arg Glu Thr Ile Phe Ala Gln Arg 210 215 220 Pro Val Phe Trp Arg Thr Ser Thr Leu Gly Thr Cys Pro Phe Val Pro 225 230 235 240 Gly Ala Pro Leu Cys Pro Lys Gly Ala Trp Leu Ser Arg Gln Arg Arg 245 250 255 Met Leu Glu Gln Val Asn Asn Leu Ala Ile Thr Gly Gly Asn Ala Arg 260 265 270 Pro Leu Asp His Glu Glu Arg Arg Ala Ile Leu Ala Val Leu Gln Thr 275 280 285 Gln Ala Ser Met Ser Trp Gly Ala Val Arg Thr Ala Leu Lys Pro Leu 290 295 300 Phe Lys Ala Arg Gly Glu Ala Gly Ala Glu Arg Arg Leu Arg Phe Asn 305 310 315 320 Leu Glu Glu Gly Gly Gly Lys Thr Leu Leu Gly Asn Pro Leu Glu Ala 325 330 335 Lys Leu Ala Arg Ile Phe Gly Glu Ala Trp Ala Thr His Pro His Arg 340 345 350 Asp Ala Ile Arg Glu Thr Ile His Asp Arg Leu Phe Ala Ala Thr Tyr 355 360 365 Asn Ala Lys Gly Ala Gln Arg Ile Val Ile Leu Pro Ala Ser Gln Arg 370 375 380 Ala Glu Arg Met Arg Gly Val Ile Ala Gly Leu Gln Ala Asp Phe Gly 385 390 395 400 Leu Ser His Glu Gln Ala Met Ala Leu Ala Glu Leu Pro Leu Thr Pro 405 410 415 Gly Trp Glu Pro Tyr Ser Ser Glu Ala Leu Arg Ala Leu Met Pro Lys 420 425 430 Leu Glu Glu Gly Val Arg Phe Gly Ala Leu Val Val Ala Pro Glu Trp 435 440 445 Glu Asp Trp Arg Glu Ala Thr Phe Pro Gln Arg Glu Arg Pro Thr Gly 450 455 460 Glu Val Leu Asp Leu Leu Pro Ser Pro Lys Cys His Asp Glu Ser Arg 465 470 475 480 Arg Gln Thr Arg Leu Arg Asn Pro Thr Val Leu Arg Thr Gln Asn Glu 485 490 495 Leu Arg Lys Val Val Asn Asn Leu Ile Arg Ala His Gly Lys Pro Asp 500 505 510 Ile Ile Arg Val Glu Val Ala Arg Glu Val Gly Leu Ser Lys Arg Glu 515 520 525 Arg Glu Asp Arg Tyr Asn Gly Met Arg Arg Gln Glu Arg Gln Arg Gln 530 535 540 Ala Ala Ile Lys Asp Leu Gln Ala Lys Gly Phe Ala Glu Pro Ser Arg 545 550 555 560 Ala Asp Val Glu Lys Trp Leu Leu Trp Lys Glu Ser Lys Glu Thr Cys 565 570 575 Pro Tyr Thr Gly Asp Lys lie Cys Phe Asp Ala Leu Phe Arg Arg Gly 580 585 590 Glu Phe Gin Val Glu His lie Trp Pro Arg Ser Arg Ser Phe Asp Asp 595 600 605 Ser Phe Arg Asn Lys Thr Leu Cys Arg Arg Asp Val Asn Leu Ala Lys 610 615 620 Gly Asn Gin Thr Pro Phe Glu Phe Phe Glu Ser Arg Pro Glu Glu Trp 625 630 635 640 Glu Ala Val Lys Arg Arg Leu Asp Gly Leu Gin Ala Lys Arg Ala Gly 645 650 655 Gly Glu Gly Met Ala Arg Gly Lys Val Lys Arg Phe Val Ala Ser Thr 660 665 670 Leu Pro Asp Asp Phe Ala Gin Arg Gin Leu Asn Asp Thr Gly Trp Ala 675 680 685 Ala Arg Glu Ala Val Ala Phe Leu Lys Arg Leu Trp Pro Asp Glu Gly 690 695 700 Gln Ala Ala Pro Val Arg Val Gin Ala Val Thr Gly Arg Val Thr Ala 705 710 715 720 Gln Leu Arg His Leu Gly Gly Leu Asp Gly Val Leu Ser Asp Gly Ala 725 730 735 Arg Lys Thr Arg Asp Asp His Arg His His Ala Val Asp Ala Leu Val 740 745 750 Val Ala Cys Thr His Pro Gly Met Thr Glu Arg Leu Ser Arg Tyr Trp 755 760 765 Gln Gln Lys Glu Asp Glu Arg Ala Glu Arg Pro Gln Leu Asp Pro Pro 770 775 780 Trp Pro Thr Ile Arg Ala Asp Ala Glu Ala Ala Lys Asp Leu Ile Val 785 790 795 800 Val Ser His Arg Val Arg Lys Lys Ile Ser Gly Pro Phe His Lys Glu 805 810 815 Thr Val Tyr Gly Ala Thr Asp Glu Arg Glu Val Thr Arg Gly Leu Glu 820 825 830 Tyr Glu Lys Phe Val Thr Arg Lys Arg Val Glu Asp Leu Thr Lys Ser 835 840 845 Met Leu Ala Asp Ile Arg Asp Asp Arg Val Arg Gln Ile Val Thr Ala 850 855 860 Trp Val Ala Glu Arg Gly Gly Asp Pro Lys Lys Ala Phe Pro Pro Tyr 865 870 875 880 Pro Thr Leu Gly Ser Ser Gly Pro Glu Ile Arg Lys Val Arg Val Leu 885 890 895 Ile Arg Arg Gin Pro Thr Leu Met Ala Arg Ala Ala Thr Gly Phe Ala 900 905 910 Asp Leu Gly Ala Asn His His Val Ala Ile Tyr Lys Thr Ala Asp Glu 915 920 925 Arg Phe Ala Phe Glu Val Val Ser Leu Leu Glu Val Ala Arg Arg Val 930 935 940 Asp Arg Gly Glu Pro Pro Val Lys Arg Gin Arg Gly Asp Glu Lys Leu 945 950 955 960 Val Met Ser Leu Ala Gin Gly Asp Leu Ile Arg Phe Ala Lys Thr Pro 965 970 975 Asp Ala Glu Ala Ala Ile Trp Arg Val Gin Lys Ile Ala Thr Lys Gly 980 985 990 Gln Ile Ser Leu Leu His His Asp Asp Ala Ser Pro Lys Glu Pro Ser 995 1000 1005 Leu Phe Glu Pro Met Val Gly Gly Leu Met Ala Arg Asn Pro Glu 1010 1015 1020 Lys Leu Ala Val Asp Pro Ile Gly Arg Val Arg Lys Ala Gly Asp 1025 1030 1035 <210> 2 <211> 1375 <212> PRT <213> Unknown <220> <223> Cas9.2 <400> 2 Met Lys Lys Glu Lys Val Tyr Met Gly Leu Asp Leu Gly Thr Asn Ser 1 5 10 15 Val Gly Trp Ala Val Thr Asp Asn Asp Tyr Lys Val Leu Lys Phe Lys 20 25 30 Arg Arg Ala Met Trp Gly Val Arg Leu Phe Asn Glu Ala Asn Pro Ala 35 40 45 Val Glu Arg Arg Val Ala Arg Ser Asn Arg Arg Arg Leu Ala Arg Lys 50 55 60 Lys Gln Arg Val Ala Trp Leu Lys Glu Ile Phe Lys Asn Ser Ile Ser 65 70 75 80 Glu Ile Asp Pro Glu Phe Phe Asp Arg Leu Glu Gin Ser Ala Leu Trp 85 90 95 Ala Glu Asp Lys Asn Val Ala Gly Lys Tyr Ser Leu Phe Asn Glu Lys 100 105 110 Lys Leu Thr Asp Lys Thr Phe Tyr Arg Lys Phe Pro Thr Val Phe His 115 120 125 Leu Lys Lys Ala Leu Met Asp Gly Lys Ile Lys Lys Pro Asp Ile Arg 130 135 140 Phe Val Tyr Leu Ala Leu Ser His Tyr Leu Gin Asn Arg Gly His Phe 145 150 155 160 Leu Leu Glu Asn Glu Leu Asn Ser Val Glu Asp lie Asp lie Arg Asp 165 170 175 Ile Phe Asn Ser Leu Asn Glu Arg lie His Val Leu lie Asp Ser Gly 180 185 190 Asp Asp Met Val Pro Ala Phe Asp Leu Thr Asn Leu Asp Asp Leu Lys 195 200 205 Gln lie Ala Thr Asp Thr Asn lie Ser Gly Lys Thr Gin Glu Lys Glu 210 215 220 Ala Phe lie Lys Thr Leu Leu Asn Gly Ala Lys Gin Pro Ala Leu Glu 225 230 235 240 Ala lie lie Lys Leu Cys Thr Gly Gly Ser Ala Asn Leu Ser Lys lie 245 250 255 Phe Gly Asp Met Phe Glu Phe Glu Ser Glu lie Lys Ser lie Ser Phe 260 265 270 Glu Lys Ala Asn Phe Glu Asp Glu lie Ala Pro Lys Leu Gin Asp Cys 275 280 285 Leu Gly Asp Tyr Tyr Gin lie lie Glu Leu Ala Gin Gin lie Tyr Ser 290 295 300 Trp Tyr Thr Leu Tyr Lys Val Cys Ser Gly Arg Pro Ser Val Ser His 305 310 315 320 Ala Lys Val Glu Asp Tyr Glu Lys His Lys Glu Gin Leu Ser His Leu 325 330 335 Lys Val Leu Val Arg Lys His Phe Ser Lys Asn Val Tyr Arg Glu Ile 340 345 350 Phe Arg Lys Glu Asp Asp Lys Ile His Asn Tyr Val Ser Tyr Ile Ser 355 360 365 Gly Lys Lys Asp Arg Asp Glu Phe Tyr Lys Tyr Leu Lys Lys Thr Leu 370 375 380 Glu Lys Lys Ser Thr Phe Lys Lys Thr Ser Glu Phe Glu Asn Ile Ser 385 390 395 400 Arg Ala Ile Glu Gin Gin Asn Tyr Leu Pro Lys Gin Arg Val Lys Asp 405 410 415 Asn Ser Val Val Pro Gin Gin Leu Tyr Lys Gin Glu Ile Val Lys Ile 420 425 430 Leu Asn Asn Leu Ser Ser His Tyr Pro Phe Leu Ser Gin Lys Thr Asp 435 440 445 Gly Ile Ser Asn Arg Glu Lys Ile Ile Lys Ile Phe Glu Tyr Arg Ile 450 455 460 Pro Tyr Tyr Val Gly Pro Leu Cys Asp Ile His Arg Ala Gly Asp Asp 465 470 475 480 Gly Phe Ser Trp Leu Val Arg Asp Cys Ser Lys Lys Ile Thr Pro Trp 485 490 495 Asn Phe Glu Gln Val Val Asp Ile Pro Gln Ser Ala Glu Asn Phe Ile 500 505 510 Lys Asn Met Thr Arg Lys Cys Thr Tyr Leu Lys Gln Tyr Asn Val Leu 515 520 525 Pro Lys Asn Ser Leu Leu Tyr Ser Glu Tyr Ser Val Leu Asn Glu Leu 530 535 540 Asn Asn Val Arg Ile Lys Thr Lys Lys Leu Thr Pro Lys Leu Lys Glu 545 550 555 560 Lys Met Leu Asn Thr Leu Phe Arg Gln Lys Lys Asn Ile Ser Ile Thr 565 570 575 Ser Leu Ile His Trp Leu Val Ser Glu Gly Val Tyr Glu Lys Gly Glu 580 585 590 Ile Glu Lys Ser Asp Val Ser Gly Val Asp Ser Asn Phe Thr Ser Ser 595 600 605 Leu Ser Ala Ala Ile Ser Phe Asp Arg Ile Ile Gly Glu Lys Met Lys 610 615 620 Asn Lys Lys Thr Gin Lys Met Val Glu Glu He He Asn Trp Leu Ala 625 630 635 640 Leu Phe Ser Asp Lys Lys He Leu Gin Gin Lys He Val Glu Lys Tyr 645 650 655 Gln Asp Lys Val Ser Gin Glu Gin He Gly Lys He Leu Arg Leu Asn 660 665 670 Leu Ser Gly Trp Gly Arg Leu Ser Ser Glu Phe Leu Gin Leu Lys Asn 675 680 685 Ser Gin Pro Gly Glu His Asp Gly Lys Thr Leu He Asn He Met Arg 690 695 700 Gln Thr Gin Met Asn Leu Met Glu He He His Ser Pro Gin Phe Ser 705 710 715 720 Phe Asn Thr Val He Glu Thr Glu Ala Lys Lys Gin Leu Thr Gly His 725 730 735 He Thr His Ser His Val Glu Ala Leu Tyr Cys Ser Pro Val Val Lys 740 745 750 Lys Gin He Trp Gin Ala Leu Gin He Ala Leu Glu Leu Lys Lys Thr 755 760 765 Leu Lys Lys Asp Pro Asn Lys lie Phe Val Glu Thr Thr Arg His Glu 770 775 780 Gly Glu Lys Lys Arg Thr Thr Ser Arg His Lys Gin Leu Leu Glu Leu 785 790 795 800 Tyr Gin Ala Ala Lys Ser His Leu Pro Asp Leu Thr Lys Ser lie Lys 805 810 815 Glu Leu Asn Asp Ala Leu Lys Asp Thr Glu Pro Glu Lys Met Lys Arg 820 825 830 Lys Lys Leu Phe His Tyr Tyr Lys Gin Leu Gly Arg Cys Met Tyr Thr 835 840 845 Gly Arg Pro lie Ser Leu Glu Asp Leu Phe Thr Asn Lys Tyr Asp lie 850 855 860 Asp His lie Tyr Pro Gin Ser Leu Thr Lys Asp Asp Ser Phe Thr Asn 865 870 875 880 Thr Val Leu Val Glu Arg Leu Ser Asn Ala Glu Lys Ser Asp Ala Phe 885 890 895 Pro Leu Asp Ser Lys Thr Arg Lys Asp Arg Gin Gly Leu Trp Arg Cys 900 905 910 Leu Arg Arg Asn Gly Leu lie Thr Lys Glu Lys Tyr Tyr Arg Leu Thr 915 920 925 Arg Glu Thr Pro Leu Ser Glu Glu Glu Lys Ala Ala Phe Ile Arg Arg 930 935 940 Gln Leu Val Glu Thr Ser Gln Thr Thr Lys Glu Val Ile Arg Phe Leu 945 950 955 960 Ala Thr Leu Phe Pro Lys Ser Lys Val Val Tyr Val Lys Ser Gly Asn 965 970 975 Val Ser Asp Phe Arg Arg Asp Phe Ser Pro Ser Leu Pro Glu Asn Lys 980 985 990 Thr Asn Gly Lys Asp Pro Lys Gly Ile Thr Asp Tyr Ser Met Ile Lys 995 1000 1005 Val Arg Glu Ile Asn Asp Leu His His Ala Lys Asp Ala Tyr Leu 1010 1015 1020 Asn Ile Val Val Gly Asn Val Tyr Asp Thr Lys Phe Arg Tyr Arg 1025 1030 1035 Gly Lys Asp Leu Thr Ala Ile Val Arg Glu Lys Ala Arg Gln Tyr 1040 1045 1050 His Leu Ser Arg Leu Phe Leu Tyr Ser Thr Asp Gly Ala Trp Ile 1055 1060 1065 Gly Ala Ala Asp Glu Asn Arg Gly Lys Gln Arg Pro Ser Ile Glu 1070 1075 1080 Thr Val lie Ala Glu Met Arg Arg Asn Ser Cys Gin Val Thr Trp 1085 1090 1095 Glu Ala Val Phe Lys Lys Gly Gin Leu Trp Asp Met Asn Ala Lys 1100 1105 1110 Ser Lys Arg Pro Gly Leu Leu Pro lie Lys Lys Glu Leu Ser Asp 1115 1120 1125 Thr Ala Lys Tyr Gly Gly Tyr Gin Gly Lys Thr Ala Ser Tyr Phe 1130 1135 1140 Val Val Val Glu Tyr Glu Asn Lys Lys Gly Glu Arg Glu Lys Lys 1145 1150 1155 Leu Glu Ser Val Pro lie Tyr Val Lys Ala Leu Ser Lys Gin Lys 1160 1165 1170 Pro Asp Ala Val Asn Ser Phe Leu Arg Asp Thr Leu Gly Leu Glu 1175 1180 1185 Lys Pro Ser Val Met Val Asp Asn lie Lys lie Gly Ser lie Val 1190 1195 1200 Glu lie Asn Gly Ala Arg Met Val Leu Thr Gly Asn Asn Glu Val 1205 1210 1215 Leu Val Phe Gly Arg lie Ala Ser Gin Leu lie Leu Asp lie Thr 1220 1225 1230 Met Ala Ala Tyr Leu Lys Arg Met Phe Lys Leu Leu Ala Asp Thr 1235 1240 1245 Ala Lys Ile Lys Glu Asn Asn Val Tyr Phe Lys Asn Cys Gly Tyr 1250 1255 1260 Leu Asp Lys Glu Thr Asn Leu Ala Val Tyr Asp Thr Phe Ile Ala 1265 1270 1275 Lys Leu Lys Leu Pro Arg Tyr Ala Gln Ile Ile Thr His Ser Leu 1280 1285 1290 Tyr Glu Lys Met Glu Ser Asn Arg Asp Val Phe Ile Asn Leu Ser 1295 1300 1305 Leu Ala Asp Gln Cys Asn Leu Leu Ala Gly Val Leu Pro Ala Leu 1310 1315 1320 Gln Cys Asn Ser Gln Asn Ala Asp Leu Ser Leu Leu Gly Glu Gly 1325 1330 1335 Lys Ala Val Gly Asn Ile Ala Phe Ser Lys Asn Ala Ile Leu Lys 1340 1345 1350 Lys Asn Gln Val Arg Leu Val Asp Cys Ser Ile Thr Gly Leu Phe 1355 1360 1365 Glu Asn Ser Arg Asn Met Ala 1370 1375 <210> 3 <211> 1254 <212> PRT <213> Unknown <220> <223> Cas12a.1 <400> 3 Met Lys Val Ser Thr Trp Asp Ser Phe Thr Asn Gln Tyr Pro Leu Thr 1 5 10 15 Lys Thr Leu Arg Phe Glu Leu Lys Pro Val Gly Lys Thr Leu Gln Lys 20 25 30 Ile Gln Asp Arg Asn Leu Ile Thr Glu Asp Glu Gln Arg Gln Lys Asp 35 40 45 Phe Asn Lys Val Lys Lys Ile Met Asp Gly Tyr Tyr Lys Gln Phe Ile 50 55 60 Glu Glu Cys Leu Glu Gly Ala Lys Ile Pro Leu Lys Lys Leu Glu Glu 65 70 75 80 Asn Asn Asn Ala Tyr Thr Lys Leu Lys Lys Asp Pro Tyr Asn Lys Lys 85 90 95 Leu Arg Glu Glu Tyr Ala Lys Leu Gln Lys Gln Leu Arg Lys Leu Ile 100 105 110 His Asp Glu Ile Asn Lys Lys Glu Glu Phe Lys Tyr Leu Phe Lys Lys 115 120 125 Glu Phe Ile Lys Lys Ile Leu Pro Glu Trp Leu Glu Lys Lys Gly Lys 130 135 140 Lys Glu Glu Leu Lys Glu lie Glu Lys Phe Asp Lys Trp Val Thr Tyr 145 150 155 160 Phe Ser Gly Phe Phe Asn Asn Arg Lys Asn Val Phe Ser Ser Asp Glu 165 170 175 Ile Ser Thr Ser Met lie Tyr Arg lie Val Asn Asp Asn Leu Pro Lys 180 185 190 Phe Leu Asp Asp Val Ser Arg Phe Gly Glu lie Thr Arg Tyr Lys Glu 195 200 205 Phe Asp Ala Asn Gin lie Glu Glu Asn Phe Glu Ser Glu Leu Asn Gly 210 215 220 Glu Lys Leu Lys Asp Phe Phe Asn Leu Lys Asn Phe Asn Asn Cys Leu 225 230 235 240 Asn Gin Glu Gly lie Glu Lys Phe Asn Leu lie lie Gly Gly Lys Ser 245 250 255 Glu Glu Gly Asn Asn Lys lie Lys Gly Leu Asn Glu Leu Val Asn Glu 260 265 270 Leu Ala Gin Lys Gin Ala Asp Lys Asn Glu Gin Lys Lys Val Arg Lys 275 280 285 Leu Lys Leu Ala Pro Leu Phe Lys Gin lie Leu Ser Asp Arg Lys Ser 290 295 300 Ser Ser Phe Ala Phe Glu Lys Phe Glu Glu Asn Thr Glu Val Phe Asp 305 310 315 320 Ala Ile Asp Glu Phe Tyr Asp Lys Ile Ser Leu Glu Thr Leu Lys Lys 325 330 335 Ile Glu Ala Thr Leu Glu Lys Leu Glu Glu Lys Asp Leu Glu Leu Val 340 345 350 Tyr Leu Lys Asn Asp Arg Cys Leu Thr Gly Ile Ser Gln Glu Val Phe 355 360 365 Gly Asp Arg Glu Arg Val Leu Gln Ala Leu Arg Glu Tyr Ala Lys Thr 370 375 380 Glu Leu Gly Leu Lys Thr Asp Lys Lys Ile Glu Lys Trp Met Lys Lys 385 390 395 400 Gly Arg Tyr Ser Ile His Glu Ile Glu Ser Gly Leu Lys Lys Ile Gly 405 410 415 Ser Thr Gly His Pro Ile Cys Asn Tyr Phe Ser Lys Leu Glu Glu Lys 420 425 430 Lys Thr Asn Leu Ile Gln Glu Ile Lys Lys Ala Arg Thr Glu Tyr Glu 435 440 445 Lys Ile Ser Asp Lys Lys Lys Lys Leu Thr Ala Glu Ser Gln Glu Pro 450 455 460 Asn Val Ala Arg Ile Lys Ala Leu Leu Asp Ser Ile Met Arg Leu Tyr 465 470 475 480 His Phe Ile Lys Pro Leu Asn Ile Asn Phe Lys Asn Lys Lys Glu Lys 485 490 495 Asp Ser Glu Ala Leu Glu Thr Asp Asn Asp Phe Tyr Asn Asp Phe Asp 500 505 510 Glu Ser Phe Ala Glu Leu Gly Asn Ile Ile Pro Leu Tyr Asn Gln Val 515 520 525 Arg Asn Tyr Val Thr Gln Lys Pro Phe Ser Thr Glu Lys Phe Lys Leu 530 535 540 Asn Phe Glu Asn Pro Lys Leu Leu Ser Gly Trp Asp Lys Asn Lys Glu 545 550 555 560 Lys Asp Tyr Tyr Ser Val Ile Leu Arg Lys Glu Glu Ser Tyr Tyr Leu 565 570 575 Ala Ile Met Thr Pro Lys Gln Lys Asn Val Phe Asp Glu Leu Glu Arg 580 585 590 Leu Pro Ala Gly Lys Asn Tyr Phe Glu Lys Ile Asp Tyr Lys Leu Leu 595 600 605 Pro Thr Pro Glu Lys Asn Leu Pro Arg Ile Leu Phe Ala Lys Lys Asn 610 615 620 Ile Ser Phe Tyr Lys Pro Ser Lys Glu Ile Glu Ala Ile Arg Asn His 625 630 635 640 Ser Ala His Thr Lys His Gly Asn Pro Gln Asn Gly Phe Lys Lys Arg 645 650 655 Asp Phe Arg Leu Ser Asp Cys His Lys Met Ile Asp Phe Tyr Lys Lys 660 665 670 Ser Ile Gln Lys His Pro Glu Trp Lys Glu Tyr Asp Phe Gln Phe Lys 675 680 685 Lys Thr Glu Asp Tyr Val Asp Ile Ser Glu Phe Tyr Lys Glu Val Ser 690 695 700 Asp Gln Gly Tyr Lys Ile Glu Phe Lys Lys Ile Ser Glu Lys Tyr Leu 705 710 715 720 Leu Asp Leu Val Glu Glu Gly Lys Leu Tyr Leu Phe Gln Ile Trp Asn 725 730 735 Lys Asp Phe Ser Lys Tyr Ser Glu Gly Arg Lys Asn Leu His Thr Ile 740 745 750 Tyr Trp Lys Glu Leu Phe Ser Lys Glu Asn Leu Ser Asp Ile Thr Tyr 755 760 765 Lys Leu Asn Gly Glu Ala Glu Ile Phe Tyr Arg Pro Lys Ser Met Glu 770 775 780 Arg Lys Val Thr His Pro Lys Asn Gln Lys Ile Glu Asn Lys Asp Pro 785 790 795 800 Ile Lys Gly Lys Lys Phe Ser Lys Phe Lys Tyr Asp Phe Ile Lys Asn 805 810 815 Lys Arg Tyr Thr Glu Asp Arg Phe Phe Phe His Cys Pro Ile Thr Leu 820 825 830 Asn Phe Gln Ala Arg Asp Gly Ser Lys Thr Ile Asn Lys Arg Val Asn 835 840 845 Asp His Ile Arg Glu Thr Lys Asp Asp Ile Phe Val Leu Ser Ile Asp 850 855 860 Arg Gly Glu Arg His Leu Ala Tyr Tyr Thr Leu Leu Asn Ser Lys Gly 865 870 875 880 Glu Ile Gln Glu Gln Gly Ser Phe Asn Val Ile Ser Asp Asp Lys Glu 885 890 895 Arg Lys Arg Asp Tyr His Glu Lys Leu Asp Glu Arg Glu Lys Glu Arg 900 905 910 Asp Lys Ala Arg Lys Ser Trp Gln Lys Ile Glu Thr Ile Lys Lys Leu 915 920 925 Lys Asp Gly Tyr Leu Ser Gin lie Val His Lys lie Ala Lys Leu Ala 930 935 940 lie Glu Lys Asn Ala lie lie Val Leu Glu Asp Leu Asn Leu Asp Phe 945 950 955 960 Lys Arg Gly Arg Leu Lys lie Glu Lys Gin Val Tyr Gin Lys Phe Glu 965 970 975 Lys Lys Leu lie Asp Lys Leu Asn Tyr Leu Val Phe Lys Glu Arg Thr 980 985 990 Glu Lys Glu Ala Gly Gly Ser Leu Asn Ala Tyr Gin Leu Thr Gly Lys 995 1000 1005 Phe Glu Gly Phe Lys Lys Leu Gly Lys Glu Thr Gly lie lie Tyr 1010 1015 1020 Tyr Val Pro Ala Ala Tyr Thr Ser Lys lie Cys Pro Lys Thr Gly 1025 1030 1035 Phe Val Asn Leu Leu Arg Pro Lys Phe Lys Asn lie Glu Lys Ala 1040 1045 1050 Lys Glu Phe Phe Lys Lys Phe Asn Tyr lie Lys Tyr Asp Ser Ser 1055 1060 1065 Glu Gly Leu Phe Glu Phe Asn Phe Asp Tyr Ser Lys Phe lie Lys 1070 1075 1080 Asn Gly Lys Lys Glu Thr Lys Ile lie Gin Asp Asn Trp Ser Val 1085 1090 1095 Tyr Ser Asn Gly Thr Lys Leu Val Gly Phe Arg Asn Lys Asn Lys 1100 1105 1110 Asn Asn Ser Trp Asp Thr Lys Glu Val Lys Pro Asn Glu Lys Leu 1115 1120 1125 Lys Ile Leu Phe Lys Glu Tyr Gly Val Ser Phe Gin Lys Asp Glu 1130 1135 1140 Asn lie lie Ser Gin lie Ala Ser Gin Asn Lys Lys Ala Phe Phe 1145 1150 1155 Glu Asn Leu lie Lys lie Phe Lys Thr lie Leu Met Leu Arg Asn 1160 1165 1170 Ser Arg Lys Asp Pro Glu Glu Asp Tyr Val Leu Ser Cys Val Lys 1175 1180 1185 Asp Glu Asn Gly Glu Phe Phe Asp Ser Arg Lys Ala Lys Asp Asn 1190 1195 1200 Glu Pro Lys Asp Ala Asp Ala Asn Gly Ala Tyr His lie Gly Leu 1205 1210 1215 Lys Gly Leu Met Leu Leu Glu Arg lie Lys Ala Asn Lys Gly Lys 1220 1225 1230 Lys Lys Leu Asp Leu Leu lie Ser Arg Asn Asp Phe lie Asn Phe 1235 1240 1245 Ala Val Glu Arg Ser Lys 1250 <210> 4 <211> 1281 <212> PRT <213> Unknown (Unknown) <220> <223> Cas12p <400> 4 Met Lys Lys Ser lie Phe Asp Gin Phe Val Asn Gin Tyr Ala Leu Ser 1 5 10 15 Lys Thr Leu Arg Phe Glu Leu Lys Pro Val Gly Glu Thr Gly Arg Met 20 25 30 Leu Glu Glu Ala Lys Val Phe Ala Lys Asp Glu Thr lie Lys Lys Lys 35 40 45 Tyr Glu Ala Thr Lys Pro Phe Phe Asn Lys Leu His Arg Glu Phe Val 50 55 60 Glu Glu Ala Leu Asn Glu Val Glu Leu Ala Gly Leu Pro Glu Tyr Phe 65 70 75 80 Glu lie Phe Lys Tyr Trp Lys Arg Tyr Lys Lys Lys Phe Glu Lys Asp 85 90 95 Leu Gin Lys Lys Glu Lys Glu Leu Arg Lys Ser Val Val Gly Phe Phe 100 105 110 Asn Ala Gln Ala Lys Glu Trp Ala Lys Lys Tyr Glu Thr Leu Gly Val 115 120 125 Lys Lys Lys Asp Val Gly Leu Leu Phe Glu Glu Asn Val Phe Ala Ile 130 135 140 Leu Lys Glu Arg Tyr Gly Asn Glu Glu Gly Ser Gln Ile Val Asp Glu 145 150 155 160 Ser Thr Gly Lys Asp Val Ser Ile Phe Asp Ser Trp Lys Gly Phe Thr 165 170 175 Gly Tyr Phe Ile Lys Phe Gln Glu Thr Arg Lys Asn Phe Tyr Lys Asp 180 185 190 Asp Gly Thr Ala Thr Ala Leu Ala Thr Arg Ile Ile Asp Gln Asn Leu 195 200 205 Lys Arg Phe Cys Asp Asn Leu Leu Ile Phe Glu Ser Ile Arg Asp Lys 210 215 220 Ile Asp Phe Ser Glu Val Glu Gln Thr Met Gly Asn Ser Ile Asp Lys 225 230 235 240 Val Phe Ser Val Ile Phe Tyr Ser Ser Cys Leu Leu Gln Glu Gly Ile 245 250 255 Asp Phe Tyr Asn Cys Val Leu Gly Gly Glu Thr Leu Pro Asn Gly Glu 260 265 270 Lys Arg Gin Gly lie Asn Glu Leu lie Asn Leu Tyr Arg Gin Lys Thr 275 280 285 Ser Glu Lys Val Pro Phe Leu Lys Leu Leu Asp Lys Gin lie Leu Ser 290 295 300 Glu Lys Glu Lys Phe Met Asp Glu lie Glu Asn Asp Glu Ala Leu Leu 305 310 315 320 Asp Thr Leu Lys lie Phe Arg Lys Ser Ala Glu Glu Lys Thr Thr Leu 325 330 335 Leu Lys Asn lie Phe Gly Asp Phe Val Met Asn Gin Gly Lys Tyr Asp 340 345 350 Leu Ala Gin lie Tyr lie Ser Arg Glu Ser Leu Asn Thr lie Ser Arg 355 360 365 Lys Trp Thr Ser Glu Thr Asp lie Phe Glu Asp Ser Leu Tyr Glu Val 370 375 380 Leu Lys Lys Ser Lys lie Val Ser Ala Ser Val Lys Lys Lys Asp Gly 385 390 395 400 Gly Tyr Ala Phe Pro Glu Phe lie Ala Leu lie Tyr Val Lys Ser Ala 405 410 415 Leu Glu Gin lie Pro Thr Glu Lys Phe Trp Lys Glu Arg Tyr Tyr Lys 420 425 430 Asn Ile Gly Asp Val Leu Asn Lys Gly Phe Leu Asn Gly Lys Glu Gly 435 440 445 Val Trp Leu Gln Phe Leu Leu Ile Phe Asp Phe Glu Phe Asn Ser Leu 450 455 460 Phe Glu Arg Glu Ile Ile Asp Glu Asn Gly Asp Lys Lys Val Ala Gly 465 470 475 480 Tyr Asn Leu Phe Ala Lys Gly Phe Asp Asp Leu Leu Asn Asn Phe Lys 485 490 495 Tyr Asp Gln Lys Ala Lys Val Val Ile Lys Asp Phe Ala Asp Glu Val 500 505 510 Leu His Ile Tyr Gln Met Gly Lys Tyr Phe Ala Ile Glu Lys Lys Arg 515 520 525 Ser Trp Leu Ala Asp Tyr Asp Ile Asp Ser Phe Tyr Thr Asp Pro Glu 530 535 540 Lys Gly Tyr Leu Lys Phe Tyr Glu Asn Ala Tyr Glu Glu Ile Ile Gln 545 550 555 560 Val Tyr Asn Lys Leu Arg Asn Tyr Leu Thr Lys Lys Pro Tyr Ser Glu 565 570 575 Asp Lys Trp Lys Leu Asn Phe Glu Asn Pro Thr Leu Ala Asp Gly Trp 580 585 590 Asp Lys Asn Lys Glu Ala Asp Asn Ser Thr Val Ile Leu Lys Lys Asp 595 600 605 Gly Arg Tyr Tyr Leu Gly Leu Met Ala Arg Gly Arg Asn Lys Leu Phe 610 615 620 Asp Asp Arg Asn Leu Pro Lys Ile Leu Glu Gly Val Glu Asn Gly Lys 625 630 635 640 Tyr Glu Lys Val Val Tyr Lys Tyr Phe Pro Asp Gln Ala Lys Met Phe 645 650 655 Pro Lys Val Cys Phe Ser Thr Lys Gly Leu Glu Phe Phe Gln Pro Ser 660 665 670 Glu Glu Val Ile Thr Ile Tyr Lys Asn Ser Glu Phe Lys Lys Gly Tyr 675 680 685 Thr Phe Asn Val Arg Ser Met Gln Arg Leu Ile Asp Phe Tyr Lys Asp 690 695 700 Cys Leu Val Arg Tyr Glu Gly Trp Gln Cys Tyr Asp Phe Arg Asn Leu 705 710 715 720 Arg Lys Thr Glu Asp Tyr Arg Lys Asn Ile Glu Glu Phe Phe Ser Asp 725 730 735 Val Ala Met Asp Gly Tyr Lys lie Ser Phe Gin Asp Val Ser Glu Ser 740 745 750 Tyr lie Lys Glu Lys Asn Gin Asn Gly Asp Leu Tyr Leu Phe Glu lie 755 760 765 Lys Asn Lys Asp Trp Asn Glu Gly Ala Asn Gly Lys Lys Asn Leu His 770 775 780 Thr lie Tyr Phe Glu Ser Leu Phe Ser Ala Asp Asn lie Ala Met Asn 785 790 795 800 Phe Pro Val Lys Leu Asn Gly Gin Ala Glu lie Phe Tyr Arg Pro Arg 805 810 815 Thr Glu Gly Leu Glu Lys Glu Arg lie lie Thr Lys Lys Gly Asn Val 820 825 830 Leu Glu Lys Gly Asp Lys Ala Phe His Lys Arg Arg Tyr Thr Glu Asn 835 840 845 Lys Val Phe Phe His Val Pro lie Thr Leu Asn Arg Thr Lys Lys Asn 850 855 860 Pro Phe Gin Phe Asn Ala Lys lie Asn Asp Phe Leu Ala Lys Asn Ser 865 870 875 880 Asp lie Asn Val lie Gly Val Asp Arg Gly Glu Lys Gin Leu Ala Tyr 885 890 895 Phe Ser Val lie Ser Gin Arg Gly Lys lie Leu Asp Arg Gly Ser Leu 900 905 910 Asn Val lie Asn Gly Val Asn Tyr Ala Glu Lys Leu Glu Glu Lys Ala 915 920 925 Arg Gly Arg Glu Gin Ala Arg Lys Asp Trp Gin Gin lie Glu Gly lie 930 935 940 Lys Asp Leu Lys Lys Gly Tyr lie Ser Gin Val Val Arg Lys Leu Ala 945 950 955 960 Asp Leu Ala lie Gin Tyr Asn Ala lie lie Val Phe Glu Asp Leu Asn 965 970 975 Met Arg Phe Lys Gin lie Arg Gly Gly lie Glu Lys Ser Val Tyr Gin 980 985 990 Gln Leu Glu Lys Ala Leu lie Asp Lys Leu Thr Phe Leu Val Glu Lys 995 1000 1005 Glu Glu Lys Asp Val Glu Lys Ala Gly His Leu Leu Lys Ala Tyr 1010 1015 1020 Gln Leu Ala Ala Pro Phe Glu Thr Phe Gin Lys Met Gly Lys Gin 1025 1030 1035 Thr Gly lie Val Phe Tyr Thr Gin Ala Ala Tyr Thr Ser Arg lie 1040 1045 1050 Asp Pro Val Thr Gly Trp Arg Pro His Leu Tyr Leu Lys Tyr Ser 1055 1060 1065 Ser Ala Glu Lys Ala Lys Ala Asp Leu Leu Lys Phe Lys Lys Ile 1070 1075 1080 Lys Phe Val Asp Gly Arg Phe Glu Phe Thr Tyr Asp Ile Lys Ser 1085 1090 1095 Phe Arg Glu Gln Lys Glu His Pro Lys Ala Thr Val Trp Thr Val 1100 1105 1110 Cys Ser Cys Val Glu Arg Phe Arg Trp Asn Arg Tyr Leu Asn Ser 1115 1120 1125 Asn Lys Gly Gly Tyr Asp His Tyr Ser Asp Val Thr Lys Phe Leu 1130 1135 1140 Val Glu Leu Phe Gln Glu Tyr Gly Ile Asp Phe Glu Arg Gly Asp 1145 1150 1155 Ile Val Gly Gln Ile Glu Val Leu Glu Thr Lys Gly Asn Glu Lys 1160 1165 1170 Phe Phe Lys Asn Phe Val Phe Phe Phe Asn Leu Ile Cys Gln Ile 1175 1180 1185 Arg Asn Thr Asn Ala Ser Glu Leu Ala Lys Lys Asp Gly Lys Asp 1190 1195 1200 Asp Phe Ile Leu Ser Pro Val Glu Pro Phe Phe Asp Ser Arg Asn 1205 1210 1215 Ser Glu Lys Phe Gly Glu Asp Leu Pro Lys Asn Gly Asp Asp Asn 1220 1225 1230 Gly Ala Phe Asn Ile Ala Arg Lys Gly Leu Val Ile Met Asp Lys 1235 1240 1245 Ile Thr Lys Phe Ala Asp Glu Asn Gly Gly Cys Glu Lys Met Lys 1250 1255 1260 Trp Gly Asp Leu Tyr Val Ser Asn Val Glu Trp Asp Asn Phe Val 1265 1270 1275 Ala Asn Lys 1280 <210> 5 <211> 1137 <212> PRT <213> Unknown <220> <223> Cas12q <400> 5 Met Ile Asn Ile Asp Glu Leu Lys Asn Leu Tyr Lys Val Gln Lys Thr 1 5 10 15 Ile Thr Phe Glu Leu Lys Asn Lys Trp Glu Asn Lys Asn Asp Glu Asn 20 25 30 Asp Arg Val Glu Phe Leu Lys Thr Gln Glu Trp Val Glu Ser Leu Phe 35 40 45 Lys Val Asp Glu Glu Asn Phe Asp Glu Lys Glu Ser lie Pro Asn Leu 50 55 60 Leu Asp Phe Gly Gin Lys lie Ala Ser Leu Phe Tyr Lys Leu Ser Glu 65 70 75 80 Asp lie Ala Asn Asn Gin lie Asp Thr Arg Val Leu Lys Val Ser Lys 85 90 95 Phe Leu Leu Glu Glu lie Asp Arg Asn Gin Tyr His Glu Lys Lys Asn 100 105 110 Lys Pro Thr Lys Val Lys Glu Met Asn Pro Asn Thr Asn Lys Ser Tyr 115 120 125 Ile Lys Glu Tyr Lys Leu Ser Asp Gin Asn Thr Leu Tyr Val Leu Leu 130 135 140 Lys lie Met Glu Asp Glu Gly Arg Gly Leu Gin Lys Phe Leu Tyr Asp 145 150 155 160 Lys Ala Asp Arg Leu Asn Leu Tyr Asn Gin Lys Val Arg Arg Asp Phe 165 170 175 Ala Leu Lys Glu Ser Asn Glu Gin Gin Lys Phe Ser Gly Asn Ala Asn 180 185 190 Tyr Tyr Gly Asn lie Lys Leu Leu lie Asp Ser Leu Glu Asp Ala Val 195 200 205 Arg Ile Ile Gly Tyr Phe Thr Phe Asp Asp Gln Ala Glu Asn Ala Gln 210 215 220 Ile Asn Glu Phe Lys Ser Val Lys Gln Glu Met Asn Asn Asn Glu Ala 225 230 235 240 Ser Tyr Gln Ala Leu Lys Asp Phe Ala Ile Asp Asn Ala Lys Lys Glu 245 250 255 Ile Glu Leu Thr Thr Leu Asn His Arg Ala Val Asn Lys Asp Pro Lys 260 265 270 Lys Ile Gln Glu Gln Ile Glu Glu Val Glu Asn Phe Glu Glu Asp Ile 275 280 285 Asn Gln Leu Lys His Gln Ile Ser Ala Leu Asn Asp Lys Lys Phe Asp 290 295 300 Val Val Ser Arg Leu Lys His Ala Leu Ile Lys Met Leu Pro Glu Leu 305 310 315 320 Asn Leu Leu Asp Ala Glu Ser Glu Gln Gly Arg Glu Val Gln Gln Ile 325 330 335 Tyr Gln Asp Lys Lys Asn Gly Leu Glu Leu Asp Asp Phe Lys Phe Asn 340 345 350 Leu Leu Lys His His Gln Trp Gln Lys Thr Ile Phe Lys Tyr Ile Lys 355 360 365 Leu Glu Gly Leu Val Leu Pro Asp Leu Tyr Ala Glu Asn Lys Gin Asp 370 375 380 Lys He Lys Val Tyr He Glu Asn Tyr Arg Gin Ser Gly Glu Arg He 385 390 395 400 Ser Lys Lys Ala Arg Glu Glu Leu Gly Lys He Asp Lys Arg Glu Glu 405 410 415 Phe Asn Gly Asn Asp Glu Leu Lys Lys Ala Trp Tyr Glu Tyr Lys Asp 420 425 430 Phe Cys Arg Asp Lys Arg Asn Lys Ser Val Glu Leu Gly Asn Lys Lys 435 440 445 Ser Leu Tyr Asn Ala He Lys Arg Glu Val Leu Arg Gin Lys Met Cys 450 455 460 Asn His Phe Ala Val Leu Val Ser Asp Gly Glu Asp Thr Ser Pro Tyr 465 470 475 480 Tyr Tyr Leu He Leu He Pro Asn Glu Asn Ser Asp Glu Met Asn Arg 485 490 495 Thr Phe Lys Glu Leu Lys Ala Ser Glu Gly Asn Trp Lys Met Leu Asp 500 505 510 Tyr Asn Arg Leu Thr Phe Lys Ala Leu Glu Lys Leu Ala Leu Leu Arg 515 520 525 Ser Ser Thr Phe Glu Ile Ala Asp Gin Glu Leu Gin Glu Glu Ala Lys 530 535 540 Lys Ile Trp Glu Glu Tyr Lys Glu Lys Ala Tyr Lys Asp Phe Lys Asn 545 550 555 560 Lys Lys Leu Leu Gin Gly Leu Ser Gly Arg Gin Arg Glu Glu Lys Lys 565 570 575 Gln Glu Leu Gin Lys Glu Ser Leu Asn Arg Val Ile Asn Tyr Leu Ile 580 585 590 Arg Cys Ile Gin Ser Leu Pro Asp Ser Gly Lys Tyr Asn Phe Asn Phe 595 600 605 Lys Glu Pro His Gin Tyr Gin Ser Leu Glu Glu Phe Ala Glu Glu Ile 610 615 620 Asp Arg Gin Gly Tyr His Cys Ala Trp Lys Asn Val Ser Lys Asp Lys 625 630 635 640 Leu Met Glu Leu Glu Ala Met Glu Lys Ile Lys Val Phe Lys Leu His 645 650 655 Asn Lys Asp Phe Arg Lys Val Lys Leu Asn Asp Ser Lys His Asn Pro 660 665 670 Asn Leu Phe Thr Leu Tyr Trp Leu Asp Ala Met Asn Leu Asp Lys Val 675 680 685 Asn Val Arg Leu Leu Pro Glu Val Asp Leu Tyr Lys Arg Ala Lys Glu 690 695 700 Thr Gln Leu Lys Leu Phe Glu Arg Asp Val Lys Cys Asn Ile Asn Asn 705 710 715 720 Gln Lys Ile Lys Ser Ile Lys Glu Lys Asn Arg Leu Phe Gln Asp Lys 725 730 735 Leu Tyr Ala Ser Phe Lys Leu Glu Phe Tyr Pro Glu Asn Glu Gly Leu 740 745 750 Gly Phe Glu Gln Val Asn Asp Lys Val Asn Asn Phe Cys Gly Ser Asp 755 760 765 Thr Ala Tyr Tyr Leu Gly Leu Asp Arg Gly Glu Lys Glu Leu Val Thr 770 775 780 Phe Cys Leu Val Asp Ser Asp Gly Arg Leu Val Lys Asn Gly Asp Trp 785 790 795 800 Thr Lys Phe Lys Glu Val Asn Tyr Ala Asp Lys Leu Lys Gln Phe Tyr 805 810 815 Tyr Ser Lys Gly Glu Ile Glu Ser Thr Gln Gln Gln Leu Leu Glu Ala 820 825 830 Arg Asp Asn lie Lys Gin Ala Thr Asn Thr Glu Asp Lys Glu Ser Met 835 840 845 Lys Leu Asn Tyr Lys Lys Leu Glu Leu Lys Leu Lys Gin Gin Asn Leu 850 855 860 Leu Ala Gin Glu Phe lie Lys Lys Ala Tyr Cys Gly Tyr Leu lie Asp 865 870 875 880 Ser lie Asn Glu lie Leu Arg Glu Tyr Pro Asn Thr Tyr Leu Val Leu 885 890 895 Glu Asp Leu Asp lie Ala Gly Lys Ala Asp Pro Glu Ser Gly Met Thr 900 905 910 Asn Lys Glu Gin Asn Leu Asn Lys Thr Met Gly Ala Ser Val Tyr Gin 915 920 925 Ala lie Glu Asn Ala lie Val Asn Lys Phe Lys Tyr Arg Thr Val Lys 930 935 940 Leu Ser Asp lie Lys Gly Leu Gin Thr Val Pro Asn Val Val Lys Val 945 950 955 960 Glu Asp Leu Arg Glu Val Lys Glu Val Glu Asp Gly Glu His Lys Phe 965 970 975 Gly Leu lie Arg Ser Val Lys Ser Lys Asp Gin lie Gly Asn lie Leu 980 985 990 Phe Val Asp Glu Gly Glu Thr Ser Asn Thr Cys Pro Asn Cys Gly Phe 995 1000 1005 Asn Ser Asp Trp Phe Lys Arg Asp Val Asp Phe Asp Leu Glu Ile 1010 1015 1020 Val Ala Thr Val Asn Gly Gln Lys Asn Ala Val Ile Glu Gln Asn 1025 1030 1035 Asp Lys Lys Tyr Cys Phe Pro Gly Glu Ile Tyr Lys Leu Glu Ile 1040 1045 1050 Ile Asn Lys Glu Tyr Glu Thr Asn Lys Arg Asn Leu Ala Met Ile 1055 1060 1065 Phe Lys Pro Arg Ala Lys Ala Cys Arg Lys Phe Ile Asn Asn Asn 1070 1075 1080 Leu Asp Lys Asn Asp Tyr Phe Tyr Cys Pro Tyr Cys Ala Phe Ser 1085 1090 1095 Ser Lys Asn Cys Asn Asn Pro Lys Leu Gln Asn Gly Asp Phe Val 1100 1105 1110 Val Tyr Ser Gly Asp Asp Val Ala Ala Tyr Asn Val Ala Ile Arg 1115 1120 1125 Gly Ile Asn Leu Leu Asn Asn Ile Lys 1130 1135 <210> 6 <211> 36 <212> DNA <213> Unknown (Unknown) <220> <223> Cas12a.1 direct repeat <400> 6 gtttaaggcc ttgacaaaat ttctactgta gtagat 36 <210> 7 <211> 36 <212> DNA <213> Unknown (Unknown) <220> <223> Cas12p direct repeat <400> 7 atctacaaaa gtagaaatct aatagggata ttcgag 36 <210> 8 <211> 36 <212> DNA <213> Unknown (Unknown) <220> <223> Cas12q direct repeat <400> 8 atctacaaaa gtagaaatta aataggtcta tttgag 36 <210> 9 <211> 36 <212> DNA <213> Unknown (Unknown) <220> <223> Cas9.1 direct repeat <400> 9 actgtagcaa gacgaagggc cggcgcaatc cgcagc 36 <210> 10 <211> 1031 <212> PRT <213> Unknown (Unknown) <220> <223> Cas9.3 <400> 10 Met Ile Phe Gly Leu Asp Val Gly Thr Thr Ser Ile Gly Phe Ala Leu 1 5 10 15 Ile Ser Leu Asp Glu Asp Lys Glu Thr Gly Cys Ile Val His Ser Gly 20 25 30 Cys Arg Val Phe Pro Glu Gly Val Thr Glu Asp Lys Lys Glu Ser Arg 35 40 45 Asn Lys Ala Arg Arg Glu Ala Arg Leu Arg Arg Arg Gln Leu Arg Arg 50 55 60 Lys Lys Glu Asn Arg Lys Arg Leu Ala Gln Phe Leu His Glu Thr Ser 65 70 75 80 Leu Leu Pro Val Phe Gly Ser Thr Glu Trp Lys Asn Leu Met Asp Asn 85 90 95 Thr His Ser Asn Pro Tyr Glu Leu Arg Ser Ala Ala Leu Lys Lys Gln 100 105 110 Leu Gln Pro Phe Glu Leu Gly Lys Val Ile Tyr His Leu Ala Lys His 115 120 125 Arg Gly Phe Lys Ala Thr Lys Leu Asp Glu Leu Met Ala Glu Ser Asp 130 135 140 Glu Lys Lys Glu Leu Gly Val Val Lys Asp Gly Ile Lys Glu Leu Asp 145 150 155 160 His Lys Leu Gly Asp Gin Thr Leu Gly Val Tyr Leu Ala Ser lie Pro 165 170 175 Pro Ser Glu Lys Lys Arg Gly Arg Tyr Leu Gly Arg Tyr Met lie Gin 180 185 190 Glu Glu Leu Glu Gin lie Leu Glu Tyr Gin Lys His Tyr Asn Pro Glu 195 200 205 Leu lie Thr Ser Thr Phe Lys Lys His Leu Asn Ser Leu lie Phe Ser 210 215 220 Gln Arg Pro Thr Phe Trp Arg Leu Asn Thr Leu Gly Thr Cys Ser Leu 225 230 235 240 Glu Gin Asn Glu Ser Val Cys Pro Lys His Ser Trp lie Gly Gin Gin 245 250 255 Phe lie Met Met Gin Lys Val Asn Asp Leu Arg lie Val Glu Pro His 260 265 270 Pro Arg His Leu Thr Met Glu Glu Arg Thr Gin Leu lie Gin Gly Leu 275 280 285 Cys Lys Gin Lys lie Met Ser Phe Gly Gly lie Arg Lys Leu Leu His 290 295 300 Leu Pro Lys Gly Thr Val Phe Asn Phe Glu Thr Tyr Gin Asp Lys Glu 305 310 315 320 Asp Lys Arg Gly Leu Pro Gly Asn Ala Ile Glu Ala Ala Leu Ser Thr 325 330 335 Ile Phe Gly Ser Glu Trp Lys His Leu Pro His Lys Asp Ala Ile Arg 340 345 350 Ser Ser Leu Ser Asn Arg Ile Trp Ser Ile Ser Tyr Asn Arg Val Gly 355 360 365 Asn Lys Arg Ile Glu Ile Arg Ala Asp Glu Ser Tyr Gln Asn Gln Arg 370 375 380 Gln Thr Val Lys Gln Glu Met Met Lys Asp Trp Asn Ile Ala Glu Asp 385 390 395 400 Gln Ala Glu Gln Leu Val Gln Leu Pro Ile Pro Pro Gln Trp Leu Arg 405 410 415 Phe Ser Glu Lys Ala Ile Gln Lys Leu Leu Pro Asp Leu Glu Ser Gly 420 425 430 Val Pro Leu Gln Thr Ala Ile Lys Glu His Tyr Pro Glu Thr Leu Lys 435 440 445 Ser Ser Glu Val Glu His Glu Leu Leu Pro Ser Ser Pro His Leu Val 450 455 460 Pro Glu Leu Arg Asn Pro Thr Val Asn Arg Ala Leu Asn Glu Leu Arg 465 470 475 480 Lys Val Val Asn Asn Ile Ile Arg Ser Tyr Gly Lys Pro Asp Ile Ile 485 490 495 Arg Ile Glu Leu Ala Arg Asp Leu Lys Leu Gly Lys Lys Lys Lys Leu 500 505 510 Glu Ile Thr Lys Lys Asn Arg Gln Arg Glu Gln Glu Arg Lys Glu Ala 515 520 525 Lys Asn Gln Leu Glu Lys Glu Gly Val Lys Pro Thr Gly Met Asn Ile 530 535 540 Glu Lys Phe Leu Leu Trp Gln Glu Ser Asp Gly Leu Asp Leu Tyr Thr 545 550 555 560 Gly Gln Lys Ile Ser Phe Ala Ala Leu Phe Lys Gln Thr Glu Tyr Asp 565 570 575 Ile Glu His Ile Ile Pro Arg Ser Arg Ser Phe Asn Asn Thr Phe Phe 580 585 590 Asn Lys Thr Leu Ala His Asn Glu Ile Asn Arg Gln Lys Gly Asn Met 595 600 605 Ile Pro Lys Glu Phe Phe Gly Asp Gly Glu Thr Trp His Ala Phe Val 610 615 620 Thr Arg Val Asn Gln Ser Lys Leu Pro Leu Glu Lys Lys Glu Lys Leu 625 630 635 640 Leu Ile Pro His Tyr Asp Ala Ile Ala Ser Glu Glu Met Thr Glu Arg 645 650 655 Gln Leu Arg Asp Thr Ala Tyr Ile Ala Thr Glu Ala Lys Thr Tyr Leu 660 665 670 Gln Thr Leu Gly Ile Pro Val Gln Pro Thr Asn Gly Arg Ala Thr Ala 675 680 685 Ser Leu Arg Arg Val Trp Gly Ile Asn Ser Ile Trp Ala Thr Glu Phe 690 695 700 Gly Leu Glu Glu Glu Ser Lys Lys Ala Ala Gly Glu Lys Ile Arg Asp 705 710 715 720 Asp His Arg His His Ala Val Asp Ala Ala Val Val Ala Leu Thr Ser 725 730 735 Pro Gly Arg Ile Lys Arg Leu Ser Thr Phe Tyr Gln Tyr Arg Lys Glu 740 745 750 Met Lys Pro Asp Asp Phe Pro Leu Pro Trp Glu Thr Phe Arg Ala Asp 755 760 765 Leu Ile Thr Ser Leu His Lys Ile Ile Ile Ser His Arg Val Gln Arg 770 775 780 Lys lie Ser Gly Pro Leu His Glu Glu Thr Ala Tyr Gly Phe Thr Lys 785 790 795 800 Lys Lys Ser Glu Thr Asp Pro Thr Ala Tyr Tyr Phe Val Thr Arg Lys 805 810 815 Thr Leu Asp Lys Asp Phe Lys Pro Asn Lys Val Lys Asp lie Val Asp 820 825 830 Pro Ala Val Arg His Leu lie Gly Glu His Leu Gin Lys Phe Asp Asn 835 840 845 Asn Pro Ala Val Ala Phe Ala Pro Glu Asn Arg Pro His Met Pro Leu 850 855 860 Arg Lys Gly Gly Trp Gly Pro Pro lie Lys Lys Val Arg lie Gin lie 865 870 875 880 Ala Arg Asn Pro Gin Phe Met Val Ser Arg Gin Lys Asn Pro lie Ser 885 890 895 Tyr Tyr Asp Ser Gly Asp Asn His His Met Ala lie Tyr Gly Thr His 900 905 910 Leu Asp Asp Gly Thr Val Asp Pro Glu Thr Val Ser Phe Glu Val Val 915 920 925 Ser Arg Phe Glu Val Asn Gin Arg Ala Ser Lys Asn Glu Pro Leu Val 930 935 940 Lys Pro Gin Asn Glu Asn Gly Val Pro Leu Leu Phe Thr Leu Val Lys 945 950 955 960 Asn Asn Val Leu lie Trp Asn Glu Pro Gly Glu Glu Glu Gin Met His 965 970 975 Leu Val Arg Trp Thr Thr Ala Asn Lys Gly Arg lie Phe His Lys Pro 980 985 990 Leu Trp Met Ser Gly Thr Pro Pro lie Glu lie Ser lie Ser Val Lys 995 1000 1005 Asn Leu lie Ser Tyr Gly Gly Arg Lys Val Ser Val Asp Pro lie 1010 1015 1020 Gly Asn lie Phe Pro Cys Asn Asp 1025 1030 <210> 11 <211> 1308 <212> PRT <213> Unknown <220> <223> Cas9.4 <400> 11 Met Lys Lys lie Leu Gly Leu Asp Leu Gly Thr Asp Ser lie Gly Trp 1 5 10 15 Thr lie Val Gin Gin Asn Glu Glu Lys Lys Phe Lys Leu lie Asp Lys 20 25 30 Gly Val Arg lie Phe Gin Lys Gly Val Gly Glu Glu Lys Asn Asn Glu 35 40 45 Phe Ser Leu Ala Lys Glu Arg Thr Thr His Arg Asn Thr Arg Lys Lys 50 55 60 Tyr Arg Arg Thr Lys Gln Arg Lys Val Arg Leu Leu Arg Glu Leu Ile 65 70 75 80 Lys His Gly Met Cys Pro Leu Ser Phe Asp Glu Leu Glu Leu Trp Ser 85 90 95 Lys Tyr Arg Lys Gly Lys Pro Tyr Ile Tyr Pro Leu Ser Asn Lys Gly 100 105 110 Phe Thr Gln Trp Leu Lys Leu Asn Pro Tyr Asp Leu Arg Glu Arg Ala 115 120 125 Ile Lys Pro Asp Glu Lys Leu Thr Pro Leu Glu Leu Gly Arg Ile Phe 130 135 140 Tyr His Ile Thr Gln Arg Arg Gly Phe Lys Ser Asn Arg Lys Asp Asn 145 150 155 160 Ser Glu Asp Ser Glu Gly Val Val Lys Thr Ser Ile Ser Gln Leu Arg 165 170 175 Glu Glu Met Glu Gly Lys Thr Leu Gly Gln Phe Phe Asn Asn Glu Leu 180 185 190 Light Light Gly Asn Light Val Arg Light Light Tyr Thr Ala Arg Glu Asp Tyr 195 200 205 His His Glu Phe Asn Glu Ile Cys Asn Ile Gln Lys Ile Asp Asn Lys 210 215 220 Thr Lys Ala Ala Leu Glu Arg Glu Ile Phe Phe Gln Arg Ala Leu Lys 225 230 235 240 Ser Gln Arg His Leu Val Gly Lys Cys Thr Leu Glu Pro Lys Lys Pro 245 250 255 Arg Cys Pro Leu Ser Ala Ile Pro Tyr Glu Glu Phe Arg Ala Leu Gln 260 265 270 Phe Ile Asn Ser Ile Arg Ile Lys Asp Ala Glu Glu Asn Leu Met Pro 275 280 285 Leu Thr Gln Lys Glu Arg Glu Val Ile Gln Ser Leu Phe Phe Arg Lys 290 295 300 Ser Lys Pro Ser Phe Pro Phe Asn Asp Ile Lys Lys Ile Leu Glu Lys 305 310 315 320 His Asn Gly Gln Arg Leu Thr Phe Asn Tyr Pro Glu Lys Leu Gln Ile 325 330 335 Ile Gly Ser Pro Thr Ile Ala Leu Leu Lys Ser Val Phe Gly Glu Glu 340 345 350 Trp Ala Ser Leu Ser Val Ala Tyr Thr Lys Lys Asp Gly Thr Thr Gly 355 360 365 Thr Ile Asn Ser Glu Asp Val Trp His Ala Leu Phe Glu Phe Glu His 370 375 380 Asn Asp Lys Leu Glu Asp Phe Leu Lys Gln Arg Leu Lys Leu Ser Asp 385 390 395 400 Asp Asn Ile Gln Lys Leu Ile Lys Gly Asn Leu Lys Gln Gly Tyr Ala 405 410 415 Ser Leu Ser Arg Lys Ala Ile Asn Asn Ile Leu Pro Phe Leu Lys Asp 420 425 430 Gly His Ile Tyr Thr His Ala Val Phe Leu Ala Lys Ile Pro Glu Ile 435 440 445 Ile Gly Arg Lys Gln Trp Leu His Ser Lys Asp Gln Ile Val Asn Trp 450 455 460 Phe Leu Lys Ser Ala Glu Glu Leu Pro Leu Lys Asn Arg Leu Cys Lys 465 470 475 480 Ile Val Asn Asn Leu Ile Thr Glu Phe Asn Glu Thr Tyr Ala Asn Ala 485 490 495 Asp Pro Lys Tyr Ile Leu Asp Asp Ser Asp Lys Lys Ser Ile Asn Arg 500 505 510 Ser Leu Gin His Asp Phe Gly Pro Lys Thr Trp Asn Lys Phe Ser Ser 515 520 525 Glu Lys Lys Asp Glu Leu Gin Lys Glu Thr Glu Arg Leu Phe Leu Ser 530 535 540 Gln lie Asn Lys Gly Asn Ala Ser Ala Pro Tyr lie Lys Pro Tyr Arg 545 550 555 560 Gln Asp Glu Glu Leu Lys Gin Tyr Leu lie Asp Asn Phe Asn lie Lys 565 570 575 Gln Glu Glu Ala Glu Arg lie Tyr His Pro Ser Ala lie Asp lie Phe 580 585 590 Asp Glu Ala Pro Tyr Asn Asp Asp Gly lie Lys Leu Leu Gin Ser Pro 595 600 605 Arg Thr Pro Ser Ala Arg Asn Pro Met Ala Met Arg Ala Leu His Glu 610 615 620 Leu Arg Tyr Leu Leu Asn Gin Leu Leu Ser Gin Arg Gly lie Asp Glu 625 630 635 640 His Thr Val lie His Leu Glu Met Ser Arg Glu Leu Asn Asn Gin Asn 645 650 655 Lys Arg Leu Ala lie Gin Arg Tyr Gin Gin Ala Arg Asn Glu Glu His 660 665 670 Gln Glu Tyr Ala Lys Glu Ile Lys Lys Ile Phe Lys Glu Gln Thr Gln 675 680 685 Lys Glu Ile Glu Pro Thr Glu Ala Asp Ile Leu Lys Tyr Arg Leu Trp 690 695 700 Lys Glu Gln Glu His Asn Cys Leu Tyr Thr Gly Arg Lys Ile Gly Ile 705 710 715 720 Ala Asp Phe Ile Gly Asp Asn Ser Asn Val Asp Ile Glu His Thr Trp 725 730 735 Pro Arg Ser Lys Ser Phe Asp Asn Ser Thr Ala Asn Lys Thr Leu Cys 740 745 750 Asp Ser His Tyr Asn Arg Asn Ile Lys Lys Asn Lys Ile Pro Tyr Asp 755 760 765 Leu Pro Asn Phe Lys Glu Ser Ala Ile Ile Glu Gly Lys Gln Tyr Asp 770 775 780 Pro Ile Lys Ala Arg Leu Lys Asp Trp Glu Glu Lys Cys Asn His Leu 785 790 795 800 Lys Glu Leu Ala Ala Lys Tyr Arg Tyr Asn Ala Lys Arg Ala Ser Thr 805 810 815 Lys Glu Gln Lys Asp Lys Ala Leu Gln Asn Ala His Phe Tyr Gln Met 820 825 830 His His Glu Tyr Trp Lys Asp Lys Ile Phe Arg Phe Thr Gly Lys Glu 835 840 845 Ile Arg Asn Ser Phe Lys Asn Ser Gln Leu Val Asp Thr Gly Ile Ile 850 855 860 Asn Lys Tyr Ala Arg Ala Tyr Leu Gln Thr Val Phe Asn Lys Val Phe 865 870 875 880 Thr Ile Lys Gly Thr Leu Thr Ala Asp Phe Arg Lys Ala Trp Gly Ile 885 890 895 Gln Asn Pro Asp Thr Ser Lys Ser Arg Gln Arg His Thr His His Ala 900 905 910 Ile Asp Ala Ala Val Val Ala Cys Leu Thr Arg Asp Arg Tyr Asp Phe 915 920 925 Leu Thr Gln Trp Tyr Arg Ala Glu Glu Lys Gly Asn Glu Arg Lys Lys 930 935 940 His Ile Ile Gln Glu Arg Met Lys Pro Trp Thr Thr Phe Val Gln Asp 945 950 955 960 Ile Lys Ala Phe Glu Asn Ser Ile Leu Val Ser His His Thr Arg Lys 965 970 975 Thr Ser Ala Lys Gin Thr Arg Lys Arg Leu Arg Glu Asn Gly Lys He 980 985 990 Val Lys Asp Pro Asn Gly Asn Pro He Tyr Ser Lys Gly Asp Thr Phe 995 1000 1005 Arg Asn Arg Leu His Lys Asp Thr Phe Tyr Gly Ala He Leu Arg 1010 1015 1020 Pro Gin He Asp Lys Glu Gly Lys Thr Val Thr Asp Glu Asn Gly 1025 1030 1035 Asn Pro Lys Leu Thr Thr Gin Tyr Val Val Lys Lys Pro Val Thr 1040 1045 1050 Asp Leu Lys Glu Thr Asp He Lys Asn He Val Asp Ser Lys He 1055 1060 1065 Lys Ser Leu Phe Glu Ser Lys Lys Leu Asn Glu He Gin Lys Glu 1070 1075 1080 Gly He Ser He Pro Pro Ser Lys Pro Glu Gly Lys Glu Thr Pro 1085 1090 1095 Ile Lys Ser Val Arg Leu Lys Gin Pro Phe Asn Pro He Pro Leu 1100 1105 1110 Arg Glu His Thr His Leu Ser Gin Lys Pro His Lys Gin Tyr Tyr 1115 1120 1125 His Val Gin Asn Glu Gly Asn Phe Leu Met Ala lie Tyr Glu Glu 1130 1135 1140 Thr Ser Ala Ser Lys Lys Pro Glu Lys Thr Phe Glu Leu lie Ser 1145 1150 1155 Asn Leu Gin Ala Ala Asp Tyr Tyr Lys Ala Ser Asn Lys Glu Asn 1160 1165 1170 Arg Glu Gin Tyr Pro lie Val Pro Glu Arg Lys Phe lie Thr Lys 1175 1180 1185 Arg Asn Lys Glu lie Glu Leu Pro Leu Lys Gin lie lie Tyr lie 1190 1195 1200 Gly Gin Met Val Met Leu Tyr Glu Asn Ser Pro Glu Glu Leu Lys 1205 1210 1215 Ser Lys Asn Glu Glu Glu Leu Phe Lys Cys Leu Tyr Lys lie Val 1220 1225 1230 Gly lie Thr Ser Met Thr lie Gin Ala Lys Tyr Glu Tyr Gly Val 1235 1240 1245 Phe lie Leu Lys His His Ala lie Ser Thr Pro Tyr Ser Glu Leu 1250 1255 1260 Lys Pro Lys Asp Gly Asp Phe Ser Trp Glu Gly Asn lie Glu Ala 1265 1270 1275 Met Arg Lys Gin Leu His Ser Arg lie Lys Val Val lie Gin Asn 1280 1285 1290 Leu Asp Phe Lys lie Thr Pro Thr Gly Lys lie Gin Trp Leu Phe 1295 1300 1305 <210> 12 <211> 25 <212> DNA <213> Unknown (Unknown) <220> <223> Cas12q spacer <400> 12 gcttggaaat atgtcttatt tatca 25 <210> 13 <211> 3765 <212> DNA <213> Unknown (Unknown) <220> <223> Cas12a.1 <400> 13 atgaaggtct cgacttggga ttcgtttaca aaccaatacc ccctaacgaa aactctacgc 60 ttcgaattaa agccagtcgg caaaacactg cagaaaattc aagatcgcaa cctgattaca 120 gaagacgaac aacgccaaaa agatttcaac aaagtcaaaa aaataatgga cggatactac 180 aagcaattca tagaagaatg cttggaaggt gccaagatac cgttaaaaaa attggaagaa 240 aacaacaacg cttacacgaa actgaaaaaa gacccttaca acaaaaaatt aagggaagaa 300 tacgcaaaac tccaaaaaca attaaggaaa ctaattcacg acgaaataaa taaaaaagaa 360 gaattcaaat acttgttcaa gaaagaattc atcaaaaaaa tattgccgga atggctcgaa 420 aaaaaaggga aaaaagagga actcaaagaa atcgaaaaat tcgataaatg ggttacctac 480 tttagcggtt tttttaacaa ccgcaaaaac gttttttcaa gcgatgaaat ttcgacgtca 540 atgatttaca ggatagtcaa cgacaaccta ccgaaattcc tagatgacgt ttcacgcttc 600 ggagaaataa ccagatacaa ggaatttgac gccaaccaaa tagaagaaaa ctttgaaagc 660 gagttgaacg gagagaaatt aaaagatttt ttcaacttga aaaacttcaa caactgcctt 720 aaccaagaag gaatagaaaa attcaactta atcataggag gcaaaagcga agaaggcaac 780 aataaaataa agggcttaaa cgaattagtc aacgaactcg cccaaaaaca agcggacaaa 840 aacgagcaaa aaaaggttag aaaattaaaa ctcgcgccgt tattcaagca aatcttaagt 900 gaccgcaaat cctcctcgtt cgcattcgaa aaattcgagg aaaatacgga ggtattcgat 960 gcaatagacg aattttacga taaaataagc ttggaaacac tcaaaaaaat agaagcgacc 1020 ctcgaaaagc tagaagaaaa agatttggaa ttagtttact tgaaaaacga tagatgccta 1080 acaggaattt cacaagaagt attcggggat cgggaaagag tacttcaagc cctaagggaa 1140 tacgcgaaaa ccgaactcgg cctcaaaacc gacaaaaaaa tagaaaaatg gatgaaaaaa 1200 ggcaggtatt caatccacga aatagagagc ggcctcaaaa aaatcggttc aaccggacac 1260 ccgatatgta attatttctc aaaactagaa gaaaaaaaga caaacttgat tcaagaaata 1320 aaaaaagcgc gcactgaata tgaaaaaata agtgacaaaa aaaagaaatt aactgctgaa 1380 agccaagagc ccaacgtcgc aagaataaag gcgttactgg actcaataat gcggctatac 1440 cacttcataa aacccctcaa catcaacttc aaaaacaaga aagaaaagga ttcagaggca 1500 cttgaaaccg ataacgattt ctataacgat ttcgacgaat cgtttgcgga actagggaat 1560 ataatcccac tatacaatca agtcagaaac tatgttacgc aaaaaccgtt cagcaccgaa 1620 aaattcaagt taaactttga aaatcccaaa ctcctaagcg gctgggacaa aaacaaggaa 1680 aaagactatt attctgttat attgagaaaa gaggagtcat actacttagc cattatgacc 1740 ccaaaacaaa aaaacgtttt tgacgaactg gaacggcttc cggctgggaaa aaattatttt 1800 gaaaaaatag actacaaatt attgcctacc ccagaaaaaa atctacctag aatattattt 1860 gcaaaaaaaa acatttcatt ttacaagcca tcaaaagaaa tcgaagcgat tcgtaatcac 1920 tctgcccaca ccaagcatgg aaacccacaa aacgggttca aaaaaaggga tttccgatta 1980 agcgattgcc ataaatgat tgacttttac aaaaagagca ttcaaaaaca ccccgaatgg 2040 aaagaatacg atttccaatt caaaaaaacg gaagattacg tcgacatatc agaattttat 2100 aaagaagtat ccgaccaagg ctataaaata gaattcaaaa aaataagcga aaaatatttg 2160 cttgacttgg tcgaagaagg aaaactttac ttattccaaa tttggaacaa ggacttttcg 2220 aagtattcgg aaggccgtaa aaacctgcac acaatttact ggaaagaact attctccaaa 2280 gaaaaccttt cagacataac ttacaaatta aacggcgaag ccgaaatatt ctaccgccca 2340 aagtcaatgg aaaggaaagt aactcaccca aaaaaccaaa aaatagaaaa caaagacccg 2400 attaaaggga aaaaattcag taaattcaaa tacgacttta taaaaaacaa aaggtacacc 2460 gaagaccgtt tcttcttcca ctgcccgata accttgaact tccaggcgcg cgatggcagc 2520 aaaacgatta acaagcgggt caacgaccac atacgcgaaa caaaagatga cattttcgtg 2580 ttaagcattg accgcgggga aaggcacttg gcgtactaca cgctattgaa ttcaaaagga 2640 gaaatccaag aacaaggctc tttcaacgta atctcggacg acaaagaaag aaaacgtgat 2700 taccacgaaa aactggatga acgcgaaaaa gaacgcgaca aagcaaggaa aagctggcag 2760 aaaatcgaga ccataaagaa attgaaggat ggctacctat cccaaatcgt acacaaaatc 2820 gctaaactcg caatagaaaa aaacgcgata atcgtcttgg aagacctgaa cttagacttc 2880 aagcgcggga gattaaaaat cgagaagcaa gtataccaaa agttcgagaa aaaactaata 2940 gacaaactca attacttggt tttcaaggaa agaaccgaaa aagaagccgg cggatcccta 3000 aacgcatacc aactaaccgg aaaatttgaa ggatttaaga aactcggaaa agaaacaggt 3060 ataatatact acgttcccgc ggcgtacacc tcgaagattt gcccgaaaac aggcttcgta 3120 aatctgttaa gacctaaatt caagaacata gaaaaagcta aggaattctt caaaaaattc 3180 aactacatca aatacgattc gagcgaaggc ttattcgaat tcaacttcga ctactccaaa 3240 ttcattaaaa acggaaaaaa agaaacaaaa ataattcaag acaattggtc ggtttactcg 3300 aacggaacga aactagtcgg cttcagaaac aagaataaaa acaattcatg ggatacaaag 3360 gaagtcaaac cgaacgaaaa actaaaaata ttgttcaaag aatacggggt ttccttccaa 3420 aaagacgaaa atattataag ccaaatagcc agccaaaaca aaaaagcttt ctttgaaaac 3480 ctcattaaaa tctttaaaac gattttaatg ttacgcaact caagaaaaga ccccgaagaa 3540 gattacgtac tttcctgcgt aaaagacgaa aacggcgaat tcttcgactc aagaaaagct 3600 aaagacaacg agcccaagga cgccgacgcg aacggcgctt accacatagg gttgaaagga 3660 ttaatgctct tggaaagaat aaaggccaac aaaggaaaga aaaaactcga tttactaatc 3720 agcaggaacg acttcatcaa cttcgcagtt gaacggagca agtaa 3765 <210> 14 <211> 3846 <212> DNA <213> Unknown <220> <223> Cas12p <400> 14 atgaaaaaat ctatttttga tcagtttgta aatcagtatg ctctttctaa aacgttgcgg 60 tttgaattga agccggtggg ggagacgggg aggatgcttg aggaggcgaa ggtttttgct 120 aaagatgaaa caatcaagaa aaaatatgag gcaaccaagc ctttttttaa taaattgcat 180 cgtgaatttg tagaggaggc tttaaatgag gtggaattag ctggtttgcc tgaatatttt 240 gaaatattta aatattggaa aaggtataaa aagaagtttg aaaaggattt gcagaagaaa 300 gaaaaagaat tgcggaaatc agttgtaggt ttttttaatg cacaggcaaa ggaatgggcg 360 aaaaaatatg aaactttggg tgtgaagaaa aaagatgtgg gacttttatt tgaagaaaat 420 gtttttgcta tattgaagga aaggtacgga aatgaggagg gatcacaaat tgttgatgaa 480 agtacaggaa aagatgtttc gatatttgat agttggaagg gctttacagg gtattttatt 540 aaattccagg aaactcgtaa gaatttttat aaggatgatg gcacggctac tgctttggct 600 acaaggatta ttgatcaaaa tttgaagcgt ttttgtgata atttactaat atttgaaagt 660 attagagata aaattgattt ttcagaggta gaacaaacta tgggaaactc tattgataag 720 tttttttcag taatttttta tagttcctgt ttacttcagg aaggaattga tttttataat 780 tgtgttttag gtggggagac tctgccaaat ggtgaaaaga gacagggaat aaatgagctt 840 attaatctct ataggcaaaa aactagtgag aaagtacctt ttttaaagtt gcttgataag 900 cagattttga gtgaaaaaga gaagtttatg gatgaaattg aaaatgatga ggctctcttg 960 gatactctta aaatatttag aaaatcggct gaagaaaaaa ccactttgtt aaaaaatatt 1020 tttggtgatt ttgttatgaa tcagggtaag tatgatttag cgcagattta tatttccaga 1080 gaatctttaa atactatttc acggaaatgg accagtgaaa cagatatatt tgaggattca 1140 ttatatgaag tgttaaagaa atcaaaaata gtttctgcct ctgtaaaaaa gaaagatgga 1200 gggtacgctt tccctgagtt tattgcgctt atttatgtga aaagtgctct tgaacaaatt 1260 cctactgaaa aattttggaa ggagcgatat tataaaaata ttggagatgt tttgaataaa 1320 gggtttttga atggtaagga aggtgtctgg ttacaatttt tattgatttt tgattttgaa 1380 tttaattctc tttttgaaag agaaataatt gatgaaaatg gagacaagaa agtggccgga 1440 tataatttgt ttgccaaggg ttttgatgat cttttgaata actttaaata tgatcaaaaa 1500 gctaaggttg ttattaagga ttttgcagat gaggttttac atatttatca gatgggaaaa 1560 tattttgcta ttgaaaagaa acgttcttgg ttggctgatt atgatattga ttcattttat 1620 actgatcctg aaaaaggtta tttgaagttt tatgaaaatg cgtatgaaga gattattcaa 1680 gtttataata aattgcgaaa ttacctaacg aagaaacctt atagtgagga taaatggaaa 1740 ctlaattttg agaatccaac tttagctgat gggtgggaca aaaataaaga agctgataat 1800 tctacagtta ttttgaaaaa ggatggtcgc tattatttag ggttgatggc tcgcgggcga 1860 aataaacttt ttgatgatag aaatttacca aaaattttgg agggcgttga gaatgggaaa 1920 tatgagaaag ttgtatataa gtattttccg gatcaggcaa aaatgtttcc aaaagtttgt 1980 ttttcaacta aaggtttgga gtttttccaa ccttcggagg aagtcattac tatttacaaa 2040 aattctgaat tcaaaaaagg gtatactttt aatgtaagga gtatgcagag gcttattgat 2100 ttttataaag attgtcttgt tagatatgag gggtggcaat gttatgattt tagaaatttg 2160 agaaagacag aagattatcg gaagaatatt gaagagtttt tcagcgatgt tgctatggat 2220 gggtataaaa tatcctttca ggatgtctcg gaaagttata ttaaagagaa aaatcagaat 2280 ggggatttat atttatttga gataaaaaat aaagattgga atgaaggcgc aaatggaaag 2340 aaaaatttgc acactatata ttttgaatct cttttttcgg ctgataatat tgccatgaat 2400 tttcccgtta agttgaatgg acaagcggaa attttttatc ggccaagaac agaggggctg 2460 gagaaagaaa ggataatcac taaaaagggt aatgttttgg aaaaaggaga taaagctttt 2520 cataaaagaa ggtatacgga aaacaaagtt ttttttcatg ttccgattac acttaatcga 2580 acaaaaaaaa atccatttca atttaatgca aaaattaatg attttttggc taaaaattct 2640 gatataaatg ttattggggt cgatcgtggg gagaagcaat tagcatattt ttctgttatt 2700 tcacagagag gcaaaatttt ggataggggt agtttaaatg tgataaatgg agttaattat 2760 gcagagaaat tagaagaaaa agctagaggg cgtgagcagg cgcgtaagga ttggcagcag 2820 attgaaggta ttaaagattt aaagaaggga tatatttctc aggtagttag aaagctagcc 2880 gatttagcaa ttcagtataa tgcgattatt gtttttgaag atttgaacat gcggtttaag 2940 cagattcgtg gaggtattga aaaaagtgtt tatcagcagt tggagaaggc tttgattgat 3000 aaattaactt ttttggttga aaaggaagaa aaagatgtag aaaaggcagg tcatttgtta 3060 aaagcttacc agcttgctgc tccgtttgag acttttcaga aaatgggtaa acaaacgggg 3120 attgtttttt atacacaggc tgcatatact tcacgaattg atcctgttac aggttggcgg 3180 cctcacttgt atttgaaata ttccagtgcg gagaaggcaa aggcggattt attaaaattt 3240 aaaaagataa agtttgtgga tggccggttt gagtttactt atgatattaa gagttttcgt 3300 gaacaaaagg aacatccaaa ggcgactgtc tggacggtgt gttcttgcgt ggagagattt 3360 cgttggaata gatatttaaa tagcaataaa ggtggttatg accattacag tgatgtgacg 3420 aagttcttgg tagagctttt tcaagagtat gggattgatt ttgaaagagg ggatattgtc 3480 gggcaaattg aggttttgga aacgaaggga aatgaaaaat tttttaagaa tttcgttttt 3540 ttctttaatt tgatttgtca gataagaaat actaatgcgt cggagttggc aaaaaaagat 3600 ggaaaagatg attttattct ttcaccggtg gaaccgtttt ttgatagcag aaattcggag 3660 aagtttgggg aggatttgcc aaaaaatggg gatgataatg gggcatttaa tattgcgagg 3720 aaagggcttg ttattatgga taaaattaca aaatttgcag atgagaatgg tgggtgcgag 3780 aagatgaagt ggggagattt gtatgtttct aatgtggagt gggataattt tgtagctaat 3840 aaatga 3846 <210> 15 <211> 3414 <212> DNA <213> Unknown <220> <223> Cas12q <400> 15 atgataaata ttgacgaatt aaaaaattta tataaagttc aaaaaacaat tacttttgaa 60 ttaaaaaata aatgggaaaa taagaatgat gaaaatgata gagttgagtt tttaaagact 120 caagaatggg tggaatcttt attcaaagtt gatgaggaga attttgatga aaaggagtca 180 attccgaact tgttagattt cggccaaaag attgcgagtc ttttttataa gttgagtgaa 240 gatatcgcta ataatcaaat tgatacacgg gttttaaaag tgagcaagtt tttgttggag 300 gagatcgata gaaatcaata tcatgagaaa aaaaataaac caacaaaggt taaggagatg 360 aatccaaata caaataagag ttatattaag gagtataagt tatcagatca aaatacattg 420 tatgttctgt tgaagataat ggaagatgaa gggcggggtt tacaaaaatt tttatatgat 480 aaggcagaca gattaaattt atataatcag aaggtaagaa gagatttcgc tttaaaagaa 540 agtaacgaac agcagaagtt ttcgggtaac gctaattatt acggaaacat aaaattgttg 600 attgattcat tggaagacgc tgttcgtatt attggttatt tcacgtttga tgatcaagca 660 gaaaatgctc aaataaatga attcaagagc gttaagcagg aaatgaataa caatgaagct 720 tcgtatcagg ctttgaaaga ttttgctatt gataacgcaa aaaaagaaat tgaacttaca 780 actctaaatc atagggctgt taacaaggat ccaaaaaaga tacaagaaca gattgaagaa 840 gtggaaaatt ttgaagaaga tataaatcaa ttgaagcacc aaatttctgc gcttaatgat 900 aaaaaatttg atgtagtgtc aagattaaag catgcattaa ttaaaatgtt accggagttg 960 aatttgttag atgctgaaag cgagcaaggt agagaggttc agcaaatata tcaagataaa 1020 aagaatggtt tggaattaga cgattttaag ttcaatttgc ttaaacatca tcaatggcag 1080 aaaaccattt taaatacat taaattagag gttttggttt tacctgattt atatgccgaa 1140 aaaaaaaag agaaggataa agtgtatatt gaaaatttac haaagcgg agaaaggata 1200 agtaaaaagg cacgcgagga gttggggcaag atcgataaaa gagaggaatt taatggtaat 1260 gatgaactaa agaagcgtg gtacgaatac aagattttt gcagagacaa gcgtataa 1320 tccgtggaat tgggcaataa gaatcactg tacaatgcca tcaagcgtga ggttttaagg 1380 cagaaaatgt gtaatcattt tgccgtattg gtgagtgatg gggaagatac atcgccttat 1440 tattattga tattaatttcc caatgaaac agtgatgaaa tgacaggac attcaagag 1500 cttaaagcat ccgaaggaaa ttggaagatg ctcgattata acagattaac ttttaaagct 1560 ttggaaaaat tggcattatt gcgcagctct acatttgaaa ttgcagacca agaactaca 1620 gagaagcta aaaaatttg gagaatat aaaaaagg cgtataaaga ttttaagaat 1680 aaaaaattat tacaagggct atccggtcgc aaagagaag aaaaaaaca agaattgca 1740 aaagaaagtt taaatcgagt tataattat taattcgtt gcattcagtc gttgccggat 1800 AGGCTGCATGCTGTAATCATGGTTGTTATGTTGGAAATGTTGGAAATGTTGGAAATG 120 GCGGAAGAAATTGATAGACAGGTTATCATTGCGCTTGGAAGAATGTAAGCAAGACA 180 CTTATGGAGCTGGAGGCGATGGAAAAAATTAAAGTATTCAAATTGCATAATAAGGAT 240 TTCGAAACACAATCCGAACTTTTTACTTTATATTGGCTT2400 GACGCATATGAATTTGGATAAAGTCAATGTTTGTATTGCCTGAGGTGATTTATATAA 2100 AGAGCCAAAGAAACGCAACTAAAATTATCGAAAGAGATGTAAGTGCAATATTAATAA 2160 CAAAAAATAAATCAATTAAAGAAAAAAATAGATTATTTCAAGATAAACTTTACGCT 2220 TTC A AGCTGGAATT TT ATCC AG AAAACG AAGGTTTGGGTTT TG A AC A AGT CAATGA 2280 TGTGAATAATTTTTGCGGAAGTGATACAGCGTATTATTTGGGTTTGGATAGGGGTGA 2340 GAA TTG GTTACGTTTTGCTTG GTTGA TTCT GATGGGCGGTTG GTTAAAGACGGAGAT 2400 TG A AGTTT A AGAGGTT A ACTATGC GGAT AAATTA AGCA TTTTATTATTC A A AGGT 2460 GAAATAGAATCTACTCAACACAACCTTTTGGAAGCTCGAGACAATATTAACAAGCTAC 2520 aacacggagg ataaagaatc gatgaaatta aactataaaa aattagagtt gaaactaaaa 2580 caacagaatt tgttagcgca ggagtttatt aaaaaagctt attgcggtta tttgatagat 2640 tcaataaatg aaatattacg ggaatatcca aatacgtatc ttgtattaga ggatttggat 2700 atagcaggta aagctgaccc cgaaagcggc atgaccaata aagaacaaaa tttaaataaa 2760 acaatgggtg ccagcgttta tcaagctatt gaaaatgcca tagtaaataa gtttaaatac 2820 cgtactgtta aattatccga tatcaaaggt ttgcaaactg taccgaatgt agtgaaggtg 2880 gaagatttgc gcgaagttaa ggaagtggaa gatggtgagc ataaatttgg tttgataaga 2940 tccgtgaaat caaaggatca aattggcaat attctgtttg tggatgaagg agaaacatct 3000 aatacttgcc cgaattgcgg atttaacagc gattggttta agcgggatgt tgattttgat 3060 ttggagattg tggctactgt aaacggtcag aaaaatgcgg ttatagaaca aaacgacaaa 3120 aagtactgtt ttcccggtga aatttataag ttagaaataa ttaataaaga atacgaaaca 3180 aataaacgga atttagccat gatttttaaa ccgcgcgcaa aagcttgtag aaaatttata 3240 aataataatt tggataagaa tgactatttt tattgcccgt attgcgcttt ttctagcaag 3300 aactgcaata atccaaaatt gcaaaacggt gattttgtgg tatattcggg tgatgatgtg 3360 gcggcataca atgtagcgat cagaggtatt aaccttttaa acaatataaa atag 3414 <210> 16 <211> 3822 <212> DNA <213> Artificial Sequence <220> <223> Cas12a.1 Codon Optimized Version <400> 16 atggggtcaa gtcatcacca ccaccaccac tcaagtggac tagtaccccg tggcagcatg 60 aaagttagca cctgggatag cttcaccaac cagtacccgc tgaccaagac cctgcgtttt 120 gagctgaagc cggtgggtaa aaccctgcag aagatccaag accgtaacct gattaccgag 180 gacgaacagc gtcaaaagga tttcaacaag gttaagaaaa tcatggatgg ttactacaag 240 cagttcatcg aggaatgcct ggaaggcgcg aagatcccgc tgaagaaact ggaggaaaac 300 aacaacgcgt acaccaaact gaagaaagac ccgtataaca agaaactgcg tgaggaatac 360 gcgaagctgc agaaacaact gcgtaaactg atccacgatg agattaacaa gaaagaggaa 420 ttcaagtacc tgtttaagaa agaattcatc aagaaaattc tgccggaatg gctggagaag 480 aaaggtaaga aagaggaact gaaagagatc gaaaagttcg acaaatgggt gacctacttt 540 agcggcttct ttaacaaccg taagaacgtt ttcagcagcg acgagattag caccagcatg 600 atctatcgta ttgtgaacga taacctgccg aaattcctgg acgatgttag ccgttttggt 660 gaaattaccc gttacaagga gttcgacgcg aaccagatcg aggaaaactt tgagagcgaa 720 ctgaacggtg aaaaactgaa ggatttcttt aacctgaaaa acttcaacaa ctgcctgaac 780 caagaaggca ttgagaaatt taacctgatc attggtggca agagcgagga aggtaataac 840 aaaatcaagg gcctgaacga actggtgaac gagctggcgc aaaaacaagc ggacaagaac 900 gagcagaaga aagttcgtaa actgaagctg gcgccgctgt tcaaacaaat cctgagcgat 960 cgtaagagca gcagctttgc gttcgaaaaa tttgaggaaa acaccgaggt gttcgacgcg 1020 atcgatgaat tttatgacaa gattagcctg gagaccctga agaaaatcga agcgaccctg 1080 gagaaactgg aggaaaagga cctggaactg gtttacctga aaaacgatcg ttgcctgacc 1140 ggtatcagcc aggaagtgtt cggcgaccgt gagcgtgttc tgcaagcgct gcgtgaatac 1200 gcgaaaaccg agctgggtct gaagaccgat aagaaaatcg agaagtggat gaagaaaggt 1260 cgttatagca tccacgagat tgaaagcggc ctgaagaaaa tcggtagcac cggccacccg 1320 1380 aaagcgcgta ccgagtatga aaagatcagc gacaagaaaa agaaactgac cgcggaaagc 1440 caagagccga acgtggcgcg tatcaaagcg ctgctggata gcattatgcg tctgtatcac 1500 ttcatcaagc cgctgaacat caacttcaag aacaagaaag agaaggacag cgaagcgctg 1560 gagaccgaca acgattttta caacgacttc gatgaaagct ttgcggagct gggcaacatc 1620 attccgctgt aaaccaagt gcgtaactat gttacccaaa aaccgttcag caccgagaaa 1680 ttcaagctga actttgaaaa cccgaagctg ctgagcggtt gggacaaaaa caaggaaaaa 1740 gattactata gcgtgattct gcgtaaagag gaagctact atctggcgat catgaccccg 1800 aagcagaaaa acgtttttcga cgagctgggaa cgtctgccgg cgggcaaaaa ttacttcgag 1860 aagatcgatt acaagctgct gccgaccccg gaaaagaacc tgccgcgtat cctgttcgcg 1920 aagaaaaaca ttagctttta caagccgagc aaagagatcg aagcgattcg taaccacagc 1980 gcgcacacca aacacggtaa cccgcagaac ggcttcaaga aacgtgactt tcgtctgagc 2040 gattgccaca agatgatcga cttctacaag aaaagcattc agaaacaccc ggaatggaag 2100 gagtatgatt ttcaattcaa gaaaaccgag gactacgtgg atatcagcga attctataaa 2160 gaggtttctg accagggtta caagatcgaa ttcaagaaaa ttagcgagaa atacctgctg 2220 gacctggtgg aggaaggtaa actgtacctg ttccaaatct ggaacaagga tttcagcaag 2280 tacagcgaag gccgtaaaaa cctgcacacc atctattgga aagaactgtt cagcaaggag 2340 aacctgagcg atattaccta taagctgaac ggcgaggcgg aaatctttta ccgtccgaaa 2400 agcatggagc gtaaggttac ccacccgaag aaccagaaaa tcgaaaacaa agacccgatc 2460 aagggtaaga aattcagcaa gttcaagtat gacttcatca agaacaagcg ttacaccgag 2520 gatcgtttct ttttccactg cccgatcacc ctgaactttc aagcgcgtga cggcagcaaa 2580 accatcaaca agcgtgtgaa cgatcacatt cgtgagacca aagacgatat cttcgttctg 2640 agcattgatc gtggtgaacg tcacctggcg tactataccc tgctgaacag caagggtgaa 2700 attcaggagc aaggcagctt taacgtgatc agcgacgata aggagcgtaa acgtgactat 2760 cacgaaaaac tggatgagcg tgaaaaggag cgtgacaagg cgcgtaaaag ctggcagaaa 2820 atcgagacca ttaagaaact gaaggatggc tacctgagcc aaatcgtgca caagattgcg 2880 aaactggcga tcgagaaaaa cgcgatcatt gttctggaag acctgaacct ggatttcaag 2940 cgtggtcgtc tgaagattga gaaacaggtg taccaaaaat tcgaaaagaa actgatcgac 3000 aagctgaact atctggtttt taaagaacgt accgaaaaag aggcgggtgg tagcctgaac 3060 gcgtatcagc tgaccggtaa attcgagggc tttaagaaac tgggcaagga aaccggcatc 3120 atttactatg tgccggcggc gtacaccagc aaaatctgcc cgaagaccgg cttcgttaac 3180 ctgctgcgtc cgaagttcaa gaacatcgaa aaggcgaagg agtttttcaa gaagttcaac 3240 tacatcaagt acgacagcag cgaaggtctg tttgagttca acttcgatta cagcaagttc 3300 GAGAAGAAGA AGAAGAAGAA GAAGAAGAAG AAGAAGAAGA AGAAG 48 GTAACCAAGC TGGTTGGCTT CCCTAACAAG AACAAAAACA ACACTGGGAT ACCAAGGAA 3420 GTGAAACCGA ACGAGAAGCT GAAAATTCTG TTCAAAGAGT ACGGTGTTAG CTTTCAAAG 3480 GACGAAAACA TCATTAGCCA GATCGCGAGC CAAAACAAGA AGCGTTTTTC GAGAACC 3540 ATCAAGATTT CAAGACCATT CTGATGCTGC GTAACAGCC GCAAAGACCC GGAGGAAGAT 3600 TACGTGCTGA GCTGCgttaa ggacgaaaac ggcgagtttt tcgacagccg taaggcgaaa 3660 GATAACGAGC CGAAAGACGC GGATGCGAAC GGCgcgtacc acattggtct gaagggcctg 3720 ATGCTGCTGG AACGTATCAA GGCgaacaaa ggtaagaaaa agctggacct gctgatcagc 3780 CGTAACGATT TCATTAACTT TGCggttgag cgtagcaagt aa 3822 <210> 17 <211> 3903 <212> DNA <213> Artificial Sequence <220> <223> Codon-optimized Cas12p <400> 17 ATGGGATCAA GTCATCACCA CCACCACCAC TCAAGTGGAC TAGTACCCAG GGGAAGCATG 60 aagaagagca ttttcgatca gttcgttaac cagtacgcgc tgagcaagac cctgcgtttc 120 gagctgaaac cggtgggtga aaccggccgt atgctggagg aagcgaaggt tttcgcgaag 180 gatgaaacca ttaagaaaaa gtacgaagcg accaagccgt tctttaacaa actgcaccgt 240 gaattcgtgg aggaagcgct gaacgaggtt gaactggcgg gcctgccgga gtacttcgaa 300 atcttcaagt actggaagcg ttacaaaaag aaattcgaga aggacctgca gaagaaagag 360 aaggaactgc gtaaaagcgt ggttggtttc tttaacgcgc aagcgaagga gtgggcgaag 420 aaatatgaaa ccctgggcgt gaagaaaaag gatgttggtc tgctgttcga ggaaaacgtg 480 tttgcgattc tgaaagaacg ttacggtaac gaggaaggca gccagattgt ggacgagagc 540 accggcaagg atgttagcat cttcgacagc tggaagggtt ttaccggcta tttcatcaaa 600 tttcaggaaa cccgtaagaa cttctacaaa gatgatggta ccgcgaccgc gctggcgacc 660 cgtatcattg atcaaaacct gaaacgtttc tgcgacaacc tgctgatctt tgagagcatt 720 cgtgataaga tcgacttcag cgaggttgaa cagaccatgg gcaacagcat cgataaggtg 780 ttcagcgtta tcttttatag cagctgcctg ctgcaagaag gtatcgactt ttacaactgc 840 gtgctgggtg gtgaaaccct gccgaacggt gaaaagcgtc agggcattaa cgaactgatc 900 aacctgtacc gtcaaaagac cagcgagaaa gttccgttcc tgaagctgct ggacaaacag 960 attctgagcg agaaggaaaa atttatggat gagatcgaaa acgacgaggc gctgctggat 1020 accctgaaga ttttccgtaa agcgcggag gaaaagacca ccctgctgaa aaacatcttc 1080 ggcgattttg tgatgaacca gggtaaatat gacctggcgc aaatctacat tagccgtgaa 1140 agcctgaaca ccattagccg taagtggacc agcgaaaccg atatcttcga agacagcctg 1200 tacgaggtgc tgaaaaagag caaaatcgtg agcgcgagcg ttaaaaagaa agacggtggc 1260 tacgcgttcc cggagtttat cgcgctgatt tatgttaaaa gcgcgctgga acagattccg 1320 accgagaagt tctggaaaga acgttactat aagaacatcg gcgatgtgct gaacaagggt 1380 ttcctgaacg gtaaagaagg cgtttggctg caatttctgc tgatctttga cttcgaattt 1440 aacagcctgt tcgagcgtga aatcattgat gagaacggcg acaagaaagt ggcgggttat 1500 aacctgttcg cgaagggttt tgacgatctg ctgaacaact tcaaatacga ccagaaggcg 1560 aaagtggtta ttaaggattt tgcggacgaa gttctgcaca tttatcaaat gggcaaatac 1620 ttcgcgatcg agaagaaacg tagctggctg gcggactatg atattgacag cttctacacc 1680 gatccggaga agggttacct gaaattttat gaaaacgcgt acgaggaaat cattcaggtt 1740 tataacaagc tgcgtaacta cctgaccaag aaaccgtata gcgaggacaa gtggaaactg 1800 aacttcgaaa acccgaccct ggcggatggt tgggacaaga acaaagaggc ggataacagc 1860 accgtgattc tgaagaaaga cggtcgttac tatctgggcc tgatggcgcg tggtcgtaac 1920 aagctgttcg acgatcgtaa cctgccgaaa atcctggagg gtgttgaaaa cggcaagtac 1980 gaaaaggtgg tttacaagta cttcccggat caggcgaaga tgttcccgaa agtgtgcttt 2040 agcaccaaag gcctggaatt ctttcaaccg agcgaggaag ttatcaccat ttacaagaac 2100 agcgagttca agaaaggtta tacctttaac gtgcgtagca tgcagcgtct gattgatttc 2160 tataaagact gcctggttcg ttacgaaggt tggcaatgct atgattttcg taacctgcgt 2220 aagaccgagg actaccgtaa aaacatcgag gaattcttta gcgatgtggc gatggacggc 2280 tacaagatta gcttccagga cgttagcgag agctatatca aggagaagaa ccaaaacggt 2340 gatctgtacc tgtttgagat caagaacaaa gactggaacg aaggtgcgaa cggcaagaaa 2400 aacctgcaca ccatttattt cgagagcctg tttagcgcgg ataacatcgc gatgaacttc 2460 ccggtgaaac tgaacggcca ggcggagatc ttttaccgtc cgcgtaccga aggtctggag 2520 aaggaacgta tcattaccaa gaaaggcaac gttctggaaa agggtgacaa agcgttccac 2580 aagcgtcgtt acaccgagaa caaagtgttc tttcacgttc cgattaccct gaaccgtacc 2640 aagaaaaacc cgttccaatt taacgcgaag atcaacgact tcctggcgaa aaacagcgat 2700 atcaacgtga ttggtgttga ccgtggcgag aaacagctgg cgtattttag cgtgattagc 2760 caacgtggca agatcctgga ccgtggtagc ctgaacgtga tcaacggcgt taactacgcg 2820 gagaagctgg aggaaaaagc gcgtggtcgt gaacaggcgc gtaaggattg gcagcaaatc 2880 gagggcatta aagacctgaa gaaaggttat attagccagg tggttcgtaa actggcggat 2940 CTGGCGATCCAATAC ACGCGATCATTGTGTTCGAGGA CCTGAACATGCGTTT TAAGCAA 3000 ATTCTGTTGGCGGCGTTCGAGAAAGCGTTTATCAGCAACTG GAAAAGGCGCTGATCGATAAA 3060 CTGACCTTCTGGTGGAGAAGGAAGAAAAGGACGTTGAAAAGGCGGTCACCTGCTGAAA 3120 GCATCCAGCTGGCGGCGCCGTTTGAAACCTTTCAGAAGATGGGTAACA AACC GGCATT 3180 GTGTTTTATACCCAAGCGGCGTACACCAGCCGTATCGATCCGGTTACCGGCTGGCGTCCG 3240 CACCTGTACCTGAAATATAGCAGCGGAAAAGGCGAAAGC GGACCTGCTGAAGTTCAAG 3300 AAAATTAAGTTTGTGGATGGTCGTTCGAGTTTACCTACGACATCAAGAGCTTCCGTGAG 3360 CAGAAGGAACACCCGAAAGCGACCGTGTGGACCGTTTGCA GCTGC GTTGAGC GTTTTCGT 3420 TGGAACC GTTATCTGAACAGCAACAAAGGTGGCTACGATC ACTATAGCGACGTGACCAAG 3480 TTCCTGGTTGAGCTGTTTCAGGAATACGGCATCGACTTCGAACGTGGGTATATTGTGGGC 3540 CAAAATCGAGGTTCTGGAACCAAGGGTAACGAGAAGTTCTT TAAGAAGTTCTGTTTCTTT 3600 TTCAACCTGATCTGCCAGATTCGTAACACCAACGCGAGCGA ACTGGCGAAGAAAGACGGC 3660 aaggacgatt tcattctgag cccggttgag ccgtttttcg atagccgtaa cagcgagaag 3720 ttcggcgaag acctgccgaa aaacggtgac gataacggcg cgtttaacat cgcgcgtaaa 3780 ggtctggtta ttatggataa gatcaccaaa ttcgcggacg agaacggtgg ctgcgaaaag 3840 atgaaatggg gtgacctgta tgtgagcaat gtggagtggg ataactttgt ggcgaataaa 3900 taa 3903 <210> 18 <211> 3471 <212> DNA <213> Artificial Sequence <220> <223> Codon-optimized Cas12q <400> 18 atggggtcct cccatcatca ccaccaccac tcttcaggct tggtaccgcg tggttccatg 60 atcaacatag acgaattgaa aaatttatat aaggtgcaaa agaccatcac tttcgaactt 120 aagaacaagt gggagaacaa aaatgatgag aacgacagag tagagttctt gaagactcag 180 gagtgggtcg aaagcctttt caaggtcgat gaagagaact ttgatgagaa agagtctatc 240 cctaacttgt tagacttcgg acagaagatt gcgtccttgt tttacaagct gagcgaggac 300 atagcgaaca accaaattga tacgcgggta ttgaaagtct cgaaattcct tttagaggaa 360 attgatagaa atcaatacca cgagaaaaaa aacaagccca caaaggtaaa agaatgaat 420 cccaacacaa acaaaagtta tataaaagaa tataagctgt ccgaccaaaa cacactgtac 480 gtgttattaa agataatgga agatgaaggt cggggattac aaaaattttt gtacgataaa 540 gcggaccggt taaacctgta caatcaaaaa gttcggagag acttcgcctt aaaggaatca 600 aatgagcaac aaaaattctc tggaaatgcc aactactatg ggaataaaa gctccttata 660 gatagcttag aagatgcagt ccggatcatt gggtatttca ctttcgacga tcaagcagaa 720 aacgcacaaa tcaatgaatt taagtccgtt aaacaggaaa tgaataataa tgaagcgtct 780 taccaagcac tgaaagactt cgctattgat aacgcaaaaa aagatagaga attgacgacg 840 ttgaaccacc gggcggtcaa caaggatcca aaaaagattc aagaacagat tgaggaagtc 900 960 aagtttgacg tggttagcag attaaagcac gctctttaa aaatgttacc agaactgaat 1020 ttttggatg ctgagtcgga acagggccgt gaagtccagc agatatatca agacaaaaaa 1080 aacgggttgg agcttgatga ctttaaattt aaccttttaa aacatcatca atggcaaaaa 1140 acgatcttca agtatattaa gcttgagggc ttagttctgc cagaccttta cgcggaaaac 1200 aaacaagata aaatcaaggt ttatattgag aattatagac agagtggtga gcgtatttct 1260 aagaaggcga gagaggaatt aggaaaaatc gataaacgcg aagagttcaa tggaaatgac 1320 gaacttaaga aggcatggta tgagtataag gacttctgta gagacaaacg taataagagc 1380 gtggaacttg gcaataagaa gtcgctgtac aatgccataa agcgcgaagt tttgcggcaa 1440 aaaatgtgca accatttcgc tgtgctggtg tccgacggtg aagatacttc cccttattat 1500 tatctgatat taatcccgaa cgagaactcc gatgaaatga atagaacgtt caaggaattg 1560 aaggcctccg aggggaattg gaagatgttg gattacaatc gtctgacctt caaagccttg 1620 gagaaattgg ccctgttacg gtcgtctacc ttcgagatag cggatcagga actgcaagaa 1680 gaggcaaaaa agatctggga ggagtacaag gaaaaggcgt acaaagactt caaaaacaaa 1740 aagttattac agggtttatc gggaagacag cgggaggaga aaaagcaaga attgcaaaag 1800 gagagcctga atagagtaat caattacttg atcagatgca ttcagtcatt gcccgacagc 1860 ggaaaataca actttaactt taaagagcct catcaatacc aatcgcttga agagtttgcc 1920 gaggagattg atcggcaagg ttatcactgt gcttggaaaa acgtttctaa agataaactg 1980 atggaattgg aagcgatgga aaagattaag gttttcaaac ttcataacaa agactttcgc 2040 aaggtaaaac tgaacgactc caagcacaac cctaatcttt ttactttgta ctggttagac 2100 gccatgaatt tggataaggt taacgtccgc ctgttaccgg aagttgacct ttacaagaga 2160 gctaaggaaa cacagctgaa attgttcgaa cgtgatgtga aatgcaatat caataaccaa 2220 aagattaaat ctatcaagga gaagaataga ctgtttcagg acaagttgta tgctagtttt 2280 aagttagagt tttatccaga aaacgaagga ttaggtttcg agcaggtaaa tgacaaggtc 2340 aataacttct gcggtagcga tacggcctat tatcttgggc ttgatcgtgg agagaaagag 2400 cttgttacat tctgcctggt ggactctgat ggccgcctgg taaaaaacgg agactggacc 2460 aagtttaaag aggtgaacta tgccgacaaa ctgaagcaat tctactactc aaaaggcgaa 2520 atagagagta cccaacaaca gctgttagaa gcccgggaca atattaaaca agcgaccaac 2580 acggaagata aggagtccat gaaactgaat tataagaaac tggaactgaa gttaaaacaa 2640 cagaatttgc tggcgcaaga attcataaaa aaagcgtact gcggctacct tatcgatagc 2700 attaatgaga ttctgagaga atatccaaat acttatcttg tcttagagga tttggatatc 2760 gcgggtaaag cggatccaga gtcggggatg actaataaag agcagaactt aaacaaaacg 2820 atgggggctt cagtatacca ggccattgag aatgcgatcg taaataaatt caaatatcgc 2880 accgtgaaat tgtccgatat caagggcctt cagactgtac ctaatgtagt gaaggtcgaa 2940 gacttacggg aagtgaaaga ggttgaagat ggggaacaca agttcgggtt aataagatca 3000 gttaagagca aggatcaaat cggtaacata ctttttgtcg acgaggggga gaccagtaac 3060 acttgtccga attgcggttt taatagtgat tggtttaaac gcgatgttga ttttgactta 3120 gaaatagtcg ctactgtaaa cgggcaaaag aatgccgtga ttgagcaaaa tgacaaaaaa 3180 TACTGTTTCCC GGGCGAAAT ATATAAATTG GAAATCATT AATAAAGAGT ACGAAACAAA C 3240 AAGCGTAATC TTGCCATGAT TTCTAACCTC GGGCCAAAGC GTGCCGTAAT TTATCAAT 3300 AATAATTTAG ATAAGAACGA TTATTTCCTA TTGTCCCTAC TGCGCCTTCT CGTCGAAGAA T 3360 TGTAACAACC CGAAACTGCA GAACGGCGAT TTCGTGGTAT ATTCAAGAGA C GATGTTCCT 3420 GCTTACAATG TTGCTATCAG AGGAATTAAC CTGCTGAACA ATATTAATAG 3471 <210> 19 <211> 36 <212> DNA <213> Unknown (Unknown) <220> <223> Cas12a.1 repeat <400> 19 GTTTAAGGCC TTGACAAAAT TTCTACTGTA GTAGAT 36 <210> 20 <211> 36 <212> DNA <213> Unknown (Unknown) <220> <223> Cas12p repeat <400> 20 CTCGAATATC CCTATTAGAT TTCTACTTTT GTAGAT 36 <210> 21 <211> 36 <212> DNA <213> Unknown (Unknown) <220> <223> Cas12q repeat <400> 21 atctacaaaa gtagaaatta aataggtcta tttgag 36 <210> 22 <211> 30 <212> DNA <213> Unknown (Unknown) <220> <223> Cas12a.1 spacer <400> 22 cgtcgtattg agtgctagta ctggtttgag 30 <210> 23 <211> 32 <212> DNA <213> Unknown (Unknown) <220> <223> Cas12p spacer <400> 23 attaaattac ataatgagcc aacacggcga cc 32 <210> 24 <211> 33 <212> DNA <213> Unknown (Unknown) <220> <223> Cas12q spacer <400> 24 ctagctcctc tacgtcttta ttttcaccct cat 33 <210> 25 <211> 28 <212> DNA <213> Unknown (Unknown) <220> <223> Cas12a.1 - Cas12p target <400> 25 gtggcagctc aaaaattggc tacaaaac 28 <210> 26 <211> 36 <212> DNA <213> Unknown (Unknown) <220> <223> Cas12p direct repeat <400> 26 ctcgaatatc cctattagat ttctactttt gtagat 36 <210> 27 <211> 36 <212> DNA <213> Unknown (Unknown) <220> <223> Cas12q direct repeat <400> 27 ctcaaataga cctatttaat ttctactttt gtagat 36 <210> 28 <211> 36 <212> DNA <213> Unknown (Unknown) <220> <223> Cas12a.1 direct repeat <400> 28 gtttaaggcc ttgacaaaat ttccactgta gtggat 36 <210> 29 <211> 37 <212> DNA <213> Unknown (Unknown) <220> <223> Cas12a.1 direct repeat <400> 29 ggtttaaggc cttgacaaaa tttctcctgt aggagat 37 <210> 30 <211> 36 <212> DNA <213> Unknown (Unknown) <220> <223> Cas12a.1 direct repeat <400> 30 gtttaaggcc ttgacaaaat ttcccctgta ggggat 36 <210> 31 <211> 36 <212> DNA <213> Unknown (Unknown) <220> <223> Cas12p direct repeat <400> 31 atctacaaaa gtagaagtct aatagggaca ttcgag 36 <210> 32 <211> 36 <212> DNA <213> Unknown (Unknown) <220> <223> Cas12p direct repeat <400> 32 atctacaaaa gtagaaagct aatagggcta ttcgag 36 <210> 33 <211> 36 <212> DNA <213> Unknown (Unknown) <220> <223> Cas12p direct repeat <400> 33 atctacaaaa gtagaaggct aatagggcca ttcgag 36 <210> 34 <211> 36 <212> DNA <213> Unknown (Unknown) <220> <223> Cas12p direct repeat <400> 34 ctcgaatatc cctattagat ttcgactttt gtcgat 36 <210> 35 <211> 36 <212> DNA <213> Unknown (Unknown) <220> <223> Cas12p direct repeat <400> 35 ctcgaatatc cctattagat ttctcctttt ggagat 36 <210> 36 <211> 36 <212> DNA <213> Unknown (Unknown) <220> <223> Cas12p direct repeat <400> 36 ctcgaatatc cctattagat ttcggctttt gccgat 36 <210> 37 <211> 36 <212> DNA <213> Unknown (Unknown) <220> <223> Cas12q direct repeat <400> 37 atctacaaaa gtagaaattg aataggtcta ttcgag 36 <210> 38 <211> 36 <212> DNA <213> Unknown (Unknown) <220> <223> Cas12q direct repeat <400> 38 atctacaaaa gtagaaatta aagaggtctc tttgag 36 <210> 39 <211> 36 <212> DNA <213> Unknown (Unknown) <220> <223> Cas12q direct repeat <400> 39 atctacaaaa gtagaaattg ggtaggtcta cccgag 36 <210> 40 <211> 36 <212> DNA <213> Unknown (Unknown) <220> <223> Cas12q direct repeat <400> 40 ctcaaataga cctatttaat ttccactttt gtggat 36 <210> 41 <211 > 36 <212> DNA <213> Unknown (Unknown) <220> <223> Cas12q direct repeat <400> 41 ctcaaataga cctatttaat ttctcctttt ggagat 36 <210> 42 <211 > 36 <212> DNA <213> Unknown (Unknown) <220> <223> Cas12q direct repeat <400> 42 ctcaaataga cctatttaat ttcccctttt ggggat 36 <210> 43 <211 > 487 <212> DNA <213> Artificial Sequence (Artificial Sequence) <220> <223> KPC gB template 1 sequence <400> 43 ttcaagggct ttcttgctgc cgctgtgctg gctcgcagcc agcagcaggc cggcttgctg 60 gacacaccca tccgttacgg caaaaatgcg ctggttccgt ggtcacccat ctcggaaaaa 120 TATCTGACAACAGGCATGACGGTGGCGGAGCTGTCCGCGGCCGCCGTGCAATACAGTGAT 180 AACGCCGCCGCCAATTGTGTGCTGAAGGAGTTGGGCGGCCCGGCCGGGCTGACGGCCTTC 240 ATGCGCTCTATCGGCATACCAACGTTCCTCTGGACCCTGGGAGCTGGAGCTGAAC TCC 300 GCCATCCCAGGCATGCACGCGACACCTCATCGCCGCCGCCGTGACGGAAAGCTTA CAA 360 AAACTGACACTGGGCTCTGC ACTGGCTGC GCCGCAGCGGC AGCAGTTTGT TGATTGGCTA 420 AAGGGAAACACGACC GGCACCACC GC ATCCGC GCGGGTGCCGGCAG ACTGGGCAGTC 480 GGAGACA 487 <210> 44 <211> 480 <212> DNA <213> Artificial Sequence <220> <223> NDM gB Template 1 Sequence <400> 44 CCAAATTAAGATCATCTATTTACTAGGCCTC GCATTTGCGGGGTTTTAA TGCTGAATAA 60 AAGGAAAAC TTGATGGAATTGCCCAATATT ATGCACCCGGTCGC GAAGCTGAGC ACCGCA 120 TTAGCCGCTGCATTGATGCTGAGCGG GTGCATGCCC GGTGAAATCCGCCC GACGATTGGC 180 240. cagcaaatgg aaactggcga ccaacggttt ggcgatctgg ttttccgcca gctcgcaccg aatgtctggc agcacacttc ctatctcgac atgccggggtt tcggggcagt cgcttccaac ggttttgatcg tcagggatgg cggccgcgtg ctggtggtcg ataccgcctg gaccgatgac 360 cagaccgccc agatcctcaa ctggatcaag caggagatca acctgccggt cgcgctggcg gtggtgactc acgcgcatca ggacaagatg ggcggtatgg acgcgctgca tgcggcgggg 480 <210> 45 <211> 125 <212> DNA <213> Artificial Sequence <220> <223> OXA gBlock 1-year-old <400> 45 cgaagccaat ggtgactata ttattcgggc taaaactgga tactcgacta gaatcgaacc taagattggc tggtgggtcg gttgggttga acttgatgat aatgtgtggt tttttgcgat 125 gates <210> 46 <211> 490 <212> DNA <213> Artificial Sequence <220> <223> MecA gBlock1processor <400> 46 TACAACCTTCAACCTAGGTTC AACTCAA AAAAT ATTAACAGCA ATGATTGGGTT AAATAACAA 60 AACATTAGAC GATAAAACAA GTTATAAAAT CGATGGTAAA GGTTGGCAAA AAGATAAATC 120 TTGGGGTG GTTACAACGT TACAAGATAT GAAGTGGTAA ATGGTAATAT CGACTTAAAC A 180 AGCAATAGAA T CATCAGATA ACATTTTCTT TGCTAGAGTA GC ACTC GAAT TAGGCAGTA A 240 GAAATTTGAA AAAGG CATGAAAAACTAGGTGTTGGTGAA GATATACC AAGTGATTA TCC 300 ATTTTATAAT GCTCAAATTT CAAACAAAAA TT TAGATAAT GAAATATTAT TAGCTGATTC 360 AGGTTACGGA CAAGGTGAAA TACTGATTAACCAGTACAG ATCCTTTCAA TCTATAGCGC 420 ATTAGAAAAT ATGGCAATATT AACGCACCTC ACTTATTA AAAGACACGA AAAACAAAGT 480 TTGGAAGAAA 490 <210> 47 <211> 232 <212> DNA <213> Artificial Sequence <220> <223> hHPRT1 template sequence <400> 47 CTCTGTATGT TATATGTCAC ATTTTGTAAT TAACAGCTTG CTG GT GAAAAAGAC CCCACG 60 aagtgttgga tataagccag actgtaagtg aattactttt tttgtcaatc atttaaccat 120 ctttaaccta aaagagtttt atgtgaaatg gcttataatt gcttagagaa tatttgtaga 180 gaggcacatt tgccagtatt agatttaaaa gtgatgtttt ctttatctaa at 232 <210> 48 <211> 100 <212> RNA <213> Artificial Sequence <220> <223> DENV ssRNA target template sequence <400> 48 ugacgaagac caugcucacu ggacagaagc aaaaaugcug cuggacaaca ucaacacacc 60 agaagggauu auaccagcuc ucuuugaacc agaaagggag 100 <210> 49 <211> 100 <212> RNA <213> Artificial Sequence <220> <223> ZIK ssRNA target template sequence <400> 49 ccacacugga acaacaaaga agcacuggua gaguucaagg acgcacaugc caaaaggcaa 60 acugucgugg uucuagggag ucaagaagga gcaguucaca 100 <210> 50 <211> 100 <212> RNA <213> Artificial Sequence <220> <223> HANT ssRNA target template sequence <400> 50 agaggcaacu ugcagauuug guggcagcuc aaaaauuggc uacaaaacca guugauccaa 60 cagggcuuga gccugaugau caucuaaagg aaaaaucauc 100 <210> 51 <211> 23 <212> DNA <213> Artificial Sequence <220> <223> Target sequence KPC 1 <400> 51 ttgctgaagg agttgggcgg ccc 23 <210> 52 <211> 23 <212> DNA <213> Artificial Sequence <220> <223> Target sequence NDM 1 <400> 52 gcgatctggt tttccgccag ctc 23 <210> 53 <211> 21 <212> DNA <213> Artificial Sequence <220> <223> Target sequence Ctrol + hHPRT1 1 <400> 53 ggttaaagat ggttaaatga t 21 <210> 54 <211> 23 <212> DNA <213> Artificial Sequence <220> <223> Target sequence S16 cntl Escherichia coli 1 <400> 54 cagtagttat ccccctccat cag 23 <210> 55 <211> 23 <212> DNA <213> Artificial Sequence (Artificial Sequence) <220> <223> Target sequence DENV1 <400> 55 cttctgtcca gtgagcatgg tct 23 <210> 56 <211> 23 <212> DNA <213> Artificial Sequence (Artificial Sequence) <220> <223> Target sequence DENV2 <400> 56 tggttcaaag agagctggta taa 23 <210> 57 <211> 23 <212> DNA <213> Artificial Sequence (Artificial Sequence) <220> <223> Target sequence ZIK1 <400> 57 ggcatgtgcg tccttgaact cta 23 <210> 58 <211> 23 <212> DNA <213> Artificial Sequence (Artificial Sequence) <220> <223> Target sequence ZIK2 <400> 58 ccttttggca tgtgcgtcct tga 23 <210> 59 <211> 23 <212> DNA <213> Artificial Sequence <220> <223> Target sequence OXA1 <400> 59 agcccgaata atatagtcac cat 23 <210> 60 <211> 23 <212> DNA <213> Artificial Sequence <220> <223> Target sequence OXA1b <400> 60 agcccgaata atatagtcgc cat 23 <210> 61 <211> 23 <212> DNA <213> Artificial Sequence <220> <223> Target sequence HANTAndes1 <400> 61 gtggcagctc aaaaattggc tac 23 <210> 62 <211> 23 <212> DNA <213> Artificial Sequence <220> <223> Target sequence HANTAndes2 <400> 62 gatgatcatc aggctcaagc cct 23 <210> 63 <211> 23 <212> DNA <213> Artificial Sequence <220> <223> Target sequence MecA1 <400> 63 tctttttgcc aacctttacc atc 23 <210> 64 <211> 9165 <212> DNA <213> Artificial Sequence <220> <223> Cas12a.1 expression vector sequence <400> 64 tggcgaatgg gacgcgccct gtagcggcgc attaagcgcg gcgggtgtgg tggttacgcg 60 cagcgtgacc gctacacttg ccagcgccct agcgcccgct cctttcgctt tcttcccttc 120 ctttctcgcc acgttcgccg gctttccccg tcaagctcta aatcgggggc tccctttagg 180 gttccgattt agtgctttac ggcacctcga ccccaaaaaa cttgattagg gtgatggttc 240 acgtagtggg ccatcgccct gatagacggt ttttcgccct ttgacgttgg agtccacgtt 300 ctttaatagt ggactcttgt tccaaactgg aacaacactc aaccctatct cggtctattc 360 ttttgattta taagggattt tgccgatttc ggcctattgg ttaaaaaatg agctgattta 420 acaaaaattt aacgcgaatt ttaacaaaat attaacgttt acaatttcag gtggcacttt 480 tcggggaaat gtgcgcggaa cccctatttg tttatttttc taaatacatt caaatatgta 540 tccgctcatg aattaattct tagaaaaact catcgagcat caaatgaaac tgcaatttat 600 tcatatcagg attatcaa...
Claims
1. An engineered system, the engineered system comprising: a. Two types of type V CRISPR-Cas RNA-guided endonuclease proteins; and b. Single guide RNA (gRNA) The gRNA and the two types of CRISPR-Cas RNA-guided endonuclease proteins are not naturally co-occurring. The gRNA is capable of heterozygous for target sequences in the target DNA. The gRNA is capable of forming a complex with the two types of CRISPR-Cas RNA-guided endonuclease proteins. The two types of CRISPR-Cas RNA-guided endonuclease proteins have paracleaning activity and can paraclean polynucleotides in the absence of tracrRNA. The type 2 V CRISPR-Cas RNA-guided endonuclease protein described therein contains the amino acid sequence of SEQ ID NO: 4, and the gRNA described therein contains the scaffold sequence of SEQ ID NO:
117.
2. The engineered system of claim 1, wherein the type V CRISPR-Cas RNA-guided endonuclease protein is capable of paracleaving single-stranded RNA.
3. The engineered system of claim 1, wherein the type V CRISPR-Cas RNA-guided endonuclease protein is capable of paracleaving single-stranded DNA / RNA hybrids.
4. An engineered system, said engineered system comprising: a. The Cas12p protein or the nucleic acid encoding the Cas12p protein; as well as b. Cas12p gRNA or nucleic acid encoding Cas12p gRNA, The Cas12p gRNA and the Cas12p protein are not naturally coexisting. The Cas12p gRNA is capable of heterozygosing with the target sequence in the target DNA, and the Cas12p gRNA is capable of forming a complex with the Cas12p protein. The Cas12p protein contains the amino acid sequence of SEQ ID NO: 4, and the Cas12p gRNA contains the scaffold sequence of SEQ ID NO:
117.
5. The engineered system of claim 4, wherein the system comprises: a. Cas12p protein; as well as b. Cas12p gRNA.
6. The engineered system of claim 4, wherein the system comprises: a. The nucleic acid encoding the Cas12p protein; as well as b. Nucleic acid encoding Cas12p gRNA.
7. The engineered system of claim 4, wherein the target sequence is a human sequence.
8. The engineered system of claim 4, wherein the target sequence is a sequence of a non-human primate.
9. The engineered system of claim 4, wherein the target sequence is a bacterial or viral sequence.
10. The engineered system of claim 4, wherein the Cas12p protein is a catalytically active Cas12p protein.
11. The engineered system of claim 4, wherein the Cas12p protein is a catalytically inactivated Cas12p protein.
12. The engineered system of claim 4, wherein the Cas12p protein contains nicking enzyme activity.
13. An engineered single-molecule gRNA comprising a scaffold sequence of SEQ ID NO: 117 and a spacer sequence capable of heterozygous for a target sequence in a target DNA.
14. The engineered single-molecule gRNA of claim 13, wherein the target DNA comprises viral DNA, plant DNA, fungal DNA, or bacterial DNA.
15. The engineered single-molecule gRNA of claim 13, wherein the target sequence is a target sequence, wherein the target is selected from carbapenem-hydrolyzing class A β-lactamases, metallo-β-lactamases, oxacillin-hydrolyzing class D β-lactamases, PBP2a family β-lactam resistance peptidoglycan transpeptidases, vancomycin resistance genes vanA / B, dengue virus, Zika virus, chikungunya virus, adenovirus, coronavirus, human metapneumovirus, human rhinovirus, enterovirus, influenza A virus, influenza B virus, parainfluenza virus, respiratory syncytial virus, Bordetella parapertussis (… Bordetella parapertussis Bordetella pertussis ( ), Bordetella pertussis Mycoplasma pneumoniae ( ) Mycoplasma pneumoniae Human immunodeficiency virus, herpes simplex virus 1, herpes simplex virus 2, hepatitis A virus, hepatitis B virus, hepatitis C virus, Treponema pallidum ( Treponema pallidum Chlamydia ( ) Chlamydia spp. ), Neisseria gonorrhoeae ( Neisseria gonorrhoeae ), poxvirus, parvovirus, Geminiviridae, Dwarfviridae, Algal DNAviridae, Hantavirus, sex-determining gene, hypoxanthine phosphoribosyltransferase 1, and 16S Escherichia coli ( Escherichia coli ).
16. The engineered single-molecule gRNA of claim 13, wherein the target is a coronavirus.
17. The engineered single-molecule gRNA of claim 13, wherein the target is the SARS-CoV-2 virus.
18. The engineered single-molecule gRNA of claim 13, wherein the target DNA is cDNA and has been obtained by reverse transcription.
19. The use of the system of claim 4 in the preparation of a kit or kit reagent for modifying target DNA, wherein the modification comprises contacting the target DNA with the system of claim 4, wherein the gRNA is hybridized with the target sequence, thereby causing modification of the target DNA, wherein the target DNA is ex vivo DNA.
20. The application according to claim 19, wherein the target DNA is extrachromosomal DNA.
21. The application according to claim 19, wherein the target DNA is part of a chromosome.
22. The application according to claim 19, wherein the target DNA is part of an in vitro chromosome.
23. The application according to claim 19, wherein the target DNA is extracellular.
24. The application according to claim 19, wherein the target DNA is intracellular.
25. The application of claim 24, wherein the target DNA comprises a gene and / or its regulatory region.
26. The application according to claim 23 or 24, wherein the cell is selected from the group consisting of: archaea cells, bacterial cells, and / or eukaryotic cells, wherein the eukaryotic cell is a eukaryotic unicellular organism, plant cell, algal cell, and / or animal cell, wherein the animal cell is an invertebrate cell and / or vertebrate cell, wherein the vertebrate cell is a fish cell, frog cell, bird cell, and / or mammalian cell, wherein the mammalian cell is a pig cell, cow cell, goat cell, sheep cell, rodent cell, non-human primate cell, or human cell, wherein the rodent cell is a rat cell or mouse cell, and wherein the invertebrate cell and / or vertebrate cell is a somatic cell, germ cell, or stem cell.
27. The application of claim 19, wherein the modification comprises introducing double-strand breaks in the target DNA.
28. The application of claim 19, wherein the contact occurs under conditions that allow for non-homologous end bonding or homologous directional repair.
29. The application of claim 19, wherein the modification comprises contacting the target DNA with a donor polynucleotide, wherein the donor polynucleotide, a portion of the donor polynucleotide, a copy of the donor polynucleotide, or a portion of a copy of the donor polynucleotide is integrated into the target DNA.
30. The application of claim 26, wherein the application does not involve contacting the cell with a donor polynucleotide, or wherein the target DNA is modified such that nucleotides are missing from the target DNA.
31. The application of Cas12p protein, Cas12p gRNA, and labeled detectors in the preparation of reagents or diagnostic kits for detecting target DNA in samples, wherein: i. The Cas12p protein contains the amino acid sequence of SEQ ID NO: 4; ii. The Cas12p gRNA contains a spacer sequence capable of heterozygous for a target sequence in the target DNA, and the Cas12p gRNA contains the scaffold sequence of SEQ ID NO: 117; iii. The labeled detector is not heterozygous for the spacer sequence of the gRNA; iv. The labeled detector generates a detectable signal upon cleavage by the Cas12p protein; and v. The labeled detector contains a pair of fluorescent emitting dyes.
32. The application of claim 31, wherein the labeled detector comprises labeled single-stranded DNA.
33. The application of claim 31, wherein the labeled detector comprises labeled RNA.
34. The application according to claim 33, wherein the labeled RNA is a single-stranded RNA.
35. The application of claim 31, wherein the labeled detector comprises a labeled single-stranded DNA / RNA chimera.
36. The application of claim 31, wherein the labeled detector comprises one or more modified nucleotides.
37. The application of claim 31, comprising contacting the sample with a precursor gRNA array, wherein the Cas12p protein cleaves the precursor gRNA array to produce the gRNA.
38. The application according to claim 31, wherein the target DNA is single-stranded.
39. The application according to claim 31, wherein the target DNA is double-stranded.
40. The application according to claim 31, wherein the target DNA is viral DNA, plant DNA, fungal DNA, or bacterial DNA.
41. The application according to claim 40, wherein the target DNA is a target DNA sequence, wherein the target is selected from carbapenem-hydrolyzing class A β-lactamases, metallo-β-lactamases, oxacillin-hydrolyzing class D β-lactamases, PBP2a family β-lactam resistance peptidoglycan transpeptidases, vancomycin resistance genes vanA / B, dengue virus, Zika virus, chikungunya virus, adenovirus, coronavirus, human metapneumovirus, human rhinovirus, enterovirus, influenza A virus, influenza B virus. Viruses, parainfluenza viruses, respiratory syncytial virus, Bordetella parapertussis, Bordetella pertussis, Mycoplasma pneumoniae, human immunodeficiency virus, herpes simplex virus 1, herpes simplex virus 2, hepatitis A virus, hepatitis B virus, hepatitis C virus, Treponema pallidum, Chlamydia, Neisseria gonorrhoeae, poxvirus, parvovirus, Geminiviridae, Dwarfviridae, Algal DNAviridae, Hantavirus, sex-determining gene, hypoxanthine phosphoribosyltransferase 1, and 16S Escherichia coli.
42. The application of claim 41, wherein the target is a coronavirus.
43. The application according to claim 42, wherein the target is the SARS-CoV-2 virus.
44. The application according to claim 31, wherein the target DNA is cDNA and has been obtained by reverse transcription.
45. The application according to claim 31, wherein the target DNA is derived from human cells.
46. The application according to claim 45, wherein the target DNA is human fetal or cancer cell DNA.
47. The application of claim 31, wherein the sample comprises DNA from cell lysate.
48. The application according to claim 31, wherein the sample comprises cells.
49. The application according to claim 31, wherein the sample is a urine sample, blood sample, serum sample, plasma sample, lymph sample, cerebrospinal fluid sample, saliva sample, nasopharyngeal sample, oropharyngeal sample, nasopharyngeal / oropharyngeal sample, aspirate sample, or biopsy sample.
50. The application of claim 31, wherein the detectable signal determines the amount of the target DNA present in the sample.
51. The application of claim 50, wherein the detectable signal comprises one or more of the following: vision-based detection, sensor-based detection, color detection, gold nanoparticle-based detection, fluorescence polarization, colloidal phase transition / dispersion, electrochemical detection, and semiconductor-based sensing.
52. The application according to claim 31, wherein the labeled detector comprises modified nucleobases, modified sugar moieties, and / or modified nucleic acid linkages.
53. The application according to claim 31 further includes the preparation of a kit or diagnostic kit for detecting positive control target DNA in a positive control sample using positive control gRNA, wherein: i. The Cas12p protein contains the amino acid sequence of SEQ ID NO: 4; ii. The positive control gRNA comprises: a region that binds to the Cas12p protein and a positive control spacer sequence that is heterozygous for the positive control target DNA; iii. The labeled detector is not heterozygous for the positive control spacer sequence of the positive control gRNA; as well as iv. The detectable signal is generated by cleaving the labeled detector by the Cas12p protein, wherein the detectable signal detects the positive control target DNA.
54. The application of claim 53, wherein the detectable signal is detectable for less than 240 minutes.
55. The application of claim 53 further comprises amplifying the target DNA in the sample by means of: loop-mediated isothermal amplification (LAMP), helicase-dependent amplification (HDA), recombinase polymerase amplification (RPA), strand substitution amplification (SDA), sequence-based amplification (NASBA), transcription-mediated amplification (TMA), nicking enzyme amplification reaction (NEAR), rolling circle amplification (RCA), multiple substitution amplification (MDA), branching (RAM), circular helicase-dependent amplification (cHDA), single primer isothermal amplification (SPIA), signal-mediated RNA amplification technology (SMART), self-sustaining sequence replication (3SR), genomic exponential amplification reaction (GEAR), or isothermal multiple substitution amplification (IMDA).
56. The application of claim 53, wherein the target DNA in the sample is present at a concentration of less than 100 μM.
57. A pharmaceutical composition comprising any one of the engineered systems of claims 1 to 12, and optionally a pharmaceutically acceptable carrier.
58. A composition comprising any one of the engineered systems of claims 1 to 12, and optionally comprising a nucleic acid stabilizing buffer and / or a protein stabilizing buffer.
59. A pharmaceutical composition comprising any one of the single-molecule gRNAs of claims 13 to 18, and optionally a pharmaceutically acceptable carrier.
60. A composition comprising any one of the single-molecule gRNAs of claims 13 to 18, and optionally a nucleic acid stabilization buffer and / or a protein stabilization buffer.
61. A DNA polynucleotide comprising a nucleotide sequence encoding the following: a gRNA according to any one of claims 1 to 3, a Cas12p gRNA according to any one of claims 4 to 6, or an engineered single-molecule gRNA according to any one of claims 13 to 18.
62. A recombinant expression vector comprising the DNA polynucleotide of claim 61.
63. The recombinant expression vector of claim 62, wherein the nucleotide sequence encoding a single gRNA is operatively linked to a promoter.
64. A host cell comprising the DNA polynucleotide of claim 61.
65. A kit comprising one or more components of any one of the engineered systems of claims 1 to 12.
66. The kit of claim 65, wherein one or more components are lyophilized.
67. The kit of claim 65, wherein one or more components comprise Cas12p, a labeled RNA reporter, and gRNA targeting SARS-CoV-2, wherein the gRNA comprises the scaffold sequence SEQ ID NO: 117.
Citation Information
Patent Citations
Type V CRISPR / Cas effector proteins for cleaving ssDNAs and detecting target DNAs
US10253365B1
Novel crispr enzymes and systems
US20160208243A1
CRISPR enzymes and systems
US9790490B2
Compositions and methods for modifying genomes
CN109312316A
Nucleic acid composition for detecting novel coronavirus COVID-19 and application
CN111363860A