Systems, methods and compositions comprising micro-CRISPR nucleases for gene editing and for programmable gene activation and inhibition
Patent Information
- Application Number
- JP2023577655
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-06-17
- Filing Date
- 2022-06-16
- Publication Date
- 2025-06-23
AI Technical Summary
Existing CRISPR nucleases, such as Cas9 and Cas12, are too large for efficient delivery and use in mouse models and mammalian genome editing, limiting their application in gene editing, activation, and inhibition due to size constraints and the need for precise assembly in eukaryotic cells.
Development of miniature CRISPR nucleases, such as Cas12m and Cas12f, with lengths less than 1000 amino acids, guided by specific guide RNAs, allowing for compact delivery and versatile applications in gene editing and activation/inhibition, including fusion with effector domains.
The miniature CRISPR nucleases enable efficient gene editing and activation/inhibition in mammalian cells, facilitating delivery via AAV vectors and enhancing the versatility of genome editing tools for various cellular contexts.
Smart Images

Figure 00000141_0000 
Figure 00000141_0001 
Figure 00000142_0000
Abstract
Description
[Technical field]
[0001] This application claims the benefit of priority to U.S. Provisional Patent Application No. 63 / 211,610, filed June 17, 2021, the entirety of which is incorporated herein by reference.
[0002] The subject matter disclosed herein is generally directed to systems, methods and compositions comprising miniature CRISPR nucleases for gene editing and for programmable gene activation and inhibition. [Background technology]
[0003] Cluster Regularly Interspaced Short Palindromic Repeat (CRISPR)-associated (Cas) nuclease systems are widely used as genome editing tools. Cas9 and Cas12 are two examples of nucleases frequently used in CRISPR-Cas systems to edit genomes. These nucleases are generally over 1000 amino acids long and can be guided by guide RNAs to edit single- or double-stranded DNA targets that are near short sequences called protospacer adjacent motifs (PAMs). However, although these nucleases offer great flexibility, their size poses a considerable barrier to their use. For example, gene editing techniques based on these nucleases, as well as programmable gene activation and inhibition techniques, cannot generally be introduced into mouse models using common methods such as adeno-associated vectors (AAV) due to the large size of the nucleases. Furthermore, the development of effective gene and cell therapies requires genome editing tools that can meet the demands for reduced payload size and effective integration of diverse and large sequences, regardless of cell type or active repair pathway. CRISPR-associated transposases, such as Cas12k or type I-F directed Tn7 systems, allow programmable integration into bacteria without the need for repair pathway-dependent editing, but still need to be reconstituted in eukaryotic cells for mammalian genome editing. The difficulty in reconstituting these systems may be due to the vast number of proteins (4-7 proteins) that must be precisely expressed and delivered to the nucleus for accurate assembly and DNA targeting. Optimal editing has also been reported by programmable gene editing independent of DNA repair pathways, but is limited to base substitutions or small deletions and insertions (approximately <50 bp). Summary of the Invention [Problem to be solved by the invention]
[0004] Thus, there is a need for smaller, more compact CRISPR nucleases for gene editing, programmable gene activation and inhibition, and new applications. Smaller, more compact CRISPR nucleases can simplify delivery and broaden applications, and the additional space provided by such nucleases can allow fusion with effector domains. [Means for solving the problem]
[0005] (Abstract) The present disclosure provides systems, methods and compositions comprising miniature CRISPR nucleases for gene editing and for programmable gene activation and inhibition.
[0006] In one aspect, the disclosure relates to a composition comprising a target-specific nuclease comprising an amino acid sequence having 70% identity to an amino acid sequence selected from the group consisting of SEQ ID NOs: 1-9 and a guide RNA (gRNA), wherein the target comprises a DNA target. In some embodiments, the DNA target can be single-stranded DNA. In some embodiments, the DNA target can be double-stranded DNA. In some embodiments, the target-specific nuclease can have a length of less than about 1000 amino acids. In some embodiments, the target-specific nuclease can have a length of less than about 900 amino acids. In some embodiments, the target-specific nuclease can have a length of less than about 800 amino acids. In some embodiments, the amino acid sequence can be SEQ ID NO: 1. In some embodiments, the target-specific nuclease can comprise an amino acid sequence having 90% identity to the amino acid sequence of SEQ ID NO: 1, or an amino acid sequence having 95% identity to the amino acid sequence of SEQ ID NO: 1, or an amino acid sequence having 98% identity to the amino acid sequence of SEQ ID NO: 1, or an amino acid sequence having 99% identity to the amino acid sequence of SEQ ID NO: 1. In some embodiments, the nuclease can be the amino acid sequence of SEQ ID NO:1.
[0007] In some embodiments, the target-specific nuclease can be selected from the group consisting of Cas12m, Cas12f, and variants thereof, and optionally the target-specific nuclease can be PsaCas12f.
[0008] In some embodiments, the gRNA can be a single guide RNA (sgRNA) or a dual guide (dgRNA). In some embodiments, the gRNA can be an sgRNA, and the sgRNA can comprise a nucleic acid sequence having 75% identity to a nucleic acid sequence selected from the group consisting of SEQ ID NOs: 20-43 and 61-79. In some embodiments, the gRNA can have a spacer region whose sequence comprises a length of about 17 to about 53 nucleotides (nt), optionally, the sequence can comprise a length of about 29 to about 53 nt, optionally, the sequence can comprise a length of about 40 to about 50 nt, or optionally, the sequence can comprise a length of about 22 nt. In some embodiments, the gRNA can have a direct repeat region whose sequence comprises a length of about 20 to about 29 nt. In some embodiments, the gRNA can have a tracrRNA region whose sequence comprises a length of about 27 to about 35 nt.
[0009] In some embodiments, the DNA target may be within a cell. In some embodiments, the cell may be a prokaryotic cell. In some embodiments, the cell may be a eukaryotic cell. In some embodiments, the eukaryotic cell may be a mammalian cell. In some embodiments, the mammalian cell may be a human cell.
[0010] In some embodiments, the amino acid sequence can specifically bind to a protospacer adjacent motif (PAM). In some embodiments, the PAM can be selected from the group consisting of NNNNGATT, NNNNGNNN, NNG, NG, NGAN, NGNG, NGAG, NGCG, NAAG, NGN, NRN, NNGRRN, NNNRRT, TTTN, TTTV, TYCV, TATV, TYCV, TATV, TTN, KYTV, TYCV, TATV, TBN, and any variants thereof, and any combination thereof.
[0011] In another aspect, a nucleic acid molecule encoding a target specific nuclease is contemplated.
[0012] In another embodiment, a nucleic acid molecule encoding a guide RNA is contemplated.
[0013] In another embodiment, one or more vectors comprising a nucleic acid molecule encoding a target-specific nuclease and / or a guide RNA are contemplated.
[0014] In another aspect, a target-specific nuclease comprising an amino acid sequence having 70% identity to an amino acid sequence selected from the group consisting of SEQ ID NOs: 1-19 (wherein the target comprises DNA) and a guide RNA, or a cell comprising a nucleic acid molecule encoding the target-specific nuclease, or a cell comprising a nucleic acid molecule encoding a gRNA, or a cell comprising one or more vectors comprising a nucleic acid molecule encoding a target-specific nuclease and / or a guide RNA are contemplated. In some embodiments, the cell may be a prokaryotic cell. In some embodiments, the cell may be a eukaryotic cell. In some embodiments, the eukaryotic cell may be a mammalian cell. In some embodiments, the mammalian cell may be a human cell.
[0015] In another aspect, a method of inserting or deleting one or more base pairs in DNA is contemplated, comprising cleaving the DNA with a target specific nuclease at a target site, where the cleavage creates overhangs at both ends of the DNA, inserting nucleotides complementary to the overhanging nucleotides at both ends of the dsDNA or removing the overhanging nucleotides at both ends of the DNA, and ligating the dsDNA ends together, thereby inserting or deleting one or more base pairs into or from the dsDNA, wherein the nuclease comprises an amino acid sequence having 70% identity to an amino acid sequence selected from the group consisting of SEQ ID NOs: 1-19, and the target specificity of the target specific nuclease is provided by a guide RNA (gRNA). In some embodiments, the target specific nuclease can have a length of less than about 1000 amino acids. In some embodiments, the target specific nuclease can have a length of less than about 900 amino acids. In some embodiments, the target specific nuclease can have a length of less than about 800 amino acids. In some embodiments, the amino acid sequence can be SEQ ID NO: 1.
[0016] In some embodiments, the target specific nuclease can comprise an amino acid sequence that is 90% identical to the amino acid sequence of SEQ ID NO:1. In some embodiments, the target specific nuclease can comprise an amino acid sequence that is 95% identical to the amino acid sequence of SEQ ID NO:1. In some embodiments, the target specific nuclease can comprise an amino acid sequence that is 98% identical to the amino acid sequence of SEQ ID NO:1. In some embodiments, the target specific nuclease can comprise an amino acid sequence that is 99% identical to the amino acid sequence of SEQ ID NO:1. In some embodiments, the nuclease can be the amino acid sequence of SEQ ID NO:1.
[0017] In some embodiments, the target-specific nuclease can be selected from the group consisting of Cas12f, Cas12m, and variants thereof, and optionally the target-specific nuclease can be PsaCas12f.
[0018] In some embodiments, the gRNA can be a single guide RNA (sgRNA) or a dual guide RNA (dgRNA). In some embodiments, the gRNA can be an sgRNA comprising a nucleic acid sequence having 70% identity to a nucleic acid sequence selected from the group consisting of SEQ ID NOs: 20-43 and 61-79. In some embodiments, the gRNA comprises a spacer region having a sequence of about 20 to about 30 nucleotides (nt), a length of about 22 nt, or the gRNA comprises a spacer region having a sequence of about 20 to about 53 nt, or about 29 to about 53 nt, or about 40 to about 50 nt.
[0019] In some embodiments, the DNA target may be within a cell. In some embodiments, the cell may be a prokaryotic cell. In some embodiments, the cell may be a eukaryotic cell. In some embodiments, the eukaryotic cell may be a mammalian cell. In some embodiments, the mammalian cell may be a human cell.
[0020] In some embodiments, the amino acid sequence can specifically bind to a protospacer adjacent motif (PAM). In some embodiments, the PAM can be selected from the group consisting of NNNNGATT, NNNNGNNN, NNG, NG, NGAN, NGNG, NGAG, NGCG, NAAG, NGN, NRN, NNGRRN, NNNRRT, TTTN, TTTV, TYCV, TATV, TYCV, TATV, TTN, KYTV, TYCV, TATV, TBN, and any variants thereof, and any combination thereof.
[0021] In another aspect, a method of detecting a DNA target is contemplated, comprising coupling the DNA target with a reporter to form a DNA reporter complex, mixing the DNA reporter complex with a target-specific nuclease and a guide RNA (gRNA), cleaving the DNA reporter complex, and measuring a signal from the reporter, thereby detecting the DNA target. In some embodiments, the target-specific nuclease can be selected from the group consisting of Cas12f, Cas12m, and variants thereof, and optionally the target-specific nuclease can be PsaCas12f. In some embodiments, the target-specific nuclease can be complexed with a crRNA. In some embodiments, the reporter can be a fluorescent reporter.
[0022] In another aspect, a method of activating or inhibiting expression of a gene is contemplated, comprising combining a composition with one or more transcription factors, wherein the composition comprises a target-specific nuclease comprising an amino acid sequence having 70% identity to an amino acid sequence selected from the group consisting of SEQ ID NOs: 1-19, a DNA target, and a guide RNA (gRNA), wherein the target-specific nuclease lacks endonuclease capability, and the target DNA comprises the gene, thereby activating the gene.
[0023] In another aspect, a method of editing a nucleobase is contemplated, comprising mixing compositions, the composition comprising a target-specific nuclease comprising an amino acid sequence having 70% identity to an amino acid sequence selected from the group consisting of SEQ ID NOs: 1-19, a DNA target, and a guide RNA (gRNA), wherein the target-specific nuclease is a nickase or a nuclease coupled to a deaminase, thereby editing a nucleobase from a target DNA.
[0024] In another aspect, a method of activating or inhibiting expression of a gene is contemplated, comprising mixing a composition comprising a target-specific nuclease comprising an amino acid sequence 70% identical to an amino acid sequence selected from the group consisting of SEQ ID NOs: 1-19, and a guide RNA (gRNA) (wherein the target comprises a DNA target) with one or more epigenetic modifiers, wherein the target-specific nuclease lacks endonuclease activity, the target DNA comprises the gene, and modifies one or more histones associated with the target DNA or the target DNA, thereby activating or inhibiting the gene. In some embodiments, the epigenetic modifiers can include KRAB, DNMT3a, DNMT1, DNMT3b, DNMT3L, TET1, p300, and any variants thereof, and any combination thereof.
[0025] These aspects and embodiments, as well as others, are disclosed in further detail herein.
[0026] Aspects, features, benefits, and advantages of the embodiments described herein will be apparent from the following description, the appended claims, and the accompanying drawings. [Brief description of the drawings]
[0027] [Figure 1A] 1 shows a schematic illustrating the computational identification of novel micro-CRISPR nucleases in metagenomic samples according to an embodiment of the present teachings. [Figure 1B] 1 shows a simulated tree of Cas orthologs according to an embodiment of the present teachings. [Figure 1C] 1 shows the size distribution of Cas12a orthologs according to an embodiment of the present teachings. [Figure 1D] 1 shows the size distribution of CasM orthologs according to an embodiment of the present teachings. [Figure 1E] 1 shows secondary structure prediction of PasCas12f tandem repeats according to an embodiment of the present teachings. [Figure 1F] 1 shows secondary structure prediction of putative PasCas12tracrRNA according to an embodiment of the present teachings. [Diagram 2] 1 shows a schematic illustrating screening of smaller CRISPR nucleases for functional activity via LASSO and TXTL according to an embodiment of the present teachings. [Figure 3A] 1 shows a vector map depicting a smaller CRISPR nuclease-enabled single vector activator, base editor, or homology directed repair (HDR) according to an embodiment of the present teachings. [Figure 3B] FIG. 1 shows a schematic illustrating in vivo modification via a single vector activator, base editor, or HDR using AAV according to an embodiment of the present teachings. [Figure 3C] 1 shows the optimization of small CRISPR effectors for mammalian single vector delivery according to an embodiment of the present teachings. [Figure 4] 1 shows testing of PsaCas12f sgRNA constructs in human mammalian cells according to an embodiment of the present teachings. [Figure 5A] 1 shows testing of a PsaCas12f NLS construct according to an embodiment of the present teachings. [Figure 5B] 1 shows editing with PsaCas12f (NLS14) with sgRNA13 according to an embodiment of the present teachings. [Figure 5C] 1 shows editing with PsaCas12f (NLS14) with a non-targeted guide according to an embodiment of the present teachings. [Figure 5D] 1 shows editing with PsaCas12f (no NLS) with sgRNA14 according to an embodiment of the present teachings. [Figure 5E] 1 shows editing with PsaCas12f (no NLS) with a non-targeted guide according to an embodiment of the present teachings. [Figure 6A] 1 illustrates a process for optimal guide RNA prediction according to an embodiment of the present teachings. [Figure 6B] 1 shows the predicted energy landscapes of different RNA designs according to embodiments of the present teachings. [Figure 6C] 1 shows in vitro cleavage with PsaCas12f using different sgRNA scaffolds generated by in silico optimization, according to an embodiment of the present teachings. [Figure 7A] FIG. 1 shows a diagram of a luciferase indel reporter for engineering novel CRISPR effectors such as PsaCas12f for editing mammalian genomes in accordance with an embodiment of the present teachings. [Figure 7B] 1 shows genome editing data of PasCas12f in HEK293FT cells, showing approximately 0.05% indel activity, 100-fold higher than background detection, according to an embodiment of the present teachings, where activity is detected by N-terminal NLS Cas12f expression and the native guide scaffold. [Figure 7C] 1 shows a bar graph of gene editing using PasCas12f in HEK293FT cells according to an embodiment of the present teachings. [Figure 7D] 1 shows an allelic plot of Cas12f EMX1 cleavage showing indels in the target, according to an embodiment of the present teachings. [Figure 7E] 1 shows a bar graph of sgRNA and DR / tracr optimization of Cas12f according to an embodiment of the present teachings, where the luciferase reporter of indels reveals the main sgRNA and tracrRNA / DR combinations with indel activity in HEK293FT cells. [Figure 8A] 1 shows a schematic of the PsaCas12f expression locus according to an embodiment of the present teachings. [Figure 8B] 1 shows PsaCas12f PAM determined by in vitro cleavage according to an embodiment of the present teachings. [Figure 8C] 1 shows predicted crRNAs determined by small RNA sequencing according to an embodiment of the present teachings. [Figure 8D] 1 shows validation of PasCas12f PAM in vitro cleavage using recombinant protein according to an embodiment of the present teachings. [Figure 9A]1 shows PsaCas12f coupled to mini-VPR for CRISPR activation (CRISPRa) using inactive (dead) PsaCas12f according to an embodiment of the present teachings. [Figure 9B] 1 shows a bar graph of RLU of PsaCas12f coupled to VPR and miniVPR, demonstrating that gene activation using miniVPR and VPR can be achieved with catalytically inactive PsaCas12f according to an embodiment of the present teachings, where the pDF235 reporter and the EXM1v2 reporter are different luciferase reporters that measure gene activation. [Figure 9C] 1 shows a bar graph of the RLU of PsaCas12f coupled to small linker sequences (5-10 aa) at six different positions according to an embodiment of the present teachings. [Figure 9D] 1 shows a bar graph of PsaCas12f fluorescence based on target-specific collagen activity that can be used for diagnosis, according to an embodiment of the present teachings. [Figure 10A] Illustrated is the resulting sgRNA secondary structure derived from in silico secondary structure determination; the boxed stem-loops (SL1-3) were predicted using http: / / rna.tbi.univie.ac.at / . Stem-loop 4 (SL4, interacts with crRNA) and stem-loop 5 (SL5) were derived from Takeda et al., Mol Cell, 81(3):558-570 (2021). [Figure 10B-1] Shows annotated stem-loop sequences of sgRNA stem-loop variants mutated to analyze the impact of gene editing efficiency. Red indicates the introduced nucleobase change, orange indicates the nucleobases that form the stem, and purple indicates the loop added to allow recruitment of MS2 coat / protein. [Figure 10B-2] Shows annotated stem-loop sequences of sgRNA stem-loop variants mutated to analyze the impact of gene editing efficiency. Red indicates the introduced nucleobase change, orange indicates the nucleobases that form the stem, and purple indicates the loop added to allow recruitment of MS2 coat / protein. [Figure 10C-1] 1 shows a bar graph of RLU using PsaCas12f with different sgRNA stem-loop variants demonstrating that modifications to the secondary structure of the sgRNA affect gene editing efficiency. [Figure 10C-2] 1 shows a bar graph of RLU using PsaCas12f with different sgRNA stem-loop variants demonstrating that modifications to the secondary structure of the sgRNA affect gene editing efficiency. [Figure 11A] A bar graph of RLU using PsaCas12f with a panel of sgRNA variants, each with a combination of modifications derived from a single modified sgRNA stem-loop variant. [Figure 11B-1] A bar graph of percent indel formation at the EMX1 genomic locus using PsaCas12f with a panel of sgRNA variants, each with a combination of modifications derived from a single sgRNA stem-loop variant (left panel: 4x combinations and right panel: 2x combinations). [Figure 11B-2] A bar graph of percent indel formation at the EMX1 genomic locus using PsaCas12f with a panel of sgRNA variants, each with a combination of modifications derived from a single sgRNA stem-loop variant (left panel: 4x combinations and right panel: 2x combinations). [Figure 11C] The two best sgRNA combination stem-loop variants (referred to as scaffold version 3.1 and scaffold version 3.2) show a bar graph of RLU using a panel of 30 mutated PsaCas12f demonstrating the robustness of sgRNA scaffold version 3.2. [Figure 12A] Schematic of the sgRNA scaffold, designated version 3.2, highlighting the location of the spacer sequence at the 3' end. [Figure 12B] A bar graph of RLU using PsaCas12f with a panel of version 3.2 sgRNA scaffolds with various spacer lengths (2, 3, 18, 19, 20, 21, 22, 23, 24 and 25 base pairs) is shown. [Figure 13] 1 shows the percent indel formation at two different positions (HBB g1, HBB h2, RNF g4, and RNF g6) within the HBB and RNF genomic loci using either PsaCas12f with sgRNA scaffold version 3.2 or Un1Cas12f1 with nbt scaffold. [Figure 14] A bar graph of the percent indel formation at the EMX genomic locus using a panel of PsaCas12f variants (intra-protein NLS constructs 1 to 6) in which an NLS sequence derived from SV40 was fused at random positions in the PsaCas12f sequence (shown in the bottom schematic). [Figure 15] A bar graph of percent indel formation at the RUNX1 genomic locus using PsaCas12f with an sgRNA scaffold (with a flanking SV40 NLS) delivered to cells by AAV particles is shown. [Figure 16A] A bar graph of RLU using a panel of 12 circularly permuted PsaCas12f mutants (referred to as cpPsaCas12_1-12) is shown. The bottom schematic depicts how the PsaCas12f sequence can be split at different positions by the insertion of (GGS)6 peptide linkers to create new N- and C-termini. [Figure 16B] A bar graph of percent indel formation at the RUNX1 genomic locus using a panel of 12 circularly permuted PsaCas12f mutants (cpPsaCas12_1-12) is shown. [Figure 17] Figure 1 shows a bar graph of percent indel formation at the RNF2 genomic locus using a panel of PasCas12f variants derived from a machine learning model that predicted point mutations that could result in higher gene editing efficiency. A PasCas12f variant with a point mutation at position 333 dramatically increased the cutting efficiency. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0028] It is understood that the following disclosure describes various aspects of the embodiments for clarity. It should be noted that a particular embodiment is not intended as an exhaustive description or as a limitation on the broad aspects discussed herein. An aspect described with a particular embodiment is not necessarily limited to that embodiment and may be practiced by any other embodiment. References throughout this application to "one embodiment," "an embodiment," or "an exemplary embodiment" mean that the particular feature, structure, or characteristic described with the embodiment is included in at least one embodiment of the invention. Thus, the appearance of the phrases "in some embodiments," "in an embodiment," or "an exemplary embodiment" in various places throughout this application are not necessarily all referring to the same embodiment, although they may. Furthermore, particular features, structures, or characteristics can be combined in any suitable manner in one or more embodiments, as would be apparent to one of ordinary skill in the art from this disclosure. Furthermore, although some embodiments described herein include some features and not other features included in other embodiments, combinations of features of different embodiments are intended to be within the scope of the invention. For example, in the appended claims, any of the claimed embodiments may be used in any combination.
[0029] definition Unless otherwise defined, technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains. Definitions of common terms and techniques in molecular biology can be found in Molecular Cloning: A Laboratory Manual, 2nd Edition (1989) (Sambrook, Fritsch, and Maniatis); Molecular Cloning: A Laboratory Manual, 4th Edition (2012) (Green and Sambrook); Current Protocols in Molecular Biology (1987) (eds. FMA Usubel et al.); the series Methods in Enzymology (Academic Press, Inc.): PCR 2: A Practical Approach (1995) (eds. MJ MacPherson, BD Hames, and GR Taylor); Antibodies, A Laboratory Manual (1988) (eds. Harlow and Lane); Antibodies A Laboratory Manual, 2nd Edition 2013 (ed. E. A. Greenfield); Animal Cell Culture (1987) (ed. R. I. Freshney); Benjamin Lewin, Genes IX, Jones and Bartlet Publishing, 2008 (ISBN 0763752223); Kendrew et al. (eds.), The Encyclopedia of Molecular Biology, Blackwell Science Ltd., 1994 (ISBN 0632021829); Robert A. Meyers (ed.), Molecular Biology and Biotechnology: a Comprehensive Desk Reference, VCH Publishers, Inc., 1995 (ISBN 9780471185710); Singleton et al., Dictionary of Microbiology and Molecular Biology 2nd ed., J. Wiley & Sons, New York, NY(March 1994), Advanced Organic Chemistry Reactions, Mechanisms and Structure, 4th ed., John Wiley & Sons (New York, NY 1992); and Marten H. Hofker and Jan van Deursen, Transgenic Mouse Methods and Protocols, 2nd ed. (2011).
[0030] As used herein, the singular forms "a," "an," and "the" include both singular and plural referents unless the context clearly dictates otherwise. Thus, for example, a reference to "a cell" includes a plurality of such cells.
[0031] As used herein, the term "optionally" or "optionally" means that the subsequently described event, circumstance or substituent may or may not occur, and that the description includes cases where the event or circumstance occurs and cases where it does not occur.
[0032] The recitation of numerical ranges by endpoints includes not only the recited endpoints but also all numbers and fractions subsumed within the respective ranges.
[0033] As used herein, the term "about" or "approximately" refers to a measurable value, e.g., a parameter, amount, temporal duration, etc., that is intended to encompass variations from the specified value, e.g., variations of + / - 10% or less, + / - 5% or less, + / - 1% or less, + / - 0.5% or less, and + / - 0.1% or less from the specified value, so long as such variations are appropriate for the practice of the disclosed invention. It should be understood that the value to which the modifier "about" or "approximately" refers is itself disclosed.
[0034] As used herein, the term "polypeptide" and the like refers to an amino acid sequence that includes a plurality of contiguous polymerized amino acid residues (e.g., at least about 2 contiguous polymerized amino acid residues). "Polypeptide" refers to an amino acid sequence, oligopeptide, peptide, protein, enzyme, nuclease, or portion thereof, and the terms "polypeptide," "oligopeptide," "peptide," "protein," "enzyme," and "nuclease" are used interchangeably.
[0035] Polypeptides, as used herein, also include polypeptides having various amino acid additions, deletions, or substitutions relative to the native amino acid sequence of the polypeptide of the present disclosure. In some embodiments, polypeptides that are homologs of the polypeptide of the present disclosure contain non-conservative changes of certain amino acids relative to the native sequence of the polypeptide of the present disclosure. In some embodiments, polypeptides that are homologs of the polypeptide of the present disclosure contain conservative changes of certain amino acids relative to the native sequence of the polypeptide of the present disclosure, and thus can be referred to as conservatively modified variants. Conservatively modified variants may include discrete substitutions, deletions, or additions in the polypeptide sequence, thereby resulting in the replacement of amino acids with chemically similar amino acids. Conservative substitution tables providing functionally similar amino acids are well known in the art. Such conservatively modified variants are in addition to, and do not exclude, the polymorphic variants, interspecies homologs, and alleles of the present disclosure. The following eight groups contain amino acids that are conservatively substituted for one another: 1) alanine (A), glycine (G); 2) aspartic acid (D), glutamic acid (E); 3) asparagine (N), glutamine (Q); 4) arginine (R), lysine (K); 5) isoleucine (I), leucine (L), methionine (M), valine (V); 6) phenylalanine (F), tyrosine (Y), tryptophan (W); 7) serine (S), threonine (T); and 8) cysteine (C), methionine (M) (see, e.g., Creighton, Proteins (1984)). Modifications of amino acids that produce chemically similar amino acids can be referred to as analogous amino acids.
[0036] The term "variant" as used herein refers to a polypeptide or nucleotide sequence that differs in amino acid or nucleic acid sequence by addition (e.g., insertion), deletion, or conservative substitution of amino acids or nucleotides, but retains some or all of the biological activity of the given polypeptide (e.g., a variant nucleic acid can still encode the same or similar amino acid sequence). It is recognized in the art that conservative substitutions of amino acids, i.e., replacement of an amino acid with an amino acid of different similar properties (e.g., hydrophilicity, and degree and distribution of charged regions), typically involve minor changes. These minor changes can be ascertained in part by consideration of the hydropathic index of the amino acid, and are understood in the art (see, e.g., Kyte et al., J. Mol. Biol., 157:105-132 (1982)). The hydropathic index of an amino acid is based on consideration of the hydrophobicity and charge of the amino acid. It is known in the art that amino acids of similar hydropathic indexes can be substituted and still retain protein function. The present disclosure provides amino acids having hydropathic indexes of ±2 that may be substituted. The hydrophilicity of an amino acid may be used to identify substitutions that result in a protein that retains some or all of its biological function. Consideration of the hydrophilicity of an amino acid in the context of a peptide allows for calculation of the maximum local average hydrophilicity of the peptide, which is a useful measure that has been reported to correlate well with antigenicity and immunogenicity (see, e.g., U.S. Pat. No. 4,554,101). Substitution of amino acids with similar hydrophilicity values may result in peptides that retain some or all of their biological activity, such as immunogenicity, and is understood in the art. The present disclosure provides substitutions that may be made with amino acids having hydrophilicity values within ±2 of each other. Both the hydropathic index and hydrophilicity value of an amino acid are influenced by the particular side chain of the amino acid.Consistent with this observation, amino acid substitutions that are compatible with biological function are understood to depend on the relative similarity of the amino acids, particularly the side chains of those amino acids, as manifested by hydrophobicity, hydrophilicity, charge, size, and other characteristics.
[0037] The term "variant" can also be used to describe a polypeptide or fragment thereof that has been differentially processed, for example, by proteolysis, phosphorylation, or other post-translational modification, but still retains some or all of its biological and / or antigenic reactivity. The use of "variant" herein is intended to encompass fragments of variants unless the context indicates otherwise. The term "protospacer adjacent motif" as used herein refers to a DNA sequence that immediately follows a DNA sequence that is targeted by a nuclease. Examples of protospacer adjacent motifs include, but are not limited to, NNNNGATT, NNNNGNNN, NNG, NG, NGAN, NGNG, NGAG, NGCG, NAAG, NGN, NRN, NNGRRN, NNNRRT, TTTN, TTTV, TYCV, TATV, TYCV, TATV, TTN, KYTV, TYCV, TATV, TBN, any variants thereof, and any combinations thereof.
[0038] Alternatively or additionally, a "variant" should be understood to be a polynucleotide or protein that differs in one or more changes in length or sequence compared to the polynucleotide or protein from which it is derived. The polypeptide or polynucleotide from which a protein or nucleic acid variant is derived is also known as the parent polypeptide or polynucleotide. The term "variant" includes "fragments" or "derivatives" of the parent molecule. Typically, a "fragment" is smaller in length or size than the parent molecule, while a "derivative" exhibits one or more differences in sequence compared to the parent molecule. Also included are modified molecules, such as, but not limited to, post-translationally modified proteins (e.g., glycosylated, biotinylated, phosphorylated, ubiquitinated, palmitoylated or proteolytically cleaved proteins) and modified nucleic acids, such as methylated DNA. Also included in the term "variant" are mixtures of different molecules, such as, but not limited to, RNA-DNA hybrids. Typically, variants are artificially constructed by genetic engineering means, while the parent polypeptide or polynucleotide is a wild-type protein or polynucleotide. However, it should be understood that naturally occurring variants are also encompassed by the term "variant" as used herein. Furthermore, variants usable in the present disclosure may also be derived from homologs, orthologs, or paralogs of the parent molecule, or artificially constructed variants, provided that the variant exhibits at least one biological activity of the parent molecule, i.e., is functionally active.
[0039] Alternatively or additionally, a "variant" as used herein can be characterized by a degree of sequence identity to the parent polypeptide or polynucleotide from which it is derived. More specifically, a protein variant in the context of this disclosure exhibits at least 80% sequence identity to the parent polypeptide. A polynucleotide variant in the context of this disclosure exhibits at least 70% sequence identity to the parent polynucleotide. The term "at least 70% sequence identity" and the like is used throughout this application with respect to comparisons of polypeptide and polynucleotide sequences. This term refers to at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity to a corresponding reference polypeptide or a corresponding reference polynucleotide.
[0040] Nucleotide and amino acid sequence similarity, i.e., the percentage of sequence identity, can be determined by sequence alignment. Such alignments can be carried out by algorithms known in the art, by the mathematical algorithm of Karlin and Altschul (Karlin & Altschul (1993) Proc. Natl. Acad. Sci. USA 90:5873-5877), by hmmalign (HMMER package, hmmer.wustl.edu / ), or by the CLUSTAL algorithm (Thompson, JD, Higgins, DG and Gibson, TJ (1994) Nucleic Acids Res. 22:4673-80), available for example at www.ebi.ac.uk / Tools / clustalw / or on www.ebi.ac.uk / Tools / clustalw2 / index.html or on npsa-pbil.ibcp.fr / cgi-bin / npsa_automat.pl?page= / NPSA / npsa_clustalw.html. Some parameters used are default parameters and are set at www.ebi.ac.uk / Tools / clustalw / or www.ebi.ac.uk / Tools / clustalw2 / index.html. The grade of sequence identity (sequence matching) can be calculated, for example, using BLAST, BLAT or BlastZ (or BlastX). A similar algorithm is incorporated into the BLASTN and BLASTP programs of Altschul et al. (1990) J. Mol. Biol. 215:403-410. To obtain gapped alignments for comparison purposes, Gapped BLAST is used as described in Altschul et al. (1997) Nucleic Acids Res. 25:3389-3402. When using BLAT and Gapped BLAST programs, the default parameters of the corresponding programs can be used.Sequence matching analysis may be supplemented with established homology mapping techniques such as Shuffle-LAGAN (Brudno M., Bioinformatics 2003b, vol. 19 Suppl. 1: I54-I62) or Markov Random Fields. When percentages of sequence identity are referred to in this application, these percentages are calculated with respect to the full length of the longer sequence, unless specifically indicated.
[0041] As used herein, the term "mini-CRISPR nuclease" refers to a "target-specific nuclease" having a compact structure with a small number of amino acids.
[0042] As used herein, the term "target-specific nuclease" and the like refers to a nuclease that targets DNA and is directed to a target nucleic acid sequence of the DNA by a guide RNA (gRNA). The DNA can be single-stranded DNA or double-stranded DNA.
[0043] As used herein, the term "guide RNA" (gRNA) and the like refers to an RNA that guides the editing, activation or inhibition of one or more genes of interest or one or more nucleic acid sequences of interest to a target genome. The gRNA can target a nuclease to a target nucleic acid or sequence in a genome. The gRNA can also refer to prime editing guide RNA (pegRNA), nicking guide RNA (ngRNA), single guide RNA (sgRNA), i.e., a fusion of two non-coding RNAs, synthetic CRISPR RNA (crRNA) and trans-activating CRISPR RNA (tracrRNA), and dual guide RNA (dgRNA). In some embodiments, the term "gRNA molecule" and the like refers to a nucleic acid that encodes a gRNA. In some embodiments, the gRNA molecule is of non-natural origin. In some embodiments, the gRNA molecule is a synthetic gRNA molecule.
[0044] As used herein, the term "target" or the like refers to a polynucleotide or polypeptide that is targeted. In some embodiments, the target is a DNA target. In some embodiments, the DNA target is associated with one or more histones. In some embodiments, the DNA target is a double-stranded DNA target. In other embodiments, the DNA target is a single-stranded DNA target.
[0045] As used herein, the terms "circular permutation," "circularly permuted," and "(CP)" refer to the conceptual process of taking a linear protein or its cognate nucleic acid sequence, fusing the natural N- and C-termini (either directly or via a linker using protein or recombinant DNA methods) to form a circular molecule, and then cleaving the circular molecule at a different location to form a new linear protein or cognate nucleic acid molecule with termini different from those of the original molecule. Thus, circular permutation preserves the sequence, structure, and function of the protein (except for the optional linker), while generating new C- and N-termini in different locations, thereby resulting in improved orientation to fuse a desired polypeptide fusion partner compared to the original ligand, according to one aspect of the invention. Circular permutation also includes any process that results in a circularly permuted linear molecule, as defined herein. In general, circularly permuted molecules are expressed de novo as linear molecules and do not formally go through the circularization and circular opening steps.
[0046] It is noted that all publications and references cited herein are expressly incorporated herein by reference in their entirety. The publications discussed herein are provided solely for their disclosure prior to the filing date of this application. Nothing herein is to be construed as an admission that the present invention is not entitled to antedate such publications. Further, the dates of publication provided may be different from the actual publication dates, which may need to be independently verified.
[0047] Overview The embodiments disclosed herein provide non-naturally occurring or modified systems, methods and compositions comprising mini-CRISPR nucleases for gene editing and for programmable gene activation and inhibition. Mini-CRISPR nucleases are target-specific nucleases with compact structures with a small number of amino acids. Target-specific nucleases target single-stranded or double-stranded DNA and are directed to the target nucleic acid sequence of DNA by a guide RNA (gRNA). The gRNA can be a single guide RNA, i.e., a fusion of two non-coding RNAs, a synthetic CRISPR RNA (crRNA) and a trans-activating CRISPR RNA (tracrRNA). The crRNA and tracrRNA help to direct the target-specific nuclease to the target nucleic acid sequence, and these RNA molecules can be specifically modified to target a particular nucleic acid sequence. Certain aspects of the present teachings concern target-specific nucleases that exhibit DNA cleavage activity and are directed to the target nucleic acid of DNA by a gRNA. Certain embodiments of the present teachings relate to target-specific nucleases that do not exhibit DNA cleavage activity and that are directed to a DNA target nucleic acid by a gRNA molecule. Certain embodiments of the present teachings relate to target-specific nucleases for diagnostic applications.
[0048] Micro CRISPR nuclease Some embodiments disclosed herein are directed to non-naturally occurring or modified CRISPR-Cas (clustered regularly interspaced palindromic repeats associated protein) systems. In conflicts with viruses associated with bacterial hosts, CRISPR-Cas systems provide an adaptive defense mechanism that exploits programmed immune memory. CRISPR-Cas systems provide defense at three stages: an adaptation stage that incorporates short nucleic acid sequences into CRISPR arrays that act as memories of past infections, an expression stage that transcribes CRISPR arrays into pre-crRNA (CRISPR RNA) transcripts and processes the pre-crRNA into functional crRNA species that target foreign nucleic acids, and an interference stage that programs CRISPR effectors with crRNA to cleave the nucleic acid of the foreign threat. In all CRISPR-Cas systems, these basic stages represent a huge variation, including the identity of the target nucleic acid (either RNA, DNA or both) and the identity of the diverse domains and proteins involved in the effector ribonucleoprotein complexes of the system.
[0049] CRISPR-Cas systems can be broadly divided into two classes based on the structure of the effector modules involved in pre-crRNA processing and interference. Class 1 systems have multisubunit effector complexes composed of many proteins, while class 2 systems rely on a single effector protein with multiple domains capable of crRNA binding and interference, with class 2 effectors often also providing pre-crRNA processing activity. Class 1 systems include three types (types I, III, and IV) and 33 subtypes, including RNA- and DNA-targeted type III systems. Class 2 CRISPR family encompasses three types of systems (types II, V, and VI) and 17 subtypes, including the RNA-guided DNases Cas9 and Cas12, and the RNA-guided RNase Cas13. Continuous sequencing of novel bacterial genomes and metagenomes has revealed new diversity and innovative relationships of CRISPR-Cas systems, so experiments are needed to clarify the functionality of these systems and develop new tools.
[0050] The CRISPR-Cas system disclosed herein comprises a mini-CRISPR nuclease. The mini-CRISPR nuclease is a target-specific nuclease that has a compact structure with a small number of amino acids and targets DNA. The target-specific nuclease disclosed herein may be, for example, but not limited to, Cas12f, Cas12m and variants thereof, and optionally the target-specific nuclease may be PsaCas12f. In some embodiments, the target-specific nuclease is a nuclease that edits single-stranded or double-stranded DNA. In some embodiments, the target-specific nuclease is a nuclease that edits single-stranded DNA (ssDNA). In some embodiments, the target-specific nuclease is a nuclease that edits double-stranded DNA. In some embodiments, the target-specific nuclease is a nuclease that edits DNA of the genome of a cell.
[0051] The CRISPR-Cas systems disclosed herein can include one or more epigenetic modifiers. Examples of epigenetic modifiers include, but are not limited to, KRAB, DNMT3a, DNMT1, DNMT3b, DNMT3L, TET1, p300, any variants thereof, and any combinations thereof.
[0052] The target-specific nuclease can comprise an amino acid sequence selected from the group consisting of SEQ ID NOs: 1-19. For example, the target-specific nuclease comprises an amino acid sequence having at least 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identity to an amino acid sequence selected from the group consisting of SEQ ID NOs: 1-19.
[0053] In some embodiments, the target-specific nuclease comprises a tag, such as, but not limited to, 3xFlag, a nuclear localization sequence (NLS), and a combination of 3xFlag and NLS.
[0054] The CRISPR-Cas system disclosed herein comprises a guide RNA (gRNA). The gRNA directs the target-specific nuclease to a target nucleic acid sequence of single-stranded or double-stranded DNA targeted by the nuclease. In some embodiments, the gRNA is a single guide RNA (gRNA). In some embodiments, the gRNA comprises a CRISPR RNA (crRNA), a trans-activating CRISPR RNA (tracrRNA), or a combination thereof. The crRNA and tracrRNA help direct the target-specific nuclease to the target nucleic acid sequence, and these RNA molecules can be specifically engineered to target a particular nucleic acid sequence.
[0055] In general, the guide sequence of a gRNA is any polynucleotide sequence that has sufficient complementarity with a target polynucleotide to hybridize with the target sequence and direct sequence-specific binding of a target-specific nuclease to the target sequence. In some embodiments, the degree of complementarity between a guide sequence and a corresponding target sequence is about or greater than about 50%, 52%, 54%, 56%, 58%, 60%, 62%, 64%, 66%, 68%, 70%, 72%, 74%, 76%, 78%, 80%, 82%, 84%, 86%, 88%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more when optimally aligned using a suitable alignment algorithm. Optimal alignment can be determined by use of any suitable algorithm to align sequences, non-limiting examples of which include the Smith-Waterman algorithm, the Needleman-Wunsch algorithm, algorithms based on the Burrows-Wheeler Transform (e.g., Burrows Wheeler Aligner), ClustalW, ClustalX, BLAT, Novoalign (Novocraft Technologies, ELAND (Illumina, San Diego, Calif.), SOAP (available at soap.genomics.org.cn), and Maq (available at maq.sourceforge.net). In some embodiments, the guide sequences are about 5 or more than about 5, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, 55, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 13, 14, 15, 16, 17 In some embodiments, the guide sequence is less than about 75, 50, 45, 40, 35, 30, 25, 20, 15, 12, or fewer nucleotides in length. In some embodiments, the guide RNA has a spacer region whose sequence is about 17 to about 53 nucleotides (nt), about 25 nt to about 53 nt, about 29 nt to about 53 nt, or about 40 nt to about 50 nt in length.In some embodiments, the guide RNA has a spacer region whose sequence has a length of about 20 nt, about 21 nt, about 22 nt, about 23 nt, about 24 nt, about 25 nt, about 26 nt, about 27 nt, about 28 nt, about 29 nt, about 30 nt, about 31 nt, about 32 nt, about 33 nt, about 34 nt, about 35 nt, about 36 nt, about 37 nt, about 38 nt, about 39 nt, about 40 nt, about 41 nt, about 42 nt, about 43 nt, about 44 nt, about 45 nt, about 46 nt, about 47 nt, about 48 nt, about 49 nt, about 50 nt, or any range made up of any two or more points in the above list. In some embodiments, the guide RNA has a tandem repeat region whose sequence has a length of about 15 nt, about 16 nt, about 17 nt, about 18 nt, about 19 nt, about 20 nt, about 21 nt, about 22 nt, about 23 nt, about 24 nt, about 25 nt, about 26 nt, about 27 nt, about 28 nt, about 29 nt, about 30 nt, about 31 nt, about 32 nt, about 33 nt, about 34 nt, about 35 nt, about 36 nt, about 37 nt, about 38 nt, about 39 nt, about 40 nt, about 41 nt, about 42 nt, about 43 nt, about 44 nt, about 45 nt, about 46 nt, about 47 nt, about 48 nt, about 49 nt, about 50 nt or any range made up of any two or more points in the above list. In some embodiments, the guide RNA has a tracrRNA region having a sequence of about 15 nt, about 16 nt, about 17 nt, about 18 nt, about 19 nt, about 20 nt, about 21 nt, about 22 nt, about 23 nt, about 24 nt, about 25 nt, about 26 nt, about 27 nt, about 28 nt, about 29 nt, about 30 nt, about 31 nt, about 32 nt, about 33 nt, about 34 nt, about 35 nt, about 36 nt, about 37 nt, about 38 nt, about 39 nt, about 40 nt, about 41 nt, about 42 nt, about 43 nt, about 44 nt, about 45 nt, about 46 nt, about 47 nt, about 48 nt, about 49 nt, about 50 nt, or any range made up of any two or more points in the above list. The ability of the guide sequence to direct sequence-specific binding of a target-specific nuclease to a target sequence can be assessed by any suitable assay.
[0056] In some embodiments, the gRNA comprises a nucleic acid sequence selected from the group consisting of SEQ ID NOs: 20-43 and 61-79. For example, the sgRNA may comprise a nucleic acid sequence that is at least 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to an amino acid sequence selected from the group consisting of SEQ ID NOs: 20-43 and 61-79.
[0057] Discovery of tiny CRISPR nucleases The main challenge for in vivo genome modification is the size of the tools, which is a hindrance to viral delivery in applications such as base editing, activation, inhibition, and HDR, among others. The most commonly used Cas9 orthologue is Streptococcus pyogenes SpCas9, a large, 1368 amino acid long protein. Small CRISPR nucleases, less than about 1000 amino acids long, can yield base editors and transcriptional activators that can fit within the 4.7 kb limit of AAV vectors. Small CRISPR nucleases can be discovered through metagenomic mining and innovative screening methods. Protein and guide RNA modifications can be used to boost the activity of these smaller nucleases for robust mammalian cell applications.
[0058] Cas12f and Cas12h nucleases are the smallest DNA-targeting Cas12 family members characterized to date, with Cas12f having about 400 to about 700 residues and Cas12h having about 870 to about 933 residues. However, these enzymes have not been engineered for high performance genome editing, and the unquantified editing rate of Cas12f and genome editing in mammalian cells has not yet been demonstrated with Cas12h.
[0059] Cas12f, Cas12h and novel Cas12 systems can be mined in various prokaryotic genomes to identify shorter proteins. The family of known Cas12f / h orthologs can be used to seed Hidden Markov Model (HMM) alignment algorithms to search the NCBI and JGI databases of prokaryotic genomes and metagenomes to discover new enzymes. The computational identification of novel minute CRISPR nucleases from metagenomic samples is illustrated in Figure 1A. The JGI database is particularly suitable for this search, as it contains more than about 100,000 genomes and metagenomes, and more than about 54 billion protein-coding genes, and is continuously growing rapidly.
[0060] Single effector CRISPR enzyme families lacking homology to classified enzymes can be found by searching CRISPR arrays in the aggregate genome and in CRISPRs selected near single effector proteins, which may represent new putative subtypes of class 2 CRISPR systems. Additional data sources from novel metagenomic sources can be used to complement this approach, including urban sampled metagenomes from diverse metros and microbiomes from non-Western cohorts, which have been demonstrated to have numerous additional uncharacterized genes.
[0061] Using CRISPR arrays as seed markers, genes can be selected within the proximity of these arrays to develop a neighborhood of CRISPR-associated genes. HMM profiles of CRISPR-associated proteins can be generated from the literature, and these profiles can be applied to screen from known systems. All remaining genes in the dataset can be clustered by a linear time clustering algorithm, such as LinClust. To select single effectors, the simultaneous association of different protein clusters with each other can be investigated, and clusters that are only associated with CRISPR arrays or only associated with known CRISPR adaptation mechanisms, such as, but not limited to, Cas1, Cas2, and Cas4, can be selected. These putative single effector clusters can then be functionally annotated by HMM-based alignment to the constructed pfam. Clusters can be initially selected based on the presence or similarity to known nuclease domains, such as, but not limited to, RuvC and HNH, if they are less than about 800 residues in length. These candidates can be iteratively searched through the integrated database to ensure that the "shorter" CRISPR nuclease is not a misannotated truncation of the larger nuclease due to loss of sequencing coverage or homologs of the truncated and inactivated larger nuclease. The results of panning the small CRISPR nucleases are shown in Figures 1B-1D and described in Example 1 below.
[0062] Characterization of miniature CRISPR nucleases Small CRISPR nuclease systems found by computational discovery can be screened in vitro and in vivo. DNA synthesis can allow large-scale synthesis of primers to clone gene clusters from metagenomic samples. For candidate selection, the corresponding CRISPR effector genes and any auxiliary RNA can be synthesized to test activity. This approach can be scaled up to tens of orthologs, but for screening, complementary approaches are needed to screen hundreds to thousands of potential orthologs. Next-generation DNA synthesis can allow large-scale synthesis of primers to clone gene clusters from metagenomic samples. Small CRISPR nucleases can be amplified from city sample metagenomes by isolation or context of neighboring genes and cloned into plasmids for bulk biochemical sampling using transcription-translation (TXTL) in microfluidic droplets. Biochemical assays can profile the sequence constraints or cleavage activity of CRISPR enzymes. Profiling can allow modification of these qualities for later use in mammalian cells.
[0063] Small CRISPR nucleases can be cloned using covalently linked primers (Long Adapter Single-Stranded Oligonucleotide, or LASSO) generated via pooled DNA synthesis, allowing for the cloning of hundreds of thousands of gene candidates. Because these enzymes are selected for their small size, they can be easily reconstituted into TXTL systems, allowing for the rapid screening of millions of candidates in an unpurified controlled biochemical setting. If small RNAs can be expressed into TXTLs, the orientation of the crRNA needs to be determined for each CRISPR system, so the pooled candidate library can be first expressed via RNA sequencing to determine the orientation and processing of the crRNA. A second set of LASSO primers that amplify the candidate systems that can then be synthesized, and a synthetic CRISPR array targeting the synthetic target sites can be added to the plasmid with gene-specific barcodes. A pool of these constructs can be cloned into a vector containing the target sites of the synthetic CRISPR array flanked by randomized sequences to accommodate all possible PAMs. In the TXTL system, a successful cleavage event can generate a double-stranded break next to the PAM sequence, which can be captured by adapter ligation. Subsequent PCR amplification can generate an amplicon containing both the cleaved PAM sequence and the gene-specific barcode. Pool sequencing of this library can reveal the best candidates capable of cleavage and their corresponding sequence choices. In addition, pool TXTL assays can be performed at different time points to profile cleavage kinetics and select the orthologs with the best activity. Once the best candidates are identified, the enzymes can be cloned individually and cleavage activity can be tested in separate TXTL reactions with fixed PAM targets. The candidates with the most active and least restrictive optimal PAM can then be validated.
[0064] Existing orthologues of Cas12f / h can also be screened to maximize the success of identifying smaller nucleases for genome editing. This can cause problems with candidate nuclease expression in the TXTL system. For example, base expression bias can limit expression. If unsatisfactory results are found in the TXTL assay, pooled LASSO can be used to assay the heterogeneity of the construct in E. coli cells. Candidates can be screened by targeting synthetic guides to ccdB toxin plasmids with degenerate PAM libraries, allowing positive selection of active gene candidates and easy sequencing of candidate barcode and PAM sequences by picking surviving clones. Examples of protospacer adjacent motifs include, but are not limited to, NNNNGATT, NNNNGNNN, NNG, NG, NGAN, NGNG, NGAG, NGCG, NAAG, NGN, NRN, NNGRRN, NNNRRT, TTTN, TTTV, TYCV, TATV, TYCV, TATV, TTN, KYTV, TYCV, TATV, TBN, any variant thereof, and any combination thereof.
[0065] Discovery of guide RNA for miniature CRISPR nucleases Some embodiments disclosed herein require a gRNA that includes a tracrRNA. Small RNA sequencing studies can be performed to determine the molecular identity of the tracrRNA and associated crRNA. However, further optimization of the small RNA is often required to reach the activity level required for DNA cleavage and genome editing in mammalian cells. These designs can be informed by secondary structure algorithms that predict both optimal hybridization and tracrRNA structures with ideal hairpins for protein binding. In vitro cleavage assays can be performed on a panel of both crRNAs with various DR and spacer lengths and tracrRNAs of different structures. These models can be further optimized in the in silico design space by simulating the progressive shortening and folding of the putative tracrRNA or crRNA, generating energy landscapes that can be validated by in vitro cleavage reactions (Figures 6A and 6B). Once good candidates are found, the crRNAs and tracrRNAs can be combined into single guide RNAs (sgRNAs) by using potential loop and linker combinations to find optimal sgRNA designs. For Cas12 orthologues without tracrRNA, one can simply screen the crRNA designs to find the optimal design. As an example, PsaCas12f was tested with different crRNA / tracrRNA designs and is disclosed in Example 4 and Figure 6C.
[0066] With the optimal crRNA and sgRNA design, mutagenesis tests can be performed to find mutations that can optimally stabilize proteins and boost cleavage activity. It has been found that mutations, insertions and deletions can dramatically change the editing activity of CRISPR enzymes. In vitro cleavage screening can be performed to find optimal sgRNA and crRNA variants for efficient enzyme activity. The best designs are then tested in bacteria to confirm the cellular DNA cleavage activity of these best orthologs.
[0067] Characterization of genome editing with micro-CRISPR nucleases Mini-CRISPR nucleases can serve as a rich addition to a new toolbox of easily deliverable genome modification tools. Their small size allows for delivery by AAV, so they can be used for genome editing in vivo. Furthermore, the additional space allowed by these mini-proteins can allow fusion with multiple effectors, including transcription activators, repressors, and deaminases, enabling single-vector HDR delivery (Figure 3A). Mini-CRISPR nucleases can be engineered for mammalian genome editing, and editing efficiency can be improved through multiplex optimization of proteins. Mini-editors can be fused with transcription activators to create mini-programmable activators capable of in vivo delivery by AAV constructs. These mini-activators can be used to demonstrate selective gene activators that activate the Pdx1 gene in vivo and treat mouse models of type I diabetes.
[0068] First, a set of miniature CRISPR nucleases can be engineered to allow genome editing, drawing on both new nucleases and previously characterized Cas12 members. The novel nucleases can be human codon optimized and cloned into mammalian expression constructs for genome editing of luciferase reporter constructs in HEK297FT cells. In this model, indels can inactivate the luciferase gene, allowing quantification of editing efficiency by loss of luciferase signal (Figure 7A). As localization of CRISPR enzymes can be a significant factor for these efficiencies, the best candidates can be selected and nuclear localization signals (NLS) can be fused to the N-terminus, C-terminus, or both to determine the effect on editing efficiency. Localization can be further verified by tagging the constructs with a small HA epitope tag, which can then be examined using immunofluorescence microscopy. In addition to providing evidence of localization, the availability of these tags can provide insight into the availability of the N- and C-termini of proteins, which can inform alterations in activation.
[0069] Furthermore, since sgRNA expression and localization may differ in vitro in mammalian context, the best sgRNA designs can be compared to further fine-tune the efficiency of editing. Flexible insertions into sgRNA can also be modified, and the effect on cleavage efficiency can be tested to determine potential areas where binding loops can be inserted. Constructs with high efficiency can be verified with disease-related endogenous gene EMX1. For example, editing tests of the PsaCas12f family for indel generation in EMX1 were performed as disclosed in Example 5 and Figure 7B. Optimization of PsaCas12f with respect to codons, differential expression, stabilization and localization can allow further increase in mammalian activity.
[0070] It is essential that genome editing tools, such as CRISPR nucleases, are active in a variety of contexts. Once enzymes and sgRNA constructs optimized for mammalian editing are determined, these constructs can be tested for robust editing of a panel of cell lines, as well as additional endogenous genes TRAC, VEGF, and Pdx1. Because the specificity of these enzymes is a key factor for their use, both as basic research tools and potential future therapeutics, an unbiased method of profiling genome-wide specificity can be used. The best performing candidates can be subjected to the GUIDE-Seq genome-wide processing pipeline. After knowing that these enzymes are effective and specific, they can be further modified for activation-based applications.
[0071] Engineering miniature CRISPR nucleases for programmable gene activation and inhibition To convert the miniature CRISPR nucleases into programmable binding platforms for applications such as editing, catalytic inactivation is required. To this end, conserved catalytic residues can be mutated in the RuvC domain of V-type effectors and loss of cleavage can be tested. Maintenance of binding activity can be verified by fusing an HA tag to the effector and determining the binding location by CHIP-Seq. If binding is still maintained in these catalytically inactivated mutants, the CHIP signal should correspond to the location targeted by the sgRNA. Upon validation of binding in mammalian cells, this minimal programmable binding platform can be used to generate programmable activators.
[0072] To reconstitute programmable activators from minimal CRISPR nucleases in mammalian cells, two parallel synergistic approaches to recruit transcriptional activators can be taken. First, a set of transcriptional activators can be fused to effector proteins at either the N- or C-terminus. These fusions are drawn from a set of known effectors, including VP64, p65, HSF1, and RTA, and these effectors can be tested alone or in combinations of up to three effectors. In parallel, sgRNAs can be modified to contain MS2 hairpin loops, which can bind to MCP proteins. MS2 loops can then be inserted into potential predetermined accessible areas. These loops can bind to MCP activator fusions, such as MCP-VP64 or p65. These constructs can then be tested alone or in combination with fusion activators to optimize activation potency. To conserve construct size and avoid the need for a second promoter, a P2A fusion linker can be used to express both the minimal CRISPR nuclease and the MCP activator from a single promoter.
[0073] Transcriptional activation candidates were tested with luciferase reporter constructs in HEK293FT cells with secreted luciferase downstream of a minimal promoter. This assay can screen different activator constructs in throughput over multiple rounds to determine the most active constructs. Importantly, constructs resulting from these rounds of optimization can be selected that are small enough to be packed into an AAV. The activity of these constructs can be verified with endogenous genes by RT-qPCR. Because recruitment of transcriptional activators and the resulting transcription machinery can be cell-state dependent, optimal constructs can be tested in various cell types to ensure robust activation in vivo. Finally, the specificity of this activation system can be profiled by targeting the HBG gene in HEK293FT cells and measuring transcriptome-wide gene expression. If the activator is specific, activation of HBG should be observed and no off-target activation should be observed. If the activator construct is specific, it can be prepared for in vivo delivery.
[0074] The transcriptional activator of the present disclosure may be targeted to a specific target nucleic acid to induce activation / expression of the target nucleic acid. In some embodiments, the transcriptional activator polypeptide targets the target nucleic acid via a heterologous DNA-binding domain. In this sense, the target nucleic acid of the present disclosure is targeted based on a specific nucleotide sequence in the target nucleic acid that is recognized by the targeting portion of the DNA-binding domain. In some embodiments, the transcriptional activator activates expression of the target nucleic acid by being targeted to the nucleic acid with the help of a guide RNA (via CRISPR-based targeting). With CRISPR-based targeting, the target nucleic acid of the present disclosure may be targeted based on a specific nucleotide sequence in the target nucleic acid that is recognized by the targeting portion of the crRNA or guide RNA used by the method of the present disclosure.
[0075] Various types of nucleic acids can be targeted for activation of expression. The target nucleic acid can be located within the coding region of the target gene or upstream or downstream thereof. Furthermore, the target nucleic acid can be endogenously resident in the target gene or can be inserted into, for example, a heterologous gene using techniques such as, for example, homologous recombination. For example, the target gene of the present disclosure can be operably linked to a regulatory region, such as, for example, a promoter, that contains a sequence that can be recognized by, for example, the crRNA / tracrRNA and / or guide RNA of the present disclosure, such that the transcriptional activator of the present disclosure can target the sequence. In some embodiments, the target nucleic acid is not a target naturally associated with and / or is not naturally associated with a naturally occurring transcriptional activator polypeptide.
[0076] The target-specific nucleases disclosed herein can be used with a variety of CRISPR gene activation methods (see, for example, Konermann S, Brigham MD, Trevino AE, Joung J, Abudayyeh OO, Barcena C, Hsu PD, Habib N, , Gootenberg JS, Nishimasu H, Nureki O, Zhang F. Genome-scale transcriptional activation by an engineered CRISPR-Cas9 complex. Nature. 2015 Jan 29; 517(7536):583-8. doi:10.1038 / nature14136. Epub 2014 Dec 10. PMID:25494202; PMCID:PMC4420636; David Bikar, Wenyan Jiang, Poulami Samai, Ann Hochschild, Feng Zhang, Luciano A. Marraffini, Programmable repression and activation of bacterial gene expression using an engineered CRISPR-Cas system, Nucleic Acids Research, Volume 41, Issue 15, August 1, 2013, Pages 7429-7437, doi.org / 10.1093 / nar / gkt520; Perez-Pinera, P., Kocak, D., Vockley, C. et al., RNA-guided gene activation by CRISPR-Cas9-based transcription factors.Nat Methods 10, 973-976 (2013).doi.org / 10.1038 / nmeth.2600;Marvin E. Tanenbaum, Luke A. Gilbert, Lei S. Qi, Jonathan S. Weissman, Ronald D.Vale, "A Protein-Tagging System for Signal Amplification in Gene Expression and Fluorescence Imaging", RESOURCE|Volume 159, Issue 3, Pages 635-646, October 23, 2014, DOI:doi.org / 10.1016 / j.cell.2014.09.039;Konermann S, Brigham MD, Trevino AE, Joung J, Abudayyeh OO, Barcena C, Hsu PD, Habib N, Gootenberg JS, Nishimasu H, Nureki O, Zhang F.Genome-scale transcriptional activation by an engineered CRISPR-Cas9 complex.Nature.2015 Jan 29; 517(7536):583-8.doi:10.1038 / nature14136.Epub December 10, 2014 PMID:25494202;PMCID:PMC4420636;Chavez, A., Schheiman, J., Vora, S. et al., Highly efficient Cas9-mediated transcriptional programming. Nat. Methods 12, 326-328 (2015). doi.org / 10.1038 / nmeth.3312 Chavez, A., Tuttle, M., Pruitt, B. et al., Comparison of Cas9 activators in multiple species. Nat Methods 13, 563-567 (2016). doi.org / 10.1038 / nmeth.3871; and Sajwan, S., Mannervik, M. Gene activation by dCas9-CBP and the SAM system differ in target preference. Sci Rep 9, 18104 (2019). See doi.org / 10.1038 / s41598-019-54179-x.
[0077] Examples of CRISPR gene activation methods include, but are not limited to, dCas9-CBP CRISPR gene activation method, SPH CRISPR gene activation method, Synergistic Activation Mediator (SAM) CRISPR gene activation method, Sun Tag CRISPR gene activation method, VPR CRISPR gene activation method, and any alternative CRISPR gene activation method. The dCas9-VP64 CRISPR gene activation method uses a nuclease that lacks endonuclease ability and is fused to VP64, a powerful transcription activation domain. Guided by the nuclease, VP64 recruits the transcription machinery to specific sequences, causing targeted gene regulation. This can be used to activate transcription in either initiation or elongation depending on which sequence is targeted. The SAM-CRISPR gene activation method uses modified sgRNA to increase transcription, which is done by creating a nuclease / VP64 fusion protein modified with an aptamer that binds to the MS2 protein. These MS2 proteins then recruit additional activation domains (HS1 and p65) to then activate the gene. Instead of using a single copy of VP64 per nuclease, the Sun Tag CRISPR gene activation method uses a repeat peptide array to fuse with multiple copies of VP64. Having multiple copies of VP64 at each locus of interest allows for the recruitment of many transcriptional machinery per target gene. The VPR-CRISPR gene activation method uses a ternary complex fused with a nuclease to activate transcription. This complex consists of the VP64 activator used in other CRISPR activation methods and two other strong transcriptional activators (p65 and Rta). These transcriptional activators act in tandem to recruit transcription factors.
[0078] The target-specific nucleases disclosed herein can be used as base editors for base editing (see, for example, Anzalone, AV, Koblan, LW & Liu, eds. DRGenome, CRISPR-Cas nucleases, base editors, transposases and prime editors. Nat Biotechnol 38:824-844 (2020), which is incorporated by reference in its entirety). There are broadly three classes of base editors: cytosine base editors (CBEs), adenine base editors (ABEs) and dual deaminase editors (also called synchronously tuned programmable adenine and cytosine editors, SPACEs). Base editing requires a nickase or nuclease fused or coupled to a deaminase that performs the edit, a gRNA that targets the nuclease to a specific locus, and a target base to edit within the editing window specified by the nuclease.
[0079] Cytosine base editors (CBEs) use cytidine deaminase coupled to an inactive nuclease. These fusions convert cytosine to uracil without cleaving DNA. Uracil is then converted to thymine via DNA replication or repair. Fusing an inhibitor of uracil DNA glycosylase (UGI) to the nuclease prevents base excision repair, which reverts U to a C mutation. To increase base editing efficiency, a nuclease nickase can be used instead of a nuclease to force the cell to use the deaminated DNA strand as a template. The resulting editor can make a nick in the unmodified DNA strand so that it appears to the cell as "newly synthesized." Thus, the cell repairs the DNA using the U-containing strand as a template and copies the base edit.
[0080] Adenine base editors (ABEs) can convert adenine to inosine, resulting in an A to G change. Creating an adenine base editor requires an additional step since there is no known DNA adenine deaminase. Directed evolution can be used to create one of the RNA adenine deaminases, TadA. While cytosine base editors often produce a mixed population of edits, some ABEs do not show significant A to non-G conversions at the target locus. Removal of inosine from DNA is potentially rare, thus preventing the induction of base excision repair. With regard to off-target effects, ABEs are also generally favorable compared to other methods.
[0081] Suitable target nucleic acids are readily apparent to those skilled in the art depending on the particular need or outcome. The target nucleic acid may be within a region of euchromatin (e.g., highly expressed genes) or the target nucleic acid may be within a region of heterochromatin (e.g., centromeric DNA). The use of transcriptional activators according to the methods disclosed herein to induce transcriptional activation within regions of heterochromatin or other highly methylated regions of the plant genome may be particularly useful in certain embodiments. The target nucleic acids of the present disclosure may be methylated or unmethylated.
[0082] The target gene can be any target gene used and / or known in the art. Exemplary target genes include, but are not limited to, Pdx1 and any variants thereof.
[0083] Delivery of micro-CRISPR nucleases In some embodiments, the target-specific nuclease and / or peptide sequence is introduced into the cell as a nucleic acid encoding the respective protein. The nucleic acid introduced into the eukaryotic cell is a plasmid DNA or a viral vector. In some embodiments, the target-specific nuclease and / or peptide sequence is introduced into the cell via a ribonucleoprotein (RNP).
[0084] Delivery is in the form of a vector, which may be a viral vector, e.g., lentiviral or baculoviral or adenoviral / adeno-associated viral vector, although other delivery methods are known and provided (e.g., yeast systems, microvesicles, vectors coupled to gene guns / gold nanoparticles). Viral vectors may be selected from various families / genera of viruses, including Myoviridae, Siphoviridae, Podoviridae, Corticoviridae, Lipothrixviridae, Poxviridae, Iridoviridae, Adenoviridae, Polyomaviridae, and the like. idae), Papillomaviridae, Mimiviridae, Pandoravirusa, Salterprovirsa, Inoviridae, Microviridae, Parvoviridae, Circoviridae, Hepadnaviridae, Caulimobilidae ulimoviridae, Retroviridae, Cystoviridae, Reoviridae, Birnaviridae, Totiviridae, Partitiviridae, Filoviridae, Orthomyxoviridae, Deltavirusa, Leviviridae viridae, Picornaviridae, Marnaviridae, Secoviridae, Potyviridae, Caliciviridae, Hepeviridae, Astroviridae, Nodaviridae, Tetraviridae, Luteoviridae,These may include, but are not limited to, Tombusviridae, Coronaviridae, Arteriviridae, Flaviviridae, Togaviridae, Virgaviridae, Bromoviridae, Tymoviridae, Alphaflexiviridae, Sobemovirusa, or Idaeovirus.
[0085] Vectors can be viral or yeast systems (e.g., where the nucleic acid of interest can be operably linked to a promoter and under the control of the promoter (e.g., for expression to ultimately provide processed RNA)), but also methods for direct delivery of nucleic acids to host cells. For example, baculoviruses can be used for expression in insect cells. These insect cells can then be useful for the production of large amounts of further vectors, such as AAV or lentiviruses, adapted for delivery of the present invention. Also envisaged are methods for delivery of target-specific nucleases and / or peptide sequences, including delivery of mRNA encoding each to cells.
[0086] One of the values of the micro transcription activator is its ability to be packaged into AAV.To this end, the best activator found can be cloned into AAV packaging vector, and AAV2 containing the minimal activator can be purified.The activity of these AAVs can be confirmed by delivery to HepG2 cells to confirm both liver targeting and activity.The titer or expression is found to be low, and various liver-specific promoters, including albumin promoter and TGB promoter, can be tested to find the minimal promoter with high expression to optimize delivery.
[0087] After confirming delivery of the minimal constructs to cell culture, expression can be assessed in mice by hydrodynamic injection of promoterless luciferase constructs, followed by tail vein injection of minimal activator AAV targeting the upstream region of these luciferase constructs. Luciferase expression can only be induced in the liver where there is successful activation, which can be measured by bioluminescence imaging.
[0088] To test activation in a less perturbed model, Pdx1 can be activated. Pdx1 is the target of in vivo activation performed by Cas9 activator in Cas9 mouse model (see PMC5732045). Overexpression of Pdx1 in liver transdifferentiates hepatocytes in vivo to generate insulin-secreting cells. Pdx1 activation can be tested in cell cultures using Hepa1-6 cells and expression can be measured by RT-qPCR to determine optimal guides. These optimal Pdx1-targeted guides can be injected into mice via tail vein injection. These mice can be harvested 2 weeks after injection to determine changes in Pdx1 expression and in genes downstream of Pdx1, such as, but not limited to, insulin and Pcsk1. To verify the phenotypic effect of Pdx1, mice can be treated with streptozotocin to produce hyperglycemia. Introduction of a Pdx1 activator can be tested to determine whether it can reduce blood glucose levels and increase serum insulin, as seen with the Cas9 activator in the Cas9 mouse model.
[0089] Combinations of transcriptional activators can result in successful activation. However, these combinations can be too large. In such cases, the activators can be truncated to find the essential domains that allow activation but with reduced size. Truncation of guide RNAs can also be assessed to regulate the binding of novel Cas effectors and quantitatively fine-tune gene activation.
[0090] In some embodiments, expression of the nucleic acid sequence encoding the target-specific nuclease and / or peptide sequence may be driven by a promoter. In some embodiments, the target-specific nuclease is Cas. In some embodiments, a single promoter drives expression of the nucleic acid sequence encoding Cas and one or more guide sequences. In some embodiments, the Cas and guide sequences are operably linked to and expressed from the same promoter. In some embodiments, the CRISPR enzyme and the guide sequence are expressed from different promoters. For example, the promoter may be, but is not limited to, a UBC promoter, a PGK promoter, an EF1A promoter, a CMV promoter, an EFS promoter, an SV40 promoter, and a TRE promoter. The promoter may be a weak or strong promoter. The promoter may be a constitutive promoter or an inducible promoter. In some embodiments, the promoter may also be an AAV ITR, which may have the advantage of eliminating the need for additional promoter elements that may take up space in the vector. The additional space freed up by the use of the AAV ITRs may be used to drive expression of additional elements, such as guide sequences. In some embodiments, the promoter may be a tissue-specific promoter.
[0091] In some embodiments, the enzyme coding sequence encoding the target-specific nuclease and / or peptide sequence is codon-optimized for expression in a particular cell, e.g., a eukaryotic cell. The eukaryotic cell may be of or derived from a particular organism, such as a mammal, including human, mouse, rat, rabbit, dog, or non-human primate. In general, codon optimization refers to the process of modifying a nucleic acid sequence for enhanced expression in a host cell of interest by replacing at least one codon (e.g., about one or more than about one, 2, 3, 4, 5, 10, 15, 20, 25, 50 or more codons) of the native sequence with a codon that is more or most frequently used in the host cell's genes, while maintaining the native amino acid sequence. Different species exhibit specific biases for certain codons of certain amino acids. Codon bias (differences in codon usage by organisms) often correlates with the efficiency of translation of messenger RNA (mRNA), which in turn is believed to depend, among other things, on the properties of the codon being translated and the availability of certain transfer RNA (tRNA) molecules. The predominance of selected tRNAs in a cell is largely a reflection of the codons most frequently used in peptide synthesis. Thus, genes can be tailored for optimal gene expression in a given organism based on codon optimization. Codon usage tables are readily available, e.g., the "Codon Usage Database," and these tables can be adapted in a number of ways. See Nakamura, Y. et al., "codon usage tabulated from the international DNA sequence databases: status for the year 2000," Nucl. Acids Res. 28:292 (2000). Computer algorithms are also available that codon-optimize a particular sequence for expression in a particular host cell, e.g., Gene Forge (Aptagen: Jacobus, Pa).In some embodiments, one or more codons (e.g., 1, 2, 3, 4, 5, 10, 15, 20, 25, 50 or more, or all codons) in a sequence encoding a Cas protein correspond to the most frequently used codon for a particular amino acid.
[0092] In some embodiments, the vector encodes a target-specific nuclease and / or peptide sequence that includes one or more nuclear localization sequences (NLSs), such as about or more than about one, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more NLSs. In some embodiments, the Cas protein includes about or more than one, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more NLSs at or near the amino terminus, about or more than one, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more NLSs at or near the carboxy terminus, or combinations thereof (e.g., one or more NLSs at the amino terminus and one or more NLSs at the carboxy terminus). When more than one NLS is present, each may be selected independently of the other, such that a single NLS may be present in more than one copy and / or a combination with one or more other NLSs may be present in one or more copies. In some embodiments, an NLS is considered to be near the N- or C-terminus if the nearest amino acid of the NLS is within about 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 40, 50 or more amino acids along the polypeptide chain of the N- or C-terminus. Typically, an NLS consists of one or more short sequences of positively charged lysines or arginines exposed on the protein surface, although other types of NLS are known. In some embodiments, an NLS is between two domains, for example, between a Cas12 protein and a viral protein. An NLS may also be between two functional domains separated or flanked by a glycine-serine linker.
[0093] In general, one or more NLSs are strong enough to drive accumulation of detectable amounts of the target-specific nuclease and / or peptide sequence in the nucleus of a eukaryotic cell. In general, the strength of the nuclear localization activity can derive from the number of NLSs in the target-specific nuclease and / or other peptide sequence, the specific NLS used, or a combination of these elements. Detection of accumulation in the nucleus can be performed by any suitable technique. For example, a detectable marker can be fused to the target-specific nuclease and / or peptide sequence so that the location within the cell can be visualized, such as by a combination of methods to detect the location of the nucleus (e.g., a nuclear-specific stain such as DAPI). Examples of detectable markers include fluorescent proteins (e.g., green fluorescent protein, or GFP, RFP, CFP) and epitope tags (HA tag, FLA tag, SNAP tag). Cell nuclei can also be isolated from the cells and their contents can then be analyzed by any suitable method of detecting proteins, such as immunohistochemistry, Western blot, or enzyme activity assays. Accumulation in the nucleus can also be determined indirectly.
[0094] In some aspects, the invention provides methods that include delivering one or more polynucleotides, such as one or more vectors described herein, one or more transcripts thereof, and / or one or more proteins transcribed therefrom, to a host cell. In some aspects, the invention further provides cells produced by such methods, and organisms (e.g., animals, plants, or fungi) that contain or are produced from such cells. In some embodiments, a Cas protein in combination with a guide sequence (optionally complexed) is delivered to the cell. Conventional viral and non-viral based gene transfer methods can be used to introduce nucleic acids into mammalian cells or target tissues. Such methods can be used to administer nucleic acids encoding target-specific nucleases and / or blunting enzymes to cells in culture or in a host organism. Non-viral vector delivery systems include DNA plasmids, RNA (e.g., transcripts of vectors described herein), naked nucleic acids, nucleic acids complexed with a delivery vehicle, such as liposomes, and ribonucleoproteins. Viral vector delivery systems include DNA and RNA viruses, which have episomal or integrated genomes after delivery to a cell.Reviews of gene therapy procedures include Anderson, Science 256:808-8313 (1992); Navel and Felgner, TIBTECH 11:211-217 (1993); Mitani and Caskey, TIBTECH 11:162-166 (1993); Dillon, TIBTECH 11:167-175 (1993); Miller, Nature 357:455-460 (1992); Van Brunt, Biotechnology 6(10):1149-1154 (1988); Vigne, Restorative Neurology and Neuroscience 8:35-36 (1995); Kremer and Perricaudet, British Medical Bulletin 51(1):31-44 (1995); Haddada et al., in Current Topics in Microbiology and Immunology, Doerfler and Bohm (eds.) (1995); and Yu et al., Gene Therapy 1:13-26 (1994).
[0095] The target-specific nuclease and / or peptide sequence may be delivered using adeno-associated virus (AAV), lentivirus, adenovirus, or other viral vector types, or combinations thereof. In some embodiments, the Cas protein and one or more guide RNAs may be packaged in one or more viral vectors. In some embodiments, the targeted trans-splicing system is delivered via AAV as a split intein system similar to Levy et al. (Nature Biomedical Engineering, 2020, DOI:doi.org / 10.1038 / s41551-019-0501-5). In other embodiments, the target-specific nuclease and / or peptide sequence may be delivered via AAV as a trans-splicing system similar to Lai et al. (Nature Biotechnology, 2005, DOI:10.1038 / nbt1153). In some embodiments, the viral vector is delivered to the tissue of interest, for example, via intramuscular injection, while at other times the virus is delivered via intravenous, transdermal, intranasal, oral, mucosal, intrathecal, intracranial or other delivery methods. Such delivery can be via either a single dose or multiple doses. Those skilled in the art will appreciate that the actual dosage delivered herein can vary widely depending on a variety of factors, such as the vector selected, the target cell, organism or tissue, the general condition of the subject being treated, the degree of transformation / modification being pursued, the route of administration, the mode of administration, the type of transformation / modification being pursued, etc.
[0096] The use of RNA or DNA virus-based systems for delivery of nucleic acids takes advantage of the highly evolved process of targeting viruses to specific cells in the body and transporting the viral payload to the nucleus. Viral vectors can be administered directly to the patient (in vivo) or used to treat cells in vitro, and the modified cells can optionally be administered to the patient (ex vivo). Traditional virus-based systems can include retroviral, lentiviral, adenoviral, adeno-associated, and herpes simplex virus vectors for gene transfer. Integration into the host genome is possible with retroviral, lentiviral, and adeno-associated viral gene transfer methods, often resulting in long-term expression of the inserted transgene. In addition, high transduction efficiency has been observed in many different cell types and target tissues. Viral-mediated in vivo delivery of Cas13 and guide RNA provides a rapid and powerful technique to achieve precise mRNA perturbation in cells, especially in postmitotic cells and tissues.
[0097] In certain embodiments, delivery of target-specific nuclease and / or peptide sequences to cells is non-viral, hi certain embodiments, the non-viral delivery system is selected from ribonucleoproteins, cationic lipid vehicles, electroporation, nucleofection, calcium phosphate transfection, transfection via membrane disruption using mechanical shear forces, mechanical transfection, and nanoparticle delivery.
[0098] In some embodiments, host cells are transiently or non-transiently transfected with one or more vectors described herein. In some embodiments, cells are transfected as they naturally occur in a subject. In some embodiments, transfected cells are removed from a subject. In some embodiments, cells are derived from cells, e.g., cell lines, removed from a subject. Cell lines are available from a variety of sources known to those skilled in the art (see, e.g., American Type Culture Collection (ATCC), Manassas, VA). In some embodiments, cells transfected with one or more vectors described herein are used to establish new cell lines that contain one or more vector-derived sequences.
[0099] diagnosis The present disclosure provides target-specific nucleases for diagnostic applications. Diagnostic applications include, for example, but are not limited to, molecular, amino acid, nucleic acid, and derivatives thereof diagnostic agents (see, e.g., Harrington LB, Burstein D, Chen JS, Paez-Espino D, Ma E, Witte IP, Cofsky JC, Kyrpides NC, Banfield JF, Doudna JA. Programmed DNA destruction by miniature CRISPR-Cas14 enzymes. Science. 2018 Nov. 16; 362(6416):839-842. doi:10.1126 / science.aav4294. Epub 2018 Oct. 18. PMID:30337455; PMCID:PMC6659742; and Xiang X, Qian K, Zhang Z, Lin F, Xie Y, Liu Y, Yang Z. CRISPR-cas systems based molecular diagnostic tool for infectious diseases, which are incorporated by reference in their entirety. (See J Drug Target. 2020 Aug-Sep; 28(7-8):727-731. doi:10.1080 / 1061186X.2020.1769637. Epub 2020 May 26. PMID:32401064; PMCID:PMC7265108). In one example, target-specific nucleases can be used with DETECTR, a DNA endonuclease-targeted CRISPR transreporter technology for molecular diagnostics. This technology achieves high sensitivity in DNA detection by combining non-specific single-stranded deoxyribonuclease activation of Cas12 ssDNase with isothermal amplification to enable rapid and specific detection of biological agents such as viruses. In this assay, the crRNA-Cas12a complex binds to target DNA and induces indiscriminate cleavage of ssDNA that is coupled to a fluorescent reporter.In another example, target-specific nucleases can be combined with a fluorescence-based point-of-care (POC) device, in which Cas12a / crRNA detects and binds to the targeting DNA, and then the Cas12a / crRNA / DNA complex is activated and degrades the fluorescent ssDNA reporter, generating a signal.
[0100] kit The present disclosure provides kits for carrying out the methods. The present disclosure provides kits including any one or more of the elements disclosed in the methods and compositions above. In some embodiments, the kit includes a vector system and instructions for using the kit. In some embodiments, the kit includes a vector system including a regulatory element and a polynucleotide encoding a target-specific nuclease and / or peptide sequence. In some embodiments, the kit includes a viral delivery system of the target-specific nuclease and / or peptide sequence. In some embodiments, the kit includes a non-viral delivery system of the target-specific nuclease and / or peptide sequence. The elements may be provided individually or in combination and may be provided in any suitable container, e.g., a vial, bottle, or tube. In some embodiments, the kit includes instructions in one or more languages, e.g., in more than one language.
[0101] In some embodiments, the kit includes one or more reagents utilized in a process utilizing one or more of the elements described herein. The reagents may be provided in any suitable container. For example, the kit may include one or more reaction or storage buffers. The reagents may be provided in a form that is ready for use in a particular assay or that requires the addition of one or more other components prior to use (e.g., concentrated or lyophilized). The buffer may be any buffer, including, but not limited to, sodium carbonate buffer, sodium bicarbonate buffer, borate buffer, Tris buffer, MOPS buffer, HEPES buffer, and combinations thereof. In some embodiments, the buffer is alkaline. In some embodiments, the buffer has a pH of about 7 to about 10. In some embodiments, the kit includes one or more oligonucleotides corresponding to the guide sequence for insertion into a vector so as to be operably linked to the guide sequence and the regulatory element.
[0102] array The sequences of the target-specific nuclease, guide and nuclear localization signal (NLS) can be found in Table 1 below.
[0103] [Table 1] TIFF2024522764000002.tif227169TIFF2024522764000003.tif231168TIFF2024522 764000004.tif229166TIFF2024522764000005.tif231167TIFF2024522764000006.t if228168TIFF2024522764000007.tif230168TIFF2024522764000008.tif230168TIF F2024522764000009.tif224168TIFF2024522764000010.tif229167TIFF20245227640 00011.tif228167TIFF2024522764000012.tif225168TIFF2024522764000013.tif22 7167TIFF2024522764000014.tif225167TIFF2024522764000015.tif227168TIFF202 4522764000016.tif228166TIFF2024522764000017.tif229168TIFF20245227640000 18.tif229167TIFF2024522764000019.tif229167TIFF2024522764000020.tif229167
[0104] The percent identity of Cas12m with other Cas12 orthologs can be found in Tables 2-13 below.
[0105] [Table 2] TIFF2024522764000022.tif236161TIFF2024522764000023.tif228161TIFF2024522764000024.tif233161TIFF2024522764000025.tif81164
[0106] [Table 3] TIFF2024522764000027.tif 232168 TIFF2024522764000028.tif 236168 TIFF2024522764000029.tif 234168 TIFF2024522764000030.tif 83168
[0107]
Table 4
[0108]
Table 5
[0109]
Table 6
[0110]
Table 7
[0111]
Table 8
[0112]
Table 9
[0113]
Table 10
[0114]
Table 11
[0115]
Table 12
[0116]
Table 13
[0117]
Table 14
[0118]
Table 15
[0119]
Table 16
[0120]
Table 17
[0121] [Table 18] TIFF2024522764000097.tif223165TIFF2024522764000098.tif215166
[0122] [Table 19]
[0123] [Table 20] TIFF2024522764000101.tif218165TIFF2024522764000102.tif223166TIFF2024522764000103.tif121164
[0124] [Table 21] TIFF2024522764000105.tif224169TIFF2024522764000106.tif219168TIFF2024522764000107.tif20165 EXAMPLES
[0125] Several examples are contemplated, however, these examples are intended to be non-limiting.
[0126] [Example 1] Computational discovery of tiny CRISPR nucleases Computational discovery of miniature CRISPR nucleases was performed (Figures 1A-1D).
[0127] Novel small CRISPR nucleases from metagenomic samples were identified by computational discovery (Figure 1A). Initial panning of small CRISPR nucleases yielded orthologs, including 30 novel Cas12f orthologs, 20 novel Cas12j orthologs, and 45 novel Cas12m orthologs (Figure 1B). These orthologs contain a C-terminal RuvC domain indicative of a Cas12 system and a CRISPR array of two or more spacers with tandem repeats that fold with the appropriate secondary structure (Figure 1E). The Cas12f and as12m systems have readily identifiable putative tracrRNAs, found by homology searches of DRs against surrounding loci and secondary structure modeling / prediction to identify tracrRNA sequences with the best folding energy into the crRNA. The Cas12j system does not have any identifiable tracrRNA, and the Cas12m system has a identifiable tracrRNA. The new subclasses of Cas12 may or may not require tracrRNA.
[0128] FIG. 1C shows the size distribution of Cas12a, and FIG. 1D shows the size distribution of CasM orthologs.
[0129] [Example 2] PsaCas12f sgRNA construct The PasCas12f sgRNA construct was tested in human mammalian cells (Figure 4).
[0130] A panel of 24 sgRNA designs was tested against the pUC19 previously reported plasmid harboring PsaCasf. The sgRNA designs are disclosed in Table 1 and achieve editing down to approximately 0.5%. Experiments were performed with plasmid expression in HEK293FT for 48-72 h.
[0131] [Example 3] PsaCas12f sgRNA design based on sgRNA secondary structure The secondary structure of the sgRNA is important to enable specific and efficient recognition between Cas9 and the target sequence. To further improve the cleavage efficiency of the PsaCas12f-sgRNA complex, sgRNA variants were designed to contain genetic mutations that affect the secondary structure of the sgRNA as well as its interaction with the sgRNA-protein complex.
[0132] The predicted sgRNA secondary structure was obtained by using in silico structure determination. Stem loop 1-3 (SL1-3) were predicted via http: / / rna.tbi.univie.ac.at / . Stem loop 4 (SL4, interacting with crRNA) and stem loop 5 (SL5) were informed by Takeda et al., Mol Cell, 81(3):558-570 (2021). Figure 10A illustrates the obtained sgRNA secondary structure with SL1-SL3 marked by blue, red and green boxes, respectively.
[0133] Using this predicted sgRNA secondary structure, genetic mutations were engineered into SL1, SL2, SL3, SL4 or SL5. Figure 10B lists and annotates all the sgRNA variants designed (see also sequence listing in Table 14). Red indicates the introduced nucleobase changes, orange indicates the nucleobases that form the stem, and purple indicates the loop that was added to allow recruitment of the MS2 coat / protein.
[0134] We subsequently tested sgRNA variants using an in vitro luciferase reporter assay to assess whether secondary structure modifications of SL1-SL5 affected cleavage efficiency. Briefly, HEK293T cells were seeded and transfected with 25 ng of luciferase reporter, 100 ng of the different annotated CRISPR guides listed above, and 300 ng of PsaCas12f expression plasmid. 72 hours after transfection, media was collected from the cells and analyzed for luciferase expression.
[0135] The corresponding bar graphs in Figure 10C show the results of the reporter assay. Notably, certain modifications to SL1, SL2, SL3, SL4, or SL5 increased cleavage efficiency compared to the controls (control sgRNA constructs previously optimized using a different strategy and labeled "5pr_truncated4-7" and "best guide v2").
[0136] [Example 4] PsaCas12f sgRNA combination micromutation stem loop construct The sgRNA variants in Example 3 each targeted a different stem-loop region (SL1, SL2, SL3, SL4 or SL5). It was hypothesized that each stem-loop region could affect different functions (e.g., hairpin stability, transcription efficiency, protein interaction) and that a combination of the single stem-loop variants designed in Example 3 would further improve cleavage efficiency. Therefore, sgRNA variants were designed that contained a combination of modifications from sgRNA variants and single modifications of specific stem-loop regions (also called "combination constructs"). The purpose of the sgRNA combination stem-loop variants was to increase folding and Cas12f interaction (e.g., increasing GC content, correcting sgRNA truncations / mismatches in stem-loops, removing premature termination signals).
[0137] The combination constructs are presented in Table 16. Figure 11A shows the resulting performance of the combination constructs relative to the control by in vitro luciferase reporter assay. Surprisingly, certain combinations, e.g., the construct labeled "SL1_modified_+increased interaction with crRNA_22", produced enhanced cleavage efficiency (about 0.035% RLU cleavage) relative to the single modified construct labeled "SL1_modified_1" (about 0.025% RLU cleavage), comparing Figure 10C with Figure 11A.
[0138] Combinatorial constructs were then tested for cleavage efficiency at the EMX1 (empty spiracles-like protein 1) locus, either as double variants with modifications of stem loops 1 and 2 (labeled 2x combination in Figure 11B) or as quadruple variants with modifications of stem loops 1, 2, 3, and 5 (labeled 4x combination in Figure 11B).
[0139] Briefly, to measure the cleavage efficiency at the EMX1 locus, 100ng of different CRISPR guides annotated in Table 16 above and 300ng of PsaCas12f expression plasmid were transfected into HEK293FT cells. 72 hours after transfection, cells were harvested for genomic DNA, and primers that amplify the EMX1 genomic locus were used to amplify the genomic region of this locus. Next generation sequencing (NGS) was then performed on these amplified gDNAs, and the insertion / deletion profiles caused by Cas12f with different guides were analyzed with CRISPResso.
[0140] FIG. 11B shows the results of the editing efficiency at the EMX1 locus of the combination constructs shown above. Notably, in the 4× combination constructs tested, the construct labeled “SL5_4+cr21+SL1_8” had a greater editing efficiency at the EMX1 locus than the control constructs with either a single stem-loop modification or no stem-loop modification. It is not entirely clear why certain combination constructs work better than other combinations. For example, compare “SL2_4+SL1_1” with “SL2_4+SL1_3” for EMX1 editing efficiency at 2× combination. One hypothesis is that certain base pair combinations do not provide optimal sgRNA folding / sgRNA-protein interactions, and these occurrences are difficult to predict in silico.
[0141] The best sgRNA combination mutant stem-loop constructs, referred to in Figures 11A and 11B as (1) scaffold "version 2", (2) "version 3.1, SL1_modification_8 + increased interaction with crRNA_21 or SEQ ID NO: 203" and (3) "v.3.2, SEQ ID NO: 198", were subsequently tested with 30 different PsaCas12f mutants against controls by in vitro luciferase reporter assay to test the robustness of the sgRNA scaffolds as shown in Figure 11C. Notably, the scaffold "v.3.2" containing the mutation combination "SL1_8" and "interaction with cRNA_22" modifications performed well in the panel of PsaCas12f mutants tested, demonstrating the robustness of "v.3.2" as an sgRNA scaffold.
[0142] [Example 5] Spacer optimization of PsaCas12f sgRNA scaffold version 3.2 The sgRNA spacer sequence can affect the degree of target specificity and off-target activity. Figure 12A is a schematic of the sgRNA scaffold version 3.2 with the location of the spacer sequence at the 3' end highlighted. This experiment was designed to test the cleavage efficiency of the sgRNA v.3.2 scaffold of Example 4 by varying the nucleotide length of the sgRNA spacer sequence. To test the spacer length, the version 3.2 sgRNA scaffold was tested by in vitro luciferase reporter assay with spacer sequence lengths of 2, 3, 18, 19, 20, 21, 22, 23, 24 and 25 base pairs against the control. Figure 12B shows that using the v3.2 sgRNA scaffold with PsaCas12f, the highest cleavage efficiency was achieved using a 21 bp spacer sequence for this particular target. Although 22 bp, 20 bp, 19 bp and 18 bp also worked, 21 bp showed the best gene editing. Thus, for PsaCas12f version 3.2 sgRNA, 20 bp or 21 bp is sufficient to allow sufficient base pairing before cleavage.
[0143] [Example 6] PsaCas12f with sgRNA scaffold version 3.2 is more effective than UnCas12f (Cas14a1) PsaCas12f with sgRNA scaffold version 3.2 described in Example 4 was then compared to a different Cas12f protein called Un1Cas12F1 (also called Cas14a1) that is similarly small and has good on-target efficiency at either the HBB (hemoglobin subunit beta) or RNF2 (ring finger protein 2) genomic loci. Un1Cas12F1 is a protein identified in uncultured archaea (Un1).
[0144] Briefly, 100ng of different CRISPR guides based on scaffold version 2 with different spacer lengths (e.g., staggered 24 indicates a spacer length of 24nt) according to the description annotated in Table 17 and 300ng of PsaCas12f expression plasmid are transfected into HEK293FT cells. Two spacer sequences targeting either RNF2 or HBB genomic loci were designed in the sgRNA v.3.2 scaffold. 72 hours after transfection, cells were harvested for genomic DNA, and the gDNA of this locus was amplified using primers that amplify the corresponding genomic locus. Next-generation sequencing (NGS) was then performed on these amplified gDNAs, and the insertion / deletion profiles caused by Cas12f with different guides were analyzed with CRISPResso.
[0145] Figure 13 shows that PsaCas12f with sgRNA scaffold version 3.2 outperformed Un1Cas12f with nbt scaffold in terms of indel activity (insertion / deletion formation) at both sites tested in the HBB locus (g1 and g2) and at one site in the RNF locus (g4). Thus, PsaCas12f with sgRNA scaffold version 3.2 enables efficient indel formation and may be a useful tool in a wide range of genome engineering applications.
[0146] Example 7 PsaCas12f NLS construct The PsaCas12f nuclear localization signal (NLS) construct was tested in HEK293FT human mammalian cells (Figures 5A-5D).
[0147] A panel of 15 NLS designs were tested using the best two guide sequences from Example 2 fused to PsaCas12f on the pUC19 reported plasmid. The NLS designs are disclosed in Table 1, achieving editing down to about 0.1% (Figure 5A). Experiments were performed with plasmid expression in HEK293FT for 48-72 hours. Sequencing traces show bona fide editing as illustrated in Figures 5B-5E. Editing with PsaCas12f (NLS14) with sgRNA (Figure 5B) or non-targeted guide (Figure 5C) shows obvious deletions (purple) and insertions (red). Editing with PsaCas12f (no NLS) with sgRNA (Figure 5D) or non-targeted guide (Figure 5E) also shows obvious deletions (purple) and insertions (red).
[0148] Internal NLS signals can allow for better design of proteins to be delivered via virus-like particles (Banskota et al., Cell 185(2):250-265 (2022)) or allow for inducible NLS signals following conformational changes (Saleh et al., Exp Cell Res 260(1):105-115 (2000)). Thus, an internal protein NLS sequence from SV40 (simian virus 40) was fused to random positions of PsaCas12f as shown in FIG. 14 and annotated in Table 18. These constructs were tested for indel activity in the EMX genomic locus.
[0149] Briefly, 72 hours after transfection, cells were harvested for genomic DNA, and gDNA from the corresponding EMX genomic locus was amplified using primers that amplify this locus. Next generation sequencing (NGS) was then performed on these amplified gDNAs, and insertion / deletion profiles were analyzed with CRISPResso. The internal NLS signals labeled "NSL_2", "NSL_3", "NSL_5", and "NSL_6" had higher indel-inducing potential at the EMX locus than wild-type PsaCas12f (labeled "pDF0106") flanked by two NLS sequences at the N- and C-termini, as shown in FIG. 14. Thus, the internal NLS signal can provide an alternative localization to the flanking NLS signal and still maintain optimal gene editing activity. The internal NLS signal can be advantageous, for example, when the N- or C-terminal NLS fusion interferes with protein function.
[0150] Example 8: CRISPR editing with PsaCas12f and guide RNA delivered by adeno-associated virus (AAV) Adeno-associated virus (AAV) is a safe vehicle approved by the US Food and Drug Administration for gene therapy, so AAV-loadable CRISPR tools are advantageous. AAV has a limited payload size of <4.7kb, which prevents most CRISPR tools from clinical application. Therefore, this example validates the AAV delivery of PsaCas12f-sgRNA.
[0151] Briefly, PsaCas12f with the best NLS conformation (flanking SV40NLS) was cloned into AAV ITR together with a guide targeting the RUNX1 (runt-related transcription factor 1) genomic locus. The plasmid was then transfected into HEK293FT cells with an AAV helper plasmid to generate AAV particles. The AAV particles in the medium of the producer cell line were collected and subsequently added to HEK293FT cells. Four days after transduction, the indel profile at the RUNX1 locus was analyzed by NGS.
[0152] As shown in Figure 15, AAV loaded with PsaCas12f+ guide had an indel frequency of approximately 10-14% at the RUNX1 locus, which increased proportionally with the amount (1, 5 or 25 μl) transduced into HEK293 cells. This experiment demonstrates that PsaCas12f can be efficiently expressed from AAV particles while retaining the ability to induce cleavage at the genomic target.
[0153] [Example 9] PsaCas12f with guide crRNA / tracrRNA PsaCas12f with crRNA / tracrRNA guides was screened at different local free energy minima (Figure 6). The results of PsaCas12f indicate that many crRNA / tracrRNA designs must be screened at various local minimum free energies to find the optimal combination of activity in bacterial or mammalian protein lysates. It was found that a 20nt DR and a 90nt tracrRNA provide optimal activity for dsDNA cleavage and that these can be combined for sgRNA. These designs showed that computational and experimental RNA screening yields optimal designs and that the sgRNA has a significant effect on activity.
[0154] [Example 10] Genome editing using Cas12f family members Cas12f family members were tested for genome editing (Figure 7). Testing of Cas12f family members for indel generation in EMX1 results in editing efficiencies above background.
[0155] Example 11 Screening of a panel of 12 Cas12f orthologs A panel of 12 novel Cas12f orthologs ranging in size from 400 to 800 amino acids were screened. To maintain the correct small RNA species in these orthologs, noncoding regions from the surrounding loci were cloned along with the Cas12f gene (Figure 8A). Generation of lysates of these samples allowed for testing of in vitro cleavage of the degenerate PAM library, enriching for cleavage fragments, and determining the PAM. In all 12 proteins, one of the orthologs, Cas12f from Pseudomonas aeruginos (a proteobacteria), a 586-residue protein, had substantial cleavage activity as determined by this high-throughput PAM screen. PAM characterization determined the motif of PsaCas12f to be TTR (Figure 8B). In addition, small RNA sequencing of these purified proteins could determine the mature isoforms of processed crRNA and tracrRNA (Figure 8C), yielding a native DR length of 31 nt and a tracrRNA length of 97 nt. Finally, PAM of PsaCas12f of the fixed sequence target was verified, demonstrating detectable in vitro cleavage by gel reading (Figure 8D). Characterization of PsaCas12f and the corresponding RNA species, as well as other effectors selected from high-throughput screening, can be optimized for activity by guide RNA engineering.
[0156] [Example 12] PsaCas12f circular permutation Although Cas nucleases have not evolved function into modular DNA-binding scaffolds that optimize Cas nucleases by fusing them to functional protein domains, the use of linkers can allow for controlled nuclease activity and broaden the use of Cas nucleases as genetic tools. Oakes et al., Cell, 176(2):254-267 (2019). One method to alter CRISPR structures to allow fusion to other protein domains is protein circular permutation (CP). Ibid. CP is a topological rearrangement of a protein's primary sequence that connects the N- and C-termini with a peptide linker while simultaneously splitting the sequence at a different position to create new adjacent N- and C-termini. Yu and Lutz, Trends Biotechnol, 28:18-25 (2011).
[0157] To test whether the PsaCas12f protein described above can undergo circular permutation without compromising functional activity, the PsaCas12f sequence was split at different positions using (GGS)6 peptide linkers to create new adjacent N- and C-termini, as shown in Table 15 (see also the bottom schematic of Figure 16A).
[0158] The circularly permuted constructs listed in Table 21 were then tested for editing efficiency by using the in vitro luciferase reporter assay described above or by testing for indel formation at the RUNX1 genomic locus, as shown in Figures 16A and 16B, respectively.
[0159] Briefly, for in vitro luciferase reporter assay, 25ng of Gluc reporter, 100ng of CRISPR guide and 300ng of normal PsaCas12f expression plasmid (control, labeled pDF0106) or different circular permutations of protein encoding plasmid were transfected into HEK293FT cells. 72 hours after transfection, media was collected from cells and analyzed for luciferase expression. For assessment of indel formation at the RUNX1 genomic locus, the same panel of circular permutations of PsaCas12f protein was tested with guides targeting the genomic RUNX1 locus. Cell transfection conditions were the same as for in vitro luciferase, PCR was used to amplify the genomic locus with RUNX1, and indel efficacy was estimated by CRISPResso.
[0160] Of note, some circular permutations of PsaCas12f are functional, allowing different positioning of the N- and C-termini. Interestingly, the editing efficiency varies depending on the guide used (compare the editing efficiencies in Figures 16A and 16B).
[0161] Example 13: PsaCas12f sequence optimization via machine learning The wild-type PsaCas12f sequence was sent to a machine learning model (Facebook Evolutionary Scale Modeling (ESM), https: / / github.com / facebookresearch / esm) to predict protein point mutations that could result in high editing efficiency. That is, the original WT sequence was used as the input of the ESM model. The output of the ESM model was a single vector (1x1280), which was successfully used as the input of a linear regression model to predict the output, which was the indel formation rate. The new mutation model of the protein was modeled in a similar manner to predict indels and subsequently tested in vitro.
[0162] 48 different point mutations were compared to the integration of the best-guide v3.2 scaffold described above and spacer-targeted RNF2 (tatgagttacaacgaacacctc) (see Table 18) targeting the genomic RNF2 locus. 72 hours after transfection of a panel of PsaCas12f variants containing single point mutations (+sgRNA), the RNF2 locus was PCR amplified and subjected to NGS. Indel profiles were quantified by CRISPResso in all mutants.
[0163] In the panel of point mutations tested, a point mutation at position 333 of PsaCas12f from lysine to valine dramatically increased the cleavage efficiency of PsaCas12f, as shown in Figure 17.
[0164] Those skilled in the art will appreciate further features and advantages of the present invention based on the above-described embodiments. Accordingly, the present invention should not be limited by what has been particularly shown and described, except as indicated by the appended claims. All publications and references cited herein are expressly incorporated herein by reference in their entirety.
Claims
1. (a) a target-specific nuclease comprising an amino acid sequence having 70% identity to an amino acid sequence selected from the group consisting of SEQ ID NO: 1, and (b) a guide RNA (gRNA) comprising a nucleic acid sequence having 70% identity to a nucleic acid sequence selected from the group consisting of SEQ ID NOs: 20-43, 61-79, 145-198, and 246-272 A composition comprising: The target-specific nuclease and the gRNA form a complex that can bind to a DNA target having a certain degree of complementarity to the gRNA sequence, nick, and / or cleave.
2. The composition according to claim 1, wherein the DNA target is single-stranded DNA or double-stranded DNA.
3. The composition according to claim 1, wherein the degree of complementarity between the DNA target and the gRNA ranges from about 50% to about 99% or more.
4. The composition according to claim 1, wherein the gRNA is a single guide RNA (sgRNA) or a dual guide (dgRNA).
5. The composition according to claim 1, wherein the DNA target is in a prokaryotic cell or a eukaryotic cell.
6. The composition according to claim 1, further comprising an adeno-associated virus vector (AAV) for delivery in vivo.
7. The composition according to claim 1, wherein the gRNA specifically binds to a protospacer adjacent motif (PAM) in the target DNA.
8. The composition according to claim 1, wherein the target-specific nuclease and the gRNA are encoded by a nucleic acid molecule.
9. The composition according to claim 1, further comprising one or more spacers in the gRNA sequence.
10. The composition according to claim 1, wherein the DNA target is in a mammalian cell.
11. The composition according to claim 1, further comprising an epigenetic modifier. **Claim 12**: The composition according to claim 1, wherein the DNA target further comprises a protospacer adjacent motif comprising NNNNGATT, NNNNGNNN, NNG, NG, NGAN, NGGNG, NGAG, NGCGA, NAAG, NGN, NRN, NNGRRN, NNNRRT, TTTN, TTTW, TYCV, TATV, TYCV, TTN, KYTV, TYCV, TBN, any variant thereof, or any combination thereof. **Claim 13**: The composition according to claim 1, wherein the target-specific nuclease is fused to one or more nuclear localization signals (NLSs) at either the 3′ end, 5′ end, or internally. **Claim 14**: The composition according to claim 1, further comprising one or more adeno-associated virus (AAV) vectors. **Claim 15**: The composition according to claim 1, wherein the target-specific nuclease is circularly permuted with a peptide linker.