Type v crispr effector RNA-guided endonuclease variants
Engineered Type V CRISPR nuclease variants with optimized guide RNAs enhance specificity and on-target activity, addressing the limitations of existing Type V CRISPR nucleases for precise genome editing.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- INTEGRATED DNA TECHNOLOGIES INC
- Filing Date
- 2025-11-18
- Publication Date
- 2026-05-28
AI Technical Summary
Existing Type V CRISPR nucleases exhibit lower specificity and higher off-target activity, limiting their effectiveness in genome editing, particularly in AT-rich regions.
Development of engineered Type V CRISPR nuclease variants with enhanced specificity and on-target activity through amino acid substitutions and optimized guide RNAs, utilizing a 42mer crRNA variant for precise genome editing.
The engineered Type V CRISPR nuclease variants demonstrate improved on-target potency with reduced off-target effects, offering a safer and more efficient genome editing tool for research and therapeutic applications.
Smart Images

Figure US2025055856_28052026_PF_FP_ABST
Abstract
Description
Attorney Docket No. 6391-0035W0011Type V CRISPR Effector RNA-guided Endonuclease VariantsThis application claims benefit of U.S. Serial No. 63 / 722,441, filed November 19, 2024, and U.S. Serial No. 63 / 860,031, filed August 8, 2025, the entireties of which are incorporated herein by reference.INCORPORATION'S Y-REFERENCE OF MATERIAL SUBMITTED ELECTRONICALLY
[0001] Incorporated by reference in its entirety herein is a computer-readable nucleotide / amino acid sequence listing submitted concurrently herewith and identified as follows: One 236,200 Bytes Extensible Markup Language (XML) file named “6391- 0035W001” created on November 18, 2025.BACKGROUND
[0002] Clustered regularly interspaced short palindromic repeats (CRISPR)-Cas9 enzyme (Cas9) has been found to be a powerful genome editing tool. Class II type V CRISPR enzyme (Cpfl) is a relatively new RNA guided endonuclease that has emerged an alternative to Cas9. Type V CRISPR nuclease works in a different manner than Cas9 and, accordingly, is a better tool than Cas9 in certain genome editing situations. In some situations, “wild type” Type V CRISPR nuclease nuclease, from Acidaminococcus sp. BV3L6, or Lachnospiraceae bacterium ND2006, can experience lower specificity than desired and / or higher than desirable off-target activity.
[0003] The Type V CRISPR effector RNA-guided DNA endonuclease family of proteins has utility in synthetic biology, diagnostics, and genome engineering. As an alternative to the commonly used Streptococcus pyogenes Cas9 (SpCas9), this nuclease recognizes TTTV (V = A / G / C) PAM sequences, which permits editing in AT-rich regions of the human genome. While this family of targeted CRISPR nucleases demonstrates high-fidelity genome editing with low off target cleavage, it suffers from poor on-target potency with certain targets, cell types, and delivery modalities. Because of this deficiency, we searched in metagenomic databases, which house many divergent Cas systems that have yet to be characterized, for a type V CRISPR effector alternative. After generating editing data from ~75 bioprospected nucleases, 4 were found to be active in mammalian cells. The Type V nuclease with the mostAttorney Docket No. 6391-0035W0012 potent editing in human cells was then engineered using end-to-end saturation mutagenesis with a bacterial selection and enrichment scheme to isolate more highly active variants. Substitutions in many amino acid positions were found to significantly increase on-target potency. Stacking of the most efficacious point mutations reveals an alternative Type V CRISPR effector with high genome editing efficiency with plasmid, mRNA, and RNP delivery modalities. We believe that this engineered Type V nuclease variant, which is highly efficient and uses a short 42mer crRNA variant, is an excellent solution for basic research, pre-clinical research, and future therapeutic applications.
[0004] Accordingly, there exists an unmet need for Type V CRISPR nucleases with high specificity (i.e., low off-target effects) and high on-target activity. Further, guide RNAs (gRNAs) are needed to guide the Type V CRISPR nucleases to the genomic target sites.
[0005] BRIEF SUMMARY
[0006] An aspect of the disclosure provides a Type V CRISPR nuclease, wherein the Type V CRISPR nuclease comprises an amino acid sequence that has at least 80% sequence identity to at least one of SEQ ID NOs: 63-66, SEQ ID NO: 114, SEQ ID NO: 116, and SEQ ID NO: 118.
[0007] A further aspect of the disclosure provides a Type V CRISPR nuclease ribonucleoprotein complex, wherein the Type V CRISPR nuclease ribonucleoprotein complex comprises (a) at least one guide RNA, and (b) a Type V CRISPR nuclease of an aspect of the disclosure.
[0008] An additional aspect of the disclosure provides methods for manipulating DNA in a cell, the method comprising contacting the DNA in the cell with at least one Type V CRISPR nuclease ribonucleoprotein complex of an aspect of the disclosure.
[0009] Another aspect of the disclosure provides an isolated or purified nucleic acid comprising a nucleotide sequence encoding a protein according to any aspect of the present disclosure.
[0010] An aspect of the disclosure provides a recombinant expression vector comprising a nucleic acid according to an aspect of the present disclosure.
[0011] A further aspect of the disclosure provides kits comprising (a) at least one guide RNA, and (b) at least one Type V CRISPR nuclease of an aspect of the disclosure.[00.12] Another aspect of the disclosure provides a nucleic acid guided nuclease system, the system comprising (a) a Type V CRISPR nuclease nucleic acid having at least 80% identity to least one of SEQ ID NOs: 59-62, 113, 115, and 117 that encodes a Type V CRISPR nucleaseAttorney Docket No. 6391-0035W0013 amino acid sequence having at least 80% sequence identity to at least one of SEQ ID NOs: 63- 66, 114, 116, and 118, and (b) a guide nucleic acid capable of complexing with the nucleic acid guided nuclease, wherein the guide nucleic acid comprises a constant region comprising at least 80% sequence identity to at least one of SEQ ID NOs: 90-112.BRIEF DESCRIPTION OF THE SEVERAL VIEWS OF THE DRAWINGS[0013 j FIG. 1 is a graph showing the relative editing efficiency of Type V CRISPR nuclease variants of an aspect of the disclosure in human cells over 2 target sites using guides specific to either AsCasl2a or LbCasl2a. Plasmids encoding each variant were delivered by lipofection, and the editing efficiency was measured by Next Generation Sequencing (NGS) 3 days post-delivery.
[0014] FIG. 2 is a graph showing the editing efficiency of Type V CRISPR nuclease variants of an aspect of the disclosure in human cells at the HPRT 38115 site. Guides were paired with specific constant regions (SEQ ID NOs: 90-112) to the specific Type V CRISPR nuclease variant. Plasmids encoding each variant were delivered by lipofection, and the editing efficiency was measured by T7EI assay 3 days post-delivery.
[0015] FIG. 3 is a graph showing the editing efficiency of Type V CRISPR nuclease variants of an aspect of the disclosure in human cells at the TRACI site. Guides were paired with specific constant regions (SEQ ID NOs: 90-112) to the specific Type V CRISPR nuclease variant. Plasmids encoding each variant were delivered by lipofection, and the editing efficiency was measured by T7EI assay 3 days post-delivery.
[0016] FIG. 4A-4D are a set of graphs showing the editing efficiency of Type V CRISPR nuclease mutants of an aspect of an disclosure in human cells at the HPRT 38115 site. Guides were paired with gRNA (SEQ ID NO: 30) having the constant region 10851pred_3 (SEQ ID NO: 92). Plasmids encoding each variant were delivered by lipofection, and the editing efficiency was measured ty T7EI assay 3 days post-delivery.[00.17] FIG. 5A-5D are a set of graphs showing the editing efficiency of Type V CRISPR nuclease mutants of an aspect of an disclosure in human cells at the TRACI site. Guides were paired with gRNA (SEQ ID NO: 30) having the constant region 10851pred_3 (SEQ ID NO: 92). Plasmids encoding each variant were delivered by lipofection, and the editing efficiency was measured ty T7EI assay 3 days post-delivery.
[0018] FIG. 6A shows schematic diagrams of the bioprospecting process. Bioprospecting is the study of the structure and function of entire nucleotide sequences isolated and analyzedAttorney Docket No. 6391-0035W0014 from all the organisms in a bulk sample, usually microbes. The development of data mining methods and genetic and analytical tools has enabled the discovery of an array of previously overlooked natural products. The analysis starts with a metagenomic database, which is publicly available. As an initial step in bioprospecting, search parameters are set, followed by multiple sequence alignments, leading to the visualization of the variants of interest, which can be pulled from for closer examination FIG. 6B are schematics showing that the bioprospecting process found several unique clusters of Cas9 and Type V CRISPR sequences. The long branches in the Type V CRISPR group suggest significant variation, which is good for the potential of discovering unique Type V CRISPR variants that may have future utility in CRISPR applications. FIG. 6C is a chart showing the cleavage efficiencies of Type V CRISPR variants found through bioprospecting was measured utilizing NGS; the HPRT 38115 site was targeted using guides specific for AsCasl2a, including the constant region. The top hit from the bioprospecting is referred to as TEFO. FIG. 6D is a chart showing the cleavage efficiencies of Type V CRISPR variants found through bioprospecting was measured utilizing NGS; the HPRT 38330 site was targeted using guides specific for AsCasl2a, including the constant region. The top hit from the bioprospecting is referred to as TEFO.
[0019] FIG. 7A is a chart showing percent editing of TEFO targeting the HPRT 38115 site using several constant region patterns specific to TEFO. The percent editing was measured through T7 cleavage following transfection of plasmid with the TEFO specific guides in HEK293 cells. FIG. 7B is a chart showing percent editing of TEFO targeting the TRACI site using several guide patterns specific to TEFO. The percent editing was measured through T7 cleavage following transfection of plasmid with the TEFO specific guides in HEK293 cells.
[0020] FIG. 8 is a diagram showing a screening overview used to identify mutations in TEFO that improve activity. Saturated mutagenic plasmid libraries encoding TEFO with single amino acid changes at every location were generated and delivered into bacteria alongside guide RNAs targeting an arabinose inducible toxin expression plasmid. Bacteria that express a variant of TEFO that is more efficient at cleaving the toxin plasmid have a survival advantage under arabinose induction. Plasmid DNA was collected and purified after every subsequent screening round, followed by NGS sequencing of pre- and post-selection DNA on a NextSeq 2000® platform (Illumina®). Mutations showing large enrichment, when compared to the prescreened library, were selected and cloned into a mammalian expression plasmid backbone for additional testing.Attorney Docket No. 6391-0035W0015
[0021] FIG. 9A is a chart showing a comparison of the editing percentage of wildtype (WT) TEFO to the single mutations found through the bacterial screening at the HPRT 38115 site. The percent editing was measured through NGS following Lonza nucleofection of plasmids and TEFO specific guide in HEK293 cells. FIG. 9B is a chart showing a comparison of the editing percentage of wildtype (WT) TEFO to the single mutations found through the bacterial screening at the TRAC 1 locus. The percent editing was measured through NGS following Lonza nucleofection of plasmids and TEFO specific guide in HEK293 cells.
[0022] FIG. 10 is a chart showing a comparison of the editing percentage of wildtype (WT) TEFO to the stacked mutations found through the bacterial screening at the TRACI locus. The percent editing was measured through NGS following Lonza nucleofection of plasmids and TEFO specific guide in HEK293 cells.
[0023] FIG. 11 is a chart showing the results of testing the cleavage ability of the Top eleven 10851 mutations found in the bacterial screen at the HPRT 38115 site in HEK293 cells using the guide specific to 10851 (Guide 3). Nucleofection of250ng of plasmid and 2uM guide, processed after 72 hours for T7 Assay.DETAILED DESCRIPTION OF THE INVENTION
[0024] An aspect of the disclosure provides a Type V CRISPR nuclease, wherein the Type V CRISPR nuclease comprises an amino acid sequence that has at least 80% sequence identity (i.e., at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identity) to at least one of SEQ ID NOs: 63-66, 114, 116, and 118.
[0025] As compared to CRISPR / Cas9 systems, Type V CRISPR nuclease systems do not require tracer RNA (tRNA). This allows for a simpler, cheaper CRISPR system that is smaller and therefore easier to insert into a cell. In addition, Type V CRISPR nucleases can identity a T-rich protospacer adjacent motif (PAM) site (e.g., TTTV), unlike Cas9 which requires a G- rich PAM site. Further, Type V CRISPR nuclease provides a genomic cut that is off-set, allowing for more precise placement of insertions (Cas9 nuclease produces a double stranded break with blunt ends). Type V CRISPR nucleases also cut further from the PAM site as compared to Cas9 nucleases, thereby not disrupting the PAM site and allowing for additional rounds of DNA cleavage at the desired location.Attorney Docket No. 6391-0035W0016
[0026] Although there are many advantages that a Type V CRISPR nuclease system has over a CRISPR / Cas9 system, there is still room for further improvement in the Type V CRISPR nuclease system. For example, the Type V CRISPR nucleases of the present disclosure are able to provide increased on-target activity with decreased off-target activity as compared to wild type Type V CRISPR nucleases, therefore providing a safe and effective editing tool. Further, the gRNAs of the present disclosure can be used to precisely guide the Type V CRISPR nucleases to the targeted region of the DNA. Due to abundant options and the unpredictability in the field, it was surprising that the gRNAs of an aspect of the disclosure are so effective with the Type V CRISPR nucleases of an aspect of the disclosure in the systems. The gRNAs are from a different species than the Type V CRISPR nucleases and would not be expected to so effectively cleave the DNA. Accordingly, the Type V CRISPR nucleases, gRNAs, and systems of the present disclosure provide an advancement in gene editing.
[0027] In an aspect of the disclosure, the Type V CRISPR nuclease comprises an amino acid sequence that has at least 95% sequence identity (i.e., at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identity) to SEQ ID NO: 63. In an aspect of the disclosure, the Type V CRISPR nuclease comprises an amino acid sequence that has at least 95% sequence identity (i.e., at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identity) to SEQ ID NO: 64. In an aspect of the disclosure, the Type V CRISPR nuclease comprises an amino acid sequence that has at least 80% sequence identity (i.e., at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identity) to SEQ ID NO: 65. In an aspect of the disclosure, the Type V CRISPR nuclease comprises an amino acid sequence that has at least 80% sequence identity (i.e., at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identity) to SEQ ID NO: 66. In an aspect of the disclosure, wherein the Type V CRISPR nuclease comprises an amino acid sequence that has at least 80% sequence identity (i.e., at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identity) to SEQ ID NO: 114. In an aspect of the disclosure, wherein the Type V CRISPR nuclease comprises an amino acid sequence that has at least 80% sequence identity (i.e., at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identity) to SEQ ID NO: 116. In an aspect of the disclosure, wherein the Type V CRISPRAttorney Docket No. 6391-0035W0017 nuclease comprises an amino acid sequence that has at least 80% sequence identity (i.e., at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identity) to SEQ ID NO: 118.
[0028] In an aspect of the disclosure, the amino acid sequence that has at least 80% sequence identity (i.e., at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identity) to at least one of SEQ ID NOs: 63-66, 114, 116, and 118 is human codon pair optimized. Without being bound to any particular theory or mechanism, it is believed that codon optimization of the sequence increases the translation efficiency of the mRNA transcripts. Codon optimization of the nucleotide sequence may involve substituting a native codon for another codon that encodes the same amino acid, but can be translated by tRNA that is more readily available within a cell, thus increasing translation efficiency. Optimization of the nucleotide sequence may also reduce secondary mRNA structures that would interfere with translation, thus increasing translation efficiency.
[0029] An aspect of the disclosure provides a Type V CRISPR nuclease ribonucleoprotein complex, wherein the Type V CRISPR nuclease ribonucleoprotein complex comprises: (a) at least one guide RNA (gRNA); and (b) a Type V CRISPR nuclease of an aspect of the disclosure.
[0030] The at least one gRNA can be any suitable gRNA that directs the Type V CRISPR nucleases of an aspect of the disclosure to the target sequence / site in the genome. The gRNAs have two regions: (a) the constant region (loop domain, at the 5’ end) and (b) the target specific region (protospacer domain, at the 3' end of the gRNA). Accordingly, an aspect of the disclosure provides a gRNA comprising a constant region and a target specific region. In an aspect of the disclosure, the gRNAs have a constant region with at least 80% sequence identity (i.e., at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identity) to the sequences of SEQ ID NOs: 67-89. Table 1 below shows the alignment of SEQ ID NOs: 67, 71, 77, and 82 to the constant region for A.s. Casl2a (SEQ ID NO: 119). Table 2 below shows the SEQ ID NOs: 67-89 for DNA constant regions of an aspect of the disclosure. Table 3 below shows SEQ ID NOs: 90-112 for r RNA constant regions of an aspect of the disclosure. Table 4 belwo shows the sequences of Type V CRISPR nuclease variants. Table 5 below shows the sequcne of a 10851 Guide RNA, 3PS design, which works well with plasmid and mRNA.Attorney Docket No. 6391-0035W0018Table 1Table 2Attorney Docket No. 6391-0035W0019Table 3 (RNA)Attorney Docket No. 6391-0035W00110Attorney Docket No. 6391-0035W00111Table 4. Type V CRISPR Nuclease VariantsAttorney Docket No. 6391-0035W00112Attorney Docket No. 6391-0035W00113Attorney Docket No. 6391-0035W00114Attorney Docket No.6391-0035W00115Attorney Docket No. 6391-0035W00116Attorney Docket No. 6391-0035W00117Attorney Docket No. 6391-0035W00118Attorney Docket No. 6391-0035W00119Attorney Docket No. 6391-0035W00120Attorney Docket No.6391-0035W00121Attorney Docket No. 6391-0035W00122Attorney Docket No. 6391-0035W00123Attorney Docket No.6391-0035W00124Attorney Docket No. 6391-0035W00125Attorney Docket No. 6391-0035W00126Attorney Docket No. 6391-0035W00127Attorney Docket No. 6391-0035W00128Attorney Docket No. 6391-0035W00129Attorney Docket No. 6391-0035W00130Attorney Docket No.6391-0035W00131Attorney Docket No. 6391-0035W00132Attorney Docket No.6391-0035W00133Attorney Docket No. 6391-0035W00134Attorney Docket No. 6391-0035W00135Attorney Docket No. 6391-0035W00136Attorney Docket No. 6391-0035W00137Attorney Docket No. 6391-0035W00138Attorney Docket No.6391-0035W00139Attorney Docket No. 6391-0035W00140Tabel 5. Guide RNA(underlined is constant region, * is phosphorothioate linkages)
[0031] In an aspect of the disclosure, the Type V CRISPR nuclease ribonucleoprotein complex is active in a Type V CRISPR nuclease system. The Type V CRISPR nuclease system can be used to cleave target double-stranded DNA (dsDNA) (cis-activity) and indiscriminately degrade single-stranded DNA (ssDNA) (trans-activity). Type V CRISPR nuclease comprises three structural regions: the recognition (Rec) lobe, the nuclease (Nuc) lobe (including the RuvC, bridge helix (BH), and Nuc domains), and a connection region (including the Wedge (WED), and PAM-interacting (PI) domains) between the two lobes. It has been shown that Type V CRISPR nucleases use a single active site located in the RuvC domain to cleave both strands of the target dsDNA in cis-activity as well as non-specific ssDNA in trans-activity, following a well-defined order in the reaction steps. When Type V CRISPR nuclease binds to the target sequence, Type V CRISPR nuclease undergo a closed-to-open conformational change and the protospacer region of the gRNA, or crRNA, forms Watson-Crick base pairs with its complementary target strand (TS) in the target DNA to form an R-loop. Upon R-loop formation, the constant region or non-target strand (NTS) is immediately accessed by the active site in the RuvC domain, nicked, and trimmed by five nucleotides, which enables the TS to reach the active site and to be cleaved. Specifically, the target sequence comprises approximately 21 nucleotides and is adjacent to a PAM sequence. The Type V CRISPR nucleases recognizes a unique PAM sequence of TTTV (wherein V is A, C, or G). Type V CRISPR nucleases cut the strand which contains the PAM at a location 18-19 bases from theAttorney Docket No. 6391-0035W001413’ end of the PAM. The cut on the opposite strand of the DNA is 23 bases from the PAM, resulting in a 5' overhang of 4-5 bases.
[0032] The Type V CRISPR nuclease ribonucleoprotein complex described herein may also be referred to as an effector complex.
[0033] In an aspect of the disclosure, the Type V CRISPR nuclease endonuclease system using the Type V CRISPR nucleases and gRNAs of an aspect of the disclosure provides improved on-target editing activity as compared to a Type V CRISPR nuclease system utilizing Type V CRISPR nuclease from Acidaminococcus sp. BV3L6, under similar conditions. In an aspect of the disclosure, the Type V CRISPR nuclease system using the Type V CRISPR nucleases and gRNAs of an aspect of the disclosure provides improved on-target editing activity as compared to a Type V CRISPR nuclease system utilizing Type V CRISPR nuclease from Lachnospiraceae bacterium ND2006, under similar conditions. In a further aspect, the Type V CRISPR nuclease system provides improved on-target editing activity as compared to a endonuclease system utilizing type II CRISPR RNA-guided endonuclease Cas9 (NCBI Gene ID: 69900935; NZ_LS483338.1). In an aspect of the disclosure, the guide RNA comprises a nucleotide sequence that has at least 80% sequence identity (i.e., at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identity) to at least one of SEQ ID NOs: 9-54 and has at least 80% sequence identity (i.e., at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identity) to the constant regions of SEQ ID NOs: 90-112. In a further aspect, the guide RNA comprises a nucleotide sequence that has at least 80% sequence identity (i.e., at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identity) to at least one of SEQ ID NOs: 1-4.
[0034] One skilled in the art can compare the Type V CRISPR system of an aspect of the disclosure to another CRISPR / Cas System and determine which system has more on-target activity (and less off-target activity) using known methods. In an aspect of the disclosure, the Type V CRISPR system of an aspect of the disclosure has at least 5 % (i.e., at least about 10%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 100%, at least aboutAttorney Docket No. 6391-0035W00142125%, at least about 150%, at least about 175%, at least about 200%, at least about 250%, at least about 300%, at least about 400%, at least about 500%, etc.) more on-target activity as compared to a currently commercially available CRISPR system, a CRISPR / Cas9 system, or a Type V CRISPR nuclease wild type system. In an aspect of the disclosure, the Type V CRISPR nuclease system of an aspect of the disclosure has at least 5 % (i.e., at least about 10%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 100%, at least about 125%, at least about 150%, at least about 175%, at least about 200%, at least about 250%, at least about 300%, at least about 400%, at least about 500%, etc.) less off-target activity as compared to a currently commercially available CRISPR system, a CRISPR / Cas9 system, or a Type V CRISPR nuclease wild type system. Another aspect of the disclosure provides methods for manipulating DNA in a cell, the method comprising contacting the DNA in the cell with at least one Type V CRISPR nuclease ribonucleoprotein complex of an aspect of the disclosure. The DNA manipulation can be any suitable modification, including deletions, mutations, inversions, and substitutions.
[0035] An aspect of the disclosure provides an isolated or purified nucleic acid comprising a nucleotide sequence encoding the protein according to an aspect of the present disclosure. In an aspect of the disclosure, the nucleic acid comprises a nucleotide sequence that has at least 95% sequence identity (i.e., at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identity) to at least one of SEQ ID NOs: 59-62 (SEQ ID NO: 59 is the nucleotide sequence of 14043; SEQ ID NO: 60 is the nucleotide sequence of 12723; SEQ ID NO: 61 is the nucleotide sequence of 10851; and SEQ ID NO: 62 is the nucleotide sequence of 13809), 113 (the nucleotide sequence of 10851 K147R), 115 (the nucleotide sequence of 10851 F797L), and 117 (the nucleotide sequence of 10851 K147R / F797L).
[0036] In a further aspect of the disclosure, the nucleic acid comprises a nucleotide sequence that has at least 80% sequence identity (i.e., at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identity) to SEQ ID NO: 61 (nucleotide sequence of 10851).Attorney Docket No. 6391-0035W00143
[0037] In a further aspect of the disclosure, the nucleic acid comprises a nucleotide sequence that has at least 95% sequence identity (i.e., at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identity) to at least one of SEQ ID NOs: 113 (the nucleotide sequence of 10851 K147R), 115 (the nucleotide sequence of 10851 F797L), and 117 (the nucleotide sequence of 10851 K147R / F797L).
[0038] “Nucleic acid,” as used herein, includes “polynucleotide,” “oligonucleotide,” and “nucleic acid molecule,” and can contain natural, non-natural or altered nucleotides, and which can contain a natural, non-natural or altered internucleotide linkage, such as a phosphoroamidate linkage or a phosphorothioate linkage, instead of the phosphodiester found between the nucleotides of an unmodified oligonucleotide. In an aspect of the disclosure, the nucleic acid does not comprise any insertions, deletions, inversions, and / or substitutions. However, it may be suitable in some instances, as discussed herein, for the nucleic acid to comprise one or more insertions, deletions, inversions, and / or substitutions.
[0039] Preferably, the nucleic acids of an aspect of the disclosure are recombinant. As used herein, the term “recombinant” refers to (i) molecules that are constructed outside living cells by joining natural or synthetic nucleic acid segments to nucleic acid molecules that can replicate in a living cell, or (ii) molecules that result from the replication of those described in (i) above. For purposes herein, the replication can be in vitro replication or in vivo replication.
[0040] The nucleic acids can be constructed based on chemical synthesis and / or enzymatic ligation reactions using procedures known in the art. See, for example, Green and Sambrook et al., supra. For example, a nucleic acid can be chemically synthesized using naturally occurring nucleotides or variously modified nucleotides designed to increase the biological stability of the molecules or to increase the physical stability of the duplex formed upon hybridization (e.g., phosphorothioate derivatives and acridine substituted nucleotides). Examples of modified nucleotides that can be used to generate the nucleic acids include, but are not limited to, 5-fluorouracil, 5 -bromouracil, 5-chlorouracil, 5-iodouracil, hypoxanthine, xanthine, 4-acetylcytosine, 5 -(carboxyhydroxymethyl) uracil, 5-carboxymethylaminomethyl- 2-thiouridine, 5-carboxymethylaminomethyluracil, dihydrouracil, beta-D-galactosylqueosine, inosine, N6-isopentenyladenine, 1-methylguanine, 1 -methylinosine, 2,2-dimethylguanine, 2- methyladenine, 2-methylguanine, 3 -methylcytosine, 5-methylcytosine, N6-substituted adenine, 7-methylguanine, 5-methylaminomethyluracil, 5-methoxyaminomethyl-2-thiouracil, beta-D- mannosylqueosine, 5'-methoxycarboxymethyluracil, 5-methoxyuracil, 2-methylthio-N6-Attorney Docket No. 6391-0035W00144 isopentenyladenine, uracil-5-oxyacetic acid (v), wybutoxosine, pseudouracil, queosine, 2- thiocytosine, 5-methyl-2-thiouracil, 2-thiouracil, 4-thiouracil, 5-methyluracil, uracil-5- oxyacetic acid methylester, 3-(3-amino-3-N-2-carboxypropyl) uracil, and 2,6-diaminopurine. Alternatively, one or more of the nucleic acids of the disclosure can be purchased from companies, such as Macromolecular Resources (Fort Collins, CO) and Synthegen (Houston, TX).
[0041] The nucleic acids of the disclosure can be incorporated into a recombinant expression vector. In this regard, the disclosure provides a recombinant expression vector comprising any of the nucleic acids of the disclosure.
[0042] For purposes herein, the term “recombinant expression vector” means a genetically- modified oligonucleotide or polynucleotide construct that permits the expression of a RNA, protein, polypeptide, or peptide by a host cell, when the construct comprises a nucleotide sequence encoding the RNA, protein, polypeptide, or peptide, and the vector is contacted with the cell under conditions sufficient to have the RNA, protein, polypeptide, or peptide expressed within the cell. The vectors of the disclosure are not naturally-occurring as a whole. However, parts of the vectors can be naturally-occurring. The inventive recombinant expression vectors can comprise any type of nucleotide, including, but not limited to DNA and RNA, which can be single-stranded or double-stranded, synthesized or obtained in part from natural sources, and which can contain natural, non-natural or altered nucleotides. The recombinant expression vectors can comprise naturally-occurring, non-naturally-occurring internucleotide linkages, or both types of linkages. Preferably, the non-naturally occurring or altered nucleotides or internucleotide linkages do not hinder the transcription or replication of the vector.
[0043] The recombinant expression vector of the disclosure can be any suitable recombinant expression vector, and can be used to transform or transfect any suitable host cell. Suitable vectors include those designed for propagation and expansion or for expression or both, such as plasmids and viruses. The vector can be selected from the group consisting of the pUC series (Fermentas Life Sciences), the pBluescript series (Stratagene, LaJolla, CA), the pET series (Novagen, Madison, WI), the pGEX series (Pharmacia Biotech, Uppsala, Sweden), and the pEX series (Clontech, Palo Alto, CA). Bacteriophage vectors, such as 1GT10, 1GT11, IZapII (Stratagene), 1EMBL4, and ANM1149, also can be used. Examples of plant expression vectors include pBIOl, pBI101.2, pBI101.3, pBH21 and pBIN19 (Clontech). Examples of animal expression vectors include pEUK-Cl, pMAM and pMAMneo (Clontech). Preferably, the recombinant expression vector is a viral vector, e.g., a retroviral vector.Attorney Docket No. 6391-0035W00145
[0044] The recombinant expression vectors of the disclosure can be prepared using standard recombinant DNA techniques described in, for example, Green and Sambrook et al., supra. Constructs of expression vectors, which are circular or linear, can be prepared to contain a replication system functional in a prokaryotic or eukaryotic host cell. Replication systems can be derived, e.g., from ColEl, 2 p plasmid, , SV40, bovine papillomavirus, and the like.
[0045] Desirably, the recombinant expression vector comprises regulatory sequences, such as transcription and translation initiation and termination codons, which are specific to the type of host cell (e.g., bacterium, fungus, plant, or animal) into which the vector is to be introduced, as appropriate and taking into consideration whether the vector is DNA- or RNA-based.
[0046] The recombinant expression vector can include one or more marker genes, which allow for selection of transformed or transfected host cells. Marker genes include biocide resistance, e.g., resistance to antibiotics, heavy metals, etc., complementation in an auxotrophic host cell to provide prototrophy, and the like. Suitable marker genes for the inventive expression vectors include, for instance, neomycin / G418 resistance genes, hygromycin resistance genes, histidinol resistance genes, tetracycline resistance genes, and ampicillin resistance genes.
[0047] The recombinant expression vector can comprise a native or nonnative promoter operably linked to the nucleotide sequence encoding the protein. The selection of promoters, e.g., strong, weak, inducible, tissue-specific and developmental-specific, is within the ordinary skill of the artisan. Similarly, the combining of a nucleotide sequence with a promoter is also within the skill of the artisan. The promoter can be a non-viral promoter or a viral promoter, e.g., a cytomegalovirus (CMV) promoter, an SV40 promoter, an RSV promoter, and a promoter found in the long-terminal repeat of the murine stem cell virus.
[0048] The inventive recombinant expression vectors can be designed for either transient expression, for stable expression, or for both. Also, the recombinant expression vectors can be made for constitutive expression or for inducible expression.
[0049] Further, the recombinant expression vectors can be made to include a suicide gene. As used herein, the term “suicide gene” refers to a gene that causes the cell expressing the suicide gene to die. The suicide gene can be a gene that confers sensitivity to an agent, e.g., a drug, upon the cell in which the gene is expressed, and causes the cell to die when the cell is contacted with or exposed to the agent. Suicide genes are known in the art and include, for example, the Herpes Simplex Virus (HSV) thymidine kinase (TK) gene, cytosine deaminase, purine nucleoside phosphorylase, nitroreductase, and the inducible caspase 9 gene system.Attorney Docket No. 6391-0035W00146
[0050] An aspect of the disclosure provides a kit comprising a gRNA and a Type V CRISPR nuclease of an aspect of the disclosure. The kit can contain instructions, containers for holding the gRNAs and Type V CRISPR nucleases, and any other components necessary or useful for performing the kit.
[0051] An aspect of the disclosure provides a nucleic acid guided nuclease system, the system comprising: (a) a Type V CRISPR nuclease nucleic acid having at least 80% sequence identity (i.e., at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identity) to least one of SEQ ID NOs: 59-62, 113, 115, and 117 that encodes a Type V CRISPR nuclease amino acid sequence having at least 95% sequence identity to at least one of SEQ ID NOs: 63-66, 114, 116, and 118, and (b) a guide nucleic acid (e.g., RNA) capable of complexing with the nucleic acid guided nuclease, wherein the guide nucleic acid comprises a constant region comprising at least 95% sequence identity to at least one of SEQ ID NOs: 90-112.
[0052] In an aspect of the disclosure, the nucleic acid guided nuclease system using the Type V CRISPR nucleases and gRNAs of an aspect of the disclosure provides improved on- target editing activity as compared to a nucleic acid guided nuclease system utilizing Type V CRISPR nuclease from Acidaminococcus sp. BV3L6, under similar conditions. In an aspect of the disclosure, the nucleic acid guided nuclease system using the Type V CRISPR nucleases and gRNAs of an aspect of the disclosure provides improved on-target editing activity as compared to a nucleic acid guided nuclease system utilizing Type V CRISPR nuclease from Lachnospiraceae bacterium ND2006, under similar conditions. In a further aspect, the nucleic acid guided nuclease system provides improved on-target editing activity as compared to a endonuclease system utilizing type II CRISPR RNA-guided endonuclease Cas9 (NCBI Gene ID: 69900935; NZ_LS483338.1).
[0053] One skilled in the art can compare the nucleic acid guided nuclease system of an aspect of the disclosure to another nucleic acid guided nuclease system and determine which system has more on-target activity (and less off-target activity) using known methods. In an aspect of the disclosure, the nucleic acid guided nuclease system of an aspect of the disclosure has at least 5 % (i.e., at least about 10%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 100%, at least about 125%, at least about 150%, at least about 175%, at leastAttorney Docket No. 6391-0035W00147 about 200%, at least about 250%, at least about 300%, at least about 400%, at least about 500%, etc.) more on-target activity as compared to a currently commercially available CRISPR system, a CRISPR / Cas9 system, or a Type V CRISPR nuclease wild type system. In an aspect of the disclosure, the nucleic acid guided nuclease system of an aspect of the disclosure has at least 5 % (i.e., at least about 10%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 100%, at least about 125%, at least about 150%, at least about 175%, at least about 200%, at least about 250%, at least about 300%, at least about 400%, at least about 500%, etc.) less off-target activity as compared to a currently commercially available CRISPR system, a CRISPR / Cas9 system, or a Type V CRISPR nuclease wild type system.
[0054] In an aspect of the disclosure, the nucleic acid guided nuclease system is administered to at least one cell. In a further aspect of the disclosure, administration of the system to the at least one cell results in an edit or modification (e.g., deletion, mutation, inversion, or substitution) to the genome of the at least one cell in a target sequence in the genome of the least one cell.
[0055] The at least one cell can be any suitable cell. In an aspect of the disclosure, the cell is a mammal cell. The mammal referred to in the inventive methods and systems can be any mammal. As used herein, the term “mammal” refers to any mammal, including, but not limited to, mammals of the order Rodentia, such as mice and hamsters, and mammals of the order Logomorpha, such as rabbits, mammals from the order Carnivora, including Felines (cats) and Canines (dogs), order Artiodactyla, including Bovines (cows) and Swines (pigs) or of the order Perssodactyla, including Equines (horses), or order Primates, Ceboids, or Simoids (monkeys) or of the order Anthropoids (humans and apes). In an aspect of the disclosure, the mammal is a human. In a further aspect, the at least one cell is in subject. In another aspect, the subject is human.
[0056] The target sequence can be any suitable sequence in the genome of the cell. In an aspect, the at least one cell is a eukaryotic cell. In an aspect of the disclosure, the target sequence is a viral sequence present in a eukaryotic cell. In another aspect of the disclosure, the target sequence is a proto-oncogene or an oncogene. In an aspect of the disclosure, the target sequence is HPRT 38115 or TRACI.Attorney Docket No. 6391-0035W00148
[0057] In an aspect of the disclosure, the Type V CRISPR nuclease amino acid sequence has at least 80% sequence identity (i.e., at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identity) to SEQ ID NO: 63. In a further aspect of the disclosure, the guide nucleic acid comprises a constant region having at least 80% (i.e., at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identity) sequence identity to at least one of SEQ ID NOs: 100-104. In an aspect of the disclosure, the Type V CRISPR nuclease amino acid sequence has at least 80% sequence identity (i.e., at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identity) to SEQ ID NO: 63 and the guide nucleic acid comprises a constant region having at least 80% sequence identity (i.e., at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identity) to at least one of SEQ ID NOs: 100-104.
[0058] In an aspect on the disclosure, the Type V CRISPR nuclease amino acid sequence has at least 80% sequence identity (i.e., at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identity) to SEQ ID NO: 64. In a further aspect of the disclosure, the guide nucleic acid comprises a constant region having at least 80% sequence identity (i.e., at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identity) to at least one of SEQ ID NOs: 94-99. In an aspect of the disclosure, the Type V CRISPR nuclease amino acid sequence has at least 80% sequence identity to SEQ ID NO: 64 and the guide nucleic acid comprises a constant region having at least 80% sequence identity to at least one of SEQ ID NOs: 94-99.
[0059] In an aspect of the disclosure, the Type V CRISPR nuclease amino acid sequence has at least 80% sequence identity (i.e., at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identity) to SEQ ID NO: 65. In a further aspect of the disclosure, the guide nucleic acid comprises a constant region having at least 95% sequence identity (i.e., at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identity) to at least one of SEQ ID NOs: 90-94. In an aspect of theAttorney Docket No. 6391-0035W00149 disclosure, the Type V CRISPR nuclease amino acid sequence has at least 80% sequence identity (i.e., at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identity) to SEQ ID NO: 65 and the guide nucleic acid comprises a constant region having at least 80% sequence identity (i.e., at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identity) to at least one of SEQ ID NOs: 90-94.
[0060] In an aspect of the disclosure, the Type V CRISPR nuclease amino acid sequence has at least 80% sequence identity (i.e., at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identity) to SEQ ID NO: 66. In a further aspect of the disclosure, the guide nucleic acid comprises a constant region having at least 80% sequence identity (i.e., at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identity) to at least one of SEQ ID NOs: 105-112. In an aspect of the disclosure, the Type V CRISPR nuclease amino acid sequence has at least 80% sequence identity (i.e., at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identity) to SEQ ID NO: 66 and the guide nucleic acid comprises a constant region having at least 80% sequence identity (i.e., at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identity) to at least one of SEQ ID NOs: 105-112.
[0061] In an aspect of the disclosure, the Type V CRISPR nuclease amino acid sequence has at least 80% sequence identity (i.e., at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identity) to SEQ ID NO: 114. In a further aspect of the disclosure, the guide nucleic acid comprises a constant region having at least 80% sequence identity (i.e., at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identity) to SEQ ID NO: 92. In an aspect of the disclosure, the Type V CRISPR nuclease amino acid sequence has at least 80% sequence identity (i.e., at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identity) to SEQ ID NO: 114 and the guide nucleic acidAttorney Docket No. 6391-0035W00150 comprises a constant region having at least 80% sequence identity (i.e., at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identity) to SEQ ID NO: 92.
[0062] In an aspect of the disclosure, the Type V CRISPR nuclease amino acid sequence has at least 80% sequence identity (i.e., at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identity) to SEQ ID NO: 116. In a further aspect of the disclosure, the guide nucleic acid comprises a constant region having at least 80% sequence identity (i.e., at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identity) to SEQ ID NO: 92. In an aspect of the disclosure, the Type V CRISPR nuclease amino acid sequence has at least 80% sequence identity (i.e., at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identity) to SEQ ID NO: 116 and the guide nucleic acid comprises a constant region having at least 80% sequence identity (i.e., at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identity) to SEQ ID NO: 92.
[0063] In an aspect of the disclosure, the Type V CRISPR nuclease amino acid sequence has at least 80% sequence identity (i.e., at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identity) to SEQ ID NO: 118. In a further aspect of the disclosure, the guide nucleic acid comprises a constant region having at least 80% sequence identity (i.e., at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identity) to SEQ ID NO: 92. In an aspect of the disclosure, the Type V CRISPR nuclease amino acid sequence has at least 80% sequence identity (i.e., at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identity) to SEQ ID NO: 118 and the guide nucleic acid comprises a constant region having at least 80% sequence identity (i.e., at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identity) to SEQ ID NO: 92.Attorney Docket No. 6391-0035W00151
[0064] The Type V CRISPR nucleases of an aspect of the disclosure can be any suitable length, i.e., can comprise any number of amino acids, provided that the nucleases retain their biological activity.
[0065] The guide nucleic acids or gRNAs of an aspect of the disclosure can be any suitable length, i.e., can comprise any number of nucleotides, provided that the guide nucleic acids or gRNAs retain their biological activity. In an aspect of the disclosure, the guide nucleic acids or gRNAs are from about 35 to about 50 bases (e.g., from about 36 to about 49 bases, from about 37 to about 48 bases, from about 38 to about 47 bases, from about 39 to about 46 bases, from about 40 to about 44 bases, or about 41 bases.
[0066] In an aspect of the disclosure, the guide nucleic acids or gRNAs comprise a constant region that is from about 15 to about 25 bases, from about 16 to about 24 bases, from about 17 to about 23 bases, from about 18 to about 22 bases, from about 19 to about 21 bases, or about 20 bases.
[0067] In an aspect of the disclosure, the guide nucleic acids or gRNAs comprise a target specific region that is from about 17 to about 25 bases, from about 18 to about 24 bases, from about 19 to about 20 bases, from about 19 to about 21 bases, from about 20 to about 22 bases, or about 21 bases.
[0068] The nucleases described herein may include functional variants thereof. The functional variant can, for example, comprise the sequence of the parent nuclease / protein with at least one conservative amino acid substitution. Conservative amino acid substitutions are known in the art, and include amino acid substitutions in which one amino acid having certain physical and / or chemical properties is exchanged for another amino acid that has the same chemical or physical properties. For instance, the conservative amino acid substitution can be an acidic amino acid substituted for another acidic amino acid (e.g., Asp or Glu), an amino acid with a nonpolar side chain substituted for another amino acid with a nonpolar side chain (e.g., Ala, Gly, Vai, He, Leu, Met, Phe, Pro, Trp, Vai, etc.), a basic amino acid substituted for another basic amino acid (Lys, Arg, etc.), an amino acid with a polar side chain substituted for another amino acid with a polar side chain (Asn, Cys, Gin, Ser, Thr, Tyr, etc.), etc.
[0069] Alternatively or additionally, the functional variants can comprise the amino acid sequence of the parent protein with at least one non-conservative amino acid substitution. In this case, it is preferable for the non-conservative amino acid substitution to not interfere with or inhibit the biological activity of the functional variant. Preferably, the non-conservativeAttorney Docket No. 6391-0035W00152 amino acid substitution enhances the biological activity of the functional variant, such that the biological activity of the functional variant is increased as compared to the parent protein.
[0070] The Type V CRISPR nucleases and guide nucleic acids described herein can comprise, consist essentially of, or consist of the specified sequences described herein, such that other components of the sequences, e.g., other nucleic acids, do not materially change the biological activity of the Type V CRISPR nucleases or guide nucleic acids.
[0071] The Type V CRISPR nucleases and guide nucleic acid (e.g., gRNAs) of an aspect of the disclosure can be produced via any suitable means. For example, the Type V CRISPR nucleases and guide nucleic acid can be synthetic or recombinant.
[0072] The Type V CRISPR nucleases and guide nucleic acids (e.g., gRNAs) of an aspect of the disclosure can be isolated and / or purified. The Type V CRISPR nucleases and guide nucleic acids of an aspect of the disclosure may be purified using any suitable means. The term “isolated” as used herein means having been removed from its natural environment. The term “purified” as used herein means having been increased in purity, wherein “purity” is a relative term, and not to be necessarily construed as absolute purity. For example, the purity can be at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, or can be about 100%.
[0073] Also provided herein are compositions comprising at least one Type V CRISPR nuclease and / or at least one guide nucleic acid (e.g., gRNA) of an aspect of the disclosure.
[0074] Also provided herein are compositions comprising at least one Type V CRISPR nuclease and / or at least one constant region of a guide nucleic acid (e.g., gRNA) of an aspect of the disclosure.
[0075] The compositions may comprise any suitable excipient. In an aspect of the disclosure, the compositions comprise a buffer. The excipient can be any of those conventionally used for the particular inventive Type V CRISPR nucleases or guide nucleic acids (or gRNAs, or constant regions thereof) under consideration.
[0076] The compositions can comprise several Type V CRISPR nucleases. Further, the compositions can comprise several guide nucleic acids (or gRNAs, or constant regions thereof). Further, the compositions can comprise several constant regions of the guide nucleic acids (e.g., gRNAs).
[0077] The Type V CRISPR nucleases, guide nucleic acids (gRNAs or constant regions thereof), or compositions thereof, can be administered to the cell or subject using any suitable means. The amount or dose of the Type V CRISPR nucleases and guide nucleic acids shouldAttorney Docket No. 6391-0035W00153 be sufficient to effect, e.g., a therapeutic or prophylactic response, in the cell or subject over a reasonable time frame.
[0078] The following examples further illustrate the disclosure but, of course, should not be construed as in any way limiting its scope.EXAMPLE 1
[0079] This example demonstrates the production of Type V CRISPR nucleases of an aspect of the disclosure.
[0080] The CRISPRCasFinder program (Couvin, et al., Nucleic Acids Res., 46 (Wl): W246-W251 (2018)) was used with default settings to prospect for Cas-sy stems within a selected metagenomic data set. This mining resulted in the identification of a large number of protein sequences. A subset of protein sequences was selected and then the protein sequences were modified by adding nucleases (e.g., AsCasl2a). The proteins were further modified by human codon optimization using standard techniques and E. Coli. The modified proteins were used for further testing.
[0081] Seventy-six modified Type V CRISPR nuclease protein sequences were selected to be synthesized and cloned into a typical human protein expression plasmid. The following protocol for 10851 is shown as an example.
[0082] Plasmid pTWIST-10851 was used with various site directed mutagenesis (SDMs) to introduce either the Lb or As Ultra Mutations: 10815 No Change; 10851 K147R (DNA sequence SEQ ID NO: 113; protein sequence SEQ ID NO: 114); 10851 F797L (DNA sequence SEQ ID NO: 115; protein sequence SEQ ID NO: 116); and 10851 K147R / F797L (DNA sequence SEQ ID NO: 117; protein sequence SEQ ID NO: 118).
[0083] The initial tests with the potential Type V CRISPR nuclease variants used guides that target two different sites within the HPRTlocus (HPRT-38330 and HPRT-38115) utilizing constant regions specific for the Acidaminococcus and Lachnospiraceae species Type V CRISPR nuclease variants (SEQ ID NOs: 1-8). The Type V CRISPR nuclease variant plasmid and either AsCasl2a or LbCasl2a gRNAs targeting HPRT sites 38115 or 38330 were co- lipofected into HEK293 cells. The cells were incubated at 37 °C with 5% CO2 for 72 hours. The DNA was then extracted for evaluation using Next Generation Sequencing (NGS; Figure 1).Attorney Docket No. 6391-0035W00154EXAMPLE 2
[0084] This example demonstrates the development of guide RNA constant regions for increasing the editing efficiency of Type V CRISPR nucleases of an aspect of the disclosure.
[0085] Guide RNA constant regions were designed for each species variant using predicted RNA guides located on the same metagenome sequence contig as the CRISPR nuclease for 3 out of 4 variants where gRNA sequences were predicted to be present 12723, 13809, and 14043). For the remaining nuclease 10851), a predicted gRNA was not present on the same contig, therefore the gRNA constant regions were designed based off another nuclease with approximately 45% sequence identity and a predicted gRNA sequence on the same contig. Multiple gRNAs were designed for each potential Type V CRISPR nuclease variant, targeting the HPRT 38115 site and clinically relevant TRACI site (SEQ ID NO: 9-54). The 5’ end of each constant region was unpredictable, therefore many gRNAs with multiple lengths for each variant had to be tested (SEQ ID NOs: 90-112). The Type V CRISPR nuclease variant plasmid and a gRNA targeting either HPRT 38115 or TRACI were co-lipofected into HEK293 cells. The cells were incubated at 37 °C with 5% CO2 for 72 hours. The DNA was then extracted for evaluation using a standard T7 endonuclease I (T7EI) assay (T7 HPRT Primers: HPRT T7 Rev: ACACATCCATGGGACTTCTGCCTC (SEQ ID NO: 56), and HPRT T7For: AAGAATGTTGTGATAAAAGGTGATGCT (SEQ ID NO: 55); T7 TRACI Primers: TRAC1 T7 F3: CAATGGTCCTGTCTCTCAAG (SEQ ID NO: 57), and TRAC1 T7 R: GTGATGGAACAGGATGCAGT (SEQ ID NO: 58)). Cleavage efficiency was quantified using capillary electrophoresis (Figures 2 and 3).
[0086] In the T7EI assay, target genomic regions from CRISPR-modified cells are amplified by PCR. The PCR products are denatured and reannealed to allow heteroduplex formation between wild-type DNA and CRISPR-mutated DNA. Mutations are then detected using T7EI, which recognizes and cleaves mismatched DNA heterodulplexes. T7EI assay results are analyzed by visualizing cleavage products and full-length amplicons by gel or capillary electrophoresis. The relative editing efficiency of the Type V CRISPR nuclease variants with each designed constant region was compared to standard AsCasl2a or LbCasl2a gRNA. Type V CRISPR nuclease variants targeting HPRT or TRACI locus (SEQ ID NOs: 9- 54) with the designed constant regions (SEQ ID NOs: 90-112) showed improved cleavage efficiency as measured by T7EI cleavage assay when compared to standard AsCasl2a or LbCasl2a gRNA (SEQ ID NOs: 1-4).Attorney Docket No. 6391-0035W00155Example 3
[0087] This example demonstrates the development and use of the developed gRNA constant regions for increasing the editing efficiency of Type V CRISPR nuclease variants of an aspect of the disclosure.
[0088] pET vectors were generated incorporating mutations into the 10851 Type V CRISPR nuclease. These mutations were introduced into the 10851 Type V CRISPR nuclease by site directed mutagenesis. Mutations were introduced into DNA sequences at: 1) a K147R (SEQ ID NO: 113); 2) F797L (SEQ ID NO: 115); or the double mutant K147R / F797L (SEQ ID NO: 117). The mutant 10851 Type V CRISPR nucleases were then expressed and purified (SEQ ID NO: 114, SEQ ID NO: 116, or SEQ ID NO: 118).
[0089] The purified mutant 10851 Type V CRISPR nucleases were tested as ribonucleoprotein (RNP) in HEK293 cells with 10851 gRNA directed to HPRT locus (SEQ ID NO: 30) or with 10851 gRNA directed to TRAC 1 locus (SEQ ID NO: 53).
[0090] The Type V CRISPR nuclease variant plasmid and gRNAs targeting either HPRT 38115 or TRACI were co-lipofected into HEK293 cells. The cells were incubated at 37 °C with 5% CO2 for 72 hours. The DNA was then extracted for evaluation using a standard T7 endonuclease I (T7EI) assay (T7 HPRT Primers: HPRT T7 Rev: ACACATCCATGGGACTTCTGCCTC (SEQ ID NO: 56), and HPRT T7For: AAGAATGTTGTGATAAAAGGTGATGCT (SEQ ID NO: 55); T7 TRACI Primers: TRAC1 T7 F3: CAATGGTCCTGTCTCTCAAG (SEQ ID NO: 57), and TRAC1 T7 R: GTGATGGAACAGGATGCAGT (SEQ ID NO: 58)). Cleavage efficiency was quantified using capillary electrophoresis.
[0091] Figures 4A-4D show that Type V CRISPR nucleases and guide nucleic acids of an aspect of the disclosure are effective in editing human cells at the HPRT 38115 site. Figures 5A-5D show that Type V CRISPR nucleases and guide nucleic acids of an aspect of an disclosure are also effective in editing human cells at the TRACI site.Attorney Docket No. 6391-0035W00156
[0092] All references, including publications, patent applications, and patents, cited herein are hereby incorporated by reference to the same extent as if each reference were individually and specifically indicated to be incorporated by reference and were set forth in its entirety herein.
[0093] The use of the terms “a” and “an” and “the” and “at least one” and similar referents in the context of describing the disclosure (especially in the context of the following claims) are to be construed to cover both the singular and the plural, unless otherwise indicated herein or clearly contradicted by context. The use of the term “at least one” followed by a list of one or more items (for example, “at least one of A and B”) is to be construed to mean one item selected from the listed items (A or B) or any combination of two or more of the listed items (A and B), unless otherwise indicated herein or clearly contradicted by context. The terms “comprising,” “having,” “including,” and “containing” are to be construed as open-ended terms (i.e., meaning “including, but not limited to,”) unless otherwise noted. Recitation of ranges of values herein are merely intended to serve as a shorthand method of referring individually to each separate value falling within the range, unless otherwise indicated herein, and each separate value is incorporated into the specification as if it were individually recited herein. All methods described herein can be performed in any suitable order unless otherwise indicated herein or otherwise clearly contradicted by context. The use of any and all examples, or exemplary language (e.g., “such as”) provided herein, is intended merely to better illuminate the disclosure and does not pose a limitation on the scope of the disclosure unless otherwise claimed. No language in the specification should be construed as indicating any non-claimed element as essential to the practice of the disclosure.
[0094] Preferred aspects of this disclosure are described herein, including the best mode known to the inventors for carrying out the disclosure. Variations of those preferred aspects may become apparent to those of ordinary skill in the art upon reading the foregoing description. The inventors expect skilled artisans to employ such variations as appropriate, and the inventors intend for the disclosure to be practiced otherwise than as specifically described herein. Accordingly, this disclosure includes all modifications and equivalents of the subject matter recited in the claims appended hereto as permitted by applicable law. Moreover, any combination of the above-described elements in all possible variations thereof is encompassed by the disclosure unless otherwise indicated herein or otherwise clearly contradicted by context.Attorney Docket No. 6391-0035W00157
[0095] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art. For example, any nomenclatures used in connection with, and techniques of biochemistry, molecular biology, immunology, microbiology, genetics, cell and tissue culture, and protein and nucleic acid chemistry described herein are well known and commonly used in the art. In case of conflict, the present disclosure, including definitions, will control. Exemplary methods and materials are described below, although methods and materials similar or equivalent to those described herein can be used in practice or testing of the embodiments and aspects described herein.
[0096] As used herein, the terms “amino acid,” “nucleotide,” “polynucleotide,” “vector,” “polypeptide,” and “protein” have their common meanings as would be understood by a biochemist of ordinary skill in the art. Standard single letter nucleotides (A, C, G, T, U) and standard single letter amino acids (A, C, D, E, F, G, H, I, K, L, M, N, P, Q, R, S, T, V, W, or Y) are used herein.
[0097] As used herein, the terms such as “include,” “including,” “contain,” “containing,” “having,” and the like mean “comprising.” The present disclosure also contemplates other embodiments “comprising,” “consisting of,” and “consisting essentially of,” the embodiments or elements presented herein, whether explicitly set forth or not.
[0098] As used herein, the term “or” can be conjunctive or disjunctive.
[0099] As used herein, the term “substantially” means to a great or significant extent, but not completely.
[0100] As used herein, the term “about” or “approximately” as applied to one or more values of interest, refers to a value that is similar to a stated reference value, or within an acceptable error range for the particular value as determined by one of ordinary skill in the art, which will depend in part on how the value is measured or determined, such as the limitations of the measurement system. In one aspect, the term “about” refers to any values, including both integers and fractional components that are within a variation of up to ± 10% of the value modified by the term “about.” Alternatively, “about” can mean within 3 or more standard deviations, per the practice in the art. Alternatively, such as with respect to biological systems or processes, the term “about” can mean within an order of magnitude, in some embodiments within 5-fold, and in some embodiments within 2-fold, of a value. As used herein, the symbol “-’’means “about” or “approximately.”
[0101] All ranges disclosed herein include both end points as discrete values as well as all integers and fractions specified within the range. For example, a range of 0.1-2.0 includes 0.1,Attorney Docket No. 6391-0035W001580.2, 0.3, 0.4 . . . 2.0. If the end points are modified by the term “about,” the range specified is expanded by a variation of up to ± 10% of any value within the range or within 3 or more standard deviations, including the end points.
[0102] As used herein, the term “hybridization” refers to the process of combining two complementary single-stranded nucleic acid molecules and allowing them to form a single double-stranded hybrid molecule through base pairing.
[0103] As used herein, the term “target enrichment” with respect to a nucleic acid is intended to refer to increasing the relative concentration of particular nucleic acid species in the sample.
[0104] As used herein, the term “nucleic acid” may refer to DNA, RNA, dsDNA, dsRNA, ssDNA, ssRNA, or hybrids of DNA / RNA complexes or sequences obtained from any source, containing target and non-target sequences. For example, a nucleic acid sample can be obtained from artificial sources or by chemical synthesis, or from viruses, prokaryotic cells including microbes, or eukaryotic cells. Biological samples may be vertebrate, including human or excluding humans, invertebrates, plants, microbes, viruses, mycoplasma, fungi, or Archaea. A nucleic acid sample may comprise whole genomic sequences, portions of the genomic sequence, chromosomal sequences, mitochondrial sequences, PCR products, whole genome amplification products or products of other amplification protocol, such as but not limited to, cDNA sequences, mRNA sequences, whole transcriptome sequences, exons, or intronic. These examples are not to be construed as limiting the sample types applicable to aspects described herein.
[0105] By a polynucleotide having a nucleotide sequence at least, for example, 90-99% “identical” to a reference nucleotide sequence (e.g., SEQ ID NO: 65) is intended that the nucleotide sequence of the polynucleotide be identical to the reference sequence except that the polynucleotide sequence can include up to about 10 to 1 point mutations, additions, or deletions per each 100 nucleotides of the reference nucleotide sequence without affecting the function or activity of the polynucleotide sequence.
[0106] In other words, to obtain a polynucleotide having a nucleotide sequence about at least 90-99% identical to a reference nucleotide sequence, up to 10% of the nucleotides in the reference sequence can be deleted, added, or substituted, with another nucleotide, or a number of nucleotides up to 10% of the total nucleotides in the reference sequence can be inserted into the reference sequence. These mutations of the reference sequence can occur at the 5'- or 3'- terminal positions of the reference nucleotide sequence or anywhere between those terminalAttorney Docket No. 6391-0035W00159 positions, interspersed either individually among nucleotides in the reference sequence or in one or more contiguous groups within the reference sequence.
[0107] As noted above, two or more polynucleotide sequences can be compared by determining their percent identity. The percent identity of two sequences is generally described as the number of exact matches between two aligned sequences divided by the length of the shorter sequence and multiplied by 100. An approximate alignment for nucleic acid sequences is provided by the local homology algorithm of Smith and Waterman, Advances in Applied Mathematics, 2(4): 82-489 (1981).
[0100] The polynucleotides described herein include those comprising mutations, variations, substitutions, additions, deletions, and particular examples of the polynucleotides described herein. In addition, the polynucleotides can be ribonucleotides (RNA), deoxyribonucleotides (DNA), or combinations thereof, where the sequence comprise one or more ribonucleotides, or deoxyribonucleotides. The sequences can comprise ribonucleotides with for example ribothymidine nucleotides instead of uridine nucleotides to maintain consonance with the analogous DNA sequence.
[0101] It will be apparent to one of ordinary skill in the relevant art that suitable modifications and adaptations to the compositions, formulations, methods, processes, and applications described herein can be made without departing from the scope of any embodiments or aspects thereof. The compositions and methods provided are exemplary and are not intended to limit the scope of any of the specified embodiments. All of the various embodiments, aspects, and options disclosed herein can be combined in any variations or iterations. The scope of the compositions, formulations, methods, and processes described herein include all actual or potential combinations of embodiments, aspects, options, examples, and preferences herein described. The exemplary compositions and formulations described herein may omit any component, substitute any component disclosed herein, or include any component disclosed elsewhere herein. The ratios of the mass of any component of any of the compositions or formulations disclosed herein to the mass of any other component in the formulation or to the total mass of the other components in the formulation are hereby disclosed as if they were expressly disclosed. Should the meaning of any terms in any of the patents or publications incorporated by reference conflict with the meaning of the terms used in this disclosure, the meanings of the terms or phrases in this disclosure are controlling. Furthermore, the foregoing discussion discloses and describes merely exemplary embodiments.
Claims
Attorney Docket No. 6391-0035W00160CLAIM(S):
1. A Type V CRISPR nuclease, wherein the Type V CRISPR nuclease comprises an amino acid sequence that has at least 80% sequence identity to at least one of SEQ ID NOs: 63-66, 114, 116, and 118.
2. The Type V CRISPR nuclease of claim 1, wherein the Type V CRISPR nuclease comprises an amino acid sequence that has at least 80% sequence identity to SEQ ID NO: 65.
3. The Type V CRISPR nuclease of claim 1, wherein the Type V CRISPR nuclease consists of an amino acid sequence that has 100% sequence identity to one of SEQ ID NOs: 63-66, 114, 116, and 118.
4. The Type V CRISPR nuclease of any one of claims claim 1-3, wherein the amino acid sequence that has at least 80% sequence identity to at least one of SEQ ID NOs: 63-66, 114, 116, and 118 is human codon pair optimized.
5. A Type V CRISPR nuclease ribonucleoprotein complex, wherein the Cast 2a ribonucleoprotein complex comprises: (a) a guide RNA; and (b) a Type V CRISPR nuclease of any one of claims 1-4.
6. The Type V CRISPR nuclease ribonucleoprotein complex of claim 5, wherein the Type V CRISPR endonuclease ribonucleoprotein complex is active in a Type V CRISPR endonuclease system.
7. The Type V CRISPR nuclease ribonucleoprotein complex of claim 5 or 6, wherein the Type V CRISPR nuclease system provides improved on-target editing activity as compared to a Type V CRISPR nuclease system utilizing Type V CRISPR nuclease from Ackiammococcus sp. B V3L6.
8. The Type V CRISPR nuclease ribonucleoprotein complex of any one of claims 5-7, wherein the guide RNA comprises a nucleotide sequence that has at least 80% sequence identity to at least one of SEQ ID NOs: 9-54 and 67-89.Attorney Docket No. 6391-0035W001619. The Type V CRISPR nuclease ribonucleoprotein complex of any one of claims 5-7, wherein the guide RNA comprises a nucleotide sequence that has at least 80% sequence identity to at least one of SEQ ID NOs: 1-4.
10. A method for manipulating DNA in a cell, the method comprising contacting the DNA in the cell with at least one Casl2a ribonucleoprotein complex of any one of claims 5-9.
11. An isolated or purified nucleic acid comprising a nucleotide sequence encoding the protein according to any one of claims 1-4.
12. The isolated or purified nucleic acid of claim 11, wherein the nucleic acid comprises a nucleotide sequence that has at least 80% sequence identity to at least one of SEQ ID NOs: 59-62, 113, 115, and 117.
13. The isolated or purified nucleic acid of claim 11, wherein the nucleic acid comprises a nucleotide sequence that has at least 80% sequence identity to SEQ ID NO: 61.
14. A recombinant expression vector comprising the nucleic acid according to any one of claims 11-13.
15. A kit comprising: (a) at least one guide RNA; and (b) a Type V CRISPR nuclease of any one of claims 1-4.
16. A nucleic acid guided nuclease system, the system comprising: a. a Type V CRISPR nuclease nucleic acid having at least 80% identity to least one of SEQ ID NOs: 59-62, 113, 115, and 117 that encodes a Type V CRISPR nuclease amino acid sequence having at least 80% sequence identity to at least one of SEQ ID NOs: 63-66, 114, 116, and 118; and b. a guide nucleic acid capable of complexing with the nucleic acid guided nuclease, wherein the guide nucleic acid comprises a constant region comprising at least 95% sequence identity to at least one of SEQ ID NOs: 67-89.Attorney Docket No. 6391-0035W0016217. The nucleic acid guided nuclease system of claim 16, wherein the system is administered to at least one cell.
18. The nucleic acid guided nuclease system of claim 17, wherein administration of the system to the at least one cell results in an edit to the genome of the at least one cell in a targeted region of the genome of the least one cell.
19. The nucleic acid guided nuclease system of any one of claims 16-18, wherein the Type V CRISPR nuclease amino acid sequence has at least 80% sequence identity to SEQ ID NO: 63.
20. The nucleic acid guided nuclease system of claim 19, wherein the guide nucleic acid comprises a constant region having at least 95% sequence identity to at least one of SEQ ID NOs: 77-81.
21. The nucleic acid guided nuclease system of any one of claims 16-18, wherein the Type V CRISPR nuclease amino acid sequence has at least 80% sequence identity to SEQ ID NO: 64.
22. The nucleic acid guided nuclease system of claim 21, wherein the guide nucleic acid comprises a constant region having at least 80% sequence identity to at least one of SEQ ID NOs: 71-76.
23. The nucleic acid guided nuclease system of any one of claims 16-18, wherein the Type V CRISPR nuclease amino acid sequence has at least 80% sequence identity to SEQ ID NO: 65.
24. The nucleic acid guided nuclease system of claim 23, wherein the guide nucleic acid comprises a constant region having at least 80% sequence identity to at least one of SEQ ID NOs: 67-70.
25. The nucleic acid guided nuclease system of any one of claims 16-18, wherein the Type V CRISPR nuclease amino acid sequence has at least 80% sequence identity to SEQ ID NO: 66.Attorney Docket No. 6391-0035W0016326. The nucleic acid guided nuclease system of claim 25, wherein the guide nucleic acid comprises a constant region having at least 80% sequence identity to at least one of SEQ ID NOs: 82-89.
27. The nucleic acid guided nuclease system of any one of claims 16-18, wherein the Type V CRISPR nuclease amino acid sequence has at least 80% sequence identity to at least one of SEQ ID NOs: 114, 116, and 118.
28. The nucleic acid guided nuclease system of claim 27, wherein the guide nucleic acid comprises a constant region having at least 80% sequence identity to SEQ ID NO: 92.