Engineered CasX systems
By deeply mutating and evolving the CasX protein and gNA, their binding and cleavage abilities with target DNA were optimized, solving the problems of insufficient efficiency and specificity of existing CRISPR-Cas systems in gene editing, and achieving a significant improvement in gene editing efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-06-05
- Publication Date
- 2026-03-20
AI Technical Summary
Existing CRISPR-Cas systems require improvement in therapeutic, diagnostic, and research applications, particularly due to insufficient optimization of the Cas protein and guide RNA, resulting in a lack of efficiency and specificity in gene editing.
Variants of the CasX nuclease protein and guide RNA are provided. Through deep mutation evolution (DME) optimization of the CasX protein and gNA, specific modifications are introduced to improve their binding and cleavage ability with target DNA, including modifications to the non-target strand binding domain, target strand loading domain, helical I domain, helical II domain, oligonucleotide binding domain, and RuvC DNA cleavage domain, and to bind specific scaffold stem-loop structures.
It has achieved a significant improvement in gene editing efficiency, with the editing capability of some variants being more than twice that of the original system, and significantly improving the specificity and efficiency of the target DNA.
Smart Images

Figure CN114375334B_ABST
Abstract
Description
[0001] Cross Reference to Related Applications
[0002] This application claims priority to U.S. Provisional Patent Application Nos. 62 / 858,750, filed June 7, 2019, 62 / 944,892, filed December 6, 2019, and 63 / 030,838, filed May 27, 2020, the contents of each of which are incorporated herein by reference in their entirety.
[0003] INCORPORATION BY REFERENCE OF THE SEQUENCE LISTING
[0004] The instant application contains a Sequence Listing which has been submitted via EFS-WEB and is hereby incorporated by reference in its entirety. The ASCII copy, created on June 5, 2020, is named SCRB_011_03WO_SeqList_25 and is 3.63 MB in size. BACKGROUND
[0005] CRISPR-Cas systems confer adaptive immunity to bacteria and archaea against phages and viruses. Intensive research over the past decade has not covered the biochemistry of these systems. CRISPR-Cas systems consist of Cas proteins involved in acquisition, targeting, and cleavage of foreign DNA or RNA, and CRISPR arrays comprising direct repeat sequences flanking short spacer sequences that guide Cas proteins to their targets. Class 2 CRISPR-Cas are streamlined versions in which a single Cas protein bound to RNA is responsible for binding to and cleaving targeted sequences. The programmable nature of these minimal systems has facilitated their use as a versatile technology that is revolutionizing the field of genome manipulation.
[0006] To date, only some of the class 2 CRISPR / Cas systems discovered have been widely used. Thus, there is a need in the art for additional class 2 CRISPR / Cas systems (e.g., Cas protein plus guide RNA combinations) that have been optimized and / or provide improvements over earlier generation systems used in a variety of therapeutic, diagnostic, and research applications. SUMMARY
[0007] In some aspects, the present disclosure provides a variant of a reference CasX nuclease protein, wherein the CasX variant is capable of forming a complex with a guide nucleic acid (NA), and wherein the complex can bind a target DNA, wherein the target DNA comprises a non-target strand and a target strand, and wherein the CasX variant comprises at least one modification relative to a domain of a reference CasX, and exhibits one or more improved characteristics compared to the reference CasX protein. Domains of a reference CasX protein include: (a) a non-target strand binding (NTSB) domain that binds to the non-target strand of the DNA, wherein the NTSB domain comprises a four-stranded beta sheet; (b) a target strand loading (TSL) domain that positions the target DNA in the cleavage site of the CasX variant, the TSL domain comprising three positively charged amino acids, wherein the three positively charged amino acids bind to the target strand of the DNA, (c) a helix I domain that interacts with the spacer region of the target DNA and the guide NA, wherein the helix I domain comprises one or more alpha helices; (d) a helix II domain that interacts with the scaffold stem of the target DNA and the guide NA; (e) an oligonucleotide binding domain (OBD) that binds the triple helix region of the guide NA; and (f) a RuvC DNA cleavage domain.
[0008] In some aspects, the present disclosure provides a variant of a reference guide nucleic acid (gNA) that is capable of binding a CasX protein, wherein the reference guide nucleic acid comprises at least one modification in a region compared to a reference guide nucleic acid sequence, and the variant exhibits one or more improved characteristics compared to the reference guide RNA. Scaffold regions of a gNA include: (a) an extended stem loop; (b) a scaffold stem loop; (c) a triple helix; and (d) a pseudoknot. In some cases, the scaffold stem of the variant gNA further comprises a bubble. In other cases, the scaffold of the variant gNA further comprises a triple helix loop region. In other cases, the scaffold of the variant gNA further comprises a 5’ unstructured region.
[0009] In some aspects, the present disclosure provides a gene editing pair comprising a CasX protein and a gNA of any of the embodiments described herein.
[0010] In some aspects, the present disclosure provides polynucleotides and vectors encoding the CasX proteins, gNAs, and gene editing pairs described herein. In some embodiments, the vector is a viral vector, such as an adeno-associated virus (AAV) vector or a lentivirus vector. In other embodiments, the vector is a non-viral particle, such as a virus-like particle or a nanoparticle.
[0011] In some aspects, the present disclosure provides cells comprising the polynucleotides, vectors, CasX proteins, gNAs, and gene editing pairs described herein. In other aspects, the present disclosure provides cells comprising a target DNA edited by the methods of the editing embodiments described herein.
[0012] In some aspects, the present invention provides kits comprising the polynucleotides, vectors, CSX proteins, gNA, and gene editing pairs described herein.
[0013] In some aspects, the present invention provides a method for editing target DNA, comprising contacting the target DNA with one or more of the gene editing pairs described herein, wherein said contact causes editing of the target DNA.
[0014] In other respects, the present invention provides a method for treating an individual in need, comprising administering a gene-editing pair or a vector comprising or encoding a gene-editing pair according to any of the embodiments described herein.
[0015] In another respect, this article provides gene editing pairs for use as pharmaceutical agents, compositions containing gene editing pairs, or vectors containing or encoding gene editing pairs.
[0016] In another aspect, this document provides gene editing pairs for therapeutic methods, compositions comprising gene editing pairs, or vectors comprising or encoding gene editing pairs, wherein the method comprises editing or modifying target DNA; optionally wherein the editing occurs in an individual having a mutation in the alleles of a gene, wherein the mutation causes a disease or condition in the individual, preferably wherein the editing alters a mutation in the wild-type allele of the gene, or knocks out or eliminates an allele of a gene causing a disease or condition in the individual. Attached Figure Description
[0017] The novel features of the invention are set forth in detail in the appended claims. A better understanding of the features and advantages of the invention will be obtained by referring to the following detailed description of illustrative embodiments utilizing the principles of the invention and their accompanying drawings:
[0018] FIG. 1 A diagram illustrating an exemplary method for creating CasX protein and guide RNA variants of the present invention using deep mutational evolution (DME). In some exemplary embodiments, DME establishes and tests virtually every possible mutation, insertion, and deletion, and their combinations / multiple recombinations, in a biomolecule, and provides a near-comprehensive and unbiased assessment of the biomolecule's fitness landscape and a path of sequence space toward the desired outcome. As described herein, DME can be applied to CasX proteins and guide RNAs.
[0019] FIG. 2FIGURE AND EXAMPLE FLUORESCENCE ACTIVATED CELL SORTING (FACS) PLOTS TO ILLUSTRATE AN EXAMPLE METHOD OF ANALYZING THE EFFECTIVENESS OF A CASX PROTEIN OR SINGLE GUIDE RNA (sgRNA) OR VARIANTS THEREOF. A reporter (e.g., GFP reporter) that is coupled to a gRNA target sequence, complementary to a gRNA spacer, is integrated into a reporter cell line. Cells are transformed or transfected with a CasX protein and / or sgNA variant, where the spacer motif of the sgRNA is complementary to and targets the gRNA target sequence of the reporter. The ability of the CasX:sgRNA ribonucleoprotein complex to cleave the target sequence is analyzed by FACS. Cells that lose reporter expression indicate CasX:sgRNA ribonucleoprotein complex-mediated cleavage and indel formation has occurred.
[0020] FIG. 3A and FIG. 3B HEATMAP TO SHOW THE RESULTS OF EXAMPLE DME MUTAGENESIS OF A REFERENCE sgRNA ENCODED BY SEQ ID NO: 5 AS DESCRIBED IN EXAMPLE 3. FIG. 3A TO SHOW THE EFFECT OF SINGLE BASE PAIR (SINGLE BASE) SUBSTITUTIONS, DOUBLE BASE PAIR (DOUBLE BASE) SUBSTITUTIONS, SINGLE BASE PAIR INSERTIONS, SINGLE BASE PAIR DELETIONS, AND SINGLE BASE PAIR DELETIONS PLUS SINGLE BASE PAIR SUBSTITUTIONS AT EACH POSITION OF THE REFERENCE sgRNA SHOWN AT THE TOP. FIG. 3B TO SHOW THE EFFECT OF DOUBLE BASE PAIR INSERTIONS AND SINGLE BASE PAIR INSERTIONS PLUS SINGLE BASE PAIR SUBSTITUTIONS AT EACH POSITION OF THE MODIFIED REFERENCE sgRNA. THE REFERENCE sgRNA SEQUENCE OF SEQ ID NO: 5 IS SHOWN AT FIG. 3A TOP AND FIG. 3B BOTTOM. IN FIG. 3A AND FIG. 3B GRAYSCALE INDICATES THE Log2 FOLD ENRICHMENT OF VARIANTS IN THE DME POOL RELATIVE TO THE REFERENCE sgRNA AFTER SELECTION. ENRICHMENT IS A PROXY FOR ACTIVITY, WITH GREATER ENRICHMENT BEING A MORE ACTIVE MOLECULE. THE RESULTS SHOW THE REGIONS OF THE REFERENCE sgRNA THAT SHOULD NOT BE MUTATED AND THE KEY REGIONS TARGETED FOR MUTAGENESIS.
[0021] FIG. 4A TO SHOW THE RESULTS OF AN EXAMPLE DME EXPERIMENT USING A REFERENCE sgRNA AS DESCRIBED IN EXAMPLE 3. THE MODIFIED REFERENCE sgNA (sgRNA) WITH SEQUENCE SEQ ID NO: 5 IS SHOWN AT THE TOP AND GRAYSCALE INDICATES THE Log2 FOLD ENRICHMENT OF VARIANTS IN THE DME POOL RELATIVE TO THE REFERENCE sgRNA AFTER SELECTION. ENRICHMENT IS A PROXY FOR ACTIVITY, WITH GREATER ENRICHMENT BEING A MORE ACTIVE MOLECULE. THE HEATMAP SHOWS AN EXAMPLE DME EXPERIMENT THAT SHOWS FOUR REPLICATES OF THE POOL WHERE EACH BASE PAIR IN THE REFERENCE sgRNA HAS BEEN SUBSTITUTED WITH EACH POSSIBLE ALTERNATE BASE PAIR.
[0022] FIG. 4BA series of 8 plots to compare biological replicates of different DME libraries. Individual variants are plotted against each other for DME replicates relative to the reference sgRNA sequence. Plots are shown for single deletion, single insertion, and single substitution DME experiments as well as wild type controls and indicate that each replicate has a good amount of agreement.
[0023] FIG. 4C A heatmap showing an exemplary DME experiment for four replicates of a library in which each position in the reference sgRNA has undergone a single base pair insertion. The DME experiment used the reference sgRNA of SEQ ID NO: 5 (at the top) and was performed as described in Example 3. Log2 fold enrichment of variants in the DME library relative to the reference sgRNA after selection is indicated in grayscale.
[0024] FIG. 5A-FIG. 5E A series of plots showing that sgNA variants can improve gene editing by more than two-fold in an EGFP disruption assay as described in Examples 2 and 3. Editing was measured by indel formation and GFP disruption in HEK293 cells harboring a GFP reporter. FIG. 5A The fold change in editing efficiency of the CasX sgRNA reference of SEQ ID NO: 4 and variants with the reference of sequence SEQ ID NO: 5 across 10 targets is shown. When averaged across the 10 targets, the editing efficiency of the sgRNA of SEQ ID NO: 5 was improved by 176% compared to SEQ ID NO: 4. FIG. 5B It is possible to further improve the sgRNA scaffold of SEQ ID NO: 5 by exchanging the extended loop sequence with additional sequences to create scaffolds with sequences shown in Table 2. Fold change in editing efficiency is shown on the Y-axis. FIG. 5C A plot showing the fold improvement of sgNA variants produced by DME mutation of SEQ ID NO: 5 normalized to be the CasX reference sgRNA of SEQ ID NO: 5, including the variant with SEQ ID NO: 17. FIG. 5D A plot showing the fold improvement of sgNA variants of the sequences listed in Table 2 produced by appending a ribonuclease sequence to a reference sgRNA sequence normalized to be the SEQ ID NO: 5 reference sgRNA of CasX. FIG. 5E A plot showing the fold improvement of DME mutation of SEQ ID NO: 5 reference sgRNA produced by combining (stacking) the improved cleavage scaffold mutations shown, the improved cleavage DME mutations shown, and using the improved cleavage ribonuclease appendage shown. In this analysis, the resulting sgNA variants produced a 2-fold or greater improvement in cleavage compared to SEQ ID NO: 5. EGFP editing analysis was performed by the spacer target sequence of E6 and E7.
[0025] FIG. 6 Hepatitis delta virus (HDV) genome ribozyme used in exemplary gNA variants (SEQ ID NOs: 18-22) are shown.
[0026] FIG. 7A-FIG. 7I A series of heat maps showing the effect of single amino acid substitutions, single amino acid insertions, and deletions at each amino acid position in the reference CasX protein of SEQ ID NO: 2 as described in Example 4. Data was generated by DME analysis run at 37°C. The Y-axis shows various possible substitutions or insertions (from top to bottom: R, H, K, D, E, S, T, N, Q, C, G, P, A, I, L, M, F, W, Y, or V; square indicates amino acid identity to the reference protein), and the X-axis shows amino acid positions in the reference CasX protein. Indicated is the Log2 fold enrichment of CasX variant proteins in the DME library relative to the reference CasX protein of SEQ ID NO: 2 after enrichment. As used herein, “enrichment” is a proxy for activity, with greater enrichment being a more active molecule. (*) indicates active sites. FIG. 7A to FIG. 7D Effect of single amino acid substitutions is shown. FIG. 7E to FIG. 7H Effect of single amino acid insertions is shown. FIG. 7I Effect of single amino acid deletions is shown.
[0027] FIG. 8A-FIG. 8C A series of heat maps showing the effect of single amino acid substitutions, single amino acid insertions, and deletions at each amino acid position in the reference CasX protein of SEQ ID NO: 2 as described in Example 4. Data was generated by DME analysis run at 45°C. FIG. 8A Effect of single amino acid substitutions is shown. FIG. 8B Effect of single amino acid insertions is shown. FIG. 8C Effect of single amino acid deletions is shown. For all FIG. 8A-FIG. 8C , the Y-axis shows various possible substitutions or insertions (from top to bottom: R, H, K, D, E, S, T, N, Q, C, G, P, A, I, L, M, F, W, Y, or V; square indicates amino acid identity to the reference protein), and the X-axis shows amino acid positions in the reference CasX protein. Indicated in grayscale is the Log2 fold enrichment of CasX variant proteins relative to the reference CasX protein of SEQ ID NO: 2 in the DME library after enrichment, with greater enrichment being a more active molecule. (*) indicates active sites. Running this analysis at 45°C enriches different variants than running the same analysis at 37°C (see FIG. 7A-FIG. 7I ), thereby indicating which amino acid residues and changes are important for thermostability and folding.
[0028] FIG. 9Investigation of the comprehensive mutational landscape of all single mutations of the reference CasX protein of SEQ ID NO: 2. On the Y-axis, fold enrichment of CasX variants relative to the reference CasX protein for single substitutions (top), single insertions (middle), or single deletions (bottom). On the X-axis, amino acid position in the reference CasX protein. Key regions for generating improved CasX variants are the initial helical region and the region in the RuvC domain adjacent to the Target Strand Loading (TLS) domain, among others.
[0029] FIG. 10To show that the CasX variant proteins evaluated in Example 5 were greater than three-fold improved in editing relative to the reference CasX protein in an EGFP disruption assay, the ability of the test CasX proteins to cleave the EGFP reporter at 2 different target sites in human HEK293 cells was tested, and the normalized improvement in genome editing at these sites compared to the base reference CasX protein of SEQ ID NO: 2 was shown. From left to right, the variants (indicated by amino acid substitution, insertion, or deletion at the given residue number) are: Y789T, [P793], Y789D, T72S, I546V, E552A, A636D, F536S, A708K, Y797L, L792G, A739V, G791M, G661, A788W, K390R, A751S, E385A, P696, M773, G695H, AS793, AS795, C477R, C477K, C479A, C479L, I55F, K210R, C233S, D231N, Q338E, Q338R, L379R, K390R, L481Q, F495S, D600N, T886K, A739V, K460N, I199F, G492P, T153I, R591I, AS795, AS796, L889, E121D, S270W, E712Q, K942Q, E552K, K25Q, N47D, T696, L685I, N880D, Q102R, M734K, A724S, T704K, P224K, K25R, M29E, H152D, S219R, E475K, G226R, A377K, E480K, K416E, H164R, K767R, I7F, M29R, H435R, E385Q, E385K, I279F, D489S, D732N, A739T, W885R, E53K, A238T, P283Q, E292K, Q628E, R388Q, G791M, L792K, L792E, M779N, G27D, K955R, S867R, R693I, F189Y, V635M, F399L, E498K, E386S, V254G, P793S, K188E, QT945KI, T620P, T946P, TT949PP, N952T, K682E, K975R, L212P, E292R, I303K, C349E, E385P, E386N, D387K, L404K, E466H, C477Q, C477H, C479A, D659H, T806V, K808S, AS797, V959M, K975Q, W974G, A708Q, V711K, D733T, L742W, V747K,F755M, M771A, M771Q, W782Q, G791F, L792D, L792K, P793Q, P793G, Q804A, Y966N, Y723N, Y857R, S890R, S932M, L897M, R624G, S603G, N737S, L307K, I658V, PT688, SA794, S877R, N580T, V335G, T620S, W345G, T280S, L406P, A612D, A751S, E386R, V351M, K210N, D40A, E773G, H207L, T62A, T287P, T832A, A893S, V14, AG13, R11V, R12N, R13H, Y13, R12L, Q13, V15S, D17. Indicate insertions, [] indicate deletions.
[0030] FIG. 11To show that individual beneficial mutations can be combined (sometimes referred to as “stacked”) to achieve even greater improvements in gene editing activity. The ability of CasX proteins to cleave at 2 different target sites in human HEK293 cells was tested using E6 and E7 spacers targeting an EGFP reporter, as described in Example 5. From left to right, the variants are: S794R+Y797L, K416E+A708K, A708K+[P793], [P793]+P793AS, Q367K+I425S, A708K+[P793]+A793V, Q338R+A339E, Q338R+A339K, S507G+G508R, L379R+A708K+[P793], C477K+A708K+[P793], L379R+C477K+A708K+[P793], L379R+A708K+[P793]+A739V, C477K+A708K+[P793]+A739V, L379R+C477K+A708K+[P793]+A739V, L379R+A708K+[P793]+M779N, L379R+A708K+[P793]+M771N, L379R+A708K+[P793]+D489S, L379R+A708K+[P793]+A739T, L379R+A708K+[P793]+D732N, L379R+A708K+[P793]+G791M, L379R+A708K+[P793]+Y797L, L379R+C477K+A708K+[P793]+M779N, L379R+C477K+A708K+[P793]+M771N, L379R+C477K+A708K+[P793]+D489S, L379R+C477K+A708K+[P793]+A739T, L379R+C477K+A708K+[P793]+D732N, L379R+C477K+A708K+[P793]+G791M, L379R+C477K+A708K+[P793]+Y797L, L379R+C477K+A708K+[P793]+T620P, A708K+[P793]+E386S, E386R+F399L+[P793], and R4581I+A739V of SEQ ID NO: 2. [] refers to a deletion of the amino acid residue at the indicated position of SEQ ID NO: 2.
[0031] FIG. 12A and FIG. 12BA pair of graphs showing that CasX proteins and sgNA variants can improve activity more than 6-fold relative to reference sgRNA and reference CasX protein when combined. sgNA: Analyzing the ability of protein pairs to cleave a GFP reporter in HEK293 cells as described in Example 5. On the Y-axis, the fraction of cells in which expression of the GFP reporter was disrupted by CasX-mediated gene editing is shown. FIG. 12A CasX proteins and sgNA analyzed by targeting the E6 spacer of GFP. FIG. 12B CasX proteins and sgNA analyzed by targeting the E7 spacer of GFP. iGFP indicates “inducible GFP”.
[0032] FIG. 13A , FIG. 13B and FIG. 13C Manufacturing and screening DME libraries allowed the generation and identification of variants exhibiting 1- to 81-fold improvement in editing efficiency as described in Examples 1 and 3. FIG. 13A RFP+ and GFP+ reporters in E. coli cells were analyzed for inhibition of GFP by CRISPR interference with reference nuclease-dead CasX proteins and sgNA. FIG. 13B The same reporter cells were analyzed for GFP inhibition by nuclease-dead CasX variants screened from DME libraries. FIG. 13CReference, the selected CasX proteins and sgNA variants improved editing efficiency. The Y axis shows disruption of B2M staining by HLA1 antibody, indicating gene disruption by CasX editing and indel formation. In the case of guide spacer #43, the improved CasX variants improved editing of this locus up to 81-fold compared to the reference. CasX was paired with the reference sgRNA: protein pair of SEQ ID NO: 5 and SEQ ID NO: 2, and indicates a CasX variant protein of SEQ ID NO: 2 L379R + A708K + [P793], analyzed by sgNA variants with a truncated stem loop and T10C substitutions, the variants encoded by the sequence TACTGGCGCCTTTATCTCATTACTTTGAGAGCCATCACCAGCGACTATGTCGTATGGGTAA AGCGCTTACGGACTTCGGTCCGTAAGAAGCATCAAAG (SEQ ID NO: 23). The following spacer sequences were used: #9: GTGTAGTACAAGAGATAGAA (SEQ ID NO: 24); #14: TGAAGCTGACAGCATTCGGG (SEQ ID NO: 25), #20: tagATCGAGACATGTAAGCA (SEQ ID NO: 26); #37: GGCCGAGATGTCTCGCTCCG (SEQ ID NO: 27) and #43: AGGCCAGAAAGAGAGAGTAG (SEQ ID NO: 28).
[0033] FIG. 14A-FIG. 14F A series of structural models of the prototypical CasX protein, showing the positions of mutations of the CasX variant proteins of the application that exhibit improved activity. FIG. 14A Showing the deletion of P at position 793 of SEQ ID NO: 2, where a deletion in the loop can affect folding. FIG. 14B Showing the substitution of alanine (A) with lysine (K) at position 708 of SEQ ID NO: 2. This mutation faces the 5' end of the gNA plus a salt bridge to the gNA. FIG. 14C Showing the substitution of cysteine (C) with lysine (K) at position 477 of SEQ ID NO: 2. This mutation faces the gNA. There is a salt bridge to the gNA phosphate backbone (gNAbb) at approximately base 14 that can be affected. This mutation removes a surface-exposed cysteine. FIG. 14D Showing the substitution of leucine (L) with arginine (R) at position 379 of SEQ ID NO: 2. There is a salt bridge to the target DNA phosphate backbone (DNAbb) towards base pairs 22-23 that can be affected. FIG. 14EA view showing the combination of the absence of P at 793 and the A708K substitution. FIG. 14F An alternative view showing that the effects of individual mutants are additive, and that single mutants can be combined (stacked) to achieve even greater improvements. Throughout FIG. 14A-FIG. 14F , the arrow indicates the mutation position.
[0034] FIG. 15 A graph showing the identification of optimal Protophtyte CasX PAMs and spacers for genes of interest as described in Example 6. On the Y-axis, % GFP negative cells are shown, which indicates cleavage of the GFP reporter. On the X-axis, different PAM sequences and spacers: ATC PAM, CTC PAM, and TTC PAM. GTC, TTT, and CTT PAMs were also tested and are not shown to be active.
[0035] FIG. 16 A graph showing that improved CasX variants produced by DME can edit canonical and non-canonical PAMs more efficiently compared to a reference CasX protein as described in Example 6. The Y-axis shows the average fold improvement in editing relative to the reference with 2 targets, N = 6 sgRNA: protein pairs (SEQ ID NO: 2, SEQ ID NO: 5). From left to right, the protein variants for each set of bars are: A708K + [P793] + A739V; L379R + A708K + [P793]; C477K + A708K + [P793]; L379R + C477K + A708K + [P793]; L379R + A708K + [P793] + A739V; C477K + A708K + [P793] + A739V; and L379R + C477K + A708K + [P793] + A739V. The reference CasX and protein variants were analyzed by the reference sgRNA scaffold of SEQ ID NO: 5 with the following DNA encoding spacer sequences from left to right: E6 with TTC PAM (SEQ ID NO: 29); E7 with TTC PAM (SEQ ID NO: 30); GFP8 with TTC PAM (SEQ ID NO: 31); B1 with CTC PAM (SEQ ID NO: 32); and A7 with ATC PAM (SEQ ID NO: 33).
[0036] FIG. 17A to FIG. 17F A series of graphs showing a reference CasX protein and reference sgRNA scaffold pair with high specificity for a target sequence as described in Example 7. FIG. 17A and FIG. 17DThe ability of the editing template to be edited was analyzed for S. pyogenes Cas9 (SpyCas9) with two different gNA spacers and 5' PAM sites (SEQ ID NO: 34-65) and (SEQ ID NO: 136-166) with either a target sequence (arrow) complementary to the spacer sequence or with 1, 2, 3, or 4 mutations in the target sequence relative to the spacer sequence. FIG. 17B and FIG. 17E The ability of the editing template to be edited was analyzed for S. pyogenes Cas9 (SauCas9) with two different gNA spacers and 5' PAM sites (SEQ ID NO: 66-103) and (SEQ ID NO: 167-204) with either a target sequence (arrow) complementary to the spacer sequence or with 1, 2, 3, or 4 mutations in the target sequence relative to the spacer sequence. FIG. 17C and FIG. 17F The ability of the editing template to be edited was analyzed for the reference Plm CasX protein and sgNA scaffold pair with two different gNA spacers and 3' PAM sites (SEQ ID NO: 104-135) and (SEQ ID NO: 205-236) with either a target sequence (arrow) complementary to the spacer sequence or with 1, 2, 3, or 4 mutations in the target sequence relative to the spacer sequence. In all of FIG. 17A-FIG. 17F The X-axis shows the fraction of cells in which gene editing occurred at the target sequence.
[0037] FIG. 18 Illustrating the scaffold stem loop of an exemplary reference sgRNA of the disclosure (SEQ ID NO: 237).
[0038] FIG. 19 Illustrating the extended stem loop sequence of an exemplary reference sgRNA of the disclosure (SEQ ID NO: 238).
[0039] FIG. 20A-FIG. 20B A pair of graphs showing that a particular subset of changes discovered by the DME of CasX as described in Example 4 are more likely to predict improved activity. The graphs represent data from the experiments described in Figures 7 and 8. FIG. 20A Show that changing residues within 10 A of the guide RNA and hydrophobic residues (A, V, I, L, M, F, Y, W) results in proteins with significantly lower activity. FIG. 20B Show that changing residues within 10 A of the RNA and positively charged amino acids (R, H, K) can improve activity.
[0040] FIG. 21 Illustrating an alignment of two reference CasX protein sequences (SEQ ID NO: 1, top; SEQ ID NO: 2, bottom) with domains annotated.
[0041] FIG. 22 Illustration of the domain organization of the reference CasX protein of SEQ ID NO: 1. Domains have the following coordinates: Non-target strand binding (NTSB) domain: amino acids 101-191; Helical I domain: amino acids 57-100 and 192-332; Helical II domain: 333-509; Oligonucleotide binding domain (OBD): amino acids 1-56 and 510-660; RuvC DNA cleavage domain (RuvC): amino acids 551-824 and 935-986; Target strand loading (TSL) domain: amino acids 825-934. Note that the Helical I, OBD, and RuvC domains are non-contiguous.
[0042] FIG. 23 Illustration of the alignment of two CasX reference sgRNA scaffolds, SEQ ID NO: 5 (top) and SEQ ID NO: 4 (bottom).
[0043] FIG. 24 Presented is a SDS-PAGE gel of the StX2 (CasX reference of SEQ ID NO: 2) purification elution fractions observed by colloidal coomassie staining as described in Example 8. From left to right, lanes are: Pellet: insoluble fraction after cell lysis, Lysate: soluble fraction after cell lysis, Flow-Through: proteins that did not bind the heparin column, Wash: proteins eluted from the column in wash buffer, Elution: proteins eluted from the heparin column with elution buffer, Flow-Through: proteins that did not bind the StrepTactin column, Elution: proteins eluted from the StrepTactin column with elution buffer, Injected: concentrated proteins injected onto the s200 gel filtration column, Frozen: concentrated and frozen pooled elution fractions from the s200 elution.
[0044] FIG. 25 Presented is a chromatogram of the size exclusion chromatography analysis from StX2 as described in Example 8.
[0045] FIG. 26 Presented is a SDS-PAGE gel of the StX2 purification elution fractions observed by colloidal coomassie staining as described in Example 8. From left to right: injected sample, molecular weight marker, lanes 3-9: samples from the indicated elution volumes.
[0046] FIG. 27 Presented is a chromatogram of the size exclusion chromatography analysis from CasX 119 using Superdex 200 16 / 600 pg gel filtration as described in Example 8. The 67.47 mL peak corresponds to the apparent molecular weight of CasX variant 119 and contains the majority of the CasX variant 119 protein.
[0047] FIG. 28 A SDS-PAGE gel of purified fractions of CasX 119 viewed by colloidal Coomassie staining as described in Example 8 is presented. Samples from the indicated elution fractions were resolved by SDS-PAGE and stained by colloidal Coomassie. From left to right, injections: protein sample injected onto the gel filtration column, molecular weight marker, lanes 3-10: samples from the indicated elution volumes.
[0048] FIG. 29 A SDS-PAGE gel of purified fractions of CasX 438 viewed on a Bio-Rad Stain-Free TM gel. From left to right, lanes are: Pellet: insoluble fraction after cell lysis, Lysate: soluble fraction after cell lysis, Flow-through: proteins that did not bind the heparin column, Elution: proteins eluted from the heparin column with elution buffer, Flow-through: proteins that did not bind the StrepTactin column, Elution: proteins eluted from the StrepTactin column with elution buffer, Injection: concentrated protein injected onto the s200 gel filtration column, Pool: pooled CasX-containing elution fractions, Final: pooled elution fractions from s200 elution that have been concentrated and frozen.
[0049] FIG. 30 A chromatogram from size exclusion chromatography analysis of CasX 438 using Superdex 200 16 / 600 pg gel filtration as described in Example 8 is presented. The 69.13 mL peak corresponds to the apparent molecular weight of CasX variant 438 and contains most of the CasX variant 438 protein.
[0050] FIG. 31 A SDS-PAGE gel of purified fractions of CasX 438 viewed by colloidal Coomassie staining as described in Example 8 is presented. Samples from the indicated elution fractions were resolved by SDS-PAGE and stained by colloidal Coomassie. From left to right, injections: protein sample injected onto the gel filtration column, molecular weight marker, lanes 3-10: samples from the indicated elution volumes.
[0051] FIG. 32 A SDS-PAGE gel of purified fractions of CasX 438 viewed on a Bio-Rad Stain-Free TMSDS-PAGE gel of purified CasX 457 sample observed on a gel. From left to right, the channels are: Aggregate: insoluble fraction after cell lysis; Soluble fraction: soluble fraction after cell lysis; Through: proteins not bound to the heparin column; Elution: proteins eluted from the heparin column with elution buffer; Through: proteins not bound to the StrepTactin column; Elution: proteins eluted from the StrepTactin column with elution buffer; Injection: concentrated proteins injected onto the S200 gel filter column; Final: concentrated and frozen aggregate eluent from the S200 column.
[0052] FIG. 33 Present the chromatogram of size exclusion chromatography analysis of CasX 457 using a Superdex 200 16 / 600 pg gel filter as described in Example 8. The 67.52 mL peak corresponds to the apparent molecular weight of CasX variant 457 and contains most of the CasX variant 457 protein.
[0053] FIG. 34 Present the SDS-PAGE gel of the CasX 457 purified eluent as described in Example 8, observed by colloidal Coomassie staining. Samples from the specified eluents were resolved by SDS-PAGE and stained by colloidal Coomassie staining. From left to right: Injection: protein sample injected onto the gel filter column; molecular weight marker; channels 3-10: sample from the specified elution volume.
[0054] FIG. 35 A schematic diagram showing the organization of components in the pSTX34 plasmid used to assemble the CAX construct as described in Example 9.
[0055] FIG. 36 A schematic diagram illustrating the steps for generating the CasX 119 variant as described in Example 9.
[0056] FIG. 37 This is a graphical representation of the results of quantitative analysis of the activity fractions of RNPs formed from sgRNA174 and CasX variants 119 and 457 as described in Example 19. Equimolar amounts of RNP and target were co-incubated, and the amount of target cleavage was determined at specified time points. The mean and standard deviation of three independent copies are shown for each time point. A biphasic fit of the pooled copies is presented. “2” refers to the reference CasX protein of SEQ ID NO:2.
[0057] FIG. 38Graphical representation of the results of the quantitative analysis of the activity fraction of RNP formed by CasX2 and reference guide 2, modified sgRNA guides 32, 64, and 174 as described in Example 19. Equal molar amounts of RNP and target were co-incubated and the amount of target cleaved was determined at the indicated time points. The mean and standard deviation of three independent replicates are shown for each time point. Biphasic fits of the pooled replicates are shown. “2” refers to reference gRNA SEQ ID NO: 5, and modified sgRNAs are indicated by the number in Table 2.
[0058] FIG. 39 Graphical representation of the results of the quantitative analysis of the cleavage rate of RNP formed by sgRNA 174 and CasX variants 119 and 457 as described in Example 19. Target DNA was incubated with a 20-fold excess of the indicated RNP and the amount of target cleaved was determined at the indicated time points. The mean and standard deviation of three independent replicates are shown for each time point. Monophasic fits of the pooled replicates are shown.
[0059] FIG. 40 Graphical representation of the results of the quantitative analysis of the cleavage rate of RNP formed by CasX2 and sgRNA guide variants 2, 32, 64, and 174 as described in Example 19. Target DNA was incubated with a 20-fold excess of the indicated RNP and the amount of target cleaved was determined at the indicated time points. The mean and standard deviation of three independent replicates are shown for each time point. Monophasic fits of the pooled replicates are shown.
[0060] FIG. 41 Graphical representation of the results of the quantitative analysis of the initial velocity of RNP formed by CasX2 and sgRNA guide variants 2, 32, 64, and 174 as described in Example 19. The first two time points of the preceding cleavage experiment were fit to a linear model to determine the initial cleavage velocity.
[0061] FIG. 42 Graphical representation showing an example of CasX protein and scaffold DNA sequence packaged in adeno-associated virus (AAV) as described in Example 20. The DNA segment between the AAV inverted terminal repeats (ITRs) composed of DNA encoding CasX and its promoter, and DNA encoding the scaffold and its promoter becomes packaged within the AAV capsid during AAV production.
[0062] FIG. 43 Graph showing representative results of AAV titration by qPCR as described in Example 20. Flow through (FT) and consecutive elution fractions (1-6) were collected during AAV purification and titrated by qPCR. The majority of the virus (in this example, about le14 viral genomes) is found in the second elution fraction.
[0063] FIG. 44Results from AAV-mediated gene editing experiments in the SOD1-GFP reporter cell line as described in Example 21 are shown. CasX constructs with SOD1 targeting spacer 2 (CasX 119 and guide 64, ATGTTCATGAGTTTGGAGAT; SEQ ID NO: 239) and SauCas9 were packaged in AAV vectors and used to transduce SOD1-GFP reporter cells at a range of different multiplicity of infection (MOI, number of viral genomes per cell). Twelve days later, cells were analyzed for GFP disruption via FACS. In this example, CasX and SauCas9 showed equivalent levels of editing, with 1-2% of cells showing GFP disruption at the highest MOI, 1e7 or 1e6.
[0064] FIG. 45 Results from a second AAV-mediated gene editing experiment in the SOD1-GFP reporter cell line as described in Example 21 are shown. CasX construct 119.64 with SOD1 targeting spacer 2 (ATGTTCATGAGTTTGGAGAT; SEQ ID NO: 239) and SauCas9 with SOD1 targeting spacer were packaged in AAV vectors and used to transduce SOD1-GFP reporter cells at a range of different multiplicity of infection (MOI, number of viral genomes per cell). Twelve days later, cells were analyzed for GFP disruption via FACS. In this example, CasX and SauCas9 showed equivalent levels of editing at the highest MOI, with about 2-4% of cells showing GFP disruption.
[0065] FIG. 46 Results from AAV-mediated gene editing experiments in neural progenitor cells (NPCs) from the G93A mouse model of ALS as described in Example 21 are shown. CasX construct (CasX 119 and guide 64, with SOD1 targeting spacer 2, ATGTTCATGAGTTTGGAGAT; SEQ ID NO: 239) was packaged in AAV vectors and used to transduce G93A NPCs at a range of different multiplicity of infection (MOI, number of viral genomes per cell). Twelve days later, cells were analyzed for gene editing via T7E1 analysis. The agarose gel images from T7E1 analysis shown here indicate successful editing of the SOD1 locus. The double arrows show the two bands of DNA attributed to successful editing in the cells.
[0066] FIG. 47 Results from editing analysis of 6 target genes in HEK293T cells as described in Example 23 are shown. Each point represents the results using an individual spacer.
[0067] FIG. 48Results of editing analysis of 6 genes of interest in HEK293T cells as described in Example 23 are shown, where individual bars represent results obtained from individual spacers.
[0068] FIG. 49 Results of editing analysis of 4 genes of interest in HEK293T cells as described in Example 23 are shown. Each dot represents results using individual spacers, using the CTC (CTCN) PAM.
[0069] FIG. 50 A schematic of the deep mutational evolution steps used to generate the gene libraries encoding CasX variants as described in Example 24 is shown. pSTX1 backbone is minimal, consisting only of a high copy number origin and KanR resistance gene, making it compatible with the recombinantly engineered E. coli strain EcNR2. pSTX2 is a BsmbI destination plasmid for aTc inducible expression in E. coli.
[0070] FIG. 51 A dot plot showing results of CRISPRi screening for mutations in libraries D1, D2, and D3 as described in Example 24 is shown. In the absence of CRISPRi, E. coli constitutively expresses both GFP and RFP, producing intense fluorescence at both wavelengths, represented by the dots in the upper right region of the plot. CasX proteins for CRISPRi that produce GFP can reduce green fluorescence >10-fold while leaving red fluorescence unchanged, and these cells belong to the indicated sort gate 1. The total fraction of cells exhibiting CRISPRi is indicated.
[0071] FIG. 52 A photograph of colony growth in the ccdB assay as described in Example 24 is shown. 10-fold dilutions were assayed in the presence of glucose or arabinose to induce expression of the ccdB toxin, which produces approximately 1000-fold difference between functional and nonfunctional proteins. When grown in liquid culture, the resolving power is approximately 10,000-fold, as seen on the right-hand side.
[0072] FIG. 53 A plot of HEK iGFP genome editing efficiency testing CasX variants with sgRNA2 (SEQ ID NO: 5) and appropriate spacers as described in Example 24 is shown, where data is expressed as fold improvement over wild-type CasX protein (SEQ ID NO: 2) in the HEK iGFP editing assay. Individual mutations are shown at the top, and mutation sets are shown at the bottom of the plot). Error bars combine intra-assay error (SD) and inter-assay error (for those variants tested more than once, SD across replicate experiments) in assays performed at least in triplicate.
[0073] FIG. 54A scatter plot showing the results of SOD1-GFP reporter assays for CasX variants with sgRNA scaffold 2, using two different spacers as described in Example 24 for GFP.
[0074] FIG. 55 A plot showing the results of HEK293 iGFP genome editing assays comparing wild type CasX (SEQ ID NO: 2) and CasX variant 119 across four different PAM sequences, as described in Example 24; both the CasX and CasX variant used sgRNA scaffold 1 (SEQ ID NO: 4) with spacers using four different PAM sequences.
[0075] FIG. 56 A plot showing the results of genome editing activity of CasX variant 119 and sgRNA 174 compared to wild type CasX 2 and guide scaffold 1 in an iGFP lipofection assay using two different spacers, as described in Example 24.
[0076] FIG. 57 A plot showing the results of genome editing activity of CasX variant 119 and sgRNA 174 compared to wild type CasX and guide in an iGFP lentiviral transduction assay using two different spacers, as described in Example 24.
[0077] FIG. 58 A plot showing the results of genome editing in a more stringent lentiviral assay for comparing editing activity of four CasX variants (119, 438, 488, and 491) and optimized sgNA 174 and two different spacers, as described in Example 24. Results show stepwise improvement in editing efficiency achieved by introducing additional modifications and domain swaps into the starting point 119 variant.
[0078] FIG. 59A-FIG. 59B Results of NGS analysis of a library of sgRNAs, as described in Example 25, are shown. FIG. 59A Distribution of substitutions, deletions, and insertions are shown. FIG. 59B A scatter plot showing the high reproducibility of variant representation in two independent library pools following CRISPRi analysis in an unsorted, untreated cell population. (Library pool D3 is two different versions of the dCasX protein relative to D2, and represents replicates of the CRISPRi analysis).
[0079] FIG. 60A-FIG. 60B Structure of wild type CasX and RNA guide (SEQ ID NO: 4) is shown. FIG. 60ACryoEM structure depicting a L. delbrueckii CasX protein: sgRNA RNP complex (PDB ID: 6YN2) including two stem-loops, a pseudoknot, and a triple helix. FIG. 60B Depiction of the secondary structure of the sgRNA identified from the structure shown in (A) using the tool RNAPDBee 2.0 (rnapdbee.cs.put.poznan.pl / ), using the tool 3DNA / DSSR, and visualized using VARNA. RNA regions are indicated. Residues not apparent in the PDB crystal structure file are indicated by plain text letters (i.e., not circled), and are not included in the residue numbering.
[0080] FIG. 61A-FIG. 61C Depiction of a comparison between the two guide RNA scaffolds. FIG. 61A A sequence alignment between the single guide scaffold 1 (SEQ ID NO: 4) and scaffold 2 (SEQ ID NO: 5) is provided. FIG. 61B The predicted secondary structure of scaffold 1 is shown (without the 5’ ACAUCU bases, which are not in the cryoEM structure). The prediction was made using RNAfold (v 2.1.7) using the constraint derived from the base pairing observed in the cryoEM structure (see FIG. 60A-FIG. 60B ). This constraint requires that base pairs observed in the cryoEM structure be formed, and that bases involved in triple helix formation not be paired. This structure has unique base pairing from the lowest energy predicted structure at the 5’ end (i.e., the pseudoknot and triple helix loop). FIG. 61C The predicted secondary structure of scaffold 2 is shown. The prediction was made using a similar constraint based on the sequence alignment for scaffold 1.
[0081] FIG. 62 A plot showing the GFP knockdown ability of scaffold 1 relative to scaffold 2 using four different spacers utilizing different PAM sequences in a GFP-lipofectin transfection assay is shown, as described in Example 25. The results show that greater editing is conferred by using the modified scaffold 2 compared to the wild type scaffold 1; the wild type scaffold 1 is not shown edited by using spacers with GTC and CTC PAM sequences.
[0082] FIG. 63A-FIG. 63C A plot depicting enrichment across single variants of the scaffold is shown, exhibiting mutable regions, as described in Example 25. FIG. 63A Depiction of substituted bases (A, T, G, or C; top to bottom), FIG. 63B Depiction of inserted bases (A, T, G, or C; top to bottom), and FIG. 63CDepiction of deletions across scaffold 2 at individual nucleotide positions (X-axis). Average WT values were relative to average WT values. Enrichment values were averaged across the three catalytically dead CasX versions. Scaffolds with relative log2 enrichment > 0 were considered ‘enriched’ as they were more represented in the sorted population relative to the wild type scaffold represented. Error bars represent the confidence interval across three catalytically dead CasX experiments.
[0083] FIG. 64 To show the scatter plot of enrichment values obtained across different dCasX variants that are largely consistent, as described in Example 25. Library D2 and DDD have highly correlated enrichment scores, while D3 is more unique.
[0084] FIG. 65 To show a bar graph of cleavage activity at the SOD1-GFP locus of several scaffold variants in a more stringent lipofection analysis, as described in Example 25.
[0085] FIG. 66 To show a bar graph of cleavage activity of several scaffold variants using two different spacers; spacers are 8.2 and 8.4 targeting the SOD1-GFP locus (and non-targeting spacer NT) with low MOI lentiviral transduction using p34 plasmid backbone, as described in Example 25.
[0086] FIG. 67 To show a schematic of the secondary structure (top) and linear structure (bottom) of single guide 174, where lines connect those segments that are bound by base pairing or other non-covalent interactions. Scaffold stems (white, no fill) (and loops) and extended stems (gray, no fill) (and loops) are adjacent 5’ to 3’ in the sequence. However, pseudoknots and extensions are formed by strands that have intervening regions in the sequence. In the case of single guide 174, a triple helix is formed that includes nucleotides 5’-CUUUG’-3’ and 5’-CAAAG-3’ that form a base-paired double helix, and nucleotides 5’-UUU-3’ that form a triple helix region in conjunction with 5’-AAA-3’.
[0087] Figure 68 shows a comparison between the highly evolved single guide 174 and scaffolds 1 and 2 that serve as the starting point for the DME procedure described in Example 25. FIG. 68A To show a bar graph of cleavage activity of head-to-head comparison of guide scaffolds with five different spacers in a plasmid lipofection transfection analysis at the GFP locus in HEK-GFP cells. FIG. 68B To show a sequence alignment between scaffold 2 and guide 174 (SEQ ID NO: 2238). Asterisks indicate point mutations, and dashed boxes show the entire extended stem exchange.
[0088] FIG. 69A-69BPresenting the stent sequence relative to having 2 spacers (4.76) FIG. 69A ) and 4.77 FIG. 69B The scatter plot of HEK-iGFP lysis analysis of the WT scaffold is shown in Example 25.
[0089] FIG. 70 Present a scatter plot of normalized cleavage activity for several scaffolds compared to WT with two spacers (4.76 and 4.77), as described in Example 25. Error bars combine internal measurement error (SD) and inter-experimental measurement error (SD spans repeated experiments for those variants tested more than once) (orthogonal).
[0090] FIG. 71 Present a scatter plot comparing the normalized cleavage activity of multiple scaffolds relative to WT in HEK-iGFP cleavage analysis with enrichments obtained from CRISPRi comprehensive screening, as described in Example 25. Generally, scaffold mutations with high enrichment (>1.5) exhibit similar or greater activity to WT. Two variants showed high cleavage activity with low enrichment scores (C18G and T17G); interestingly, these substitutions occurred at the same locations as several highly enriched insertions (…). FIG. 63A-FIG. 63C The marker indicates a mutation in the comparison subgroup.
[0091] FIG. 72 The results of flow cytometry analysis of Cas-mediated editing at the RHO locus in APRE19 RHO-GFP cells 14 days post-transfection for CasX variant constructs 438, 499, and 491 are presented as described in Example 26. Dots represent results for individual samples, and light dashed lines represent the upper and lower quartiles.
[0092] FIG. 73 Present the quantification of cleavage rates of targets with different PAMs using RNPs formed from sgRNA174 and CasX variants. Target DNA was incubated with a 20-fold overdose of the specified RNPs, and the amount of cleaved target was determined at specified time points. Present the single-phase fit of the merged duplicates. Detailed Implementation
[0093] While exemplary embodiments have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Many variations, modifications, and substitutions will now occur to those skilled in the art without departing from the invention asserted herein. It should be understood that various alternatives to the embodiments described herein can be used to practice embodiments of the invention. The claims are intended to define the scope of the invention and therefore cover methods and structures within the scope of these claims and their equivalents.
[0094] definition
[0095] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. Although methods and materials similar or equivalent to those described herein can be used in the practice or testing of embodiments of the present application, suitable methods and materials are described below. In case of conflict, the patent specification, including definitions, will control. In addition, the materials, methods, and examples are illustrative only and are not intended to be limiting. Numerous variations, changes, and substitutions will now occur to those skilled in the art without departing from the application.
[0096] The terms "polynucleotide" and "nucleic acid" are used interchangeably herein to refer to a polymeric form of nucleotides of any length, either ribonucleotides or deoxyribonucleotides. Thus, this term includes single-stranded DNA; double-stranded DNA; multiple forms of DNA; single-stranded RNA; double-stranded RNA; multiple forms of RNA; genomic DNA; cDNA; DNA-RNA hybrids; and mixed polymers comprising ribonucleotides and deoxyribonucleotides. The polynucleotide can be isolated or recombinant.
[0097] "Hybridizable" or "complementary" are used interchangeably and mean that a nucleic acid (e.g., RNA, DNA) comprises a sequence of nucleotides that enables it to specifically bind to another nucleic acid under conditions of thermal and solution ionic strength that are permissive for sequence-specific, antiparallel (i.e., nucleic acid-specific binding to a complementary nucleic acid), noncovalent binding (i.e., Watson-Crick base pairing and / or G / U base pairing), "annealing" or "hybridization." It is understood that the sequence of a polynucleotide need not be 100% complementary to the target nucleic acid to which it is intended to hybridize specifically; it can have at least about 70%, at least about 80%, or at least about 90%, or at least about 95% sequence identity and still hybridize to the target nucleic acid. Moreover, a polynucleotide can hybridize over one or more segments so that intervening or adjacent segments are not involved in the hybridization event (e.g., looped out, "bulges," etc.).
[0098] For purposes of the present application, a "gene" includes DNA regions encoding a gene product (e.g., a protein, an RNA) as well as all DNA regions which regulate the production of the gene product, whether or not such regulatory sequences are adjacent to coding and / or transcribed sequences. A gene can therefore include regulatory sequences, including but not limited to promoters, terminators, translational control sequences (such as ribosome binding sites and internal ribosome entry sites), enhancers, silencers, insulators, boundary elements, replication origins, matrix attachment regions, and locus control regions. Coding sequences encode gene products after transcription or transcription and translation; coding sequences of the present application can comprise fragments and need not contain a full-length open reading frame. A gene can include transcribed strands, such as strands containing coding sequences, as well as complementary strands.
[0099] The term "downstream" refers to a nucleotide sequence located 3' of a reference nucleotide sequence. In certain embodiments, a downstream nucleotide sequence is associated with a sequence following the transcription start site. For example, the translation initiation codon of a gene is located downstream of the transcription start site.
[0100] The term "upstream" refers to a nucleotide sequence located 5' of a reference nucleotide sequence. In certain embodiments, an upstream nucleotide sequence is associated with a sequence located 5' of a coding region or transcription start site. For example, most promoters are located upstream of the transcription start site.
[0101] The term "regulatory element" is used interchangeably herein with the term "regulatory sequence" and is intended to include promoters, enhancers, and other expression regulating elements (e.g., transcription termination signals such as polyadenylation signals and poly-U sequences). Exemplary regulatory elements include transcription promoters such as, but not limited to, CMV, CMV+, intron A, SV40, RSV, HIV-Ltr, elongation factor 1 alpha (EFl alpha), MMLV-ltr, internal ribosome entry site (IRES) or P2A peptide to permit translation of multiple genes from a single transcript, metallothionein, transcriptional enhancer elements, transcription termination signals, polyadenylation sequences, sequences for optimizing translation initiation, and translation termination sequences. It will be appreciated that the selection of appropriate regulatory elements will depend on the encoded component (e.g., protein or RNA) to be expressed or whether the nucleic acid comprises multiple components that require different polymerases or are not intended to be expressed as a fusion protein.
[0102] The term "promoter" refers to a DNA sequence containing an RNA polymerase binding site, a transcription initiation site, a TATA box, and / or a B recognition element and which facilitates or promotes transcription and expression of an associated transcribable polynucleotide sequence and / or gene (or transgene). A promoter can be produced synthetically or can be derived from a known or naturally occurring promoter sequence or another promoter sequence. A promoter can be proximal or distal to a gene to be transcribed. A promoter can also include a chimeric promoter comprising a combination of two or more heterologous sequences to impart certain properties. A promoter of the present application can include variants of promoter sequences that are similar in composition to other promoter sequences known or provided herein, but are not identical thereto. A promoter can be classified according to standard categories related to the expression pattern of an associated coding or transcribable sequence or gene that is operably linked to the promoter, such as constitutive, developmental, tissue-specific, inducible promoters, and the like.
[0103] The term "enhancer" refers to a regulatory element DNA sequence that modulates the expression of an associated gene when bound by specific proteins called transcription factors. Enhancers can be located in introns of a gene, or at 5' or 3' of the coding sequence of a gene. Enhancers can be proximal to a gene (i.e., within tens or hundreds of base pairs (bp) of a promoter), or can be located distally (i.e., thousands of bp, hundreds of thousands of bp, or even millions of bp away from a promoter). A single gene can be modulated by more than one enhancer, all of which are contemplated to be within the scope of the present application.
[0104] As used herein, "recombinant" means that a specified nucleic acid (DNA or RNA) is the product of various combinations of cloning, restriction, and / or ligation steps, resulting in a construct having structural coding or non-coding sequences distinguishable from endogenous nucleic acids found in natural systems. In general, the DNA sequences encoding the structural coding sequences can be assembled from cDNA fragments and short oligonucleotide adapters, or from a series of synthetic oligonucleotides, to give a synthetic nucleic acid capable of being expressed from a recombinant transcriptional unit contained in a cell or in an in vitro transcription and translation system. Such sequences can be provided in the form of an open reading frame uninterrupted by internal non-translated sequences, or introns, which are typically found in eukaryotic genes. Genomic DNA containing the relevant sequences can also be used in forming a recombinant gene or transcriptional unit. Sequences of non-translated DNA can be present 5' or 3' of the open reading frame, where such sequences do not interfere with manipulation or expression of the coding region, and can in fact be used to modulate production of the desired product by various mechanisms (see "enhancer" and "promoter" above).
[0105] The term "recombinant polynucleotide" or "recombinant nucleic acid" refers to a polynucleotide or nucleic acid that does not occur in nature, e.g., made by artificial combination of two otherwise separate segments of sequence via human intervention. Such artificial combination is typically accomplished by chemical synthesis means or by artificial manipulation of separate segments of nucleic acid, such as by genetic engineering techniques. Such manipulations can be performed to replace codons with redundant codons that encode the same or conservative amino acids, while often introducing or removing sequence recognition sites. Alternatively, they are performed to join together nucleic acid segments that have a desired function to produce a functional desired combination. Such artificial combination is typically accomplished by chemical synthesis means or by artificial manipulation of separate segments of nucleic acid, such as by genetic engineering techniques.
[0106] Similarly, the term "recombinant polypeptide" or "recombinant protein" refers to a polypeptide or protein that does not occur in nature, e.g., made by artificial combination of two otherwise separate segments of amino sequence via human intervention. Thus, for example, a protein comprising a heterologous amino acid sequence is recombinant.
[0107] As used herein, the term "contacting" means establishing a physical connection between two or more entities. For example, contacting a target nucleic acid with a guide nucleic acid means that the target nucleic acid and the guide nucleic acid share a physical connection; for example, can hybridize when sequence shares sequence similarity.
[0108] "Dissociation constant" or "K d " are used interchangeably and mean the affinity between a ligand "L" and a protein "P"; that is, how tightly the ligand binds to a particular protein. It can be calculated using the formula K d = [L][P] / [LP], where [P], [L], and [LP] represent the molar concentrations of protein, ligand, and complex, respectively.
[0109] The present disclosure provides compositions and methods suitable for editing a target nucleic acid sequence. As used herein, "editing" is used interchangeably with "modifying" and includes, but is not limited to, cleaving, cutting, deleting, knocking in, knocking out, and the like.
[0110] As used herein, "homology directed repair" (HDR) refers to a form of DNA repair that occurs during the repair of double-strand breaks in cells. This method requires nucleotide sequence homology and uses a donor template to repair or knock out a target DNA and allows genetic information to be transferred from a donor (e.g., a donor template) to a target. If the donor template is different from the target DNA sequence and a portion or all of the sequence of the donor template is incorporated into the target DNA at the appropriate genomic locus, homology directed repair can cause a change in the sequence of the target nucleic acid sequence by insertion, deletion, or mutation.
[0111] As used herein, "non-homologous end joining" (NHEJ) refers to the repair of double-strand breaks in DNA by direct ligation of the break ends to each other without the need for a homologous template (as opposed to homology directed repair, which requires homologous sequences to direct repair). NHEJ often causes an indel; a loss (deletion) or insertion of nucleotide sequence near the site of the double-strand break.
[0112] As used herein, "microhomology-mediated end joining" (MMEJ) refers to a mutagenic DSB repair mechanism that invariably combines with a deletion flanking the break site without the need for a homologous template (as opposed to homology directed repair, which requires homologous sequences to direct repair). MMEJ often causes a loss (deletion) of nucleotide sequence near the site of the double-strand break.
[0113] That a polynucleotide or polypeptide (or protein) has a certain percent "sequence similarity" or "sequence identity" to another polynucleotide or polypeptide means that, when aligned, a certain percentage of the bases or amino acids are the same and are in the same relative position when the two sequences are compared. Sequence similarity (sometimes referred to as percent similarity, percent identity, or homology) can be determined in a number of different ways. To determine sequence similarity, sequences can be aligned using methods and computer programs known in the art, including BLAST available at ncbi.nlm.nih.gov / BLAST on the World Wide Web. Percent complementarity between particular stretches of nucleic acid sequence within a nucleic acid can be determined using any convenient method. Exemplary methods include the BLAST program (Basic Local Alignment Search Tool) and the PowerBLAST program (Altschul et al., J. Mol. Biol., 1990, 215, 403-410; Zhang and Madden, Genome Res., 1997, 7, 649-656) or by using the Gap program (Wisconsin Sequence Analysis Package, Version 8 for Unix, Genetics Computer Group, University Research Park, Madison Wis.), for example using default settings, which uses the algorithm of Smith and Waterman (Adv. Appl. Math., 1981, 2, 482-489).
[0114] The terms "polypeptide" and "protein" are used interchangeably herein and refer to a polymeric form of amino acids of any length, which can include coded and non-coded amino acids, chemically or biochemically modified or derivatized amino acids, and polypeptides having modified peptide backbones. The term includes fusion proteins, including but not limited to fusion proteins with heterologous amino acid sequences.
[0115] A "vector" or "expression vector" is a replicon, such as a plasmid, phage, virus, or cosmid, to which another DNA segment (i.e., "insert") can be attached so as to bring about the replication or expression of the attached segment.
[0116] The term "naturally occurring" or "unmodified" or "wild type" as used herein applied to a nucleic acid, polypeptide, cell, or organism means a nucleic acid, polypeptide, cell, or organism found in nature.
[0117] As used herein, "mutation" refers to an insertion, deletion, substitution, duplication, or inversion of one or more amino acids or nucleotides compared to a wild type or reference amino acid sequence or wild type or reference nucleotide sequence.
[0118] As used herein, the term "isolated" is intended to describe a polynucleotide, polypeptide, or cell in an environment different from that in which the polynucleotide, polypeptide, or cell naturally occurs. An isolated genetically modified host cell can exist in a mixed population of genetically modified host cells.
[0119] As used herein, "host cell" indicates a eukaryotic cell, a prokaryotic cell, or a cell from a multicellular organism (e.g., a cell line) cultured in a unicellular entity form, which serves as a recipient for a nucleic acid (e.g., an expression vector), and includes the progeny of the original cell that has been genetically modified by the nucleic acid. It is understood that progeny of a single cell can not necessarily be completely identical to the original parent cell, either upon state, or in genomic or total DNA complement, due to natural, accidental, or deliberate mutation. A "recombinant host cell" (also referred to as a "genetically modified host cell") is a host cell in which a heterologous nucleic acid, such as an expression vector, has been introduced.
[0120] The term "conservative amino acid substitution" refers to the interchangeability of amino acid residues having similar side chains. For example, one group of amino acids that have aliphatic side chains consists of glycine, alanine, valine, leucine, and isoleucine; a group of amino acids that have aliphatic-hydroxyl consists of serine and threonine; a group of amino acids that have amide includes asparagine and glutamine; a group of amino acids that have aromatic consists of phenylalanine, tyrosine, and tryptophan; a group of amino acids that have basic side chains consists of lysine, arginine, and histidine; and a group of amino acids that have sulfuric side chains consists of cysteine and methionine. Exemplary conservative amino acid substitution groups are: valine-leucine-isoleucine, phenylalanine-tyrosine, lysine-arginine, alanine-valine, and asparagine-glutamine.
[0121] As used herein, "treatment" or "treating" are used interchangeably herein, and refer to methods of obtaining beneficial or desired results, including but not limited to therapeutic benefit and / or prophylactic benefit. By therapeutic benefit is meant eradication or amelioration of the underlying disorder or disease being treated. Therapeutic benefit can also be achieved with the eradication or amelioration of one or more of the symptoms associated with the underlying disorder or disease, such that the overall impression of the individual is improved, even though the individual can still be afflicted with the underlying disorder or disease.
[0122] As used herein, the terms "therapeutically effective amount" and "therapeutically effective dose" refer to the amount of a composition, vector cell, etc. that, when administered to an individual, either as a single dose or as repeated doses, is capable of having any detectable, beneficial effect on any symptom, aspect, measured parameter, or characteristic of a disease condition or state. Such effect need not be absolute beneficial. Such effect can be transient.
[0123] As used herein, "administering" means the methods of giving an individual a dose of a composition of the present application.
[0124] As used herein, "individual" is a mammal. Mammals include, but are not limited to, domestic animals, primates, non-human primates, humans, canines, porcine animals (pigs), rabbits, mice, rats, and other rodents.
[0125] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the extent allowable by law as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference.
[0126] I. General Methods
[0127] The practice of the present application employs, unless otherwise indicated, conventional techniques of immunology, biochemistry, chemistry, molecular biology, microbiology, cell biology, genomics, and recombinant DNA, which can be found in standard textbooks of these disciplines such as Molecular Cloning: A Laboratory Manual, 3rded. (Sambrook et al., Harbor Laboratory Press 2001); Short Protocols in Molecular Biology, 4thed. (Ausubel et al., eds., John Wiley & Sons 1999); Protein Methods (Bollag et al., John Wiley & Sons 1996); Nonviral Vectors for Gene Therapy (Wagner et al., eds., Academic Press 1999); Viral Vectors (Kaplift and Loewy, eds., Academic Press 1995); Immunology Methods Manual (I. Lefkovits, ed., Academic Press 1997); and Cell and Tissue Culture: Laboratory Procedures in Biotechnology (Doyle and Griffiths, John Wiley & Sons 1998), the disclosures of which are incorporated herein by reference.
[0128] When a range of values is provided, it is understood that, unless otherwise specified, each intervening value, to the tenth of the unit of the lower limit unless the context clearly dictates otherwise, and any other specified or intervening value or intermediate value within the stated range is encompassed. The upper and lower limits of these smaller ranges can independently be included in the smaller ranges, and are also encompassed, subject to any specifically excluded limit in the stated range. Where the range includes one or both of the limits, ranges excluding either or both of those included limits are also included.
[0129] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. All publications mentioned herein are incorporated herein by reference to disclose and describe the methods and / or materials in connection with which the publications are cited.
[0130] It must be noted that, as used herein and in the appended claims, the singular form "a", "an", and "the" include plural referents unless the context clearly dictates otherwise.
[0131] It is to be understood that certain features of the application which are, for brevity, described in the context of separate embodiments, can also be provided in combination in a single embodiment. Conversely, various features of the application which are, for brevity, described in the context of a single embodiment, can also be provided separately or in any suitable
[0132] II. CasX: gNA system
[0133] In a first aspect, the present application provides CasX:gNA systems comprising a CasX protein and one or more guide nucleic acids (gNA) for modifying or editing a target nucleic acid, including coding and non-coding regions. The terms CasX protein and CasX are used interchangeably herein; the terms CasX variant protein and CasX variant are used interchangeably herein. The CasX protein and gNA of the CasX:gNA systems provided herein can each independently be a reference CasX protein, a CasX variant protein, a reference gNA, a gNA variant, or any combination of a reference CasX protein, a reference gNA, a CasX variant protein, or a gNA variant. The gNA and CasX protein, gNA variant and CasX variant, or any combination thereof, can form a complex and bind via non-covalent interactions, referred to herein as a ribonucleoprotein (RNP) complex. In some embodiments, the use of pre-complexed CasX:gNA confers an advantage in delivering components of the system to a cell or a target nucleic acid for editing of the target nucleic acid. In the RNP, the gNA can provide target specificity to the RNP complex through inclusion of a spacer sequence (targeting sequence) having a nucleotide sequence complementary to a sequence of the target nucleic acid. In the RNP, the CasX protein of the pre-complexed CasX:gNA provides site-specific activity and is directed to a target site within the target nucleic acid sequence to be modified (and further stabilized at the target site) through its association with the gNA. The CasX protein of the RNP complex provides site-specific activity of the complex, such as by CasX protein binding, cleavage, or cleaving the target sequence. Compositions and cells comprising reference CasX proteins, CasX variant proteins, reference gNAs, gNA variants, and any combination of CasX and gNA, as well as delivery modalities comprising CasX:gNA, are provided herein. In other embodiments, the present application provides vectors encoding or comprising CasX:gNA pairs, and optionally donor templates, for production and / or delivery of CasX:gNA systems. Methods of making CasX proteins and gNAs, as well as methods of using CasX and gNAs, including methods of gene editing and methods of treatment, are also provided herein. CasX protein and gNA components of CasX:gNA and features thereof, as well as delivery modalities and methods of using compositions, are more fully described below.
[0134] CasX:gNA system is designed depending on whether it is used to correct a mutation in a target gene or to insert a transgene at a different locus in the genome ("gene knock-in"), or to disrupt expression of an aberrant gene product; e.g., it contains one or more mutations that reduce expression of the gene product or render the protein dysfunctional ("gene knockdown" or "gene knockout"). In some embodiments, the donor template is a single-stranded DNA template or a single-stranded RNA template. In other embodiments, the donor template is a double-stranded DNA template. In some embodiments, the CasX:gNA system used in editing of a target nucleic acid comprises a donor template having all or at least a portion of the open reading frame of a gene in the target nucleic acid for insertion of a corrective, wild-type sequence to correct a defective protein. In other cases, the donor template comprises all or a portion of a wild-type gene for insertion at a different locus in the genome for expression of the gene product. In other cases, a portion of a gene can be inserted upstream ('5) of a mutation in the target nucleic acid, where the donor template gene portion spans to the C-terminus of the gene, causing expression of the gene product upon its insertion into the target nucleic acid. In other embodiments, the donor template can comprise one or more mutations in the coding sequence as compared to the normal, wild-type sequence of the target gene used for insertion to knock out or knock down (more fully described below) a defective target nucleic acid sequence. In other embodiments, the donor template can comprise a regulatory element, intron, or intron-exon junction having a sequence specifically designed to knock down or knock out a defective gene, or in the alternative, to knock in a corrective sequence to permit expression of a functional gene product. In some embodiments, the donor polynucleotide comprises at least about 10, at least about 20, at least about 50, at least about 100, at least about 200, at least about 300, at least about 400, at least about 500, at least about 600, at least about 700, at least about 800, at least about 900, at least about 1000, at least about 10,000, at least about 15,000, at least about 25,000, at least about 50,000, at least about 100,000, or at least about 200,000 nucleotides. With the proviso that there is an extension of DNA sequence having a sufficient number of nucleotides of sufficient homology flanking the cleavage site of the target nucleic acid sequence targeted by the CasX:gNA (i.e., 5' and 3' of the cleavage site) to support homology directed repair (the flanking regions are "homology arms"), such donor templates can be integrated into the target nucleic acid by HDR. In other cases, the donor template can be inserted by non-homologous end joining (NHEJ; which does not require homology arms) or by microhomology-mediated end joining (MMEJ; which requires short homologous regions on the 5' and 3' ends).In some embodiments, the donor template comprises homology arms on the 5' and 3' ends each having at least about 2, at least about 10, at least about 20, at least about 30, at least about 50, at least about 100, at least about 150, at least about 300, at least about 1000, at least about 1500, or more nucleotides of homology to the sequence flanking the intended cleavage site of the target nucleic acid. In some embodiments, the CasX:gNA system utilizes two or more gNAs having targeting sequences that are complementary to overlapping or different regions of the target nucleic acid, such that the defective sequence can be excised by multiple double-stranded breaks or by cleavage at positions flanking the defective sequence, and replaced by insertion of a donor template through HDR. In the foregoing, the gNAs will be designed to contain targeting sequences 5' and 3' of the individual sites or sequences to be excised. Through such appropriate selection of the targeting sequences of the gNAs, defined regions of the target nucleic acid can be edited using the CasX:gNA system described herein.
[0135] III. Guide nucleic acids of the CasX:gNA system
[0136] In other aspects, the present disclosure provides guide nucleic acids (gNAs) for use in the CasX:gNA system and that can be used to edit a target nucleic acid. The present disclosure provides gNAs that are specifically designed to have a targeting sequence (or "spacer") that is complementary to (and thus capable of hybridizing with) a target nucleic acid as a component of the gene editing CasX:gNA system. It is contemplated that in some embodiments, multiple gNAs (e.g., multiple gRNAs) are delivered by the CasX:gNA system to modify different regions of a gene, including regulatory elements, exons, introns, or intron-exon junctions. In some embodiments, the targeting sequence of the gNA is complementary to a sequence comprising one or more single nucleotide polymorphisms (SNPs) of the target nucleic acid. In other embodiments, the targeting sequence of the gNA is complementary to a sequence of an intergenic region. For example, when it is desired to delete a protein-coding gene, a pair of gNAs having targeting sequences directed to different or overlapping regions of the target nucleic acid sequence can be used in order to bind and cleave at two different sites within the gene, which can then be edited by either an indel formation or homology directed repair (HDR), in the case of HDR, which utilizes an insertion of a donor template to replace the deleted sequence to complete the edit.
[0137] a. Reference gNAs and gNA variants
[0138] In some embodiments, a gNA of the application comprises the sequence of a naturally occurring gNA (“reference gNA”). In other cases, a reference gNA of the application can be subjected to one or more mutagenesis methods, such as mutagenesis methods described herein, which can include deep mutational evolution (DME), deep mutational scanning (DMS), error-prone PCR, cassette mutagenesis, random mutagenesis, staggered extension PCR, gene shuffling, or domain swapping, in order to generate one or more gNA variants having enhanced or altered properties relative to the reference gNA. gNA variants also include variants comprising one or more exogenous sequences, such as fused to the 5’ or 3’ end, or inserted internally. The activity of a reference gNA can be used as a benchmark for comparison to the activity of a gNA variant, thereby measuring improvements in function or other properties of the gNA variant. In other embodiments, a reference gNA can be subjected to one or more intentional, targeted mutations to generate a gNA variant, such as a rationally designed variant. As used herein, the terms gNA, gRNA, and gDNA encompass naturally occurring molecules (reference molecules), as well as sequence variants.
[0139] In some embodiments, a gNA is a deoxyribonucleic acid molecule (“gDNA”); in some embodiments, a gNA is a ribonucleic acid molecule (“gRNA”), and in other embodiments, a gNA is a chimera, and comprises both DNA and RNA.
[0140] A gNA of the application comprises two segments; a targeting sequence and a protein binding segment (which constitutes a scaffold discussed herein). The targeting segment of a gNA includes a nucleotide sequence (interchangeably referred to herein as a guide sequence, a spacer, a targeting sequence, or a targeting region) that is complementary to (and thus hybridizes with) a specific sequence (target site) within a target nucleic acid sequence, such as a target ssRNA, a target ssDNA, a complementary strand of a double stranded target DNA, and the like, described more fully below.
[0141] The targeting sequence of a gNA is capable of binding to a target nucleic acid sequence, including coding sequences, complementary sequences of coding sequences, non-coding sequences, and to regulatory elements. The protein binding segment (or “protein binding sequence”) interacts with (e.g., binds to) a CasX protein. The protein binding segment is alternatively referred to herein as a “scaffold.” In some embodiments, the targeting sequence and the scaffold each comprise a complementary stretch of nucleotides that hybridize to each other to form a double-stranded duplex (e.g., a dsRNA duplex for a gRNA). Site-specific binding and / or cleavage by CasX:gNA of a target nucleic acid sequence, such as genomic DNA, can occur at one or more locations of the target nucleic acid, as determined by base-pairing complementarity between the targeting sequence of the gNA and the target nucleic acid sequence.
[0142] The gNA provides target specificity to the complex by having a nucleotide sequence that is complementary to a target sequence of a target nucleic acid. The CasX of the complex provides the site-specific activity of the complex, such as by the CasX nuclease binding to, cleaving, or cutting the target sequence of the target nucleic acid, and / or in the case of a CasX-containing fusion protein, the activity provided by the fusion partner (described below). In some embodiments, the present application provides a gene editing pair of a CasX and a gNA of any of the embodiments described herein that are capable of being bound together and thus“pre-complexed” as an RNP prior to their use for gene editing. Using pre-complexed RNPs confers an advantage in delivering the system components to a cell or target nucleic acid sequence for editing the target nucleic acid sequence. The CasX protein of the RNP provides the site-specific activity, directed to a target site within the target nucleic acid sequence (e.g., stabilized at the target site) by its association with a guide RNA comprising a targeting sequence.
[0143] In some embodiments, where the gNA is a gRNA, the term“targeter” or“targeter RNA” is used herein to refer to the crRNA-like molecule (crRNA:“CRISPR RNA”) of a CasX dual guide RNA (dgRNA). In a single guide RNA (sgRNA), the“activator” and“targeter” are linked together, such as by intervening nucleotides). Thus, for example, a guide RNA (dgRNA or sgRNA) comprises a guide sequence and a crRNA duplex-forming segment, which can also be referred to as a crRNA repeat sequence. The targeter can be modified by the user to hybridize to a desired target nucleic acid sequence due to the targeter sequence of the guide sequence hybridizing to a particular target nucleic acid sequence. In some embodiments, the sequence of the targeter can be generally a non-naturally occurring sequence. The targeter and activator each have a duplex-forming segment, where the duplex-forming segment of the targeter and the duplex-forming segment of the activator have complementarity to each other and hybridize to each other to form a double-stranded duplex (dsRNA duplex for gRNA). In some embodiments, the targeter comprises a stretch of nucleotides that form one half of the dsRNA duplex of the protein-binding segment of the CasX guide RNA and the guide sequence. The corresponding tracrRNA-like molecule (activator“trans-acting CRISPR RNA”) also comprises a duplex-forming segment of nucleotides that form the other half of the dsRNA duplex of the protein-binding segment of the CasX guide RNA. In some cases, the activator comprises one or more stem loops that can interact with the CasX protein. Thus, the targeter and activator in corresponding pairs hybridize to form a CasX dual guide NA, referred to herein as a“dual guide NA,”“dgNA,”“dual-molecule guide NA,” or“two-molecule guide NA.”
[0144] In some embodiments, the activator and the targeting subunit of a reference gNA are covalently linked to each other and comprise a single molecule, referred to herein as a "single molecule guide NA," "one molecule guide NA," "single guide NA," "single guide RNA," "single molecule guide RNA," "one molecule guide RNA," "single guide DNA," "single molecule DNA," or "one molecule guide DNA" ("sgNA," "sgRNA," or "sgDNA"). In some embodiments, a sgNA comprises an "activator" or a "targeting subunit" and thus can be "activator-RNA" and "targeting subunit-RNA," respectively.
[0145] A reference gRNA of the present application comprises four distinct regions or domains: an RNA triplex, a scaffold stem, an extension stem, and a targeting sequence (which is specific for a target nucleic acid. The RNA triplex, scaffold stem, and extension stem together are referred to as the "scaffold" of the reference gNA, from which additional gNA variants are generated.
[0146] b. RNA triplex
[0147] In some embodiments of the guide NAs provided herein, the gNA comprises an RNA triplex, and the RNA triplex comprises a UUU X (~4-15) - UUU stem loop (SEQ ID NO: 241) ending in AAAG after 2 intervening stem loops (scaffold stem loop and extension stem loop), forming a pseudoknot that can also extend through the triplex into a duplex pseudoknot. The UU-UUU-AAA sequence of the triplex is formed as a junction between the targeting sequence, the scaffold stem, and the extension stem. In exemplary gRNAs, the UUU-loop-UUU region is encoded first, followed by the scaffold stem loop, and then the extension stem loop, which is connected by a four loop, and then the closing AAAG triplex, followed by the targeting sequence.
[0148] c. Scaffold stem loop
[0149] In some embodiments of the gNA of the present invention, a scaffold stem loop follows the triple-helix region. The scaffold stem loop is the gNA region that binds to a CasX protein (such as a reference or CasX variant). In some embodiments, the scaffold stem loop is a relatively short and stable stem loop, increasing the overall stability of the gNA. In some cases, the scaffold stem loop is intolerant to many variations and some form of RNA bubble is required. In some embodiments, the scaffold stem is required for gNA function. Although the scaffold stem of the gNA may serve as an important stem loop similar to the linker stem of Cas9, in some embodiments it has a desired protrusion (RNA bubble) that differs from many other stem loops found in the CRISPR / Cas system. In some embodiments, the presence of this protrusion is conserved across the gNA interacting with different CasX proteins. An exemplary sequence of the gNA scaffold stem loop sequence comprises the sequence CCAGCGACUAUGUCGUAUGG (SEQ ID NO: 242). In other embodiments, the present invention provides gNA variants in which the scaffold stem loop is replaced by an RNA stem loop sequence from a heterologous RNA source having proximal 5' and 3' ends, such as, but not limited to, stem loop sequences selected from MS2, Qβ, U1 hairpin II, Uvsx, or PP7 stem loops. In some cases, the heterologous RNA stem loop of gNA can bind proteins, RNA structures, DNA sequences, or small molecules.
[0150] d. Extended stem ring
[0151] In some embodiments of the gNA of the application, the scaffold stem loop is followed by an extended stem loop. In some embodiments, the extended stem comprises a synthetic tracr and crRNA fusion that is largely unbound by the CasX protein. In some embodiments, the extended stem loop can be highly exapted. In some embodiments, a single guide gRNA is made by a GAAA tetraloop linker or a GAGAAA linker between the tracr and crRNA in the extended stem loop. In some cases, the targeting and activator of the sgNA are linked to each other by an intervening nucleotide, and the length of the linker can be 3 to 20 nucleotides. In some embodiments of the sgNA of the application, the extended stem is a large 32-bp loop outside of the CasX protein in the ribonucleoprotein complex. An exemplary sequence of the extended stem loop sequence of the sgNA comprises the sequence GCGCUUAUUUAUCGGAGAGAAAUCCGAUAAAU AAGAAGC (SEQ ID NO: 15). In some embodiments, the extended stem loop comprises a GAGAAA spacer sequence. In some embodiments, the application provides gNA variants in which the extended stem loop is replaced with a RNA stem loop sequence from a heterologous RNA source having a proximal 5' and 3' end, such as, but not limited to, a stem loop sequence selected from MS2, Qp, U1 hairpin II, Uvsx, or PP7 stem loop. In such cases, the heterologous RNA stem loop increases the stability of the gNA. In other embodiments, the application provides gNA variants having an extended stem loop region comprising at least 10, at least 100, at least 500, at least 1000, or at least 10,000 nucleotides.
[0152] e. Targeting sequence
[0153] In some embodiments of the gNA of the application, the extended stem loop is followed by a region that forms part of a triplex body, and then by a targeting sequence (or "spacer"). The targeting sequence can be designed to target the CasX ribonucleoprotein complex as a whole to a specific region of a nucleic acid sequence of interest. Thus, when any of the PAM sequences TTC, ATC, GTC, or CTC is located 1 nucleotide 5' to the non-target strand sequence that is complementary to the sequence of interest, the gNA targeting sequence of the gNA of the application has sequence complementarity to, and thus can hybridize with, a portion of a nucleic acid in a nucleic acid in a eukaryotic cell (e.g., a eukaryotic chromosome, a chromosomal sequence, a eukaryotic RNA, etc.), as a component of an RNP.
[0154] In some embodiments, the present application provides a gNA, wherein the targeting sequence of the gNA is complementary to a target nucleic acid sequence comprising one or more mutations compared to the sequence of a wild-type gene for the purpose of editing the sequence comprising the mutation with the CasX:gNA system of the present application. In some embodiments, the targeting sequence of the gNA is designed to be specific to an exon of a gene of the target nucleic acid. In other embodiments, the targeting sequence of the gNA is designed to be specific to an intron of a gene of the target nucleic acid. In other embodiments, the targeting sequence of the gNA is designed to be specific to an intron-exon junction of a gene of the target nucleic acid. In other embodiments, the targeting sequence of the gNA is designed to be specific to a regulatory element of a gene of the target nucleic acid. In some embodiments, the targeting sequence of the gNA is designed to be complementary to a sequence comprising one or more single nucleotide polymorphisms (SNPs) in a gene of the target nucleic acid. SNPs within coding sequences or within non-coding sequences are within the scope of the present application. In other embodiments, the targeting sequence of the gNA is designed to be complementary to a sequence of an intergenic region of a gene of the target nucleic acid.
[0155] In some embodiments, the targeting sequence of the gNA is designed to be specific to a regulatory element that regulates expression of a gene product of the target nucleic acid. Such regulatory elements include, but are not limited to, promoter regions, enhancer regions, intergenic regions, 5' untranslated regions (5' UTRs), 3' untranslated regions (3' UTRs), conserved elements, and regions containing cis-regulatory elements. A promoter region is intended to encompass nucleotides within 5 kb of the start of the coding sequence, or in the case of a gene enhancer element or a conserved element, can be thousands of bp, hundreds of thousands of bp, or even millions of bp away from the coding sequence of a gene of the target nucleic acid. In some embodiments of the foregoing, the target is one in which the encoded gene of the target is intended to be knocked out or knocked down such that the encoded protein comprising the mutation is not expressed in the cell or is expressed at a lower amount in the cell.
[0156] In some embodiments, the targeting sequence of the gNA has 14 to 35 contiguous nucleotides. In some embodiments, the targeting sequence has 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, or 35 contiguous nucleotides. In some embodiments, the targeting sequence of the gNA consists of 20 contiguous nucleotides. In some embodiments, the targeting sequence consists of 19 contiguous nucleotides. In some embodiments, the targeting sequence consists of 18 contiguous nucleotides. In some embodiments, the targeting sequence consists of 17 contiguous nucleotides. In some embodiments, the targeting sequence consists of 16 contiguous nucleotides. In some embodiments, the targeting sequence consists of 15 contiguous nucleotides. In some embodiments, the targeting sequence has 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, or 35 contiguous nucleotides, and the targeting sequence can comprise 0 to 5, 0 to 4, 0 to 3, or 0 to 2 mismatches relative to the target nucleic acid sequence and retain sufficient binding specificity such that an RNP containing the gNA comprising the targeting sequence can form a complementary bond with the target nucleic acid.
[0157] In some embodiments, the CasX:gNA system comprises a first gNA and further comprises a second (and optionally a third, fourth, fifth, or greater) gNA, wherein the second gNA or additional gNA has a targeting sequence that is different from or overlaps with a portion of the targeting sequence of the first gNA that is complementary to the target nucleic acid sequence, such that multiple points in the target nucleic acid are targeted and multiple breaks in the target nucleic acid are introduced, e.g., by CasX. It will be appreciated that in such cases, the second or additional gNA is complexed with an additional copy of the CasX protein. By selecting the targeting sequence of the gNA, a defined region of the target nucleic acid sequence harboring a mutation can be modified or edited using the CasX:gNA system described herein, including to facilitate donor template insertion.
[0158] f. gNA scaffolds
[0159] In addition to the targeting sequence region, the remaining region of the gNA is referred to herein as the scaffold. In some embodiments, the gNA scaffold is derived from a naturally occurring sequence, described below as a reference gNA. In other embodiments, the gNA scaffold is a variant of a reference gNA, wherein mutations, insertions, deletions, or domain substitutions are introduced to impart desired properties to the gNA.
[0160] In some embodiments, the reference gRNA comprises a sequence isolated or derived from Deltaproteobacteria. In some embodiments, the sequence is a CasX tracrRNA sequence. Exemplary CasX reference tracrRNA sequences isolated or derived from Deltaproteobacteria can include: ACAUCUGGCGCGUUUAUUCCAUUACUUUGGAGCCAG UCCCAGCGACUAUGUCGUAUGGACGAAGCGCUUAUUUAUCGGAGA (SEQ ID NO: 6) and ACAUCUGGCGCGUUUAUUCCAUUACUUUGGAGCCAGUCCCAGCGACUAUGUCGUA UGGACGAAGCGCUUAUUUAUCGG (SEQ ID NO: 7). Exemplary crRNA sequences isolated or derived from Deltaproteobacteria can comprise the sequence CCGAUAAGUAAAACGCAUCAAAG (SEQ ID NO: 243). In some embodiments, the reference gNA comprises a sequence that is at least 60% identical, at least 65% identical, at least 70% identical, at least 75% identical, at least 80% identical, at least 81% identical, at least 82% identical, at least 83% identical, at least 84% identical, at least 85% identical, at least 86% identical, at least 86% identical, at least 87% identical, at least 88% identical, at least 89% identical, at least 89% identical, at least 90% identical, at least 91% identical, at least 92% identical, at least 93% identical, at least 94% identical, at least 95% identical, at least 96% identical, at least 97% identical, at least 98% identical, at least 99% identical, at least 99.5% identical, or 100% identical to a sequence isolated or derived from Deltaproteobacteria.
[0161] In some embodiments, the reference guide RNA comprises a sequence isolated or derived from the phylum Planctomycetes. In some embodiments, the sequence is a CasX tracrRNA sequence. Exemplary reference tracrRNA sequences isolated or derived from the phylum Planctomycetes can include: UACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUGGGUAAAGCGCUUAUUUAUCGGAGA (SEQ ID NO: 8) and UACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUGGGUAAAGCGCUUAUUUAUCGG (SEQ ID NO: 9). Exemplary crRNA sequences isolated or derived from the phylum Planctomycetes can comprise the sequence UCUCCGAUAAAUAAGAAGCAUCAAAG (SEQ ID NO: 244). In some embodiments, the reference gNA comprises a sequence that is at least 60% identical, at least 65% identical, at least 70% identical, at least 75% identical, at least 80% identical, at least 81% identical, at least 82% identical, at least 83% identical, at least 84% identical, at least 85% identical, at least 86% identical, at least 86% identical, at least 87% identical, at least 88% identical, at least 89% identical, at least 89% identical, at least 90% identical, at least 91% identical, at least 92% identical, at least 93% identical, at least 94% identical, at least 95% identical, at least 96% identical, at least 97% identical, at least 98% identical, at least 99% identical, at least 99.5% identical, or 100% identical to a sequence isolated or derived from the phylum Planctomycetes.
[0162] In some embodiments, the reference gNA comprises a sequence isolated or derived from Candidatus Sungbacteria. In some embodiments, the sequence is a CasX tracrRNA sequence. Exemplary CasX reference tracrRNA sequences isolated or derived from Candidatus Sungbacteria can comprise the following sequences: GUUUACACACUCCCUCUCAUAGGGU (SEQ ID NO: 10), GUUUACACACUCCCUCUCAUGAGGU (SEQ ID NO: 11), UUUUACAUACCCCCUCUCAUGGGAU (SEQ ID NO: 12), and GUUUACACACUCCCUCUCAUGGGGG (SEQ ID NO: 13). In some embodiments, the reference guide RNA comprises a sequence that is at least 60% identical, at least 65% identical, at least 70% identical, at least 75% identical, at least 80% identical, at least 81% identical, at least 82% identical, at least 83% identical, at least 84% identical, at least 85% identical, at least 86% identical, at least 86% identical, at least 87% identical, at least 88% identical, at least 89% identical, at least 89% identical, at least 90% identical, at least 91% identical, at least 92% identical, at least 93% identical, at least 94% identical, at least 95% identical, at least 96% identical, at least 97% identical, at least 98% identical, at least 99% identical, at least 99.5% identical, or 100% identical to a sequence isolated or derived from Candidatus Sungbacteria.
[0163] Table 1 provides sequences for reference gRNA tracr, cr, and scaffold sequences. In some embodiments, the present application provides gNA sequences, wherein the gNA has a scaffold comprising a sequence having at least one nucleotide modification relative to a reference gNA sequence having the sequence of any one of SEQ ID NOs: 4-16 of Table 1. It will be appreciated that in those embodiments in which the vector comprises a DNA coding sequence for a gNA, or in which the gNA is a gDNA or a chimera of RNA and DNA, a thymine (T) base can be substituted for a uracil (U) base in any of the gNA sequence embodiments described herein.
[0164] Table 1. Reference gRNA tracr, cr, and scaffold sequences
[0165]
[0166] g.gNA variants
[0167] In another aspect, the present application relates to guide nucleic acid variants (alternatively referred to herein as "gNA variants" or "gRNA variants") comprising one or more modifications relative to a reference gRNA scaffold. As used herein, "scaffold" refers to all portions of a gNA required for gNA function other than the spacer sequence.
[0168] In some embodiments, the gNA variants comprise one or more nucleotide substitutions, insertions, deletions or exchanges or replacement regions relative to a reference gRNA sequence of the present application. In some embodiments, the mutations can occur in any region of the reference gRNA scaffold to produce a gNA variant. In some embodiments, the scaffold of the gNA variant sequence has at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, or at least 70%, at least 80%, at least 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to the sequence of SEQ ID NO: 4 or SEQ ID NO: 5.
[0169] In some embodiments, the gNA variant comprises one or more nucleotide changes within one or more regions of the reference gRNA scaffold that improve a property of the reference gRNA. Exemplary regions include the RNA triple helix, the pseudoknot, the scaffold stem loop, and the extended stem loop. In some cases, the variant scaffold stem further comprises a bubble. In other cases, the variant scaffold further comprises a triple helix loop region. In other cases, the variant scaffold further comprises a 5' unstructured region. In some embodiments, the gNA variant scaffold comprises a scaffold stem loop having at least 60% sequence identity, at least 70% sequence identity, at least 80% sequence identity, at least 90% sequence identity, at least 95% sequence identity, or at least 99% sequence identity to SEQ ID NO: 14. In some embodiments, the gNA variant scaffold comprises a scaffold stem loop having at least 60% sequence identity to SEQ ID NO: 14. In other embodiments, the gNA variant comprises a scaffold stem loop having the sequence CCAGCGACUAUGUCGUAGUGG (SEQ ID NO: 245). In other embodiments, the present application provides a gNA scaffold comprising a C18G substitution, a G55 insertion, a U1 deletion, and a modified extended stem loop relative to SEQ ID NO: 5, wherein the original 6 nt loop and 13 loop proximal base pairs (total of 32 nucleotides) are replaced with a Uvsx hairpin (4 nt loop and 5 loop proximal base pairs; total of 14 nucleotides), and the loop distal base pairs of the extended stem are converted to a fully base paired stem contiguous with the new Uvsx hairpin by deletion of A99 and substitution of G65U. In the foregoing embodiment, the gNA scaffold comprises the sequence ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGUGGGUAAAGCUCCCUCUUCGGAGGGAGCAUCAAAG (SEQ ID NO: 2238).
[0170] All gNA variants having one or more improved features, or adding one or more new functions, are contemplated as within the scope of the application when compared to the reference gRNAs described herein. A representative example of such a gNA variant is Guide174 (SEQ ID NO: 2238), the design of which is described in the Examples. In some embodiments, a gNA variant adds a new function to the RNP comprising the gNA variant. In some embodiments, a gNA variant has an improved feature selected from the group consisting of: improved stability; improved solubility; improved gNA transcription; improved nuclease activity resistance; increased gNA folding rate; reduced byproduct formation during folding; increased productive folding; improved binding affinity to CasX protein; improved binding affinity to target DNA when complexed with CasX protein; improved gene editing when complexed with CasX protein; improved editing specificity when complexed with CasX protein; and improved ability to utilize a larger range of one or more PAM sequences, including ATC, CTC, GTC, or TTC, in editing of a target DNA when complexed with CasX protein, and any combination thereof. In some cases, one or more of the improved features of a gNA variant is improved by at least about 1.1 to about 100,000 fold relative to the reference gNA of SEQ ID NO: 4 SEQ ID NO: 5. In other cases, one or more improved features of a gNA variant is improved by at least about 1.1, at least about 10, at least about 100, at least about 1000, at least about 10,000, at least about 100,000 fold or more relative to the reference gNA of SEQ ID NO: 4 or SEQ ID NO: 5.In other cases, one or more of the improved features of the gNA variant is improved by about 1.1- to 100,00-fold, about 1.1- to 10,00-fold, about 1.1- to 1,000-fold, about 1.1- to 500-fold, about 1.1- to 100-fold, about 1.1- to 50-fold, about 1.1- to 20-fold, about 10- to 100,00-fold, about 10- to 10,00-fold, about 10- to 1,000-fold, about 10- to 500-fold, about 10- to 100-fold, about 10- to 50-fold, about 10- to 20-fold, about 2- to 70-fold, about 2- to 50-fold, about 2- to 30-fold, about 2- to 20-fold, about 2- to 10-fold, about 5- to 50-fold, about 5- to 30-fold, about 5- to 10-fold, about 100- to 100,00-fold, about 100- to 10,00-fold, about 100- to 1,000-fold, about 100- to 500-fold, about 500- to 100,00-fold, about 500- to 10,00-fold, about 500- to 1,000-fold, about 500- to 750-fold, about 1,000- to 100,00-fold, about 10,000- to 100,00-fold, about 20- to 500-fold, about 20- to 250-fold, about 20- to 200-fold, about 20- to 100-fold, about 20- to 50-fold, about 50- to 10,000-fold, about 50- to 1,000-fold, about 50- to 500-fold, about 50- to 200-fold, or about 50- to 100-fold relative to the reference gNA of SEQ ID NO: 4 or SEQ ID NO: 5. In other cases, one or more of the improved features of the gNA variant is improved by about 1.1-fold, 1.2-fold, 1.3-fold, 1.4-fold, 1.5-fold, 1.6-fold, 1.7-fold, 1.8-fold, 1.9-fold, 2-fold, 3-fold, 4-fold, 5-fold, 6-fold, 7-fold, 8-fold, 9-fold, 10-fold, 11-fold, 12-fold, 13-fold, 14-fold, 15-fold, 16-fold, 17-fold, 18-fold, 19-fold, 20-fold, 25-fold, 30-fold, 40-fold, 45-fold, 50-fold, 55-fold, 60-fold, 70-fold, 80-fold, 90-fold, 100-fold, 110-fold, 120-fold, 130-fold, 140-fold, 150-fold, 160-fold, 170-fold, 180-fold, 190-fold, 200-fold, 210-fold, 220-fold, 230-fold, 240-fold, 250-fold, 260-fold, 270-fold, 280-fold, 290-fold, 300-fold, 310-fold, 320-fold, 330-fold, 340-fold, 350-fold, 360-fold, 370-fold, 380-fold, 390-fold, 400-fold, 425-fold, 450-fold, 475-fold, or 500-fold relative to the reference gNA of SEQ ID NO: 4 or SEQ ID NO: 5.
[0171] In some embodiments, gNA variants can be generated by subjecting a reference gNA to one or more mutagenesis methods, such as mutagenesis methods described below, which can include deep mutational evolution (DME), deep mutational scanning (DMS), error-prone PCR, cassette mutagenesis, random mutagenesis, staggered extension PCR, gene shuffling, or domain swapping, in order to generate gNA variants of the application. The activity of the reference gNA can be used as a benchmark for comparison with the activity of the gNA variants, thereby measuring improvements in the function of the gNA variants. In other embodiments, a reference gNA can be subjected to one or more intentional targeted mutations, substitutions, or domain swaps to generate gNA variants, such as rationally designed variants. Exemplary gNA variants generated by such methods are described in the Examples, and representative sequences of gNA scaffolds are presented in Table 2.
[0172] In some embodiments, a gNA variant comprises one or more modifications compared to a reference guide nucleic acid scaffold sequence, wherein the one or more modifications are selected from: at least one nucleotide substitution in a reference gNA region; at least one nucleotide deletion in a reference gNA region; at least one nucleotide insertion in a reference gNA region; a substitution of all or a portion of a reference gNA region; a deletion of all or a portion of a reference gNA region; or any combination of the foregoing. In some cases, the modification is a substitution of 1 to 15 contiguous or non-contiguous nucleotides in the reference gNA in one or more regions. In other cases, the modification is a deletion of 1 to 10 contiguous or non-contiguous nucleotides in the reference gNA in one or more regions. In other cases, the modification is an insertion of 1 to 10 contiguous or non-contiguous nucleotides into the reference gNA in one or more regions. In other cases, the modification is a substitution of a scaffold stem loop or an extended stem loop with an RNA stem loop sequence from a heterologous RNA source having proximal 5’ and 3’ ends. In some cases, a gNA variant of the application comprises two or more modifications in one region relative to a reference gRNA. In other cases, a gNA variant of the application comprises modifications in two or more regions. In other cases, a gNA variant comprises any combination of the foregoing modifications described in this paragraph. In some embodiments, exemplary modifications of gNAs of the application include the modifications of Table 24.
[0173] In some embodiments, a 5’G is added to gNA variant sequences for in vivo expression relative to a reference gRNA, as transcription from the U6 promoter is more efficient and more consistent for start sites when the +1 nucleotide is a G. In other embodiments, two 5’G’s are added to generate gNA variant sequences for in vitro transcription to improve production efficiency, as T7 polymerase strongly prefers a G in the +1 position and a purine in the +2 position. In some cases, a 5’G base is added to reference scaffolds of Table 1. In other cases, a 5’G base is added to variant scaffolds of Table 2.
[0174] Table 2 provides exemplary gNA variant scaffold sequences of the application. In Table 2, (-) indicates a deletion at the indicated position relative to the reference sequence of SEQ ID NO: 5, (+) indicates an insertion of the indicated base at the indicated position relative to SEQ ID NO: 5, (:) indicates a range of bases at the indicated start:stop coordinates of a deletion or substitution relative to SEQ ID NO: 5, and multiple insertions, deletions, or substitutions are separated by commas; e.g., A14C, T17G. In some embodiments, a gNA variant scaffold comprises any one of the sequences listed in Table 2, SEQ ID NOs: 2101-2280, or a sequence having at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% sequence identity thereto. It will be appreciated that in those embodiments in which the vector comprises a coding DNA sequence for a gNA, or in which the gNA is a gDNA or a chimera of RNA and DNA, a thymine (T) base can be substituted for a uracil (U) base in any of the gNA sequence embodiments described herein.
[0175] Table 2. Exemplary gNA variant scaffold sequences
[0176]
[0177]
[0178]
[0179]
[0180]
[0181]
[0182]
[0183]
[0184]
[0185]
[0186]
[0187]
[0188]
[0189]
[0190]
[0191] In some embodiments, the gNA variant contains a tracrRNA stem loop comprising the sequence -UUU-N4-25-UUU- (SEQ ID NO: 240). For example, the gNA variant comprises a scaffold stem loop or a substitute thereof, flanked by two triplet U motifs that promote triplex formation. In some embodiments, the scaffold stem loop or a substitute thereof comprises at least 4 nucleotides, at least 5 nucleotides, at least 6 nucleotides, at least 7 nucleotides, at least 7 nucleotides, at least 8 nucleotides, at least 9 nucleotides, at least 10 nucleotides, at least 11 nucleotides, at least 12 nucleotides, at least 13 nucleotides, at least 14 nucleotides, at least 15 nucleotides, at least 16 nucleotides, at least 17 nucleotides, at least 18 nucleotides, at least 19 nucleotides, at least 20 nucleotides, at least 21 nucleotides, at least 22 nucleotides, at least 23 nucleotides, at least 24 nucleotides, or at least 25 nucleotides.
[0192] In some embodiments, the gNA variant comprises a crRNA sequence having -AAAG- at the position 5' of the spacer. In some embodiments, the -AAAG- sequence is immediately 5' of the spacer.
[0193] In some embodiments, modifying at least one nucleotide of a reference gNA to produce a gNA variant comprises at least one nucleotide deletion in the CasX variant gNA relative to the reference gRNA. In some embodiments, the gNA variant comprises a deletion of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 contiguous or non-contiguous nucleotides relative to the reference gNA. In some embodiments, the at least one deletion comprises a deletion of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 or more contiguous nucleotides relative to the reference gNA. In some embodiments, the gNA variant comprises a deletion of 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 or more nucleotides relative to the reference gNA, and the deletion is not in contiguous nucleotides. In those embodiments in which there are two or more non-contiguous deletions in the gNA variant relative to the reference gRNA, any deletion length and any combination of deletion lengths as described herein are encompassed within the scope of the disclosure. For example, in some embodiments, the gNA variant can comprise a first deletion of one nucleotide, and a second deletion of two nucleotides, and the two deletions are not contiguous. In some embodiments, the gNA variant comprises at least two deletions in different regions of the reference gRNA. In some embodiments, the gNA variant comprises at least two deletions in the same region of the reference gRNA. For example, the region can be the extended stem loop, the scaffold stem loop, the scaffold stem bubble, the triple helix loop, the pseudoknot, the triplex body, or the 5’ end of the gNA variant. Deletion of any nucleotide in the reference gRNA is encompassed within the scope of the disclosure.
[0194] In some embodiments, the at least one nucleotide modification to the reference gRNA to generate a gNA variant comprises at least one nucleotide insertion. In some embodiments, the gNA variant comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 contiguous or non-contiguous nucleotides inserted relative to the reference gRNA. In some embodiments, the at least one nucleotide insertion comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 or more contiguous nucleotides inserted relative to the reference gRNA. In some embodiments, the gNA variant comprises 2 or more insertions relative to the reference gRNA, and the insertions are not contiguous. In those embodiments where there are two or more non-contiguous insertions in the gNA variant relative to the reference gRNA, any insertion length and any combination of insertion lengths as described herein are contemplated within the scope of the disclosure. For example, in some embodiments, the gNA variant can comprise a first insertion of one nucleotide, and a second insertion of two nucleotides, and the two insertions are not contiguous. In some embodiments, the gNA variant comprises at least two insertions in different regions of the reference gRNA. In some embodiments, the gNA variant comprises at least two insertions in the same region of the reference gRNA. For example, the region can be the extended stem loop, the scaffold stem loop, the scaffold stem bubble, the triple helix loop, the pseudoknot, the triplex body, or the 5' end of the gNA variant. Inserting any A, G, C, U (or T, in the corresponding DNA), or combinations thereof at any position in the reference gRNA is contemplated within the scope of the disclosure.
[0195] In some embodiments, the at least one nucleotide modification to the reference gRNA to generate a gNA variant comprises at least one nucleic acid substitution. In some embodiments, the gNA variant comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 or more contiguous or non-contiguous substituted nucleotides relative to the reference gRNA. In some embodiments, the gNA variant comprises 1-4 nucleotide substitutions relative to the reference gRNA. In some embodiments, the at least one substitution comprises substitution of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 or more contiguous nucleotides relative to the reference gRNA. In some embodiments, the gNA variant comprises 2 or more substitutions relative to the reference gRNA, and the substitutions are non-contiguous. In those embodiments where there are two or more non-contiguous substitutions in the gNA variant relative to the reference gRNA, any substituted nucleotide length and any combination of substituted nucleotide lengths as described herein are contemplated within the scope of the application. For example, in some embodiments, the gNA variant can comprise a first substitution of one nucleotide, and a second substitution of two nucleotides, and the two substitutions are non-contiguous. In some embodiments, the gNA variant comprises at least two substitutions in different regions of the reference gRNA. In some embodiments, the gNA variant comprises at least two substitutions in the same region of the reference gRNA. For example, the region can be a triple-helix body, an extended stem loop, a scaffold stem loop, a scaffold stem bubble, a triple-helix loop, a pseudoknot, a triple-helix body, or a 5' end of the gNA variant. Substitution of any A, G, C, U (or T, in the corresponding DNA), or combinations thereof at any position in the reference gRNA are contemplated within the scope of the application.
[0196] Any of the substitutions, insertions, and deletions described herein can be combined to generate a gNA variant of the application. For example, a gNA variant can comprise at least one substitution and at least one deletion relative to a reference gRNA, at least one substitution and at least one insertion relative to a reference gRNA, at least one insertion and at least one deletion relative to a reference gRNA, or at least one substitution, one insertion, and one deletion relative to a reference gRNA.
[0197] In some embodiments, the gNA variant comprises a scaffold region that is at least 20% identical, at least 30% identical, at least 40% identical, at least 50% identical, at least 60% identical, at least 65% identical, at least 70% identical, at least 75% identical, at least 80% identical, at least 85% identical, at least 90% identical, at least 91% identical, at least 92% identical, at least 93% identical, at least 94% identical, at least 95% identical, at least 96% identical, at least 97% identical, at least 98% identical, or at least 99% identical to any one of SEQ ID NOs: 4-16. In some embodiments, the gNA variant comprises a scaffold region that is at least 60% homologous (or identical) to any one of SEQ ID NOs: 4-16.
[0198] In some embodiments, the gNA variant comprises a tracr stem loop that is at least 60% identical, at least 65% identical, at least 70% identical, at least 75% identical, at least 80% identical, at least 85% identical, at least 90% identical, at least 91% identical, at least 92% identical, at least 93% identical, at least 94% identical, at least 95% identical, at least 96% identical, at least 97% identical, at least 98% identical, or at least 99% identical to SEQ ID NO: 14. In some embodiments, the gNA variant comprises a tracr stem loop that is at least 60% homologous (or identical) to SEQ ID NO: 14.
[0199] In some embodiments, the gNA variant comprises an extended stem loop that is at least 60% identical, at least 65% identical, at least 70% identical, at least 75% identical, at least 80% identical, at least 85% identical, at least 90% identical, at least 91% identical, at least 92% identical, at least 93% identical, at least 94% identical, at least 95% identical, at least 96% identical, at least 97% identical, at least 98% identical, or at least 99% identical to SEQ ID NO: 15. In some embodiments, the gNA variant comprises an extended stem loop that is at least 60% homologous (or identical) to SEQ ID NO: 15.
[0200] In some embodiments, the gNA variant comprises the sequence of any one of SEQ ID NOs: 412-3295. In some embodiments, the gNA variant comprises the sequence of any one of SEQ ID NOs: 2236, 2237, 2238, 2241, 2244, 2248, 2249, or 2259-2280. In some embodiments, the gNA variant comprises the sequence of any one of SEQ ID NOs: 2236, 2237, 2238, 2241, 2244, 2248, 2249, or 2259-2280.
[0201] In some embodiments, the gNA variant comprises an exogenous extended stem loop, wherein such differences from a reference gNA are described below. In some embodiments, the exogenous extended stem loop has little or no identity to a reference stem loop region disclosed herein (e.g., SEQ ID NO: 15). In some embodiments, the exogenous stem loop is at least 10 bp, at least 20 bp, at least 30 bp, at least 40 bp, at least 50 bp, at least 60 bp, at least 70 bp, at least 80 bp, at least 90 bp, at least 100 bp, at least 200 bp, at least 300 bp, at least 400 bp, at least 500 bp, at least 600 bp, at least 700 bp, at least 800 bp, at least 900 bp, at least 1,000 bp, at least 2,000 bp, at least 3,000 bp, at least 4,000 bp, at least 5,000 bp, at least 6,000 bp, at least 7,000 bp, at least 8,000 bp, at least 9,000 bp, at least 10,000 bp, at least 12,000 bp, at least 15,000 bp, or at least 20,000 bp. In some embodiments, the gNA variant contains an extended stem loop region comprising at least 10, at least 100, at least 500, at least 1000, or at least 10,000 nucleotides. In some embodiments, the exogenous stem loop increases the stability of the gNA. In some embodiments, the exogenous RNA stem loop is capable of binding a protein, an RNA structure, a DNA sequence, or a small molecule.In some embodiments, the exogenous stem loop region comprises an RNA stem loop or hairpin, e.g., a thermostable RNA such as MS2 (ACAUGAGGAUUACCCAUGU; SEQ ID NO: 4278), Qbeta (UGCAUGUCUAAGACAGCA; SEQ ID NO: 4279), U1 hairpin II (AAUCCAUUGCACUCCGGAUU; SEQ ID NO: 4280), Uvsx (CCUCUUCGGAGG; SEQ ID NO: 4281), PP7 (AGGAGUUUCUAUGGAAACCCU; SEQ ID NO: 4282), phage replication loop (AGGUGGGACGACCUCUCGGUCGUCCUAUCU; SEQ ID NO: 4283), stapled loop_a (UGCUCGCUCCGUUCGAGCA; SEQ ID NO: 4284), stapled loop_b1 (UGCUCGACGCGUCCUCGAGCA; SEQ ID NO: 4285), stapled loop_b2 (UGCUCGUUUGCGGCUACGAGCA; SEQ ID NO: 4286), G-quadruplex M3q (AGGGAGGGAGGGAGAGG; SEQ ID NO: 4287), G-quadruplex telomeric basket (GGUUAGGGUUAGGGUUAGG; SEQ ID NO: 4288), acomitatin-ricin loop (CUGCUCAGUACGAGAGGAACCGCAG; SEQ ID NO: 4289), or a pseudoknot (UACACUGGGAUCGCUGAAUUAGAGAUCGGCGUCCUUUCAUUCUAUAUACUUUGGAGUUUUAAAAUGUCUCUAAGUACA; SEQ ID NO: 4290). In some embodiments, the exogenous stem loop comprises an RNA scaffold. As used herein, an “RNA scaffold” refers to a multi-dimensional RNA structure that is capable of interacting with and organizing or positioning one or more proteins. In some embodiments, the RNA scaffold is synthetic or non-naturally occurring. In some embodiments, the exogenous stem loop comprises a long non-coding RNA (IncRNA). As used herein, an IncRNA refers to a non-coding RNA that is longer than about 200 bp in length. In some embodiments, the 5’ and 3’ ends of the exogenous stem loop base pair, i.e., interact to form a double helical RNA region. In some embodiments, the 5’ and 3’ ends of the exogenous stem loop base pair, and one or more regions between the 5’ and 3’ ends of the exogenous stem loop do not base pair.In some embodiments, the at least one nucleotide modification comprises: (a) a substitution of 1 to 15 contiguous or non-contiguous nucleotides of the gNA variant in one or more regions; (b) a deletion of 1 to 10 contiguous or non-contiguous nucleotides of the gNA variant in one or more regions; (c) an insertion of 1 to 10 contiguous or non-contiguous nucleotides of the gNA variant in one or more regions; (d) a substitution of a scaffold stem loop or an extended stem loop with a RNA stem loop sequence from a heterologous RNA source having a proximal 5' and 3' end; or any combination of (a)-(d).
[0202] In some embodiments, the gNA variant comprises the sequence of any one of SEQ ID NOs: 412-3295 or a subsequence thereof and the sequence of an exogenous stem loop. In some embodiments, the gNA variant comprises the sequence of any one of SEQ ID NOs: 2236, 2237, 2238, 2241, 2244, 2248, 2249, or 2259-2280 or a subsequence thereof and the sequence of an exogenous stem loop. In some embodiments, the gNA variant comprises the sequence of any one of SEQ ID NOs: 2236, 2237, 2238, 2241, 2244, 2248, 2249, or 2259-2280 or a subsequence thereof and the sequence of an exogenous stem loop.
[0203] In some embodiments, the gNA variant comprises a scaffold stem loop having at least 60% identity to SEQ ID NO: 14. In some embodiments, the gNA variant comprises a scaffold stem loop having at least 60% identity, at least 70% identity, at least 80% identity, at least 90% identity, at least 95% identity, at least 98% identity, or at least 99% identity to SEQ ID NO: 14. In some embodiments, the gNA variant contains a scaffold stem loop comprising SEQ ID NO: 14.
[0204] In some embodiments, the gNA variant comprises a scaffold stem loop sequence of CCAGCGACUAUGUCGUAGUGG (SEQ ID NO: 245). In some embodiments, the gNA variant comprises a scaffold stem loop sequence of CCAGCGACUAUGUCGUAGUGG (SEQ ID NO: 245) with at least 1, 2, 3, 4, or 5 mismatches thereto.
[0205] In some embodiments, the gNA variant contains an extended stem loop region comprising less than 32 nucleotides, less than 31 nucleotides, less than 30 nucleotides, less than 29 nucleotides, less than 28 nucleotides, less than 27 nucleotides, less than 26 nucleotides, less than 25 nucleotides, less than 24 nucleotides, less than 23 nucleotides, less than 22 nucleotides, less than 21 nucleotides, or less than 20 nucleotides. In some embodiments, the gNA variant contains an extended stem loop region comprising less than 32 nucleotides. In some embodiments, the gNA variant further comprises a thermostable stem loop.
[0206] In some embodiments, the sgRNA variant comprises the sequence of SEQ ID NO: 2104, 2106, SEQ ID NO: 2163, SEQ ID NO: 2107, SEQ ID NO: 2164, SEQ ID NO: 2165, SEQ ID NO: 2166, SEQ ID NO: 2103, SEQ ID NO: 2167, SEQ ID NO: 2105, SEQ ID NO: 2108, SEQ ID NO: 2112, SEQ ID NO: 2160, SEQ ID NO: 2170, SEQ ID NO: 2114, SEQ ID NO: 2171, SEQ ID NO: 2112, SEQ ID NO: 2173, SEQ ID NO: 2102, SEQ ID NO: 2174, SEQ ID NO: 2175, SEQ ID NO: 2109, SEQ ID NO: 2176, SEQ ID NO: 2238, SEQ ID NO: 2239, SEQ ID NO: 2240, or SEQ ID NO: 2241.
[0207] In some embodiments, the gNA variant comprises one or more additional alterations to the sequence of any one of SEQ ID NOs: 2201-2280. In some embodiments, the gNA variant comprises any one of SEQ ID NOs: 2236, 2237, 2238, 2241, 2244, 2248, 2249, or 2259-2280, or a sequence having at least about 80%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% identity thereto. In some embodiments, the gNA variant comprises one or more additional alterations to the sequence of any one of SEQ ID NOs: 2201-2280. In some embodiments, the gNA variant comprises the sequence of any one of SEQ ID NOs: 2236, 2237, 2238, 2241, 2244, 2248, 2249, or 2259-2280.
[0208] In some embodiments, the sgRNA variant comprises one or more additional alterations to the sequence of SEQ ID NO: 2104, SEQ ID NO: 2163, SEQ ID NO: 2107, SEQ ID NO: 2164, SEQ ID NO: 2165, SEQ ID NO: 2166, SEQ ID NO: 2103, SEQ ID NO: 2167, SEQ ID NO: 2105, SEQ ID NO: 2108, SEQ ID NO: 2112, SEQ ID NO: 2160, SEQ ID NO: 2170, SEQ ID NO: 2114, SEQ ID NO: 2171, SEQ ID NO: 2112, SEQ ID NO: 2173, SEQ ID NO: 2102, SEQ ID NO: 2174, SEQ ID NO: 2175, SEQ ID NO: 2109, SEQ ID NO: 2176, SEQ ID NO: 2238, SEQ ID NO: 2239, SEQ ID NO: 2240, or SEQ ID NO: 2241.
[0209] In some embodiments of the gNA variants of the application, the gNA variant comprises at least one modification, wherein the at least one modification to the reference guide scaffold of SEQ ID NO: 5 is selected from one or more of: (a) a C18G substitution in the triple helix loop; (b) a G55 insertion in the stem bubble; (c) a U1 deletion; (d) a modification of the extended stem loop, wherein (i) the 6nt loop and the 13 loop proximal base pairs are replaced by a Uvsx hairpin; and (ii) a deletion of A99 and a substitution of G65U creates a fully base-paired loop distal base. In such embodiments, the gNA variant comprises the sequence of any one of SEQ ID NOs: 2236, 2237, 2238, 2241, 2244, 2248, 2249, or 2259-2280.
[0210] In some embodiments, the scaffold of a gNA variant comprises the sequence of any one of SEQ ID NOs:2201-2280 of Table 2. In some embodiments, the scaffold of a gNA consists of or consists essentially of the sequence of any one of SEQ ID NOs:2201-2280. In some embodiments, the scaffold of a gNA variant is at least about 60% identical, at least about 65% identical, at least about 70% identical, at least about 75% identical, at least about 80% identical, at least about 85% identical, at least about 90% identical, at least about 91% identical, at least about 92% identical, at least about 93% identical, at least about 94% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, or at least about 99% identical to any one of SEQ ID NOs:2201-2280.
[0211] In some embodiments, the gNA variant further comprises a spacer (or targeting sequence) region more fully described supra, comprising at least 14 to about 35 nucleotides, wherein the spacer is designed to have a sequence complementary to a target DNA. In some embodiments, the gNA variant comprises a targeting sequence of at least 10 to 30 nucleotides complementary to a target DNA. In some embodiments, the targeting sequence has 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, or 35 nucleotides. In some embodiments, the gNA variant comprises a targeting sequence of 20 nucleotides. In some embodiments, the targeting sequence has 25 nucleotides. In some embodiments, the targeting sequence has 24 nucleotides. In some embodiments, the targeting sequence has 23 nucleotides. In some embodiments, the targeting sequence has 22 nucleotides. In some embodiments, the targeting sequence has 21 nucleotides. In some embodiments, the targeting sequence has 20 nucleotides. In some embodiments, the targeting sequence has 19 nucleotides. In some embodiments, the targeting sequence has 18 nucleotides. In some embodiments, the targeting sequence has 17 nucleotides. In some embodiments, the targeting sequence has 16 nucleotides. In some embodiments, the targeting sequence has 15 nucleotides. In some embodiments, the targeting sequence has 14 nucleotides.
[0212] In some embodiments, the scaffold of a gNA variant is a variant comprising one or more additional alterations to the sequence of a reference gRNA comprising SEQ ID NO:4 or SEQ ID NO:5. In those embodiments where the scaffold of the reference gRNA is derived from SEQ ID NO:4 or SEQ ID NO:5, the one or more improved or increased features of the gNA variant are improvements over the same features in SEQ ID NO:4 or SEQ ID NO:5.
[0213] In some embodiments, the scaffold of the gNA variant is part of an RNP with a reference CasX protein comprising SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 3. In other embodiments, the scaffold of the gNA variant is part of an RNP with a CasX variant protein comprising any one of the sequences of Tables 3, 8, 9, 10, and 12, or a sequence having at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto. In the foregoing embodiments, the gNA further comprises a spacer sequence.
[0214] h. Chemically modified gNA
[0215] In some embodiments, the present application provides chemically modified gNAs. In some embodiments, the present application provides chemically modified gNAs that have a guide NA function and reduced susceptibility to cleavage by a nuclease. A gNA comprising any nucleotide other than the four canonical ribonucleotides A, C, G, and U, or deoxynucleotides, is a chemically modified gNA. In some cases, a chemically modified gNA comprises any backbone or internucleotide linkage other than a natural phosphodiester internucleotide linkage. In certain embodiments, retaining functionality includes the ability of the modified gNA to bind to a CasX of any of the embodiments described herein. In certain embodiments, retaining functionality includes the ability of the modified gNA to bind to a target nucleic acid sequence. In certain embodiments, retaining functionality includes the ability of a targeting CasX protein or precomplexed RNP to bind to a target nucleic acid sequence. In certain embodiments, retaining functionality includes the ability of a CasX-gNA to cleave a target polynucleotide. In certain embodiments, retaining functionality includes the ability of a CasX-gNA to cleave a target nucleic acid sequence. In certain embodiments, retaining functionality is any other known function of a gNA in a recombination system with a CasX chimeric protein of an embodiment of the present application.
[0216] In some embodiments, the present application provides chemically modified gNAs, wherein the nucleotide sugar modification is incorporated into a gNA selected from the group consisting of: 2'-0-C 1-4 alkyl (such as 2'-0-methyl (2'-OMe)), 2'-deoxy (2'-H), 2'-0-C 1-3 alkyl-O-C 1-3alkyl (e.g., 2'-methoxyethyl ("2'-MOE")), 2'-fluoro ("2'-F"), 2'-amino ("2'-NH2"), 2'-arabinosyl ("2'-ara"), 2'-F-arabinosyl ("2'-F-ara"), 2'-locked nucleic acid ("LNA") nucleotides, 2'-unlocked nucleic acid ("ULNA") nucleotides, sugars in the L form ("L-sugars"), and 4'-thioribosyl nucleotides. In other embodiments, the internucleotide linkage modifications incorporated into the guide RNA are selected from the group consisting of: phosphorothioate "P(S)" (P(S)), phosphonocarboxylate (P(CH2)COOR) (e.g., phosphonacetic acid ester "PACE" (P(CH2COO n )) and phosphonothioate ((S)P(CH2)COOR) (e.g., phosphonothioacetic acid ester "thioPACE" ((S)P(CH2)COO - )) and phosphonothioate ((S)P(CH2)COOR) (e.g., phosphonothioacetic acid ester "thioPACE" ((S)P(CH2)COO n )) and phosphonothioate ((S)P(CH2)COOR) (e.g., phosphonothioacetic acid ester "thioPACE" ((S)P(CH2)COO n )) and phosphonothioate ((S)P(CH2)COOR) (e.g., phosphonothioacetic acid ester "thioPACE" ((S)P(CH2)COO - )) and phosphonothioate ((S)P(CH2)COOR) (e.g., phosphonothioacetic acid ester "thioPACE" ((S)P(CH2)COO 1-3 alkyl) (e.g., methylphosphonate-P(CH3), boranophosphonate (P(BH3)), and dithiophosphonate (P(S)2).
[0217] In certain embodiments, the present application provides chemically modified gNAs, wherein nucleobase ("base") modifications are incorporated into the gNA selected from the group consisting of: 2-thiouracil ("2-thio U"), 2-thiocytosine ("2-thio C"), 4-thiouracil ("4-thio U"), 6-thioguanine ("6-thio G"), 2-amino adenine ("2-amino A"), 2-amino purine, pseudouracil, hypoxanthine, 7-deaza guanine, 7-deaza-8-azaguanine, 7-deaza adenine, 7-deaza-8-azadenine, 5-methylcytosine ("5-methyl C"), 5-methyluracil ("5-methyl U"), 5-hydroxymethylcytosine, 5-hydroxymethyluracil, 5,6-dihydro uracil, 5-propynylcytosine, 5-propynyluracil, 5-propynylcytosine, 5-ethynyluracil, 5-allyluracil ("5-allyl U"), 5-allylcytosine ("5-allyl C"), 5-aminoallyluracil ("5-aminoallyl U"), 5-aminoallyl-cytosine ("5-aminoallyl C"), abasic nucleotides, Z base, P base, unstructured nucleic acid ("UNA"), iso guanine ("iso G"), iso cytosine ("iso C"), 5-methyl-2-pyrimidine, x (A, G, C, T), and y (A, G, C, T).
[0218] In other embodiments, the present application provides chemically modified gNAs, wherein one or more isotopic modifications are introduced on the nucleotide sugar, nucleobase, phosphodiester linkage, and / or phosphonucleotide, including one or more 15 N, 13 C, 14 C, deuterium, 3 H, 32 P, 125 I, 131 I atoms or other atoms or elements used as tracers.
[0219] In some embodiments, the "end" modification incorporated into the gNA is selected from the group consisting of: PEG (polyethylene glycol); hydrocarbon linkers (including: heteroatom (O, S, N) substituted hydrocarbon spacers; halogen substituted hydrocarbon spacers; ketone, carboxyl, amido, sulfinyl, carbamoyl, thiocarbamoyl containing hydrocarbon spacers); spermine linkers; dyes linked to linkers such as 6-fluorescein-hexyl, including fluorescent dyes (e.g. fluorescein, rhodamine, cyanine); quenchers (e.g. dabcyl, BHQ) and other labels (e.g. biotin, digoxigenin, acridine, streptavidin, avidin, peptides and / or proteins). In some embodiments, the "end" modification comprises binding (or linking) the gNA to another molecule, peptide, protein, sugar, oligosaccharide, steroid, lipid, folate, vitamin, and / or other molecule comprising deoxyribonucleotides and / or ribonucleotides. In certain embodiments, the present application provides chemically modified gNAs, wherein the "end" modification (described above) is positioned internally within the gNA sequence via a linker such as 2-(4-butylamidofluorescein)propane-1,3-diol bis(phosphodiester) linker, which is incorporated as a phosphodiester linkage and can be incorporated at any position between two nucleotides in the gNA.
[0220] In some embodiments, the present application provides chemically modified gNAs with end modifications comprising end functional groups such as amines, thiols (or sulfhydryls), hydroxyls, carboxyls, carbonyls, sulfinyls, thiocarbonyls, carbamoyls, thiocarbamoyls, phosphoryls, alkenes, alkynes, halogens, or functional group terminated linkers, which can be subsequently bound to a desired moiety selected from the group consisting of: fluorescent dyes, non-fluorescent labels, tags (e.g. 14 C, biotin, avidin, streptavidin, or containing isotopic labels such as 15 N, 13 C, deuterium, 3 H, 32 P, 125nucleotides and / or ribonucleotides, including aptamers), amino acids, peptides, proteins, sugars, oligosaccharides, steroids, lipids, folates, and vitamins. Conjugation employs standard chemical methods well known in the art, including but not limited to via N-hydroxysuccinimide, isothiocyanate, DCC (or DCI) coupling, and / or any other standard method as described in “Bioconjugate Techniques”, Greg T. Hermanson, Publisher Elsevier Science, 3rdEdition (2013), the contents of which are incorporated herein by reference in their entirety.
[0221] i. Form a complex with a CasX protein
[0222] In some embodiments, the gNA variant has improved ability to form a complex with a CasX protein (such as a reference CasX or CasX variant protein) when compared to a reference gRNA. In some embodiments, the gNA variant has improved affinity for a CasX protein (such as a reference or variant protein) when compared to a reference gRNA, thereby improving its ability to form a ribonucleoprotein (RNP) complex with the CasX protein, as described in the Examples. In some embodiments, improved ribonucleoprotein complex formation can increase the efficiency of assembling functional RNP. In some embodiments, greater than 90%, greater than 93%, greater than 95%, greater than 96%, greater than 97%, greater than 98%, or greater than 99% of RNP comprising the gNA variant and a spacer are competent for gene editing of a target nucleic acid.
[0223] In some embodiments, exemplary nucleotide changes that can improve the ability of a gNA variant to form a complex with a CasX protein can include replacing the scaffold stem with a thermostable stem loop. Without wishing to be bound by any theory, replacing the scaffold stem with a thermostable stem loop can increase the overall binding stability of the gNA variant to the CasX protein. Alternatively or additionally, removing a large section of the stem loop can alter the gNA variant folding kinetics and make it easier and faster for the functional folded gNA to structurally assemble, for example, by mitigating the extent to which the gNA variant can become “tangled” on itself. In some embodiments, the choice of scaffold stem loop sequence can vary with different spacers used for the gNA. In some embodiments, the scaffold sequence can be adapted to the spacer and thus to the target sequence. Biochemical assays can be used to assess the binding affinity of the CasX protein to the gNA variant to form an RNP, including the assays of the examples. For example, a person of ordinary skill can measure the change in the amount of fluorescently labeled gNA that binds to immobilized CasX protein as a response to increasing the concentration of additional unlabeled “cold competitor” gNA. Alternatively or additionally, the fluorescent signal can be monitored or viewed for how it changes as different amounts of fluorescently labeled gNA flow past the immobilized CasX protein. Alternatively, in vitro cleavage assays can be used to assess the ability to form an RNP relative to a defined target nucleic acid sequence.
[0224] j. gNA stability
[0225] In some embodiments, the gNA variant has improved stability when compared to a reference gRNA. In some embodiments, increased stability and efficient folding can increase the extent to which the gNA variant persists inside a target cell, which can in turn increase the probability of forming a functional RNP capable of performing CasX functions such as gene editing. In some embodiments, increased gNA variant stability can also allow for similar results with lower amounts of gNA delivered to a cell, which can in turn decrease the probability of off-target effects during gene editing.
[0226] In other embodiments, the present disclosure provides gNA, wherein the scaffold stem loop and / or the extended stem loop is replaced with a hairpin loop or a thermostable RNA stem loop, wherein the resulting gNA has increased stability and, depending on the choice of loop, can interact with certain cellular proteins or RNAs. In some embodiments, the replacement RNA loop is selected from the group consisting of MS2, Qβ, U1 hairpin II, Uvsx, PP7, phage replication loop, adnate loop al, adnate loop bl, G-quadruplex M3q, G-quadruplex telomere basket, aecleptin-ricin loop, and a pseudoknot. Sequences of gNA variants including such components are provided in Table 2.
[0227] Guide NA stability can be assessed in a variety of ways, including, for example, in vitro by assembling the guide, incubating in a solution that mimics the intracellular environment for varying periods of time, and then measuring functional activity via the in vitro cleavage assay described herein. Alternatively or additionally, gNAs can be harvested from cells at different time points after initial transfection / transduction of the gNA to determine how long the gNA variant is retained relative to the reference gRNA.
[0228] k. solubility
[0229] In some embodiments, the gNA variant has improved solubility when compared to the reference gRNA. In some embodiments, the gNA variant has improved CasX protein:gNA RNP solubility when compared to the reference gRNA. In some embodiments, the solubility of the CasX protein:gNA RNP is improved by adding a ribonuclease sequence to the 5’ or 3’ end of the gNA variant, for example, the 5’ or 3’ of the reference sgRNA. Some ribonucleases, such as M1 ribonucleases, can increase the solubility of a protein via RNA-mediated protein folding.
[0230] Increased solubility of the CasX RNP comprising the gNA variant as described herein can be assessed via a variety of methods known to one of skill in the art, such as by taking densitometry readings on gels of the soluble fraction of E. coli expressing the CasX and gNA variant.
[0231] l. nuclease activity resistance
[0232] In some embodiments, the gNA variant has improved nuclease activity resistance when compared to the reference gRNA. Without wishing to be bound by any theory, increased resistance to nucleases, such as those found in cells, can, for example, increase the persistence of the variant gNA in the intracellular environment, thereby improving gene editing.
[0233] Many nucleases are processive and degrade RNA in a 3’ to 5’ manner. Thus, in some embodiments, the addition of a nuclease resistant secondary structure to one or both ends of the gNA, or nucleotide changes that alter the secondary structure of the sgNA, can result in a gNA variant with increased nuclease activity resistance. Nuclease activity resistance can be assessed via a variety of methods known to one of skill in the art. For example, in vitro methods of measuring nuclease activity resistance can include, for example, contacting a reference gNA with a variant having one or more exemplary RNA nucleases and measuring degradation. Alternatively or additionally, measuring the persistence of the gNA variant in a cellular environment using the methods described herein can be indicative of the degree of nuclease resistance of the gNA variant.
[0234] m. binding affinity to a target DNA
[0235] In some embodiments, the gNA variant has an improved affinity for the target DNA relative to a reference gRNA. In some embodiments, the ribonucleoprotein complex containing the gNA variant has an improved affinity for the target DNA relative to the affinity of an RNP containing a reference gRNA. In some embodiments, the improved affinity of the RNP for the target DNA includes improved affinity for the target sequence, improved affinity for the PAM sequence, improved RNP search for DNA for the target sequence, or any combination thereof. In some embodiments, the improved affinity for the target DNA is a result of increased overall DNA-binding affinity.
[0236] Without being bound by theory, it is possible that nucleotide changes in gNA variants affecting the function of the OBD in the CasX protein could increase the affinity of the CasX variant protein for binding to the preseptal adjacent motif (PAM), and for binding to or utilizing a wider range of PAM sequences (including PAM sequences selected from the group consisting of TTC, ATC, GTC, and CTC) beyond the typical TTC PAM recognized by the reference CasX protein of SEQ ID NO:2. This would increase the affinity and diversity of the CasX variant protein for the target DNA sequence, thereby increasing its ability to edit and / or bind to the target nucleic acid sequence (compared to the reference CasX). As described more fully below, the increased editability of the target nucleic acid sequence compared to the reference CasX refers to the PAM and preseptal sequence and their orientation relative to the non-target strand. This does not mean that the non-target strand, or the PAM sequence of the non-target strand, determines cleavage or is mechanistically involved in target recognition. For example, when referring to the TTC PAM, it could actually be the complementary GAA sequence required for target cleavage, or it could be a combination of nucleotides from both strands. In the case of the CasX protein disclosed herein, PAM is located at the 5' of the anterior spacer, wherein at least a single nucleotide separates PAM from the first nucleotide of the anterior spacer. Alternatively or additionally, changes in the function of the helical I and / or helical II domains that affect the affinity of the CasX variant protein for the target DNA strand can increase the affinity of the CasX RNP containing the variant gNA for the target DNA.
[0237] n. Adding or changing gNA functionality
[0238] In some embodiments, gNA variants can comprise large structural changes relative to a reference gRNA that alter the topological structure of the gNA variant, thereby allowing for different gNA functions. For example, in some embodiments, gNA variants exchange the endogenous stem loops of the reference gRNA scaffold with previously identified stable RNA structures or stem loops that can interact with protein or RNA binding partners to recruit additional moieties to CasX to CasX to specific locations, such as inside a viral capsid with binding partners to the RNA structure. In other contexts, RNAs can complement each other (such as in a kissing loop), such that two CasX proteins can be co-localized for more efficient gene editing at a target DNA sequence. Such RNA structures can include MS2, Qp, U1 hairpin II, Uvsx, PP7, phage replication loop, kissing loop_a, kissing loop_b1, kissing loop_b2, G-quadruplex M3q, G-quadruplex telomeric basket, a ribavirin-ricin loop, or a pseudoknot.
[0239] In some embodiments, the gNA variant comprises a terminal fusion partner. The term gNA variant includes variants that contain exogenous sequences (such as terminal fusions) or internal insertions. Exemplary terminal fusions can include fusions of gRNA with self-cleaving ribozymes or protein binding motifs. As used herein, “ribozyme” refers to an RNA or segment thereof that has one or more catalytic activities analogous to protein enzymes. Exemplary ribozyme catalytic activities can include, for example, cleavage and / or ligation of RNA, cleavage and / or ligation of DNA, or peptide bond formation. In some embodiments, such fusions can improve scaffold folding or recruit DNA repair machinery. For example, in some embodiments, the gRNA can be fused with a hepatitis delta virus (HDV) anti-genome ribozyme, an HDV genome ribozyme, a hatchet ribozyme (from metagenomic data), an env25 pistol ribozyme (representative from Aliistipes putredinis), an HH15 minimal hammerhead ribozyme, a tobacco ring spot virus (TRSV) ribozyme, a WT virus hammerhead ribozyme (and rational variants), or a twisted sister 1 or RBMX recruiting motif. Hammerhead ribozymes are RNA motifs that catalyze reversible cleavage and ligation reactions at specific sites within the RNA molecule. Hammerhead ribozymes include Type I, Type II, and Type III hammerhead ribozymes. HDV, pistol, and hatchet ribozymes have self-cleavage activity. gNA variants comprising one or more ribozymes can allow for expanded gNA functionality compared to gRNA references. For example, in some embodiments, a gNA comprising a self-cleaving ribozyme can be transcribed and processed into a mature gNA as part of a polycistronic transcript. Such fusions can occur at the 5’ or 3’ end of the gNA. In some embodiments, the gNA variant comprises a fusion at both the 5’ and 3’ end, where each fusion is independently as described herein. In some embodiments, the gNA variant comprises a bacteriophage replication loop or tetraloop. In some embodiments, the gNA comprises a hairpin loop capable of binding a protein. For example, in some embodiments, the hairpin loop is a MS2, Qp, U1 hairpin II, Uvsx, or PP7 hairpin loop.
[0240] In some embodiments, the gNA variant comprises one or more RNA aptamer. As used herein, “RNA aptamer” refers to an RNA molecule that binds a target with high affinity and high specificity.
[0241] In some embodiments, the gNA variant comprises one or more riboswitch. As used herein, “riboswitch” refers to an RNA molecule that changes state upon binding a small molecule.
[0242] In some embodiments, the gNA variants further comprise one or more protein binding motifs. In some embodiments, the addition of a protein binding motif to a reference gRNA or gNA variant of the present application can allow for CasX RNP association with additional proteins, which can, for example, add the functionality of those proteins to the CasX RNP.
[0243] IV. CasX Proteins for Modifying Target Nucleic Acids
[0244] As used herein, the term “CasX protein” refers to a family of proteins and encompasses all naturally occurring CasX proteins, proteins having at least 50% identity to a naturally occurring CasX protein, and CasX variants having one or more improved characteristics relative to a reference CasX protein that is naturally occurring. Exemplary improved characteristics of CasX variant embodiments include, but are not limited to, improved variant folding, improved binding affinity for gNAs, improved binding affinity for target nucleic acids, improved ability to edit with a larger range of PAM sequences and / or to bind target DNA, improved target DNA unwinding, increased editing activity, improved editing efficiency, improved editing specificity, increased percentage of eukaryotic genomes that can be effectively edited, increased nuclease activity, increased target strand loading for double-strand cleavage, reduced target strand loading for single-strand cleavage, reduced off-target cleavage, improved binding of non-target strands of DNA, improved protein stability, improved protein:gNA (RNP) complex stability, improved protein solubility, improved protein:gNA (RNP) complex solubility, improved protein yield, improved protein expression, and improved melting characteristics, as more fully described below. In the foregoing embodiments, one or more of the improved characteristics of the CasX variant is at least about 1.1 to about 100,000 fold improved when analyzed in a similar manner relative to a reference CasX protein of SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 3. In other embodiments, the improvement is at least about 1.1 fold, at least about 2 fold, at least about 5 fold, at least about 10 fold, at least about 50 fold, at least about 100 fold, at least about 500 fold, at least about 1000 fold, at least about 5000 fold, at least about 10,000 fold, or at least about 100,000 fold improved compared to a reference CasX protein of SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 3 when analyzed in a similar manner.
[0245] The term CasX variant includes variants that are fusion proteins, i.e., CasX “fused to” a heterologous sequence. This includes CasX variants that comprise N-terminal, C-terminal, or internal fusions of a CasX variant sequence and CasX to a heterologous protein or domain thereof.
[0246] CasX proteins of the present invention comprise at least one of the following domains: a non-target strand binding (NTSB) domain, a target strand loading (TSL) domain, a helical I domain, a helical II domain, an oligonucleotide binding domain (OBD), and a RuvC DNA cleavage domain (the last of which can be modified or deleted in catalytically dead CasX variants), more fully described below. Additionally, CasX variant proteins of the present invention have an enhanced ability to effectively edit and / or bind to a target DNA using a PAM sequence selected from TTC, ATC, GTC, or CTC as compared to a wild-type reference CasX protein. In the foregoing, the PAM sequence is at least 1 nucleotide 5' to the non-target strand of the protospacer having identity to the targeting sequence of the gNA in the assay system as compared to the editing efficiency and / or binding of an RNP comprising the reference CasX protein in a similar assay system.
[0247] In some cases, the CasX protein is a naturally occurring protein (e.g., naturally occurring in and isolated from a prokaryotic cell). In other embodiments, the CasX protein is not a naturally occurring protein (e.g., the CasX protein is a CasX variant protein, a chimeric protein, etc.). Naturally occurring CasX proteins (referred to herein as “reference CasX proteins”) act as endonucleases that catalyze a double-stranded break in a target double-stranded DNA (dsDNA) at a specific sequence. Sequence specificity is provided by the targeting sequence of the cognate gNA with which it is complexed, which hybridizes to a target sequence within a target nucleic acid.
[0248] In some embodiments, a CasX protein can bind and / or modify (e.g., cleave, cut, methylate, demethylate, etc.) a target nucleic acid and / or a polypeptide associated with a target nucleic acid (e.g., methylation or acetylation of a histone tail). In some embodiments, a CasX protein is catalytically dead (dCasX), but retains the ability to bind a target nucleic acid. Exemplary catalytically dead CasX proteins include one or more mutations in the active site of the RuvC domain of a CasX protein. In some embodiments, a catalytically dead CasX protein includes a substitution at residue 672, 769, and / or 935 of SEQ ID NO: 1. In one embodiment, a catalytically dead CasX protein includes D672A, E769A, and / or D935A substitutions in a reference CasX protein of SEQ ID NO: 1. In other embodiments, a catalytically dead CasX protein includes a substitution at amino acid 659, 756, and / or 922 in a reference CasX protein of SEQ ID NO: 2. In some embodiments, a catalytically dead CasX protein includes D659A, E756A, and / or D922A substitutions in a reference CasX protein of SEQ ID NO: 2. In other embodiments, a catalytically dead CasX protein includes a deletion of all or a portion of the RuvC domain of a CasX protein. It will be appreciated that the same foregoing substitutions can be similarly introduced into a CasX variant of the application, resulting in a dCasX variant. In one embodiment, all or a portion of the RuvC domain is deleted from a CasX variant, resulting in a dCasX variant. In some embodiments, a dCasX variant protein that is catalytically inactive can be used for base editing or epigenetic modification. With a higher affinity for DNA, in some embodiments, a dCasX variant protein that is catalytically inactive can find its target nucleic acid faster, remain bound to the target nucleic acid for a longer period of time, bind the target nucleic acid in a more stable manner, or a combination thereof, relative to a catalytically active CasX, thereby improving the function of the catalytically dead CasX variant protein.
[0249] a. non-target strand binding domain
[0250] A reference CasX protein of the disclosure comprises a non-target strand binding domain (NTSBD). The NTSBD is a domain that has not been previously found in any Cas protein; for example, this domain is not present in Cas proteins such as Cas9, Casl2a / Cpfl, Casl3, Casl4, CASCADE, CSM, or CSY. Without being bound by theory or mechanism, the NTSBD in CasX allows for binding to a non-target DNA strand and can aid in the unwinding of the non-target and target strands. It is hypothesized that the NTSBD is responsible for the unwinding of the non-target DNA strand or the capture of the non-target DNA strand in an unwound state. The NTSBD is in direct contact with the non-target strand in the CryoEM model structure derived to date and can contain an atypical zinc finger domain. The NTSBD can also play a role in stabilizing the DNA during unwinding, guide RNA invasion, and R-loop formation. In some embodiments, an exemplary NTSBD comprises amino acids 101-191 of SEQ ID NO: 1 or amino acids 103-192 of SEQ ID NO: 2. In some embodiments, the NTSBD of a reference CasX protein comprises a four-stranded beta sheet.
[0251] b. Target strand loading domain
[0252] A reference CasX protein of the disclosure comprises a target strand loading (TSL) domain. The TSL domain is a domain that is not found in certain Cas proteins such as Cas9, CASCADE, CSM, or CSY. Without being bound by theory or mechanism, it is believed that the TSL domain is responsible for assisting in loading the target DNA strand into the RuvC active site of the CasX protein. In some embodiments, the TSL serves to place or capture the target strand in a folded state that places the scissile phosphate of the target strand DNA backbone into the RuvC active site. The TSL comprises a cys4 (CXXC (SEQ ID NO: 246, CXXC (SEQ ID NO: 246) zinc finger / belt domain isolated by a majority of the TSL. In some embodiments, an exemplary TSL comprises amino acids 825-934 of SEQ ID NO: 1 or amino acids 813-921 of SEQ ID NO: 2.
[0253] c. Helical I domain
[0254] A reference CasX protein of the disclosure comprises a helical I domain. Certain Cas proteins other than CasX have domains that can be named in a similar fashion. However, in some embodiments, the helical I domain of a CasX protein comprises one or more unique structural features, or comprises unique sequences, or a combination thereof, as compared to domains in other Cas proteins that can have a similar name. For example, in some embodiments, the helical I domain of a CasX protein comprises one or more unique secondary structures as compared to domains in other Cas proteins that can have a similar name. For example, in some embodiments, the helical I domain in a CasX protein comprises one or more alpha helices that are unique in arrangement, number, and length as compared to other CRISPR proteins. In certain embodiments, the helical I domain is responsible for binding DNA and spacer interactions with guide RNA. Without wishing to be bound by theory, it is believed that in some cases, the helical I domain can facilitate binding of a protospacer adjacent motif (PAM). In some embodiments, an exemplary helical I domain comprises amino acids 57-100 and 192-332 of SEQ ID NO: 1, or amino acids 59-102 and 193-333 of SEQ ID NO: 2. In some embodiments, the helical I domain of a reference CasX protein comprises one or more alpha helices.
[0255] d. Helical II domain
[0256] A reference CasX protein of the disclosure comprises a helical II domain. Certain Cas proteins other than CasX have domains that can be named in a similar fashion. However, in some embodiments, the helical II domain of a CasX protein comprises one or more unique structural features, or unique sequences, or a combination thereof, as compared to domains in other Cas proteins that can have a similar name. For example, in some embodiments, the helical II domain comprises one or more unique structural alpha helix bundles aligned along the target DNA:guide RNA channel. In some embodiments, in a CasX comprising a helical II domain, the target strand and guide RNA interact with the helical II (and in some embodiments, the helical I domain) to allow access of the RuvC domain to the target DNA. The helical II domain is responsible for binding to the guide RNA scaffold stem loop as well as the bound DNA. In some embodiments, an exemplary helical II domain comprises amino acids 333-509 of SEQ ID NO: 1, or amino acids 334-501 of SEQ ID NO: 2.
[0257] e. Oligonucleotide binding domain
[0258] Reference CasX proteins of the invention comprise an oligonucleotide binding domain (OBD). Certain Cas proteins other than CasX have domains that can be named in analogous fashion. However, in some embodiments, the OBD comprises one or more unique functional features, or comprises a sequence unique to CasX proteins, or a combination thereof. For example, in some embodiments, the bridging helix (BH), the helix I domain, the helix II domain, and the oligonucleotide binding domain (OBD) together are responsible for binding the CasX protein to the guide RNA. Thus, for example, in some embodiments, the OBD is unique to CasX proteins in that it functionally interacts with the helix I domain, or the helix II domain, or both, each of which can be unique to CasX proteins as described herein. In particular, in CasX, the OBD binds largely to the RNA triple helix of the guide RNA scaffold. The OBD can also be responsible for binding to the protospacer adjacent motif (PAM). Exemplary OBD domains comprise amino acids 1-56 and 510-660 of SEQ ID NO: 1, or amino acids 1-58 and 502-647 of SEQ ID NO: 2.
[0259] f. RuvC DNA cleavage domain
[0260] Reference CasX proteins of the invention comprise a RuvC domain, which includes 2 partial RuvC domains (RuvC-I and RuvC-II). The RuvC domain is an ancestral domain of all type 12 CRISPR proteins. The RuvC domain is derived from a TnpB (transposase B)-like transposase. Like other RuvC domains, the CasX RuvC domain has a DED catalytic triad responsible for coordinating a magnesium (Mg) ion and cleaving DNA. In some embodiments, the RuvC has a DED motif active site responsible for cleaving both strands of DNA (one after the other, most likely first the non-target strand at 11-14 nucleotides (nt) in the target sequence, and then subsequently the target strand at 2-4 nt after the target sequence). In particular, in CasX, the RuvC domain is unique in that it is also responsible for binding the guide RNA scaffold stem loop important for CasX function. Exemplary RuvC domains comprise amino acids 661-824 and 935-986 of SEQ ID NO: 1, or amino acids 648-812 and 922-978 of SEQ ID NO: 2.
[0261] g. Reference CasX proteins
[0262] The present disclosure provides reference CasX proteins. In some embodiments, the reference CasX protein is a naturally occurring protein. For example, the reference CasX protein can be isolated from a naturally occurring prokaryote, such as a Deltaproteobacteria, a Planctomycetes, or a Candidatus Accumulibacter species. The reference CasX protein (sometimes referred to herein as a reference CasX polypeptide) is a type II CRISPR / Cas endonuclease that belongs to the CasX (sometimes referred to as Casl2e) family of proteins that are capable of interacting with a guide NA to form a ribonucleoprotein (RNP) complex. In some embodiments, an RNP complex comprising a reference CasX protein can be targeted to a specific site in a target nucleic acid via base pairing between the targeting sequence (or spacer) of the gNA and a target sequence in the target nucleic acid. In some embodiments, an RNP comprising a reference CasX protein is capable of cleaving a target DNA. In some embodiments, an RNP comprising a reference CasX protein is capable of nicking a target DNA. In some embodiments, an RNP comprising a reference CasX protein is capable of editing a target DNA, for example, in those embodiments where the reference CasX protein is capable of cleaving or nicking the DNA followed by non-homologous end joining (NHEJ), homology directed repair (HDR), homology independent targeted integration (HITI), microhomology-mediated end joining (MMEJ), single-strand annealing (SSA), or base excision repair (BER). In some embodiments, the RNP comprising a CasX protein is a catalytically dead (no catalytic activity or substantially no cleavage activity) CasX protein (dCasX) but retains the ability to bind to a target DNA, as more fully described supra.
[0263] In some cases, the reference CasX protein is isolated or derived from Deltaproteobacteria. In some embodiments, the CasX protein comprises a sequence that is at least 50% identical, at least 60% identical, at least 65% identical, at least 70% identical, at least 75% identical, at least 80% identical, at least 81% identical, at least 82% identical, at least 83% identical, at least 84% identical, at least 85% identical, at least 86% identical, at least 86% identical, at least 87% identical, at least 88% identical, at least 89% identical, at least 89% identical, at least 90% identical, at least 91% identical, at least 92% identical, at least 93% identical, at least 94% identical, at least 95% identical, at least 96% identical, at least 97% identical, at least 98% identical, at least 99% identical, at least 99.5% identical, or 100% identical to:
[0264]
[0265] In some cases, the reference CasX protein is isolated or derived from the phylum Tenericutes. In some embodiments, the CasX protein comprises a sequence that is at least 50% identical, at least 60% identical, at least 65% identical, at least 70% identical, at least 75% identical, at least 80% identical, at least 81% identical, at least 82% identical, at least 83% identical, at least 84% identical, at least 85% identical, at least 86% identical, at least 86% identical, at least 87% identical, at least 88% identical, at least 89% identical, at least 89% identical, at least 90% identical, at least 91% identical, at least 92% identical, at least 93% identical, at least 94% identical, at least 95% identical, at least 96% identical, at least 97% identical, at least 98% identical, at least 99% identical, at least 99.5% identical, or 100% identical to the sequence of:
[0266]
[0267] In some embodiments, the CasX protein comprises SEQ ID NO: 2, or a sequence at least 60% similar thereto. In some embodiments, the CasX protein comprises SEQ ID NO: 2, or a sequence at least 80% similar thereto. In some embodiments, the CasX protein comprises SEQ ID NO: 2, or a sequence at least 90% similar thereto. In some embodiments, the CasX protein comprises SEQ ID NO: 2, or a sequence at least 95% similar thereto. In some embodiments, the CasX protein consists of the sequence of SEQ ID NO: 2. In some embodiments, the CasX protein comprises or consists of a sequence having at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 20, at least 30, at least 40, or at least 50 mutations relative to the sequence of SEQ ID NO: 2. The mutations can be insertions, deletions, amino acid substitutions, or any combination thereof.
[0268] In some cases, the reference CasX protein is isolated or derived from a bacterium tentatively named Songi. In some embodiments, the CasX protein comprises a sequence that is at least 50% identical, at least 60% identical, at least 65% identical, at least 70% identical, at least 75% identical, at least 80% identical, at least 81% identical, at least 82% identical, at least 83% identical, at least 84% identical, at least 85% identical, at least 86% identical, at least 86% identical, at least 87% identical, at least 88% identical, at least 89% identical, at least 89% identical, at least 90% identical, at least 91% identical, at least 92% identical, at least 93% identical, at least 94% identical, at least 95% identical, at least 96% identical, at least 97% identical, at least 98% identical, at least 99% identical, at least 99.5% identical, or 100% identical to the sequence of:
[0269]
[0270] In some embodiments, the CasX protein comprises SEQ ID NO: 3, or a sequence at least 60% similar thereto. In some embodiments, the CasX protein comprises SEQ ID NO: 3, or a sequence at least 80% similar thereto. In some embodiments, the CasX protein comprises SEQ ID NO: 3, or a sequence at least 90% similar thereto. In some embodiments, the CasX protein comprises SEQ ID NO: 3, or a sequence at least 95% similar thereto. In some embodiments, the CasX protein consists of the sequence of SEQ ID NO: 3. In some embodiments, the CasX protein comprises or consists of a sequence having at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 20, at least 30, at least 40, or at least 50 mutations relative to the sequence of SEQ ID NO: 3. The mutations can be insertions, deletions, amino acid substitutions, or any combination thereof.
[0271] h. CasX variant proteins
[0272] The present disclosure provides variants of reference CasX proteins (interchangeably referred to herein as "CasX variants" or "CasX variant proteins"), wherein the CasX variants comprise at least one modification in at least one domain, including but not limited to the sequences of SEQ ID NOs: 1-3, relative to the reference CasX protein. In some embodiments, the CasX variants exhibit at least one improved feature as compared to the reference CasX protein. All variants of CasX variant proteins that improve one or more functions or features of the reference CasX proteins described herein are contemplated as within the scope of the present disclosure. In some embodiments, the modification is a mutation in one or more amino acids of the reference CasX. In other embodiments, the modification is substitution of one or more domains of the reference CasX with one or more domains from a different CasX. In some embodiments, the insertion comprises insertion of some or all of the domains from a different CasX protein. The mutation can occur in any one or more domains of the reference CasX protein, and can include, for example, deletion of some or all of one or more domains, or substitution, deletion, or insertion of one or more amino acids in any domain of the reference CasX protein. The domains of the CasX protein include the non-target strand binding (NTSB) domain, the target strand loading (TSL) domain, the helical I domain, the helical II domain, the oligonucleotide binding domain (OBD), and the RuvC DNA cleavage domain. Any change in the amino acid sequence of the reference CasX protein that results in an improvement in a feature of the CasX protein is considered a CasX variant protein of the present disclosure. For example, the CasX variant can comprise one or more amino acid substitutions, insertions, deletions, or exchange of domains, or any combination thereof, relative to the reference CasX protein sequence.
[0273] In some embodiments, the CasX variant protein comprises at least one modification in each of at least two domains of the reference CasX protein, including the sequences of SEQ ID NOs: 1-3. In some embodiments, the CasX variant protein comprises at least one modification in at least 2 domains, at least 3 domains, at least 4 domains, or at least 5 domains of the reference CasX protein. In some embodiments, the CasX variant protein comprises two or more modifications in at least one domain of the reference CasX protein. In some embodiments, the CasX variant protein comprises at least two modifications in at least one domain of the reference CasX protein, at least three modifications in at least one domain of the reference CasX protein, or at least four modifications in at least one domain of the reference CasX protein. In some embodiments, wherein the CasX variant comprises two or more modifications as compared to the reference CasX protein, each modification is made in a domain independently selected from the group consisting of NTSBD, TSLD, helical I domain, helical II domain, OBD, and RuvC DNA cleavage domain.
[0274] In some embodiments, the at least one modification of the CasX variant protein comprises a deletion of at least a portion of one domain of a reference CasX protein. In some embodiments, the deletion is in the NTSBD, TSLD, Helical I domain, Helical II domain, OBD, or RuvC DNA cleavage domain.
[0275] Mutagenesis methods suitable for generating CasX variant proteins of the application can include, for example, deep mutational evolution (DME), deep mutational scanning (DMS), error-prone PCR, cassette mutagenesis, random mutagenesis, staggered extension PCR, gene shuffling, or domain swapping. Exemplary methods of generating CasX variants with improved characteristics are provided in the Examples below. In some embodiments, CasX variants are designed, for example, by selecting one or more desired mutations in a reference CasX. In certain embodiments, the activity of a reference CasX protein is used as a benchmark to compare the activity of one or more CasX variants, thereby measuring the functional improvement of the CasX variant. Exemplary improvements of CasX variants include, but are not limited to, improved variant folding, improved binding affinity for gNAs, improved binding affinity for target DNA, improved ability to edit with a larger range of PAM sequences and / or bind target DNA, improved target DNA unwinding, increased editing activity, improved editing efficiency, improved editing specificity, increased nuclease activity, increased target strand loading for double-strand cleavage, reduced target strand loading for single-strand cleavage, reduced off-target cleavage, improved binding of non-target strands of DNA, improved protein stability, improved CasX:gNA (RNP) complex stability, improved protein solubility, improved CasX:gNA (RNP) complex solubility, improved protein yield, improved protein expression, and improved melting characteristics, as described more fully below.
[0276] In some embodiments of the CasX variants described herein, the at least one modification comprises: (a) a substitution of 1 to 100 contiguous or non-contiguous amino acids in the CasX variant; (b) a deletion of 1 to 100 contiguous or non-contiguous amino acids in the CasX variant; (c) an insertion of 1 to 100 contiguous or non-contiguous amino acids in the CasX; or (d) any combination of (a)-(c). In some embodiments, the at least one modification comprises: (a) a substitution of 5-10 contiguous or non-contiguous amino acids in the CasX variant; (b) a deletion of 1-5 contiguous or non-contiguous amino acids in the CasX variant; (c) an insertion of 1-5 contiguous or non-contiguous amino acids in the CasX; or (d) any combination of (a)-(c).
[0277] In some embodiments, the CasX variant protein comprises or consists of a sequence having at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 20, at least 30, at least 40, or at least 50 mutations relative to the sequence of SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 3. The mutations can be insertions, deletions, amino acid substitutions, or any combination thereof.
[0278] In some embodiments, the CasX variant protein comprises at least one amino acid substitution in at least one domain of a reference CasX protein. In some embodiments, the CasX variant protein comprises 1-4 amino acid substitutions, 1-10 amino acid substitutions, 1-20 amino acid substitutions, 1-30 amino acid substitutions, 1-40 amino acid substitutions, 1-50 amino acid substitutions, 1-60 amino acid substitutions, 1-70 amino acid substitutions, 1-80 amino acid substitutions, 1-90 amino acid substitutions, 1-100 amino acid substitutions, 2-10 amino acid substitutions, 2-20 amino acid substitutions, 2-30 amino acid substitutions, 3-10 amino acid substitutions, 3-20 amino acid substitutions, 3-30 amino acid substitutions, 4-10 amino acid substitutions, 4-20 amino acid substitutions, 3-300 amino acid substitutions, 5-10 amino acid substitutions, 5-20 amino acid substitutions, 5-30 amino acid substitutions, 10-50 amino acid substitutions, or 20-50 amino acid substitutions relative to the reference CasX protein. In some embodiments, the CasX variant protein comprises at least about 100 amino acid substitutions relative to the reference CasX protein. In some embodiments, the CasX variant protein comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid substitutions relative to the reference CasX protein. In some embodiments, the CasX variant protein comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid substitutions in a single domain relative to the reference CasX protein. In some embodiments, the amino acid substitutions are conservative substitutions. In other embodiments, the substitutions are non-conservative; for example, a polar amino acid is substituted for a non-polar amino acid, or vice versa.
[0279] In some embodiments, the CasX variant protein comprises 1 amino acid substitution, 2-3 contiguous amino acid substitutions, 2-4 contiguous amino acid substitutions, 2-5 contiguous amino acid substitutions, 2-6 contiguous amino acid substitutions, 2-7 contiguous amino acid substitutions, 2-8 contiguous amino acid substitutions, 2-9 contiguous amino acid substitutions, 2-10 contiguous amino acid substitutions, 2-20 contiguous amino acid substitutions, 2-30 contiguous amino acid substitutions, 2-40 contiguous amino acid substitutions, 2-50 contiguous amino acid substitutions, 2-60 contiguous amino acid substitutions, 2-70 contiguous amino acid substitutions, 2-80 contiguous amino acid substitutions, 2-90 contiguous amino acid substitutions, 2-100 contiguous amino acid substitutions, 3-10 contiguous amino acid substitutions, 3-20 contiguous amino acid substitutions, 3-30 contiguous amino acid substitutions, 4-10 contiguous amino acid substitutions, 4-20 contiguous amino acid substitutions, 3-300 contiguous amino acid substitutions, 5-10 contiguous amino acid substitutions, 5-20 contiguous amino acid substitutions, 5-30 contiguous amino acid substitutions, 10-50 contiguous amino acid substitutions, or 20-50 contiguous amino acid substitutions, relative to the reference CasX protein. In some embodiments, the CasX variant protein comprises 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 contiguous amino acid substitutions. In some embodiments, the CasX variant protein comprises at least about 100 contiguous amino acid substitutions. As used herein, “contiguous amino acids” refers to amino acids that are contiguous in the primary sequence of a polypeptide.
[0280] In some embodiments, the CasX variant protein comprises two or more substitutions relative to the reference CasX protein, and the two or more substitutions are not in contiguous amino acids of the reference CasX sequence. For example, a first substitution can be in a first domain of the reference CasX protein, and a second substitution can be in a second domain of the reference CasX protein. In some embodiments, the CasX variant protein comprises 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 non-contiguous substitutions relative to the reference CasX protein. In some embodiments, the CasX variant protein comprises at least 20 non-contiguous substitutions relative to the reference CasX protein. Each non-contiguous substitution can be of any length of amino acids described herein, e.g., 1-4 amino acids, 1-10 amino acids, etc. In some embodiments, the two or more substitutions relative to the reference CasX protein are not the same length, e.g., one substitution is one amino acid and the second substitution is three amino acids. In some embodiments, the two or more substitutions relative to the reference CasX protein are the same length, e.g., two substitutions are two contiguous amino acids in length.
[0281] Any amino acid can be substituted for any other amino acid in the substitutions described herein. The substitution can be a conservative substitution (e.g., a basic amino acid substituted for another basic amino acid). The substitution can be a non-conservative substitution (e.g., a basic amino acid substituted for an acidic amino acid, or vice versa). For example, a proline in a CasX protein can be substituted for any one of the following to produce a CasX variant protein of the application: arginine, histidine, lysine, aspartic acid, glutamic acid, serine, threonine, asparagine, glutamine, cysteine, glycine, alanine, isoleucine, leucine, methionine, phenylalanine, tryptophan, tyrosine, or valine.
[0282] In some embodiments, the CasX variant protein comprises at least one amino acid deletion relative to a reference CasX protein. In some embodiments, the CasX variant protein comprises a deletion of 1-4 amino acids, 1-10 amino acids, 1-20 amino acids, 1-30 amino acids, 1-40 amino acids, 1-50 amino acids, 1-60 amino acids, 1-70 amino acids, 1-80 amino acids, 1-90 amino acids, 1-100 amino acids, 2-10 amino acids, 2-20 amino acids, 2-30 amino acids, 3-10 amino acids, 3-20 amino acids, 3-30 amino acids, 4-10 amino acids, 4-20 amino acids, 3-300 amino acids, 5-10 amino acids, 5-20 amino acids, 5-30 amino acids, 10-50 amino acids, or 20-50 amino acids relative to the reference CasX protein. In some embodiments, the CasX variant comprises a deletion of at least about 100 contiguous amino acids relative to the reference CasX protein. In some embodiments, the CasX variant protein comprises a deletion of at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, or 100 contiguous amino acids relative to the reference CasX protein. In some embodiments, the CasX variant protein comprises a deletion of 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 contiguous amino acids.
[0283] In some embodiments, the CasX variant protein comprises two or more deletions relative to a reference CasX protein, and the two or more deletions are not contiguous amino acids. For example, a first deletion can be in a first domain of the reference CasX protein, and a second deletion can be in a second domain of the reference CasX protein. In some embodiments, the CasX variant protein comprises 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 non-contiguous deletions relative to the reference CasX protein. In some embodiments, the CasX variant protein comprises at least 20 non-contiguous deletions relative to the reference CasX protein. Each non-contiguous deletion can be of any length of amino acids described herein, e.g., 1-4 amino acids, 1-10 amino acids, etc.
[0284] In some embodiments, the CasX variant protein comprises at least one amino acid insertion. In some embodiments, the CasX variant protein comprises an insertion of 1 amino acid, 2-3 contiguous amino acids, 2-4 contiguous amino acids, 2-5 contiguous amino acids, 2-6 contiguous amino acids, 2-7 contiguous amino acids, 2-8 contiguous amino acids, 2-9 contiguous amino acids, 2-10 contiguous amino acids, 2-20 contiguous amino acids, 2-30 contiguous amino acids, 2-40 contiguous amino acids, 2-50 contiguous amino acids, 2-60 contiguous amino acids, 2-70 contiguous amino acids, 2-80 contiguous amino acids, 2-90 contiguous amino acids, 2-100 contiguous amino acids, 3-10 contiguous amino acids, 3-20 contiguous amino acids, 3-30 contiguous amino acids, 4-10 contiguous amino acids, 4-20 contiguous amino acids, 3-300 contiguous amino acids, 5-10 contiguous amino acids, 5-20 contiguous amino acids, 5-30 contiguous amino acids, 10-50 contiguous amino acids, or 20-50 contiguous amino acids relative to the reference CasX protein. In some embodiments, the CasX variant protein comprises an insertion of 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 contiguous amino acids. In some embodiments, the CasX variant protein comprises an insertion of at least about 100 contiguous amino acids.
[0285] In some embodiments, the CasX variant protein comprises two or more insertions relative to the reference CasX protein, and the two or more insertions are not contiguous amino acids of the sequence. For example, a first insertion can be in a first domain of the reference CasX protein, and a second insertion can be in a second domain of the reference CasX protein. In some embodiments, the CasX variant protein comprises 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 non-contiguous insertions relative to the reference CasX protein. In some embodiments, the CasX variant protein comprises at least 10 to about 20 or more non-contiguous insertions relative to the reference CasX protein. Each non-contiguous insertion can be of any length of amino acids described herein, e.g., 1-4 amino acids, 1-10 amino acids, etc.
[0286] Any amino acid or combination of amino acids can be inserted as described herein. For example, proline, arginine, histidine, lysine, aspartic acid, glutamic acid, serine, threonine, asparagine, glutamine, cysteine, glycine, alanine, isoleucine, leucine, methionine, phenylalanine, tryptophan, tyrosine, or valine, or any combination thereof, can be inserted into the reference CasX protein of the application to produce a CasX variant protein.
[0287] Any permutation of the substitution, insertion, and deletion embodiments described herein can be combined to produce a CasX variant protein of the application. For example, a CasX variant protein can comprise at least one substitution and at least one deletion relative to the reference CasX protein sequence, at least one substitution and at least one insertion relative to the reference CasX protein sequence, at least one insertion and at least one deletion relative to the reference CasX protein sequence, or at least one substitution, one insertion, and one deletion relative to the reference CasX protein sequence.
[0288] In some embodiments, the CasX variant protein has at least about 60% sequence identity, at least 70% identity, at least 80% identity, at least 85% identity, at least 86% identity, at least 87% identity, at least 88% identity, at least 89% identity, at least 90% identity, at least 91% identity, at least 92% identity, at least 93% identity, at least 94% identity, at least 95% identity, at least 96% identity, at least 97% identity, at least 98% identity, at least 99% identity, at least 99.5% identity, at least 99.6% identity, at least 99.7% identity, at least 99.8% identity, or at least 99.9% identity to one of SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 3.
[0289] In some embodiments, the CasX variant protein has at least about 60% sequence similarity to SEQ ID NO:2 or a portion thereof. In some embodiments, the CasX variant protein comprises the substitution of Y789T in SEQ ID NO:2, the deletion of P793 in SEQ ID NO:2, the substitution of Y789D in SEQ ID NO:2, the substitution of T72S in SEQ ID NO:2, the substitution of I546V in SEQ ID NO:2, the substitution of E552A in SEQ ID NO:2, the substitution of A636D in SEQ ID NO:2, the substitution of F536S in SEQ ID NO:2, the substitution of A708K in SEQ ID NO:2, the substitution of Y797L in SEQ ID NO:2, the substitution of L792G in SEQ ID NO:2, the substitution of A739V in SEQ ID NO:2, the substitution of G791M in SEQ ID NO:2, the insertion of A at position 661 in SEQ ID NO:2, the substitution of A788W in SEQ ID NO:2, the substitution of K390R in SEQ ID NO:2, the substitution of A751S in SEQ ID NO:2, and the substitution of SEQ ID NO:2. Substitution of E385A in SEQ ID NO:2, insertion of P at position 696 in SEQ ID NO:2, insertion of M at position 773 in SEQ ID NO:2, substitution of G695H in SEQ ID NO:2, insertion of AS at position 793 in SEQ ID NO:2, insertion of AS at position 795 in SEQ ID NO:2, substitution of C477R in SEQ ID NO:2, substitution of C477K in SEQ ID NO:2, substitution of C479A in SEQ ID NO:2, substitution of C479L in SEQ ID NO:2, substitution of I55F in SEQ ID NO:2, substitution of K210R in SEQ ID NO:2, substitution of C233S in SEQ ID NO:2, substitution of D231N in SEQ ID NO:2, substitution of Q338E in SEQ ID NO:2, substitution of Q338R in SEQ ID NO:2, substitution of L379R in SEQ ID NO:2, SEQ ID Substitution of K390R in SEQ ID NO:2, substitution of L481Q in SEQ ID NO:2, substitution of F495S in SEQ ID NO:2, substitution of D600N in SEQ ID NO:2, substitution of T886K in SEQ ID NO:2, substitution of A739V in SEQ ID NO:2, substitution of K460N in SEQ ID NO:2, substitution of I199F in SEQ ID NO:2, substitution of G492P in SEQ ID NO:2, substitution of T153I in SEQ ID NO:2, SEQsubstitution of E121D of SEQ ID NO: 2, substitution of S270W of SEQ ID NO: 2, substitution of E712Q of SEQ ID NO: 2, substitution of K942Q of SEQ ID NO: 2, substitution of E552K of SEQ ID NO: 2, substitution of K25Q of SEQ ID NO: 2, substitution of N47D of SEQ ID NO: 2, insertion of T at position 696 of SEQ ID NO: 2, substitution of L685I of SEQ ID NO: 2, substitution of N880D of SEQ ID NO: 2, substitution of Q102R of SEQ ID NO: 2, substitution of M734K of SEQ ID NO: 2, substitution of A724S of SEQ ID NO: 2, substitution of T704K of SEQ ID NO: 2, substitution of P224K of SEQ ID NO: 2, substitution of K25R of SEQ ID NO: 2, substitution of M29E of SEQ ID NO: 2, substitution of H152D of SEQ ID NO: 2, substitution of S219R of SEQ ID NO: 2, substitution of E475K of SEQ ID NO: 2, substitution of G226R of SEQ ID NO: 2, substitution of A377K of SEQ ID NO: 2, substitution of E480K of SEQ ID NO: 2, substitution of K416E of SEQ ID NO: 2, substitution of H164R of SEQ ID NO: 2, substitution of K767R of SEQ ID NO: 2, substitution of I7F of SEQ ID NO: 2, substitution of M29R of SEQ ID NO: 2, substitution of H435R of SEQ ID NO: 2, substitution of E385Q of SEQ ID NO: 2, substitution of E385K of SEQ ID NO: 2, substitution of I279F of SEQ ID NO: 2, substitution of D489S of SEQ ID NO: 2, substitution of D732N of SEQ ID NO: 2, substitution of A739T of SEQ ID NO: 2, substitution of W885R of SEQ ID NO: 2, substitution of E53K of SEQ ID NO: 2, substitution of A238T of SEQ ID NO: 2, substitution of P283Q of SEQ ID NO: 2, substitution of E292K of SEQ ID NO: 2, substitution of Q628E of SEQ ID NO: 2, substitution of R388Q of SEQ ID NO: 2, substitution of G791M of SEQ IDsubstitution of E292R of SEQ ID NO: 2, substitution of I303K of SEQ ID NO: 2, substitution of C349E of SEQ ID NO: 2, substitution of E385P of SEQ ID NO: 2, substitution of E386N of SEQ ID NO: 2, substitution of D387K of SEQ ID NO: 2, substitution of L404K of SEQ ID NO: 2, substitution of E466H of SEQ ID NO: 2, substitution of C477Q of SEQ ID NO: 2, substitution of C477H of SEQ ID NO: 2, substitution of C479A of SEQ ID NO: 2, substitution of D659H of SEQ ID NO: 2, substitution of T806V of SEQ ID NO: 2, substitution of K808S of SEQ ID NO: 2, insertion of AS at position 797 of SEQ ID NO: 2, substitution of V959M of SEQ ID NO: 2, substitution of K975Q of SEQ ID NO: 2, substitution of W974G of SEQ ID NO: 2, substitution of A708Q of SEQ ID NO: 2, substitution of V711K of SEQ ID NO: 2, substitution of D733T of SEQ ID NO: 2, substitution of L742W of SEQ ID NO: 2, substitution of V747K of SEQ ID NO: 2, substitution of F755M of SEQ ID NO: 2, substitution of M771A of SEQ ID NO: 2, substitution of K775E of SEQ ID NO: 2, substitution of K775Q of SEQ ID NO: 2, substitution of K775R of SEQ ID NO: 2, substitution of K775T of SEQ ID NO: 2, substitution of K775V of SEQ ID NO: 2, substitution of K775Y of SEQ ID NO: 2, substitution of K775W of SEQ ID NO: 2, substitution of K775H of SEQ ID NO: 2, substitution of K775P of SEQ ID NO: 2, substitution of K775L of SEQ ID NO: 2, substitution of K775I of SEQ ID NO: 2, substitution of K775D of SEQ ID NO: 2, substitution of K775N of SEQ ID NO: 2, substitution of K775A of SEQ ID NO: 2, substitution of K775G of SEQ ID NO: 2, substitution of K775S of SEQ ID NO: 2, substitution of K775Q of SEQ ID NO: 2, substitution of K775E of SEQ ID NO: 2, substitution of K775R of SEQ ID NO: 2, substitution of K775T of SEQ ID NO: 2, substitution of K775V of SEQ ID NO: 2, substitution of K775Y of SEQ ID NO: 2, substitution of K775W of SEQ ID NO: 2, substitution of K775H of SEQ ID NO: 2, substitution of K775P of SEQ ID NO: 2, substitution of K775L of SEQ ID NO: 2, substitution of K775I of SEQ ID NO: 2, substitution of K775D of SEQ ID NO: 2, substitution of K775N of SEQ ID NO: 2, substitution of K775A of SEQ ID NO: 2, substitution of K775G of SEQ ID NO: 2, substitution of K775S of SEQ ID NO: 2, substitution of K775Q of SEQ ID NO: 2, substitution of K775E of SEQ ID NO: 2, substitution of K775R of SEQ ID NO: 2, substitution of K775T of SEQ ID NO: 2, substitution of K775V of SEQ ID NO: 2, substitution of K775Y of SEQ ID NO: 2, substitution of K775W of SEQ ID NO: 2, substitution of K775H of SEQ ID NO: 2, substitution of K775P of SEQ ID NO: 2, substitution of K775L of SEQ ID NO: 2, substitution of K775I of SEQ ID NO: 2, substitution of K775D of SEQ ID NO: 2, substitution of K775N of SEQ ID NO: 2, substitution of K775A of SEQ ID NO: 2, substitution of K775G of SEQ ID NO: 2, substitution of K775S of SEQ ID NO: 2, substitution of K775Q of SEQ ID NO: 2, substitution of K775E of SEQ ID NO: 2, substitution of K775R of SEQ ID NO: 2, substitution of K775T of SEQ ID NO: 2, substitution of K775V of SEQ ID NO: 2, substitution of K775Y of SEQ ID NO: 2, substitution of K775W of SEQ ID NO: 2, substitution of K775H of SEQ ID NO: 2, substitution of K775P of SEQ ID NO: 2, substitution of K775L of SEQ ID NO: 2, substitution of K775I of SEQ ID NO: 2, substitution of K775D of SEQ ID NO: 2, substitution of K775N of SEQ ID NO: 2, substitution of K775A of SEQ ID NO: 2, substitution of K775G of SEQ ID NO: 2,a substitution of S603G of SEQ ID NO: 2, a substitution of N737S of SEQ ID NO: 2, a substitution of L307K of SEQ ID NO: 2, a substitution of I658V of SEQ ID NO: 2, an insertion of PT at position 688 of SEQ ID NO: 2, an insertion of SA at position 794 of SEQ ID NO: 2, a substitution of S877R of SEQ ID NO: 2, a substitution of N580T of SEQ ID NO: 2, a substitution of V335G of SEQ ID NO: 2, a substitution of T620S of SEQ ID NO: 2, a substitution of W345G of SEQ ID NO: 2, a substitution of T280S of SEQ ID NO: 2, a substitution of L406P of SEQ ID NO: 2, a substitution of A612D of SEQ ID NO: 2, a substitution of A751S of SEQ ID NO: 2, a substitution of E386R of SEQ ID NO: 2, a substitution of V351M of SEQ ID NO: 2, a substitution of K210N of SEQ ID NO: 2, a substitution of D40A of SEQ ID NO: 2, a substitution of E773G of SEQ ID NO: 2, a substitution of H207L of SEQ ID NO: 2, a substitution of T62A of SEQ ID NO: 2, a substitution of T287P of SEQ ID NO: 2, a substitution of T832A of SEQ ID NO: 2, a substitution of A893S of SEQ ID NO: 2, an insertion of V at position 14 of SEQ ID NO: 2, an insertion of AG at position 13 of SEQ ID NO: 2, a substitution of R11V of SEQ ID NO: 2, a substitution of R12N of SEQ ID NO: 2, a substitution of R13H of SEQ ID NO: 2, an insertion of Y at position 13 of SEQ ID NO: 2, a substitution of R12L of SEQ ID NO: 2, an insertion of Y at position 13 of SEQ ID NO: 2, a substitution of R12L of SEQ ID NO: 2, an insertion of Y at position 13 of SEQ ID NO: 2, a substitution of R12L of SEQ ID NO: 2, an insertion of Y at position 13 of SEQ ID NO: 2, a substitution of R12L of SEQ ID NO: 2, an insertion of Y at position 13 of SEQ ID NO: 2, a substitution of R12L of SEQ ID NO: 2, an insertion of Y at position 13 of SEQ ID NO: 2, a substitution of R12L of SEQ ID NO: 2, an insertion of Y at position 13 of SEQ ID NO: 2, a substitution of R12L of SEQ ID NO: 2, an insertion of Y at position 13 of SEQ ID NO: 2, a substitution of R12L of SEQ ID NO: 2, an insertion of Y at position 13 of SEQ ID NO: 2, a substitution of R12L of SEQ ID NO: 2, an insertion of Y at position 13 of SEQ ID NO: 2, a substitution of R12L of SEQ ID NO: 2, an insertion of Y at position 13 of SEQ ID NO: 2, a substitution of R12L of SEQ ID NO: 2, an insertion of Y at position 13 of SEQ ID NO: 2, a substitution of R12L of SEQ ID NO: 2, an insertion of Y at position 13 of SEQ ID NO: 2, a substitution of R12L of SEQ ID NO: 2, an insertion of Y at position 13 of SEQ ID NO: 2, a substitution of R12L of SEQ ID NO: 2, an insertion of Y at position 13 of SEQ ID NO: 2, a substitution of R12L of SEQ ID NO: 2, an insertion of Y at position 13 of SEQ ID NO: 2, a substitution of R12L of SEQ ID NO: 2, an insertion of Y at position 13 of SEQ ID NO: 2, a substitution of R12L of SEQ ID NO: 2, an insertion of Y at position 13 of SEQ ID NO: 2, a substitution of R12L of SEQ ID NO: 2, an insertion of Y at position 13 of SEQ ID NO: 2, a substitution of R12L of SEQ ID NO: 2, an insertion of Y at position 13 of SEQ ID NO: 2, a substitution of R12L of SEQ ID NO: 2, an insertion of Y at position 13 of SEQ ID NO: 2, a substitution of R12L of SEQ ID NO: 2, an insertion of Y at position 13 of SEQ ID NO: 2, a substitution of R12L of SEQ ID NO: 2, an insertion of Y at position 13 of SEQ ID NO: 2, a substitution of R12L of SEQ ID NO: 2, an insertion of Y at position 13 of SEQ ID NO: 2, a substitution of R12L of SEQ ID NO: 2, an insertion of Y at position 13 of SEQ ID NO: 2, a substitution of R12L of SEQ ID NO: 2, an insertion of Y at position 13 of SEQ ID NO: 2, a substitution of R12L of SEQ ID NO: 2, an insertion of Y at position 13 of SEQ ID NO: 2, a substitution of R12L of SEQ ID NO: 2, an insertion of Y at position 13 of SEQ ID NO: 2, a substitution of R12L of SEQ ID NO: 2, an insertion of Y at position 13 of SEQ ID NO: 2, a substitution of R12L of SEQ ID NO: 2, an insertion of Y at position 13 of SEQ ID NO: 2, a substitution of R12L of SEQ ID NO: 2, an insertion of Y at position 13 of SEQ ID NO: 2, a substitution of R12L of SEQ ID NO: 2, an insertion of Y at position 13 of SEQ ID NO: 2, a substitution of R12L of SEQ ID NO: 2, an insertion of Y at position 13 of SEQ ID NO: 2, a substitution of R12L of SEQ ID NO: 2, an insertion of Y at position 13 of SEQ ID NO: 2, a substitution of R12L of SEQ ID NO: 2, an insertion of Y at position 13 of SEQ ID NO: 2, a substitution of R12L of SEQ ID NO: 2, an insertion of Y at position 13 of SEQ ID NO: 2, a substitution of R12L of SEQ ID NO: 2, an insertion of Y at position 13 of SEQ ID NO: 2, a substitution of R12L of SEQ ID NO: 2, an insertion of Y at position 13 of SEQ ID NO: 2, a substitution of R12L of SEQ ID NO: 2, an insertion of Y at position 13 of SEQ ID NO: 2, a substitution of R12L of SEQ ID NO: 2, an insertion of Y at position 13 of SEQ ID NO: 2, a substitution of R12L of SEQ ID NO: 2, an insertion of Y at position 13 of SEQ ID NO: 2, a substitution of R12L of SEQ ID NO: 2, an insertion of Y at position 13 of SEQ ID NO: 2, a substitution of R12L of SEQ ID NO: 2, an insertion of Y at position 13 of SEQ ID NO: 2, a substitution of R12L of SEQ ID NO: 2, an insertion of Y at position 13 of SEQ ID NO: 2, a substitution of R12L of SEQ ID NO: 2, an insertion of Y at position 13 of SEQ ID NO: 2, a substitution of R12L of SEQ ID NO: 2, an insertion of Y at position 13 of SEQ ID NO: 2, a substitution of R12L of SEQ ID NO: 2, an insertion of Y at position 13 of SEQ ID NO: 2, a substitution of R12L of SEQ ID NO: 2, an insertion of Y at position 13 of SEQ ID NO: 2, a substitution of R12L of SEQ ID NO: 2, an insertion of Y at position 13 of SEQ ID NO: 2, a substitution of R12L of SEQ ID NO: 2, an insertion of Y at position 13 of SEQ ID NO: 2, a substitution of R12L of SEQ ID NO: 2, an insertion of Y at position 13 of SEQ ID NO: 2, a substitution of R12L of SEQ ID NO: 2, an insertion of Y at position 13 of SEQ ID NO: 2, a substitution of R12L of SEQ ID NO: 2, an insertion of Y at position 13 of SEQ ID NO: 2, a substitution of R12L of SEQ ID NO: 2, an insertion of Y at position 13 of SEQ ID NO: 2, a substitution of R12L of SEQ ID NO: 2,an insertion of Q at position 13 of SEQ ID NO: 2, a substitution of V15S of SEQ ID NO: 2, an insertion of D at position 17 of SEQ ID NO: 2, or a combination thereof.
[0290] In some embodiments, the CasX variant comprises at least one modification in the NTSB domain.
[0291] In some embodiments, the CasX variant comprises at least one modification in the TSL domain. In some embodiments, the at least one modification in the TSL domain comprises an amino acid substitution of one or more of amino acids Y857, S890, or S932 of SEQ ID NO: 2.
[0292] In some embodiments, the CasX variant comprises at least one modification in the Helical I domain. In some embodiments, the at least one modification in the Helical I domain comprises an amino acid substitution of one or more of amino acids S219, L249, E259, Q252, E292, L307, or D318 of SEQ ID NO: 2.
[0293] In some embodiments, the CasX variant comprises at least one modification in the Helical II domain. In some embodiments, the at least one modification in the Helical II domain comprises an amino acid substitution of one or more of amino acids D361, L379, E385, E386, D387, F399, L404, R458, C477, or D489 of SEQ ID NO: 2.
[0294] In some embodiments, the CasX variant comprises at least one modification in the OBD domain. In some embodiments, the at least one modification in the OBD comprises an amino acid substitution of one or more of amino acids F536, E552, T620, or I658 of SEQ ID NO: 2.
[0295] In some embodiments, the CasX variant comprises at least one modification in the RuvC DNA cleavage domain. In some embodiments, the at least one modification in the RuvC DNA cleavage domain comprises an amino acid substitution of one or more of amino acids K682, G695, A708, V711, D732, A739, D733, L742, V747, F755, M771, M779, W782, A788, G791, L792, P793, Y797, M799, Q804, S819, or Y857 of SEQ ID NO: 2 or a deletion of amino acid P793.
[0296] In some embodiments, the CasX variant comprises at least one modification selected from one or more of the following, as compared to the reference CasX sequence of SEQ ID NO: 2: (a) an amino acid substitution of L379R; (b) an amino acid substitution of A708K; (c) an amino acid substitution of T620P; (d) an amino acid substitution of E385P; (e) an amino acid substitution of Y857R; (f) an amino acid substitution of I658V; (g) an amino acid substitution of F399L; (h) an amino acid substitution of Q252K; (i) an amino acid substitution of L404K; and (j) an amino acid deletion of P793.
[0297] In some embodiments, the CasX variant protein comprises at least two amino acid changes to the amino acid sequence of a reference CasX protein. The at least two amino acid changes may be substitutions, insertions, or deletions to the amino acid sequence of the reference CasX protein, or any combination thereof. Substitutions, insertions, or deletions may be any substitution, insertion, or deletion in the sequence of the reference CasX protein described herein. In some embodiments, the changes are continuous amino acid changes, discontinuous amino acid changes, or a combination of continuous and discontinuous amino acid changes to the reference CasX protein sequence. In some embodiments, the reference CasX protein is SEQ ID NO:2. In some embodiments, the CasX variant protein comprises at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 30, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, or at least 100 amino acid changes to the reference CasX protein sequence. In some embodiments, the CasX variant protein comprises 1-50, 3-40, 5-30, 5-20, 5-15, 5-10, 10-50, 10-40, 10-30, 10-20, 15-50, 15-40, 15-30, 2-25, 2-24, or 2-22 elements of the reference CasX protein sequence. 2-23, 2-22, 2-21, 2-20, 2-19, 2-18, 2-17, 2-16, 2-15, 2-14, 2-12, 2-11, 2-10, 2-9, 2-8, 2-7, 2-6, 2-5, 2-4, 2-3, 3-25, 3-24, 3-22, 3-2 3, 3-22, 3-21, 3-20, 3-19, 3-18, 3-17, 3-16, 3-15, 3-14, 3-12, 3-11, 3-10, 3-9, 3-8, 3-7, 3-6, 3-5, 3-4, 4-25, 4-24, 4-22, 4-23, 4-22 4-21, 4-20, 4-19, 4-18, 4-17, 4-16, 4-15, 4-14, 4-12, 4-11, 4-10, 4-9, 4-8, 4-7, 4-6, 4-5, 5-25, 5-24, 5-22, 5-23, 5-22, 5-21, 5-205-19, 5-18, 5-17, 5-16, 5-15, 5-14, 5-12, 5-11, 5-10, 5-9, 5-8, 5-7, or 5-6 amino acid changes. In some embodiments, the CasX variant protein comprises 15-20 changes to a reference CasX protein sequence. In some embodiments, the CasX variant protein comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 amino acid changes to a reference protein sequence. In some embodiments, at least two of the amino acid changes to the sequence of the reference CasX variant protein are selected from the group consisting of a substitution of Y789T of SEQ ID NO: 2, a deletion of P793 of SEQ ID NO: 2, a substitution of Y789D of SEQ ID NO: 2, a substitution of T72S of SEQ ID NO: 2, a substitution of I546V of SEQ ID NO: 2, a substitution of E552A of SEQ ID NO: 2, a substitution of A636D of SEQ ID NO: 2, a substitution of F536S of SEQ ID NO: 2, a substitution of A708K of SEQ ID NO: 2, a substitution of Y797L of SEQ ID NO: 2, a substitution of L792G of SEQ ID NO: 2, a substitution of A739V of SEQ ID NO: 2, a substitution of G791M of SEQ ID NO: 2, an insertion of A at position 661 of SEQ ID NO: 2, a substitution of A788W of SEQ ID NO: 2, a substitution of K390R of SEQ ID NO: 2, a substitution of A751S of SEQ ID NO: 2, a substitution of E385A of SEQ ID NO: 2, an insertion of P at position 696 of SEQ ID NO: 2, an insertion of M at position 773 of SEQ ID NO: 2, a substitution of G695H of SEQ ID NO: 2, an insertion of AS at position 793 of SEQ ID NO: 2, an insertion of AS at position 795 of SEQ ID NO: 2, a substitution of C477R of SEQ ID NO: 2, a substitution of C477K of SEQ ID NO: 2, a substitution of C479A of SEQ ID NO: 2, a substitution of C479L of SEQ ID NO: 2, a substitution of I55F of SEQ ID NO: 2, a substitution of K210R of SEQ ID NO: 2, a substitution of C233S of SEQ ID NO: 2, a substitution of D231N of SEQ ID NO: 2, a substitution of Q338E of SEQ ID NO: 2, a substitution of Q338R of SEQ ID NO: 2, a substitution of L379R of SEQ ID NO: 2,substitution of E121D of SEQ ID NO: 2, substitution of S270W of SEQ ID NO: 2, substitution of E712Q of SEQ ID NO: 2, substitution of K942Q of SEQ ID NO: 2, substitution of E552K of SEQ ID NO: 2, substitution of K25Q of SEQ ID NO: 2, substitution of N47D of SEQ ID NO: 2, insertion of T at position 696 of SEQ ID NO: 2, substitution of L685I of SEQ ID NO: 2, substitution of N880D of SEQ ID NO: 2, substitution of Q102R of SEQ ID NO: 2, substitution of M734K of SEQ ID NO: 2, substitution of A724S of SEQ ID NO: 2, substitution of T704K of SEQ ID NO: 2, substitution of P224K of SEQ ID NO: 2, substitution of K25R of SEQ ID NO: 2, substitution of M29E of SEQ ID NO: 2, substitution of H152D of SEQ ID NO: 2, substitution of S219R of SEQ ID NO: 2, substitution of E475K of SEQ ID NO: 2, substitution of G226R of SEQ ID NO: 2, substitution of A377K of SEQ ID NO: 2, substitution of E480K of SEQ ID NO: 2, substitution of K416E of SEQ ID NO: 2, substitution of H164R of SEQ ID NO: 2, substitution of K767R of SEQ ID NO: 2, substitution of I7F of SEQ ID NO: 2, substitution of M29R of SEQ ID NO: 2, substitution of H435R of SEQ ID NO: 2, substitution of E385Q of SEQ ID NO: 2, substitution of E385K of SEQ ID NO: 2, substitution of I279F of SEQ ID NO: 2, substitution of D489S of SEQ ID NO: 2,a substitution of D732N of SEQ ID NO: 2, a substitution of A739T of SEQ ID NO: 2, a substitution of W885R of SEQ ID NO: 2, a substitution of E53K of SEQ ID NO: 2, a substitution of A238T of SEQ ID NO: 2, a substitution of P283Q of SEQ ID NO: 2, a substitution of E292K of SEQ ID NO: 2, a substitution of Q628E of SEQ ID NO: 2, a substitution of R388Q of SEQ ID NO: 2, a substitution of G791M of SEQ ID NO: 2, a substitution of L792K of SEQ ID NO: 2, a substitution of L792E of SEQ ID NO: 2, a substitution of M779N of SEQ ID NO: 2, a substitution of G27D of SEQ ID NO: 2, a substitution of K955R of SEQ ID NO: 2, a substitution of S867R of SEQ ID NO: 2, a substitution of R693I of SEQ ID NO: 2, a substitution of F189Y of SEQ ID NO: 2, a substitution of V635M of SEQ ID NO: 2, a substitution of F399L of SEQ ID NO: 2, a substitution of E498K of SEQ ID NO: 2, a substitution of E386R of SEQ ID NO: 2, a substitution of V254G of SEQ ID NO: 2, a substitution of P793S of SEQ ID NO: 2, a substitution of K188E of SEQ ID NO: 2, a substitution of Q T945KI of SEQ ID NO: 2, a substitution of T620P of SEQ ID NO: 2, a substitution of T946P of SEQ ID NO: 2, a substitution of TT949PP of SEQ ID NO: 2, a substitution of N952T of SEQ ID NO: 2, a substitution of K682E of SEQ ID NO: 2, a substitution of K975R of SEQ ID NO: 2, a substitution of L212P of SEQ ID NO: 2, a substitution of E292R of SEQ ID NO: 2, a substitution of I303K of SEQ ID NO: 2, a substitution of C349E of SEQ ID NO: 2, a substitution of E385P of SEQ ID NO: 2, a substitution of E386N of SEQ ID NO: 2, a substitution of D387K of SEQ ID NO: 2, a substitution of L404K of SEQ ID NO: 2, a substitution of E466H of SEQ ID NO: 2, a substitution of C477Q of SEQ ID NO: 2, a substitution of C477H of SEQ ID NO: 2, a substitution of C479A of SEQ ID NO: 2, a substitution of D659H of SEQ ID NO: 2, a substitution of T806V of SEQ ID NO: 2, a substitution of K808S of SEQ ID NO: 2,an insertion of PT at position 688 of SEQ ID NO: 2, an insertion of SA at position 794 of SEQ ID NO: 2, a substitution of S877R of SEQ ID NO: 2, a substitution of N580T of SEQ ID NO: 2, a substitution of V335G of SEQ ID NO: 2, a substitution of T620S of SEQ ID NO: 2, a substitution of W345G of SEQ ID NO: 2, a substitution of T280S of SEQ ID NO: 2, a substitution of L406P of SEQ ID NO: 2, a substitution of A612D of SEQ ID NO: 2, a substitution of A751S of SEQ ID NO: 2, a substitution of E386R of SEQ ID NO: 2, a substitution of V351M of SEQ ID NO: 2, a substitution of K210N of SEQ ID NO: 2, a substitution of D40A of SEQ ID NO: 2, a substitution of E773G of SEQ ID NO: 2, a substitution of H207L of SEQ ID NO: 2,a substitution of T62A of SEQ ID NO: 2, a substitution of T287P of SEQ ID NO: 2, a substitution of T832A of SEQ ID NO: 2, a substitution of A893S of SEQ ID NO: 2, an insertion of V at position 14 of SEQ ID NO: 2, an insertion of AG at position 13 of SEQ ID NO: 2, a substitution of R11V of SEQ ID NO: 2, a substitution of R12N of SEQ ID NO: 2, a substitution of R13H of SEQ ID NO: 2, an insertion of Y at position 13 of SEQ ID NO: 2, a substitution of R12L of SEQ ID NO: 2, an insertion of Q at position 13 of SEQ ID NO: 2, a substitution of V15S of SEQ ID NO: 2, an insertion of D at position 17 of SEQ ID NO: 2. In some embodiments, the at least two amino acid changes to the reference CasX protein are selected from the amino acid changes disclosed in the sequences of Table 3. In some embodiments, the CasX variant comprises any combination of the preceding embodiments of this paragraph.
[0298] In some embodiments, the CasX variant protein comprises more than one substitution, insertion, and / or deletion to a reference CasX protein amino acid sequence. In some embodiments, the reference CasX protein comprises or consists essentially of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises a substitution of S794R and a substitution of Y797L of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises a substitution of K416E and a substitution of A708K of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises a substitution of A708K and a deletion of P793 of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises a deletion of P793 and an insertion of AS at position 795 of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises a substitution of Q367K and a substitution of I425S of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises a substitution of A708K, a deletion of P at position 793, and a substitution of A793V of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises a substitution of Q338R and a substitution of A339E of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises a substitution of Q338R and a substitution of A339K of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises a substitution of S507G and a substitution of G508R of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises a substitution of L379R, a substitution of A708K, and a deletion of P at position 793 of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises a substitution of C477K, a substitution of A708K, and a deletion of P at position 793 of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises a substitution of L379R, a substitution of C477K, a substitution of A708K, and a deletion of P at position 793 of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises a substitution of L379R, a substitution of A708K, a deletion of P at position 793, and a substitution of A739V of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises a substitution of C477K, a substitution of A708K, a deletion of P at position 793, and a substitution of A739V of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises a substitution of L379R, a substitution of C477K, a substitution of A708K, a deletion of P at position 793, and a substitution of A739V of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises a substitution of L379R, a substitution of A708K, a deletion of P at position 793, and a substitution of M779N of SEQ ID NO: 2.In some embodiments, the CasX variant protein comprises a substitution of L379R of SEQ ID NO: 2, a substitution of A708K, a deletion of P at position 793, and a substitution of M771N. In some embodiments, the CasX variant protein comprises a substitution of L379R of SEQ ID NO: 2, a substitution of 708K, a deletion of P at position 793, and a substitution of D489S. In some embodiments, the CasX variant protein comprises a substitution of L379R of SEQ ID NO: 2, a substitution of A708K, a deletion of P at position 793, and a substitution of A739T. In some embodiments, the CasX variant protein comprises a substitution of L379R of SEQ ID NO: 2, a substitution of A708K, a deletion of P at position 793, and a substitution of D732N. In some embodiments, the CasX variant protein comprises a substitution of L379R of SEQ ID NO: 2, a substitution of A708K, a deletion of P at position 793, and a substitution of G791M. In some embodiments, the CasX variant protein comprises a substitution of L379R of SEQ ID NO: 2, a substitution of 708K, a deletion of P at position 793, and a substitution of Y797L. In some embodiments, the CasX variant protein comprises a substitution of L379R of SEQ ID NO: 2, a substitution of C477K, a substitution of A708K, a deletion of P at position 793, and a substitution of M779N. In some embodiments, the CasX variant protein comprises a substitution of L379R of SEQ ID NO: 2, a substitution of C477K, a substitution of A708K, a deletion of P at position 793, and a substitution of M771N. In some embodiments, the CasX variant protein comprises a substitution of L379R of SEQ ID NO: 2, a substitution of C477K, a substitution of A708K, a deletion of P at position 793, and a substitution of D489S. In some embodiments, the CasX variant protein comprises a substitution of L379R of SEQ ID NO: 2, a substitution of C477K, a substitution of A708K, a deletion of P at position 793, and a substitution of A739T. In some embodiments, the CasX variant protein comprises a substitution of L379R of SEQ ID NO: 2, a substitution of C477K, a substitution of A708K, a deletion of P at position 793, and a substitution of D732N. In some embodiments, the CasX variant protein comprises a substitution of L379R of SEQ ID NO: 2, a substitution of C477K, a substitution of A708K, a deletion of P at position 793, and a substitution of G791M. In some embodiments, the CasX variant protein comprises a substitution of L379R of SEQ ID NO: 2, a substitution of C477K, a substitution of A708K, a deletion of P at position 793, and a substitution of Y797L.In some embodiments, the CasX variant protein comprises a substitution of L379R, a substitution of C477K, a substitution of A708K, a deletion of P at position 793, and a substitution of T620P of SEQ ID NO:2. In some embodiments, the CasX variant protein comprises a substitution of A708K, a deletion of P at position 793, and a substitution of E386S of SEQ ID NO:2. In some embodiments, the CasX variant protein comprises a substitution of E386R, a substitution of F399L, and a deletion of P at position 793 of SEQ ID NO:2. In some embodiments, the CasX variant protein comprises a substitution of R581I and A739V of SEQ ID NO:2. In some embodiments, the CasX variant comprises any combination of the foregoing embodiments of this paragraph.
[0299] In some embodiments, the CasX variant protein comprises more than one substitution, insertion, and / or deletion to a reference CasX protein amino acid sequence. In some embodiments, the CasX variant protein comprises a substitution of A708K, a deletion of P at position 793, and a substitution of A739V of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises a substitution of L379R, a substitution of A708K, and a deletion of P at position 793 of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises a substitution of C477K, a substitution of A708K, and a deletion of P at position 793 of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises a substitution of L379R, a substitution of C477K, a substitution of A708K, and a deletion of P at position 793 of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises a substitution of L379R, a substitution of A708K, a deletion of P at position 793, and a substitution of A739V of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises a substitution of C477K, a substitution of A708K, a deletion of P at position 793, and a substitution of A739 of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises a substitution of L379R, a substitution of C477K, a substitution of A708K, a deletion of P at position 793, and a substitution of A739V of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises a substitution of L379R, a substitution of C477K, a substitution of A708K, a deletion of P at position 793, and a substitution of T620P of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises a substitution of M771A of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises a substitution of L379R, a substitution of A708K, a deletion of P at position 793, and a substitution of D732N of SEQ ID NO: 2. In some embodiments, the CasX variant comprises any combination of the foregoing embodiments of this paragraph.
[0300] In some embodiments, the CasX variant protein comprises a substitution of W782Q of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises a substitution of M771Q of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises a substitution of R458I and a substitution of A739V of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises a substitution of L379R, a substitution of A708K, a deletion of P at position 793, and a substitution of M771N of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises a substitution of L379R, a substitution of A708K, a deletion of P at position 793, and a substitution of A739T of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises a substitution of L379R, a substitution of C477K, a substitution of A708K, a deletion of P at position 793, and a substitution of D489S of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises a substitution of L379R, a substitution of C477K, a substitution of A708K, a deletion of P at position 793, and a substitution of D732N of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises a substitution of V711K of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises a substitution of L379R, a substitution of C477K, a substitution of A708K, a deletion of P at position 793, and a substitution of Y797L of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises a substitution of L379R, a substitution of A708K, and a deletion of P at position 793 of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises a substitution of L379R, a substitution of C477K, a substitution of A708K, a deletion of P at position 793, and a substitution of M771N of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises a substitution of A708K, a substitution of P at position 793, and a substitution of E386S of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises a substitution of L379R, a substitution of C477K, a substitution of A708K, and a deletion of P at position 793 of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises a substitution of L792D of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises a substitution of G791F of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises a substitution of A708K, a deletion of P at position 793, and a substitution of A739V of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises a substitution of L379R, a substitution of A708K, a deletion of P at position 793, and a substitution of A739V of SEQ ID NO: 2.In some embodiments, the CasX variant protein comprises a substitution of C477K of SEQ ID NO: 2, a substitution of A708K, and a substitution of P at position 793. In some embodiments, the CasX variant protein comprises a substitution of L249I of SEQ ID NO: 2 and a substitution of M771N. In some embodiments, the CasX variant protein comprises a substitution of V747K of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises a substitution of L379R of SEQ ID NO: 2, a substitution of C477, a substitution of A708K, a deletion of P at position 793, and a substitution of M779N. In some embodiments, the CasX variant protein comprises a substitution of F755M. In some embodiments, the CasX variant comprises any combination of the foregoing embodiments of this paragraph.
[0301] In some embodiments, the CasX variant protein comprises at least one modification compared to the reference CasX sequence of SEQ ID NO: 2, wherein the at least one modification is selected from one or more of: an amino acid substitution of L379R; an amino acid substitution of A708K; an amino acid substitution of T620P; an amino acid substitution of E385P; an amino acid substitution of Y857R; an amino acid substitution of I658V; an amino acid substitution of F399L; an amino acid substitution of Q252K; an amino acid substitution of L404K; and an amino acid deletion of [P793]. In other embodiments, the CasX variant protein comprises any combination of the foregoing substitutions or deletions compared to the reference CasX sequence of SEQ ID NO: 2. In other embodiments, in addition to the foregoing substitutions or deletions, the CasX variant protein can further comprise substitutions from the NTSB and / or helical 1b domain of the reference CasX of SEQ ID NO: 1.
[0302] In some embodiments, the CasX variant comprises any one of SEQ ID NOs: 247-337, 3301-3493, 3498-3501, 3505-3520, 3540-3549, and 4412-4415. In some embodiments, the CasX variant comprises any one of SEQ ID NOs: 247-337, 3498-3501, 3505-3520, 3540-3549, and 4412-4415. In some embodiments, the CasX variant comprises any one of SEQ ID NOs: 3498-3501, 3505-3520, and 3540-3549.
[0303] In some embodiments, the CasX variant comprises one or more modifications to any one of SEQ ID NOs: 247-337, 3301-3493, 3498-3501, 3505-3520, 3540-3549, and 4412-4415. In some embodiments, the CasX variant comprises one or more modifications to any one of SEQ ID NOs: 247-337, 3498-3501, 3505-3520, 3540-3549, and 4412-4415. In some embodiments, the CasX variant comprises one or more modifications to any one of SEQ ID NOs: 3498-3501, 3505-3520, and 3540-3549.
[0304] In some embodiments, the CasX variant protein comprises 400 to 2000 amino acids, 500 to 1500 amino acids, 700 to 1200 amino acids, 800 to 1100 amino acids, or 900 to 1000 amino acids.
[0305] In some embodiments, the CasX variant protein comprises one or more modifications in a non-contiguous region of residues that form a channel for gNA:target DNA complex formation. In some embodiments, the CasX variant protein comprises one or more modifications comprising a non-contiguous region of residues that form an interface with a gNA. For example, in some embodiments with reference to a CasX protein, the helical I, helical II, and OBD domains all contact or are adjacent to a gNA:target DNA complex, and one or more modifications to non-contiguous residues within any of these domains can improve the function of the CasX variant protein.
[0306] In some embodiments, the CasX variant protein comprises one or more modifications in a region of noncontiguous residues that form a channel that binds to non-target strand DNA. For example, the CasX variant protein can comprise one or more modifications to noncontiguous residues of the NTSBD. In some embodiments, the CasX variant protein comprises one or more modifications in a region of noncontiguous residues that form an interface that binds to the PAM. For example, the CasX variant protein can comprise one or more modifications to noncontiguous residues of the Helical I domain or OBD. In some embodiments, the CasX variant protein contains one or more modifications comprising a region of noncontiguous surface exposed residues. As used herein, a “surface exposed residue” refers to an amino acid on the surface of a CasX protein, or an amino acid in which at least a portion, such as the backbone or a portion of the side chain, is on the surface of the protein. Surface exposed residues of a cellular protein such as CasX, which are exposed to the aqueous intracellular environment, are often selected from positively charged hydrophilic amino acids, for example arginine, asparagine, aspartic acid, glutamine, glutamic acid, histidine, lysine, serine, and threonine. Thus, for example, in some embodiments of the variants provided herein, a region of surface exposed residues comprises one or more insertions, deletions, or substitutions as compared to a reference CasX protein. In some embodiments, one or more positively charged residues are substituted for one or more other positively charged residues, or negatively charged residues, or uncharged residues, or any combination thereof. In some embodiments, one or more substituted amino acid residues are proximal to binding nucleic acids, for example residues in the RuvC domain or Helical I domain that contact target DNA, or residues in the OBD or Helical II domain that bind gNA can be substituted for one or more positively charged or polar amino acids.
[0307] In some embodiments, a CasX variant protein comprises one or more modifications in a non-contiguous region of residues that forms a core via hydrophobic packing in a domain of a reference CasX protein. Without wishing to be bound by any theory, regions that form a core via hydrophobic packing are rich in hydrophobic amino acids, such as valine, isoleucine, leucine, methionine, phenylalanine, tryptophan, and cysteine. For example, in some reference CasX proteins, the RuvC domain comprises a hydrophobic pocket adjacent to the active site. In some embodiments, 2 to 15 residues of the region are charged, polar, or base-stacked. Charged amino acids (sometimes referred to as residues herein) can include, for example, arginine, lysine, aspartic acid, and glutamic acid, and the side chains of these amino acids can form salt bridges, with the proviso that a bridging partner is also present (see FIG. 14). Polar amino acids can include, for example, glutamine, asparagine, histidine, serine, threonine, tyrosine, and cysteine. In some embodiments, polar amino acids can form hydrogen bonds in the form of a proton donor or acceptor, depending on the identity of their side chains. As used herein, “base-stacking” includes the interaction of aromatic side chains of amino acid residues (such as tryptophan, tyrosine, phenylalanine, or histidine) with stacked nucleotide bases in nucleic acids. Any modification to a non-contiguous amino acid region that is spatially proximal to form a functional portion of a CasX variant protein is contemplated to be within the scope of the present disclosure.
[0308] i. CasX variant proteins having domains from multiple source proteins
[0309] In certain embodiments, the present disclosure provides chimeric CasX proteins comprising protein domains from two or more different CasX proteins, such as two or more naturally occurring CasX proteins, or two or more CasX variant protein sequences as described herein. As used herein, a “chimeric CasX protein” refers to a CasX that contains at least two domains that are isolated or derived from different sources, such as two naturally occurring proteins, in some embodiments, the two proteins can be isolated from different species. For example, in some embodiments, a chimeric CasX protein comprises a first domain from a first CasX protein and a second domain from a different, second CasX protein. In some embodiments, the first domain can be selected from the group consisting of an NTSB, TSL, Helical I, Helical II, OBD, and RuvC domain. In some embodiments, the second domain is selected from the group consisting of an NTSB, TSL, Helical I, Helical II, OBD, and RuvC domain, wherein the second domain is different from the aforementioned first domain. For example, a chimeric CasX protein can comprise an NTSB, TSL, Helical I, Helical II, OBD domain from a CasX protein of SEQ ID NO: 2, and a RuvC domain from a CasX protein of SEQ ID NO: 1, or vice versa. As another example, a chimeric CasX protein can comprise an NTSB, TSL, Helical II, OBD, and RuvC domain from a CasX protein of SEQ ID NO: 2, and a Helical I domain from a CasX protein of SEQ ID NO: 1, or vice versa. Thus, in certain embodiments, a chimeric CasX protein can comprise an NTSB, TSL, Helical II, OBD, and RuvC domain from a first CasX protein, and a Helical I domain from a second CasX protein. In some embodiments of a chimeric CasX protein, the domains of the first CasX protein are derived from the sequence of SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 3, and the domains of the second CasX protein are derived from the sequence of SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 3, and the first and second CasX proteins are not the same. In some embodiments, the domains of the first CasX protein comprise sequences derived from SEQ ID NO: 1, and the domains of the second CasX protein comprise sequences derived from SEQ ID NO: 2. In some embodiments, the domains of the first CasX protein comprise sequences derived from SEQ ID NO: 1, and the domains of the second protein comprise sequences derived from SEQ ID NO: 3. In some embodiments, the domains of the first CasX protein comprise sequences derived from SEQ ID NO: 2, and the domains of the second protein comprise sequences derived from SEQ ID NO: 3.In some embodiments, the CasX variant is selected from the group consisting of a CasX variant having the sequence of SEQ ID NO: 328, SEQ ID NO: 3540, SEQ ID NO: 4413, SEQ ID NO: 4414, SEQ ID NO: 4415, SEQ ID NO: 329, SEQ ID NO: 3541, SEQ ID NO: 330, SEQ ID NO: 3542, SEQ ID NO: 331, SEQ ID NO: 3543, SEQ ID NO: 332, SEQ ID NO: 3544, SEQ ID NO: 333, SEQ ID NO: 3545, SEQ ID NO: 334, SEQ ID NO: 3546, SEQ ID NO: 335, SEQ ID NO: 3547, SEQ ID NO: 336, and SEQ ID NO: 3548. In some embodiments, the CasX variant comprises one or more additional modifications to any of SEQ ID NO: 328, SEQ ID NO: 3540, SEQ ID NO: 4413, SEQ ID NO: 4414, SEQ ID NO: 4415, SEQ ID NO: 329, SEQ ID NO: 3541, SEQ ID NO: 330, SEQ ID NO: 3542, SEQ ID NO: 331, SEQ ID NO: 3543, SEQ ID NO: 332, SEQ ID NO: 3544, SEQ ID NO: 333, SEQ ID NO: 3545, SEQ ID NO: 334, SEQ ID NO: 3546, SEQ ID NO: 335, SEQ ID NO: 3547, SEQ ID NO: 336, or SEQ ID NO: 3548. In some embodiments, the one or more additional modifications comprise an insertion, a substitution, or a deletion as described herein.
[0310] In some embodiments, a CasX variant protein comprises at least one chimeric domain comprising a first portion from a first CasX protein and a second portion from a different, second CasX protein. As used herein, a “chimeric domain” refers to a domain containing at least two portions that are separate or derived from different sources, such as two naturally occurring proteins, or domain portions from two reference CasX proteins. The at least one chimeric domain can be any of an NTSB, TSL, Helical I, Helical II, OBD, or RuvC domain as described herein. In some embodiments, the first portion of the CasX domain comprises the sequence of SEQ ID NO: 1 and the second portion of the CasX domain comprises the sequence of SEQ ID NO: 2. In some embodiments, the first portion of the CasX domain comprises the sequence of SEQ ID NO: 1 and the second portion of the CasX domain comprises the sequence of SEQ ID NO: 3. In some embodiments, the first portion of the CasX domain comprises the sequence of SEQ ID NO: 2 and the second portion of the CasX domain comprises the sequence of SEQ ID NO: 3. In some embodiments, the at least one chimeric domain comprises a chimeric RuvC domain. As an example of the foregoing, the chimeric RuvC domain comprises amino acids 661 to 824 of SEQ ID NO: 1 and amino acids 922 to 978 of SEQ ID NO: 2. As an alternative example of the foregoing, the chimeric RuvC domain comprises amino acids 648 to 812 of SEQ ID NO: 2 and amino acids 935 to 986 of SEQ ID NO: 1. In some embodiments, a CasX protein comprises a first domain from a first CasX protein and a second domain from a second CasX protein, and at least one chimeric domain comprising at least two portions isolated from different CasX proteins using the methods of the embodiments described in this paragraph. In the foregoing embodiments, a chimeric CasX protein having domains or domain portions derived from SEQ ID NOs: 1, 2, and 3 can further comprise the amino acid insertions, deletions, or substitutions of any of the embodiments disclosed herein.
[0311] In some embodiments, the CasX variant protein comprises a sequence set forth in Table 3, 8, 9, 10, or 12. In other embodiments, the CasX variant protein comprises a sequence that is at least 60% identical, at least 65% identical, at least 70% identical, at least 75% identical, at least 80% identical, at least 81% identical, at least 82% identical, at least 83% identical, at least 84% identical, at least 85% identical, at least 86% identical, at least 86% identical, at least 87% identical, at least 88% identical, at least 89% identical, at least 89% identical, at least 90% identical, at least 91% identical, at least 92% identical, at least 93% identical, at least 94% identical, at least 95% identical, at least 96% identical, at least 97% identical, at least 98% identical, at least 99% identical, at least 99.5% identical to a sequence set forth in Table 3, 8, 9, 10, or 12. In other embodiments, the CasX variant protein comprises a sequence set forth in Table 3, and further comprises one or more NLSs disclosed herein at the N-terminus, C-terminus, or both. It will be appreciated that in some cases, the N-terminal methionine of a CasX variant in the table is removed from the expressed CasX variant during post-translational modification.
[0312] Table 3: CasX variant sequences
[0313]
[0314]
[0315]
[0316]
[0317] Strains indicated by a number; when indicated, change is relative to SEQ ID NO: 2
[0318] In some embodiments, the CasX variant protein has one or more improved characteristics when compared to a reference CasX protein, e.g., the reference protein of SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 3. In some embodiments, the improved characteristic of the CasX variant is improved by at least about 1.1 to about 100,000 fold relative to the reference protein. In some embodiments, the improved characteristic of the CasX variant is improved by at least about 1.1 to about 10,000 fold, improved by at least about 1.1 to about 1,000 fold, improved by at least about 1.1 to about 500 fold, improved by at least about 1.1 to about 400 fold, improved by at least about 1.1 to about 300 fold, improved by at least about 1.1 to about 200 fold, improved by at least about 1.1 to about 100 fold, improved by at least about 1.1 to about 50 fold, improved by at least about 1.1 to about 40 fold, improved by at least about 1.1 to about 30 fold, improved by at least about 1.1 to about 20 fold, improved by at least about 1.1 to about 10 fold, improved by at least about 1.1 to about 9 fold, improved by at least about 1.1 to about 8 fold, improved by at least about 1.1 to about 7 fold, improved by at least about 1.1 to about 6 fold, improved by at least about 1.1 to about 5 fold, improved by at least about 1.1 to about 4 fold, improved by at least about 1.1 to about 3 fold, improved by at least about 1.1 to about 2 fold, improved by at least about 1.1 to about 1.5 fold, improved by at least about 1.5 to about 3 fold, improved by at least about 1.5 to about 4 fold, improved by at least about 1.5 to about 5 fold, improved by at least about 1.5 to about 10 fold, improved by at least about 5 to about 10 fold, improved by at least about 10 to about 20 fold, improved by at least 10 to about 30 fold, improved by at least 10 to about 50 fold, or improved by at least 10 to about 100 fold relative to the reference CasX protein. In some embodiments, the improved characteristic of the CasX variant is improved by at least about 10 to about 1000 fold relative to the reference CasX protein.
[0319] In some embodiments, the one or more improved features of a CasX variant protein are improved by at least about 5, at least about 10, at least about 20, at least about 30, at least about 40, at least about 50, at least about 60, at least about 70, at least about 80, at least about 90, at least about 100, at least about 250, at least about 500, or at least about 1000, at least about 5,000, at least about 10,000, or at least about 100,000 fold relative to a reference CasX protein. In some embodiments, the improved features of a CasX variant protein are improved by at least about 1.1, at least about 1.2, at least about 1.3, at least about 1.4, at least about 1.5, at least about 1.6, at least about 1.7, at least about 1.8, at least about 1.9, at least about 2, at least about 2.1, at least about 2.2, at least about 2.3, at least about 2.4, at least about 2.5, at least about 2.6, at least about 2.7, at least about 2.8, at least about 2.9, at least about 3, at least about 3.5, at least about 4, at least about 4.5, at least about 5, at least about 5.5, at least about 6, at least about 6.5, at least about 7.0, at least about 7.5, at least about 8, at least about 8.5, at least about 9, at least about 9.5, at least about 10, at least about 11, at least about 12, at least about 13, at least about 14, at least about 15, at least about 20, at least about 30, at least about 40, at least about 50, at least about 60, at least about 70, at least about 80, at least about 90, at least about 100, at least about 500, at least about 1,000, at least about 10,000, or at least about 100,000 fold relative to a reference CasX protein.In other cases, the one or more improved features of the CasX variant are improved by about 1.1 to 100,00 fold, about 1.1 to 10,00 fold, about 1.1 to 1,000 fold, about 1.1 to 500 fold, about 1.1 to 100 fold, about 1.1 to 50 fold, about 1.1 to 20 fold, about 10 to 100,00 fold, about 10 to 10,00 fold, about 10 to 1,000 fold, about 10 to 500 fold, about 10 to 100 fold, about 10 to 50 fold, about 10 to 20 fold, about 2 to 70 fold, about 2 to 50 fold, about 2 to 30 fold, about 2 to 20 fold, about 2 to 10 fold, about 5 to 50 fold, about 5 to 30 fold, about 5 to 10 fold, about 100 to 100,00 fold, about 100 to 10,00 fold, about 100 to 1,000 fold, about 100 to 500 fold, about 500 to 100,00 fold, about 500 to 10,00 fold, about 500 to 1,000 fold, about 500 to 750 fold, about 1,000 to 100,00 fold, about 10,000 to 100,00 fold, about 20 to 500 fold, about 20 to 250 fold, about 20 to 200 fold, about 20 to 100 fold, about 20 to 50 fold, about 50 to 10,000 fold, about 50 to 1,000 fold, about 50 to 500 fold, about 50 to 200 fold, or about 50 to 100 fold relative to the reference CasX of SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 3. In other cases, the one or more improved features of the CasX variant are improved by about 1.1 fold, 1.2 fold, 1.3 fold, 1.4 fold, 1.5 fold, 1.6 fold, 1.7 fold, 1.8 fold, 1.9 fold, 2 fold, 3 fold, 4 fold, 5 fold, 6 fold, 7 fold, 8 fold, 9 fold, 10 fold, 11 fold, 12 fold, 13 fold, 14 fold, 15 fold, 16 fold, 17 fold, 18 fold, 19 fold, 20 fold, 25 fold, 30 fold, 40 fold, 45 fold, 50 fold, 55 fold, 60 fold, 70 fold, 80 fold, 90 fold, 100 fold, 110 fold, 120 fold, 130 fold, 140 fold, 150 fold, 160 fold, 170 fold, 180 fold, 190 fold, 200 fold, 210 fold, 220 fold, 230 fold, 240 fold, 250 fold, 260 fold, 270 fold, 280 fold, 290 fold, 300 fold, 310 fold, 320 fold, 330 fold, 340 fold, 350 fold, 360 fold, 370 fold, 380 fold, 390 fold, 400 fold, 425 fold, 450 fold, 475 fold, or 500 fold relative to the reference CasX of SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 3.Exemplary features that can be improved in a CasX variant protein relative to the same feature in a reference CasX protein include, but are not limited to, improved variant folding, improved binding affinity to a gNA, improved binding affinity to a target DNA, improved ability to edit and / or bind to a target DNA using a larger range of PAM sequences, improved target DNA unwinding, increased editing activity, improved editing efficiency, improved editing specificity, increased nuclease activity, increased target strand loading for double-strand cleavage, reduced target strand loading for single-strand cleavage, reduced off-target cleavage, improved binding of a non-target strand of DNA, improved protein stability, improved CasX:gNA RNA complex stability, improved protein solubility, improved CasX:gNA RNP complex solubility, improved protein yield, improved protein expression, and improved melting characteristics. In some embodiments, a variant comprises at least one improved feature. In other embodiments, a variant comprises at least two improved features. In other embodiments, a variant comprises at least three improved features. In some embodiments, a variant comprises at least four improved features. In other embodiments, a variant comprises at least five, at least six, at least seven, at least eight, at least nine, at least ten, at least eleven, at least twelve, at least thirteen, or more improved features. These improved features are described in more detail below.
[0320] j. Protein stability
[0321] In some embodiments, the present disclosure provides CasX variant proteins having improved stability relative to a reference CasX protein. In some embodiments, the improved stability of a CasX variant protein results in higher steady state protein expression, which improves editing efficiency. In some embodiments, the improved stability of a CasX variant protein results in a larger fraction of CasX protein remaining folded in a functional conformation, and improves editing efficiency or improves purification capabilities for manufacturing purposes. As used herein, a “functional conformation” refers to a conformation of a CasX protein in which the protein is capable of binding a gNA and a target DNA. In embodiments in which a CasX variant does not harbor one or more mutations that render it catalytically dead, the CasX variant is capable of cleaving, cutting, or otherwise modifying a target DNA. For example, in some embodiments, a functional CasX variant can be used for gene editing, and a functional conformation refers to an “editing-competent” conformation. In some exemplary embodiments, including those in which a CasX variant protein produces a larger fraction of CasX protein remaining folded in a functional conformation, lower concentrations of the CasX variant are required for applications such as gene editing relative to a reference CasX protein. Thus, in some embodiments, a CasX variant having improved stability has improved efficiency in one or more gene editing contexts relative to a reference CasX.
[0322] In some embodiments, the present application provides CasX variant proteins having improved thermostability relative to a reference CasX protein. In some embodiments, the CasX variant proteins have improved CasX variant protein thermostability within a particular temperature range. Without wishing to be bound by any theory, some reference CasX proteins naturally function in organisms that reside in groundwater and sediments at the ecological niche; thus, some reference CasX proteins can have evolved to exhibit optimal function at lower or higher temperatures than can be desirable for certain applications. For example, one application of CasX variant proteins is gene editing of mammalian cells, which is typically performed at about 37°C. In some embodiments, CasX variant proteins as described herein have improved thermostability at at least 16°C, at least 18°C, at least 20°C, at least 22°C, at least 24°C, at least 26°C, at least 28°C, at least 30°C, at least 32°C, at least 34°C, at least 35°C, at least 36°C, at least 37°C, at least 38°C, at least 39°C, at least 40°C, at least 41°C, at least 42°C, at least 44°C, at least 46°C, at least 48°C, at least 50°C, at least 52°C, or higher, relative to a reference CasX protein. In some embodiments, CasX variant proteins have improved thermostability and function relative to a reference CasX protein, resulting in improved gene editing function, such as mammalian gene editing applications, which can include human gene editing applications.
[0323] In some embodiments, the present application provides CasX variant proteins having improved CasX variant protein:gNA RNP complex stability relative to a reference CasX protein:gNA complex, such that the RNP remains in a functional form. Stability improvements can include increased thermal stability; proteolytic degradation resistance; enhanced pharmacokinetic properties; stability across a range of pH conditions, salt conditions, and tensions. In some embodiments, the improved stability of the complex results in increased editing efficiency.
[0324] In some embodiments, the present application provides a CasX variant protein having improved CasX variant protein:gNA complex thermal stability relative to a reference CasX protein:gNA complex. In some embodiments, the CasX variant protein has improved thermal stability relative to a reference CasX protein. In some embodiments, the CasX variant protein:gNA RNP complex has improved thermal stability at at least 16°C, at least 18°C, at least 20°C, at least 22°C, at least 24°C, at least 26°C, at least 28°C, at least 30°C, at least 32°C, at least 34°C, at least 35°C, at least 36°C, at least 37°C, at least 38°C, at least 39°C, at least 40°C, at least 41°C, at least 42°C, at least 44°C, at least 46°C, at least 48°C, at least 50°C, at least 52°C, or more, relative to a complex comprising a reference CasX protein. In some embodiments, the CasX variant protein has improved CasX variant protein:gNA RNP complex thermal stability relative to a reference CasX protein:gNA complex, which results in improved function for gene editing applications, such as mammalian gene editing applications, which can include human gene editing applications.
[0325] In some embodiments, the improved stability and / or thermal stability of the CasX variant protein comprises faster folding kinetics of the CasX variant protein relative to a reference CasX protein, slower unfolding kinetics of the CasX variant protein relative to a reference CasX protein, greater free energy release upon folding of the CasX variant protein relative to a reference CasX protein, higher temperature at which 50% of the CasX variant protein is unfolded (Tm) relative to a reference CasX protein, or any combination thereof. These features can be improved by a wide range of values; for example, at least 1.1, at least 1.5, at least 10, at least 50, at least 100, at least 500, at least 1,000, at least 5,000, or at least 10,000 fold relative to a reference CasX protein. In some embodiments, the improved thermal stability of the CasX variant protein comprises a higher Tm of the CasX variant protein relative to a reference CasX protein. In some embodiments, the Tm of the CasX variant protein is about 20°C to about 30°C, about 30°C to about 40°C, about 40°C to about 50°C, about 50°C to about 60°C, about 60°C to about 70°C, about 70°C to about 80°C, about 80°C to about 90°C, or about 90°C to about 100°C. Thermal stability is measured by determining the “melting temperature” (Tm) of the CasX variant protein, as described herein. m) to determine, with the melting temperature defined as the temperature at which half the molecules are denatured. Methods for measuring characteristics of protein stability, such as Tm and the free energy of unfolding, are known to those of ordinary skill in the art and can be measured in vitro using standard biochemical techniques. For example, Tm can be measured using differential scanning calorimetry, a thermal analysis technique in which the heat required to increase the temperature of a sample and a reference is measured as a function of temperature (Chen et al. (2003) Pharm Res 20:1952-60; Ghirlando et al. (1999) Immunol Lett 68:47-52). Alternatively or additionally, CasX variant protein Tm can be measured using a commercially available method, such as the ThermoFisher Protein Thermal Shift System. Alternatively or additionally, circular dichroism can be used to measure the kinetics of folding and unfolding, as well as Tm (Murray et al. (2002) J. Chromatogr Sci 40:343-9). Circular dichroism (CD) relies on the unequal absorption of left- and right-handed circularly polarized light by asymmetric molecules such as proteins. Certain structures of proteins, such as alpha helices and beta sheets, have characteristic CD spectra. Thus, in some embodiments, CD can be used to determine the secondary structure of a CasX variant protein.
[0326] In some embodiments, the improved stability and / or thermostability of the CasX variant protein comprises an improved folding kinetics of the CasX variant protein relative to a reference CasX protein. In some embodiments, the folding kinetics of the CasX variant protein is improved at least about 5-fold, at least about 10-fold, at least about 50-fold, at least about 100-fold, at least about 500-fold, at least about 1,000-fold, at least about 2,000-fold, at least about 3,000-fold, at least about 4,000-fold, at least about 5,000-fold, or at least about 10,000-fold relative to a reference CasX protein. In some embodiments, the folding kinetics of the CasX variant protein is improved at least about 1 kJ / mol, at least about 5 kJ / mol, at least about 10 kJ / mol, at least about 20 kJ / mol, at least about 30 kJ / mol, at least about 40 kJ / mol, at least about 50 kJ / mol, at least about 60 kJ / mol, at least about 70 kJ / mol, at least about 80 kJ / mol, at least about 90 kJ / mol, at least about 100 kJ / mol, at least about 150 kJ / mol, at least about 200 kJ / mol, at least about 250 kJ / mol, at least about 300 kJ / mol, at least about 350 kJ / mol, at least about 400 kJ / mol, at least about 450 kJ / mol, or at least about 500 kJ / mol relative to a reference CasX protein.
[0327] Exemplary amino acid changes that can increase the stability of a CasX variant protein relative to a reference CasX protein can include, but are not limited to, the following amino acid changes: increasing the number of hydrogen bonds within the CasX variant protein, increasing the number of disulfide bridges within the CasX variant protein, increasing the number of salt bridges within the CasX variant protein, enhancing the interactions between portions of the CasX variant protein, increasing the buried hydrophobic surface area of the CasX variant protein, or any combination thereof.
[0328] k. Protein yield
[0329] In some embodiments, the present application provides CasX variant proteins having improved yield during expression and purification relative to a reference CasX protein. In some embodiments, the yield of a CasX variant protein purified from a bacterial or eukaryotic host cell is improved relative to a reference CasX protein. In some embodiments, the bacterial host cell is an E. coli cell. In some embodiments, the eukaryotic cell is a yeast, plant (e.g., tobacco), insect (e.g., Spodoptera frugiperda sf9 cells), mouse, rat, hamster, guinea pig, non-human primate, or human cell. In some embodiments, the eukaryotic host cell is a mammalian cell, including but not limited to, a HEK293 cell, a HEK293T cell, a HEK293-F cell, a Lenti-X 293T cell, a BHK cell, a HepG2 cell, a Saos-2 cell, a HuH7 cell, an A549 cell, an NS0 cell, an SP2 / 0 cell, a YO myeloma cell, a P3X63 mouse myeloma cell, a PER cell, a PER.C6 cell, a hybridoma cell, a VERO cell, a NIH3T3 cell, a COS, a WI38 cell, a MRC5 cell, a HeLa, a HT1080 cell, or a CHO cell.
[0330] In some embodiments, improved yield of the CasX variant protein is achieved by codon optimization. Cells use 64 different codons, 61 of which encode 20 standard amino acids, while the other 3 act as stop codons. In some cases, a single amino acid is encoded by more than one codon. For the same naturally occurring amino acid, different organisms exhibit a bias toward using different codons. Thus, the choice of codons in a protein, and matching the choice of codons to the organism in which the protein will be expressed, can in some cases significantly affect protein translation and thus protein expression levels. In some embodiments, the CasX variant protein is encoded by a nucleic acid that has been codon optimized. In some embodiments, the nucleic acid encoding the CasX variant protein has been codon optimized for expression in a bacterial cell, a yeast cell, an insect cell, a plant cell, or a mammalian cell. In some embodiments, the mammalian cell is a mouse, rat, hamster, guinea pig, monkey, or human. In some embodiments, the CasX variant protein is encoded by a nucleic acid that has been codon optimized for expression in a human cell. In some embodiments, the CasX variant protein is encoded by a nucleic acid that has been removed of nucleotide sequences that decrease translation rate in prokaryotes and eukaryotes. For example, runs of greater than three thymine residues in a row can decrease translation rate in certain organisms, or internal polyadenylation signals can decrease translation.
[0331] In some embodiments, the improvements in solubility and stability as described herein result in improved yield of the CasX variant protein relative to a reference CasX protein.
[0332] Protein yield during expression and purification can be assessed by methods known in the art. For example, the amount of CasX variant protein can be determined by running the protein on an SDS-page gel and comparing the CasX variant protein to a control of known amount or concentration to determine the absolute amount of protein. Alternatively or additionally, the purified CasX variant protein can be run on an SDS-page gel next to a reference CasX protein that has undergone the same purification process to determine the relative improvement in CasX variant protein yield. Alternatively or additionally, protein content can be measured using immunohistochemical methods, such as by western blot or ELISA with an antibody against CasX, or by HPLC. For proteins in solution, the concentration can be determined by measuring the intrinsic UV absorbance of the protein, or by methods that use protein-dependent color changes, such as the Lowry assay, the Smith copper / bicinchoninic acid assay, or the Bradford dye assay. Such methods can be used to calculate the total protein (such as total soluble protein) yield obtained by expression under certain conditions. For example, this can be compared to the protein yield of a reference CasX protein under similar expression conditions.
[0333] l. Protein solubility
[0334] In some embodiments, a CasX variant protein has improved solubility relative to a reference CasX protein. In some embodiments, a CasX variant protein has improved CasX:gNA ribonucleoprotein complex variant solubility relative to a ribonucleoprotein complex comprising a reference CasX protein.
[0335] In some embodiments, the improved solubility of the protein results in higher protein yields from protein purification techniques, such as purification from E. coli. In some embodiments, the improved solubility of the CasX variant protein can enable higher activity in cells because more soluble proteins are less likely to aggregate in cells. Protein aggregates can be toxic or burdensome to cells in certain embodiments, and without wishing to be bound by any theory, increasing the solubility of the CasX variant protein can improve this protein aggregation outcome. Additionally, the improved solubility of the CasX variant protein can allow for enhanced formulations, permitting delivery of more highly effective doses of functional protein, for example, in desired gene editing applications. In some embodiments, the improved solubility of the CasX variant protein relative to a reference CasX protein results in an improved yield of the CasX variant protein during purification, by at least about 5-fold, at least about 10-fold, at least about 20-fold, at least about 30-fold, at least about 40-fold, at least about 50-fold, at least about 60-fold, at least about 70-fold, at least about 80-fold, at least about 90-fold, at least about 100-fold, at least about 250-fold, at least about 500-fold, or at least about 1000-fold. In some embodiments, the improved solubility of the CasX variant protein relative to a reference CasX protein improves the activity of the CasX variant protein in cells, by at least about 1.1-fold, at least about 1.2-fold, at least about 1.3-fold, at least about 1.4-fold, at least about 1.5-fold, at least about 1.6-fold, at least about 1.7-fold, at least about 1.8-fold, at least about 1.9-fold, at least about 2-fold, at least about 2.1-fold, at least about 2.2-fold, at least about 2.3-fold, at least about 2.4-fold, at least about 2.5-fold, at least about 2.6-fold, at least about 2.7-fold, at least about 2.8-fold, at least about 2.9-fold, at least about 3-fold, at least about 3.5-fold, at least about 4-fold, at least about 4.5-fold, at least about 5-fold, at least about 5.5-fold, at least about 6-fold, at least about 6.5-fold, at least about 7.0-fold, at least about 7.5-fold, at least about 8-fold, at least about 8.5-fold, at least about 9-fold, at least about 9.5-fold, at least about 10-fold, at least about 11-fold, at least about 12-fold, at least about 13-fold, at least about 14-fold, at least about 15-fold, or at least about 20-fold.
[0336] Methods of measuring CasX protein solubility and improvements thereof in CasX variant proteins will be apparent to one of ordinary skill in the art. For example, in some embodiments, CasX variant protein solubility can be measured by taking densitometry readings on gels resolving the soluble fraction of E. coli lysates. Alternatively or additionally, improvements in CasX variant protein solubility can be measured by measuring the maintenance of soluble protein product throughout the protein purification process, including the methods of example. For example, soluble protein product can be measured at one or more steps of gel affinity purification, tag cleavage, cation exchange purification, running the protein over a size exclusion chromatography (SEC) column. In some embodiments, densitometry values for each protein band on a gel are read after each step of the purification process. In some embodiments, CasX variant proteins with improved solubility, when compared to a reference CasX protein, can maintain a higher concentration at one or more steps of the protein purification process, while insoluble protein variants can be lost at one or more steps due to buffer exchange, filtration steps, interaction with purification columns, etc.
[0337] In some embodiments, improved solubility of CasX variant proteins results in higher yields in terms of mg / L of protein during protein purification when compared to a reference CasX protein.
[0338] In some embodiments, improved solubility of CasX variant proteins enables a greater amount of editing events when assessed in editing assays, such as the EGFP disruption assay described herein, as compared to less soluble proteins.
[0339] m. Affinity for gNA
[0340] In some embodiments, CasX variant proteins have improved affinity for gNA relative to a reference CasX protein, such that ribonucleoprotein complexes are formed. Increased affinity of CasX variant proteins for gNA can, for example, result in lower Kd for RNP complex formation d which can in some cases make ribonucleoprotein complex formation more stable. In some embodiments, CasX variant proteins have a Kd for gNA that is improved relative to a reference CasX protein by at least 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 100%, or more. dat least about 1.1 fold, at least about 1.2 fold, at least about 1.3 fold, at least about 1.4 fold, at least about 1.5 fold, at least about 1.6 fold, at least about 1.7 fold, at least about 1.8 fold, at least about 1.9 fold, at least about 2 fold, at least about 3 fold, at least about 4 fold, at least about 5 fold, at least about 6 fold, at least about 7 fold, at least about 8 fold, at least about 9 fold, at least about 10 fold, at least about 15 fold, at least about 20 fold, at least about 25 fold, at least about 30 fold, at least about 35 fold, at least about 40 fold, at least about 45 fold, at least about 50 fold, at least about 60 fold, at least about 70 fold, at least about 80 fold, at least about 90 fold, or at least about 100 fold. In some embodiments, the CasX variant has an increased binding affinity for gNA of about 1.1 to about 10 fold compared to the reference CasX protein of SEQ ID NO: 2.
[0341] In some embodiments, the increased affinity of the CasX variant protein for gNA results in increased stability of the ribonucleoprotein complex when delivered to a mammalian cell, including in vivo delivery to an individual. This increased stability can affect the function and utility of the complex in the cells of the individual, as well as improve the pharmacokinetic properties in the blood when delivered to an individual. In some embodiments, the increased affinity of the CasX variant, and the resulting increased stability of the ribonucleoprotein complex, allows for lower doses of the CasX variant to be delivered to an individual or cell while still having the desired activity; e.g., in vivo or in vitro gene editing. The enhanced ability to form and maintain RNP in a stable form can be assessed using assays, such as the in vitro cleavage assay described herein. In some embodiments, the CasX variants of the application are capable of obtaining a Kd for RNP that is at least 2 fold, at least 5 fold, or at least 10 fold higher than the RNP of the reference CasX when complexed in RNP form. cleave rate.
[0342] In some embodiments, the higher affinity (more tightly bound) of the CasX variant protein for gNA when both are maintained in an RNP complex allows for a greater amount of editing events. The increased editing events can be assessed using editing assays, such as the EGFP disruption and in vitro cleavage assays described herein.
[0343] Without wishing to be bound by theory, in some embodiments, amino acid changes in the helical I domain can increase the binding affinity of the CasX variant protein for the gNA targeting sequence, while changes in the helical II domain can increase the binding affinity of the CasX variant protein for the gNA scaffold stem loop, and changes in the oligonucleotide binding domain (OBD) increase the binding affinity of the CasX variant protein for the gNA triplex.
[0344] Methods of measuring the binding affinity of CasX proteins to gNAs include in vitro methods using purified CasX proteins and gNAs. If the gNAs or CasX proteins are labeled with a fluorophore, the binding affinity of reference CasX and variant proteins can be measured by fluorescence polarization. Alternatively or additionally, the binding affinity can be measured by biolayer interferometry, electrophoretic mobility shift assay (EMSA), or filter binding. Additional standard techniques for quantifying the absolute affinity of RNA binding proteins, such as reference CasX and variant proteins of the present disclosure, for specific gNAs, such as reference gNAs and variants thereof, include, but are not limited to, isothermal titration calorimetry (ITC) and surface plasmon resonance (SPR), as well as the methods of examples.
[0345] n. affinity for a target nucleic acid
[0346] In some embodiments, the binding affinity of a CasX variant protein for a target nucleic acid is improved relative to the affinity of a reference CasX protein for the target nucleic acid. In some embodiments, a CasX variant with higher affinity for a target nucleic acid can cleave the target nucleic acid sequence more rapidly than a reference CasX protein that does not have increased affinity for the target nucleic acid.
[0347] In some embodiments, improved affinity for a target nucleic acid includes improved affinity for a target sequence or protospacer sequence of the target nucleic acid, improved affinity for a PAM sequence, improved ability to search DNA for a target sequence, or any combination thereof. Without wishing to be bound by theory, it is believed that CRISPR / Cas system proteins, such as CasX, can find their target sequences by one-dimensional diffusion along a DNA molecule. It is believed that the process includes (1) binding of the ribonucleoprotein to the DNA molecule, followed by (2) pausing at a target sequence, either of which, in some embodiments, can be affected by improved affinity of a CasX protein for a target nucleic acid sequence, thereby improving the function of a CasX variant protein compared to a reference CasX protein.
[0348] In some embodiments, CasX variant proteins with improved target nucleic acid affinity have increased overall affinity for DNA. In some embodiments, CasX variant proteins with improved target nucleic acid affinity have increased affinity for or ability to utilize specific PAM sequences that are not the canonical TTC PAM recognized by the reference CasX protein of SEQ ID NO: 2, including PAM sequences selected from the group consisting of TTC, ATC, GTC, and CTC, thereby increasing the amount of editable target DNA compared to wild-type CasX nucleases. Without wishing to be bound by theory, it is possible that these protein variants can interact more robustly with DNA overall due to the ability to utilize additional PAM sequences beyond those of wild-type reference CasX, and can have enhanced ability to access and edit sequences within target DNA, thereby allowing for a more efficient method of searching for CasX proteins for target sequences. In some embodiments, higher overall affinity for DNA can also increase the frequency with which CasX proteins can effectively initiate and complete the binding and unwinding steps, thereby facilitating target strand invasion and R-loop formation, and ultimately target nucleic acid sequence cleavage.
[0349] Without wishing to be bound by theory, it is possible that amino acid changes in the NTSBD that increase the unwinding of non-target DNA strands or the capture efficiency of non-target DNA strands in an unwound state can increase the affinity of CasX variant proteins for target DNA. Alternatively or additionally, amino acid changes in the NTSBD that increase the ability of the NTSBD to stabilize DNA during unwinding can increase the affinity of CasX variant proteins for target DNA. Alternatively or additionally, amino acid changes in the OBD can increase the affinity of CasX variant proteins to bind to protospacer adjacent motifs (PAMs), thereby increasing the affinity of CasX variant proteins for target nucleic acids. Alternatively or additionally, amino acid changes in the helical I and / or II, RuvC, and TSL domains that increase the affinity of CasX variant proteins for target nucleic acid strands can increase the affinity of CasX variant proteins for target nucleic acids.
[0350] In some embodiments, the binding affinity of a CasX variant protein of the application to a target nucleic acid molecule is increased by at least about 1.1-fold, at least about 1.2-fold, at least about 1.3-fold, at least about 1.4-fold, at least about 1.5-fold, at least about 1.6-fold, at least about 1.7-fold, at least about 1.8-fold, at least about 1.9-fold, at least about 2-fold, at least about 3-fold, at least about 4-fold, at least about 5-fold, at least about 6-fold, at least about 7-fold, at least about 8-fold, at least about 9-fold, at least about 10-fold, at least about 15-fold, at least about 20-fold, at least about 25-fold, at least about 30-fold, at least about 35-fold, at least about 40-fold, at least about 45-fold, at least about 50-fold, at least about 60-fold, at least about 70-fold, at least about 80-fold, at least about 90-fold, or at least about 100-fold relative to a reference CasX protein. In some embodiments, the binding affinity of a CasX variant protein to a target nucleic acid is increased by about 1.1 to about 100-fold relative to the reference protein of SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 3.
[0351] In some embodiments, the binding affinity of a CasX variant protein to the non-target strand of a target nucleic acid is improved. As used herein, the term “non-target strand” refers to the strand of a DNA target nucleic acid sequence that does not form Watson and Crick base pairs with the targeting sequence in the gNA and is complementary to the target DNA strand. In some embodiments, the binding affinity of a CasX variant protein to the non-target strand of a target nucleic acid is increased by about 1.1 to about 100-fold relative to the reference protein of SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 3.
[0352] Methods of measuring the affinity of a CasX protein (e.g., a reference or variant) to a target and / or non-target nucleic acid molecule can include electrophoretic mobility shift assay (EMSA), filter binding, isothermal calorimetry (ITC), and surface plasmon resonance (SPR), fluorescence polarization, and biolayer interferometry (BLI). Other methods of measuring the affinity of a CasX protein to a target include in vitro biochemical assays that measure DNA cleavage events over time.
[0353] o. Improved specificity to a target site
[0354] In some embodiments, the CasX variant protein improves specificity of the CasX variant protein for a target nucleic acid sequence relative to a reference CasX protein. As used herein, “specificity” (sometimes referred to as “on-target specificity”) refers to the degree to which a CRISPR / Cas system ribonucleoprotein complex cleaves off-target sequences that are similar to, but not identical to, the target nucleic acid sequence; for example, a CasX variant RNP with a higher degree of specificity relative to a reference CasX protein would exhibit reduced sequence off-target cleavage. Reduction in specificity and potentially deleterious off-target effects of CRISPR / Cas system proteins can be of utmost importance in order to achieve an acceptable therapeutic index for use in a mammalian subject.
[0355] In some embodiments, the CasX variant protein improves specificity of the CasX variant protein for a target site within a target sequence that is complementary to a targeting sequence of a gNA. Without wishing to be bound by theory, it is possible that amino acid changes in the helical I and II domains that increase specificity of the CasX variant protein for a target nucleic acid strand can increase overall specificity of the CasX variant protein for a target nucleic acid. In some embodiments, amino acid changes that increase specificity of the CasX variant protein for a target nucleic acid can also decrease affinity of the CasX variant protein for DNA.
[0356] Methods to test on-target specificity of a CasX protein (e.g., a variant or reference) can include directed and circularization to report cleavage effects in vitro by sequencing (CIRCLE-seq), or similar methods. Briefly, in the CIRCLE-seq technique, genomic DNA is sheared and circularized by ligation of a stem-loop adaptor, which is nicked in the stem-loop region to expose a 4-nucleotide palindromic overhang. This is followed by intramolecular ligation and degradation of the remaining linear DNA. Circular DNA molecules containing CasX cleavage sites are then linearized by CasX, and adaptors are ligated to the exposed ends, followed by high-throughput sequencing to generate paired-end reads containing information about off-target sites. Additional analyses that can be used to detect off-target events, and thus detect CasX protein specificity, include analyses for detecting and quantifying indels (insertions and deletions) formed at those selected off-target sites, such as mismatch detection nuclease assays and next-generation sequencing (NGS). An exemplary mismatch detection assay includes a nuclease assay in which genomic DNA from cells treated with CasX and sgNA is PCR-amplified, denatured, and re-hybridized to form heteroduplex DNA, which contains one wild-type strand and one strand with an indel. Mismatches are recognized and cleaved by a mismatch detection nuclease, such as Surveyor nuclease or T7 endonuclease I.
[0357] p. Protospacer and PAM sequences
[0358] Herein, a protospacer is defined as the DNA sequence that is complementary to the targeting sequence of a guide RNA, and the DNA complementary to the sequence is referred to as the target strand and the non-target strand, respectively. As used herein, a PAM is a nucleotide sequence proximal to the protospacer that binds with the targeting sequence of a gNA to aid CasX orientation and positioning for potential cleavage of the protospacer strands.
[0359] PAM sequences can be degenerate, and a particular RNP construct can have different preferred and tolerated PAM sequences that support different cleavage efficiencies. Unless otherwise noted, it is conventional for the present disclosure to refer to PAM and protospacer sequences and their orientation according to the non-target strand. This does not mean the non-target strand, but rather the PAM sequence of the non-target strand dictates cleavage or is mechanistically involved in target recognition. For example, when reference is made to a TTC PAM, it can actually be the complementary GAA sequence required for target cleavage, or it can be some combination of nucleotides from both strands. In the case of the CasX proteins disclosed herein, the PAM is located 5’ of the protospacer, with a single nucleotide separating the PAM from the first nucleotide of the protospacer. Thus, in reference to CasX, a TTC PAM should be understood to mean a sequence following the formula 5’-…NNTTCN(protospacer)NNNNNN…3’ (SEQ ID NO: 3296), where ‘N’ is a DNA nucleotide and ‘(protospacer)’ is a DNA sequence that has identity to the targeting sequence of a guide RNA. In the case of CasX variants with expanded PAM recognition, a TTC, CTC, GTC, or ATC PAM should be understood to mean a sequence following the formula: 5’-…NNTTCN(protospacer)NNNNNN…3’ (SEQ ID NO: 3296); 5’-…NNCTCN(protospacer)NNNNNN…3’ (SEQ ID NO: 3297); 5’-…NNGTCN(protospacer)NNNNNN…3’ (SEQ ID NO: 3298); or 5’-…NNATCN(protospacer)NNNNNN…3’ (SEQ ID NO: 3299). Alternatively, a TC PAM should be understood to mean a sequence following the formula: 5’-…NNNTCN(protospacer)NNNNNN…3’ (SEQ ID NO: 3300).
[0360] In some embodiments, the CasX variant with improved PAM sequence editing exhibits greater editing efficiency and / or binding of a target sequence in a target DNA when any one of the PAM sequences TTC, ATC, GTC, or CTC is located 1 nucleotide 5’ to the non-target strand of the protospacer having identity to the targeting sequence of the gNA in the cellular assay system relative to the editing efficiency and / or binding of an RNP comprising a reference CasX protein in a similar assay system. In some embodiments, the PAM sequence is TTC. In some embodiments, the PAM sequence is ATC. In some embodiments, the PAM sequence is CTC. In some embodiments, the PAM sequence is GTC.
[0361] q. DNA unwinding
[0362] In some embodiments, the CasX variant protein has improved ability to unwind DNA relative to the reference CasX protein. It has been previously shown that poor dsDNA unwinding impairs or prevents the ability of the CRISPR / Cas system proteins AnaCas9 or Cas14s to cleave DNA. Thus, without wishing to be bound by any theory, it is possible that the increased DNA cleavage activity by some of the CasX variant proteins of the present invention is at least partially due to the enhanced ability to find and unwind dsDNA at the target site. Methods of measuring the ability of a CasX protein (such as a variant or reference) to unwind DNA include, but are not limited to, in vitro assays that observe the increased association rate of a dsDNA target in fluorescence polarization or biolayer interferometry.
[0363] Without wishing to be bound by theory, it is believed that amino acid changes in the NTSB domain can result in CasX variant proteins with increased DNA unwinding characteristics. Alternatively or additionally, amino acid changes in the OBD or helical domain region that interacts with the PAM can also result in CasX variant proteins with increased DNA unwinding characteristics.
[0364] r. Catalytic activity
[0365] The ribonucleoprotein complex of the CasX:gNA system disclosed herein comprises a reference CasX protein or CasX variant protein complexed with a gNA, the gNA binds to a target nucleic acid and in some cases cleaves the target nucleic acid. In some embodiments, the CasX variant protein has improved catalytic activity relative to the reference CasX protein. Without wishing to be bound by theory, it is believed that in some cases, target strand cleavage can be a limiting factor for Cas12-like molecules to produce a dsDNA break. In some embodiments, the CasX variant protein improves the bending of the target strand of DNA and cleavage of this strand such that the overall efficiency of dsDNA cleavage by the CasX ribonucleoprotein complex is improved.
[0366] In some embodiments, a CasX variant protein has increased nuclease activity compared to a reference CasX protein. Variants with increased nuclease activity can be generated, for example, via amino acid changes in the RuvC nuclease domain. In some embodiments, amino acid substitutions in amino acid residues 708-804 of the RuvC domain can result in increased editing efficiency, as seen in FIG. 10 In some embodiments, a CasX variant comprises a nuclease domain with nickase activity. In the foregoing embodiments, the CasX nickase of the gene editing pair creates a single-strand break within 10-18 nucleotides 3' of the PAM site in the non-target strand. In other embodiments, a CasX variant comprises a nuclease domain with double-strand cleavage activity. In the foregoing, the CasX of the gene editing pair creates a double-strand break within 18-26 nucleotides 5' of the PAM site on the target strand and 10-18 nucleotides 3' on the non-target strand. Nuclease activity can be analyzed by a variety of methods, including those of the examples. In some embodiments, the K cleave nucleotides, or at least 9-fold, or at least 10-fold, compared to a reference or wild-type CasX.
[0367] In some embodiments, a CasX variant protein has increased target strand loading for double-strand cleavage. Variants with increased target strand loading activity can be generated, for example, via amino acid changes in the TLS domain. Without wishing to be bound by theory, amino acid changes in the TSL domain can result in CasX variant proteins with improved catalytic activity. Alternatively or additionally, amino acid changes around the binding channel of the RNA:DNA duplex can also improve the catalytic activity of the CasX variant protein.
[0368] In some embodiments, a CasX variant protein has increased collateral cleavage activity compared to a reference CasX protein. As used herein, “collateral cleavage activity” refers to additional non-targeted cleavage of nucleic acids following recognition and cleavage of a target nucleic acid. In some embodiments, a CasX variant protein has decreased collateral cleavage activity compared to a reference CasX protein.
[0369] In some embodiments, for example those embodiments encompassing applications where cleavage of a target nucleic acid is not a desired outcome, improving the catalytic activity of a CasX variant protein comprises altering, reducing, or eliminating the catalytic activity of the CasX variant protein. In some embodiments, a ribonucleoprotein complex comprising a dCasX variant protein binds to a target nucleic acid and does not cleave the target nucleic acid.
[0370] In some embodiments, a CasX ribonucleoprotein complex comprising a CasX variant protein binds to a target DNA, but generates a single-strand nick in the target DNA. In some embodiments, especially those in which the CasX protein is a nickase, the CasX variant protein has reduced target strand loading for a single-strand nick. Variants with reduced target strand loading can be generated, for example, via amino acid changes in the TSL domain.
[0371] Exemplary methods for characterizing the catalytic activity of a CasX protein can include, but are not limited to, in vitro cleavage assays, including those of the following examples. In some embodiments, electrophoresis of DNA products on agarose gels can interrogate the kinetics of strand cleavage.
[0372] s. Affinity for a target RNA
[0373] In some embodiments, a ribonucleoprotein complex comprising a reference CasX protein or a variant thereof binds to a target RNA and cleaves the target nucleic acid. In some embodiments, a variant of a reference CasX protein increases the specificity of the CasX variant protein for a target RNA and increases the activity of the CasX variant protein for a target RNA when compared to the reference CasX protein. For example, a CasX variant protein can show increased binding affinity for a target RNA, or increased cleavage of a target RNA when compared to a reference CasX protein. In some embodiments, a ribonucleoprotein complex comprising a CasX variant protein binds to a target RNA and / or cleaves a target RNA. In some embodiments, a CasX variant has at least about two-fold to about 10-fold increased binding affinity for a target nucleic acid compared to a reference protein of SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 3.
[0374] t. CasX fusion proteins
[0375] In some embodiments, the present disclosure provides a CasX protein comprising a heterologous protein fused to a CasX. In some cases, the CasX is a reference CasX protein. In other cases, the CasX is a CasX variant of any of the embodiments described herein.
[0376] In some embodiments, a CasX variant protein is fused to one or more proteins or domains thereof having a different activity of interest, resulting in a fusion protein. For example, in some embodiments, a CasX variant protein is fused to a protein (or domain thereof) that inhibits transcription, modifies a target nucleic acid, or modifies a polypeptide associated with a nucleic acid (e.g., histone modification).
[0377] In some embodiments, the CasX variant comprises any one of SEQ ID NOs:247-337, 3301-3493, 3498-3501, 3505-3520, 3540-3549, and 4412-4415 fused to one or more proteins or domains thereof having an activity of interest. In some embodiments, the CasX variant comprises any one of SEQ ID NOs:247-337, 3498-3501, 3505-3520, 3540-3549, and 4412-4415 fused to one or more proteins or domains thereof having an activity of interest. In some embodiments, the CasX variant comprises any one of SEQ ID NOs:3498-3501, 3505-3520, and 3540-3549 fused to one or more proteins or domains thereof having an activity of interest.
[0378] In some embodiments, a heterologous polypeptide (or heterologous amino acid, such as a cysteine residue or a non-natural amino acid) can be inserted into one or more positions within a CasX protein to produce a CasX fusion protein. In other embodiments, a cysteine residue can be inserted into one or more positions within a CasX protein, followed by binding of a heterologous polypeptide as described below. In some alternative embodiments, a heterologous polypeptide or heterologous amino acid can be added at the N-terminus or C-terminus of a reference or CasX variant protein. In other embodiments, a heterologous polypeptide or heterologous amino acid can be inserted into the interior of the sequence of a CasX protein.
[0379] In some embodiments, the reference CasX or variant fusion protein retains RNA- guided sequence-specific target nucleic acid binding and cleavage activity. In some cases, the reference CasX or variant fusion protein has (retains) 50% or greater of the activity (e.g., cleavage and / or binding activity) of the corresponding reference CasX or variant protein without the inserted heterologous protein. In some cases, the reference CasX or variant fusion protein retains at least about 60%, or at least about 70%, at least about 80%, or at least about 90%, or at least about 92%, or at least about 95%, or at least about 98%, or about 100% of the activity (e.g., cleavage and / or binding activity) of the corresponding CasX protein without the inserted heterologous protein.
[0380] In some cases, the reference CasX or CasX variant fusion protein retains (has) target nucleic acid binding activity relative to the activity of the CasX protein without the inserted heterologous amino acid or heterologous polypeptide. In some cases, the reference CasX or CasX variant fusion protein retains at least about 60%, or at least about 70%, at least about 80%, or at least about 90%, or at least about 92%, or at least about 95%, or at least about 98%, or about 100% of the binding activity of the corresponding CasX protein without the inserted heterologous protein.
[0381] In some cases, the reference CasX or CasX variant fusion protein retains (has) target nucleic acid binding and / or cleavage activity relative to the activity of the parent CasX protein without the inserted heterologous amino acid or heterologous polypeptide. For example, in some cases, the reference CasX or CasX variant fusion protein has (retains) 50% or greater of the binding and / or cleavage activity of the corresponding parent CasX protein (CasX protein without the insertion). For example, in some cases, the reference CasX or CasX variant fusion protein has (retains) 60% or greater (70% or greater, 80% or greater, 90% or greater, 92% or greater, 95% or greater, 98% or greater, or 100%) of the binding and / or cleavage activity of the corresponding parent CasX protein (CasX protein without the insertion). Methods of measuring cleavage and / or binding activity of CasX proteins and / or CasX fusion proteins will be known to those of ordinary skill in the art, and any convenient method can be used.
[0382] A variety of heterologous polypeptides are suitable for inclusion in the reference CasX or CasX variant fusion proteins of the present application. In some cases, the fusion partner can modulate transcription of the target DNA (e.g., inhibit transcription, increase transcription). For example, in some cases, the fusion partner is a protein (or domain from a protein) that inhibits transcription (e.g., a transcriptional repressor, a protein that functions by recruiting a transcriptional repressor protein, modifying the target DNA (e.g., methylation), recruiting a DNA modification agent, modulating histones associated with the target DNA, recruiting a histone modification agent (e.g., those that modify acetylation and / or methylation of histones), etc.). In some cases, the fusion partner is a protein (or domain from a protein) that increases transcription (e.g., a transcriptional activator, a protein that functions by recruiting a transcriptional activator protein, modifying the target DNA (e.g., demethylation), recruiting a DNA modification agent, modulating histones associated with the target DNA, recruiting a histone modification agent (e.g., those that modify acetylation and / or methylation of histones), etc.).
[0383] In some cases, the fusion partner has an enzymatic activity that modifies the target nucleic acid; e.g., nuclease activity, methyltransferase activity, demethylase activity, DNA repair activity, DNA damage activity, deaminase activity, glycosylase activity, alkylating activity, depurination activity, oxidizing activity, pyrimidine dimer formation activity, integrase activity, transposase activity, recombinase activity, polymerase activity, ligase activity, helicase activity, photolyase activity, or glycosidase activity.
[0384] In some cases, the fusion partner has an enzymatic activity that modifies a polypeptide associated with the target nucleic acid (e.g., a histone); e.g., a methyltransferase activity, a demethylase activity, an acetyltransferase activity, a deacetylase activity, a kinase activity, a phosphatase activity, a ubiquitin ligase activity, a deubiquitinating activity, an adenylation activity, a deadenylation activity, a SUMOylation activity, a desumoylation activity, a ribosylation activity, a deribosylation activity, a myristoylation activity, or a demyristoylation activity. In some embodiments, the CasX variant comprises any one of SEQ ID NOs: 247-337, 3301-3493, 3498-3501, 3505-3520, 3540-3549, and 4412-4415, and a polypeptide having a methyltransferase activity, a demethylase activity, an acetyltransferase activity, a deacetylase activity, a kinase activity, a phosphatase activity, a ubiquitin ligase activity, a deubiquitinating activity, an adenylation activity, a deadenylation activity, a SUMOylation activity, a desumoylation activity, a ribosylation activity, a deribosylation activity, a myristoylation activity, or a demyristoylation activity. In some embodiments, the CasX variant comprises any one of SEQ ID NOs: 247-337, 3498-3501, 3505-3520, 3540-3549, and 4412-4415, and a polypeptide having a methyltransferase activity, a demethylase activity, an acetyltransferase activity, a deacetylase activity, a kinase activity, a phosphatase activity, a ubiquitin ligase activity, a deubiquitinating activity, an adenylation activity, a deadenylation activity, a SUMOylation activity, a desumoylation activity, a ribosylation activity, a deribosylation activity, a myristoylation activity, or a demyristoylation activity. In some embodiments, the CasX variant comprises any one of SEQ ID NOs: 3498-3501, 3505-3520, and 3540-3549, and a polypeptide having a methyltransferase activity, a demethylase activity, an acetyltransferase activity, a deacetylase activity, a kinase activity, a phosphatase activity, a ubiquitin ligase activity, a deubiquitinating activity, an adenylation activity, a deadenylation activity, a SUMOylation activity, a desumoylation activity, a ribosylation activity, a deribosylation activity, a myristoylation activity, or a demyristoylation activity.
[0385] Examples of proteins (or fragments thereof) that can be used as fusion partners suitable for reference CasX or CasX variants to increase transcription include, but are not limited to: transcriptional activators, such as VP16, VP64, VP48, VP160, p65 subdomain (e.g., from NFkB), and activation domains of EDLL and / or transcription activator-like (TAL) activation domains (e.g., for activity in plants); histone lysine methyltransferases, such as SET domain-containing 1A, histone lysine methyltransferase (SET1A), SET domain-containing 1B, histone lysine methyltransferase (SET1B), lysine methyltransferase 2A (MLL1) through 5, ASCL1 (ASH1) achaete-scute family bHLH transcription factor 1 (ASH1), SET and MYND domain- containing 2 provided (SMYD2), nuclear receptor SET domain protein 1 (NSD1), and the like; histone lysine demethylases, such as lysine demethylase 3A (JHDM2a) / lysine-specific demethylase 3B (JHDM2b), lysine demethylase 6A (UTX), lysine demethylase 6B (JMJD3), and the like; histone acetyltransferases, such as lysine acetyltransferase 2A (GCN5), lysine acetyltransferase 2B (PCAF), CREB-binding protein (CBP), E1A-binding protein p30 (p300), TATA-box binding protein-associated factor 1 (TAF1), lysine acetyltransferase 5 (TIP60 / PLIP), lysine acetyltransferase 6A (MOZ / MYST3), lysine acetyltransferase 6B (MORF / MYST4), SRC proto-oncogene, non-receptor tyrosine kinase (SRC1), nuclear receptor coactivator 3 (ACTR), MYB binding protein 1a (P160), clock circadian regulator (CLOCK), and the like; and DNA demethylases, such as ten-eleven translocation (TET) dioxygenase 1 (TET1CD), tet methylcytosine dioxygenase 1 (TET1), demeter (DME), demeter-like 1 (DML1), demeter-like 2 (DML2), protein ROS1 (ROS1), and the like.
[0386] Examples of proteins (or fragments thereof) that can be used as fusion partners suitable for reference to CasX or CasX variants to reduce transcription include, but are not limited to: transcriptional repressors, such as Kruppel-associated box (KRAB or SKD); KOX1 repressor domain; Madm SIN3 interaction domain (SID); ERF repressor domain (ERD), SRDX repressor domain (e.g., for repression in plants), and the like; histone lysine methyltransferases, such as PR / SET domain-containing protein (Pr-SET) 7 / 8, lysine methyltransferase 5B (SUV4-20H1), PR / SET domain 2 (RIZ1), and the like; histone lysine demethylases, such as lysine demethylase 4A (JMJD2A / JHDM3A), lysine demethylase 4B (JMJD2B), lysine demethylase 4C (JMJD2C / GASC1), lysine demethylase 4D (JMJD2D), lysine demethylase 5A (JARID1A / RBP2), lysine demethylase 5B (JARID1B / PLU-1), lysine demethylase 5C (JARID 1C / SMCX), lysine demethylase 5D (JARID1D / SMCY), and the like; histone lysine deacetylases, such as histone deacetylase 1 (HDAC1), HDAC2, HDAC3, HDAC8, HDAC4, HDAC5, HDAC7, HDAC9, Sirtuin 1 (SIRT1), SIRT2, HDAC11, and the like; DNA methylases, such as Hhal DNA m5c-methyltransferase (M. Hhal), DNA methyltransferase 1 (DNMT1), DNA methyltransferase 3a (DNMT3a), DNA methyltransferase 3b (DNMT3b), Methyltransferase 1 (MET1), S-adenosyl-L-methionine-dependent methyltransferase superfamily protein (DRM3) (plant), DNA cytosine methyltransferase MET2a (ZMET2), chromatin methylase 1 (CMT1), chromatin methylase 2 (CMT2) (plant), and the like; and margin recruiting elements, such as lamin A, lamin B, and the like.
[0387] In some cases, the fusion partner to a reference CasX or CasX variant has an enzymatic activity that modifies a target nucleic acid (e.g., ssRNA, dsRNA, ssDNA, dsDNA). Examples of enzymatic activities that can be provided by a fusion partner include, but are not limited to: nuclease activity, such as provided by restriction enzymes (e.g., Fokl nuclease); methyltransferase activity, such as provided by methyltransferases (e.g., M. Hhal, DNA methyltransferase 1 (DNMT1), DNA methyltransferase 3a (DNMT3a), DNA methyltransferase 3b (DNMT3b), METI, DRM3 (plant), ZMET2, CMT1, CMT2 (plant), etc.); demethylase activity, such as provided by demethylases (e.g., ten-eleven translocation (TET) dioxygenase 1 (TET 1CD), TET1, DME, DML1, DML2, ROS1, etc.); DNA repair activity; DNA damage activity; deaminase activity, such as provided by deaminases (e.g., cytosine deaminases, such as APOBEC proteins, such as rat apolipoprotein B mRNA-editing enzyme, catalytic polypeptide 1 {APOBEC1}); racemase activity; alkylation activity; depurination activity; oxidation activity; pyrimidine dimer formation activity; integrase activity, such as provided by integrases and / or resolvases (e.g., Gin invertase, such as the hyperactive mutant of Gin invertase, Gin H106Y; human immunodeficiency virus type 1 integrase (IN); Tn3 resolvase; etc.); transposase activity; recombinase activity, such as provided by recombinases (e.g., catalytic domain of Gin recombinase); polymerase activity; ligase activity; helicase activity; photolyase activity; and glycosidase activity.
[0388] In some cases, a reference CasX or CasX variant of the application is fused to a polypeptide selected from the group consisting of a domain that increases transcription (e.g., a VP16 domain, a VP64 domain), a domain that decreases transcription (e.g., a KRAB domain, such as from the Koxl protein), a nuclear catalytic domain of a histone acetyltransferase (e.g., histone acetyltransferase p300), a protein / domain that provides a detectable signal (e.g., a fluorescent protein, such as GFP), a nuclease domain (e.g., a Fokl nuclease), and a base editor (further discussed below).
[0389] In some embodiments, the CasX variant comprises any one of SEQ ID NOS: 247-337, 3301-3493, 3498-3501, 3505-3520, 3540-3549, and 4412-4415 fused to a polypeptide selected from the group consisting of a domain that reduces transcription, a domain with enzymatic activity, a nucleic catalytic domain of a histone acetyltransferase, a protein / domain that provides a detectable signal, a nuclease domain, and a base editor. In some embodiments, the CasX variant comprises any one of SEQ ID NOS: 247-337, 3498-3501, 3505-3520, 3540-3549, and 4412-4415 fused to a polypeptide selected from the group consisting of a domain that reduces transcription, a domain with enzymatic activity, a nucleic catalytic domain of a histone acetyltransferase, a protein / domain that provides a detectable signal, a nuclease domain, and a base editor. In some embodiments, the CasX variant comprises any one of SEQ ID NOS: 3498-3501, 3505-3520, and 3540-3549 fused to a polypeptide selected from the group consisting of a domain that reduces transcription, a domain with enzymatic activity, a nucleic catalytic domain of a histone acetyltransferase, a protein / domain that provides a detectable signal, a nuclease domain, and a base editor.
[0390] In some cases, a reference CasX protein or C...
Claims
1. A chimeric CasX variant protein derived from a first reference CasX sequence as shown in SEQ ID NO: 2, comprising a domain derived from a second, different reference CasX protein, wherein the chimeric CasX variant protein comprises: a) Replace amino acid positions 103-192 of the NTSB domain of SEQ ID NO: 2 with amino acids from positions 101-191 of the non-targeted strand binding (NTSB) domain of the second reference CasX sequence as shown in SEQ ID NO: 1; b) A chimeric helical I domain comprising amino acids at positions 59-102 of SEQ ID NO: 2, and amino acids at positions 193-333 of the helical Ib domain of SEQ ID NO: 1 replaced by amino acids at positions 192-332 of the helical Ib domain of SEQ ID NO: 1; c) Target Stock Load (TSL) domain of SEQ ID NO:2; d) The spiral II domain of SEQ ID NO:2; e) Oligonucleotide binding domain (OBD) of SEQ ID NO:2; f) The RuvC cleavage domain of SEQ ID NO:2; and g) Compared with the reference CasX sequence of SEQ ID NO: 2, at least two or more modifications are selected from the following: (i) Amino acid substitutions in L379R; (ii) Amino acid substitutions in A708K; (iii) Amino acid substitutions in T620P; (iv) Amino acid substitutions in E385P; (v) Amino acid substitutions in Y857R; (vi) Amino acid substitutions in I658V; (vii) Amino acid substitutions in F399L; (viii) Amino acid substitutions in L404K; and (ix) Amino acid deletion of P793.
2. The chimeric CasX variant protein according to claim 1, wherein, compared to the CasX protein of SEQ ID NO: 2, the chimeric CasX variant protein exhibits one or more modified features, wherein the one or more modified features are selected from the group consisting of: improved target DNA editing, improved target DNA cleavage rate, increased formation of cleavage competency ribonucleoprotein (RNP) complex, improved utilization of preseptal neighbor motif (PAM), and improved solubility.
3. The chimeric CasX variant protein of claim 1, further comprising one or more nuclear localization signals (NLS).
4. The chimeric CasX variant protein according to claim 3, wherein the one or more NLS are selected from the group consisting of the following sequences: PKKKRKV (SEQ ID NO: 352), KRPAATKKAGQAKKKK (SEQ ID NO: 353), PAAKRVKLD (SEQ ID NO: 354), RQRRNELKRSP (SEQ ID NO: 355), NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY (SEQ ID NO: 356), RMRIZFKNKGKDTAELRRRRVEVSVELRKAKKDEQILKRRNV (SEQ ID NO: 357), VSRKRPRP (SEQ ID NO: 358), PPKKARED (SEQ ID NO: 35), PQPKKKPL (SEQ ID NO: 360), SALIKKKKKMAP (SEQ ID NO: 361), DRLRR (SEQ ID NO: 358), PPKKARED (SEQ ID NO: 35), PQPKKKPL (SEQ ID NO: 360), SALIKKKKKMAP (SEQ ID NO: 361), DRLRR (SEQ ID NO: 350 ... 362), PKQKKRK (SEQ ID NO: 363), RKLKKKIKKL (SEQ ID NO: 364), REKKKFLKRR (SEQ ID NO: 365), KRKGDEVDGVDEVAKKKSKK (SEQ ID NO: 366), RKCLQAGMNLEARKTKK (SEQ ID NO: 367), PRPRKIPR (SEQ ID NO: 368), PPRKKRTVV (SEQ ID NO: 369), NLSKKKKRKREK (SEQ ID NO: 370), RRPSRPFRKP (SEQ ID NO: 371), KRPRSPSS (SEQ ID NO: 372), KRGINDRNFWRGENERKTR (SEQ ID NO: 373), PRPPKMARYDN (SEQ ID NO: 374), KRSFSKAF (SEQ ID NO: 375), KLKIKRPVK (SEQ ID NO: 376), PKKKRKVPPPPAAKRVKLD (SEQ ID NO: 377), PKTRRRPRRSQRKRPPT (SEQ ID NO: 378), SRRRKANPTKLSENAKKLAKEVEN (SEQ ID NO: 379), KTRRRPRRSQRKRPPT (SEQ ID NO: 380), RRKKRRPRRKKRR (SEQ ID NO: 381), PKKKSRKPKKKSRK (SEQ ID NO:382), HKKKHPDASVNFSEFSK (SEQ ID NO: 383), QRPGPYDRPQRPGPYDRP (SEQ ID NO: 384), LSPSLSPLLSPSLSPL (SEQ ID NO: 385), RGKGGKGLGKGGAKRHRK (SEQ ID NO: 386), PKRGRGRPKRGRGR (SEQ ID NO: 387) and PKKKRKVPPPPKKKRKV (SEQ ID NO: 389).
5. The chimeric CasX variant protein of claim 4, wherein one or more NLS are located at or near the C-terminus of the CasX variant protein.
6. The chimeric CasX variant protein of claim 4, wherein one or more NLS are located at or near the N-terminus of the CasX variant protein.
7. The chimeric CasX variant protein of claim 4, comprising at least two NLS, wherein the at least two NLS are located at or near the N-terminus and the C-terminus of the CasX variant protein.
8. The chimeric CasX variant protein of claim 1, wherein, compared to the editing efficiency and / or binding of an RNP complex containing a reference CasX protein in a similar analytical system, the RNP complex containing the CasX variant protein exhibits greater editing efficiency and / or binding to the target sequence in the target DNA when any one of the PAM sequences TTC, ATC, GTC, or CTC is located on the 5' of the non-target strand of the prespacer that is consistent with the target sequence of the guide nucleic acid gNA in the cell analytical system and is 1 nucleotide.
9. The chimeric CasX variant of claim 8, wherein the PAM sequence is TTC.
10. The chimeric CasX variant of claim 8, wherein the PAM sequence is ATC.
11. The chimeric CasX variant of claim 8, wherein the PAM sequence is CTC.
12. The chimeric CasX variant of claim 8, wherein the PAM sequence is GTC.
13. The chimeric CasX variant protein of claim 8, wherein the improved editing efficiency and / or binding affinity of the target DNA of the RNP complex containing the chimeric CasX variant protein is improved by at least 1.1 to 100 times relative to the RNP complex containing the reference CasX.
14. The chimeric CasX variant protein of claim 1, wherein the chimeric CasX variant protein is a non-catalytically active CasX (dCasX) protein, and wherein the dCasX and the guiding nucleic acid gNA retain the ability to bind to the target DNA.
15. The chimeric CasX variant protein of claim 14, wherein the dCasX comprises a mutation at the following residues: a. D672, E769, and D935 of the CasX protein corresponding to SEQ ID NO: 1; or b. D659, E756, and D922 of the CasX protein corresponding to SEQ ID NO:
2.
16. The chimeric CasX variant protein of claim 15, wherein the mutation is a substitution of an alanine residue.
17. The chimeric CasX variant protein of claim 1, comprising a heterologous protein or a domain thereof fused to the chimeric CasX variant protein.
18. The chimeric CasX variant protein of claim 17, wherein the heterologous protein or its domain is a base editing agent.
19. The chimeric CasX variant protein of claim 18, wherein the base editing agent is adenosine deaminase, cytosine deaminase, or guanine oxidase.
20. The chimeric CasX variant protein of claim 1, wherein the chimeric CasX variant protein forms an RNP complex with guide ribonucleic acid (gRNA).
21. The chimeric CasX variant protein of claim 20, wherein the chimeric CasX variant protein comprises the sequence of any one of SEQ ID NO:333-336, and retains the sequence-specific target nucleic acid binding and cleavage activity guided by the guide nucleic acid gNA.
22. A gene editing pair comprising a chimeric CasX variant protein according to any one of claims 1 to 21 and a guide ribonucleic acid (gRNA), wherein the gRNA comprises a targeting sequence complementary to the target DNA, and wherein the CasX and the gRNA are capable of binding together in an RNP complex.
23. The gene editing pair of claim 22, wherein, compared to a gene editing pair comprising a reference CasX protein containing the sequence of SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 3 and a reference guide RNA containing the sequence of SEQ ID NO: 4 or 5, the gene editing pair has one or more modified features, wherein the one or more modified features include improved stability of the CasX:guide RNA gNA complex, a higher percentage of cleavage-competent RNPs, the ability to utilize a wider range of PAM sequences, increased solubility, increased cleavage rate, increased editing activity, improved editing efficiency, improved editing specificity, increased nuclease activity, or reduced off-target cleavage.
24. Use of the gene editing pair of claim 22, wherein the use comprises contacting target DNA with the gene editing pair, wherein the contact causes editing or modification of the target DNA, wherein the editing occurs in vitro.
25. Use of the gene editing pair of claim 22, wherein the use includes contacting target DNA with the gene editing pair, wherein the contact causes editing or modification of the target DNA, wherein the editing occurs in vitro.
26. The use of the gene editing method of claim 22 for editing or modifying target DNA in cells, wherein the editing occurs in vitro or inside cells, and wherein the cells do not include human germ cells or human fetal cells.
27. The use of the gene editing method of claim 22 for editing or modifying target DNA in cells, wherein the editing occurs ex vivo inside the cell, and wherein the cell does not include human germ cells or human fetal cells.
28. The use according to claim 26 or 27, wherein the cell is a eukaryotic cell.
29. The use according to claim 28, wherein the eukaryotic cell is a human cell.
30. The use according to claim 29, wherein the cell is a fibroblast, glial cell, neuron, myocyte, osteocyte, hepatocyte, pancreatic cell, retinal cell, cancer cell, T cell, B cell, NK cell, adipocyte, epithelial cell, endothelial cell, mesothelial cell, chondrocyte, stem cell, or monocyte.
31. The use according to claim 26 or 27, wherein the editing alters a mutation in the wild-type allele of the gene.
32. The use according to claim 26 or 27, wherein the editing knocks out or removes the allele of the gene.
33. The use according to claim 28, wherein the use comprises contacting the eukaryotic cell with a vector encoding or comprising the chimeric CasX variant protein and the gRNA, and optionally further comprising a donor template.
34. The use according to claim 33, wherein the vector is an adeno-associated virus (AAV) vector.
35. The use according to claim 34, wherein the AAV is AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV-Rh74, or AAVRh10.
36. The use according to claim 33, wherein the vector is a lentiviral vector.
37. The use according to claim 33, wherein the carrier is a non-viral particle.
38. The use according to claim 33, wherein the carrier is a virus-like particle (VLP).
39. A polynucleotide encoding a chimeric CasX variant protein according to any one of claims 1 to 21.
40. A carrier comprising the polynucleotide according to claim 39.
41. The vector according to claim 40, wherein the vector is an adeno-associated virus (AAV) vector.
42. The carrier according to claim 41, wherein the AAV is AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV-Rh74, or AAVRh10.
43. The vector according to claim 40, wherein the vector is a lentiviral vector.
44. The vector of claim 40, wherein the vector is a virus-like particle (VLP).
45. The carrier according to claim 40, wherein the carrier is a non-viral particle.
46. A cell comprising the polynucleotide of claim 39, or the carrier of any one of claims 40 to 45.
47. A composition comprising the chimeric CasX variant protein according to any one of claims 1 to 21.
48. The composition of claim 47, further comprising a gRNA variant and a targeting sequence.
49. The composition of claim 48, wherein the chimeric CasX variant protein and the gRNA variant are bound together in an RNP complex.
50. A composition comprising the gene editing pair according to claim 22 or 23.
51. The composition of claim 50, further comprising a donor template nucleic acid, wherein the donor template comprises a nucleotide sequence homologous to the target DNA.
52. The composition according to any one of claims 50, further comprising a buffer, a nuclease inhibitor, a protease inhibitor, a liposome, a therapeutic agent, a label, a chromogenic label, or any combination thereof.
53. A kit comprising a chimeric CasX variant protein and a container according to any one of claims 1 to 21.
54. The kit according to claim 53, further comprising a gRNA variant and a targeting sequence.
55. The kit of claim 54, further comprising a donor template nucleic acid, wherein the donor template comprises a nucleotide sequence homologous to a target sequence of the target DNA.
56. The kit according to claim 54, further comprising a buffer, a nuclease inhibitor, a protease inhibitor, a liposome, a therapeutic agent, a label, a chromogenic label, or any combination thereof.
57. A kit comprising the composition according to any one of claims 47 to 52.
Citation Information
Patent Citations
Navigational systems
CA485491A
Crispr compositions and methods of using the same for gene therapy
US20180258424A1
Modified site-directed modifying polypeptides and methods of use thereof
US20180363009A1
Splicing factors with a PUF protein RNA-binding domain and a splicing effector domain and uses of same
WO2010075303A1
Peptides for the specific binding of RNA targets
WO2012068627A1