Engineered CasX systems
CasX variants with modified domains and guide nucleic acids enhance gene editing efficiency by up to 81-fold, addressing the need for improved Class 2 CRISPR/Cas systems for therapeutic and research applications.
Patent Information
- Authority / Receiving Office
- AU · AU
- Patent Type
- Applications
- Current Assignee / Owner
- SCRIBE THERAPEUTICS INC
- Filing Date
- 2020-06-05
- Publication Date
- 2026-07-23
Smart Images

Figure 00000001_0000 
Figure 00000442_0000 
Figure 00000443_0000
Abstract
Description
CROSS REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority to U.S. provisional patent application numbers 62,858,750, filed on June 7, 2019, 62 / 944,892, filed on December 6, 2019 and 63 / 030,838, filed on May 27, 2020, the contents of each of which are incorporated herein by reference in their entireties. INCORPORATION BY REFERENCE OF SEQUENCE LISTING
[0002] This application contains a Sequence listing which has been submitted in ASCII format via EFS-WEB and is hereby incorporated by reference in its entirety. Said ASCII copy, created on June 5, 2020 is named SCRB01 l_03WO_SeqList_25 and is 3.63 MB in size. BACKGROUND
[0003] The CRISPR-Cas systems confer bacteria and archaea with acquired immunity against phage and viruses. Intensive research over the past decade has uncovered the biochemistry of these systems. CRISPR-Cas systems consist of Cas proteins, which are involved in acquisition, targeting and cleavage of foreign DNA or RNA, and a CRISPR array, which includes direct repeats flanking short spacer sequences that guide Cas proteins to their targets. Class 2 CRISPR-Cas are streamlined versions in which a single Cas protein bound to RNA is responsible for binding to and cleavage of a targeted sequence. The programmable nature of these minimal systems has facilitated their use as a versatile technology that is revolutionizing the field of genome manipulation.
[0004] To date, only a few Class 2 CRISPR / Cas systems have been discovered that have been widely used. Thus, there is a need in the art for additional Class 2 CRISPR / Cas systems (e.g., Cas protein plus guide RNA combinations) that have been optimized and / or offer improvements over earlier generation systems for utilization in a variety of therapeutic, diagnostic, and research applications. SUMMARY
[0005] In some aspects, the present disclosure provides variants of a reference CasX nuclease protein, wherein the CasX variant is capable of forming a complex with a guide nucleic acid (NA), and wherein the complex can bind a target DNA, wherein the target DNA comprises nontarget strand and a target strand, and wherein the CasX variant comprises at least one modification relative to a domain of the reference CasX and exhibits one or more improved characteristics as compared to the reference CasX protein. The domains of the reference CasX protein include: (a) a non-target strand binding (NTSB) domain that binds to the non-target strand of DNA, wherein the NTSB domain comprises a four-stranded beta sheet; (b) a target strand loading (TSL) domain that places the target DNA in a cleavage site of the CasX variant, the TSL domain comprising three positively charged amino acids, wherein the three positively charged amino acids bind to the target strand of DNA, (c) a helical I domain that interacts with both the target DNA and a spacer region of a guide NA, wherein the helical I domain comprises one or more alpha helices; (d) a helical II domain that interacts with both the target DNA and a scaffold stem of the guide NA; (e) an oligonucleotide binding domain (OBD) that binds a triplex region of the guide NA; and (f) a RuvC DNA cleavage domain.
[0006] In some aspects, the present disclosure provides variants of a reference guide nucleic acid (gNA) capable of binding a CasX protein, wherein the reference guide nucleic acid comprises at least one modification in a region compared to the reference guide nucleic acid sequence, and the variant exhibits one or more improved characteristics compared to the reference guide RNA. The regions of the scaffold of the gNA include: (a) an extended stem loop; (b) a scaffold stem loop; (c) a triplex; and (d) pseudoknot. In some cases, the scaffold stem of the variant gNA further comprises a bubble. In other cases, the scaffold of the variant gNA further comprises a triplex loop region. In other cases, the scaffold of the variant gNA further comprises a 5' unstructured region.
[0007] In some aspects, the present disclosure provides gene editing pairs comprising the CasX proteins and gNAs of any of the embodiments described herein.
[0008] In some aspects, the present disclosure provides polynucleotides and vectors encoding the CasX proteins, gNAs and gene editing pairs described herein. In some embodiments, the vectors are viral vectors such as an Adeno-Associated Viral (AAV) vector or a lentiviral vector. In other embodiments, the vectors are non-viral particles such as virus-like particles or nanoparticles.
[0009] In some aspects, the present disclosure provides cells comprising the polynucleotides, vectors, CasX proteins, gNAs and gene editing pairs described herein. In other aspects, the present disclosure provides cells comprising target DNA edited by the methods of editing embodiments described herein.
[0010] In some aspects, the present disclosure provides kits comprising the polynucleotides, vectors, CasX proteins, gNAs and gene editing pairs described herein.
[0011] In some aspects, the present disclosure provides methods of editing a target DNA, comprising contacting the target DNA with one or more of the gene editing pairs described herein, wherein the contacting results in editing of the target DNA.
[0012] In other aspects, the disclosure provides methods of treatment of a subject in need thereof, comprising administration of the gene editing pairs or vectors comprising or encoding the gene editing pairs of any of the embodiments described herein.
[0013] In another aspect, provided herein are gene editing pairs, compositions comprising gene editing pairs, or vectors comprising or encoding gene editing pairs, for use as a medicament.
[0014] In another aspect, provided herein are gene editing pairs, compositions comprising gene editing pairs, or vectors comprising or encoding gene editing pairs, for use in a method of treatment, wherein the method comprises editing or modifying a target DNA; optionally wherein the editing occurs in a subject having a mutation in an allele of a gene wherein the mutation causes a disease or disorder in the subject, preferably wherein the editing changes the mutation to a wild type allele of the gene or knocks down or knocks out an allele of a gene causing a disease or disorder in the subject. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The novel features of the invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention are utilized, and the accompanying drawings of which:
[0016] FIG. lisa diagram showing an exemplary method of making CasX protein and guide RNA variants of the disclosure using Deep Mutational Evolution (DME). In some exemplary embodiments, DME builds and tests nearly every possible mutation, insertion and deletion in a biomolecule and combinations / multiples thereof, and provides a near comprehensive and unbiased assessment of the fitness landscape of a biomolecule and paths in sequence space towards desired outcomes. As described herein, DME can be applied to both CasX protein and guide RNA.
[0017] FIG. 2 is a diagram and an example fluorescence activated cell sorting (FACS) plot illustrating an exemplary method for assaying the effectiveness of a reference CasX protein or single guide RNA (sgRNA), or variants thereof. A reporter (e.g. GFP reporter) coupled to a gRNA target sequence, complementary to the gRNA spacer, is integrated into a reporter cell line. Cells are transformed or transfected with a CasX protein and / or sgNA variant, with the spacer motif of the sgRNA complementary to and targeting the gRNA target sequence of the reporter. Ability of the CasX: sgRNA ribonucleoprotein complex to cleave the target sequence is assayed by FACS. Cells that lose reporter expression indicate occurrence of CasX:sgRNA ribonucleoprotein complex-mediated cleavage and indel formation.
[0018] FIG. 3 A and FIG. 3B are heat maps showing the results of an exemplary DME mutagenesis of the reference sgRNA encoded by SEQ ID NO: 5, as described in Example 3. FIG. 3 A shows the effect of single base pair (single base) substitutions, double base pair (double base) substitutions, single base pair insertions, single base pair deletions, and a single base pair deletion plus at single base pair substitution at each position of the reference sgRNA shown at top. FIG. 3B shows the effect of double base pair insertions and a single base pair insertion plus a single base pair substitution at each position of the improved reference sgRNA. The reference sgRNA sequence of SEQ ID NO: 5 is shown at the top of FIG. 3A and bottom of FIG. 3B. In FIG. 3 A and FIG. B, Log2 fold enrichment of the variant in the DME library relative to the reference sgRNA following selection is indicated in grayscale. Enrichment is a proxy for activity, where greater enrichment is a more active molecule. The results show regions of the reference sgRNA that should not be mutated and key regions that are targeted for mutagenesis.
[0019] FIG. 4A shows the results of exemplary DME experiments using a reference sgRNA, as described in Example 3. The improved reference sgNA (an sgRNA) with a sequence of SEQ ID NO: 5 is shown at top, and Log2 fold enrichment of the variant in the DME library relative to the reference sgRNA following selection is indicated in grayscale. Enrichment is a proxy for activity, where greater enrichment is a more active molecule. The heat map shows an exemplary DME experiment showing four replicates of a library where every base pair in the reference sgRNA has been substituted with every possible alternative base pair.
[0020] FIG. 4B is a series of 8 plots that compare biological replicates of different DME libraries. The Log2 fold enrichment of individual variants relative to the reference sgRNA sequence for pairs of DME replicates are plotted against each other. Shown are plots for single deletion, single insertion and single substitution DME experiments, as well as wild type controls, and the plots indicate that there is a good amount of agreement for each replicate.
[0021] FIG. 4C is a heat map of an exemplary DME experiment showing four replicates of a library where every location in the reference sgRNA has undergone a single base pair insertion. The DME experiment used a reference sgRNA of SEQ ID NO: 5 (at top), and was performed as described in Example 3. Log2 fold enrichment of the variant in the DME library relative to the reference sgRNA following selection is indicated in grayscale.
[0022] FIGS. 5 A-5E are a series of plots showing that sgNA variants can improve gene editing by greater than two fold in an EGFP disruption assay, as described in Examples 2 and 3. Editing was measured by indel formation and GFP disruption in HEK293 cells carrying a GFP reporter. FIG. 5 A shows the fold change in editing efficiency of a CasX sgRNA reference of SEQ ID NO: 4 and a variant of the reference which has a sequence of SEQ ID NO: 5, across 10 targets. When averaged across 10 targets, the editing efficiency of sgRNA SEQ ID NO: 5 improved 176% compared to SEQ ID NO: 4. FIG. 5B shows that further improvement of the sgRNA scaffold of SEQ ID NO: 5 is possible by swapping the extended stem loop sequence for additional sequences to generate the scaffolds whose sequences are shown in Table 2. Fold change in editing efficiency is shown on the Y-axis. FIG. 5C is a plot showing the fold improvement of sgNA variants (including a variant with SEQ ID NO: 17) generated by DME mutations normalized to SEQ ID NO: 5 as the CasX reference sgRNA. FIG. 5D is a plot showing the fold improvement of sgNA variants of sequences listed in Table 2, which were generated by appending ribozyme sequences to the reference sgRNA sequence, normalized to SEQ ID NO: 5 as the CasX reference sgRNA. FIG. 5E is a plot showing the fold improvement normalized to the SEQ ID NO: 5 reference sgRNA of variants created by both combining (stacking) scaffold stem mutations showing improved cleavage, DME mutations showing improved cleavage, and using ribozyme appendages showing improved cleavage. The resulting sgNA variants yield 2 fold or greater improvement in cleavage compared to SEQ ID NO: 5 in this assay. EGFP editing assays were performed with spacer target sequences of E6 and E7.
[0023] FIG. 6 shows a Hepatitis Delta Virus (HDV) genomic ribozyme used in exemplary gNA variants (SEQ ID NOs: 18-22).
[0024] FIGS. 7A-7I are a series of heat maps showing the effect of single amino acid substitutions, single amino acid insertions, and deletions at each amino acid position in a reference CasX protein of SEQ ID NO: 2, as described in Example 4. Data were generated by a DME assay run at 37°C. The Y-axis shows each possible substitution or insertion (from top to bottom: R, H, K, D, E, S, T, N, Q, C, G, P, A, I, L, M, F, W, Y or V; boxes indicate the amino acid identity of the reference protein), the X-axis shows the amino acid position in the reference CasX protein. Log2 fold enrichment of the CasX variant protein relative to the reference CasX protein of SEQ ID NO: 2 in a DME library following enrichment is indicated. As used herein, “enrichment” is a proxy for activity, where greater enrichment is a more active molecule. (*)s indicate active sites. FIGS. 7A-7D show the effect of single amino acid substitutions. FIGS. 7E-7H show the effect of single amino acid insertions. FIG. 71 shows the effect of single amino acid deletions.
[0025] FIGS. 8A-8C are a series of heat maps showing the effect of single amino acid substitutions, single amino acid insertions and deletions at each amino acid position in a reference CasX protein of SEQ ID NO: 2, as described in Example 4. Data were generated by a DME assay run at 45°C. FIG. 8A shows the effect of single amino acid substitutions. FIG. 8B shows the effect of single amino acid insertions. FIG. 8C shows the effect of single amino acid deletions. For all of FIGS. 8A- 8C, The Y-axis shows each possible substitution or insertion (from top to bottom: R, H, K, D, E, S, T, N, Q, C, G, P, A, I, L, M, F, W, Y or V; boxes indicate the amino acid identity of the reference protein), the X-axis shows the amino acid position in the reference CasX protein. Log2 fold enrichment of the CasX variant protein relative to the reference CasX protein of SEQ ID NO: 2 in a DME library following enrichment is indicated in grayscale, where greater enrichment is a more active molecule. (*)s indicate active sites. Running this assay at 45 °C enriches for different variants than running the same assay at 37 °C (see FIGS. 7A-7I), thereby indicating which amino acid residues and changes are important for thermostability and folding.
[0026] FIG. 9 shows a survey of the comprehensive mutational landscape of all single mutations of a reference CasX protein of SEQ ID NO: 2. On the Y-axis, fold enrichment of CasX variants relative to the reference CasX protein for single substitutions (top), single insertions (middle) or single deletions (bottom). On the X-axis, amino acid position in the reference CasX protein. Key regions that yield improved CasX variants are the initial helix region and regions in the RuvC domain bordering the target strand loading (TLS) domain, as well as others.
[0027] FIG. 10 is a plot showing that the evaluated CasX variant proteins improved editing greater than three-fold relative to a reference CasX protein in the EGFP disruption assay, as described in Example 5. CasX proteins were tested for their ability to cleave an EGFP reporter at 2 different target sites in human HEK293 cells, and the normalized improvement in genome editing at these sites over the basic reference CasX protein of SEQ ID NO: 2 is shown. Variants, from left to right (indicated by the amino acid substitution, insertion or deletion at the given residue number) are: Y789T, [P793], Y789D, T72S, I546V, E552A, A636D, F536S, A708K, Y797L, L792G, A739V, G791M, AG661, A788W, K390R, A751S, E385A, AP696, AM773, G695H, AAS793, AAS795, C477R, C477K, C479A, C479L, I55F, K210R, C233S, D231N, Q338E, Q338R, L379R, K390R, L481Q, F495S, D600N, T886K, A739V, K460N, I199F, G492P, T153I, R591I, AAS795, AAS796, AL889, E121D, S270W, E712Q, K942Q, E552K, K25Q, N47D, AT696, L685I, N880D, Q102R, M734K, A724S, T704K, P224K, K25R, M29E, H152D, S219R, E475K, G226R, A377K, E480K, K416E, H164R, K767R, I7F, M29R, H435R, E385Q, E385K, I279F, D489S, D732N, A739T, W885R, E53K, A238T, P283Q, E292K, Q628E, R388Q, G791M, L792K, L792E, M779N, G27D, K955R, S867R, R693I, F189Y, V635M, F399L, E498K, E386S, V254G, P793S, K188E, QT945KI, T620P, T946P, TT949PP, N952T, K682E, K975R, L212P, E292R, I303K, C349E, E385P, E386N, D387K, L404K, E466H, C477Q, C477H, C479A, D659H, T806V, K808S, AAS797, V959M, K975Q, W974G, A708Q, V711K, D733T, L742W, V747K, F755M, M771A, M771Q, W782Q, G791F, L792D, L792K, P793Q, P793G, Q804A, Y966N, Y723N, Y857R, S890R, S932M, L897M, R624G, S603G, N737S, L307K, I658V APT688, ASA794, S877R, N580T, V335G, T620S, W345G, T280S, L406P, A612D, A751S, E386R, V351M, K210N, D40A, E773G, H207L, T62A, T287P, T832A, A893S, AV14, AAG13, RI IV, R12N, R13H, AY13, R12L, AQ13,V15S,AD17. A indicate insertions, [] indicate deletions.
[0028] FIG. 11 is a plot showing individual beneficial mutations can be combined (sometimes referred to as “stacked”) for even greater improvements in gene editing activity. CasX proteins were tested for their ability to cleave at 2 different target sites in human HEK293 cells using the E6 and E7 spacers targeting an EGFP reporter, as described in Example 5. The variants, from left to right, are: S794R + Y797L, K416E+A708K, A708K+[P793], [P793]+P793AS, Q367K+I425S, A708K+[P793]+A793V, Q338R+A339E, Q338R+A339K, S507G+G508R, L379R+A708K+[P793], C477K+A708K+[P793], L379R+C477K+A708K+[P793], L379R+A708K+[P793]+A739V, C477K+A708K+[P793]+A739V, L379R+C477K+A708K+[P793]+A739V, L379R+A708K+[P793]+M779N, L379R+A708K+[P793]+M771N, L379R+A708K+[P793]+D489S, L379R+A708K+[P793]+A739T, L379R+A708K+[P793]+D732N, L379R+A708K+[P793]+G791M, L379R+A708K+[P793]+Y797L, L379R+C477K+A708K+[P793]+M779N, L379R+C477K+A708K+[P793]+M771N, L379R+C477K+A708K+[P793]+D489S, L379R+C477K+A708K+[P793]+A739T, L379R+C477K+A708K+[P793]+D732N, L379R+C477K+A708K+[P793]+G791M, L379R+C477K+A708K+[P793]+Y797L, L379R+C477K+A708K+[P793]+T620P, A708K+[P793]+E386S, E386R+F399L+[P793] and R4581I+A739V of the reference CasX protein of SEQ ID NO: 2. [] refer to deleted amino acid residues at the specified position of SEQ ID NO: 2.
[0029] FIG. 12 A and FIG. 12B are a pair of plots showing that CasX protein and sgNA variants when combined, can improve activity more than 6-fold relative to a reference sgRNA and reference CasX protein pair. sgNA:protein pairs were assayed for their ability to cleave a GFP reporter in HEK293 cells, as described in Example 5. On the Y-axis, the fraction of cells in which expression of the GFP reporter was disrupted by CasX mediated gene editing are shown. FIG. 12A shows CasX protein and sgNAs that were assayed with the E6 spacer targeting GFP. FIG. 12B shows CasX protein and sgNAs that were assayed with the E7 spacer targeting GFP. iGFP stands for “inducible GFP.”
[0030] FIG. 13 A, FIG. 13B and FIG. 13C show that making and screening DME libraries has allowed for generation and identification of variants that exhibit a 1 to 81-fold improvement in editing efficiency, as described in Examples 1 and 3. FIG. 13 A shows an RFP+ and GFP+ reporter in E. coli cells assayed for CRISPR interference repression of GFP with a reference nuclease dead CasX protein and sgNA. FIG. 13B shows the same reporter cells assayed for GFP repression with nuclease dead CasX variants screened from a DME library. FIG. 13C shows improved editing efficiency of a selected CasX protein and sgNA variant compared to the reference with 5 spacers targeting the endogenous B2M locus in HEK 293 human cells. The Y axis shows disruption in B2M staining by HLA1 antibody indicating gene disruption via CasX editing and indel formation. The improved CasX variants improved editing of this locus up to 81-fold over the reference in the case of guide spacer # 43. CasX pairs with the reference sgRNA: protein pair of SEQ ID NO: 5 and SEQ ID NO: 2, and CasX variant protein of L379R+A708K+[P793] of SEQ ID NO: 2, assayed with the sgNA variant with a truncated stem loop and a T10C substitution, which is encoded by a sequence of TACTGGCGCCTTTATCTCATTACTTTGAGAGCCATCACCAGCGACTATGTCGTATGG GTAAAGCGCTTACGGACTTCGGTCCGTAAGAAGCATCAAAG (SEQ ID NO: 23), are indicated. The following spacer sequences were used: #9: GTGTAGTACAAGAGATAGAA (SEQ ID NO: 24); #14: TGAAGCTGACAGCATTCGGG (SEQ ID NO: 25), #20: tagATCGAGACATGTAAGCA (SEQ ID NO: 26); #37: GGCCGAGATGTCTCGCTCCG (SEQ ID NO: 27) and #43: AGGCCAGAAAGAGAGAGTAG (SEQ ID NO: 28).
[0031] FIGS. 14A-14F are a series of structural models of a prototypic CasX protein showing the location of mutations in CasX variant proteins of the disclosure which exhibit improved activity. FIG. 14A shows a deletion of P at 793 of SEQ ID NO: 2, with a deletion in a loop that may affect folding. FIG. 14B shows a replacement of Alanine (A) by Lysine (K) at position 708 of SEQ ID NO: 2. This mutation is facing the gNA 5’ end plus a salt bridge to the gNA. FIG. 14C shows a replacement of Cysteine (C) by Lysine (K) at position 477 of SEQ ID NO: 2. This mutation is facing the gNA. There is salt bridge to the gNAbb (gNA phosphase backbone) at approximately base 14 that may be affected. This mutation removes a surface exposed cysteine. FIG. 14D shows a replacement of Leucine (L) with Arginine (R) at position 379 of SEQ ID NO: 2. There is a salt bridge to the target DNAbb (DNA phosphate backbone) towards base pairs 2223 that may be affected. FIG. 14E shows one view of a combination of the deletion of P at 793 and the A708K substitution. FIG. 14F shows an alternate view, that shows that the effects of individual mutants are additive and single mutants can be combined (stacked) for even greater improvements. Arrows indicate the locations of mutations throughout FIG. 14A-14F.
[0032] FIG. 15 is a plot showing the identification of optimal Planctomycetes CasX PAM and spacers for genes of interest, as described in Example 6. On the Y-axis, percent GFP negative cells, indicating cleavage of a GFP reporter, is shown. On the X-axis, different PAM sequences and spacers: ATC PAM, CTC PAM and TTC PAM. GTC, TTT and CTT PAMs were also tested and showed no activity.
[0033] FIG. 16 is a plot showing that improved CasX variants generated by DME can edit both canonical and non-canonical PAMs more efficiently than reference CasX proteins, as described in Example 6. The Y-axis shows the average fold improvement in editing relative to a reference sgRNA: protein pair (SEQ ID NO:2, SEQ ID NO: 5) with 2 targets, N= 6. Protein variants, from left to right for each set of bars were: A708K+[P793]+ A739V; L379R+A708K+[P793]; C477K+A708K+[P793]; L379R+C477K+A708K+[P793]; L379R+A708K+[P793]+A739V; C477K+A708K+[P793]+A739V; and L379R+C477K+A708K+[P793]+A739V. Reference CasX and protein variants were assayed with a reference sgRNA scaffold of SEQ ID NO: 5 with DNA encoding spacer sequences of, from left to right, E6 (SEQ ID NO: 29) with a TTC PAM; E7 (SEQ ID NO: 30) with a TTC PAM; GFP8 (SEQ ID NO: 31) with a TTC PAM; Bl (SEQ ID NO: 32) with a CTC PAM and A7 (SEQ ID NO: 33) with an ATC PAM.
[0034] FIGS. 17A-17F are a series of plots showing that a reference CasX protein and a reference sgRNA scaffold pair is highly specific for the target sequence, as described in Example 7. FIG. 17A and FIG. 17D, Streptococcus pyogenes Cas9 (SpyCas9) was assayed with two different gNA spacers and a 5’ PAM site (SEQ ID NOs: 34-65) and (SEQ ID NOs: 136166) for its ability to edit templates with a target sequence complementary to the spacer sequence (arrow), or with 1, 2, 3 or 4 mutations in the target sequence relative to the spacer sequence. FIG. 17B and FIG. 17E, Staphylococcus aureus Cas9 (SauCas9) was assayed with two different gNA spacers and a 5’ PAM site (SEQ ID NOs: 66-103) and (SEQ ID NOs: 167-204) for its ability to edit templates with a target sequence complementary to the spacer sequence (arrow), or with 1, 2, 3 or 4 mutations in the target sequence relative to the spacer sequence. FIG. 17C and FIG. 17F, the reference Pim CasX protein and sgNA scaffold pair was assayed with two different gNA spacers and a 3’ PAM site (SEQ ID NOs: 104-135) and (SEQ ID NOs: 205-236) for its ability to edit templates with a target sequence complementary to the spacer sequence (arrow), or with 1, 2, 3 or 4 mutations in the target sequence relative to the spacer sequence. In all of FIG. 17A-17F, the X-axis shows the fraction of cells where gene editing at the target sequence occurred.
[0035] FIG. 18 illustrates a scaffold stem loop of an exemplary reference sgRNA of the disclosure (SEQ ID NO: 237).
[0036] FIG. 19 illustrates an extended stem loop sequence of an exemplary reference sgRNA of the disclosure (SEQ ID NO: 238).
[0037] FIGS. 20A-20B are a pair of plots that demonstrate that specific subsets of changes discovered by DME of the CasX are more likely to predict improvements of activity, as described in Example 4. The plots represent data from the experiments described in FIG.7 and FIG. 8. FIG 20A shows that changing amino acids within a distance of 10 Angstroms (A) of the guide RNA to hydrophobic residues (A, V, I, L, M, F, Y, W) results in a significantly less active protein. FIG. 20B demonstrates that, in contrast, changing a residue within 10 A of the RNA to a positively charged amino acid (R, H, K) is likely to improve activity.
[0038] FIG. 21 illustrates an alignment of two reference CasX protein sequences (SEQ ID NO: 1, top; SEQ ID NO: 2, bottom), with domains annotated.
[0039] FIG. 22 illustrates the domain organization of a reference CasX protein of SEQ ID NO: 1. The domains have the following coordinates: non-target strand binding (NTSB) domain: amino acids 101-191; Helical I domain: amino acids 57-100 and 192-332; Helical II domain: 333-509; oligonucleotide binding domain (OBD): amino acids 1-56 and 510-660; RuvC DNA cleavage domain (RuvC): amino acids 551-824 and 935-986; target strand loading (TSL) domain: amino acids 825-934. Note that the Helical I, OBD and RuvC domains are noncontiguous.
[0040] FIG. 23 illustrates an alignment of two CasX reference sgRNA scaffolds SEQ ID NO: 5 (top) and SEQ ID NO: 4 (bottom).
[0041] FIG. 24 shows an SDS-PAGE gel of StX2 (CasX reference of SEQ ID NO: 2) purification fractions visualized by colloidal Coomassie staining, as described in Example 8. The lanes, from left to right, are: Pellet: insoluble portion following cell lysis, Lysate: soluble portion following cell lysis, Flow Thru: protein that did not bind the heparin column, Wash: protein that eluted from the column in wash buffer, Elution: protein eluted from the heparin column with elution buffer, Flow Thru: Protein that did not bind the StrepTactin column, Elution: protein eluted from the StrepTactin column with elution buffer, Injection: concentrated protein injected onto the s200 gel filtration column, Frozen: pooled fractions from the s200 elution that have been concentrated and frozen.
[0042] FIG. 25 shows the chromatogram from a size exclusion chromatography assay of the StX2, as described in Example 8.
[0043] FIG. 26 shows an SDS-PAGE gel of StX2 purification fractions visualized by colloidal Coomassie staining, as described in Example 8. From right to left: Injection sample, molecular weight markers, lanes 3 -9: samples from the indicated elution volumes.
[0044] FIG. 27 shows the chromatogram from a size exclusion chromatography assay of the CasX 119, using of Superdex 200 16 / 600 pg gel filtration, as described in Example 8. The 67.47 mL peak corresponds to the apparent molecular weight of CasX variant 119 and contained the majority of CasX variant 119 protein.
[0045] FIG. 28 shows an SDS-PAGE gel of CasX 119 purification fractions visualized by colloidal Coomassie staining, as described in Example 8. Samples from the indicated fractions were resolved by SDS-PAGE and stained with colloidal Coomassie. From right to left, Injection: sample of protein injected onto the gel filtration column, molecular weight markers, lanes 3-10: samples from the indicated elution volumes.
[0046] FIG. 29 shows an SDS-PAGE gel of purification samples of CasX 438, visualized on a Bio-Rad Stain-Free™ gel. The lanes, from left to right, are: Pellet: insoluble portion following cell lysis, Lysate: soluble portion following cell lysis, Flow Thru: protein that did not bind the heparin column, Elution: protein eluted from the heparin column with elution buffer, Flow Thru: Protein that did not bind the StrepTactin column, Elution: protein eluted from the StrepTactin column with elution buffer, Injection: concentrated protein injected onto the s200 gel filtration column, Pool: pooled CasX-containing fractions, Final: pooled fractions from the s200 elution that have been concentrated and frozen.
[0047] FIG. 30 shows the chromatogram from a size exclusion chromatography assay of the CasX 438, using of Superdex 200 16 / 600 pg gel filtration, as described in Example 8. The 69.13 mL peak corresponds to the apparent molecular weight of CasX variant 438 and contained the majority of CasX variant 438 protein.
[0048] FIG. 31 shows an SDS-PAGE gel of CasX 438 purification fractions visualized by colloidal Coomassie staining, as described in Example 8. Samples from the indicated fractions were resolved by SDS-PAGE and stained with colloidal Coomassie. From right to left, Injection: sample of protein injected onto the gel filtration column, molecular weight markers, lanes 3-10: samples from the indicated elution volumes.
[0049] FIG. 32 shows an SDS-PAGE gel of purification samples of CasX 457, visualized on a Bio-Rad Stain-Free™ gel. The lanes, from left to right, are: Pellet: insoluble portion following cell lysis, Lysate: soluble portion following cell lysis, Flow Thru: protein that did not bind the heparin column, Wash, Elution: protein eluted from the heparin column with elution buffer, Flow Thru: Protein that did not bind the StrepTactin column, Elution: protein eluted from the StrepTactin column with elution buffer, Injection: concentrated protein injected onto the s200 gel filtration column, Final: pooled fractions from the s200 elution that have been concentrated and frozen.
[0050] FIG. 33 shows the chromatogram from a size exclusion chromatography assay of the CasX 457, using of Superdex 200 16 / 600 pg gel filtration, as described in Example 8. The 67.52 mL peak corresponds to the apparent molecular weight of CasX variant 457 and contained the majority of CasX variant 457 protein.
[0051] FIG. 34 shows an SDS-PAGE gel of CasX 457 purification fractions visualized by colloidal Coomassie staining, as described in Example 8. Samples from the indicated fractions were resolved by SDS-PAGE and stained with colloidal Coomassie. From right to left, Injection: sample of protein injected onto the gel filtration column, molecular weight markers, lanes 3-10: samples from the indicated elution volumes.
[0052] FIG. 35 is a schematic showing the organization of the components in the pSTX34 plasmid used to assemble the CasX constructs, as described in Example 9.
[0053] FIG. 36 is a schematic showing the steps of generating the CasX 119 variant, as described in Example 9.
[0054] FIG. 37 is a graph of the results of an assay for the quantification of active fractions of RNP formed by sgRNA174 and the CasX variants 119 and 457, as described in Example 19. Equimolar amounts of RNP and target were co-incubated and the amount of cleaved target was determined at the indicated timepoints. Mean and standard deviation of three independent replicates are shown for each timepoint. The biphasic fit of the combined replicates is shown. “2” refers to the reference CasX protein of SEQ ID NO: 2.
[0055] FIG. 38 is a graph of the results of an assay for quantification of active fractions of RNP formed by CasX2 and reference guide 2 the modified sgRNA guides 32, 64, and 174, as described in Example 19. Equimolar amounts of RNP and target were co-incubated and the amount of cleaved target was determined at the indicated timepoints. Mean and standard deviation of three independent replicates are shown for each timepoint. The biphasic fit of the combined replicates is shown. “2” refers to reference gRNAs SEQ ID NO: 5, respectively, and the identifying number of modified sgRNAs are indicated in Table 2.
[0056] FIG. 39 is a graph of the results of an assay for quantification of cleavage rates of RNP formed by sgRNA174 and the CasX variants 119 and 457, as described in Example 19. Target DNA was incubated with a 20-fold excess of the indicated RNP and the amount of cleaved target was determined at the indicated time points. Mean and standard deviation of three independent replicates are shown for each timepoint. The monophasic fit of the combined replicates is shown.
[0057] FIG. 40 is a graph of the results of an assay for quantification of cleavage rates of RNP formed by CasX2 and the sgRNA guide variants 2, 32, 64 and 174, as described in Example 19. Target DNA was incubated with a 20-fold excess of the indicated RNP and the amount of cleaved target was determined at the indicated time points. Mean and standard deviation of three independent replicates are shown for each timepoint. The monophasic fit of the combined replicates is shown.
[0058] FIG. 41 is a graph of the results of an assay for quantification of initial velocities of RNP formed by CasX2 and the sgRNA guide variants 2, 32, 64 and 174, as described in Example 19. The first two time-points of the previous cleavage experiment were fit with a linear model to determine the initial cleavage velocity.
[0059] FIG. 42 is a schematic showing an example of CasX protein and scaffold DNA sequence for packaging in adeno-associated virus (AAV), as described in Example 20. The DNA segment between the AAV inverted terminal repeats (ITRs), comprised of a CasX-encoding DNA and its promoter, and scaffold-encoding DNA and its promoter gets packaged within an AAV capsid during AAV production.
[0060] FIG. 43 is a graph showing representative results of AAV titering by qPCR, as described in Example 20. During AAV purification, flow through (FT) and consecutive eluent fractions (1-6) are collected and titered by qPCR. Most virus, ~lel4 viral genomes in this example, is found in the second elution fraction.
[0061] FIG. 44 shows the results of an AAV-mediated gene editing experiment in the SOD1-GFP reporter cell line, as described in Example 21. CasX constructs (CasX 119 and guide 64 with SOD1 targeting spacer 2, ATGTTCATGAGTTTGGAGAT; SEQ ID NO: 239) and SauCas9 with SOD1 targeting spacer were packaged in AAV vectors and used to transduce SOD1-GFP reporter cells at a range of different multiplicity of infection (MOIs, no. of viral genomes / cell). Twelve days later, cells were assayed for GFP disruption via FACS. In this example, CasX and SauCas9 shows equivalent levels of editing, where 1-2% of the cells show GFP disruption at the highest MOIs, le7 or le6.
[0062] FIG. 45 shows the results of a second AAV-mediated gene editing experiment in the SOD1-GFP reporter cell line, as described in Example 21. CasX constructs 119.64 with SOD1 targeting spacer (2, ATGTTCATGAGTTTGGAGAT; SEQ ID NO: 239) and SauCas9 with SOD1 targeting spacer were packaged in AAV vectors and used to transduce SOD 1-GFP reporter cells at a range of different multiplicity of infection (MOIs, no. of viral genomes / cell). Twelve days later, cells were assayed for GFP disruption via FACS. In this example, CasX and SauCas9 shows equivalent levels of editing at the highest MOI, where -2-4% of the cells show GFP disruption.
[0063] FIG. 46 shows the results of an AAV-mediated gene editing experiment in neural progenitor cells (NPCs) from the G93A mouse model of ALS, as described in Example 21. CasX constructs (CasX 119 and guide 64 with SOD1 targeting spacer 2, ATGTTCATGAGTTTGGAGAT; SEQ ID NO: 239) was packaged in an AAV vector and used to transduce G93 A NPCs at a range of different multiplicity of infection (MOIs, no. of viral genomes / cell). Twelve days later, cells were assayed for gene editing via T7E1 assay. Agarose gel image from the T7E1 assay shown here demonstrates successful editing of the SOD1 locus. Double arrows show the two DNA bands as a result of successful editing in cells.
[0064] FIG. 47 shows the results of an editing assay of 6 target genes in HEK293T cells, as described in Example 23. Each dot represents results using an individual spacer.
[0065] FIG. 48 shows the results of an editing assay of 6 target genes in HEK293T cells, with individual bars representing the results obtained with individual spacers, as described in Example 23.
[0066] FIG. 49 shows the results of an editing assay of 4 target genes in HEK293T cells, as described in Example 23. Each dot represents results using an individual spacer utilizing a CTC (CTCN) PAM.
[0067] FIG. 50 is a schematic showing the steps of Deep Mutational Evolution used to create libraries of genes encoding CasX variants, as described in Example 24. The pSTXl backbone is minimal, composed of only a high-copy number origin and KanR resistance gene, making it compatible with the recombineering E. coli strain EcNR2. pSTX2 is a BsmbI destination plasmid for aTc-inducible expression in E. coli.
[0068] FIG. 51 are dot plot graphs showing the results of CRISPRi screens for mutations in libraries DI, D2, and D3, as described in Example 24. In the absence of CRISPRi, E. coli constitutively express both GFP and RFP, resulting in intense fluorescence in both wavelengths, represented by dots in the upper-right region of the plot. CasX proteins resulting in CRISPRi of GFP can reduce green fluorescence by > 10-fold, while leaving red fluorescence unaltered, and these cells fall within the indicated Sort Gate 1. The total fraction of cells exhibiting CRISPRi is indicated.
[0069] FIG. 52 are photographs of colonies grown in the ccdB assay, as described in Example 24. 10-fold dilutions were assayed in the presence of glucose or arabinose to induce expression of the ccdB toxin, resulting in approximately a 1000-fold difference between functional and nonfunctional proteins. When grown in liquid culture, the resolving power was approximately 10,000-fold, as seen on the right-hand side.
[0070] FIG. 53 is a graph of HEK iGFP genome editing efficiency testing CasX variants with sgRNA 2 (SEQ ID NO :5), with appropriate spacers, with data expressed as fold-improvement over the wild-type CasX protein (SEQ ID NO: 2) in the HEK iGFP editing assay, as described in Example 24. Single mutations are shown at the top, with groups of mutations shown at the bottom of the graph). Error bars combine internal measurement error (SD) and interexperimental measurement error (SD across replicate experiments for those variants tested more than once), in at least triplicate assays.
[0071] FIG. 54 is a scatterplot showing results of the SOD1-GFP reporter assay for CasX variants with sgRNA scaffold 2 utilizing two different spacers for GFP, as described in Example 24.
[0072] FIG. 55 is a graph showing the results of the HEK293 iGFP genome editing assay assessing editing across four different PAM sequences comparing wild-type CasX (SEQ ID NO: 2) and CasX variant 119; both utilizing sgRNA scaffold 1 (SEQ ID NO: 4), with spacers utilizing four different PAM sequences, as described in Example 24.
[0073] FIG. 56 is a graph showing the results of genome editing activity of CasX variant 119 and sgRNA 174 compared to wild-type CasX 2 and guide scaffold 1 in the iGFP lipofection assay utilizing two different spacers, as described in Example 24.
[0074] FIG. 57 is a graph showing the results of genome editing activity of CasX variant 119 and sgRNA 174 compared to wild-type CasX and guide in the iGFP lentiviral transduction assay, using two different spacers, as described in Example 24.
[0075] FIG. 58 is a graph showing the results of genome editing in the more stringent lentiviral assay to compare the editing activity of four CasX variants (119, 438, 488 and 491) and the optimized sgNA 174 and two different spacers, as described in Example 24. The results show the step-wise improvement in editing efficiency achieved by the additional modifications and domain swaps introduced to the starting-point 119 variant.
[0076] FIGS. 59A- 59B shows the results of NGS analyses of the libraries of sgRNA, as described in Example 25. FIG. 59A shows the distribution of substitutions, deletions and insertions. FIG. 59B is a scatterplot showing the high reproducibility of variant representation in two separate library pools after the CRISPRi assay in the unsorted, naive population of cells. (Library pool D3 vs D2 are two different versions of the dCasX protein, and represent replicates of the CRISPRi assay.)
[0077] FIGS. 60A-60B shows the structure of wild-type CasX and RNA guide (SEQ ID NO:4). FIG. 60A depicts the CryoEM structure of Deltaproteobacteria CasX protein:sgRNA RNP complex (PDB id: 6YN2), including two stem loops, a pseudoknot, and a triplex. FIG. 60B depicts the secondary structure of the sgRNA was identified from the structure shown in (A) using the tool RNAPDBee 2.0 (rnapdbee.cs.put.poznan.pl / , using the tools 3DNA / DSSR, and using the VARNA visualization tool). RNA regions are indicated. Residues that were not evident in the PDB crystal structure file are indicated by plain-text letters (i.e., not encircled), and are not included in residue numbering.
[0078] FIGS. 61A-61C depicts comparisons between two guide RNA scaffolds. FIG. 61A provides the sequence alignment between the single guide scaffold 1 (SEQ ID NO: 4) and scaffold 2 (SEQ ID NO: 5). FIG. 61B shows the predicted secondary structure of scaffold 1 (without the 5’ ACAUCU bases which were not in the cryoEM structure). Prediction was done using RNAfold (v 2.1.7), using a constraint that was derived from the base-pairing observed in the cryoEM structure (see FIGS. 60A-60B). This constraint required the base pairs observed in the cryoEM structure to be formed, and required the bases involved in triplex formation to be unpaired. This structure has distinct base pairing from the lowest-energy predicted structure at the 5’ end (i.e., the pseudoknot and triplex loop). FIG. 61C shows the predicted secondary structure of scaffold 2. Prediction was done for scaffold 1, using a similar constraint based on the sequence alignment.
[0079] FIG. 62 shows a graph comparing GFP-knockdown capability of scaffold 1 versus scaffold 2 in GFP-lipofection assay, using four different spacers utilizing different PAM sequences, as described in Example 25. The results demonstrate the greater editing imparted by use of the modified scaffold 2 compared to the wild-type scaffold 1; the latter showing no editing with spacers utilizing GTC and CTC PAM sequences.
[0080] FIGS. 63 A-63C shows graphs depicting the enrichment of single variants across the scaffold, revealing mutable regions, as described in Example 25. FIG. 63 A depicts substituted bases (A, T, G, or C; top to bottom), FIG. 63B depicts inserted bases (A, T, G, or C; top to bottom), and FIG. 63C depicts deletions at the individual nucleotide position (X-axis) across scaffold 2. Enrichment values were averaged across the three dead CasX versions, relative to the average WT value. Scaffolds with relative log2 enrichment > 0 are considered ‘enriched’, as they were more represented in the sorted population relative to the naive population than the wildtype scaffold was represented. Error bars represent the confidence interval across the three catalytically dead CasX experiments.
[0081] FIG. 64 are scatterplots showing that the enrichment values obtained across different dCasX variants are largely consistent, as described in Example 25. Libraries D2 and DDD have highly correlated enrichment scores, while D3 is more distinct.
[0082] FIG. 65 shows a bar graph of cleavage activity of several scaffold variants in a more stringent lipofection assay at the SOD1-GFP locus, as described in Example 25.
[0083] FIG. 66 shows a bar graph of cleavage activity for several scaffold variants using two different spacers; 8.2 and 8.4 that target SOD1-GFP locus (and a non-targeting spacer NT), with low-MOI lentiviral transduction using a p34 plasmid backbone, as described in Example 25.
[0084] FIG. 67 is a schematic showing the secondary structure of single guide 174 on top and the linear structure on the bottom, with lines joining those segments associating by base-pairing or other non-covalent interactions. The scaffold stem (white, no fill) (and loop) and the extended stem (grey, no fill) (and loop) are adjacent from 5’ to 3’ in the sequence. However, the pseudoknot and extended stems are formed from strands that have intervening regions in the sequence. The triplex is formed, in the case of single guide 174, comprising nucleotides 5’ -CUUUG’-3’ AND 5’-CAAAG-3’ that form a base-paired duplex and nucleotides 5’-UUU-3’ that associates with the 5’-AAA-3’ to form the triplex region.
[0085] FIG. 68 shows comparisons between the highly-evolved single guide 174 and the scaffolds 1 and 2 that served as the starting points for the DME procedures described in Example 25. FIG. 68 A shows a bar graph of cleavage activity of head-to-head comparisons of cleavage activity of the guide scaffolds with five different spacers in a plasmid lipofection assay at the GFP locus in HEK-GFP cells. FIG. 68B shows the sequence alignment between scaffold 2 and guide 174 (SEQ ID NO: 2238). Asterisks indicate point mutations, and the dotted box shows the entire extended stem swap.
[0086] FIGS. 69A-69B shows scatterplots of HEK-iGFP cleavage assay for scaffolds sequences relative to WT scaffold with 2 spacers; 4.76 (FIG. 69A) and 4.77 (FIG. 69B), as described in Example 25.
[0087] FIG. 70 shows a scatterplot comparing the normalized cleavage activity of several scaffolds relative to WT with 2 spacers (4.76 and 4.77), as described in Example 25. Error bars combine internal measurement error (SD) and inter-experimental measurement error (SD across replicate experiments for those variants tested more than once), in quadrature.
[0088] FIG. 71 shows a scatterplot comparing the normalized cleavage activity of multiple scaffolds relative to WT in the HEK-iGFP cleavage assay to the enrichments obtained from the CRISPRi comprehensive screen, as described in Example 25. Generally, scaffold mutations with high enrichment (>1.5) have cleavage activity comparable to or greater than WT. Two variants have high cleavage activity with low enrichment scores (C18G and T17G); interestingly, these substitutions are at the same position as several highly enriched insertions (FIGS. 63 A-63C). Labels indicate the mutations for a subset of the comparisons.
[0089] FIG. 72 shows the results of flow cytometry analysis of Cas-mediated editing at the RHO locus in APRE19 RHO-GFP cells 14 days post-transfection for the CasX variant constructs 438, 499 and 491, as described in Example 26. The points are the results of individual samples and the light dashed lines are upper and lower quartiles.
[0090] FIG. 73 shows the quantification of cleavage rates of RNP formed by sgRNA174 and the CasX variants on targets with different PAMs. Target DNA was incubated with a 20-fold excess of the indicated RNP and the amount of cleaved target was determined at the indicated time points. The monophasic fit of the combined replicates is shown. DETAILED DESCRIPTION
[0091] While exemplary embodiments have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions will now occur to those skilled in the art without departing from the inventions claimed herein. It should be understood that various alternatives to the embodiments described herein may be employed in practicing the embodiments of the disclosure. It is intended that the claims define the scope of the invention and that methods and structures within the scope of these claims and their equivalents be covered thereby. Defintions
[0092] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Although methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present embodiments, suitable methods and materials are described below. In case of conflict, the patent specification, including definitions, will control. In addition, the materials, methods, and examples are illustrative only and not intended to be limiting. Numerous variations, changes, and substitutions will now occur to those skilled in the art without departing from the invention.
[0093] The terms "polynucleotide" and "nucleic acid," used interchangeably herein, refer to a polymeric form of nucleotides of any length, either ribonucleotides or deoxyribonucleotides. Thus, terms "polynucleotide" and "nucleic acid" encompass single-stranded DNA; doublestranded DNA; multi-stranded DNA; single-stranded RNA; double-stranded RNA; multistranded RNA; genomic DNA; cDNA; DNA-RNA hybrids; and a polymer comprising purine and pyrimidine bases or other natural, chemically or biochemically modified, non-natural, or derivatized nucleotide bases.
[0094] "Hybridizable" or "complementary" are used interchangeably to mean that a nucleic acid (e.g., RNA, DNA) comprises a sequence of nucleotides that enables it to non-covalently bind, i.e., form Watson-Crick base pairs and / or G / U base pairs, "anneal", or "hybridize," to another nucleic acid in a sequence-specific, antiparallel, manner (i.e., a nucleic acid specifically binds to a complementary nucleic acid) under the appropriate in vitro and / or in vivo conditions of temperature and solution ionic strength. It is understood that the sequence of a polynucleotide need not be 100% complementary to that of its target nucleic acid to be specifically hybridizable; it can have at least about 70%, at least about 80%, or at least about 90%, or at least about 95% sequence identity and still hybridize to the target nucleic acid. Moreover, a polynucleotide may hybridize over one or more segments such that intervening or adjacent segments are not involved in the hybridization event (e.g., a loop structure or hairpin structure, a 'bulge', ‘bubble’ and the like).
[0095] A “gene,” for the purposes of the present disclosure, includes a DNA region encoding a gene product (e.g., a protein, RNA), as well as all DNA regions which regulate the production of the gene product, whether or not such regulatory sequences are adjacent to coding and / or transcribed sequences. Accordingly, a gene may include regulatory sequences including, but not necessarily limited to, promoter sequences, terminators, translational regulatory sequences such as ribosome binding sites and internal ribosome entry sites, enhancers, silencers, insulators, boundary elements, replication origins, matrix attachment sites and locus control regions. Coding sequences encode a gene product upon transcription or transcription and translation; the coding sequences of the disclosure may comprise fragments and need not contain a full-length open reading frame. A gene can include both the strand that is transcribed, e.g. the strand containing the coding sequence, as well as the complementary strand.
[0096] The term "downstream" refers to a nucleotide sequence that is located 3' to a reference nucleotide sequence. In certain embodiments, downstream nucleotide sequences relate to sequences that follow the starting point of transcription. For example, the translation initiation codon of a gene is located downstream of the start site of transcription.
[0097] The term "upstream" refers to a nucleotide sequence that is located 5' to a reference nucleotide sequence. In certain embodiments, upstream nucleotide sequences relate to sequences that are located on the 5' side of a coding region or starting point of transcription. For example, most promoters are located upstream of the start site of transcription.
[0098] The term “regulatory element” is used interchangeably herein with the term “regulatory sequence,” and is intended to include promoters, enhancers, and other expression regulatory elements (e.g. transcription termination signals, such as polyadenylation signals and poly-U sequences). Exemplary regulatory elements include a transcription promoter such as, but not limited to, CMV, CMV+intron A, SV40, RSV, HIV-Ltr, elongation factor 1 alpha (EFla), MMLV-ltr, internal ribosome entry site (IRES) or P2A peptide to permit translation of multiple genes from a single transcript, metallothionein, a transcription enhancer element, a transcription termination signal, polyadenylation sequences, sequences for optimization of initiation of translation, and translation termination sequences. It will be understood that the choice of the appropriate regulatory element will depend on the encoded component to be expressed (e.g., protein or RNA) or whether the nucleic acid comprises multiple components that require different polymerases or are not intended to be expressed as a fusion protein.
[0099] The term "promoter" refers to a DNA sequence that contains an RNA polymerase binding site, transcription start site, TATA box, and / or B recognition element and assists or promotes the transcription and expression of an associated transcribable polynucleotide sequence and / or gene (or transgene). A promoter can be synthetically produced or can be derived from a known or naturally occurring promoter sequence or another promoter sequence. A promoter can be proximal or distal to the gene to be transcribed. A promoter can also include a chimeric promoter comprising a combination of two or more heterologous sequences to confer certain properties. A promoter of the present disclosure can include variants of promoter sequences that are similar in composition, but not identical to, other promoter sequence(s) known or provided herein. A promoter can be classified according to criteria relating to the pattern of expression of an associated coding or transcribable sequence or gene operably linked to the promoter, such as constitutive, developmental, tissue-specific, inducible, etc.
[00100] The term “enhancer” refers to regulatory element DNA sequences that, when bound by specific proteins called transcription factors, regulate the expression of an associated gene. Enhancers may be located in the intron of the gene, or 5’ or 3’ of the coding sequence of the gene. Enhancers may be proximal to the gene (i.e., within a few tens or hundreds of base pairs (bp) of the promoter), or may be located distal to the gene (i.e., thousands of bp, hundreds of thousands of bp, or even millions of bp away from the promoter). A single gene may be regulated by more than one enhancer, all of which are envisaged as within the scope of the instant disclosure.
[00101] "Recombinant," as used herein, means that a particular nucleic acid (DNA or RNA) is the product of various combinations of cloning, restriction, and / or ligation steps resulting in a construct having a structural coding or non-coding sequence distinguishable from endogenous nucleic acids found in natural systems. Generally, DNA sequences encoding the structural coding sequence can be assembled from cDNA fragments and short oligonucleotide linkers, or from a series of synthetic oligonucleotides, to provide a synthetic nucleic acid which is capable of being expressed from a recombinant transcriptional unit contained in a cell or in a cell-free transcription and translation system. Such sequences can be provided in the form of an open reading frame uninterrupted by internal non-translated sequences, or introns, which are typically present in eukaryotic genes. Genomic DNA comprising the relevant sequences can also be used in the formation of a recombinant gene or transcriptional unit. Sequences of non-translated DNA may be present 5’ or 3’ from the open reading frame, where such sequences do not interfere with manipulation or expression of the coding regions, and may indeed act to modulate production of a desired product by various mechanisms (see "enhancers” and “promoters", above).
[00102] The term "recombinant polynucleotide" or "recombinant nucleic acid" refers to one which is not naturally occurring, e.g., is made by the artificial combination of two otherwise separated segments of sequence through human intervention. This artificial combination is often accomplished by either chemical synthesis means, or by the artificial manipulation of isolated segments of nucleic acids, e.g., by genetic engineering techniques. Such can be done to replace a codon with a redundant codon encoding the same or a conservative amino acid, while typically introducing or removing a sequence recognition site. Alternatively, it is performed to join together nucleic acid segments of desired functions to generate a desired combination of functions. This artificial combination is often accomplished by either chemical synthesis means, or by the artificial manipulation of isolated segments of nucleic acids, e.g., by genetic engineering techniques.
[00103] Similarly, the term "recombinant polypeptide” or “recombinant protein” refers to a polypeptide or protein which is not naturally occurring, e.g., is made by the artificial combination of two otherwise separated segments of amino sequence through human intervention. Thus, e.g., a protein that comprises a heterologous amino acid sequence is recombinant.
[00104] As used herein, the term "contacting" means establishing a physical connection between two or more entities. For example, contacting a target nucleic acid with a guide nucleic acid means that the target nucleic acid and the guide nucleic acid are made to share a physical connection; e.g., can hybridize if the sequences share sequence similarity.
[00105] “Dissociation constant”, or “Ka”, are used interchangeably and mean the affinity between a ligand “L” and a protein “P”; i.e., how tightly a ligand binds to a particular protein. It can be calculated using the formula Kd=[L] [P] / [LP], where [P], [L] and [LP] represent molar concentrations of the protein, ligand and complex, respectively.
[00106] The disclosure provides compositions and methods useful for editing a target nucleic acid sequence. As used herein “editing” is used interchangeably with “modifying” and includes but is not limited to cleaving, nicking, deleting, knocking in, knocking out, and the like.
[00107] As used herein, "homology-directed repair" (HDR) refers to the form of DNA repair that takes place during repair of double-strand breaks in cells. This process requires nucleotide sequence homology, and uses a donor template to repair or knock-out a target DNA, and leads to the transfer of genetic information from the donor (e.g., such as the donor template) to the target. Homology-directed repair can result in an alteration of the sequence of the target nucleic acid sequence by insertion, deletion, or mutation if the donor template differs from the target DNA sequence and part or all of the sequence of the donor template is incorporated into the target DNA at the correct genomic locus.
[00108] As used herein, "non-homologous end joining" (NHEJ) refers to the repair of doublestrand breaks in DNA by direct ligation of the break ends to one another without the need for a homologous template (in contrast to homology-directed repair, which requires a homologous sequence to guide repair). NHEJ often results in indels; the loss (deletion) or insertion of nucleotide sequence near the site of the double- strand break.
[00109] As used herein “micro-homology mediated end joining” (MMEJ) refers to a mutagenic DSB repair mechanism, which always associates with deletions flanking the break sites without the need for a homologous template (in contrast to homology-directed repair, which requires a homologous sequence to guide repair). MMEJ often results in the loss (deletion) of nucleotide sequence near the site of the double- strand break.
[00110] A polynucleotide or polypeptide (or protein) has a certain percent "sequence similarity" or "sequence identity" to another polynucleotide or polypeptide, meaning that, when aligned, that percentage of bases or amino acids are the same, and in the same relative position, when comparing the two sequences. Sequence similarity (sometimes referred to as percent similarity, percent identity, or homology) can be determined in a number of different manners. To determine sequence similarity, sequences can be aligned using the methods and computer programs that are known in the art, including BLAST, available over the world wide web at ncbi.nlm.nih.gov / BLAST. Percent complementarity between particular stretches of nucleic acid sequences within nucleic acids can be determined using any convenient method. Example methods include BLAST programs (basic local alignment search tools) and PowerBLAST programs (Altschul et al., J. Mol. Biol., 1990, 215, 403-410; Zhang and Madden, Genome Res., 1997, 7, 649-656) or by using the Gap program (Wisconsin Sequence Analysis Package, Version 8 for Unix, Genetics Computer Group, University Research Park, Madison Wis.), e.g., using default settings, which uses the algorithm of Smith and Waterman (Adv. Appl. Math., 1981, 2, 482-489).
[00111] The terms "polypeptide," and "protein" are used interchangeably herein, and refer to a polymeric form of amino acids of any length, which can include coded and non-coded amino acids, chemically or biochemically modified or derivatized amino acids, and polypeptides having modified peptide backbones. The term includes fusion proteins, including, but not limited to, fusion proteins with a heterologous amino acid sequence.
[00112] A "vector" or "expression vector" is a replicon, such as plasmid, phage, virus, or cosmid, to which another DNA segment, i.e., an "insert", may be attached so as to bring about the replication or expression of the attached segment in a cell.
[00113] The term "naturally-occurring" or "unmodified" or "wild-type" as used herein as applied to a nucleic acid, a polypeptide, a cell, or an organism, refers to a nucleic acid, polypeptide, cell, or organism that is found in nature.
[00114] As used herein, a "mutation" refers to an insertion, deletion, substitution, duplication, or inversion of one or more amino acids or nucleotides as compared to a wild-type or reference amino acid sequence or to a wild-type or reference nucleotide sequence.
[00115] As used herein the term "isolated" is meant to describe a polynucleotide, a polypeptide, or a cell that is in an environment different from that in which the polynucleotide, the polypeptide, or the cell naturally occurs. An isolated genetically modified host cell may be present in a mixed population of genetically modified host cells.
[00116] A "host cell," as used herein, denotes a eukaryotic cell, a prokaryotic cell, or a cell from a multicellular organism (e.g., a cell line) cultured as a unicellular entity, which cells are used as recipients for a nucleic acid (e.g., an expression vector), and include the progeny of the original cell which has been genetically modified by the nucleic acid. It is understood that the progeny of a single cell may not necessarily be completely identical in morphology or in genomic or total DNA complement as the original parent, due to natural, accidental, or deliberate mutation. A "recombinant host cell" (also referred to as a "genetically modified host cell") is a host cell into which has been introduced a heterologous nucleic acid, e.g., an expression vector.
[00117] The term "conservative amino acid substitution" refers to the interchangeability in proteins of amino acid residues having similar side chains. For example, a group of amino acids having aliphatic side chains consists of glycine, alanine, valine, leucine, and isoleucine; a group of amino acids having aliphatic-hydroxyl side chains consists of serine and threonine; a group of amino acids having amide-containing side chains consists of asparagine and glutamine; a group of amino acids having aromatic side chains consists of phenylalanine, tyrosine, and tryptophan; a group of amino acids having basic side chains consists of lysine, arginine, and histidine; and a group of amino acids having sulfur-containing side chains consists of cysteine and methionine. Exemplary conservative amino acid substitution groups are: valine-leucine-isoleucine, phenylalanine-tyrosine, lysine-arginine, alanine-valine, and asparagine-glutamine.
[00118] As used herein, "treatment" or "treating," are used interchangeably herein and refer to an approach for obtaining beneficial or desired results, including but not limited to a therapeutic benefit and / or a prophylactic benefit. By therapeutic benefit is meant eradication or amelioration of the underlying disorder or disease being treated. A therapeutic benefit can also be achieved with the eradication or amelioration of one or more of the symptoms or an improvement in one or more clinical parameters associated with the underlying disease such that an improvement is observed in the subject, notwithstanding that the subject may still be afflicted with the underlying disorder.
[00119] The terms "therapeutically effective amount" and "therapeutically effective dose", as used herein, refer to an amount of a composition, vector, cells, etc., that is capable of having any detectable, beneficial effect on any symptom, aspect, measured parameter or characteristics of a disease state or condition when administered in one or repeated doses to a subject. Such effect need not be absolute to be beneficial. Such effect can be transient.
[00120] As used herein, "administering" is meant as a method of giving a dosage of a composition of the disclosure to a subject.
[00121] As used herein, a "subject" is a mammal. Mammals include, but are not limited to, domesticated animals, primates, non-human primates, humans, dogs, porcine (pigs), rabbits, mice, rats and other rodents.
[00122] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. I. General Methods
[00123] The practice of the present invention employs, unless otherwise indicated, conventional techniques of immunology, biochemistry, chemistry, molecular biology, microbiology, cell biology, genomics and recombinant DNA, which can be found in such standard textbooks as Molecular Cloning: A Laboratory Manual, 3rd Ed. (Sambrook et al., Harbor Laboratory Press 2001); Short Protocols in Molecular Biology, 4th Ed. (Ausubel et al. eds., John Wiley & Sons 1999); Protein Methods (Bollag et al., John Wiley & Sons 1996); Nonviral Vectors for Gene Therapy (Wagner et al. eds., Academic Press 1999); Viral Vectors (Kaplift & Loewy eds., Academic Press 1995); Immunology Methods Manual (I. Lefkovits ed., Academic Press 1997); and Cell and Tissue Culture: Laboratory Procedures in Biotechnology (Doyle & Griffiths, John Wiley & Sons 1998), the disclosures of which are incorporated herein by reference.
[00124] Where a range of values is provided, it is understood that endpoints are included and that each intervening value, to the tenth of the unit of the lower limit unless the context clearly dictates otherwise, between the upper and lower limit of that range and any other stated or intervening value in that stated range, is encompassed. The upper and lower limits of these smaller ranges may independently be included in the smaller ranges, and are also encompassed, subject to any specifically excluded limit in the stated range. Where the stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included .
[00125] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. All publications mentioned herein are incorporated herein by reference to disclose and describe the methods and / or materials in connection with which the publications are cited.
[00126] It must be noted that as used herein and in the appended claims, the singular forms “a,” “an,” and “the” include plural referents unless the context clearly dictates otherwise.
[00127] It will be appreciated that certain features of the disclosure, which are, for clarity, described in the context of separate embodiments, may also be provided in combination in a single embodiment. In other cases, various features of the disclosure, which are, for brevity, described in the context of a single embodiment, may also be provided separately or in any suitable sub-combination. It is intended that all combinations of the embodiments pertaining to the disclosure are specifically embraced by the present disclosure and are disclosed herein just as if each and every combination was individually and explicitly disclosed. In addition, all subcombinations of the various embodiments and elements thereof are also specifically embraced by the present disclosure and are disclosed herein just as if each and every such sub-combination was individually and explicitly disclosed herein. II. CasX: gNA Systems
[00128] In a first aspect, the present disclosure provides CasX:gNA systems comprising a CasX protein and one or more guide nucleic acids (gNA) for use in modifying or editing a target nucleic acid, inclusive of coding and non-coding regions. The terms CasX protein and CasX are used interchangeably herein; the terms CasX variant protein and CasX variant are used interchangeably herein. The CasX protein and gNA of the CasX:gNA systems provided herein each independently may be a reference CasX protein, a CasX variant protein, a reference gNA, a gNA variant, or any combination of a reference CasX protein, reference gNA, CasX variant protein, or gNA variant. A gNA and a CasX protein, a gNA variant and CasX variant, or any combination thereof can form a complex and bind via non-covalent interactions, referred to herein as a ribonucleoprotein (RNP) complex. In some embodiments, the use of a precomplexed CasX:gNA confers advantages in the delivery of the system components to a cell or target nucleic acid for editing of the target nucleic acid. In the RNP, the gNA can provide target specificity to the RNP complex by including a spacer sequence (targeting sequence) having a nucleotide sequence that is complementary to a sequence of a target nucleic acid. In the RNP, the CasX protein of the pre-complexed CasX:gNA provides the site-specific activity and is guided to a target site (and further stabilized at a target site) within a target nucleic acid sequence to be modified by virtue of its association with the gNA. The CasX protein of the RNP complex provides the site-specific activities of the complex such as binding, cleavage, or nicking of the target sequence by the CasX protein. Provided herein are compositions and cells comprising the reference CasX proteins, CasX variant proteins, reference gNAs, gNA variants, and CasX:gNA gene editing pairs of any combination of CasX and gNA, as well as delivery modalities comprising the CasX:gNA. In other embodiments, the disclosure provides vectors encoding or comprising the CasX:gNA pair and, optionally, donor templates for the production and / or delivery of the CasX:gNA systems. Also provided herein are methods of making CasX proteins and gNA, as well as methods of using the CasX and gNA, including methods of gene editing and methods of treatment. The CasX proteins and gNA components of the CasX:gNA and their features, as well as the delivery modalities and the methods of using the compositions are described more fully, below.
[00129] The donor templates of the CasX:gNA systems are designed depending on whether they are utilized to correct mutations in a target gene or insert a transgene at a different locus in the genome (a “knock-in”), or are utilized to disrupt the expression of a gene product that is aberrant; e.g., it comprises one or more mutations reducing expression of the gene product or rendering the protein dysfunctional (a “knock-down” or “knock-out”). In some embodiments, the donor template is a single stranded DNA template or a single stranded RNA template. In other embodiments, the donor template is a double stranded DNA template. In some embodiments, the CasX:gNA systems utilized in the editing of the target nucleic acid comprises a donor template having all or at least a portion of an open reading frame of a gene in the target nucleic acid for insertion of a corrective, wild-type sequence to correct a defective protein. In other cases, the donor template comprises all or a portion of a wild-type gene for insertion at a different locus in the genome for expression of the gene product. In still other cases, a portion of the gene can be inserted upstream (‘5) of the mutation in the target nucleic acid, wherein the donor template gene portion spans to the C-terminus of the gene, resulting, upon its insertion into the target nucleic acid, in expression of the gene product. In other embodiments, the donor template can comprise one or more mutations in an encoding sequence compared to a normal, wild-type sequence of the target gene utilized for insertion for either knocking out or knocking down (described more fully, below) the defective target nucleic acid sequence. In other embodiments, the donor template can comprise regulatory elements, an intron, or an intron-exon junction having sequences specifically designed to knock-down or knock-out a defective gene or, in the alternative, to knock-in a corrective sequence to permit the expression of a functional gene product. In some embodiments, the donor polynucleotide comprises at least about 10, at least about 20, at least about 50, at least about 100, at least about 200, at least about 300, at least about 400, at least about 500, at least about 600, at least about 700, at least about 800, at least about 900, at least about 1000, at least about 10,000, at least about 15,000, at least about 25,000, at least about 50,000, at least about 100,000 or at least about 200,000 nucleotides. Provided that there are stretches of DNA sequence with sufficient numbers of nucleotides having sufficient homology flanking the cleavage site(s) of the target nucleic acid sequence targeted by the CasX:gNA (i.e., 5’ and 3’ to the cleavage site) to support homology-directed repair (the flanking regions being “homologous arms”), use of such donor templates can result in its integration into the target nucleic acid by HDR. In other cases, the donor template can be inserted by non-homologous end joining (NHEJ; which does not require homologous arms) or by microhomology-mediated end joining (MMEJ; which requires short regions of homology on the 5’ and 3’ ends). In some embodiments, the donor template comprises homologous arms on the 5’ and 3’ ends, each having at least about 2, at least about 10, at least about 20, at least about 30, at least about 50, at least about 100, at least about 150, at least about 300, at least about 1000, at least about 1500 or more nucleotides having homology with the sequences flanking the intended cleave site(s) of the target nucleic acid. In some embodiments, the CasX:gNA systems utilize two or more gNA with targeting sequences complementary to overlapping or different regions of the target nucleic acid such that the defective sequence can be excised by multiple doublestranded breaks or by nicking in locations flanking the defective sequence and the donor template inserted by HDR to replace the excised sequence. In the foregoing, the gNA would be designed to contain targeting sequences that are 5' and 3' to the individual site or sequence to be excised. By such appropriate selection of the targeting sequences of the gNA, defined regions of the target nucleic acid can be edited using the CasX:gNA systems described herein. III. Guide Nucleic Acids of the CasX:gNA Systems
[00130] In other aspects, the disclosure provides guide nucleic acids (gNA) utilized in the CasX:gNA systems, and have utility in editing of a target nucleic acid. The present disclosure provides specifically-designed gNAs with targeting sequences (or "spacers") that are complementary to (and are therefore able to hybridize with) the target nucleic acid as a component of the gene editing CasX:gNA systems. It is envisioned that in some embodiments, multiple gNAs (e.g., multiple gRNAs) are delivered by the CasX:gNA system for the modification of different regions of a gene, including regulatory elements, an exon, an intron, or an intron-exon junction. In some embodiments, the targeting sequence of the gNA is complementary to a sequence comprising one or more single nucleotide polymorphisms (SNPs) of the target nucleic. In other embodiments, the targeting sequence of the gNA is complementary to a sequence of an intergenic region. For example, when a deletion of a protein-encoding gene is desired, a pair of gNAs with targeting sequences to different or overlapping regions of the target nucleic acid sequence can be used in order to bind and cleave at two different sites within the gene that can then be edited by indel formation or homology-directed repair (HDR), which, in the case of HDR, utilizes a donor template that is inserted to replace the deleted sequence to complete the editing. a. Reference gNA and gNA variants
[00131] In some embodiments, a gNA of the present disclosure comprises a sequence of a naturally-occurring gNA (“reference gNA”). In other cases, a reference gNA of the disclosure may be subjected to one or more mutagenesis methods, such as the mutagenesis methods described herein, which may include Deep Mutational Evolution (DME), deep mutational scanning (DMS), error prone PCR, cassette mutagenesis, random mutagenesis, staggered extension PCR, gene shuffling, or domain swapping, in order to generate one or more gNA variants with enhanced or varied properties relative to the reference gNA. gNA variants also include variants comprising one or more exogenous sequences, for example fused to either the 5’ or 3’ end, or inserted internally. The activity of reference gNAs may be used as a benchmark against which the activity of gNA variants are compared, thereby measuring improvements in function or other characteristics of the gNA variants. In other embodiments, a reference gNA may be subjected to one or more deliberate, targeted mutations in order to produce a gNA variant, for example a rationally-designed variant. As used herein, the terms gNA, gRNA, and gDNA cover naturally-occurring molecules (reference molecules), as well as sequence variants.
[00132] In some embodiments, the gNA is a deoxyribonucleic acid molecule (“gDNA”); in some embodiments, the gNA is a ribonucleic acid molecule (“gRNA”), and in other embodiments, the gNA is a chimera, and comprises both DNA and RNA.
[00133] The gNAs of the disclosure comprise two segments; a targeting sequence and a protein-binding segment (which constitutes the scaffold, discussed herein). The targeting segment of a gNA includes a nucleotide sequence (referred to interchangeably herein as a guide sequence, a spacer, a targeting sequence, or a targeting region) that is complementary to (and therefore hybridizes with) a specific sequence (a target site) within the target nucleic acid sequence (e.g., a target ssRNA, a target ssDNA, the complementary strand of a double stranded target DNA, etc.), described more fully below.
[00134] The targeting sequence of a gNA is capable of binding to a target nucleic acid sequence, including a coding sequence, a complement of a coding sequence, a non-coding sequence, and to regulatory elements. The protein-binding segment (or “protein-binding sequence”) interacts with (e.g., binds to) a CasX protein. The protein-binding segment is alternatively referred to herein as a “scaffold”. In some embodiments, the targeting sequence and scaffold each include complementary stretches of nucleotides that hybridize to one another to form a double stranded duplex (e.g. dsRNA duplex for a gRNA). Site-specific binding and / or cleavage of a target nucleic acid sequence (e.g., genomic DNA) by the CasX:gNA can occur at one or more locations of a target nucleic acid, determined by base-pairing complementarity between the targeting sequence of the gNA and the target nucleic acid sequence.
[00135] The gNA provides target specificity to the complex by having a nucleotide sequence that is complementary to a target sequence of a target nucleic acid. The CasX of the complex provides the site-specific activities of the complex such as binding, cleavage, or nicking of the target sequence of the target nucleic acid by the CasX nuclease and / or an activity provided by a fusion partner in case of a CasX containing fusion protein, described below. In some embodiments, the disclosure provides gene editing pairs of a CasX and gNA of any of the embodiments described herein that are capable of being bound together prior to their use for gene editing and, thus, are “pre-complexed” as the RNP. The use of a pre-complexed RNP confers advantages in the delivery of the system components to a cell or target nucleic acid sequence for editing of the target nucleic acid sequence. The CasX protein of the RNP provides the site-specific activity that is guided to a target site (e.g., stabilized at a target site) within a target nucleic acid sequence by virtue of its association with the guide RNA comprising a targeting sequence.
[00136] In some embodiments, wherein the gNA is a gRNA, the term “targeter” or “targeter RNA” is used herein to refer to a crRNA-like molecule (crRNA: "CRISPR RNA") of a CasX dual guide RNA (dgRNA). In a single guide RNA (sgRNA), the “activator" and the "targeter” are linked together, e.g., by intervening nucleotides). Thus, for example, a guide RNA (dgRNA or sgRNA) comprises a guide sequence and a duplex-forming segment of a crRNA, which can also be referred to as a crRNA repeat. Because the targeter sequence of a guide sequence hybridizes with a specific target nucleic acid sequence, a targeter can be modified by a user to hybridize with a desired target nucleic acid sequence. In some embodiments, the sequence of a targeter may often be a non-naturally occurring sequence. The targeter and the activator each have a duplex-forming segment, where the duplex forming segment of the targeter and the duplex-forming segment of the activator have complementarity with one another and hybridize to one another to form a double stranded duplex (dsRNA duplex for a gRNA). In some embodiments, a targeter comprises both the guide sequence of the CasX guide RNA and a stretch of nucleotides that forms one half of the dsRNA duplex of the protein-binding segment of the gNA. A corresponding tracrRNA-like molecule (the activator “trans-acting CRISPR RNA”) also comprises a duplex-forming stretch of nucleotides that forms the other half of the dsRNA duplex of the protein-binding segment of the CasX guide RNA. In some cases the activator comprises one or more stem loops that can interact with CasX protein. Thus, a targeter and an activator, as a corresponding pair, hybridize to form a CasX dual guide NA, referred to herein as a “dual guide NA”, a “dgNA”, a “double-molecule guide NA”, or a “two-molecule guide NA”.
[00137] In some embodiments, the activator and targeter of the reference gNA are covalently linked to one another and comprise a single molecule, referred to herein as a “single-molecule guide NA,” “one-molecule guide NA,” “single guide NA”, “single guide RNA”, a “singlemolecule guide RNA,” a “one-molecule guide RNA”, a “single guide DNA”, a “single-molecule DNA,” or a “one-molecule guide DNA”, (“sgNA”, “sgRNA”, or a “sgDNA”). In some embodiments, the sgNA includes an “activator” or a “targeter” and thus can be an “activator-RNA” and a “targeter-RNA,” respectively.
[00138] The reference gRNAs of the disclosure comprise four distinct regions, or domains: the RNA triplex, the scaffold stem, the extended stem, and the targeting sequence (specific for a target nucleic acid. The RNA triplex, the scaffold stem, and the extended stem, together, are referred to as the “scaffold” of the reference gNA, based upon which further gNA variants are generated. b. RNA triplex
[00139] In some embodiments of the guide NAs provided herein, the gNA comprises an RNA triplex, and the RNA triplex comprises the sequence of a UUU--Nx(~4-15)--UUU stem loop (SEQ ID NO: 241) that ends with an AAAG after 2 intervening stem loops (the scaffold stem loop and the extended stem loop), forming a pseudoknot that may also extend past the triplex into a duplex pseudoknot. The UU-UUU-AAA sequence of the triplex forms as a nexus between the targeting sequence, scaffold stem, and extended stem. In exemplary gRNAs, the UUU-loop-UUU region is coded for first, then the scaffold stem loop, and then the extended stem loop, which is linked by the tetraloop, and then an AAAG closes off the triplex before becoming the targeting sequence. c. Scaffold Stem Loop
[00140] In some embodiments of gNAs of the disclosure, the triplex region is followed by the scaffold stem loop. The scaffold stem loop is a region of the gNA that is bound by CasX protein (such as a reference or CasX variant protein). In some embodiments, the scaffold stem loop is a fairly short and stable stem loop, and increases the overall stability of the gNA. In some cases, the scaffold stem loop does not tolerate many changes, and requires some form of an RNA bubble. In some embodiments, the scaffold stem is necessary for gNA function. While it is perhaps analogous to the nexus stem of Cas9 as being a critical stem loop, the scaffold stem of a gNA, in some embodiments, has a necessary bulge (RNA bubble) that is different from many other stem loops found in CRISPR / Cas systems. In some embodiments, the presence of this bulge is conserved across gNA that interact with different CasX proteins. An exemplary sequence of a scaffold stem loop sequence of a gNA comprises the sequence CCAGCGACUAUGUCGUAUGG (SEQ ID NO: 242). In other embodiments, the disclosure provides gNA variants wherein the scaffold stem loop is replaced with an RNA stem loop sequence from a heterologous RNA source with proximal 5’ and 3’ ends, such as, but not limited to stem loop sequences selected from MS2, QP, UI hairpin II, Uvsx, or PP7 stem loops. In some cases, the heterologous RNA stem loop of the gNA is capable of binding a protein, an RNA structure, a DNA sequence, or a small molecule. d. Extended Stem Loop
[00141] In some embodiments of the gNAs of the disclosure, the scaffold stem loop is followed by the extended stem loop. In some embodiments, the extended stem comprises a synthetic tracr and crRNA fusion that is largely unbound by the CasX protein. In some embodiments, the extended stem loop can be highly malleable. In some embodiments, a single guide gRNA is made with a GAAA tetraloop linker or a GAGAAA linker between the tracr and crRNA in the extended stem loop. In some cases, the targeter and activator of a sgNA are linked to one another by intervening nucleotides and the linker can have a length of from 3 to 20 nucleotides. In some embodiments of the sgNAs of the disclosure, the extended stem is a large 32-bp loop that sits outside of the CasX protein in the ribonucleoprotein complex. An exemplary sequence of an extended stem loop sequence of a sgNA comprises the sequence GCGCUUAUUUAUCGGAGAGAAAUCCGAUAAAUAAGAAGC (SEQ ID NO: 15). In some embodiments, the extended stem loop comprises a GAGAAA spacing sequence. In some embodiments, the disclosure provides gNA variants wherein the extended stem loop is replaced with an RNA stem loop sequence from a heterologous RNA source with proximal 5’ and 3’ ends, such as, but not limited to stem loop sequences selected from MS2, QP, UI hairpin n, Uvsx, or PP7 stem loops. In such cases, the heterologous RNA stem loop increases the stability of the gNA. In other embodiments, the disclosure provides gNA variants having an extended stem loop region comprising at least 10, at least 100, at least 500, at least 1000, or at least 10,000 nucleotides. e. Targeting Sequence
[00142] In some embodiments of the gNAs of the disclosure, the extended stem loop is followed by a region that forms part of the triplex, and then the targeting sequence (or "spacer"). The targeting sequence can be designed to target the CasX ribonucleoprotein holo complex to a specific region of the target nucleic acid sequence. Thus, the gNA targeting sequences of the gNAs of the disclosure have sequences complementarity to, and therefore can hybridize to, a portion of the target nucleic acid in a nucleic acid in a eukaryotic cell, (e.g., a eukaryotic chromosome, chromosomal sequence, a eukaryotic RNA, etc.) as a component of the RNP when any one of the PAM sequences TTC, ATC, GTC, or CTC is located 1 nucleotide 5’ to the nontarget strand sequence complementary to the target sequence.
[00143] In some embodiments, the disclosure provides a gNA wherein the targeting sequence of the gNA is complementary to a target nucleic acid sequence comprising one or more mutations compared to a wild-type gene sequence for purposes of editing the sequence comprising the mutations with the CasX:gNA systems of the disclosure. In some embodiments, the targeting sequence of a gNA is designed to be specific for an exon of the gene of the target nucleic acid. In other embodiments, the targeting sequence of a gNA is designed to be specific for an intron of the gene of the target nucleic acid. In other embodiments, the targeting sequence of the gNA is designed to be specific for an intron-exon junction of the gene of the target nucleic acid. In other embodiments, the targeting sequence of the gNA is designed to be specific for a regulatory element of the gene of the target nucleic acid. In some embodiments, the targeting sequence of the gNA is designed to be complementary to a sequence comprising one or more single nucleotide polymorphisms (SNPs) in a gene of the target nucleic acid. SNPs that are within the coding sequence or within non-coding sequences are both within the scope of the instant disclosure. In other embodiments, the targeting sequence of the gNA is designed to be complementary to a sequence of an intergenic region of the gene of the target nucleic acid.
[00144] In some embodiments, the targeting sequence of a gNA is designed to be specific for a regulatory element that regulates expression of the gene product of the target nucleic acid. Such regulatory elements include, but are not limited to promoter regions, enhancer regions, intergenic regions, 5' untranslated regions (5' UTR), 3' untranslated regions (3' UTR), conserved elements, and regions comprising cis-regulatoiy elements. The promoter region is intended to encompass nucleotides within 5 kb of the initiation point of the encoding sequence or, in the case of gene enhancer elements or conserved elements, can be thousands of bp, hundreds of thousands of bp, or even millions of bp away from the encoding sequence of the gene of the target nucleic acid. In some embodiments of the foregoing, the targets are those in which the encoding gene of the target is intended to be knocked out or knocked down such that the encoded protein comprising mutations is not expressed or is expressed at a lower level in a cell.
[00145] In some embodiments, the targeting sequence of a gNA has between 14 and 35 consecutive nucleotides. In some embodiments, the targeting sequence has 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34 or 35 consecutive nucleotides. In some embodiments, the targeting sequence of the gNA consists of 20 consecutive nucleotides. In some embodiments, the targeting sequence consists of 19 consecutive nucleotides. In some embodiments, the targeting sequence consists of 18 consecutive nucleotides. In some embodiments, the targeting sequence consists of 17 consecutive nucleotides. In some embodiments, the targeting sequence consists of 16 consecutive nucleotides. In some embodiments, the targeting sequence consists of 15 consecutive nucleotides. In some embodiments, the targeting sequence has 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34 or 35 consecutive nucleotides and the targeting sequence can comprise 0 to 5, 0 to 4, 0 to 3, or 0 to 2 mismatches relative to the target nucleic acid sequence and retain sufficient binding specificity such that the RNP comprising the gNA comprising the targeting sequence can form a complementary bond with respect to the target nucleic acid.
[00146] In some embodiments, the CasX:gNA system comprises a first gNA and further comprises a second (and optionally a third, fourth, fifth, or more) gNA, wherein the second gNA or additional gNA has a targeting sequence complementary to a different or overlapping portion of the target nucleic acid sequence compared to the targeting sequence of the first gNA such that multiple points in the target nucleic acid are targeted, and for example, multiple breaks are introduced in the target nucleic acid by the CasX. It will be understood that in such cases, the second or additional gNA is complexed with an additional copy of the CasX protein. By selection of the targeting sequences of the gNA, defined regions of the target nucleic acid sequence bracketing a mutation can be modified or edited using the CasX:gNA systems described herein, including facilitating the insertion of a donor template. f. gNA scaffolds
[00147] With the exception of the targeting sequence region, the remaining regions of the gNA are referred to herein as the scaffold. In some embodiments, the gNA scaffolds are derived from naturally-occurring sequences, described below as reference gNA. In other embodiments, the gNA scaffolds are variants of reference gNA wherein mutations, insertions, deletions or domain substitutions are introduced to confer desirable properties on the gNA.
[00148] In some embodiments, a reference gRNA comprises a sequence isolated or derived from Deltaproteobacteria. In some embodiments, the sequence is a CasX tracrRNA sequence. Exemplary CasX reference tracrRNA sequences isolated or derived from Deltaproteobacteria may include: ACAUCUGGCGCGUUUAUUCCAUUACUUUGGAGCCAGUCCCAGCGACUAUGUCGU AUGGACGAAGCGCUUAUUUAUCGGAGA (SEQ ID NO: 6) and ACAUCUGGCGCGUUUAUUCCAUUACUUUGGAGCCAGUCCCAGCGACUAUGUCGU AUGGACGAAGCGCUUAUUUAUCGG (SEQ ID NO: 7). Exemplary crRNA sequences isolated or derived from Deltaproteobacteria may comprise a sequence of CCGAUAAGUAAAACGCAUCAAAG (SEQ ID NO: 243). In some embodiments, a reference gNA comprises a sequence at least 60% identical, at least 65% identical, at least 70% identical, at least 75% identical, at least 80% identical, at least 81% identical, at least 82% identical, at least 83% identical, at least 84% identical, at least 85% identical, at least 86% identical, at least 86% identical, at least 87% identical, at least 88% identical, at least 89% identical, at least 89% identical, at least 90% identical, at least 91% identical, at least 92% identical, at least 93% identical, at least 94% identical, at least 95% identical, at least 96% identical, at least 97% identical, at least 98% identical, at least 99% identical, at least 99.5% identical or 100% identical to a sequence isolated or derived from Deltaproteobacteria.
[00149] In some embodiments, a reference guide RNA comprises a sequence isolated or derived from Planctomycetes. In some embodiments, the sequence is a CasX tracrRNA sequence. Exemplary reference tracrRNA sequences isolated or derived from Planctomycetes may include: UACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUA UGGGUAAAGCGCUUAUUUAUCGGAGA (SEQ ID NO: 8) and UACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUA UGGGUAAAGCGCUUAUUUAUCGG (SEQ ID NO: 9). Exemplary crRNA sequences isolated or derived from Planctomycetes may comprise a sequence of UCUCCGAUAAAUAAGAAGCAUCAAAG (SEQ ID NO: 244). In some embodiments, a reference gNA comprises a sequence at least 60% identical, at least 65% identical, at least 70% identical, at least 75% identical, at least 80% identical, at least 81% identical, at least 82% identical, at least 83% identical, at least 84% identical, at least 85% identical, at least 86% identical, at least 86% identical, at least 87% identical, at least 88% identical, at least 89% identical, at least 89% identical, at least 90% identical, at least 91% identical, at least 92% identical, at least 93% identical, at least 94% identical, at least 95% identical, at least 96% identical, at least 97% identical, at least 98% identical, at least 99% identical, at least 99.5% identical or 100% identical to a sequence isolated or derived from Planctomycetes.
[00150] In some embodiments, a reference gNA comprises a sequence isolated or derived from Candidatus Sungbacteria. In some embodiments, the sequence is a CasX tracrRNA sequence. Exemplary CasX reference tracrRNA sequences isolated or derived from Candidatus Sungbacteria may comprise sequences of: GUUUACACACUCCCUCUCAUAGGGU (SEQ ID NO: 10), GUUUACACACUCCCUCUCAUGAGGU (SEQ ID NO: 11), UUUUACAUACCCCCUCUCAUGGGAU (SEQ ID NO: 12) and GUUUACACACUCCCUCUCAUGGGGG (SEQ ID NO: 13). In some embodiments, a reference guide RNA comprises a sequence at least 60% identical, at least 65% identical, at least 70% identical, at least 75% identical, at least 80% identical, at least 81% identical, at least 82% identical, at least 83% identical, at least 84% identical, at least 85% identical, at least 86% identical, at least 86% identical, at least 87% identical, at least 88% identical, at least 89% identical, at least 89% identical, at least 90% identical, at least 91% identical, at least 92% identical, at least 93% identical, at least 94% identical, at least 95% identical, at least 96% identical, at least 97% identical, at least 98% identical, at least 99% identical, at least 99.5% identical or 100% identical to a sequence isolated or derived from Candidatus Sungbacteria.
[00151] Table 1 provides the sequences of reference gRNA tracr, cr and scaffold sequences. In some embodiments, the disclosure provides gNA sequences wherein the gNA has a scaffold comprising a sequence having at least one nucleotide modification relative to a reference gNA sequence having a sequence of any one of SEQ ID NOS: 4-16 of Table 1. It will be understood that in those embodiments wherein a vector comprises a DNA encoding sequence for a gNA, or where a gNA is a gDNA or a chimera of RNA and DNA, that thymine (T) bases can be substituted for the uracil (U) bases of any of the gNA sequence embodiments described herein. Table 1. Reference gRNA tracr, cr and scaffold sequences SEQ ID NO. Nucleotide Sequence 4 ACAUCUGGCGCGUUUAUUCCAUUACUUUGGAGCCAGUCCCAGCGACUAUGUCGU AUGGACGAAGCGCUUAUUUAUCGGAGAGAAACCGAUAAGUAAAACGCAUCAAAG 5 UACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUA UGGGUAAAGCGCUUAUUUAUCGGAGAGAAAUCCGAUAAAUAAGAAGCAUCAAAG 6 ACAUCUGGCGCGUUUAUUCCAUUACUUUGGAGCCAGUCCCAGCGACUAUGUCGU AUGGACGAAGCGCUUAUUUAUCGGAGA 7 ACAUCUGGCGCGUUUAUUCCAUUACUUUGGAGCCAGUCCCAGCGACUAUGUCGU AUGGACGAAGCGCUUAUUUAUCGG 8 UACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUA UGGGUAAAGCGCUUAUUUAUCGGAGA 9 UACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUA UGGGUAAAGCGCUUAUUUAUCGG 10 GUUUACACACUCCCUCUCAUAGGGU 11 GUUUACACACUCCCUCUCAUGAGGU 12 UUUUACAUACCCCCUCUCAUGGGAU 13 GUUUACACACUCCCUCUCAUGGGGG 14 CCAGCGACUAUGUCGUAUGG 15 GCGCUUAUUUAUCGGAGAGAAAUCCGAUAAAUAAGAAGC 16 GGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUGGG UAAAGCGCUUAUUUAUCGGA g. gNA Variants
[00152] In another aspect, the disclosure relates to guide nucleic acid variants (referred to herein alternatively as “gNA variant” or “gRNA variant”), which comprise one or more modifications relative to a reference gRNA scaffold. As used herein, “scaffold” refers to all parts to the gNA necessary for gNA function with the exception of the spacer sequence.
[00153] In some embodiments, a gNA variant comprises one or more nucleotide substitutions, insertions, deletions, or swapped or replaced regions relative to a reference gRNA sequence of the disclosure. In some embodiments, a mutation can occur in any region of a reference gRNA scaffold to produce a gNA variant. In some embodiments, the scaffold of the gNA variant sequence has at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, or at least 70%, at least 80%, at least 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to the sequence of SEQ ID NO: 4 or SEQ ID NO: 5.
[00154] In some embodiments, a gNA variant comprises one or more nucleotide changes within one or more regions of the reference gRNA scaffold that improve a characteristic of the reference gRNA. Exemplary regions include the RNA triplex, the pseudoknot, the scaffold stem loop, and the extended stem loop. In some cases, the variant scaffold stem further comprises a bubble. In other cases, the variant scaffold further comprises a triplex loop region. In still other cases, the variant scaffold further comprises a 5' unstructured region. In some embodiments, the gNA variant scaffold comprises a scaffold stem loop having at least 60% sequence identity, at least 70% sequence identity, at least 80% sequence identity, at least 90% sequence identity, at least 95% sequence identity, or at least 99% sequence identity to SEQ ID NO: 14. In some embodiments, the gNA variant scaffold comprises a scaffold stem loop having at least 60% sequence identity to SEQ ID NO: 14. In other embodiments, the gNA variant comprises a scaffold stem loop having the sequence of CCAGCGACUAUGUCGUAGUGG (SEQ ID NO: 245). In other embodiments, the disclosure provides a gNA scaffold comprising, relative to SEQ ID NO:5, a C18G substitution, a G55 insertion, a UI deletion, and a modified extended stem loop in which the original 6 nt loop and 13 most-loop-proximal base pairs (32 nucleotides total) are replaced by a Uvsx hairpin (4 nt loop and 5 loop-proximal base pairs; 14 nucleotides total) and the loop-distal base of the extended stem was converted to a fully base-paired stem contiguous with the new Uvsx hairpin by deletion of the A99 and substitution of G65U. In the foregoing embodiment, the gNA scaffold comprises the sequence ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGUGGGUAAAGCUCCCUCUUCGGAG GGAGCAUCAAAG (SEQ ID NO: 2238).
[00155] All gNA variants that have one or more improved characteristics, or add one or more new functions when the variant gNA is compared to a reference gRNA described herein, are envisaged as within the scope of the disclosure. A representative example of such a gNA variant is guide 174 (SEQ ID NO: 2238), the design of which is described in the Examples. In some embodiments, the gNA variant adds a new function to the RNP comprising the gNA variant. In some embodiments, the gNA variant has an improved characteristic selected from: improved stability; improved solubility; improved transcription of the gNA; improved resistance to nuclease activity; increased folding rate of the gNA; decreased side product formation during folding; increased productive folding; improved binding affinity to a CasX protein; improved binding affinity to a target DNA when complexed with a CasX protein; improved gene editing when complexed with a CasX protein; improved specificity of editing when complexed with a CasX protein; and improved ability to utilize a greater spectrum of one or more PAM sequences, including ATC, CTC, GTC, or TTC, in the editing of target DNA when complexed with a CasX protein, and any combination thereof. In some cases, the one or more of the improved characteristics of the gNA variant is at least about 1.1 to about 100,000-fold improved relative to the reference gNA of SEQ ID NO: 4 or SEQ ID NO: 5. In other cases, the one or more improved characteristics of the gNA variant is at least about 1.1, at least about 10, at least about 100, at least about 1000, at least about 10,000, at least about 100,000-fold or more improved relative to the reference gNA of SEQ ID NO: 4 or SEQ ID NO: 5. In other cases, the one or more of the improved characteristics of the gNA variant is about 1.1 to 100,00-fold, about 1.1 to 10,00-fold, about 1.1 to 1,000-fold, about 1.1 to 500-fold, about 1.1 to 100-fold, about 1.1 to 50fold, about 1.1 to 20-fold, about 10 to 100,00-fold, about 10 to 10,00-fold, about 10 to 1,000fold, about 10 to 500-fold, about 10 to 100-fold, about 10 to 50-fold, about 10 to 20-fold, about 2 to 70-fold, about 2 to 50-fold, about 2 to 30-fold, about 2 to 20-fold, about 2 to 10-fold, about 5 to 50-fold, about 5 to 30-fold, about 5 to 10-fold, about 100 to 100,00-fold, about 100 to 10,00fold, about 100 to 1,000-fold, about 100 to 500-fold, about 500 to 100,00-fold, about 500 to 10,00-fold, about 500 to 1,000-fold, about 500 to 750-fold, about 1,000 to 100,00-fold, about 10,000 to 100,00-fold, about 20 to 500-fold, about 20 to 250-fold, about 20 to 200-fold, about 20 to 100-fold, about 20 to 50-fold, about 50 to 10,000-fold, about 50 to 1,000-fold, about 50 to 500-fold, about 50 to 200-fold, or about 50 to 100-fold, improved relative to the reference gNA of SEQ ID NO: 4 or SEQ ID NO: 5. In other cases, the one or more improved characteristics of the gNA variant is about 1.1-fold, 1.2-fold, 1.3-fold, 1.4-fold, 1.5-fold, 1.6-fold, 1.7-fold, 1.8-fold, 1.9-fold, 2-fold, 3-fold, 4-fold, 5-fold, 6-fold, 7-fold, 8-fold, 9-fold, 10-fold, 11-fold, 12fold, 13-fold, 14-fold, 15-fold, 16-fold, 17-fold, 18-fold, 19-fold, 20-fold, 25-fold, 30-fold, 40- fold, 45-fold, 50-fold, 55-fold, 60-fold, 70-fold, 80-fold, 90-fold, 100-fold, 110-fold, 120-fold, 130-fold, 140-fold, 150-fold, 160-fold, 170-fold, 180-fold, 190-fold, 200-fold, 210-fold, 220fold, 230-fold, 240-fold, 250-fold, 260-fold, 270-fold, 280-fold, 290-fold, 300-fold, 310-fold, 320-fold, 330-fold, 340-fold, 350-fold, 360-fold, 370-fold, 380-fold, 390-fold, 400-fold, 425fold, 450-fold, 475-fold, or 500-fold improved relative to the reference gNA of SEQ ID NO: 4 or SEQIDNO: 5.
[00156] In some embodiments, a gNA variant can be created by subjecting a reference gNA to a one or more mutagenesis methods, such as the mutagenesis methods described herein, below, which may include Deep Mutational Evolution (DME), deep mutational scanning (DMS), error prone PCR, cassette mutagenesis, random mutagenesis, staggered extension PCR, gene shuffling, or domain swapping, in order to generate the gNA variants of the disclosure. The activity of reference gNAs may be used as a benchmark against which the activity of gNA variants are compared, thereby measuring improvements in function of gNA variants. In other embodiments, a reference gNA may be subjected to one or more deliberate, targeted mutations, substitutions, or domain swaps in order to produce a gNA variant, for example a rationally designed variant. Exemplary gNA variants produced by such methods are described in the Examples and representative sequences of gNA scaffolds are presented in Table 2.
[00157] In some embodiments, the gNA variant comprises one or more modifications compared to a reference guide nucleic acid scaffold sequence, wherein the one or more modification is selected from: at least one nucleotide substitution in a region of the reference gNA at least one nucleotide deletion in a region of the reference gNA; at least one nucleotide insertion in a region of the reference gNA ; a substitution of all or a portion of a region of the reference gNA; a deletion of all or a portion of a region of the reference gNA; or any combination of the foregoing. In some cases, the modification is a substitution of 1 to 15 consecutive or non-consecutive nucleotides in the reference gNA in one or more regions. In other cases, the modification is a deletion of 1 to 10 consecutive or non-consecutive nucleotides in the reference gNA in one or more regions. In other cases, the modification is an insertion of 1 to 10 consecutive or non-consecutive nucleotides in the reference gNA in one or more regions. In other cases, the modification is a substitution of the scaffold stem loop or the extended stem loop with an RNA stem loop sequence from a heterologous RNA source with proximal 5' and 3' ends. In some cases, a gNA variant of the disclosure comprises two or more modifications in one region relative to a reference gRNA. In other cases, a gNA variant of the disclosure comprises modifications in two or more regions. In other cases, a gNA variant comprises any combination of the foregoing modifications described in this paragraph. In some embodiments, exemplary modifications of gNA of the disclosure include the modifications of Table 24.
[00158] In some embodiments, a 5' G is added to a gNA variant sequence, relative to a reference gRNA, for expression in vivo, as transcription from a U6 promoter is more efficient and more consistent with regard to the start site when the +1 nucleotide is a G. In other embodiments, two 5' Gs are added to generate a gNA variant sequence for in vitro transcription to increase production efficiency, as T7 polymerase strongly prefers a Gin the +1 position and a purine in the +2 position. In some cases, the 5’ G bases are added to the reference scaffolds of Table 1. In other cases, the 5’ G bases are added to the variant scaffolds of Table 2.
[00159] Table 2 provides exemplary gNA variant scaffold sequences of the disclosure. In Table 2, (-) indicates a deletion at the specified position(s) relative to the reference sequence of SEQ ID NO: 5, (+) indicates an insertion of the specified base(s) at the position indicated relative to SEQ ID NO: 5, (:) indicates the range of bases at the specified start:stop coordinates of a deletion or substitution relative to SEQ ID NO: 5, and multiple insertions, deletions or substitutions are separated by commas; e.g., A14C, T17G. In some embodiments, the gNA variant scaffold comprises any one of the sequences listed in Table 2, SEQ ID NOS: 2101-2280, or a sequence having at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% sequence identity thereto. It will be understood that in those embodiments wherein a vector comprises a DNA encoding sequence for a gNA, or where a gNA is a gDNA or a chimera of RNA and DNA, that thymine (T) bases can be substituted for the uracil (U) bases of any of the gNA sequence embodiments described herein. Table 2. Exemplary gNA Variant Scaffold Sequences SEQ ID NO: NAME or Modification NUCLEOTIDE SEQUENCE 2101 phage replication stable UACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUG GGUAAAGCGCAGGUGGGACGACCUCUCGGUCGUCCUAUCUGAAGCAUCAAAG 2102 Kissing loopbl UACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUG GGUAAAGCGCUGCUCGACGCGUCCUCGAGCAGAAGCAUCAAAG 2103 Kissing loopa UACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUG GGUAAAGCGCUGCUCGCUCCGUUCGAGCAGAAGCAUCAAAG 2104 32, uvsX hairpin GUACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAU GGGUAAAGCGCCCUCUUCGGAGGGAAGCAUCAAAG 2105 PP7 UACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUG SEQ ID NO: NAME or Modification NUCLEOTIDE SEQUENCE GGUAAAGCGCAGGAGUUUCUAUGGAAACCCUGAAGCAUCAAAG 2106 64, trip mut, extended stem truncation GUACUGGCGCCUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAU GGGUAAAGCGCUUACGGACUUCGGUCCGUAAGAAGCAUCAAAG 2107 hyperstable tetraloop UACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUG GGUAAAGCGCUGCGCUUGCGCAGAAGCAUCAAAG 2108 C18G UACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUG GGUAAAGCGCUUAUUUAUCGGAGAGAAAUCCGAUAAAUAAGAAGCAUCAAAG 2109 T17G UACUGGCGCUUUUAUCGCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUG GGUAAAGC GCUUAUUUAU C GGAGAGAAAU C C GAUAAAUAAGAAGCAU CAAAG 2110 CUUCGG loop UACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUG GGUAAAGCGCUUAUUUAUCGGAGACUUCGGUCCGAUAAAUAAGAAGCAUCAAAG 2111 MS2 UACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUG GGUAAAGCGCACAUGAGGAUUACCCAUGUGAAGCAUCAAAG 2112 -1, A2G, -78, G77T GCUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUGG GUAAAGC GCUUAUUUAU C GU GAGAAAU C C GAUAAAUAAGAAGCAU CAAAG 2113 QB UACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUG GGUAAAGCGCUGCAUGUCUAAGACAGCAGAAGCAUCAAAG 2114 45,44 hairpin UACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUG GGUAAAGCGCAGGGCUUCGGCCGAAGCAUCAAAG 2115 U1A UACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUG GGUAAAGCGCAAUCCAUUGCACUCCGGAUUGAAGCAUCAAAG 2116 A14C, TUG UACUGGCGCUUUUCUCGCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUG GGUAAAGC GCUUAUUUAU C GGAGAGAAAU C C GAUAAAUAAGAAGCAU CAAAG 2117 CUUCGG loop modified UACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUG GGUAAAGCGCUUAUUUAUCGGACUUCGGUCCGAUAAAUAAGAAGCAUCAAAG 2118 Kissing loop_b2 UACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUG GGUAAAGCGCUGCUCGUUUGCGGCUACGAGCAGAAGCAUCAAAG 2119 -76:78, -83:87 UACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUG GGUAAAGCGCUUAUUUAUCGAGAGAUAAAUAAGAAGCAUCAAAG 2120 -4 UACGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUGG GUAAAG C GCUUAUUUAU C G GAGAGAAAU C C GAUAAAUAAGAAG CAU CAAAG 2121 extended stem truncation UACUGGCGCCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAU GGGUAAAGCGCUUACGGACUUCGGUCCGUAAGAAGCAUCAAAG 2122 C55 UACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUC GGUAAAGC GCUUAUUUAU C GGAGAGAAAU C C GAUAAAUAAGAAGCAU CAAAG 2123 trip mut UACUGGCGCCUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUG GGUAAAGCGCUUAUUUAUCGGACUUCGGUCCGAUAAAUAAGAAGCAUCAAAG 2124 -76:78 UACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUG GGUAAAGCGCUUAUUUAUCGAGAAAUCCGAUAAAUAAGAAGCAUCAAAG 2125 -1:5 GCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUGGGUAA AG C GCUUAUUUAU C G GAGAGAAAU C C GAUAAAUAAGAAG CAU CAAAG 2126 -83:87 UACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUG GGUAAAGCGCUUAUUUAUCGGAGAGAGAUAAAUAAGAAGCAUCAAAG 2127 =+G28, A82T, -84, UACUGGCGCUUUUAUCUCAUUACUUUGGAGAGCCAUCACCAGCGACUAUGUCGUAU GGGUAAAGCGCUUAUUUAUCGGAGAGUAUCCGAUAAAUAAGAAGCAUCAAAG 2128 =+5 IT UACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUUCGUAU GGGUAAAGCGCUUAUUUAUCGGAGAGAAAUCCGAUAAAUAAGAAGCAUCAAAG 2129 -1:4, +G5A, +G86, AGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUGGGUA AAGCGCUUAUUUAUCGGAGAGAAAUGCCGAUAAAUAAGAAGCAUCAAAG 2130 =+A94 UACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUG GGUAAAGCGCUUAUUUAUCGGAGAGAAAUCCGAUAAAAUAAGAAGCAUCAAAG 2131 =+G72 UACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUG GGUAAAGCGCUUAUUGUAUCGGAGAGAAAUCCGAUAAAUAAGAAGCAUCAAAG SEQ ID NO: NAME or Modification NUCLEOTIDE SEQUENCE 2132 shorten front, CUUCGG loop modified, extend extended GCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUGGGUAA AGCGCUUAUUUAUCGGACUUCGGUCCGAUAAAUAAGCGCAUCAAAG 2133 A14C UACUGGCGCUUUUCUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUG GGUAAAGCGCUUAUUUAUCGGAGAGAAAUCCGAUAAAUAAGAAGCAUCAAAG 2134 -1:3, +G3 GUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUGGG UAAAG C G CUUAUUUAU C G GAGAGAAAU C C GAUAAAUAAGAAG GAU CAAAG 2135 =+C45, +T46 UACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACCUUAUGUCGUA UGGGUAAAGCGCUUAUUUAUCGGAGAGAAAUCCGAUAAAUAAGAAGCAUCAAAG 2136 CUUCGG loop modified, fun start GAUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUGG GUAAAGCGCUUAUUUAUCGGACUUCGGUCCGAUAAAUAAGAAGCAUCAAAG 2137 -93:94 UACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUG GGUAAAGCGCUUAUUUAUCGGAGAGAAAUCCGAUAAAAGAAGCAUCAAAG 2138 =+T45 UACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGAUCUAUGUCGUAU GGGUAAAGCGCUUAUUUAUCGGAGAGAAAUCCGAUAAAUAAGAAGCAUCAAAG 2139 -69, -94 UACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUG GGUAAAGGCUUAUUUAUCGGAGAGAAAUCCGAUAAAAAGAAGCAUCAAAG 2140 -94 UACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUG GGUAAAGCGCUUAUUUAUCGGAGAGAAAUCCGAUAAAAAGAAGCAUCAAAG 2141 modified CUUCGG, minus T in 1st triplex UACUGGCGCUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUGG GUAAAGCGCUUAUUUAUCGGACUUCGGUCCGAUAAAUAAGAAGCAUCAAAG 2142 -1:4, +C4, A14C, T17G, +G72, -76:78, -83:87 CGGCGCUUUUCUCGCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUGGGU AAAG C G CUUAUU GUAU C GAGAGAUAAAUAAGAAG CAU CAAAG 2143 TIC, -73 CACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUG GGUAAAGCGCUUAUUUUCGGAGAGAAAUCCGAUAAAUAAGAAGCAUCAAAG 2144 Scaffold uuCG, stemuuCG. Stem swap, t shorten UACUGGCGCUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUUCGGUCGUAUG GGUAAAGCGCUUAUGUAUCGGCUUCGGCCGAUACAUAAGAAGCAUCAAAG 2145 Scaffold uuCG, stemuuCG. Stem swap UACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUUCGGUCGUAU GGGUAAAGCGCUUAUGUAUCGGCUUCGGCCGAUACAUAAGAAGCAUCAAAG 2146 =+G60 UACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUG GGUGAAAGCGCUUAUUUAUCGGAGAGAAAUCCGAUAAAUAAGAAGCAUCAAAG 2147 no stem Scaffold uuCG UACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUUCGGUCGUAU GGGUAAAG 2148 no stem Scaffold uuCG, fun start GAUGGGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUUCGGUCGUAUGG GUAAAG 2149 Scaffold uuCG, stem uuCG, fun start GAUGGGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUUCGGUCGUAUGG GUAAAGCGCUUAUUUAUCGGCUUCGGCCGAUAAAUAAGAAGCAUCAAAG 2150 Pseudoknots UACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUG GGUAAAGCGCUACACUGGGAUCGCUGAAUUAGAGAUCGGCGUCCUUUCAUUCUAUA UACUUU GGAGUUUUAAAAU GU CU CUAAGUACAGAAGCAU CAAAG 2151 Scaffold uuCG, stem uuCG GGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUUCGGUCGUAUGGGU AAAGCGCUUAUUUAUCGGCUUCGGCCGAUAAAUAAGAAGCAUCAAAG 2152 Scaffold uuCG, stem uuCG, no start GCUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUUCGGUCGUAUG GGUAAAGCGCUUAUUUAUCGGCUUCGGCCGAUAAAUAAGAAGCAUCAAAG SEQ ID NO: NAME or Modification NUCLEOTIDE SEQUENCE 2153 Scaffold uuCG UACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUUCGGUCGUAU GGGUAAAGCGCUUAUUUAUCGGAGAGAAAUCCGAUAAAUAAGAAGCAUCAAAG 2154 =+GCTC36 UACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUGCUCCACCAGCGACUAUGUCG UAUGGGUAAAGCGCUUAUUUAUCGGAGAGAAAUCCGAUAAAUAAGAAGCAUCAAAG 2155 G quadriplex telomere basket+ ends UACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUG GGUAAAGCGGGGUUAGGGUUAGGGUUAGGGAAGCAUCAAAG 2156 G quadriplex M3q UACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUG GGUAAAGCGGAGGGAGGGAGGGAGAGGGAAAGCAUCAAAG 2157 G quadriplex telomere basket no ends UACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUG GGUAAAGCGUUGGGUUAGGGUUAGGGUUAGGGAAAAGCAUCAAAG 2158 45,44 hairpin (old version) UACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUG GGUAAAGCGC--------AGGGCUUCGGCCG---------GAAGCAUCAAAG 2159 Sarcin-ricin loop UACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUG GGUAAAGCGCCUGCUCAGUACGAGAGGAACCGCAGGAAGCAUCAAAG 2160 uvsX, C18G UACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUG GGUAAAGCGCCCUCUUCGGAGGGAAGCAUCAAAG 2161 truncated stem loop, C18G, trip mut (T10C) UACUGGCGCCUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUG GGUAAAGCGCUUACGGACUUCGGUCCGUAAGAAGCAUCAAAG 2162 short phage rep, C18G UACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUG GGUAAAGCGCGGACGACCUCUCGGUCGUCCGAAGCAUCAAAG 2163 phage rep loop, C18G UACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUG GGUAAAGCGCAGGUGGGACGACCUCUCGGUCGUCCUAUCUGAAGCAUCAAAG 2164 =+G18, stacked onto 64 UACUGGCGCCUUUAUCUGCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAU GGGUAAAGCGCUUACGGACUUCGGUCCGUAAGAAGCAUCAAAG 2165 truncated stem loop, C18G, -1 A2G GCUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUGG GUAAAGCGCUUACGGACUUCGGUCCGUAAGAAGCAUCAAAG 2166 phage rep loop, C18G, trip mut (T10C) UACUGGCGCCUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUG GGUAAAGCGCAGGUGGGACGACCUCUCGGUCGUCCUAUCUGAAGCAUCAAAG 2167 short phage rep, C18G, trip mut (T10C) UACUGGCGCCUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUG GGUAAAGCGCGGACGACCUCUCGGUCGUCCGAAGCAUCAAAG 2168 uvsX, trip mut (T10C) UACUGGCGCCUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUG GGUAAAGCGCCCUCUUCGGAGGGAAGCAUCAAAG 2169 truncated stem loop UACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUG GGUAAAGCGCUUACGGACUUCGGUCCGUAAGAAGCAUCAAAG 2170 =+A17, stacked onto 64 UACUGGCGCCUUUAUCAUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAU GGGUAAAGCGCUUACGGACUUCGGUCCGUAAGAAGCAUCAAAG 2171 3' HDV genomic ribozyme UACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUG GGUAAAGCGCUUAUUUAUCGGAGAGAAAUCCGAUAAAUAAGAAGCAUCAAAGGGCC GGCAUGGUCCCAGCCUCCUCGCUGGCGCCGGCUGGGCAACAUUCCGAGGGGACCGU CCCCUCGGUAAUGGCGAAUGGGACCC 2172 phage rep loop, trip mut (T10C) UACUGGCGCCUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUG GGUAAAGCGCAGGUGGGACGACCUCUCGGUCGUCCUAUCUGAAGCAUCAAAG 2173 -79:80 UACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUG GGUAAAGCGCUUAUUUAUCGGAGAAAUCCGAUAAAUAAGAAGCAUCAAAG 2174 short phage rep, trip mut (T10C) UACUGGCGCCUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUG GGUAAAGCGCGGACGACCUCUCGGUCGUCCGAAGCAUCAAAG 2175 extra truncated UACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUG SEQ ID NO: NAME or Modification NUCLEOTIDE SEQUENCE stem loop GGUAAAGCGCCGGACUUCGGUCCGGAAGCAUCAAAG 2176 T17G, C18G UACUGGCGCUUUUAUCGGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUG GGUAAAGC GCUUAUUUAU C GGAGAGAAAU C C GAUAAAUAAGAAGCAU CAAAG 2177 short phage rep UACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUG GGUAAAGCGCGGACGACCUCUCGGUCGUCCGAAGCAUCAAAG 2178 uvsX, C18G, -1 A2G GCUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUGG GUAAAGCGCCCUCUUCGGAGGGAAGCAUCAAAG 2179 uvsX, C18G, trip mut (T10C), -1 A2G, HDV -99 G65U GCUGGCGCCUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUGG GUAAAGCUCCCUCUUCGGAGGGAGCAUCAAAG 2180 3'HDV antigenomic ribozyme UACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUG GGUAAAGCGCUUAUUUAUCGGAGAGAAAUCCGAUAAAUAAGAAGCAUCAAAGGGGU CGGCAUGGCAUCUCCACCUCCUCGCGGUCCGACCUGGGCAUCCGAAGGAGGACGCA CGUCCACUCGGAUGGCUAAGGGAGAGCCA 2181 uvsX, C18G, trip mut (T10C), -1 A2G, HDV AA(98:99)C GCUGGCGCCUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUGG GUAAAGCGCCCUCUUCGGAGGGCGCAUCAAAG 2182 3' HDV ribozyme (Lior Nissim, Timothy Lu) UACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUG GGUAAAGC GCUUAUUUAU C GGAGAGAAAU C C GAUAAAUAAGAAGCAU CAAAGUUUU GGCCGGCAUGGUCCCAGCCUCCUCGCUGGCGCCGGCUGGGCAACAUGCUUCGGCAU GGCGAAUGGGACCCCGGG 2183 TAC(1:3)GA, stacked onto 64 GAUGGCGCCUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUGG GUAAAGCGCUUACGGACUUCGGUCCGUAAGAAGCAUCAAAG 2184 uvsX, -1 A2G GCUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUGG GUAAAGCGCCCUCUUCGGAGGGAAGCAUCAAAG 2185 truncated stem loop, C18G, trip mut (T10C), -1 A2G, HDV -99 G65U GCUGGCGCCUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUGG GUAAAGCU CUUAC GGACUU C GGU C C GUAAGAGCAU CAAAG 2186 short phage rep, C18G, trip mut (T10C), -1 A2G, HDV -99 G65U GCUGGCGCCUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUGG GUAAAGCUCGGACGACCUCUCGGUCGUCCGAGCAUCAAAG 2187 3'sTRSVWT viral Hammerhead ribozyme UACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUG GGUAAAGCGCUUAUUUAUCGGAGAGAAAUCCGAUAAAUAAGAAGCAUCAAAGCCUG UCACCGGAUGUGCUUUCCGGUCUGAUGAGUCCGUGAGGACGAAACAGG 2188 short phage rep, C18G, -1 A2G GCUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUGG GUAAAGCGCGGACGACCUCUCGGUCGUCCGAAGCAUCAAAG 2189 short phage rep, C18G, trip mut (T10C), -1 A2G, 3' genomic HDV GCUGGCGCCUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUGG GUAAAGCGCGGACGACCUCUCGGUCGUCCGAAGCAUCAAAG 2190 phage rep loop, C18G, trip mut (T10C), -1 A2G, HDV -99 G65U GCUGGCGCCUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUGG GUAAAGCUCAGGUGGGACGACCUCUCGGUCGUCCUAUCUGAGCAUCAAAG 2191 3' HDV ribozyme (Owen Ryan, Jamie Cate) UACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUG GGUAAAGCGCUUAUUUAUCGGAGAGAAAUCCGAUAAAUAAGAAGCAUCAAAGGAUG GCCGGCAUGGUCCCAGCCUCCUCGCUGGCGCCGGCUGGGCAACACCUUCGGGUGGC GAAUGGGAC SEQ ID NO: NAME or Modification NUCLEOTIDE SEQUENCE 2192 phage rep loop, C18G, -1 A2G GCUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUGG GUAAAGCGCAGGUGGGACGACCUCUCGGUCGUCCUAUCUGAAGCAUCAAAG 2193 0.14 UACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUACU GGGUAAAGCGCUUAUUUAUCGGAGAGAAAUCCGAUAAAUAAGAAGCAUCAAAG 2194 -78, G77T UACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUG G GUAAAG C G CUUAUUUAU C GU GAGAAAU C C GAUAAAUAAGAAG GAU CAAAG 2195 GUACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAU GGGUAAAGCGCUUAUUUAUCGGAGAGAAAUCCGAUAAAUAAGAAGCAUCAAAG 2196 short phage rep, -1 A2G GCUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUGG GUAAAGCGCGGACGACCUCUCGGUCGUCCGAAGCAUCAAAG 2197 truncated stem loop, C18G, trip mut (T10C), -1 A2G GCUGGCGCCUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUGG GUAAAGCGCUUACGGACUUCGGUCCGUAAGAAGCAUCAAAG 2198 -1, A2G GCUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUGG GUAAAG C G CUUAUUUAU C G GAGAGAAAU C C GAUAAAUAAGAAG CAU CAAAG 2199 truncated stem loop, trip mut (T10C), -1 A2G GCUGGCGCCUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUGG GUAAAGCGCUUACGGACUUCGGUCCGUAAGAAGCAUCAAAG 2200 uvsX, C18G, trip mut (T10C), -1 A2G GCUGGCGCCUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUGG GUAAAGCGCCCUCUUCGGAGGGAAGCAUCAAAG 2201 phage rep loop, -1 A2G GCUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUGG GUAAAGCGCAGGUGGGACGACCUCUCGGUCGUCCUAUCUGAAGCAUCAAAG 2202 phage rep loop, trip mut (T10C), -1 A2G GCUGGCGCCUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUGG GUAAAGCGCAGGUGGGACGACCUCUCGGUCGUCCUAUCUGAAGCAUCAAAG 2203 phage rep loop, C18G, trip mut (T10C), -1 A2G GCUGGCGCCUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUGG GUAAAGCGCAGGUGGGACGACCUCUCGGUCGUCCUAUCUGAAGCAUCAAAG 2204 truncated stem loop, C18G UACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUG GGUAAAGCGCUUACGGACUUCGGUCCGUAAGAAGCAUCAAAG 2205 uvsX, trip mut (T10C), -1 A2G GCUGGCGCCUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUGG GUAAAGCGCCCUCUUCGGAGGGAAGCAUCAAAG 2206 truncated stem loop, -1 A2G GCUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUGG GUAAAGCGCUUACGGACUUCGGUCCGUAAGAAGCAUCAAAG 2207 short phage rep, trip mut (T10C), -1 A2G GCUGGCGCCUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUGG GUAAAGCGCGGACGACCUCUCGGUCGUCCGAAGCAUCAAAG 2208 5'HDV ribozyme (Owen Ryan, Jamie Cate) GAUGGCCGGCAUGGUCCCAGCCUCCUCGCUGGCGCCGGCUGGGCAACACCUUCGGG UGGCGAAUGGGACUACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCG ACUAUGUCGUAUGGGUAAAGCGCUUAUUUAUCGGAGAGAAAUCCGAUAAAUAAGAA GCAUCAAAG 2209 5'HDV genomic ribozyme GGCCGGCAUGGUCCCAGCCUCCUCGCUGGCGCCGGCUGGGCAACAUUCCGAGGGGA CCGUCCCCUCGGUAAUGGCGAAUGGGACCCUACUGGCGCUUUUAUCUCAUUACUUU GAGAGCCAUCACCAGCGACUAUGUCGUAUGGGUAAAGCGCUUAUUUAUCGGAGAGA AAU C C GAUAAAUAAGAAG CAU CAAAG 2210 truncated stem loop, C18G, trip mut (T10C), -1 A2G, HDV AA(98:99)C GCUGGCGCCUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUGG GUAAAGCGCUUACGGACUUCGGUCCGUAAGCGCAUCAAAG 2211 5'env25 pistol ribozyme (with CGUGGUUAGGGCCACGUUAAAUAGUUGCUUAAGCCCUAAGCGUUGAUCUUCGGAUC AGGUGCAAUACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAU SEQ ID NO: NAME or Modification NUCLEOTIDE SEQUENCE an added CUUCGG loop) GUCGUAUGGGUAAAGCGCUUAUUUAUCGGAGAGAAAUCCGAUAAAUAAGAAGCAUC AAAG 2212 5HDV antigenomic ribozyme GGGUCGGCAUGGCAUCUCCACCUCCUCGCGGUCCGACCUGGGCAUCCGAAGGAGGA CGCACGUCCACUCGGAUGGCUAAGGGAGAGCCAUACUGGCGCUUUUAUCUCAUUAC UUUGAGAGCCAUCACCAGCGACUAUGUCGUAUGGGUAAAGCGCUUAUUUAUCGGAG AGAAAU C C GAUAAAUAAGAAG GAU CAAAG 2213 3' Hammerhead ribozyme (Lior Nissim, Timothy Lu) guide scaffold scar UACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUG GGUAAAGCGCUUAUUUAUCGGAGAGAAAUCCGAUAAAUAAGAAGCAUCAAAGCCAG UACUGAUGAGUCCGUGAGGACGAAACGAGUAAGCUCGUCUACUGGCGCUUUUAUCU GAU 2214 =+A27, stacked onto 64 UACUGGCGCCUUUAUCUCAUUACUUUAGAGAGCCAUCACCAGCGACUAUGUCGUAU GGGUAAAGCGCUUACGGACUUCGGUCCGUAAGAAGCAUCAAAG 2215 5 Hammerhead ribozyme (Lior Nissim, Timothy Lu) smaller scar C GACUACU GAU GAGU C C GU GAGGAC GAAAC GAGUAAGCU C GU CUAGU C GUACU GGC GCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUGGGUAAAG C GCUUAUUUAU C GGAGAGAAAU C C GAUAAAUAAGAAGGAU CAAAG 2216 phage rep loop, C18G, trip mut (T10C), -1 A2G, HDV AA(98:99)C GCUGGCGCCUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUGG GUAAAGCGCAGGUGGGACGACCUCUCGGUCGUCCUAUCUGCGCAUCAAAG 2217 -27, stacked onto 64 UACUGGCGCCUUUAUCUCAUUACUUUAGAGCCAUCACCAGCGACUAUGUCGUAUGG GUAAAGCGCUUACGGACUUCGGUCCGUAAGAAGCAUCAAAG 2218 3' Hatchet UACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUG GGUAAAGCGCUUAUUUAUCGGAGAGAAAUCCGAUAAAUAAGAAGCAUCAAAGCAUU CCUCAGAAAAUGACAAACCUGUGGGGCGUAAGUAGAUCUUCGGAUCUAUGAUCGUG CAGAC GUUAAAAU CAG GU 2219 3' Hammerhead ribozyme (Lior Nissim, Timothy Lu) UACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUG GGUAAAGCGCUUAUUUAUCGGAGAGAAAUCCGAUAAAUAAGAAGCAUCAAAGCGAC UACUGAUGAGUCCGUGAGGACGAAACGAGUAAGCUCGUCUAGUCGCGUGUAGCGAA GCA 2220 5 Hatchet CAUUCCUCAGAAAAUGACAAACCUGUGGGGCGUAAGUAGAUCUUCGGAUCUAUGAU CGUGCAGACGUUAAAAUCAGGUUACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCA UCACCAGCGACUAUGUCGUAUGGGUAAAGCGCUUAUUUAUCGGAGAGAAAUCCGAU AAAUAAGAAGCAUCAAAG 2221 5HDV ribozyme (Lior Nissim, Timothy Lu) UUUUGGCCGGCAUGGUCCCAGCCUCCUCGCUGGCGCCGGCUGGGCAACAUGCUUCG GCAUGGCGAAUGGGACCCCGGGUACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCA UCACCAGCGACUAUGUCGUAUGGGUAAAGCGCUUAUUUAUCGGAGAGAAAUCCGAU AAAUAAGAAGCAUCAAAG 2222 5 Hammerhead ribozyme (Lior Nissim, Timothy Lu) C GACUACU GAU GAGU C C GU GAGGAC GAAAC GAGUAAGCU C GU CUAGU C GC GU GUAG CGAAGCAUACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUG UCGUAUGGGUAAAGCGCUUAUUUAUCGGAGAGAAAUCCGAUAAAUAAGAAGCAUCA AAG 2223 3'HH15 Minimal Hammerhead ribozyme UACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUG GGUAAAGCGCUUAUUUAUCGGAGAGAAAUCCGAUAAAUAAGAAGCAUCAAAGGGGA GCCCCGCUGAUGAGGUCGGGGAGACCGAAAGGGACUUCGGUCCCUACGGGGCUCCC 2224 5'RBMX recruiting motif CCACCCCCACCACCACCCCCACCCCCACCACCACCCUACUGGCGCUUUUAUCUCAU UACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUGGGUAAAGCGCUUAUUUAUCG GAGAGAAAU C C GAUAAAUAAGAAG CAU CAAAG 2225 3' Hammerhead ribozyme (Lior Nissim, Timothy Lu) smaller scar UACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUG GGUAAAGCGCUUAUUUAUCGGAGAGAAAUCCGAUAAAUAAGAAGCAUCAAAGCGAC UACU GAU GAGU C C GU GAG GAC GAAAC GAGUAAGCU CGUCUAGUCG SEQ ID NO: NAME or Modification NUCLEOTIDE SEQUENCE 2226 3' env25 pistol ribozyme (with an added CUUCGG loop) UACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUG GGUAAAGCGCUUAUUUAUCGGAGAGAAAUCCGAUAAAUAAGAAGCAUCAAAGCGUG GUUAGGGCCACGUUAAAUAGUUGCUUAAGCCCUAAGCGUUGAUCUUCGGAUCAGGU GCAA 2227 3' Env-9 Twister UACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUG GGUAAAGCGCUUAUUUAUCGGAGAGAAAUCCGAUAAAUAAGAAGCAUCAAAGGGCA AUAAAGCGGUUACAAGCCCGCAAAAAUAGCAGAGUAAUGUCGCGAUAGCGCGGCAU UAAUGCAGCUUUAUUG 2228 =+ATTATCTCA TTACT25 UACUGGCGCUUUUAUCUCAUUACUAUUAUCUCAUUACUUUGAGAGCCAUCACCAGC GACUAUGUCGUAUGGGUAAAGCGCUUAUUUAUCGGAGAGAAAUCCGAUAAAUAAGA AGCAUCAAAG 2229 5Env-9 Twister GGCAAUAAAGCGGUUACAAGCCCGCAAAAAUAGCAGAGUAAUGUCGCGAUAGCGCG GCAUUAAUGCAGCUUUAUUGUACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUC ACCAGCGACUAUGUCGUAUGGGUAAAGCGCUUAUUUAUCGGAGAGAAAUCCGAUAA AUAAGAAGCAUCAAAG 2230 3' Twisted Sister 1 UACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUG GGUAAAGC GCUUAUUUAU C GGAGAGAAAU C C GAUAAAUAAGAAGCAU CAAAGAC C C GCAAGGCCGACGGCAUCCGCCGCCGCUGGUGCAAGUCCAGCCGCCCCUUCGGGGGC GGGCGCUCAUGGGUAAC 2231 no stem UACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUG GGUAAAG 2232 5111115 Minimal Hammerhead ribozyme GGGAGCCCCGCUGAUGAGGUCGGGGAGACCGAAAGGGACUUCGGUCCCUACGGGGC UCCCUACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCG UAUGGGUAAAGCGCUUAUUUAUCGGAGAGAAAUCCGAUAAAUAAGAAGCAUCAAAG 2233 5 Hammerhead ribozyme (Lior Nissim, Timothy Lu) guide scaffold scar CCAGUACUGAUGAGUCCGUGAGGACGAAACGAGUAAGCUCGUCUACUGGCGCUUUU AUCUCAUUACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUG UCGUAUGGGUAAAGCGCUUAUUUAUCGGAGAGAAAUCCGAUAAAUAAGAAGCAUCA AAG 2234 5Twisted Sister 1 ACCCGCAAGGCCGACGGCAUCCGCCGCCGCUGGUGCAAGUCCAGCCGCCCCUUCGG GGGCGGGCGCUCAUGGGUAACUACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAU CACCAGCGACUAUGUCGUAUGGGUAAAGCGCUUAUUUAUCGGAGAGAAAUCCGAUA AAUAAGAAGCAUCAAAG 2235 5'sTRSV WT viral Hammerhead ribozyme CCUGUCACCGGAUGUGCUUUCCGGUCUGAUGAGUCCGUGAGGACGAAACAGGUACU GGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUGGGUA AAG C GCUUAUUUAU C G GAGAGAAAU C C GAUAAAUAAGAAG CAU CAAAG 2236 148, =+G55, stacked onto 64 GUACUGGCGCCUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAG UGGGUAAAGCGCUUACGGACUUCGGUCCGUAAGAAGCAUCAAAG 2237 158, 103+148(+G55) -99, G65U GUACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAG UGGGUAAAGCUCCCUCUUCGGAGGGAGCAUCAAAG 2238 174, Uvsx Extended stem with [A99] G65U), C18G,AG55, [GT-1] ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGUG GGUAAAGCUCCCUCUUCGGAGGGAGCAUCAAAG 2239 175, extended stem truncation, T10C, [GT-1] ACUGGCGCCUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUGG GUAAAGCGCUUACGGACUUCGGUCCGUAAGAAGCAUCAAAG 2240 176, 174 with A1G substitution forT7 GCUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGUG GGUAAAGCUCCCUCUUCGGAGGGAGCAUCAAAG SEQ ID NO: NAME or Modification NUCLEOTIDE SEQUENCE transcription 2241 177, 174 with bubble (+G55) removed ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUGG GUAAAGCUCCCUCUUCGGAGGGAGCAUCAAAG 2242 181, stem 42 (truncated stem loop); T10C,C18G,[GT -1] (95+[GT-l]) ACUGGCGCCUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUGG GUAAAGCGCUUACGGACUUCGGUCCGUAAGAAGCAUCAAAG 2243 182, stem 42 (truncated stem loop); C18G,[GT-1] ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUGG GUAAAGCGCUUACGGACUUCGGUCCGUAAGAAGCAUCAAAG 2244 183, stem 42 (truncated stem loop); C18G,AG55,[GT-1] ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGUG GGUAAAGCGCUUACGGACUUCGGUCCGUAAGAAGCAUCAAAG 2245 184, stem 48 (uvsx, -99 g65t); C18G,AT55,[GT- 1] ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUUG GGUAAAGCUCCCUCUUCGGAGGGAGCAUCAAAG 2246 185, stem 42 (truncated stem loop); C18GAT55,[GT-1] ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUUG GGUAAAGCGCUUACGGACUUCGGUCCGUAAGAAGCAUCAAAG 2247 186, stem 42 (truncated stem loop); T10C AA17,[GT-1] ACUGGCGCCUUUAUCAUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUG GGUAAAGCGCUUACGGACUUCGGUCCGUAAGAAGCAUCAAAG 2248 187, stem 46 (uvsx); C18GAG55,[GT- 1] ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGUG GGUAAAGCGCCCUCUUCGGAGGGAAGCAUCAAAG 2249 188, stem 50 (ms2 U15C, -99, g65t); C18GAG55,[GT-1] ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGUG G GUAAAG CU GAGAU GAG GAU CAC C CAU GU GAG CAU CAAAG 2250 189, 174 + G8A;T15C;T35A ACU GGCACUUUUAC CU GAUUACUUU GAGAGC CAACAC CAGC GACUAU GU C GUAGU G GGUAAAGCUCCCUCUUCGGAGGGAGCAUCAAAG 2251 190, 174 + G8A ACUGGCACUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGUG GGUAAAGCUCCCUCUUCGGAGGGAGCAUCAAAG 2252 191, 174 + G8C ACUGGCCCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGUG GGUAAAGCUCCCUCUUCGGAGGGAGCAUCAAAG 2253 192, 174 + T15C ACUGGCGCUUUUACCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGUG GGUAAAGCUCCCUCUUCGGAGGGAGCAUCAAAG 2254 193, 174 + T35A ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAACACCAGCGACUAUGUCGUAGUG GGUAAAGCUCCCUCUUCGGAGGGAGCAUCAAAG 2255 195, 175 + C18G + G8A;T15C;T35A ACUGGCACCUUUACCUGAUUACUUUGAGAGCCAACACCAGCGACUAUGUCGUAUGG GUAAAGCGCUUACGGACUUCGGUCCGUAAGAAGCAUCAAAG SEQ ID NO: NAME or Modification NUCLEOTIDE SEQUENCE 2256 196, 175 + C18G + G8A ACUGGCACCUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUGG GUAAAGCGCUUACGGACUUCGGUCCGUAAGAAGCAUCAAAG 2257 197, 175 + C18G + G8C ACUGGCCCCUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUGG GUAAAGCGCUUACGGACUUCGGUCCGUAAGAAGCAUCAAAG 2258 198, 175 + C18G + T35A ACUGGCGCCUUUAUCUGAUUACUUUGAGAGCCAACACCAGCGACUAUGUCGUAUGG GUAAAGCGCUUACGGACUUCGGUCCGUAAGAAGCAUCAAAG 2259 199, 174 +A2G (test G transcription at start; ccGCT...) GCUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGUG GGUAAAGCUCCCUCUUCGGAGGGAGCAUCAAAG 2260 200, 174 + AG1 (ccGACT...) GACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGU GGGUAAAGCUCCCUCUUCGGAGGGAGCAUCAAAG 2261 201, 174 + T10C;AG28 ACUGGCGCCUUUAUCUGAUUACUUUGGAGAGCCAUCACCAGCGACUAUGUCGUAGU GGGUAAAGCUCCCUCUUCGGAGGGAGCAUCAAAG 2262 202, 174 + T10A;A28T ACUGGCGCAUUUAUCUGAUUACUUUGUGAGCCAUCACCAGCGACUAUGUCGUAGUG GGUAAAGCUCCCUCUUCGGAGGGAGCAUCAAAG 2263 203, 174 + T10C ACUGGCGCCUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGUG GGUAAAGCUCCCUCUUCGGAGGGAGCAUCAAAG 2264 204, 174 + AG28 ACUGGCGCUUUUAUCUGAUUACUUUGGAGAGCCAUCACCAGCGACUAUGUCGUAGU GGGUAAAGCUCCCUCUUCGGAGGGAGCAUCAAAG 2265 205, 174 + T10A ACUGGCGCAUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGUG GGUAAAGCUCCCUCUUCGGAGGGAGCAUCAAAG 2266 206, 174 + A28T ACUGGCGCUUUUAUCUGAUUACUUUGUGAGCCAUCACCAGCGACUAUGUCGUAGUG GGUAAAGCUCCCUCUUCGGAGGGAGCAUCAAAG 2267 207, 174 + AT15 ACUGGCGCUUUUAUUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGU GGGUAAAGCUCCCUCUUCGGAGGGAGCAUCAAAG 2268 208, 174 + [T4] ACGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGUGG GUAAAGCUCCCUCUUCGGAGGGAGCAUCAAAG 2269 209, 174 + C16A ACUGGCGCUUUUAUAUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGUG GGUAAAGCUCCCUCUUCGGAGGGAGCAUCAAAG 2270 210, 174 + AT17 ACUGGCGCUUUUAUCUUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGU GGGUAAAGCUCCCUCUUCGGAGGGAGCAUCAAAG 2271 211, 174 + T35G (compare with 174 + T35A above) ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAGCACCAGCGACUAUGUCGUAGUG GGUAAAGCUCCCUCUUCGGAGGGAGCAUCAAAG 2272 212, 174+U11G, A105G (A86G), U26C ACUGGCGCUGUUAUCUGAUUACUUCGAGAGCCAUCACCAGCGACUAUGUCGUAGUG GGUAAAGCUCCCUCUUCGGAGGGAGCAUCGAAG 2273 213, 174+U11C, A105G (A86G), U26C ACUGGCGCUCUUAUCUGAUUACUUCGAGAGCCAUCACCAGCGACUAUGUCGUAGUG GGUAAAGCUCCCUCUUCGGAGGGAGCAUCGAAG 2274 214, 174+U12G; A106G (A87G), U25C ACUGGCGCUUGUAUCUGAUUACUCUGAGAGCCAUCACCAGCGACUAUGUCGUAGUG GGUAAAGCUCCCUCUUCGGAGGGAGCAUCAGAG 2275 215, 174+U12C; A106G (A87G), U25C ACUGGCGCUUCUAUCUGAUUACUCUGAGAGCCAUCACCAGCGACUAUGUCGUAGUG GGUAAAGCUCCCUCUUCGGAGGGAGCAUCAGAG 2276 216, 174_tx_ll.G,87. G,22.C ACUGGCGCUUUGAUCUGAUUACCUUGAGAGCCAUCACCAGCGACUAUGUCGUAGUG GGUAAAGCUCCCUCUUCGGAGGGAGCAUCAAGG 2277 217, 174 tx ll.C,87. ACUGGCGCUUUCAUCUGAUUACCUUGAGAGCCAUCACCAGCGACUAUGUCGUAGUG GGUAAAGCUCCCUCUUCGGAGGGAGCAUCAAGG SEQ ID NO: NAME or Modification NUCLEOTIDE SEQUENCE G,22.C 2278 218, 174+U11G ACUGGCGCUGUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGUG GGUAAAGCUCCCUCUUCGGAGGGAGCAUCAAAG 2279 219, 174 +A105G (A86G) ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGUG GGUAAAGCUCCCUCUUCGGAGGGAGCAUCGAAG 2280 220, 174 +U26C ACUGGCGCUUUUAUCUGAUUACUUCGAGAGCCAUCACCAGCGACUAUGUCGUAGUG GGUAAAGCUCCCUCUUCGGAGGGAGCAUCAAAG
[00160] In some embodiments, the gNA variant comprises a tracrRNA stem loop comprising the sequence -UUU-N4-25-UUU- (SEQ ID NO: 240). For example, the gNA variant comprises a scaffold stem loop or a replacement thereof, flanked by two triplet U motifs that contribute to the triplex region. In some embodiments, the scaffold stem loop or replacement thereof comprises at least 4 nucleotides, at least 5 nucleotides, at least 6 nucleotides, at least 7 nucleotides, at least 7 nucleotides, at least 8 nucleotides, at least 9 nucleotides, at least 10 nucleotides, at least 11 nucleotides, at least 12 nucleotides, at least 13 nucleotides, at least 14 nucleotides, at least 15 nucleotides, at least 16 nucleotides, at least 17 nucleotides, at least 18 nucleotides, at least 19 nucleotides, at least 20 nucleotides, at least 21 nucleotides, at least 22 nucleotides, at least 23 nucleotides, at least 24 nucleotides, or at least 25 nucleotides.
[00161] In some embodiments, the gNA variant comprises a crRNA sequence with -AAAG- in a location 5’ to the spacer region. In some embodiments, the -AAAG- sequence is immediately 5’ to the spacer region.
[00162] In some embodiments, the at least one nucleotide modification to a reference gNA to produce a gNA variant comprises at least one nucleotide deletion in the CasX variant gNA relative to the reference gRNA. In some embodiments, a gNA variant comprises a deletion of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 consecutive or non-consecutive nucleotides relative to a reference gNA. In some embodiments, the at least one deletion comprises a deletion of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 or more consecutive nucleotides relative to a reference gNA. In some embodiments, the gNA variant comprises 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 or more nucleotide deletions relative to the reference gNA, and the deletions are not in consecutive nucleotides. In those embodiments where there are two or more non-consecutive deletions in the gNA variant relative to the reference gRNA, any length of deletions, and any combination of lengths of deletions, as described herein, are contemplated as within the scope of the disclosure. For example, in some embodiments, a gNA variant may comprise a first deletion of one nucleotide, and a second deletion of two nucleotides and the two deletions are not consecutive. In some embodiments, a gNA variant comprises at least two deletions in different regions of the reference gRNA. In some embodiments, a gNA variant comprises at least two deletions in the same region of the reference gRNA. For example, the regions may be the extended stem loop, scaffold stem loop, scaffold stem bubble, triplex loop, pseudoknot, triplex, or a 5’ end of the gNA variant. The deletion of any nucleotide in a reference gRNA is contemplated as within the scope of the disclosure.
[00163] In some embodiments, the at least one nucleotide modification of a reference gRNA to generate a gNA variant comprises at least one nucleotide insertion. In some embodiments, a gNA variant comprises an insertion of 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 consecutive or non-consecutive nucleotides relative to a reference gRNA. In some embodiments, the at least one nucleotide insertion comprises an insertion of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 or more consecutive nucleotides relative to a reference gRNA. In some embodiments, the gNA variant comprises 2 or more insertions relative to the reference gRNA, and the insertions are not consecutive. In those embodiments where there are two or more non-consecutive insertions in the gNA variant relative to the reference gRNA, any length of insertions, and any combination of lengths of insertions, as described herein, are contemplated as within the scope of the disclosure. For example, in some embodiments, a gNA variant may comprise a first insertion of one nucleotide, and a second insertion of two nucleotides and the two insertions are not consecutive. In some embodiments, a gNA variant comprises at least two insertions in different regions of the reference gRNA. In some embodiments, a gNA variant comprises at least two insertions in the same region of the reference gRNA. For example, the regions may be the extended stem loop, scaffold stem loop, scaffold stem bubble, triplex loop, pseudoknot, triplex, or a 5’ end of the gNA variant. Any insertion of A, G, C, U (or T, in the corresponding DNA) or combinations thereof at any location in the reference gRNA is contemplated as within the scope of the disclosure.
[00164] In some embodiments, the at least one nucleotide modification of a reference gRNA to genereate a gNA variant comprises at least one nucleic acid substitution. In some embodiments, a gNA variant comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 or more consecutive or non-consecutive substituted nucleotides relative to a reference gRNA. In some embodiments, a gNA variant comprises 1-4 nucleotide substitutions relative to a reference gRNA. In some embodiments, the at least one substitution comprises a substitution of 1, 2, 3, 4, 5,6,7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 or more consecutive nucleotides relative to a reference gRNA. In some embodiments, the gNA variant comprises 2 or more substitutions relative to the reference gRNA, and the substitutions are not consecutive. In those embodiments where there are two or more non-consecutive substitutions in the gNA variant relative to the reference gRNA, any length of substituted nucleotides, and any combination of lengths of substituted nucleotides, as described herein, are contemplated as within the scope of the disclosure. For example, in some embodiments, a gNA variant may comprise a first substitution of one nucleotide, and a second substitution of two nucleotides and the two substitutions are not consecutive. In some embodiments, a gNA variant comprises at least two substitutions in different regions of the reference gRNA. In some embodiments, a gNA variant comprises at least two substitutions in the same region of the reference gRNA. For example, the regions may be the triplex, the extended stem loop, scaffold stem loop, scaffold stem bubble, triplex loop, pseudoknot, triplex, or a 5’ end of the gNA variant. Any substitution of A, G, C, U (or T, in the corresponding DNA) or combinations thereof at any location in the reference gRNA is contemplated as within the scope of the disclosure.
[00165] Any of the substitutions, insertions and deletions described herein can be combined to generate a gNA variant of the disclosure. For example, a gNA variant can comprise at least one substitution and at least one deletion relative to a reference gRNA, at least one substitution and at least one insertion relative to a reference gRNA, at least one insertion and at least one deletion relative to a reference gRNA, or at least one substitution, one insertion and one deletion relative to a reference gRNA.
[00166] In some embodiments, the gNA variant comprises a scaffold region at least 20% identical, at least 30% identical, at least 40% identical, at least 50% identical, at least 60% identical, at least 65% identical, at least 70% identical, at least 75% identical, at least 80% identical, at least 85% identical, at least 90% identical, at least 91% identical, at least 92% identical, at least 93% identical, at least 94% identical, at least 95% identical, at least 96% identical, at least 97% identical, at least 98% identical, or at least 99% identical to any one of SEQ ID NOS: 4-16. In some embodiments, the gNA variant comprises a scaffold region at least 60% homologous (or identical) to any one of SEQ ID NOS: 4-16.
[00167] In some embodiments, the gNA variant comprises a tracr stem loop at least 60% identical, at least 65% identical, at least 70% identical, at least 75% identical, at least 80% identical, at least 85% identical, at least 90% identical, at least 91% identical, at least 92% identical, at least 93% identical, at least 94% identical, at least 95% identical, at least 96% identical, at least 97% identical, at least 98% identical, or at least 99% identical to SEQ ID NO: 14. In some embodiments, the gNA variant comprises a tracr stem loop at least 60% homologous (or identical) to SEQ ID NO: 14.
[00168] In some embodiments, the gNA variant comprises an extended stem loop at least 60% identical, at least 65% identical, at least 70% identical, at least 75% identical, at least 80% identical, at least 85% identical, at least 90% identical, at least 91% identical, at least 92% identical, at least 93% identical, at least 94% identical, at least 95% identical, at least 96% identical, at least 97% identical, at least 98% identical, or at least 99% identical to SEQ ID NO: 15. In some embodiments, the gNA variant comprises an extended stem loop at least 60% homologous (or identical) to SEQ ID NO: 15.
[00169] In some embodiments, a gNA variant comprises a sequence of any one of SEQ ID NOs: 412-3295. In some embodiments, a gNA variant comprises a sequence of any one of SEQ ID NOS: 2236, 2237, 2238, 2241, 2244, 2248, 2249, or 2259-2280. In some embodiments, a gNA variant comprises a sequence of any one of SEQ ID NOS: 2236, 2237, 2238, 2241, 2244, 2248, 2249, or 2259-2280.
[00170] In some embodiments, the gNA variant comprises an exogenous extended stem loop, with such differences from a reference gNA described as follows. In some embodiments, an exogenous extended stem loop has little or no identity to the reference stem loop regions disclosed herein (e.g., SEQ ID NO: 15). In some embodiments, an exogenous stem loop is at least 10 bp, at least 20 bp, at least 30 bp, at least 40 bp, at least 50 bp, at least 60 bp, at least 70 bp, at least 80 bp, at least 90 bp, at least 100 bp, at least 200 bp, at least 300 bp, at least 400 bp, at least 500 bp, at least 600 bp, at least 700 bp, at least 800 bp, at least 900 bp, at least 1,000 bp, at least 2,000 bp, at least 3,000 bp, at least 4,000 bp, at least 5,000 bp, at least 6,000 bp, at least 7,000 bp, at least 8,000 bp, at least 9,000 bp, at least 10,000 bp, at least 12,000 bp, at least 15,000 bp or at least 20,000 bp. In some embodiments, the gNA variant comprises an extended stem loop region comprising at least 10, at least 100, at least 500, at least 1000, or at least 10,000 nucleotides. In some embodiments, the heterologous stem loop increases the stability of the gNA. In some embodiments, the heterologous RNA stem loop is capable of binding a protein, an RNA structure, a DNA sequence, or a small molecule. In some embodiments, an exogenous stem loop region comprises an RNA stem loop or hairpin, for example a thermostable RNA such as MS2 (ACAUGAGGAUUACCCAUGU; SEQ ID NO: 4278), Qp (UGCAUGUCUAAGACAGCA; SEQ ID NO: 4279), UI hairpin II (AAUCCAUUGCACUCCGGAUU; SEQ ID NO:4280), Uvsx (CCUCUUCGGAGG; SEQ ID NO: 4281), PP7 (AGGAGUUUCUAUGGAAACCCU; SEQ ID NO: 4282), Phage replication loop (AGGUGGGACGACCUCUCGGUCGUCCUAUCU; SEQ ID NO: 4283), Kissing loop a (UGCUCGCUCCGUUCGAGCA; SEQ ID NO: 4284), Kissing loop bl (UGCUCGACGCGUCCUCGAGCA; SEQ ID NO: 4285), Kissing loop_b2 (UGCUCGUUUGCGGCUACGAGCA; SEQ ID NO: 4286), Gquadriplex M3q (AGGGAGGGAGGGAGAGG; SEQ ID NO: 4287), G quadriplex telomere basket (GGUUAGGGUUAGGGUUAGG; SEQ ID NO: 4288), Sarcin-ricin loop (CUGCUCAGUACGAGAGGAACCGCAG; SEQ ID NO: 4289) or Pseudoknots (UACACUGGGAUCGCUGAAUUAGAGAUCGGCGUCCUUUCAUUCUAUAUACUUUGG AGUUUUAAAAUGUCUCUAAGUACA; SEQ ID NO: 4290). In some embodiments, an exogenous stem loop comprises an RNA scaffold. As used herein, an “RNA scaffold” refers to a multi-dimensional RNA structure capable of interacting with and organizing or localizing one or more proteins. In some embodiments, the RNA scaffold is synthetic or non-naturally occurring. In some embodiments, an exogenous stem loop comprises a long non-coding RNA (IncRNA). As used herein, a IncRNA refers to a non-coding RNA that is longer than approximately 200 bp in length. In some embodiments, the 5’ and 3’ ends of the exogenous stem loop are base paired, i.e., interact to form a region of duplex RNA. In some embodiments, the 5’ and 3’ ends of the exogenous stem loop are base paired, and one or more regions between the 5’ and 3’ ends of the exogenous stem loop are not base paired. In some embodiments, the at least one nucleotide modification comprises: (a) substitution of 1 to 15 consecutive or non-consecutive nucleotides in the gNA variant in one or more regions; (b) a deletion of 1 to 10 consecutive or non-consecutive nucleotides in the gNA variant in one or more regions; (c) an insertion of 1 to 10 consecutive or non-consecutive nucleotides in the gNA variant in one or more regions; (d) a substitution of the scaffold stem loop or the extended stem loop with an RNA stem loop sequence from a heterologous RNA source with proximal 5' and 3' ends; or any combination of (a)-(d).
[00171] In some embodiments, a gNA variant comprises a sequence or subsequence of any one of SEQ ID NOs: 412-3295 and an a sequence of an exogenous stem loop. In some embodiments, a gNA variant comprises a sequence or subsequence of any one of SEQ ID NOS: 2236, 2237, 2238, 2241, 2244, 2248, 2249, or 2259-2280 and a sequence of an exogenous stem loop. In some embodiments, a gNA variant comprises a sequence or subsequence of any one of SEQ ID NOS: 2236, 2237, 2238, 2241, 2244, 2248, 2249, or 2259-2280 and a sequence of an exogenous stem loop.
[00172] In some embodiments, the gNA variant comprises a scaffold stem loop having at least 60% identity to SEQ ID NO: 14. In some embodiments, the gNA variant comprises a scaffold stem loop having at least 60% identity, at least 70% identity, at least 80% identity, at least 90% identity, at least 95% identity, at least 98% identity or at least 99% identity to SEQ ID NO: 14. In some embodiments, the gNA variant comprises a scaffold stem loop comprising SEQ ID NO: 14.
[00173] In some embodiments, the gNA variant comprises a scaffold stem loop sequence of CCAGCGACUAUGUCGUAGUGG (SEQ ID NO: 245). In some embodiments, the gNA variant comprises a scaffold stem loop sequence of CCAGCGACUAUGUCGUAGUGG (SEQ ID NO: 245) with at least 1, 2, 3, 4, or 5 mismatches thereto.
[00174] In some embodiments, the gNA variant comprises an extended stem loop region comprising less than 32 nucleotides, less than 31 nucleotides, less than 30 nucleotides, less than 29 nucleotides, less than 28 nucleotides, less than 27 nucleotides, less than 26 nucleotides, less than 25 nucleotides, less than 24 nucleotides, less than 23 nucleotides, less than 22 nucleotides, less than 21 nucleotides, or less than 20 nucleotides. In some embodiments, the gNA variant comprises an extended stem loop region comprising less than 32 nucleotides. In some embodiments, the gNA variant further comprises a thermostable stem loop.
[00175] In some embodiments, a sgRNA variant comprises a sequence of SEQ ID NO: 2104, 2106, SEQ ID NO: 2163, SEQ ID NO: 2107, SEQ ID NO: 2164, SEQ ID NO: 2165, SEQ ID NO: 2166, SEQ ID NO: 2103, SEQ ID NO: 2167, SEQ ID NO: 2105, SEQ ID NO: 2108, SEQ ID NO: 2112, SEQ ID NO: 2160, SEQ ID NO: 2170, SEQ ID NO: 2114, SEQ ID NO: 2171, SEQ ID NO: 2112, SEQ ID NO: 2173, SEQ ID NO: 2102, SEQ ID NO: 2174, SEQ ID NO: 2175, SEQ ID NO: 2109, SEQ ID NO: 2176, SEQ ID NO: 2238, SEQ ID NO: 2239, SEQ ID NO: 2240, or SEQ ID NO: 2241.
[00176] In some embodiments, the gNA variant comprises one or more additional changes to a sequence of any one of SEQ ID NOs: 2201-2280. In some embodiments, the gNA variant comprises a sequence of any one of SEQ ID NOS: 2236, 2237, 2238, 2241, 2244, 2248, 2249, or 2259-2280, or having at least about 80%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% identity thereto. In some embodiments, the gNA variant comprises one or more additional changes to a sequence of any one of SEQ ID NOs: 2201-2280. In some embodiments, the gNA variant comprises the sequence of any one of SEQ ID NOS: 2236, 2237, 2238, 2241, 2244, 2248, 2249, or 2259-2280.
[00177] In some embodiments, a sgRNA variant comprises one or more additional changes to a sequence of SEQ ID NO: 2104, SEQ ID NO: 2163, SEQ ID NO: 2107, SEQ ID NO: 2164, SEQ ID NO: 2165, SEQ ID NO: 2166, SEQ ID NO: 2103, SEQ ID NO: 2167, SEQ ID NO: 2105, SEQ ID NO: 2108, SEQ ID NO: 2112, SEQ ID NO: 2160, SEQ ID NO: 2170, SEQ ID NO: 2114, SEQ ID NO: 2171, SEQ ID NO: 2112, SEQ ID NO: 2173, SEQ ID NO: 2102, SEQ ID NO: 2174, SEQ ID NO: 2175, SEQ ID NO: 2109, SEQ ID NO: 2176, SEQ ID NO: 2238, SEQ ID NO: 2239, SEQ ID NO: 2240, or SEQ ID NO: 2241.
[00178] In some embodiments of the gNA variants of the disclosure, the gNA variant comprises at least one modification, wherein the at least one modification compared to the reference guide scaffold of SEQ ID NO: 5 is selected from one or more of: (a) a C18G substitution in the triplex loop; (b) a G55 insertion in the stem bubble; (c) a UI deletion; (d) a modification of the extended stem loop wherein (i) a 6 nt loop and 13 loop-proximal base pairs are replaced by a Uvsx hairpin; and (ii) a deletion of A99 and a substitution of G65U that results in a loop-distal base that is fully base-paired. In such embodiments, the gNA variant comprises the sequence of any one of SEQ ID NOS: 2236, 2237, 2238, 2241, 2244, 2248, 2249, or 22592280.
[00179] In some embodiments, the scaffold of the gNA variant comprises the sequence of any one of SEQ ID NOS: 2201-2280 of Table 2. In some embodiments, the scaffold of the gNA consists or consists essentially of the sequence of any one of SEQ ID NOS: 2201-2280. In some embodiments, the scaffold of the gNA variant sequence is at least about 60% identical, at least about 65% identical, at least about 70% identical, at least about 75% identical, at least about 80% identical, at least about 85% identical, at least about 90% identical, at least about 91% identical, at least about 92% identical, at least about 93% identical, at least about 94% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical or at least about 99% identical to any one of SEQ ID NOS: 2201 to 2280.
[00180] In some embodiments, the gNA variant further comprises a spacer (or targeting sequence) region, described more fully, supra, which comprises at least 14 to about 35 nucleotides wherein the spacer is designed with a sequence that is complementary to a target DNA. In some embodiments, the gNA variant comprises a targeting sequence of at least 10 to 30 nucleotides complementary to a target DNA. In some embodiments, the targeting sequence has 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34 or 35 nucleotides. In some embodiments, the gNA variant comprises a targeting sequence having 20 nucleotides. In some embodiments, the targeting sequence has 25 nucleotides. In some embodiments, the targeting sequence has 24 nucleotides. In some embodiments, the targeting sequence has 23 nucleotides. In some embodiments, the targeting sequence has 22 nucleotides. In some embodiments, the targeting sequence has 21 nucleotides. In some embodiments, the targeting sequence has 20 nucleotides. In some embodiments, the targeting sequence has 19 nucleotides. In some embodiments, the targeting sequence has 18 nucleotides. In some embodiments, the targeting sequence has 17 nucleotides. In some embodiments, the targeting sequence has 16 nucleotides. In some embodiments, the targeting sequence has 15 nucleotides. In some embodiments, the targeting sequence has 14 nucleotides.
[00181] In some embodiments, the scaffold of the gNA variant is a variant comprising one or more additional changes to a sequence of a reference gRNA that comprises SEQ ID NO: 4 or SEQ ID NO: 5. In those embodiments where the scaffold of the reference gRNA is derived from SEQ ID NO: 4 or SEQ ID NO: 5, the one or more improved or added characteristics of the gNA variant are improved compared to the same characteristic in SEQ ID NO: 4 or SEQ ID NO: 5.
[00182] In some embodiments, the scaffold of the gNA variant is part of an RNP with a reference CasX protein comprising SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 3. In other embodiments, the scaffold of the gNA variant is part of an RNP with a CasX variant protein comprising any one of the sequences of Tables 3, 8, 9, 10 and 12, or a sequence having at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto. In the foregoing embodiments, the gNA further comprises a spacer sequence. h. Chemically Modified gNAs
[00183] In some embodiments, the disclosure provides chemically-modified gNAs. In some embodiments, the present disclosure provides a chemically-modified gNA that has guide NA functionality and has reduced susceptibility to cleavage by a nuclease. A gNA that comprises any nucleotide other than the four canonical ribonucleotides A, C, G, and U, or a deoxynucleotide, is a chemically modified gNA. In some cases, a chemically-modified gNA comprises any backbone or intemucleotide linkage other than a natural phosphodiester intemucleotide linkage. In certain embodiments, the retained functionality includes the ability of the modified gNA to bind to a CasX of any of the embodiments described herein. In certain embodiments, the retained functionality includes the ability of the modified gNA to bind to a target nucleic acid sequence. In certain embodiments, the retained functionality includes targeting a CasX protein or the ability of a pre-complexed RNP to bind to a target nucleic acid sequence. In certain embodiments, the retained functionality includes the ability to nick a target polynucleotide by a CasX-gNA. In certain embodiments, the retained functionality includes the ability to cleave a target nucleic acid sequence by a CasX-gNA. In certain embodiments, the retained functionality is any other known function of a gNA in a recombinant system with a CasX chimera protein of the embodiments of the disclosure.
[00184] In some embodiments, the disclosure provides a chemically-modified gNA in which a nucleotide sugar modification is incorporated into the gNA selected from the group consisting of 2'-0—Ci.4alkyl such as 2'-O-methyl (2'-0Me), 2'-deoxy (2'-H), 2'-0—C,.3alkyl-0—C,.3alkyl such as 2'-methoxyethyl (“2'-M0E”), 2'-fluoro (“2'-F”), 2'-amino (“2'-NH2”), 2'-arabinosyl (“2'-arabino”) nucleotide, 2'-F-arabinosyl (“2'-F-arabino”) nucleotide, 2'-locked nucleic acid (“LNA”) nucleotide, 2'-unlocked nucleic acid (“ULNA”) nucleotide, a sugar in L form (“L-sugar”), and 4'-thioribosyl nucleotide. In other embodiments, an intemucleotide linkage modification incorporated into the guide RNA is selected from the group consisting of: phosphorothioate “P(S)” (P(S)), phosphonocarboxylate (P(CH2)„COOR) such as phosphonoacetate “PACE” (P(CH2COO)), thiophosphonocarboxylate ((S)P(CH2)„COOR) such as thiophosphonoacetate “thioPACE” ((S)P(CH2)nCOO )), alkylphosphonate (P(Ci.3alkyl) such as methylphosphonate —P(CH3), boranophosphonate (P(BH3)), and phosphorodithioate (P(S)2).
[00185] In certain embodiments, the disclosure provides a chemically-modified gNA in which a nucleobase (“base”) modification is incorporated into the gNA selected from the group consisting of: 2-thiouracil (“2-thioU”), 2-thiocytosine (“2-thioC”), 4-thiouracil (“4-thioU’), 6-thioguanine (“6-thioG”), 2-aminoadenine (“2-aminoA”), 2-aminopurine, pseudouracil, hypoxanthine, 7-deazaguanine, 7-deaza-8-azaguanine, 7-deazaadenine, 7-deaza-8-azaadenine, 5-methylcytosine (“5-methylC”), 5-methyluracil (“5-methylU”), 5-hydroxymethylcytosine, 5-hydroxymethyluracil, 5,6-dehydrouracil, 5-propynylcytosine, 5-propynyluracil, 5-ethynylcytosine, 5-ethynyluracil, 5-allyluracil (“5-allylU”), 5-allylcytosine (“5-allylC”), 5-aminoallyluracil (“5-aminoallylU”), 5-aminoallyl-cytosine (“5-aminoallylC”), an abasic nucleotide, Z base, P base, Unstructured Nucleic Acid (“UNA”), isoguanine (“isoG”), isocytosine (“isoC”), 5-methyl-2-pyrimidine, x(A,G,C,T) and y(A,G,C,T).
[00186] In other embodiments, the disclosure provides a chemically-modified gNA in which one or more isotopic modifications are introduced on the nucleotide sugar, the nucleobase, the phosphodiester linkage and / or the nucleotide phosphates, including nucleotides comprising one or more 15N, 13C, 14C, deuterium, 3H, 32P, 125I, 13’I atoms or other atoms or elements used as tracers.
[00187] In some embodiments, an “end” modification incorporated into the gNA is selected from the group consisting of: PEG (polyethyleneglycol), hydrocarbon linkers (including: heteroatom (O,S,N)-substituted hydrocarbon spacers; halo-substituted hydrocarbon spacers; keto-, carboxyl-, amido-, thionyl-, carbamoyl-, thionocarbamaoyl-containing hydrocarbon spacers), spermine linkers, dyes including fluorescent dyes (for example fluoresceins, rhodamines, cyanines) attached to linkers such as, for example 6-fluorescein-hexyl, quenchers (for example dabcyl, BHQ) and other labels (for example biotin, digoxigenin, acridine, streptavidin, avidin, peptides and / or proteins). In some embodiments, an “end” modification comprises a conjugation (or ligation) of the gNA to another molecule comprising an oligonucleotide of deoxynucleotides and / or ribonucleotides, a peptide, a protein, a sugar, an oligosaccharide, a steroid, a lipid, a folic acid, a vitamin and / or other molecule. In certain embodiments, the disclosure provides a chemically-modified gNA in which an “end” modification (described above) is located internally in the gNA sequence via a linker such as, for example, a 2-(4-butylamidofluorescein)propane-l,3-diol bis(phosphodiester) linker, which is incorporated as a phosphodiester linkage and can be incorporated anywhere between two nucleotides in the gNA.
[00188] In some embodiments, the disclosure provides a chemically-modified gNA having an end modification comprising a terminal functional group such as an amine, a thiol (or sulfhydryl), a hydroxyl, a carboxyl, carbonyl, thionyl, thiocarbonyl, a carbamoyl, a thiocarbamoyl, a phoshoryl, an alkene, an alkyne, an halogen or a functional group-terminated linker that can be subsequently conjugated to a desired moiety selected from the group consisting of a fluorescent dye, a non-fluorescent label, a tag (for 14C, example biotin, avidin, streptavidin, or moiety containing an isotopic label such as 15N, 13C, deuterium, 3H, 32P, 125I and the like), an oligonucleotide (comprising deoxynucleotides and / or ribonucleotides, including an aptamer), an amino acid, a peptide, a protein, a sugar, an oligosaccharide, a steroid, a lipid, a folic acid, and a vitamin. The conjugation employs standard chemistry well-known in the art, including but not limited to coupling via N-hydroxysuccinimide, isothiocyanate, DCC (or DCI), and / or any other standard method as described in “Bioconjugate Techniques” by Greg T. Hermanson, Publisher Eslsevier Science, 3'ded. (2013), the contents of which are incorporated herein by reference in its entirety. i. Complex Formation with CasX Protein
[00189] In some embodiments, a gNA variant has an improved ability to form a complex with a CasX protein (such as a reference CasX or a CasX variant protein) when compared to a reference gRNA. In some embodiments, a gNA variant has an improved affinity for a CasX protein (such as a reference or variant protein) when compared to a reference gRNA, thereby improving its ability to form a ribonucleoprotein (RNP) complex with the CasX protein, as described in the Examples. Improving ribonucleoprotein complex formation may, in some embodiments, improve the efficiency with which functional RNPs are assembled. In some embodiments, greater than 90%, greater than 93%, greater than 95%, greater than 96%, greater than 97%, greater than 98% or greater than 99% of RNPs comprising a gNA variant and a spacer are competent for gene editing of a target nucleic acid.
[00190] Exemplary nucleotide changes that can improve the ability of gNA variants to form a complex with CasX protein may, in some embodiments, include replacing the scaffold stem with a thermostable stem loop. Without wishing to be bound by any theory, replacing the scaffold stem with a thermostable stem loop could increase the overall binding stability of the gNA variant with the CasX protein. Alternatively, or in addition, removing a large section of the stem loop could change the gNA variant folding kinetics and make a functional folded gNA easier and quicker to structurally-assemble, for example by lessening the degree to which the gNA variant can get “tangled” in itself. In some embodiments, choice of scaffold stem loop sequence could change with different spacers that are utilized for the gNA. In some embodiments, scaffold sequence can be tailored to the spacer and therefore the target sequence. Biochemical assays can be used to evaluate the binding affinity of CasX protein for the gNA variant to form the RNP, including the assays of the Examples. For example, a person of ordinary skill can measure changes in the amount of a fluorescently tagged gNA that is bound to an immobilized CasX protein, as a response to increasing concentrations of an additional unlabeled “cold competitor” gNA. Alternatively, or in addition, fluorescence signal can be monitored to or seeing how it changes as different amounts of fluorescently labeled gNA are flowed over immobilized CasX protein. Alternatively, the ability to form an RNP can be assessed using in vitro cleavage assays against a defined target nucleic acid sequence. j. gNA Stability
[00191] In some embodiments, a gNA variant has improved stability when compared to a reference gRNA. Increased stability and efficient folding may, in some embodiments, increase the extent to which a gNA variant persists inside a target cell, which may thereby increase the chance of forming a functional RNP capable of carrying out CasX functions such as gene editing. Increased stability of gNA variants may also, in some embodiments, allow for a similar outcome with a lower amount of gNA delivered to a cell, which may in turn reduce the chance of off-target effects during gene editing.
[00192] In other embodiments, the disclosure provides gNA in which the scaffold stem loop and / or the extended stem loop is replaced with a hairpin loop or a thermostable RNA stem loop in which the resulting gNA has increased stability and, depending on the choice of loop, can interact with certain cellular proteins or RNA. In some embodiments, the replacement RNA loop is selected from MS2, QP, UI hairpin II, Uvsx, PP7, Phage replication loop, Kissing loop a, Kissing loop bl, Kissing loop_b2, G quadriplex M3q, G quadriplex telomere basket, Sarcin-ricin loop and Pseudoknots. Sequences of gNA variants including such components are provided in Table 2.
[00193] Guide NA stability can be assessed in a variety of ways, including for example in vitro by assembling the guide, incubating for varying periods of time in a solution that mimics the intracellular environment, and then measuring functional activity via the in vitro cleavage assays described herein. Alternatively, or in addition, gNAs can be harvested from cells at varying time points after initial transfection / transduction of the gNA to determine how long gNA variants persist relative to reference gRNAs. k. Solubility
[00194] In some embodiments, a gNA variant has improved solubility when compared to a reference gRNA. In some embodiments, a gNA variant has improved solubility of the CasX protein:gNARNP when compared to a reference gRNA. In some embodiments, solubility of the CasX protein:gNA RNP is improved by the addition of a ribozyme sequence to a 5’ or 3’ end of the gNA variant, for example the 5’ or 3’ of a reference sgRNA. Some ribozymes, such as the Ml ribozyme, can increase solubility of proteins through RNA mediated protein folding.
[00195] Increased solubility of CasX RNPs comprising a gNA variant as described herein can be evaluated through a variety of means known to one of skill in the art, such as by taking densitometry readings on a gel of the soluble fraction of lysed E. coli in which the CasX and gNA variants are expressed. 1. Resistance to Nuclease Activity
[00196] In some embodiments, a gNA variant has improved resistance to nuclease activity compared to a reference gRNA. Without wishing to be bound by any theory, increased resistance to nucleases, such as nucleases found in cells, may for example increase the persistence of a variant gNA in an intracellular environment, thereby improving gene editing.
[00197] Many nucleases are processive, and degrade RNA in a 3’ to 5’ fashion. Therefore, in some embodiments the addition of a nuclease resistant secondary structure to one or both termini of the gNA, or nucleotide changes that change the secondary structure of a sgNA, can produce gNA variants with increased resistance to nuclease activity. Resistance to nuclease activity may be evaluated through a variety of methods known to one of skill in the art. For example, in vitro methods of measuring resistance to nuclease activity may include for example contacting reference gNA and variants with one or more exemplary RNA nucleases and measuring degradation. Alternatively, or in addition, measuring persistence of a gNA variant in a cellular environment using the methods described herein can indicate the degree to which the gNA variant is nuclease resistant. m. Binding Affinity to a Target DNA
[00198] In some embodiments, a gNA variant has improved affinity for the target DNA relative to a reference gRNA. In certain embodiments, a ribonucleoprotein complex comprising a gNA variant has improved affinity for the target DNA, relative to the affinity of an RNP comprising a reference gRNA. In some embodiments, the improved affinity of the RNP for the target DNA comprises improved affinity for the target sequence, improved affinity for the PAM sequence, improved ability of the RNP to search DNA for the target sequence, or any combinations thereof. In some embodiments, the improved affinity for the target DNA is the result of increased overall DNA binding affinity.
[00199] Without wishing to be bound by theory, it is possible that nucleotide changes in the gNA variant that affect the function of the OBD in the CasX protein may increase the affinity of CasX variant protein binding to the protospacer adjacent motif (PAM), as well as the ability to bind or utilize an increased spectrum of PAM sequences other than the canonical TTC PAM recognized by the reference CasX protein of SEQ ID NO: 2, including PAM sequences selected from the group consisting of TTC, ATC, GTC, and CTC, thereby increasing the affinity and diversity of the CasX variant protein for target DNA sequences, thereby increasing the target nucleic acid sequences that can be edited and / or bound, compared to a reference CasX. As described more fully, below, increasing the sequences of the target nucleic acid that can be edited, compared to a reference CasX, refers to both the PAM and the protospacer sequence and their directionality according to the orientation of the non-target strand. This does not imply that the PAM sequence of the non-target strand, rather than the target strand, is determinative of cleavage or mechanistically involved in target recognition. For example, when reference is to a TTC PAM, it may in fact be the complementary GAA sequence that is required for target cleavage, or it may be some combination of nucleotides from both strands. In the case of the CasX proteins disclosed herein, the PAM is located 5’ of the protospacer with at least a single nucleotide separating the PAM from the first nucleotide of the protospacer. Alternatively, or in addition, changes in the gNA that affect function of the helical I and / or helical II domains that increase the affinity of the CasX variant protein for the target DNA strand can increase the affinity of the CasX RNP comprising the variant gNA for target DNA. n. Adding or Changing gNA Function
[00200] In some embodiments, gNA variants can comprise larger structural changes that change the topology of the gNA variant with respect to the reference gRNA, thereby allowing for different gNA functionality. For example, in some embodiments a gNA variant has swapped an endogenous stem loop of the reference gRNA scaffold with a previously identified stable RNA structure or a stem loop that can interact with a protein or RNA binding partner to recruit additional moieties to the CasX or to recruit CasX to a specific location, such as the inside of a viral capsid, that has the binding partner to the said RNA structure. In other scenarios the RNAs may be recruited to each other, as in Kissing loops, such that two CasX proteins can be colocalized for more effective gene editing at the target DNA sequence. Such RNA structures may include MS2, QP, UI hairpin II, Uvsx, PP7, Phage replication loop, Kissing loop a, Kissing loop bl, Kissing loop_b2, G quadriplex M3q, G quadriplex telomere basket, Sarcin-ricin loop, or a Pseudoknot.
[00201] In some embodiments, a gNA variant comprises a terminal fusion partner. The term gNA variant is inclusive of variants that include exogenous sequences such as terminal fusions, or internal insertions. Exemplary terminal fusions may include fusion of the gRNA to a self cleaving ribozyme or protein binding motif. As used herein, a “ribozyme” refers to an RNA or segment thereof with one or more catalytic activities similar to a protein enzyme. Exemplary ribozyme catalytic activities may include, for example, cleavage and / or ligation of RNA, cleavage and / or ligation of DNA, or peptide bond formation. In some embodiments, such fusions could either improve scaffold folding or recruit DNA repair machinery. For example, a gRNA may in some embodiments be fused to a hepatitis delta virus (HDV) antigenomic ribozyme, HDV genomic ribozyme, hatchet ribozyme (from metagenomic data), env25 pistol ribozyme (representative from Aliistipes putredinis), HH15 Minimal Hammerhead ribozyme, tobacco ringspot virus (TRSV) ribozyme, WT viral Hammerhead ribozyme (and rational variants), or Twisted Sister 1 or RBMX recruiting motif. Hammerhead ribozymes are RNA motifs that catalyze reversible cleavage and ligation reactions at a specific site within an RNA molecule. Hammerhead ribozymes include type I, type II and type III hammerhead ribozymes. The HDV, pistol, and hatchet ribozymes have self-cleaving activities. gNA variants comprising one or more ribozymes may allow for expanded gNA function as compared to a gRNA reference. For example, gNAs comprising self-cleaving ribozymes can, in some embodiments, be transcribed and processed into mature gNAs as part of polycistronic transcripts. Such fusions may occur at either the 5’ or the 3’ end of the gNA. In some embodiments, a gNA variant comprises a fusion at both the 5’ and the 3’ end, wherein each fusion is independently as described herein. In some embodiments, a gNA variant comprises a phage replication loop or a tetraloop. In some embodiments, a gNA comprises a hairpin loop that is capable of binding a protein. For example, in some embodiments the hairpin loop is an MS2, QP, UI hairpin II, Uvsx, or PP7 hairpin loop.
[00202] In some embodiments, a gNA variant comprises one or more RNA aptamers. As used herein, an “RNA aptamer” refers to an RNA molecule that binds a target with high affinity and high specificity.
[00203] In some embodiments, a gNA variant comprises one or more riboswitches. As used herein, a “riboswitch” refers to an RNA molecule that changes state upon binding a small molecule.
[00204] In some embodiments, the gNA variant further comprises one or more protein binding motifs. Adding protein binding motifs to a reference gRNA or gNA variant of the disclosure may, in some embodiments, allow a CasX RNP to associate with additional proteins, which can for example add the functionality of those proteins to the CasX RNP. IV. CasX Proteins for Modifying a Target Nucleic Acid
[00205] The term “CasX protein”, as used herein, refers to a family of proteins, and encompasses all naturally occurring CasX proteins, proteins that share at least 50% identity to naturally occurring CasX proteins, as well as CasX variants possessing one or more improved characteristics relative to a naturally-occurring reference CasX protein. Exemplary improved characteristics of the CasX variant embodiments include, but are not limited to improved folding of the variant, improved binding affinity to the gNA, improved binding affinity to the target nucleic acid, improved ability to utilize a greater spectrum of PAM sequences in the editing and / or binding of target DNA, improved unwinding of the target DNA, increased editing activity, improved editing efficiency, improved editing specificity, increased percentage of a eukaryotic genome that can be efficiently edited, increased activity of the nuclease, increased target strand loading for double strand cleavage, decreased target strand loading for single strand nicking, decreased off-target cleavage, improved binding of the non-target strand of DNA, improved protein stability, improved protein:gNA (RNP) complex stability, improved protein solubility, improved protein:gNA (RNP) complex solubility, improved protein yield, improved protein expression, and improved fusion characteristics, as described more fully, below. In the foregoing embodiments, the one or more of the improved characteristics of the CasX variant is at least about 1.1 to about 100,000-fold improved relative to the reference CasX protein of SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 3 when assayed in a comparable fashion. In other embodiments, the improvement is at least about 1.1-fold, at least about 2-fold, at least about 5fold, at least about 10-fold, at least about 50-fold, at least about 100-fold, at least about 500-fold, at least about 1000-fold, at least about 5000-fold, at least about 10,000-fold, or at least about 100,000-fold compared to the reference CasX protein of SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 3 when assayed in a comparable fashion.
[00206] The term CasX variant is inclusive of variants that are fusion proteins, i.e. the CasX is “fused to” a heterologous sequence. This includes CasX variants comprising CasX variant sequences and N-terminal, C-terminal, or internal fusions of the CasX to a heterologous protein or domain thereof.
[00207] CasX proteins of the disclosure comprise at least one of the following domains: a nontarget strand binding (NTSB) domain, a target strand loading (TSL) domain, a helical I domain, a helical II domain, an oligonucleotide binding domain (OBD), and a RuvC DNA cleavage domain (the last of which may be modified or deleted in a catalytically dead CasX variant), described more fully, below. Additionally, the CasX variant proteins of the disclosure have an enhanced ability to efficiently edit and / or bind target DNA utilizing PAM sequences selected from TTC, ATC, GTC, or CTC, compared to wild-type reference CasX proteins. In the foregoing, the PAM sequence is located at least 1 nucleotide 5’ to the non-target strand of the protospacer having identity with the targeting sequence of the gNA in a assay system compared to the editing efficiency and / or binding of an RNP comprising a reference CasX protein in a comparable assay system.
[00208] In some cases, the CasX protein is a naturally-occurring protein (e.g., naturally occurs in and is isolated from prokaryotic cells). In other embodiments, the CasX protein is not a naturally-occurring protein (e.g., the CasX protein is a CasX variant protein, a chimeric protein, and the like). A naturally-occurring CasX protein (referred to herein as a “reference CasX protein”) functions as an endonuclease that catalyzes a double strand break at a specific sequence in a targeted double-stranded DNA (dsDNA). The sequence specificity is provided by the targeting sequence of the associated gNA to which it is complexed, which hybridizes to a target sequence within the target nucleic acid.
[00209] In some embodiments, a CasX protein can bind and / or modify (e.g., cleave, nick, methylate, demethylate, etc.) a target nucleic acid and / or a polypeptide associated with target nucleic acid (e.g., methylation or acetylation of a histone tail). In some embodiments, the CasX protein is catalytically dead (dCasX) but retains the ability to bind a target nucleic acid. An exemplary catalytically dead CasX protein comprises one or more mutations in the active site of the RuvC domain of the CasX protein. In some embodiments, a catalytically dead CasX protein comprises substitutions at residues 672, 769 and / or 935 of SEQ ID NO: 1. In one embodiment, a catalytically dead CasX protein comprises substitutions of D672A, E769A and / or D935A in a reference CasX protein of SEQ ID NO: 1. In other embodiments, a catalytically dead CasX protein comprises substitutions at amino acids 659, 756 and / or 922 in a reference CasX protein of SEQ ID NO: 2. In some embodiments, a catalytically dead CasX protein comprises D659A, E756A and / or D922A substitutions in a reference CasX protein of SEQ ID NO: 2. In further embodiments, a catalytically dead CasX protein comprises deletions of all or part of the RuvC domain of the CasX protein. It will be understood that the same foregoing substitutions can similarly be introduced into the CasX variants of the disclosure, resulting in a dCasX variant. In one embodiment, all or a portion of the RuvC domain is deleted from the CasX variant, resulting in a dCasX variant. Catalytically inactive dCasX variant proteins can, in some embodiments, be used for base editing or epigenetic modifications. With a higher affinity for DNA, in some embodiments, catalytically inactive dCasX variant proteins can, relative to catalytically active CasX, find their target nucleic acid faster, remain bound to target nucleic acid for longer periods of time, bind target nucleic acid in a more stable fashion, or a combination thereof, thereby improving the function of the catalytically dead CasX variant protein. a. Non-Target Strand Binding Domain
[00210] The reference CasX proteins of the disclosure comprise a non-target strand binding domain (NTSBD). The NTSBD is a domain not previously found in any Cas proteins; for example this domain is not present in Cas proteins such as Cas9, Casl2a / Cpfl, Casl3, Casl4, CASCADE, CSM, or CSY. Without being bound to theory or mechanism, a NTSBD in a CasX allows for binding to the non-target DNA strand and may aid in unwinding of the non-target and target strands. The NTSBD is presumed to be responsible for the unwinding, or the capture, of a non-target DNA strand in the unwound state. The NTSBD is in direct contact with the non-target strand in CryoEM model structures derived to date and may contain a non-canonical zinc finger domain. The NTSBD may also play a role in stabilizing DNA during unwinding, guide RNA invasion and R-loop formation. In some embodiments, an exemplary NTSBD comprises amino acids 101-191 of SEQ ID NO: 1 or amino acids 103-192 of SEQ ID NO: 2. In some embodiments, the NTSBD of a reference CasX protein comprises a four-stranded beta sheet. b. Target Strand Loading Domain
[00211] The reference CasX proteins of the disclosure comprise a Target Strand Loading (TSL) domain. The TSL domain is a domain not found in certain Cas proteins such as Cas9, CASCADE, CSM, or CSY. Without wishing to be bound by theory or mechanism, it is thought that the TSL domain is responsible for aiding the loading of the target DNA strand into the RuvC active site of a CasX protein. In some embodiments, the TSL acts to place or capture the target-strand in a folded state that places the scissile phosphate of the target strand DNA backbone in the RuvC active site. The TSL comprises a cys4 (CXXC (SEQ ID NO: 246, CXXC (SEQ ID NO: 246) zinc finger / ribbon domain that is separated by the bulk of the TSL. In some embodiments, an exemplary TSL comprises amino acids 825-934 of SEQ ID NO: 1 or amino acids 813-921 of SEQ ID NO: 2. c. Helical I Domain
[00212] The reference CasX proteins of the disclosure comprise a helical I domain. Certain Cas proteins other than CasX have domains that may be named in a similar way. However, in some embodiments, the helical I domain of a CasX protein comprises one or more unique structural features, or comprises a unique sequence, or a combination thereof, compared to non-CasX proteins. For example, in some embodiments, the helical I domain of a CasX protein comprises one or more unique secondary structures compared to domains in other Cas proteins that may have a similar name. For example, in some embodiments the helical I domain in a CasX protein comprises one or more alpha helices of unique structure and sequence in arrangement, number and length compared to other CRISPR proteins. In certain embodiments, the helical I domain is responsible for interacting with the bound DNA and spacer of the guide RNA. Without wishing to be bound by theory, it is thought that in some cases the helical I domain may contribute to binding of the protospacer adjacent motif (PAM). In some embodiments, an exemplary helical I domain comprises amino acids 57-100 and 192-332 of SEQ ID NO: 1, or amino acids 59-102 and 193-333 of SEQ ID NO: 2. In some embodiments, the helical I domain of a reference CasX protein comprises one or more alpha helices. d. Helical II Domain
[00213] The reference CasX proteins of the disclosure comprise a helical II domain. Certain Cas proteins other than CasX have domains that may be named in a similar way. However, in some embodiments, the helical II domain of a CasX protein comprises one or more unique structural features, or a unique sequence, or a combination thereof, compared to domains in other Cas proteins that may have a similar name. For example, in some embodiments, the helical II domain comprises one or more unique structural alpha helical bundles that align along the target DNA:guide RNA channel. In some embodiments, in a CasX comprising a helical II domain, the target strand and guide RNA interact with helical II (and the helical I domain, in some embodiments) to allow RuvC domain access to the target DNA. The helical II domain is responsible for binding to the guide RNA scaffold stem loop as well as the bound DNA. In some embodiments, an exemplary helical II domain comprises amino acids 333-509 of SEQ ID NO: 1, or amino acids 334-501 of SEQ ID NO: 2. e. Oligonucleotide Binding Domain
[00214] The reference CasX proteins of the disclosure comprise an Oligonucleotide Binding Domain (OBD). Certain Cas proteins other than CasX have domains that may be named in a similar way. However, in some embodiments, the OBD comprises one or more unique functional features, or comprises a sequence unique to a CasX protein, or a combination thereof. For example, in some embodiments the bridged helix (BH), helical I domain, helical II domain, and Oligonucleotide Binding Domain (OBD) together are responsible for binding of a CasX protein to the guide RNA. Thus, for example, in some embodiments the OBD is unique to a CasX protein in that it interacts functionally with a helical I domain, or a helical II domain, or both, each of which may be unique to a CasX protein as described herein. Specifically, in CasX the OBD largely binds the RNA triplex of the guide RNA scaffold. The OBD may also be responsible for binding to the protospacer adjacent motif (PAM). An exemplary OBD domain comprises amino acids 1-56 and 510-660 of SEQ ID NO: 1, or amino acids 1-58 and 502-647 of SEQ ID NO: 2. f. RuvC DNA Cleavage Domain
[00215] The reference CasX proteins of the disclosure comprise a RuvC domain, that includes 2 partial RuvC domains (RuvC-I and RuvC-II). The RuvC domain is the ancestral domain of all type 12 CRISPR proteins. The RuvC domain originates from a TNPB (transposase B) like transposase. Similar to other RuvC domains, the CasX RuvC domain has a DED catalytic triad that is responsible for coordinating a magnesium (Mg) ion and cleaving DNA. In some embodiments, the RuvC has a DED motif active site that is responsible for cleaving both strands of DNA (one by one, most likely the non-target strand first at 11-14 nucleotides (nt) into the targeted sequence and then the target strand next at 2-4 nucleotides after the target sequence). Specifically in CasX, the RuvC domain is unique in that it is also responsible for binding the guide RNA scaffold stem loop that is critical for CasX function. An exemplary RuvC domain comprises amino acids 661-824 and 935-986 of SEQ ID NO: 1, or amino acids 648-812 and 922-978 of SEQ ID NO: 2. g. Reference CasX Proteins
[00216] The disclosure provides reference CasX proteins. In some embodiments, a reference CasX protein is a naturally-occurring protein. For example, reference CasX proteins can be isolated from naturally occurring prokaryotes, such as Dehaproteobacteria. Planctomycetes, or Candidatus Sungbacteria species. A reference CasX protein (sometimes referred to herein as a reference CasX polypeptide) is a type II CRISPR / Cas endonuclease belonging to the CasX (sometimes referred to as Casl2e) family of proteins that is capable of interacting with a guide NA to form a ribonucleoprotein (RNP) complex. In some embodiments, the RNP complex comprising the reference CasX protein can be targeted to a particular site in a target nucleic acid via base pairing between the targeting sequence (or spacer) of the gNA and a target sequence in the target nucleic acid. In some embodiments, the RNP comprising the reference CasX protein is capable of cleaving target DNA. In some embodiments, the RNP comprising the reference CasX protein is capable of nicking target DNA. In some embodiments, the RNP comprising the reference CasX protein is capable of editing target DNA, for example in those embodiments where the reference CasX protein is capable of cleaving or nicking DNA, followed by non-homologous end joining (NHEJ), homology-directed repair (HDR), homology-independent targeted integration (HITI), micro-homology mediated end joining (MMEJ), single strand annealing (SSA) or base excision repair (BER). In some embodiments, the RNP comprising the CasX protein is a catalytically dead (is catalytically inactive or has substantially no cleavage activity) CasX protein (dCasX), but retains the ability to bind the target DNA, described more fully, supra.
[00217] In some cases, a reference CasX protein is isolated or derived from Deltaproteobacteria. In some embodiments, a CasX protein comprises a sequence at least 50% identical, at least 60% identical, at least 65% identical, at least 70% identical, at least 75% identical, at least 80% identical, at least 81% identical, at least 82% identical, at least 83% identical, at least 84% identical, at least 85% identical, at least 86% identical, at least 86% identical, at least 87% identical, at least 88% identical, at least 89% identical, at least 89% identical, at least 90% identical, at least 91% identical, at least 92% identical, at least 93% identical, at least 94% identical, at least 95% identical, at least 96% identical, at least 97% identical, at least 98% identical, at least 99% identical, at least 99.5% identical or 100% identical to a sequence of: 1 MEKRINKIRK KLSADNATKP VSRSGPMKTL LVRVMTDDLK KRLEKRRKKP EVMPQVISNN 61 AANNLRMLLD DYTKMKEAIL QVYWQEFKDD HVGLMCKFAQ PASKKIDQNK LKPEMDEKGN 121 LTTAGFACSQ CGQPLFVYKL EQVSEKGKAY TNYFGRCNVA EHEKLILLAQ LKPEKDSDEA 181 VTYSLGKFGQ RALDFYSIHV TKESTHPVKP LAQIAGNRYA SGPVGKALSD ACMGTIASFL 241 SKYQDIIIEH QKWKGNQKR LESLRELAGK ENLEYPSVTL PPQPHTKEGV DAYNEVIARV 301 RMWVNLNLWQ KLKLSRDDAK PLLRLKGFPS FPWERRENE VDWWNTINEV KKLIDAKRDM 361 GRVFWSGVTA EKRNTILEGY NYLPNENDHK KREGSLENPK KPAKRQFGDL LLYLEKKYAG 421 DWGKVFDEAW ERIDKKIAGL TSHIEREEAR NAEDAQSKAV LTDWLRAKAS FVLERLKEMD 481 EKEFYACEIQ LQKWYGDLRG NPFAVEAENR WDISGFSIG SDGHSIQYRN LLAWKYLENG 541 KREFYLLMNY GKKGRIRFTD GTDIKKSGKW QGLLYGGGKA KVIDLTFDPD DEQLIILPLA 601 FGTRQGREFI WNDLLSLETG LIKLANGRVI EKTIYNKKIG RDEPALFVAL TFERREWDP 661 SNIKPVNLIG VDRGENIPAV IALTDPEGCP LPEFKDSSGG PTDILRIGEG YKEKQRAIQA 721 AKEVEQRRAG GYSRKFASKS RNLADDMVRN SARDLFYHAV THDAVLVFEN LSRGFGRQGK 781 RTFMTERQYT KMEDWLTAKL AYEGLTSKTY LSKTLAQYTS KTCSNCGFTI TTADYDGMLV 841 RLKKTSDGWA TTLNNKELKA EGQITYYNRY KRQTVEKELS AELDRLSEES GNNDISKWTK 901 GRRDEALFLL KKRFSHRPVQ EQFVCLDCGH EVHADEQAAL NIARSWLFLN SNSTEFKSYK 961 SGKQPFVGAW QAFYKRRLKE VWKPNA (SEQ ID NO: 1).
[00218] In some cases, a reference CasX protein is isolated or derived from Planctomycetes. In some embodiments, a CasX protein comprises a sequence at least 50% identical, at least 60% identical, at least 65% identical, at least 70% identical, at least 75% identical, at least 80% identical, at least 81% identical, at least 82% identical, at least 83% identical, at least 84% identical, at least 85% identical, at least 86% identical, at least 86% identical, at least 87% identical, at least 88% identical, at least 89% identical, at least 89% identical, at least 90% identical, at least 91% identical, at least 92% identical, at least 93% identical, at least 94% identical, at least 95% identical, at least 96% identical, at least 97% identical, at least 98% identical, at least 99% identical, at least 99.5% identical or 100% identical to a sequence of: 1 MQEIKRINKI RRRLVKDSNT KKAGKTGPMK TLLVRVMTPD LRERLENLRK KPENIPQPIS 61 NTSRANLNKL LTDYTEMKKA ILHVYWEEFQ KDPVGLMSRV AQPAPKNIDQ RKLIPVKDGN 121 ERLTSSGFAC SQCCQPLYVY KLEQVNDKGK PHTNYFGRCN VSEHERLILL SPHKPEANDE 181 LVTYSLGKFG QRALDFYSIH VTRESNHPVK PLEQIGGNSC ASGPVGKALS DACMGAVASF 241 LTKYQDIILE HQKVIKKNEK RLANLKDIAS ANGLAFPKIT LPPQPHTKEG IEAYNNWAQ 301 IVIWVNLNLW QKLKIGRDEA KPLQRLKGFP SFPLVERQAN EVDWWDMVCN VKKLINEKKE 361 DGKVFWQNLA GYKRQEALLP YLSSEEDRKK GKKFARYQFG DLLLHLEKKH GEDWGKVYDE 421 AWERIDKKVE GLSKHIKLEE ERRSEDAQSK AALTDWLRAK ASFVIEGLKE ADKDEFCRCE 481 LKLQKWYGDL RGKPFAIEAE NSILDISGFS KQYNCAFIWQ KDGVKKLNLY LIINYFKGGK 541 LRFKKIKPEA FEANRFYTVI NKKSGEIVPM EVNFNFDDPN LIILPLAFGK RQGREFIWND 601 LLSLETGSLK LANGRVIEKT LYNRRTRQDE PALFVALTFE RREVLDSSNI KPMNLIGIDR 661 GENIPAVIAL TDPEGCPLSR FKDSLGNPTH ILRIGESYKE KQRTIQAAKE VEQRRAGGYS 721 RKYASKAKNL ADDMVRNTAR DLLYYAVTQD AMLIFENLSR GFGRQGKRTF MAERQYTRME 781 DWLTAKLAYE GLPSKTYLSK TLAQYTSKTC SNCGFTITSA DYDRVLEKLK KTATGWMTTI 841 NGKELKVEGQ ITYYNRYKRQ NWKDLSVEL DRLSEESVNN DISSWTKGRS GEALSLLKKR 901 FSHRPVQEKF VCLNCGFETH ADEQAALNIA RSWLFLRSQE YKKYQTNKTT GNTDKRAFVE 961 TWQSFYRKKL KEVWKPAV (SEQ ID NO: 2).
[00219] In some embodiments, the CasX protein comprises the sequence of SEQ ID NO: 2, or at least 60% similarity thereto. In some embodiments, the CasX protein comprises the sequence of SEQ ID NO: 2, or at least 80% similarity thereto. In some embodiments, the CasX protein comprises the sequence of SEQ ID NO: 2, or at least 90% similarity thereto. In some embodiments, the CasX protein comprises the sequence of SEQ ID NO: 2, or at least 95% similarity thereto. In some embodiments, the CasX protein consists of the sequence of SEQ ID NO: 2. In some embodiments, the CasX protein comprises or consists of a sequence that has at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 20, at least 30, at least 40 or at least 50 mutations relative to the sequence of SEQ ID NO: 2. These mutations can be insertions, deletions, amino acid substitutions, or any combinations thereof.
[00220] In some cases, a reference CasX protein is isolated or derived from Candidatus Sungbacteria. In some embodiments, a CasX protein comprises a sequence at least 50% identical, at least 60% identical, at least 65% identical, at least 70% identical, at least 75% identical, at least 80% identical, at least 81% identical, at least 82% identical, at least 83% identical, at least 84% identical, at least 85% identical, at least 86% identical, at least 86% identical, at least 87% identical, at least 88% identical, at least 89% identical, at least 89% identical, at least 90% identical, at least 91% identical, at least 92% identical, at least 93% identical, at least 94% identical, at least 95% identical, at least 96% identical, at least 97% identical, at least 98% identical, at least 99% identical, at least 99.5% identical or 100% identical to a sequence of 1 MDNANKPSTK SLVNTTRISD HFGVTPGQVT RVFSFGIIPT KRQYAIIERW FAAVEAARER 61 LYGMLYAHFQ ENPPAYLKEK FSYETFFKGR PVLNGLRDID PTIMTSAVFT ALRHKAEGAM 121 AAFHTNHRRL FEEARKKMRE YAECLKANEA LLRGAADIDW DKIVNALRTR LNTCLAPEYD 181 AVIADFGALC AFRALIAETN ALKGAYNHAL NQMLPALVKV DEPEEAEESP RLRFFNGRIN 241 DLPKFPVAER ETPPDTETII RQLEDMARVI PDTAEILGYI HRIRHKAARR KPGSAVPLPQ 301 RVALYCAIRM ERNPEEDPST VAGHFLGEID RVCEKRRQGL VRTPFDSQIR ARYMDIISFR 361 ATLAHPDRWT EIQFLRSNAA SRRVRAETIS APFEGFSWTS NRTNPAPQYG MALAKDANAP 421 ADAPELCICL SPSSAAFSVR EKGGDLIYMR PTGGRRGKDN PGKEITWVPG SFDEYPASGV 481 ALKLRLYFGR SQARRMLTNK TWGLLSDNPR VFAANAELVG KKRNPQDRWK LFFHMVISGP 541 PPVEYLDFSS DVRSRARTVI GINRGEVNPL AYAWSVEDG QVLEEGLLGK KEYIDQLIET 601 RRRISEYQSR EQTPPRDLRQ RVRHLQDTVL GSARAKIHSL IAFWKGILAI ERLDDQFHGR 661 EQKIIPKKTY LANKTGFMNA LSFSGAVRVD KKGNPWGGMI EIYPGGISRT CTQCGTVWLA 721 RRPKNPGHRD AMWIPDIVD DAAATGFDNV DCDAGTVDYG ELFTLSREWV RLTPRYSRVM 781 RGTLGDLERA IRQGDDRKSR QMLELALEPQ PQWGQFFCHR CGFNGQSDVL AATNLARRAI 841 SLIRRLPDTD TPPTP (SEQ ID NO: 3).
[00221] In some embodiments, the CasX protein comprises the sequence of SEQ ID NO: 3, or at least 60% similarity thereto. In some embodiments, the CasX protein comprises the sequence of SEQ ID NO: 3, or at least 80% similarity thereto. In some embodiments, the CasX protein comprises the sequence of SEQ ID NO: 3, or at least 90% similarity thereto. In some embodiments, the CasX protein comprises the sequence of SEQ ID NO: 3, or at least 95% similarity thereto. In some embodiments, the CasX protein consists of the sequence of SEQ ID NO: 3. In some embodiments, the CasX protein comprises or consists of a sequence that has at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 20, at least 30, at least 40 or at least 50 mutations relative to the sequence of SEQ ID NO: 3. These mutations can be insertions, deletions, amino acid substitutions, or any combinations thereof. h. CasX Variant Proteins
[00222] The present disclosure provides variants of a reference CasX protein (interchangeably referred to herein as “CasX variant” or “CasX variant protein”), wherein the CasX variants comprise at least one modification in at least one domain relative to the reference CasX protein, including but not limited to the sequences of SEQ ID NOS:1-3. In some embodiments, the CasX variant exhibits at least one improved characteristic compared to the reference CasX protein. All variants that improve one or more functions or characteristics of the CasX variant protein when compared to a reference CasX protein described herein are envisaged as being within the scope of the disclosure. In some embodiments, the modification is a mutation in one or more amino acids of the reference CasX. In other embodiments, the modification is a substitution of one or more domains of the reference CasX with one or more domains from a different CasX. In some embodiments, insertion includes the insertion of a part or all of a domain from a different CasX protein. Mutations can occur in any one or more domains of the reference CasX protein, and may include, for example, deletion of part or all of one or more domains, or one or more amino acid substitutions, deletions, or insertions in any domain of the reference CasX protein. The domains of CasX proteins include the non-target strand binding (NTSB) domain, the target strand loading (TSL) domain, the helical I domain, the helical II domain, the oligonucleotide binding domain (OBD), and the RuvC DNA cleavage domain. Any change in amino acid sequence of a reference CasX protein that leads to an improved characteristic of the CasX protein is considered a CasX variant protein of the disclosure. For example, CasX variants can comprise one or more amino acid substitutions, insertions, deletions, or swapped domains, or any combinations thereof, relative to a reference CasX protein sequence.
[00223] In some embodiments, the CasX variant protein comprises at least one modification in at least each of two domains of the reference CasX protein, including the sequences of SEQ ID NOS: 1-3. In some embodiments, the CasX variant protein comprises at least one modification in at least 2 domains, in at least 3 domains, at least 4 domains or at least 5 domains of the reference CasX protein. In some embodiments, the CasX variant protein comprises two or more modifications in at least one domain of the reference CasX protein. In some embodiments, the CasX variant protein comprises at least two modifications in at least one domain of the reference CasX protein, at least three modifications in at least one domain of the reference CasX protein or at least four modifications in at least one domain of the reference CasX protein. In some embodiments, wherein the CasX variant comprises two or more modifications compared to a reference CasX protein, each modification is made in a domain independently selected from the group consisting of a NTSBD, TSLD, helical I domain, helical II domain, OBD, and RuvC DNA cleavage domain.
[00224] In some embodiments, the at least one modification of the CasX variant protein comprises a deletion of at least a portion of one domain of the reference CasX protein. In some embodiments, the deletion is in the NTSBD, TSLD, helical I domain, helical II domain, OBD, or RuvC DNA cleavage domain.
[00225] Suitable mutagenesis methods for generating CasX variant proteins of the disclosure may include, for example, Deep Mutational Evolution (DME), deep mutational scanning (DMS), error prone PCR, cassette mutagenesis, random mutagenesis, staggered extension PCR, gene shuffling, or domain swapping. Exemplary methods for the generation of CasX variants with improved characteristics are provided in the Examples, below. In some embodiments, the CasX variants are designed, for example by selecting one or more desired mutations in a reference CasX. In certain embodiments, the activity of a reference CasX protein is used as a benchmark against which the activity of one or more CasX variants are compared, thereby measuring improvements in function of the CasX variants. Exemplary improvements of CasX variants include, but are not limited to, improved folding of the variant, improved binding affinity to the gNA, improved binding affinity to the target DNA, improved ability to utilize a greater spectrum of PAM sequences in the editing or binding of target DNA, improved unwinding of the target DNA, increased editing activity, improved editing efficiency, improved editing specificity, increased activity of the nuclease, increased target strand loading for double strand cleavage, decreased target strand loading for single strand nicking, decreased off-target cleavage, improved binding of the non-target strand of DNA, improved protein stability, improved CasX:gNA (RNP) complex stability, improved protein solubility, improved CasX:gNA (RNP) complex solubility, improved protein yield, improved protein expression, and improved fusion characteristics, as described more fully, below.
[00226] In some embodiments of the CasX variants described herein, the at least one modification comprises: (a) a substitution of 1 to 100 consecutive or non-consecutive amino acids in the CasX variant; (b) a deletion of 1 to 100 consecutive or non-consecutive amino acids in the CasX variant; (c) an insertion of 1 to 100 consecutive or non-consecutive amino acids in the CasX; or (d) any combination of (a)-(c). In some embodiments, the at least one modification comprises: (a) a substitution of 5-10 consecutive or non-consecutive amino acids in the CasX variant; (b) a deletion of 1-5 consecutive or non-consecutive amino acids in the CasX variant; (c) an insertion of 1-5 consecutive or non-consecutive amino acids in the CasX; or (d) any combination of (a)-(c).
[00227] In some embodiments, the CasX variant protein comprises or consists of a sequence that has at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 20, at least 30, at least 40 or at least 50 mutations relative to the sequence of SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 3. These mutations can be insertions, deletions, amino acid substitutions, or any combinations thereof.
[00228] In some embodiments, the CasX variant protein comprises at least one amino acid substitution in at least one domain of a reference CasX protein. In some embodiments, the CasX variant protein comprises at least about 1-4 amino acid substitutions, 1-10 amino acid substitutions, 1-20 amino acid substitutions, 1-30 amino acid substitutions, 1-40 amino acid substitutions, 1-50 amino acid substitutions, 1-60 amino acid substitutions, 1-70 amino acid substitutions, 1-80 amino acid substitutions, 1-90 amino acid substitutions, 1-100 amino acid substitutions, 2-10 amino acid substitutions, 2-20 amino acid substitutions, 2-30 amino acid substitutions, 3-10 amino acid substitutions, 3-20 amino acid substitutions, 3-30 amino acid substitutions, 4-10 amino acid substitutions, 4-20 amino acid substitutions, 3-300 amino acid substitutions, 5-10 amino acid substitutions, 5-20 amino acid substitutions, 5-30 amino acid substitutions, 10-50 amino acid substitutions, or 20-50 amino acid substitutions, relative to a reference CasX protein. In some embodiments, the CasX variant protein comprises at least about 100 amino acid substitutions relative to a reference CasX protein. In some embodiments, the CasX variant protein comprises 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 amino acid substitutions relative to a reference CasX protein. In some embodiments, the CasX variant protein comprises 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 amino acid substitutions in a single domain relative to the reference CasX protein. In some embodiments, the amino acid substitutions are conservative substitutions. In other embodiments, the substitutions are non-conservative; e.g., a polar amino acid is substituted for a non-polar amino acid, or vice versa.
[00229] In some embodiments, a CasX variant protein comprises 1 amino acid substitution, 2-3 consecutive amino acid substitutions, 2-4 consecutive amino acid substitutions, 2-5 consecutive amino acid substitutions, 2-6 consecutive amino acid substitutions, 2-7 consecutive amino acid substitutions, 2-8 consecutive amino acid substitutions, 2-9 consecutive amino acid substitutions, 2-10 consecutive amino acid substitutions, 2-20 consecutive amino acid substitutions, 2-30 consecutive amino acid substitutions, 2-40 consecutive amino acid substitutions, 2-50 consecutive amino acid substitutions, 2-60 consecutive amino acid substitutions, 2-70 consecutive amino acid substitutions, 2-80 consecutive amino acid substitutions, 2-90 consecutive amino acid substitutions, 2-100 consecutive amino acid substitutions, 3-10 consecutive amino acid substitutions, 3-20 consecutive amino acid substitutions, 3-30 consecutive amino acid substitutions, 4-10 consecutive amino acid substitutions, 4-20 consecutive amino acid substitutions, 3-300 consecutive amino acid substitutions, 5-10 consecutive amino acid substitutions, 5-20 consecutive amino acid substitutions, 5-30 consecutive amino acid substitutions, 10-50 consecutive amino acid substitutions or 20-50 consecutive amino acid substitutions relative to a reference CasX protein. In some embodiments, a CasX variant protein comprises 2,3,4, 5,6,7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 consecutive amino acid substitutions. In some embodiments, a CasX variant protein comprises a substitution of at least about 100 consecutive amino acids. As used herein “consecutive amino acids” refer to amino acids that are contiguous in the primary sequence of a polypeptide.
[00230] In some embodiments, a CasX variant protein comprises two or more substitutions relative to a reference CasX protein, and the two or more substitutions are not in consecutive amino acids of the reference CasX sequence. For example, a first substitution may be in a first domain of the reference CasX protein, and a second substitution may be in a second domain of the reference CasX protein. In some embodiments, a CasX variant protein comprises 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 non-consecutive substitutions relative to a reference CasX protein. In some embodiments, a CasX variant protein comprises at least 20 non-consecutive substitutions relative to a reference CasX protein. Each non-consecutive substitution may be of any length of amino acids described herein, e.g., 1-4 amino acids, 1-10 amino acids, and the like. In some embodiments, the two or more substitutions relative to the reference CasX protein are not the same length, for example, one substitution is one amino acid and a second substitution is three amino acids. In some embodiments, the two or more substitutions relative to the reference CasX protein are the same length, for example both substitutions are two consecutive amino acids in length.
[00231] Any amino acid can be substituted for any other amino acid in the substitutions described herein. The substitution can be a conservative substitution (e.g., a basic amino acid is substituted for another basic amino acid). The substitution can be a non-conservative substitution (e.g., a basic amino acid is substituted for an acidic amino acid or vice versa). For example, a proline in a reference CasX protein can be substituted for any of arginine, histidine, lysine, aspartic acid, glutamic acid, serine, threonine, asparagine, glutamine, cysteine, glycine, alanine, isoleucine, leucine, methionine, phenylalanine, tryptophan, tyrosine or valine to generate a CasX variant protein of the disclosure.
[00232] In some embodiments, a CasX variant protein comprises at least one amino acid deletion relative to a reference CasX protein. In some embodiments, a CasX variant protein comprises a deletion of 1-4 amino acids, 1-10 amino acids, 1-20 amino acids, 1-30 amino acids, 1-40 amino acids, 1-50 amino acids, 1-60 amino acids, 1-70 amino acids, 1-80 amino acids, 1-90 amino acids, 1-100 amino acids, 2-10 amino acids, 2-20 amino acids, 2-30 amino acids, 3-10 amino acids, 3-20 amino acids, 3-30 amino acids, 4-10 amino acids, 4-20 amino acids, 3-300 amino acids, 5-10 amino acids, 5-20 amino acids, 5-30 amino acids, 10-50 amino acids or 20-50 amino acids relative to a reference CasX protein. In some embodiments, a CasX variant comprises a deletion of at least about 100 consecutive amino acids relative to a reference CasX protein. In some embodiments, a CasX variant protein comprises a deletion of at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50 or 100 consecutive amino acids relative to a reference CasX protein. In some embodiments, a CasX variant protein comprises a deletion of 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 consecutive amino acids.
[00233] In some embodiments, a CasX variant protein comprises two or more deletions relative to a reference CasX protein, and the two or more deletions are not consecutive amino acids. For example, a first deletion may be in a first domain of the reference CasX protein, and a second deletion may be in a second domain of the reference CasX protein. In some embodiments, a CasX variant protein comprises 2, 3, 4, 5,6,7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 non-consecutive deletions relative to a reference CasX protein. In some embodiments, a CasX variant protein comprises at least 20 non-consecutive deletions relative to a reference CasX protein. Each non-consecutive deletion may be of any length of amino acids described herein, e.g., 1-4 amino acids, 1-10 amino acids, and the like.
[00234] In some embodiments, the CasX variant protein comprises at least one amino acid insertion. In some embodiments, a CasX variant protein comprises an insertion of 1 amino acid, an insertion of 2-3 consecutive amino acids, 2-4 consecutive amino acids, 2-5 consecutive amino acids, 2-6 consecutive amino acids, 2-7 consecutive amino acids, 2-8 consecutive amino acids, 2-9 consecutive amino acids, 2-10 consecutive amino acids, 2-20 consecutive amino acids, 2-30 consecutive amino acids, 2-40 consecutive amino acids, 2-50 consecutive amino acids, 2-60 consecutive amino acids, 2-70 consecutive amino acids, 2-80 consecutive amino acids, 2-90 consecutive amino acids, 2-100 consecutive amino acids, 3-10 consecutive amino acids, 3-20 consecutive amino acids, 3-30 consecutive amino acids, 4-10 consecutive amino acids, 4-20 consecutive amino acids, 3-300 consecutive amino acids, 5-10 consecutive amino acids, 5-20 consecutive amino acids, 5-30 consecutive amino acids, 10-50 consecutive amino acids or 20-50 consecutive amino acids relative to a reference CasX protein. In some embodiments, the CasX variant protein comprises an insertion of 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 consecutive amino acids. In some embodiments, a CasX variant protein comprises an insertion of at least about 100 consecutive amino acids.
[00235] In some embodiments, a CasX variant protein comprises two or more insertions relative to a reference CasX protein, and the two or more insertions are not consecutive amino acids of the sequence. For example, a first insertion may be in a first domain of the reference CasX protein, and a second insertion may be in a second domain of the reference CasX protein. In some embodiments, a CasX variant protein comprises 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 non-consecutive insertions relative to a reference CasX protein. In some embodiments, a CasX variant protein comprises at least 10 to about 20 or more non-consecutive insertions relative to a reference CasX protein. Each non-consecutive insertion may be of any length of amino acids described herein, e.g., 1-4 amino acids, 1-10 amino acids, and the like.
[00236] Any amino acid, or combination of amino acids, can be inserted as described herein. For example, a proline, arginine, histidine, lysine, aspartic acid, glutamic acid, serine, threonine, asparagine, glutamine, cysteine, glycine, alanine, isoleucine, leucine, methionine, phenylalanine, tryptophan, tyrosine or valine or any combination thereof can be inserted into a reference CasX protein of the disclosure to generate a CasX variant protein.
[00237] Any permutation of the substitution, insertion and deletion embodiments described herein can be combined to generate a CasX variant protein of the disclosure. For example, a CasX variant protein can comprise at least one substitution and at least one deletion relative to a reference CasX protein sequence, at least one substitution and at least one insertion relative to a reference CasX protein sequence, at least one insertion and at least one deletion relative to a reference CasX protein sequence, or at least one substitution, one insertion and one deletion relative to a reference CasX protein sequence.
[00238] In some embodiments, the CasX variant protein has at least about 60% sequence similarity, at least 70% similarity, at least 80% similarity, at least 85% similarity, at least 86% similarity, at least 87% similarity, at least 88% similarity, at least 89% similarity, at least 90% similarity, at least 91% similarity, at least 92% similarity, at least 93% similarity, at least 94% similarity, at least 95% similarity, at least 96% similarity, at least 97% similarity, at least 98% similarity, at least 99% similarity, at least 99.5% similarity, at least 99.6% similarity, at least 99.7% similarity, at least 99.8% similarity or at least 99.9% similarity to one of SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 3.
[00239] In some embodiments, the CasX variant protein has at least about 60% sequence similarity to SEQ ID NO: 2 or a portion thereof. In some embodiments, the CasX variant protein comprises a substitution of Y789T of SEQ ID NO: 2, a deletion of P793 of SEQ ID NO: 2, a substitution of Y789D of SEQ ID NO: 2, a substitution of T72S of SEQ ID NO: 2, a substitution of I546V of SEQ ID NO: 2, a substitution of E552A of SEQ ID NO: 2, a substitution of A636D of SEQ ID NO: 2, a substitution of F536S of SEQ ID NO:2, a substitution of A708K of SEQ ID NO: 2, a substitution of Y797L of SEQ ID NO: 2, a substitution of L792G SEQ ID NO: 2, a substitution of A739V of SEQ ID NO: 2, a substitution of G791M of SEQ ID NO: 2, an insertion of A at position 661 of SEQ ID NO: 2, a substitution of A788W of SEQ ID NO: 2, a substitution of K390R of SEQ ID NO: 2, a substitution of A751S of SEQ ID NO: 2, a substitution of E385A of SEQ ID NO: 2, an insertion of P at position 696 of SEQ ID NO: 2, an insertion of M at position 773 of SEQ ID NO: 2, a substitution of G695H of SEQ ID NO: 2, an insertion of AS at position 793 of SEQ ID NO: 2, an insertion of AS at position 795 of SEQ ID NO: 2, a substitution of C477R of SEQ ID NO: 2, a substitution of C477K of SEQ ID NO: 2, a substitution of C479A of SEQ ID NO: 2, a substitution of C479L of SEQ ID NO: 2, a substitution of I55F of SEQ ID NO: 2, a substitution of K210R of SEQ ID NO: 2, a substitution of C233S of SEQ ID NO: 2, a substitution of D23 IN of SEQ ID NO: 2, a substitution of Q338E of SEQ ID NO: 2, a substitution of Q338R of SEQ ID NO: 2, a substitution of L379R of SEQ ID NO: 2, a substitution of K390R of SEQ ID NO: 2, a substitution of L481Q of SEQ ID NO: 2, a substitution of F495S of SEQ ID NO:2, a substitution of D600N of SEQ ID NO: 2, a substitution of T886K of SEQ ID NO: 2, a substitution of A739V of SEQ ID NO: 2, a substitution of K460N of SEQ ID NO: 2, a substitution of I199F of SEQ ID NO: 2, a substitution of G492P of SEQ ID NO: 2, a substitution of T153I of SEQ ID NO: 2, a substitution of R591I of SEQ ID NO: 2, an insertion of AS at position 795 of SEQ ID NO: 2, an insertion of AS at position 796 of SEQ ID NO:2, an insertion of L at position 889 of SEQ ID NO: 2, a substitution of E121D of SEQ ID NO: 2, a substitution of S270W of SEQ ID NO: 2, a substitution of E712Q of SEQ ID NO: 2, a substitution of K942Q of SEQ ID NO: 2, a substitution of E552K of SEQ ID NO:2, a substitution of K25Q of SEQ ID NO: 2, a substitution of N47D of SEQ ID NO: 2, an insertion of T at position 696 of SEQ ID NO: 2, a substitution of L685I of SEQ ID NO: 2, a substitution of N880D of SEQ ID NO: 2, a substitution of Q102R of SEQ ID NO: 2, a substitution of M734K of SEQ ID NO: 2, a substitution ofA724S of SEQ ID NO: 2, a substitution of T704K of SEQ ID NO: 2, a substitution of P224K of SEQ ID NO: 2, a substitution of K25R of SEQ ID NO: 2, a substitution of M29E of SEQ ID NO: 2, a substitution of H152D of SEQ ID NO: 2, a substitution of S219R of SEQ ID NO: 2, a substitution of E475K of SEQ ID NO: 2, a substitution of G226R of SEQ ID NO: 2, a substitution of A377K of SEQ ID NO: 2, a substitution of E480K of SEQ ID NO: 2, a substitution of K416E of SEQ ID NO: 2, a substitution of H164R of SEQ ID NO: 2, a substitution of K767R of SEQ ID NO: 2, a substitution of I7F of SEQ ID NO: 2, a substitution of M29R of SEQ ID NO: 2, a substitution of H435R of SEQ ID NO: 2, a substitution of E385Q of SEQ ID NO: 2, a substitution of E385K of SEQ ID NO: 2, a substitution of I279F of SEQ ID NO: 2, a substitution of D489S of SEQ ID NO: 2, a substitution of D732N of SEQ ID NO: 2, a substitution of A739T of SEQ ID NO: 2, a substitution of W885R of SEQ ID NO: 2, a substitution of E53K of SEQ ID NO: 2, a substitution of A238T of SEQ ID NO: 2, a substitution of P283Q of SEQ ID NO: 2, a substitution of E292K of SEQ ID NO: 2, a substitution of Q628E of SEQ ID NO: 2, a substitution of R388Q of SEQ ID NO: 2, a substitution of G791M of SEQ ID NO: 2, a substitution of L792K of SEQ ID NO: 2, a substitution of L792E of SEQ ID NO: 2, a substitution of M779N of SEQ ID NO: 2, a substitution of G27D of SEQ ID NO: 2, a substitution of K955R of SEQ ID NO: 2, a substitution of S867R of SEQ ID NO: 2, a substitution of R693I of SEQ ID NO: 2, a substitution of F189Y of SEQ ID NO: 2, a substitution of V635M of SEQ ID NO: 2, a substitution of F399L of SEQ ID NO: 2, a substitution of E498K of SEQ ID NO: 2, a substitution of E386R of SEQ ID NO: 2, a substitution of V254G of SEQ ID NO: 2, a substitution of P793S of SEQ ID NO: 2, a substitution of K188E of SEQ ID NO: 2, a substitution of QT945KI of SEQ ID NO: 2, a substitution of T620P of SEQ ID NO: 2, a substitution of T946P of SEQ ID NO: 2, a substitution of TT949PP of SEQ ID NO: 2, a substitution of N952T of SEQ ID NO: 2, a substitution of K682E of SEQ ID NO: 2, a substitution of K975R of SEQ ID NO: 2, a substitution of L212P of SEQ ID NO: 2, a substitution of E292R of SEQ ID NO: 2, a substitution of I303K of SEQ ID NO: 2, a substitution of C349E of SEQ ID NO: 2, a substitution of E385P of SEQ ID NO: 2, a substitution of E386N of SEQ ID NO: 2, a substitution of D387K of SEQ ID NO: 2, a substitution of L404K of SEQ ID NO: 2, a substitution of E466H of SEQ ID NO: 2, a substitution of C477Q of SEQ ID NO: 2, a substitution of C477H of SEQ ID NO: 2, a substitution of C479A of SEQ ID NO: 2, a substitution of D659H of SEQ ID NO: 2, a substitution of T806V of SEQ ID NO: 2, a substitution of K808S of SEQ ID NO: 2, an insertion of AS at position 797 of SEQ ID NO: 2, a substitution of V959M of SEQ ID NO: 2, a substitution of K975Q of SEQ ID NO: 2, a substitution of W974G of SEQ ID NO: 2, a substitution of A708Q of SEQ ID NO: 2, a substitution of V71 IK of SEQ ID NO: 2, a substitution of D733T of SEQ ID NO: 2, a substitution of L742W of SEQ ID NO: 2, a substitution of V747K of SEQ ID NO: 2, a substitution of F755M of SEQ ID NO: 2, a substitution of M771A of SEQ ID NO: 2, a substitution of M771Q of SEQ ID NO: 2, a substitution of W782Q of SEQ ID NO: 2, a substitution of G791F, of SEQ ID NO: 2 a substitution of L792D of SEQ ID NO: 2, a substitution of L792K of SEQ ID NO: 2, a substitution of P793Q of SEQ ID NO: 2, a substitution of P793G of SEQ ID NO: 2, a substitution of Q804A of SEQ ID NO: 2, a substitution of Y966N of SEQ ID NO: 2, a substitution of Y723N of SEQ ID NO: 2, a substitution of Y857R of SEQ ID NO: 2, a substitution of S890R of SEQ ID NO: 2, a substitution of S932M of SEQ ID NO: 2, a substitution of L897M of SEQ ID NO: 2, a substitution of R624G of SEQ ID NO: 2, a substitution of S603G of SEQ ID NO: 2, a substitution of N737S of SEQ ID NO: 2, a substitution of L307K of SEQ ID NO: 2, a substitution of I658V of SEQ ID NO: 2, an insertion of PT at position 688 of SEQ ID NO: 2, an insertion of SA at position 794 of SEQ ID NO: 2, a substitution of S877R of SEQ ID NO: 2, a substitution of N580T of SEQ ID NO: 2, a substitution of V335G of SEQ ID NO: 2, a substitution of T620S of SEQ ID NO: 2, a substitution of W345G of SEQ ID NO: 2, a substitution of T280S of SEQ ID NO: 2, a substitution of L406P of SEQ ID NO: 2, a substitution of A612D of SEQ ID NO: 2, a substitution of A751S of SEQ ID NO: 2, a substitution of E386R of SEQ ID NO: 2, a substitution of V351M of SEQ ID NO: 2, a substitution of K210N of SEQ ID NO: 2, a substitution of D40A of SEQ ID NO: 2, a substitution of E773G of SEQ ID NO: 2, a substitution of H207L of SEQ ID NO: 2, a substitution of T62A SEQ ID NO: 2, a substitution of T287P of SEQ ID NO: 2, a substitution of T832A of SEQ ID NO: 2, a substitution of A893S of SEQ ID NO: 2, an insertion of V at position 14 of SEQ ID NO: 2, an insertion of AG at position 13 of SEQ ID NO: 2, a substitution of RI IV of SEQ ID NO: 2, a substitution of R12N of SEQ ID NO: 2, a substitution of R13H of SEQ ID NO: 2, an insertion of Y at position 13 of SEQ ID NO: 2, a substitution of R12L of SEQ ID NO: 2, an insertion of Q at position 13 of SEQ ID NO: 2, an substitution of V15S of SEQ ID NO: 2, an insertion of D at position 17 of SEQ ID NO: 2 or a combination thereof.
[00240] In some embodiments, the CasX variant comprises at least one modification in the NTSB domain.
[00241] In some embodiments, the CasX variant comprises at least one modification in the TSL domain. In some embodiments, the at least one modification in the TSL domain comprises an amino acid substitution of one or more of amino acids Y857, S890, or S932 of SEQ ID NO: 2.
[00242] In some embodiments, the CasX variant comprises at least one modification in the helical I domain. In some embodiments, the at least one modification in the helical I domain comprises an amino acid substitution of one or more of amino acids S219, L249, E259, Q252, E292, L307, or D318 of SEQ ID NO: 2.
[00243] In some embodiments, the CasX variant comprises at least one modification in the helical II domain. In some embodiments, the at least one modification in the helical II domain comprises an amino acid substitution of one or more of amino acids D361, L379, E385, E386, D387, F399, L404, R458, C477, or D489 of SEQ ID NO: 2.
[00244] In some embodiments, the CasX variant comprises at least one modification in the OBD domain. In some embodiments, the at least one modification in the OBD comprises an amino acid substitution of one or more of amino acids F536, E552, T620, or 1658 of SEQ ID NO: 2.
[00245] In some embodiments, the CasX variant comprises at least one modification in the RuvC DNA cleavage domain. In some embodiments, the at least one modification in the RuvC DNA cleavage domain comprises an amino acid substitution of one or more of amino acids K682, G695, A708, V711, D732, A739, D733, L742, V747, F755, M771, M779, W782, A788, G791, L792, P793, Y797, M799, Q804, S819, or Y857 or a deletion of amino acid P793 of SEQ ID NO: 2.
[00246] In some embodiments, the CasX variant comprises at least one modification compared to the reference CasX sequence of SEQ ID NO: 2 is selected from one or more of: (a) an amino acid substitution of L379R; (b) an amino acid substitution of A708K; (c) an amino acid substitution of T620P; (d) an amino acid substitution of E385P; (e) an amino acid substitution of Y857R; (f) an amino acid substitution of I658V; (g) an amino acid substitution of F399L; (h) an amino acid substitution of Q252K; (i) an amino acid substitution of L404K; and (j) an amino acid deletion of P793.
[00247] In some embodiments, a CasX variant protein comprises at least two amino acid changes to a reference CasX protein amino acid sequence. The at least two amino acid changes can be substitutions, insertions, or deletions of a reference CasX protein amino acid sequence, or any combination thereof. The substitutions, insertions or deletions can be any substitution, insertion or deletion in the sequence of a reference CasX protein described herein. In some embodiments, the changes are contiguous, non-contiguous, or a combination of contiguous and non-contiguous amino acid changes to a reference CasX protein sequence. In some embodiments, the reference CasX protein is SEQ ID NO: 2. In some embodiments, a CasX variant protein comprises at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 30, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95 or at least 100 amino acid changes to a reference CasX protein sequence. In some embodiments, a CasX variant protein comprises 1-50, 3-40, 5-30, 5-20, 5-15, 5-10, 10-50, 10-40, 10-30, 10-20, 15-50, 15-40, 15-30, 2-25, 2-24, 2-22, 2-23, 2-22, 2-21, 2-20, 2-19, 2-18, 2-17, 2-16, 2-15, 2-14, 2-12, 2-11, 2-10, 2-9, 2-8, 2-7, 2-6, 25, 2-4, 2-3, 3-25, 3-24, 3-22, 3-23, 3-22, 3-21, 3-20, 3-19, 3-18, 3-17, 3-16, 3-15, 3-14, 3-12, 311, 3-10, 3-9, 3-8, 3-7, 3-6, 3-5, 3-4, 4-25, 4-24, 4-22, 4-23, 4-22, 4-21, 4-20, 4-19, 4-18, 4-17, 4-16, 4-15, 4-14, 4-12, 4-11, 4-10, 4-9, 4-8, 4-7, 4-6, 4-5, 5-25, 5-24, 5-22, 5-23, 5-22, 5-21, 520, 5-19, 5-18, 5-17, 5-16, 5-15, 5-14, 5-12, 5-11, 5-10, 5-9, 5-8, 5-7 or 5-6 amino acid changes to a reference CasX protein sequence. In some embodiments, a CasX variant protein comprises 15-20 changes to a reference CasX protein sequence. In some embodiments, a CasX variant protein comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29 or 30 amino acid changes to a reference CasX protein sequence. In some embodiments, the at least two amino acid changes to the sequence of a reference CasX variant protein are selected from the group consisting of: a substitution of Y789T of SEQ ID NO: 2, a deletion of P793 of SEQ ID NO: 2, a substitution of Y789D of SEQ ID NO: 2, a substitution of T72S of SEQ ID NO: 2, a substitution of I546V of SEQ ID NO: 2, a substitution of E552A of SEQ ID NO: 2, a substitution of A636D of SEQ ID NO: 2, a substitution of F536S of SEQ ID NO:2, a substitution of A708K of SEQ ID NO: 2, a substitution of Y797L of SEQ ID NO: 2, a substitution of L792G SEQ ID NO: 2, a substitution of A739V of SEQ ID NO: 2, a substitution of G791M of SEQ ID NO: 2, an insertion of A at position 661of SEQ ID NO: 2, a substitution of A788W of SEQ ID NO: 2, a substitution of K390R of SEQ ID NO: 2, a substitution of A751S of SEQ ID NO: 2, a substitution of E385A of SEQ ID NO: 2, an insertion of P at position 696 of SEQ ID NO: 2, an insertion of M at position 773 of SEQ ID NO: 2, a substitution of G695H of SEQ ID NO: 2, an insertion of AS at position 793 of SEQ ID NO: 2, an insertion of AS at position 795 of SEQ ID NO: 2, a substitution of C477R of SEQ ID NO: 2, a substitution of C477K of SEQ ID NO: 2, a substitution of C479A of SEQ ID NO: 2, a substitution of C479L of SEQ ID NO: 2, a substitution of I55F of SEQ ID NO: 2, a substitution of K210R of SEQ ID NO: 2, a substitution of C233S of SEQ ID NO: 2, a substitution of D23 IN of SEQ ID NO: 2, a substitution of Q338E of SEQ ID NO: 2, a substitution of Q338R of SEQ ID NO: 2, a substitution of L379R of SEQ ID NO: 2, a substitution of K390R of SEQ ID NO: 2, a substitution of L481Q of SEQ ID NO: 2, a substitution of F495S of SEQ ID NO:2, a substitution of D600N of SEQ ID NO: 2, a substitution of T886K of SEQ ID NO: 2, a substitution of A739V of SEQ ID NO: 2, a substitution of K460N of SEQ ID NO: 2, a substitution of I199F of SEQ ID NO: 2, a substitution of G492P of SEQ ID NO: 2, a substitution of T153I of SEQ ID NO: 2, a substitution of R591I of SEQ ID NO: 2, an insertion of AS at position 795 of SEQ ID NO: 2, an insertion of AS at position 796 of SEQ ID NO :2, an insertion of L at position 889 of SEQ ID NO: 2, a substitution of E121D of SEQ ID NO: 2, a substitution of S270W of SEQ ID NO: 2, a substitution of E712Q of SEQ ID NO: 2, a substitution of K942Q of SEQ ID NO: 2, a substitution of E552K of SEQ ID NO:2, a substitution of K25Q of SEQ ID NO: 2, a substitution of N47D of SEQ ID NO: 2, an insertion of T at position 696 of SEQ ID NO: 2, a substitution of L685I of SEQ ID NO: 2, a substitution of N880D of SEQ ID NO: 2, a substitution of Q102R of SEQ ID NO: 2, a substitution of M734K of SEQ ID NO: 2, a substitution of A724S of SEQ ID NO: 2, a substitution of T704K of SEQ ID NO: 2, a substitution of P224K of SEQ ID NO: 2, a substitution of K25R of SEQ ID NO: 2, a substitution of M29E of SEQ ID NO: 2, a substitution of H152D of SEQ ID NO: 2, a substitution of S219R of SEQ ID NO: 2, a substitution of E475K of SEQ ID NO: 2, a substitution of G226R of SEQ ID NO: 2, a substitution of A377K of SEQ ID NO: 2, a substitution of E480K of SEQ ID NO: 2, a substitution of K416E of SEQ ID NO: 2, a substitution of H164R of SEQ ID NO: 2, a substitution of K767R of SEQ ID NO: 2, a substitution of I7F of SEQ ID NO: 2, a substitution of M29R of SEQ ID NO: 2, a substitution of H435R of SEQ ID NO: 2, a substitution of E385Q of SEQ ID NO: 2, a substitution of E385K of SEQ ID NO: 2, a substitution of I279F of SEQ ID NO: 2, a substitution of D489S of SEQ ID NO: 2, a substitution of D732N of SEQ ID NO: 2, a substitution of A739T of SEQ ID NO: 2, a substitution of W885R of SEQ ID NO: 2, a substitution of E53K of SEQ ID NO: 2, a substitution of A238T of SEQ ID NO: 2, a substitution of P283Q of SEQ ID NO: 2, a substitution of E292K of SEQ ID NO: 2, a substitution of Q628E of SEQ ID NO: 2, a substitution of R388Q of SEQ ID NO: 2, a substitution of G791M of SEQ ID NO: 2, a substitution of L792K of SEQ ID NO: 2, a substitution of L792E of SEQ ID NO: 2, a substitution of M779N of SEQ ID NO: 2, a substitution of G27D of SEQ ID NO: 2, a substitution of K955R of SEQ ID NO: 2, a substitution of S867R of SEQ ID NO: 2, a substitution of R693I of SEQ ID NO: 2, a substitution of F189Y of SEQ ID NO: 2, a substitution of V635M of SEQ ID NO: 2, a substitution of F399L of SEQ ID NO: 2, a substitution of E498K of SEQ ID NO: 2, a substitution of E386R of SEQ ID NO: 2, a substitution of V254G of SEQ ID NO: 2, a substitution of P793S of SEQ ID NO: 2, a substitution of K188E of SEQ ID NO: 2, a substitution of QT945KI of SEQ ID NO: 2, a substitution of T620P of SEQ ID NO: 2, a substitution of T946P of SEQ ID NO: 2, a substitution of TT949PP of SEQ ID NO: 2, a substitution of N952T of SEQ ID NO: 2, a substitution of K682E of SEQ ID NO: 2, a substitution of K975R of SEQ ID NO: 2, a substitution of L212P of SEQ ID NO: 2, a substitution of E292R of SEQ ID NO: 2, a substitution of I303K of SEQ ID NO: 2, a substitution of C349E of SEQ ID NO: 2, a substitution of E385P of SEQ ID NO: 2, a substitution of E386N of SEQ ID NO: 2, a substitution of D387K of SEQ ID NO: 2, a substitution of L404K of SEQ ID NO: 2, a substitution of E466H of SEQ ID NO: 2, a substitution of C477Q of SEQ ID NO: 2, a substitution of C477H of SEQ ID NO: 2, a substitution of C479A of SEQ ID NO: 2, a substitution of D659H of SEQ ID NO: 2, a substitution of T806V of SEQ ID NO: 2, a substitution of K808S of SEQ ID NO: 2, an insertion of AS at position 797 of SEQ ID NO: 2, a substitution of V959M of SEQ ID NO: 2, a substitution of K975Q of SEQ ID NO: 2, a substitution of W974G of SEQ ID NO: 2, a substitution of A708Q of SEQ ID NO: 2, a substitution of V71 IK of SEQ ID NO: 2, a substitution of D733T of SEQ ID NO: 2, a substitution of L742W of SEQ ID NO: 2, a substitution of V747K of SEQ ID NO: 2, a substitution of F755M of SEQ ID NO: 2, a substitution of M771A of SEQ ID NO: 2, a substitution of M771Q of SEQ ID NO: 2, a substitution of W782Q of SEQ ID NO: 2, a substitution of G791F, of SEQ ID NO: 2 a substitution of L792D of SEQ ID NO: 2, a substitution of L792K of SEQ ID NO: 2, a substitution of P793Q of SEQ ID NO: 2, a substitution of P793G of SEQ ID NO: 2, a substitution of Q804A of SEQ ID NO: 2, a substitution of Y966N of SEQ ID NO: 2, a substitution of Y723N of SEQ ID NO: 2, a substitution of Y857R of SEQ ID NO: 2, a substitution of S890R of SEQ ID NO: 2, a substitution of S932M of SEQ ID NO: 2, a substitution of L897M of SEQ ID NO: 2, a substitution of R624G of SEQ ID NO: 2, a substitution of S603G of SEQ ID NO: 2, a substitution of N737S of SEQ ID NO: 2, a substitution of L307K of SEQ ID NO: 2, a substitution of I658V of SEQ ID NO: 2, an insertion of PT at position 688 of SEQ ID NO: 2, an insertion of SA at position 794 of SEQ ID NO: 2, a substitution of S877R of SEQ ID NO: 2, a substitution of N580T of SEQ ID NO: 2, a substitution of V335G of SEQ ID NO: 2, a substitution of T620S of SEQ ID NO: 2, a substitution of W345G of SEQ ID NO: 2, a substitution of T280S of SEQ ID NO: 2, a substitution of L406P of SEQ ID NO: 2, a substitution of A612D of SEQ ID NO: 2, a substitution of A751S of SEQ ID NO: 2, a substitution of E386R of SEQ ID NO: 2, a substitution of V351M of SEQ ID NO: 2, a substitution of K210N of SEQ ID NO: 2, a substitution of D40A of SEQ ID NO: 2, a substitution of E773G of SEQ ID NO: 2, a substitution of H207L of SEQ ID NO: 2, a substitution of T62A SEQ ID NO: 2, a substitution of T287P of SEQ ID NO: 2, a substitution of T832A of SEQ ID NO: 2, a substitution of A893S of SEQ ID NO: 2, an insertion of V at position 14 of SEQ ID NO: 2, an insertion of AG at position 13 of SEQ ID NO: 2, a substitution of RI IV of SEQ ID NO: 2, a substitution of R12N of SEQ ID NO: 2, a substitution of R13H of SEQ ID NO: 2, an insertion of Y at position 13 of SEQ ID NO: 2, a substitution of R12L of SEQ ID NO: 2, an insertion of Q at position 13 of SEQ ID NO: 2, an substitution of VI5S of SEQ ID NO: 2 and an insertion of D at position 17 of SEQ ID NO: 2. In some embodiments, the at least two amino acid changes to a reference CasX protein are selected from the amino acid changes disclosed in the sequences of Table 3. In some embodiments, a CasX variant comprises any combination of the foregoing embodiments of this paragraph.
[00248] In some embodiments, a CasX variant protein comprises more than one substitution, insertion and / or deletion of a reference CasX protein amino acid sequence. In some embodiments, the reference CasX protein comprises or consists essentially of SEQ ID NO: 2. In some embodiments, a CasX variant protein comprises a substitution of S794R and a substitution of Y797L of SEQ ID NO: 2. In some embodiments, a CasX variant protein comprises a substitution of K416E and a substitution of A708K of SEQ ID NO: 2. In some embodiments, a CasX variant protein comprises a substitution of A708K and a deletion of P793 of SEQ ID NO: 2. In some embodiments, a CasX variant protein comprises a deletion of P793 and an insertion of AS at position 795 SEQ ID NO: 2. In some embodiments, a CasX variant protein comprises a substitution of Q367K and a substitution of I425S of SEQ ID NO: 2. In some embodiments, a CasX variant protein comprises a substitution of A708K, a deletion of P position 793 and a substitution A793V of SEQ ID NO: 2. In some embodiments, a CasX variant protein comprises a substitution of Q338R and a substitution of A339E of SEQ ID NO: 2. In some embodiments, a CasX variant protein comprises a substitution of Q338R and a substitution of A339K of SEQ ID NO: 2. In some embodiments, a CasX variant protein comprises a substitution of S507G and a substitution of G508R of SEQ ID NO: 2. In some embodiments, a CasX variant protein comprises a substitution of L379R, a substitution of A708K and a deletion of P at position 793 of SEQ ID NO: 2. In some embodiments, a CasX variant protein comprises a substitution of C477K, a substitution of A708K and a deletion of P at position 793 of SEQ ID NO: 2. In some embodiments, a CasX variant protein comprises a substitution of L379R, a substitution of C477K, a substitution of A708K and a deletion of P at position of 793 of SEQ ID NO: 2. In some embodiments, a CasX variant protein comprises a substitution of L379R, a substitution of A708K, a deletion of P at position 793 and a substitution A739V of SEQ ID NO: 2. In some embodiments, a CasX variant protein comprises a substitution of C477K, a substitution of A708K, a deletion of P at position 793 and a substitution of A739V of SEQ ID NO: 2. In some embodiments, a CasX variant protein comprises a substitution of L379R, a substitution of C477K, a substitution of A708K, a deletion of P at position 793 and a substitution of A739V of SEQ ID NO: 2. In some embodiments, a CasX variant protein comprises a substitution of L379R, a substitution of A708K, a deletion of P at position 793 and a substitution of M779N of SEQ ID NO: 2. In some embodiments, a CasX variant protein comprises a substitution of L379R, a substitution of A708K, a deletion of P at position 793 and a substitution of M771N of SEQ ID NO: 2. In some embodiments, a CasX variant protein comprises a substitution of L379R, a substitution of 708K, a deletion of P at position 793 and a substitution of D489S of SEQ ID NO: 2. In some embodiments, a CasX variant protein comprises a substitution of L379R, a substitution of A708K, a deletion of P at position 793 and a substitution of A739T of SEQ ID NO: 2. In some embodiments, a CasX variant protein comprises a substitution of L379R, a substitution of A708K, a deletion of P at position 793 and a substitution of D732N of SEQ ID NO: 2. In some embodiments, a CasX variant protein comprises a substitution of L379R, a substitution of A708K, a deletion of P at position 793 and a substitution of G791M of SEQ ID NO: 2. In some embodiments, a CasX variant protein comprises a substitution of L379R, a substitution of 708K, a deletion of P at position 793 and a substitution of Y797L of SEQ ID NO: 2. In some embodiments, a CasX variant protein comprises a substitution of L379R, a substitution of C477K, a substitution of A708K, a deletion of P at position 793 and a substitution of M779N of SEQ ID NO: 2. In some embodiments, a CasX variant protein comprises a substitution of L379R, a substitution of C477K, a substitution of A708K, a deletion of P at position 793 and a substitution of M771N of SEQ ID NO: 2. In some embodiments, a CasX variant protein comprises a substitution of L379R, a substitution of C477K, a substitution of A708K, a deletion of P at position 793 and a substitution of D489S of SEQ ID NO: 2. In some embodiments, a CasX variant protein comprises a substitution of L379R, a substitution of C477K, a substitution of A708K, a deletion of P at position 793 and a substitution of A739T of SEQ ID NO: 2. In some embodiments, a CasX variant protein comprises a substitution of L379R, a substitution of C477K, a substitution of A708K, a deletion of P at position 793 and a substitution of D732N of SEQ ID NO: 2. In some embodiments, a CasX variant protein comprises a substitution of L379R, a substitution of C477K, a substitution of A708K, a deletion of P at position 793 and a substitution of G791M of SEQ ID NO: 2. In some embodiments, a CasX variant protein comprises a substitution of L379R, a substitution of C477K, a substitution of A708K, a deletion of P at position 793 and a substitution of Y797L of SEQ ID NO: 2. In some embodiments, a CasX variant protein comprises a substitution of L379R, a substitution of C477K, a substitution of A708K, a deletion of P at position 793 and a substitution of T620P of SEQ ID NO: 2. In some embodiments, a CasX variant protein comprises a substitution of A708K, a deletion of P at position 793 and a substitution of E386S of SEQ ID NO: 2. In some embodiments, a CasX variant protein comprises a substitution of E386R, a substitution of F399L and a deletion of P at position 793 of SEQ ID NO: 2. In some embodiments, a CasX variant protein comprises a substitution of R581I and A739V of SEQ ID NO: 2. In some embodiments, a CasX variant comprises any combination of the foregoing embodiments of this paragraph.
[00249] In some embodiments, a CasX variant protein comprises more than one substitution, insertion and / or deletion of a reference CasX protein amino acid sequence. In some embodiments, a CasX variant protein comprises a substitution of A708K, a deletion of P at position 793 and a substitution of A739V of SEQ ID NO: 2. In some embodiments, a CasX variant protein comprises a substitution of L379R, a substitution of A708K and a deletion of P at position 793 of SEQ ID NO: 2. In some embodiments, a CasX variant protein comprises a substitution of C477K, a substitution of A708K and a deletion of P at position 793 of SEQ ID NO: 2. In some embodiments, a CasX variant protein comprises a substitution of L379R, a substitution of C477K, a substitution of A708K and a deletion of P at position 793 of SEQ ID NO: 2. In some embodiments, a CasX variant protein comprises a substitution of L379R, a substitution of A708K, a deletion of P at position 793 and a substitution of A739V of SEQ ID NO: 2. In some embodiments, a CasX variant protein comprises a substitution of C477K, a substitution of A708K, a deletion of P at position 793 and a substitution of A739 of SEQ ID NO: 2. In some embodiments, a CasX variant protein comprises a substitution of L379R, a substitution of C477K, a substitution of A708K, a deletion of P at position 793 and a substitution of A739V of SEQ ID NO: 2. In some embodiments, a CasX variant protein comprises a substitution of L379R, a substitution of C477K, a substitution of A708K, a deletion of P at position 793 and a substitution of T620P of SEQ ID NO: 2. In some embodiments, a CasX variant protein comprises a substitution of M771A of SEQ ID NO: 2. In some embodiments, a CasX variant protein comprises a substitution of L379R, a substitution of A708K, a deletion of P at position 793 and a substitution of D732N of SEQ ID NO: 2. In some embodiments, a CasX variant comprises any combination of the foregoing embodiments of this paragraph.
[00250] In some embodiments, a CasX variant protein comprises a substitution of W782Q of SEQ ID NO: 2. In some embodiments, a CasX variant protein comprises a substitution of M771Q of SEQ ID NO: 2. In some embodiments, a CasX variant protein comprises a substitution of R458I and a substitution of A739V of SEQ ID NO: 2. In some embodiments, a CasX variant protein comprises a substitution of L379R, a substitution of A708K, a deletion of P at position 793 and a substitution of M771N of SEQ ID NO: 2. In some embodiments, a CasX variant protein comprises a substitution of L379R, a substitution of A708K, a deletion of P at position 793 and a substitution of A739T of SEQ ID NO: 2. In some embodiments, a CasX variant protein comprises a substitution of L379R, a substitution of C477K, a substitution of A708K, a deletion of P at position 793 and a substitution of D489S of SEQ ID NO: 2. In some embodiments, a CasX variant protein comprises a substitution of L379R, a substitution of C477K, a substitution of A708K, a deletion of P at position 793 and a substitution of D732N of SEQ ID NO: 2. In some embodiments, a CasX variant protein comprises a substitution of V71 IK of SEQ ID NO: 2. In some embodiments, a CasX variant protein comprises a substitution of L379R, a substitution of C477K, a substitution of A708K, a deletion of P at position 793 and a substitution of Y797L of SEQ ID NO: 2. In some embodiments, a CasX variant protein comprises a substitution of L379R, a substitution of A708K and a deletion of P at position 793 of SEQ ID NO: 2. In some embodiments, a CasX variant protein comprises a substitution of L379R, a substitution of C477K, a substitution of A708K, a deletion of P at position 793 and a substitution of M771N of SEQ ID NO: 2. In some embodiments, a CasX variant protein comprises a substitution of A708K, a substitution of P at position 793 and a substitution of E386S of SEQ ID NO: 2. In some embodiments, a CasX variant protein comprises a substitution of L379R, a substitution of C477K, a substitution of A708K and a deletion of P at position 793 of SEQ ID NO: 2. In some embodiments, a CasX variant protein comprises a substitution of L792D of SEQ ID NO: 2. In some embodiments, a CasX variant protein comprises a substitution of G791F of SEQ ID NO: 2. In some embodiments, a CasX variant protein comprises a substitution of A708K, a deletion of P at position 793 and a substitution of A739V of SEQ ID NO: 2. In some embodiments, a CasX variant protein comprises a substitution of L379R, a substitution of A708K, a deletion of P at position 793 and a substitution of A739V of SEQ ID NO: 2. In some embodiments, a CasX variant protein comprises a substitution of C477K, a substitution of A708K and a substitution of P at position 793 of SEQ ID NO: 2. In some embodiments, a CasX variant protein comprises a substitution of L249I and a substitution of M771N of SEQ ID NO: 2. In some embodiments, a CasX variant protein comprises a substitution of V747K of SEQ ID NO: 2. In some embodiments, a CasX variant protein comprises a substitution of L379R, a substitution of C477, a substitution of A708K, a deletion of P at position 793 and a substitution of M779N of SEQ ID NO: 2. In some embodiments, a CasX variant protein comprises a substitution of F755M. In some embodiments, a CasX variant comprises any combination of the foregoing embodiments of this paragraph.
[00251] In some embodiments, a CasX variant protein comprises at least one modification compared to the reference CasX sequence of SEQ ID NO: 2, wherein the at least one modification is selected from one or more of: an amino acid substitution of L379R; an amino acid substitution of A708K; an amino acid substitution of T620P; an amino acid substitution of E385P; an amino acid substitution of Y857R; an amino acid substitution of I658V; an amino acid substitution of F399L; an amino acid substitution of Q252K; an amino acid substitution of L404K; and an amino acid deletion of [P793], In other embodiments, a CasX variant protein comprises any combination of the foregoing substitutions or deletions compared to the reference CasX sequence of SEQ ID NO: 2. In other embodiments, the CasX variant protein can, in addition to the foregoing substitutions or deletions, further comprise a substitution of an NTSB and / or a helical lb domain from the reference CasX of SEQ ID NO: 1.
[00252] In some embodiments, a CasX variant comprises any one of SEQ ID NOS: 247-337, 3301-3493, 3498-3501, 3505-3520, 3540-3549 and 4412-4415. In some embodiments, a CasX variant comprises any one of SEQ ID NOS: 247-337, 3498-3501, 3505-3520, 3540-3549 and 4412-4415. In some embodiments, a CasX variant comprises any one of SEQ ID NOS: 34983501, 3505-3520 and 3540-3549.
[00253] In some embodiments, a CasX variant comprises one or modifications to any one of SEQ ID NOS: 247-337, 3301-3493, 3498-3501, 3505-3520, 3540-3549 and 4412-4415. In some embodiments, a CasX variant comprises one or modifications to any one of SEQ ID NOS: 247337, 3498-3501, 3505-3520, 3540-3549 and 4412-4415. In some embodiments, a CasX variant comprises one or modifications to any one of SEQ ID NOS: 3498-3501, 3505-3520 and 35403549.
[00254] In some embodiments, the CasX variant protein comprises between 400 and 2000 amino acids, between 500 and 1500 amino acids, between 700 and 1200 amino acids, between 800 and 1100 amino acids or between 900 and 1000 amino acids.
[00255] In some embodiments, the CasX variant protein comprises one or more modifications in a region of non-contiguous residues that form a channel in which gNA:target DNA complexing occurs. In some embodiments, the CasX variant protein comprises one or more modifications comprising a region of non-contiguous residues that form an interface which binds with the gNA. For example, in some embodiments of a reference CasX protein, the helical I, helical II and OBD domains all contact or are in proximity to the gNA:target DNA complex, and one or more modifications to non-contiguous residues within any of these domains may improve function of the CasX variant protein.
[00256] In some embodiments, the CasX variant protein comprises one or more modifications in a region of non-contiguous residues that form a channel which binds with the non-target strand DNA. For example, a CasX variant protein can comprise one or more modifications to non-contiguous residues of the NTSBD. In some embodiments, the CasX variant protein comprises one or more modifications in a region of non-contiguous residues that form an interface which binds with the PAM. For example, a CasX variant protein can comprise one or more modifications to non-contiguous residues of the helical I domain or OBD. In some embodiments, the CasX variant protein comprises one or more modifications comprising a region of non-contiguous surface-exposed residues. As used herein, “surface-exposed residues” refers to amino acids on the surface of the CasX protein, or amino acids in which at least a portion of the amino acid, such as the backbone or a part of the side chain is on the surface of the protein. Surface exposed residues of cellular proteins such as CasX, which are exposed to an aqueous intracellular environment, are frequently selected from positively charged hydrophilic amino acids, for example arginine, asparagine, aspartate, glutamine, glutamate, histidine, lysine, serine, and threonine. Thus, for example, in some embodiments of the variants provided herein, a region of surface exposed residues comprises one or more insertions, deletions, or substitutions compared to a reference CasX protein. In some embodiments, one or more positively charged residues are substituted for one or more other positively charged residues, or negatively charged residues, or uncharged residues, or any combinations thereof. In some embodiments, one or more amino acids residues for substitution are near bound nucleic acid, for example residues in the RuvC domain or helical I domain that contact target DNA, or residues in the OBD or helical II domain that bind the gNA, can be substituted for one or more positively charged or polar amino acids.
[00257] In some embodiments, the CasX variant protein comprises one or more modifications in a region of non-contiguous residues that form a core through hydrophobic packing in a domain of the reference CasX protein. Without wishing to be bound by any theory, regions that form cores through hydrophobic packing are rich in hydrophobic amino acids such as valine, isoleucine, leucine, methionine, phenylalanine, tryptophan, and cysteine. For example, in some reference CasX proteins, RuvC domains comprise a hydrophobic pocket adjacent to the active site. In some embodiments, between 2 to 15 residues of the region are charged, polar, or basestacking. Charged amino acids (sometimes referred to herein as residues) may include, for example, arginine, lysine, aspartic acid, and glutamic acid, and the side chains of these amino acids may form salt bridges provided a bridge partner is also present (see FIGS. 14). Polar amino acids may include, for example, glutamine, asparagine, histidine, serine, threonine, tyrosine, and cysteine. Polar amino acids can, in some embodiments, form hydrogen bonds as proton donors or acceptors, depending on the identity of their side chains. As used herein, “base-stacking” includes the interaction of aromatic side chains of an amino acid residue (such as tryptophan, tyrosine, phenylalanine, or histidine) with stacked nucleotide bases in a nucleic acid. Any modification to a region of non-contiguous amino acids that are in close spatial proximity to form a functional part of the CasX variant protein is envisaged as within the scope of the disclosure. i. CasX Variant Proteins with Domains from Multiple Source Proteins
[00258] In certain embodiments, the disclosure provides a chimeric CasX protein comprising protein domains from two or more different CasX proteins, such as two or more naturally occurring CasX proteins, or two or more CasX variant protein sequences as described herein. As used herein, a “chimeric CasX protein” refers to a CasX containing at least two domains isolated or derived from different sources, such as two naturally occurring proteins, which may, in some embodiments, be isolated from different species. For example, in some embodiments, a chimeric CasX protein comprises a first domain from a first CasX protein and a second domain from a second, different CasX protein. In some embodiments, the first domain can be selected from the group consisting of the NTSB, TSL, helical I, helical II, OBD and RuvC domains. In some embodiments, the second domain is selected from the group consisting of the NTSB, TSL, helical I, helical II, OBD and RuvC domains with the second domain being different from the foregoing first domain. For example, a chimeric CasX protein may comprise an NTSB, TSL, helical I, helical II, OBD domains from a CasX protein of SEQ ID NO: 2, and a RuvC domain from a CasX protein of SEQ ID NO: 1, or vice versa. As a further example, a chimeric CasX protein may comprise an NTSB, TSL, helical II, OBD and RuvC domain from CasX protein of SEQ ID NO: 2, and a helical I domain from a CasX protein of SEQ ID NO: 1, or vice versa. Thus, in certain embodiments, a chimeric CasX protein may comprise an NTSB, TSL, helical II, OBD and RuvC domain from a first CasX protein, and a helical I domain from a second CasX protein. In some embodiments of the chimeric CasX proteins, the domains of the first CasX protein are derived from the sequences of SEQ ID NO: 1, SEQ ID NO: 2 or SEQ ID NO: 3, and the domains of the second CasX protein are derived from the sequences of SEQ ID NO: 1, SEQ ID NO: 2 or SEQ ID NO: 3, and the first and second CasX proteins are not the same. In some embodiments, domains of the first CasX protein comprise sequences derived from SEQ ID NO: 1 and domains of the second CasX protein comprise sequences derived from SEQ ID NO: 2. In some embodiments, domains of the first CasX protein comprise sequences derived from SEQ ID NO: 1 and domains of the second CasX protein comprise sequences derived from SEQ ID NO: 3. In some embodiments, domains of the first CasX protein comprise sequences derived from SEQ ID NO: 2 and domains of the second CasX protein comprise sequences derived from SEQ ID NO: 3. In some embodiments, the CasX variant is selected of group consisting of CasX variants with sequences of SEQ ID NO: 328, SEQ ID NO: 3540, SEQ ID NO: 4413, SEQ ID NO: 4414, SEQ ID NO: 4415, SEQ ID NO: 329, SEQ ID NO: 3541, SEQ ID NO: 330, SEQ ID NO: 3542, SEQ ID NO: 331, SEQ ID NO: 3543, SEQ ID NO: 332, SEQ ID NO: 3544, SEQ ID NO: 333, SEQ ID NO: 3545, SEQ ID NO: 334, SEQ ID NO: 3546, SEQ ID NO: 335, SEQ ID NO: 3547, SEQ ID NO: 336 and SEQ ID NO: 3548. In some embodiments, the CasX variant comprises one or more additional modifications to any one of SEQ ID NO: 328, SEQ ID NO: 3540, SEQ ID NO: 4413, SEQ ID NO: 4414, SEQ ID NO: 4415, SEQ ID NO: 329, SEQ ID NO: 3541, SEQ ID NO: 330, SEQ ID NO 332, SEQ ID NO: 3544, SEQ ID NO 3546, SEQ ID NO: 335, SEQ ID NO 3542, SEQ ID NO: 331, SEQ ID NO: 3543, SEQ ID NO: 333, SEQ ID NO: 3545, SEQ ID NO: 334, SEQ ID NO: 3547, SEQ ID NO: 336 or SEQ ID NO: 3548. In some embodiments, the one or more additional modifications comprises an insertion, substitution or deletion as described herein.
[00259] In some embodiments, a CasX variant protein comprises at least one chimeric domain comprising a first part from a first CasX protein and a second part from a second, different CasX protein. As used herein, a “chimeric domain” refers to a domain containing at least two parts isolated or derived from different sources, such as two naturally occurring proteins or portions of domains from two reference CasX proteins. The at least one chimeric domain can be any of the NTSB, TSL, helical I, helical II, OBD or RuvC domains as described herein. In some embodiments, the first portion of a CasX domain comprises a sequence of SEQ ID NO: 1 and the second portion of a CasX domain comprises a sequence of SEQ ID NO: 2. In some embodiments, the first portion of the CasX domain comprises a sequence of SEQ ID NO: 1 and the second portion of the CasX domain comprises a sequence of SEQ ID NO: 3. In some embodiments, the first portion of the CasX domain comprises a sequence of SEQ ID NO: 2 and the second portion of the CasX domain comprises a sequence of SEQ ID NO: 3. In some embodiments, the at least one chimeric domain comprises a chimeric RuvC domain. As an example of the foregoing, the chimeric RuvC domain comprises amino acids 661 to 824 of SEQ ID NO: 1 and amino acids 922 to 978 of SEQ ID NO: 2. As an alternative example of the foregoing, a chimeric RuvC domain comprises amino acids 648 to 812 of SEQ ID NO: 2 and amino acids 935 to 986 of SEQ ID NO: 1. In some embodiments, a CasX protein comprises a first domain from a first CasX protein and a second domain from a second CasX protein, and at least one chimeric domain comprising at least two parts isolated from different CasX proteins using the approach of the embodiments described in this paragraph. In the foregoing embodiments, the chimeric CasX proteins having domains or portions of domains derived from SEQ ID NOS: 1, 2 and 3, can further comprise amino acid insertions, deletions, or substitutions of any of the embodiments disclosed herein.
[00260] In some embodiments, a CasX variant protein comprises a sequence set forth in Tables 3, 8, 9, 10 or 12. In other embodiments, a CasX variant protein comprises a sequence at least 60% identical, at least 65% identical, at least 70% identical, at least 75% identical, at least 80% identical, at least 81% identical, at least 82% identical, at least 83% identical, at least 84% identical, at least 85% identical, at least 86% identical, at least 86% identical, at least 87% identical, at least 88% identical, at least 89% identical, at least 89% identical, at least 90% identical, at least 91% identical, at least 92% identical, at least 93% identical, at least 94% identical, at least 95% identical, at least 96% identical, at least 97% identical, at least 98% identical, at least 99% identical, at least 99.5% identical to a sequence set forth in Tables 3, 8, 9, 10 or 12. In other embodiments, a CasX variant protein comprises a sequence set forth in Table 3, and further comprises one or more NLS disclosed herein on either the N-terminus, the C-terminus, or both. It will be understood that in some cases, the N-terminal methionine of the CasX variants of the Tables is removed from the expressed CasX variant during post-translational modification. Table 3: CasX Variant Sequences Description* Amino Acid Sequence TSL, Helical I, Helical II, OBD and RuvC domains from SEQ ID NO: 2 and an NTSB domain from SEQ ID NO: 1 SEQ ID NO: 247 NTSB, Helical I, Helical II, OBD and RuvC domains from SEQ ID NO: 2 and a TSL domain from SEQ ID NO: 1. SEQ ID NO: 248 TSL, Helical I, Helical II, OBD and RuvC domains from SEQ ID NO: 1 and an NTSB domain from SEQ ID NO: 2 SEQ ID NO: 249 NTSB, Helical I, Helical II, OBD and RuvC domains from SEQ ID NO: 1 and an TSL domain from SEQ ID NO: 2. SEQ ID NO: 250 NTSB, TSL, Helical I, Helical II and OBD domains SEQ ID NO: 2 and an exogenous RuvC domain or a portion thereof from a second CasX protein. SEQ ID NO: 251 No description SEQ ID NO: 252 NTSB, TSL, Helical II, OBD and RuvC domains from SEQ ID NO: 2 and a Helical I domain from SEQ ID NO: 1 SEQ ID NO: 253 NTSB, TSL, Helical I, OBD and RuvC domains from SEQ ID NO: 2 and a Helical II domain from SEQ ID NO: 1 SEQ ID NO: 254 NTSB, TSL, Helical I, Helical II and RuvC domains from a first CasX protein and an exogenous OBD or a part thereof from a second CasX protein SEQ ID NO: 255 No description SEQ ID NO: 256 No description SEQ ID NO: 257 substitution of L379R, a substitution of C477K, a substitution of A708K, a deletion of P at position 793 and a substitution of T620P of SEQ ID NO: 2 SEQ ID NO: 258 Description* Amino Acid Sequence substitution of M771A of SEQ ID NO: 2. SEQ ID NO: 259 substitution of L379R, a substitution of A708K, a deletion of P at position 793 and a substitution of D732N of SEQ ID NO: 2. SEQ ID NO: 260 substitution of W782Q of SEQ ID NO: 2. SEQ ID NO: 261 substitution of M771Q of SEQ ID NO: 2 SEQ ID NO: 262 substitution of R458I and a substitution of A739V of SEQ ID NO: 2. SEQ ID NO: 263 L379R, a substitution of A708K, a deletion of P at position 793 and a substitution of M771N of SEQ ID NO: 2 SEQ ID NO: 264 substitution of L379R, a substitution of A708K, a deletion of P at position 793 and a substitution of A739T of SEQ ID NO: 2 SEQ ID NO: 265 substitution of L379R, a substitution of C477K, a substitution of A708K, a deletion of P at position 793 and a substitution of D489S of SEQ ID NO: 2. SEQ ID NO: 266 substitution of L379R, a substitution of C477K, a substitution of A708K, a deletion of P at position 793 and a substitution of D732N of SEQ ID NO: 2. SEQ ID NO: 267 substitution of V71 IK of SEQ ID NO: 2. SEQ ID NO: 268 substitution of L379R, a substitution of C477K, a substitution of A708K, a deletion of P at position 793 and a substitution of Y797L of SEQ ID NO: 2. SEQ ID NO: 269 119, substitution of L379R, a substitution of A708K and a deletion of P at position 793 of SEQ ID NO: 2. SEQ ID NO: 270 substitution of L379R, a substitution of C477K, a substitution of A708K, a deletion of P at position 793 and a substitution of M771N of SEQ ID NO: 2. SEQ ID NO: 271 substitution of A708K, a deletion of P at position 793 and a substitution of E386S of SEQ ID NO: 2. SEQ ID NO: 272 substitution of L379R, a substitution of C477K, a substitution of A708K and a deletion of P at position 793 of SEQ ID NO: 2. SEQ ID NO: 273 substitution of L792D of SEQ ID NO: 2. SEQ ID NO: 274 substitution of G791F of SEQ ID NO: 2. SEQ ID NO: 275 substitution of A708K, a deletion of P at position 793 and a substitution of A739V of SEQ ID NO: 2. SEQ ID NO: 276 substitution of L379R, a substitution of A708K, a deletion of P at position 793 and a substitution of A739V of SEQ ID NO: 2. (SEQ ID NO: 277 substitution of C477K, a substitution of A708K and a deletion of P at position 793 of SEQ ID NO: 2. SEQ ID NO: 278 substitution of L249I and a substitution of M771N of SEQ ID NO: 2. SEQ ID NO: 279 substitution of V747K of SEQ ID NO: 2. SEQ ID NO: 280 substitution of L379R, a substitution of C477K, a substitution of A708K, a deletion of SEQ ID NO: 281 Description* Amino Acid Sequence P at position 793 and a substitution of M779N of SEQ ID NO: 2. L379R, F755M SEQ ID NO: 282 429, L379R, A708K, P793_, Y857R SEQ ID NO: 283 430, L379R, A708K, P793_, Y857R, I658V SEQ ID NO: 284 431, L379R, A708K, P793_, Y857R, I658V, E386N SEQ ID NO: 285 432, L379R, A708K, P793_, Y857R, I658V, L404K SEQ ID NO: 286 433, L379R, A708K, P793_, Y857R, I658V, AV192 SEQ ID NO: 287 434, L379R, A708K, P793_, Y857R, I658V, L404K, E386N SEQ ID NO: 288 435, L379R, A708K, P793_, Y857R, I658V, F399L SEQ ID NO: 289 436, L379R, A708K, P793_, Y857R, I658V, F399L, E386N SEQ ID NO: 290 437, L379R, A708K, P793_, Y857R, I658V, F399L, C477S SEQ ID NO: 291 438, L379R, A708K, P793_, Y857R, I658V, F399L, L404K SEQ ID NO: 292 439, L379R, A708K, P793_, Y857R, I658V, F399L, E386N, C477S, L404K SEQ ID NO: 293 440, L379R, A708K, P793_, Y857R, I658V, F399L, Y797L SEQ ID NO: 294 441, L379R, A708K, P793_, Y857R, I658V, F399L, Y797L, E386N SEQ ID NO: 295 442, L379R, A708K, P793_, Y857R, I658V, F399L, Y797L, E386N, C477S, L404K SEQ ID NO: 296 443, L379R, A708K, P793_, Y857R, I658V, Y797L SEQ ID NO: 297 444, L379R, A708K, P793_, Y857R, I658V, Y797L, L404K SEQ ID NO: 298 445, L379R, A708K, P793_, Y857R, I658V, Y797L, E386N SEQ ID NO: 299 446, L379R, A708K, P793_, Y857R, I658V, Y797L, E386N, C477S, L404K SEQ ID NO: 300 447, L379R, A708K, P793_, Y857R, E386N SEQ ID NO: 301 448, L379R, A708K, P793_, Y857R, E386N, L404K SEQ ID NO: 302 449, L379R, A708K, P793_, D732N, E385P, Y857R SEQ ID NO: 303 450, L379R, A708K, P793_, D732N, E385P, Y857R, I658V SEQ ID NO: 304 451, L379R, A708K, P793_, D732N, E385P, Y857R, I658V, F399L SEQ ID NO: 305 452, L379R, A708K, P793_, D732N, E385P, Y857R, I658V, E386N SEQ ID NO: 306 453, L379R, A708K, P793_, D732N, E385P, Y857R, I658V, L404K SEQ ID NO: 307 454, L379R, A708K, P793_, T620P, E385P, Y857R, Q252K SEQ ID NO: 308 455, L379R, A708K, P793_, T620P, E385P, Y857R, I658V, Q252K SEQ ID NO: 309 456, L379R, A708K, P793_, T620P, E385P, Y857R, I658V, E386N, Q252K SEQ ID NO: 310 Description* Amino Acid Sequence 457, L379R, A708K, P793_, T620P, E385P, Y857R, I658V, F399L, Q252K SEQ ID NO: 311 458, L379R, A708K, P793_, T620P, E385P, Y857R, I658V, L404K, Q252K SEQ ID NO: 312 459, L379R, A708K, P793_, T620P, Y857R, I658V, E386N SEQ ID NO: 313 460, L379R, A708K, P793_, T620P, E385P, Q252K SEQ ID NO: 314 278 SEQ ID NO: 315 279 SEQ ID NO: 316 280 SEQ ID NO: 317 285 SEQ ID NO: 318 286 SEQ ID NO: 319 287 SEQ ID NO: 320 288 SEQ ID NO: 321 290 SEQ ID NO: 322 291 SEQ ID NO: 323 293 SEQ ID NO: 324 300 SEQ ID NO: 325 492 SEQ ID NO: 326 493 SEQ ID NO: 327 387, NTSB swap from SEQ ID NO: 1 SEQ ID NO: 328 395, Helical IB swap from SEQ ID NO: 1 SEQ ID NO: 329 485, Helical IB swap from SEQ ID NO: 1 SEQ ID NO: 330 486, Helical IB swap from SEQ ID NO: 1 SEQ ID NO: 331 487, Helical IB swap from SEQ ID NO: 1 SEQ ID NO: 332 488, NTSB and Helical IB swap from SEQ ID NO: 1 SEQ ID NO: 333 489, NTSB and Helical IB swap from SEQ ID NO: 1 SEQ ID NO: 334 490, NTSB and Helical IB swap from SEQ ID NO: 1 SEQ ID NO: 335 491, NTSB and Helical IB swap from SEQ ID NO: 1 SEQ ID NO: 336 494, NTSB swap from SEQ ID NO: 1 SEQ ID NO: 337 328, S867G SEQ ID NO: 4412 388, L379R+A708K+ [P793] + XI Helical2 swap SEQ ID NO: 4413 389, L379R+A708K+ [P793] + XI RuvCl swap SEQ ID NO: 4414 Description* Amino Acid Sequence 390, L379R+A708K+ [P793] + XI RuvC2 swap SEQ ID NO: 4415 * Strain indicated numerically; changes, where indicated, are relative to SEQ ID NO: 2
[00261] In some embodiments, the CasX variant protein has one or more improved characteristics when compared to a reference CasX protein, for example a reference protein of SEQ ID NO: 1, SEQ ID NO: 2 or SEQ ID NO: 3. In some embodiments, an improved characteristic of the CasX variant is at least about 1.1 to about 100,000-fold improved relative to the reference protein. In some embodiments, an improved characteristic of the CasX variant is at least about 1.1 to about 10,000-fold improved, at least about 1.1 to about 1,000-fold improved, at least about 1.1 to about 500-fold improved, at least about 1.1 to about 400-fold improved, at least about 1.1 to about 300-fold improved, at least about 1.1 to about 200-fold improved, at least about 1.1 to about 100-fold improved, at least about 1.1 to about 50-fold improved, at least about 1.1 to about 40-fold improved, at least about 1.1 to about 30-fold improved, at least about 1.1 to about 20-fold improved, at least about 1.1 to about 10-fold improved, at least about 1.1 to about 9-fold improved, at least about 1.1 to about 8-fold improved, at least about 1.1 to about 7fold improved, at least about 1.1 to about 6-fold improved, at least about 1.1 to about 5-fold improved, at least about 1.1 to about 4-fold improved, at least about 1.1 to about 3-fold improved, at least about 1.1 to about 2-fold improved, at least about 1.1 to about 1.5-fold improved, at least about 1.5 to about 3-fold improved, at least about 1.5 to about 4-fold improved, at least about 1.5 to about 5-fold improved, at least about 1.5 to about 10-fold improved, at least about 5 to about 10-fold improved, at least about 10 to about 20-fold improved, at least 10 to about 30-fold improved, at least 10 to about 50-fold improved or at least 10 to about 100-fold improved than the reference CasX protein. In some embodiments, an improved characteristic of the CasX variant is at least about 10 to about 1000-fold improved relative to the reference CasX protein.
[00262] In some embodiments, the one or more improved characteristics of the CasX variant protein is at least about 5, at least about 10, at least about 20, at least about 30, at least about 40, at least about 50, at least about 60, at least about 70, at least about 80, at least about 90, at least about 100, at least about 250, at least about 500, or at least about 1000, at least about 5,000, at least about 10,000, or at least about 100,000-fold improved relative to a reference CasX protein. In some embodiments, an improved characteristics of the CasX variant protein is at least about 1.1, at least about 1.2, at least about 1.3, at least about 1.4, at least about 1.5, at least about 1.6, at least about 1.7, at least about 1.8, at least about 1.9, at least about 2, at least about 2.1, at least about 2.2, at least about 2.3, at least about 2.4, at least about 2.5, at least about 2.6, at least about 2.7, at least about 2.8, at least about 2.9, at least about 3, at least about 3.5, at least about 4, at least about 4.5, at least about 5, at least about 5.5, at least about 6, at least about 6.5, at least about 7.0, at least about 7.5, at least about 8, at least about 8.5, at least about 9, at least about 9.5, at least about 10, at least about 11, at least about 12, at least about 13, at least about 14, at least about 15, at least about 20, at least about 30, at least about 40, at least about 50, at least about 60, at least about 70, at least about 80, at least about 90 at least about 100, at least about 500, at least about 1,000, at least about 10,000, or at least about 100,000-fold improved relative to a reference CasX protein. In other cases, the one or more improved characteristics of the CasX variant is about 1.1 to 100,00-fold, about 1.1 to 10,00-fold, about 1.1 to 1,000-fold, about 1.1 to 500-fold, about 1.1 to 100-fold, about 1.1 to 50-fold, about 1.1 to 20-fold, about 10 to 100,00-fold, about 10 to 10,00-fold, about 10 to 1,000-fold, about 10 to 500-fold, about 10 to 100-fold, about 10 to 50-fold, about 10 to 20-fold, about 2 to 70-fold, about 2 to 50-fold, about 2 to 30-fold, about 2 to 20-fold, about 2 to 10-fold, about 5 to 50-fold, about 5 to 30-fold, about 5 to 10-fold, about 100 to 100,00-fold, about 100 to 10,00-fold, about 100 to 1,000-fold, about 100 to 500-fold, about 500 to 100,00-fold, about 500 to 10,00-fold, about 500 to 1,000-fold, about 500 to 750-fold, about 1,000 to 100,00-fold, about 10,000 to 100,00-fold, about 20 to 500-fold, about 20 to 250fold, about 20 to 200-fold, about 20 to 100-fold, about 20 to 50-fold, about 50 to 10,000-fold, about 50 to 1,000-fold, about 50 to 500-fold, about 50 to 200-fold, or about 50 to 100-fold, improved relative to the reference CasX of SEQ ID NO: 1, SEQ ID NO: 2 or SEQ ID NO: 3. In other cases, the one or more improved characteristics of the CasX variant is about 1.1-fold, 1.2-fold, 1.3-fold, 1.4-fold, 1.5-fold, 1.6-fold, 1.7-fold, 1.8-fold, 1.9-fold, 2-fold, 3-fold, 4-fold, 5fold, 6-fold, 7-fold, 8-fold, 9-fold, 10-fold, 11-fold, 12-fold, 13-fold, 14-fold, 15-fold, 16-fold, 17-fold, 18-fold, 19-fold, 20-fold, 25-fold, 30-fold, 40-fold, 45-fold, 50-fold, 55-fold, 60-fold, 70-fold, 80-fold, 90-fold, 100-fold, 110-fold, 120-fold, 130-fold, 140-fold, 150-fold, 160-fold, 170-fold, 180-fold, 190-fold, 200-fold, 210-fold, 220-fold, 230-fold, 240-fold, 250-fold, 260fold, 270-fold, 280-fold, 290-fold, 300-fold, 310-fold, 320-fold, 330-fold, 340-fold, 350-fold, 360-fold, 370-fold, 380-fold, 390-fold, 400-fold, 425-fold, 450-fold, 475-fold, or 500-fold or more improved relative to the reference CasX of SEQ ID NO: 1, SEQ ID NO: 2 or SEQ ID NO: 3. Exemplary characteristics that can be improved in CasX variant proteins relative to the same characteristics in reference CasX proteins include, but are not limited to, improved folding of the variant, improved binding affinity to the gNA, improved binding affinity to the target DNA, improved ability to utilize a greater spectrum of PAM sequences in the editing and / or binding of target DNA, improved unwinding of the target DNA, increased editing activity, improved editing efficiency, improved editing specificity, increased activity of the nuclease, increased target strand loading for double strand cleavage, decreased target strand loading for single strand nicking, decreased off-target cleavage, improved binding of the non-target strand of DNA, improved protein stability, improved CasX:gNA RNA complex stability, improved protein solubility, improved CasX:gNA RNP complex solubility, improved protein yield, improved protein expression, and improved fusion characteristics. In some embodiments, the variant comprises at least one improved characteristic. In other embodiments, the variant comprises at least two improved characteristics. In further embodiments, the variant comprises at least three improved characteristics. In some embodiments, the variant comprises at least four improved characteristics. In still further embodiments, the variant comprises at least five, at least six, at least seven, at least eight, at least nine, at least ten, at least eleven, at least twelve, at least thirteen, or more improved characteristics. These improved characteristics are described in more detail below. j. Protein Stability
[00263] In some embodiments, the disclosure provides a CasX variant protein with improved stability relative to a reference CasX protein. In some embodiments, improved stability of the CasX variant protein results in expression of a higher steady state of protein, which improves editing efficiency. In some embodiments, improved stability of the CasX variant protein results in a larger fraction of CasX protein that remains folded in a functional conformation and improves editing efficiency or improves purifiability for manufacturing purposes. As used herein, a “functional conformation” refers to a CasX protein that is in a conformation where the protein is capable of binding a gNA and target DNA. In embodiments wherein the CasX variant does not carry one or more mutations rendering it catalytically dead, the CasX variant is capable of cleaving, nicking, or otherwise modifying the target DNA. For example, a functional CasX variant can, in some embodiments, be used for gene-editing, and a functional conformation refers to an “editing-competent” conformation. In some exemplary embodiments, including those embodiments where the CasX variant protein results in a larger fraction of CasX protein that remains folded in a functional conformation, a lower concentration of CasX variant is needed for applications such as gene editing compared to a reference CasX protein. Thus, in some embodiments, the CasX variant with improved stability has improved efficiency compared to a reference CasX in one or more gene editing contexts.
[00264] In some embodiments, the disclosure provides a CasX variant protein having improved thermostability relative to a reference CasX protein. In some embodiments, the CasX variant protein has improved thermostability of the CasX variant protein at a particular temperature range. Without wishing to be bound by any theory, some reference CasX proteins natively function in organisms with niches in groundwater and sediment; thus, some reference CasX proteins may have evolved to exhibit optimal function at lower or higher temperatures that may be desirable for certain applications. For example, one application of CasX variant proteins is gene editing of mammalian cells, which is typically carried out at about 37°C. In some embodiments, a CasX variant protein as described herein has improved thermostability compared to a reference CasX protein at a temperature of at least 16°C, at least 18°C, at least 20°C, at least 22°C, at least 24°C, at least 26°C, at least 28°C, at least 30°C, at least 32°C, at least 34°C, at least 35°C, at least 36°C, at least 37°C, at least 38°C, at least 39°C, at least 40°C, at least 41 °C, at least 42°C, at least 44°C, at least 46°C, at least 48°C, at least 50°C, at least 52°C, or greater. In some embodiments, a CasX variant protein has improved thermostability and functionality compared to a reference CasX protein that results in improved gene editing functionality, such as mammalian gene editing applications, which may include human gene editing applications.
[00265] In some embodiments, the disclosure provides a CasX variant protein having improved stability of the CasX variant protein:gNA RNP complex relative to the reference CasX protein:gNA complex such that the RNP remains in a functional form. Stability improvements can include increased thermostability, resistance to proteolytic degradation, enhanced pharmacokinetic properties, stability across a range of pH conditions, salt conditions, and tonicity. Improved stability of the complex may, in some embodiments, lead to improved editing efficiency.
[00266] In some embodiments, the disclosure provides a CasX variant protein having improved thermostability of the CasX variant protein:gNA complex relative to the reference CasX protein:gNA complex. In some embodiments, a CasX variant protein has improved thermostability relative to a reference CasX protein. In some embodiments, the CasX variant protein:gNARNP complex has improved thermostability relative to a complex comprising a reference CasX protein at temperatures of at least 16°C, at least 18°C, at least 20°C, at least 22°C, at least 24°C, at least 26°C, at least 28°C, at least 30°C, at least 32°C, at least 34°C, at least 35°C, at least 36°C, at least 37°C, at least 38°C, at least 39°C, at least 40°C, at least 41°C, at least 42°C, at least 44°C, at least 46°C, at least 48°C, at least 50°C, at least 52°C, or greater. In some embodiments, a CasX variant protein has improved thermostability of the CasX variant protein:gNARNP complex compared to a reference CasX protein:gNA complex, which results in improved function for gene editing applications, such as mammalian gene editing applications, which may include human gene editing applications.
[00267] In some embodiments, the improved stability and / or thermostability of the CasX variant protein comprises faster folding kinetics of the CasX variant protein relative to a reference CasX protein, slower unfolding kinetics of the CasX variant protein relative to a reference CasX protein, a larger free energy release upon folding of the CasX variant protein relative to a reference CasX protein, a higher temperature at which 50% of the CasX variant protein is unfolded (Tm) relative to a reference CasX protein, or any combination thereof. These characteristics may be improved by a wide range of values; for example, at least 1.1, at least 1.5, at least 10, at least 50, at least 100, at least 500, at least 1,000, at least 5,000, or at least a 10,000fold improved, as compared to a reference CasX protein. In some embodiments, improved thermostability of the CasX variant protein comprises a higher Tm of the CasX variant protein relative to a reference CasX protein. In some embodiments, the Tm of the CasX variant protein is between about 20°C to about 30°C, between about 30°C to about 40°C, between about 40°C to about 50°C, between about 50°C to about 60°C, between about 60°C to about 70°C, between about 70°C to about 80°C, between about 80°C to about 90°C or between about 90°C to about 100°C. Thermal stability is determined by measuring the “melting temperature” (Tm), which is defined as the temperature at which half of the molecules are denatured. Methods of measuring characteristics of protein stability such as Tm and the free energy of unfolding are known to persons of ordinary skill in the art, and can be measured using standard biochemical techniques in vitro. For example, Tm may be measured using Differential Scanning Calorimetry, a thermo-analytical technique in which the difference in the amount of heat required to increase the temperature of a sample and a reference is measured as a function of temperature (Chen et al (2003) Pharm Res 20:1952-60; Ghirlando et al (1999) Immunol Lett 68:47-52). Alternatively, or in addition, CasX variant protein Tm may be measured using commercially available methods such as the ThermoFisher Protein Thermal Shift system. Alternatively, or in addition, circular dichroism may be used to measure the kinetics of folding and unfolding, as well as the Tm (Murray et al. (2002) J. Chromatogr Sci 40:343-9). Circular dichroism (CD) relies on the unequal absorption of left-handed and right-handed circularly polarized light by asymmetric molecules such as proteins. Certain structures of proteins, for example alpha-helices and betasheets, have characteristic CD spectra. Accordingly, in some embodiments, CD may be used to determine the secondary structure of a CasX variant protein.
[00268] In some embodiments, improved stability and / or thermostability of the CasX variant protein comprises improved folding kinetics of the CasX variant protein relative to a reference CasX protein. In some embodiments, folding kinetics of the CasX variant protein are improved relative to a reference CasX protein by at least about 5, at least about 10, at least about 50, at least about 100, at least about 500, at least about 1,000, at least about 2,000, at least about 3,000, at least about 4,000, at least about 5,000, or at least about a 10,000-fold improvement. In some embodiments, folding kinetics of the CasX variant protein are improved relative to a reference CasX protein by at least about 1 kJ / mol, at least about 5 kJ / mol, at least about 10 kJ / mol, at least about 20 kJ / mol, at least about 30 kJ / mol, at least about 40 kJ / mol, at least about 50 kJ / mol, at least about 60 kJ / mol, at least about 70 kJ / mol, at least about 80 kJ / mol, at least about 90 kJ / mol, at least about 100 kJ / mol, at least about 150 kJ / mol, at least about 200 kJ / mol, at least about 250 kJ / mol, at least about 300 kJ / mol, at least about 350 kJ / mol, at least about 400 kJ / mol, at least about 450 kJ / mol, or at least about 500 kJ / mol.
[00269] Exemplary amino acid changes that can increase the stability of a CasX variant protein relative to a reference CasX protein may include, but are not limited to, amino acid changes that increase the number of hydrogen bonds within the CasX variant protein, increase the number of disulfide bridges within the CasX variant protein, increase the number of salt bridges within the CasX variant protein, strengthen interactions between parts of the CasX variant protein, increase the buried hydrophobic surface area of the CasX variant protein, or any combinations thereof. k. Protein Yield
[00270] In some embodiments, the disclosure provides a CasX variant protein having improved yield during expression and purification relative to a reference CasX protein. In some embodiments, the yield of CasX variant proteins purified from bacterial or eukaryotic host cells is improved relative to a reference CasX protein. In some embodiments, the bacterial host cells are Escherichia coli cells. In some embodiments, the eukaryotic cells are yeast, plant (e.g. tobacco), insect (e.g. Spodoptera frugiperda sf9 cells), mouse, rat, hamster, guinea pig, non human primate, or human cells. In some embodiments, the eukaryotic host cells are mammalian cells, including, but not limited to HEK293 cells, HEK293T cells, HEK293-F cells, Lenti-X 293T cells, BHK cells, HepG2 cells, Saos-2 cells, HuH7 cells, A549 cells, NSO cells, SP2 / 0 cells, YO myeloma cells, P3X63 mouse myeloma cells, PER cells, PER.C6 cells, hybridoma cells, VERO cells, NIH3T3 cells, COS, WI38 cells, MRC5 cells, HeLa, HT1080 cells, or CHO cells.
[00271] In some embodiments, improved yield of the CasX variant protein is achieved through codon optimization. Cells use 64 different codons, 61 of which encode the 20 standard amino acids, while another 3 function as stop codons. In some cases, a single amino acid is encoded by more than one codon. Different organisms exhibit bias towards use of different codons for the same naturally occurring amino acid. Therefore, the choice of codons in a protein, and matching codon choice to the organism in which the protein will be expressed, can, in some cases, significantly affect protein translation and therefore protein expression levels. In some embodiments, the CasX variant protein is encoded by a nucleic acid that has been codon optimized. In some embodiments, the nucleic acid encoding the CasX variant protein has been codon optimized for expression in a bacterial cell, a yeast cell, an insect cell, a plant cell, or a mammalian cell. In some embodiments, the mammal cell is a mouse, a rat, a hamster, a guinea pig, a monkey, or a human. In some embodiments, the CasX variant protein is encoded by a nucleic acid that has been codon optimized for expression in a human cell. In some embodiments, the CasX variant protein is encoded by a nucleic acid from which nucleotide sequences that reduce translation rates in prokaryotes and eukaryotes have been removed. For example, runs of greater than three thymine residues in a row can reduce translation rates in certain organisms or internal polyadenylation signals can reduce translation.
[00272] In some embodiments, improvements in solubility and stability, as described herein, result in improved yield of the CasX variant protein relative to a reference CasX protein.
[00273] Improved protein yield during expression and purification can be evaluated by methods known in the art. For example, the amount of CasX variant protein can be determined by running the protein on an SDS-page gel, and comparing the CasX variant protein to a control whose amount or concentration is known in advance to determine an absolute level of protein. Alternatively, or in addition, a purified CasX variant protein can be run on an SDS-page gel next to a reference CasX protein undergoing the same purification process to determine relative improvements in CasX variant protein yield. Alternatively, or in addition, levels of protein can be measured using immunohistochemical methods such as Western blot or ELISA with an antibody to CasX, or by HPLC. For proteins in solution, concentration can be determined by measuring of the protein's intrinsic UV absorbance, or by methods which use protein-dependent color changes such as the Lowry assay, the Smith copper / bicinchoninic assay or the Bradford dye assay. Such methods can be used to calculate the total protein (such as, for example, total soluble protein) yield obtained by expression under certain conditions. This can be compared, for example, to the protein yield of a reference CasX protein under similar expression conditions. 1. Protein Solubility
[00274] In some embodiments, a CasX variant protein has improved solubility relative to a reference CasX protein. In some embodiments, a CasX variant protein has improved solubility of the CasX:gNA ribonucleoprotein complex variant relative to a ribonucleoprotein complex comprising a reference CasX protein.
[00275] In some embodiments, an improvement in protein solubility leads to higher yield of protein from protein purification techniques such as purification from E. coll. Improved solubility of CasX variant proteins may, in some embodiments, enable more efficient activity in cells, as a more soluble protein may be less likely to aggregate in cells. Protein aggregates can in certain embodiments be toxic or burdensome on cells, and, without wishing to be bound by any theory, increased solubility of a CasX variant protein may ameliorate this result of protein aggregation. Further, improved solubility of CasX variant proteins may allow for enhanced formulations permitting the delivery of a higher effective dose of functional protein, for example in a desired gene editing application. In some embodiments, improved solubility of a CasX variant protein relative to a reference CasX protein results in improved yield of the CasX variant protein during purification of at least about 5, at least about 10, at least about 20, at least about 30, at least about 40, at least about 50, at least about 60, at least about 70, at least about 80, at least about 90, at least about 100, at least about 250, at least about 500, or at least about 1000fold greater yield. In some embodiments, improved solubility of a CasX variant protein relative to a reference CasX protein improves activity of the CasX variant protein in cells by at least about 1.1, at least about 1.2, at least about 1.3, at least about 1.4, at least about 1.5, at least about 1.6, at least about 1.7, at least about 1.8, at least about 1.9, at least about 2, at least about 2.1, at least about 2.2, at least about 2.3, at least about 2.4, at least about 2.5, at least about 2.6, at least about 2.7, at least about 2.8, at least about 2.9, at least about 3, at least about 3.5, at least about 4, at least about 4.5, at least about 5, at least about 5.5, at least about 6, at least about 6.5, at least about 7.0, at least about 7.5, at least about 8, at least about 8.5, at least about 9, at least about 9.5, at least about 10, at least about 11, at least about 12, at least about 13, at least about 14, at least about 15-fold, or at least about 20-fold greater activity.
[00276] Methods of measuring CasX protein solubility, and improvements thereof in CasX variant proteins, will be readily apparent to the person of ordinary skill in the art. For example, CasX variant protein solubility can in some embodiments be measured by taking densitometry readings on a gel of the soluble fraction of lysed E.coli. Alternatively, or addition, improvements in CasX variant protein solubility can be measured by measuring the maintenance of soluble protein product through the course of a full protein purification, including the methods of the Examples. For example, soluble protein product can be measured at one or more steps of gel affinity purification, tag cleavage, cation exchange purification, running the protein on a size exclusion chromatography (SEC) column. In some embodiments, the densitometry of every band of protein on a gel is read after each step in the purification process. CasX variant proteins with improved solubility may, in some embodiments, maintain a higher concentration at one or more steps in the protein purification process when compared to the reference CasX protein, while an insoluble protein variant may be lost at one or more steps due to buffer exchanges, filtration steps, interactions with a purification column, and the like.
[00277] In some embodiments, improving the solubility of CasX variant proteins results in a higher yield in terms of mg / L of protein during protein purification when compared to a reference CasX protein.
[00278] In some embodiments, improving the solubility of CasX variant proteins enables a greater amount of editing events compared to a less soluble protein when assessed in editing assays such as the EGFP disruption assays described herein. m. Affinity for the gNA
[00279] In some embodiments, a CasX variant protein has improved affinity for the gNA relative to a reference CasX protein, leading to the formation of the ribonucleoprotein complex. Increased affinity of the CasX variant protein for the gNA may, for example, result in a lower Ka for the generation of a RNP complex, which can, in some cases, result in a more stable ribonucleoprotein complex formation. In some embodiments, the Ka of a CasX variant protein for a gNA is increased relative to a reference CasX protein by a factor of at least about 1.1, at least about 1.2, at least about 1.3, at least about 1.4, at least about 1.5, at least about 1.6, at least about 1.7, at least about 1.8, at least about 1.9, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 15, at least about 20, at least about 25, at least about 30, at least about 35, at least about 40, at least about 45, at least about 50, at least about 60, at least about 70, at least about 80, at least about 90, or at least about 100. In some embodiments, the CasX variant has about 1.1 to about 10-fold increased binding affinity to the gNA compared to the reference CasX protein of SEQ ID NO: 2.
[00280] In some embodiments, increased affinity of the CasX variant protein for the gNA results in increased stability of the ribonucleoprotein complex when delivered to mammalian cells, including in vivo delivery to a subject. This increased stability can affect the function and utility of the complex in the cells of a subject, as well as result in improved pharmacokinetic properties in blood, when delivered to a subject. In some embodiments, increased affinity of the CasX variant protein, and the resulting increased stability of the ribonucleoprotein complex, allows for a lower dose of the CasX variant protein to be delivered to the subject or cells while still having the desired activity; for example in vivo or in vitro gene editing. The increased ability to form RNP and keep them in stable form can be assessed using assays such as the in vitro cleavage assays described herein. In some embodiments, the CasX variants of the disclosure are able to achieve a Kcieave rate when complexed as an RNP that is at last 2-fold, at least 5-fold, or at least 10-fold higher compared to RNP of reference CasX.
[00281] In some embodiments, a higher affinity (tighter binding) of a CasX variant protein to a gNA allows for a greater amount of editing events when both the CasX variant protein and the gNA remain in an RNP complex. Increased editing events can be assessed using editing assays such as the EGFP disruption and in vitro cleavage assays described herein.
[00282] Without wishing to be bound by theory, in some embodiments amino acid changes in the helical I domain can increase the binding affinity of the CasX variant protein with the gNA targeting sequence, while changes in the helical II domain can increase the binding affinity of the CasX variant protein with the gNA scaffold stem loop, and changes in the oligonucleotide binding domain (OBD) increase the binding affinity of the CasX variant protein with the gNA triplex.
[00283] Methods of measuring C...
Claims
1. A chimeric CasX variant, comprising a sequence with at least 90% sequence identity to SEQ ID NO: 2, wherein the chimeric CasX variant comprises:a. a substitution of the non-target strand binding (NTSB) domain of SEQ ID NO: 2 with the NTSB domain from SEQ ID NO: 1, or a sequence with at least 90% sequence identity thereto; andb. a substitution of the helical 1b domain of SEQ ID NO: 2 with the helical 1b domain of SEQ ID NO: 1, or a sequence with at least 90% sequence identity thereto.
2. The chimeric CasX variant of claim 1, wherein the chimeric CasX variant has an improved characteristic relative to that of SEQ ID NO: 2, wherein the improved characteristic is selected from the group consisting of: improved folding of the chimeric CasX variant; improved binding affinity to a guide nucleic acid (gNA); improved binding affinity to a target DNA; improved ability to utilize a greater spectrum of one or more PAM sequences, including ATC, CTC, GTC, or TTC, in the editing of target DNA; improved unwinding of the target DNA; increased editing activity; improved editing efficiency; improved editing specificity; increased nuclease activity; increased target strand loading for double strand cleavage; decreased target strand loading for single strand nicking; decreased off-target cleavage; improved binding of nontarget DNA strand; improved protein stability; improved protein solubility; improved protein:gNA complex (RNP) stability; improved protein:gNA complex solubility; improved protein yield; improved protein expression; improved fusion characteristics or a combination thereof.
3. The chimeric CasX variant of claim 1 or claim 2, further comprising one or more nuclear localization signals (NLS), preferably:(a) wherein the one or more NLS are selected from the group of sequences consisting of PKKKRKV (SEQ ID NO: 352), KRPAATKKAGQAKKKK (SEQ ID NO: 353), PAAKRVKLD (SEQ ID NO: 354), RQRRNELKRSP (SEQ ID NO: 355), NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY (SEQ ID NO: 356), RMRIZFKNKGKDTAELRRRRVEVSVELRKAKKDEQILKRRNV (SEQ ID NO: 357), VSRKRPRP (SEQ ID NO: 358), PPKKARED (SEQ ID NO: 35), PQPKKKPL (SEQ ID NO:2020289591 30 Jun 2026360), SALIKKKKKMAP (SEQ ID NO: 361), DRLRR (SEQ ID NO: 362), PKQKKRK (SEQ ID NO: 363), RKLKKKIKKL (SEQ ID NO: 364), REKKKFLKRR (SEQ ID NO: 365), KRKGDEVDGVDEVAKKKSKK (SEQ ID NO: 366), RKCLQAGMNLEARKTKK (SEQ ID NO: 367), PRPRKIPR (SEQ ID NO: 368), PPRKKRTVV (SEQ ID NO: 369), NLSKKKKRKREK (SEQ ID NO: 370), RRPSRPFRKP (SEQ ID NO: 371), KRPRSPSS (SEQ ID NO: 372), KRGINDRNFWRGENERKTR (SEQ ID NO: 373), PRPPKMARYDN (SEQ ID NO: 374), KRSFSKAF (SEQ ID NO: 375), KLKIKRPVK (SEQ ID NO: 376), PKKKRKVPPPPAAKRVKLD (SEQ ID NO: 377), PKTRRRPRRSQRKRPPT (SEQ ID NO: 378), SRRRKANPTKLSENAKKLAKEVEN (SEQ ID NO: 379), KTRRRPRRSQRKRPPT (SEQ ID NO: 380), RRKKRRPRRKKRR (SEQ ID NO: 381), PKKKSRKPKKKSRK (SEQ ID NO: 382), HKKKHPDASVNFSEFSK (SEQ ID NO: 383), QRPGPYDRPQRPGPYDRP (SEQ ID NO: 384), LSPSLSPLLSPSLSPL (SEQ ID NO: 385), RGKGGKGLGKGGAKRHRK (SEQ ID NO: 386), PKRGRGRPKRGRGR (SEQ ID NO: 387), and PKKKRKVPPPPKKKRKV (SEQ ID NO: 389); or(b) comprising a sequence of any one of SEQ ID NOS: 3540-3549, preferably:(i) wherein the one or more NLS are positioned at or near the C-terminus of the chimeric CasX variant;(ii) wherein the one or more NLS are positioned at or near at the N-terminus of the chimeric CasX variant; or(iii) wherein the at least two NLS are positioned at or near the N-terminus and at or near the C-terminus of the chimeric CasX variant.
4. The chimeric CasX variant of any one of claims 1-3, wherein:(a) one or more of the improved characteristics of the chimeric CasX variant is at least about 1.1 to about 100-fold or more improved relative to the reference CasX protein of SEQ ID NO: 2;(b) one or more of the improved characteristics of the chimeric CasX variant is at least about 1.1, at least about 2, at least about 10, at least about 100-fold or more improved relative to the reference CasX protein of SEQ ID NO: 2; and / or2020289591 30 Jun 2026(c) the improved characteristic comprises editing efficiency, and the chimeric CasX variant comprises a 1.1 to 100-fold improvement in editing efficiency compared to the reference CasX protein of SEQ ID NO: 2.
5. The chimeric CasX variant of any one of claims 1-4, wherein the RNP comprising the chimeric CasX variant exhibits greater editing efficiency and / or binding of a target sequence in the target DNA when any one of the PAM sequences TTC, ATC, GTC, or CTC is located 1 nucleotide 5’ to the non-target strand of the protospacer having identity with the targeting sequence of the gNA in a cellular assay system compared to the editing efficiency and / or binding of an RNP comprising a reference CasX protein in a comparable assay system, preferably:(a) wherein the PAM sequence is TTC;(b) wherein the PAM sequence is ATC;(c) wherein the PAM sequence is CTC;(d) wherein the PAM sequence is GTC; or(e) wherein the improved editing efficiency and / or binding to the target DNA of the RNP comprising the chimeric CasX variant is at least about 1.1 to about 100-fold improved relative to the RNP comprising the reference CasX of SEQ ID NO: 2.
6. The chimeric CasX variant of any one of claims 1-5, wherein the chimeric CasX variant protein:(a) comprises between 900 and 1000 amino acids;(b) comprises a nuclease domain having nickase activity;(c) comprises a nuclease domain having double-stranded cleavage activity; or(d) is a catalytically inactive CasX (dCasX) protein, and wherein the dCasX and the gNA retain the ability to bind to the target DNA, preferably wherein the dCasX comprises a mutation at residues:a. D672, and / or E769, and / or D935 corresponding to the CasX protein of SEQ ID NO:1; orb. D659, and / or E756, and / or D922 corresponding to the CasX protein of SEQ ID NO: 2,preferably wherein the mutation is a substitution of alanine for the residue.2020289591 30 Jun 20267. The chimeric CasX variant of any one of claims 1-6, wherein the chimeric CasX variant is SEQ ID NO: 333 or SEQ ID NO: 336.
8. The chimeric CasX variant of any one of claims 1-6, comprising SEQ ID NO: 333 or SEQ ID NO: 336.
9. The chimeric CasX variant of any one of claims 1-8, comprising a heterologous protein or domain thereof fused to the chimeric CasX variant, preferably wherein the heterologous protein or domain thereof is a base editor, preferably wherein the base editor is an adenosine deaminase, a cytosine deaminase or a guanine oxidase.
10. A gene editing pair comprising a chimeric CasX variant of any one of claims 1-9 and a first gNA, preferably wherein:(a) the chimeric CasX variant and the gNA are capable of associating together in a ribonuclear protein complex (RNP); and / or(b) the chimeric CasX variant and the gNA are associated together in a ribonuclear protein complex (RNP).
11. The gene editing pair of claim 10, wherein the gene editing pair of the chimeric CasX variant and the gNA variant has one or more improved characteristics compared to a gene editing pair comprising a reference CasX protein of SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 3 and a reference guide nucleic acid of SEQ ID NOS: 4 or 5, preferably wherein the one or more improved characteristics comprises improved CasX:gNA (RNP) complex stability, improved binding affinity between the chimeric CasX variant and gNA, improved kinetics of RNP complex formation, higher percentage of cleavage-competent RNP, improved RNP binding affinity to a target DNA, ability to utilize an increased spectrum of PAM sequences, improved unwinding of the target DNA, increased editing activity, improved editing efficiency, improved editing specificity, increased nuclease activity, increased target strand loading for double strand cleavage, decreased target strand loading for single strand nicking, decreased off-target cleavage, improved binding of the non-target strand of DNA, or improved resistance to nuclease activity.2020289591 30 Jun 202612. The gene editing pair of claim 11, wherein:(a) the at least one or more of the improved characteristics is at least about 1.1 to about 100-fold or more improved relative to a gene editing pair of the reference CasX protein and the reference guide nucleic acid;(b) wherein one or more of the improved characteristics of the chimeric CasX variant is at least about 1.1, at least about 2, at least about 10, or at least about 100-fold or more improved relative to a gene editing pair of the reference CasX protein and the reference guide nucleic acid; or(c) wherein the improved characteristic comprises a 4 to 9 fold increase in editing activity compared to a reference editing pair of SEQ ID NO: 2 and SEQ ID NO: 5, preferably comprising a chimeric CasX variant of SEQ ID NO: 333 or SEQ ID NO: 336.
13. A composition comprising the gene editing pair of any one of claims 10-12, further comprising a second gene editing pair comprising:a. the chimeric CasX variant of any one of claims 1-9; andb. a second reference guide nucleic acid, wherein the second gNA variant or the second reference guide nucleic acid has a targeting sequence complementary to a different or overlapping portion of the target DNA compared to the targeting sequence of the first gNA.
14. The gene editing pair or composition of any one of claims 10-13, wherein the RNP of the chimeric CasX variant and the gNA variant:(a) has a higher percentage of cleavage-competent RNP compared to an RNP of a reference CasX protein and a reference guide nucleic acid;(b) is capable of binding and cleaving a target DNA;(c) is capable of binding a target DNA but is not capable of cleaving the target DNA; or(d) is capable of binding a target DNA and generating one or more single-stranded nicks in the target DNA.
15. A method of editing a target DNA in vitro outside of a cell, in vitro inside of a cell or ex vivo, comprising contacting the target DNA with a gene editing pair or composition of any one2020289591 30 Jun 2026of claims 10-14, wherein the contacting results in editing or modification of the target DNA, preferably:(a) comprising contacting the target DNA with a plurality of gNAs comprising targeting sequences complementary to different or overlapping regions of the target DNA;(b) wherein the contacting by the gene editing pair comprises binding the target DNA and results in introducing a mutation, an insertion, or a deletion in the target DNA;(c) wherein the contacting introduces one or more single-stranded breaks in the target DNA and wherein the editing comprises introducing a mutation, an insertion, or a deletion in the target DNA;(d) wherein the contacting comprises introducing one or more double-stranded breaks in the target DNA and wherein the editing comprises introducing a mutation, an insertion, or a deletion in the target DNA; and / or(e) further comprising contacting the target DNA with a nucleotide sequence of a donor template nucleic acid wherein the donor template comprises a nucleotide sequence having homology to the target DNA, preferably:(i) wherein the donor template comprises homologous arms on the 5’ and 3’ ends of the donor template;(ii) wherein the donor template is inserted in the target DNA at the break site by homology-directed repair; or(iii) wherein the donor template is inserted in the target DNA at the break site by non-homologous end joining (NHEJ) or micro-homology end joining (MMEJ).
16. The method of claim 15, wherein the cell is a eukaryotic cell, preferably wherein the eukaryotic cell is selected from the group consisting of a plant cell, a fungal cell, a protist cell, a mammalian cell, a reptile cell, an insect cell, an avian cell, a fish cell, a parasite cell, an arthropod cell, a cell of an invertebrate, a cell of a vertebrate, a rodent cell, a mouse cell, a rat cell, a primate cell, and a non-human primate cell, preferably wherein the eukaryotic cell is a human cell, preferably wherein the cell is an embryonic stem cell, an induced pluripotent stem cell, a germ cell, a fibroblast, an oligodendrocyte, a glial cell, a hematopoietic stem cell, a neuron progenitor cell, a neuron, a muscle cell, a bone cell, a hepatocyte, a pancreatic cell, a retinal cell, a cancer cell, a T-cell, a B-cell, an NK cell, a fetal cardiomyocyte, a myofibroblast, a mesenchymal stem cell, an autotransplated expanded cardiomyocyte, an adipocyte, a totipotent2020289591 30 Jun 2026cell, a pluripotent cell, a blood stem cell, a myoblast, an adult stem cell, a bone marrow cell, a mesenchymal cell, a parenchymal cell, an epithelial cell, an endothelial cell, a mesothelial cell, fibroblasts, osteoblasts, chondrocytes, exogenous cell, endogenous cell, stem cell, hematopoietic stem cell, bone-marrow derived progenitor cell, myocardial cell, skeletal cell, fetal cell, undifferentiated cell, multi-potent progenitor cell, unipotent progenitor cell, a monocyte, a cardiac myoblast, a skeletal myoblast, a macrophage, a capillary endothelial cell, a xenogenic cell, an allogenic cell, or a post-natal stem cell.
17. The method of claim 16, wherein greater editing of a target sequence in the target DNA is achieved in a cellular assay system comprising an RNP comprising the chimeric CasX variant when any one of the PAM sequences TTC, ATC, GTC, or CTC is located 1 nucleotide 5’ to the non-target strand of the protospacer having identity with the targeting sequence of the gNA in a cellular assay system, compared to the editing efficiency of an RNP comprising a reference CasX protein in a comparable assay system.
18. The method of claim 16 or 17, wherein the method comprises contacting the eukaryoticcell with a vector encoding or comprising the chimeric CasX variant and the gNA, and optionally further comprising the donor template, preferably wherein the vector; (a) is an Adeno-Associated Viral (AAV) vector, preferably AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV-Rh74, or AAVRh10;(b) is a lentiviral vector;(c) is a non-viral particle;(d) is a virus-like particle (VLP),preferably wherein the vector is administered at a dose of at least about 1 x 109 vector genomes (vg), at least about 1 x 1010 vg, at least about 1 x 1011 vg, at least about 1 x 1012 vg, at least about 1 x 1013 vg, at least about 1 x 1014 vg, at least about 1 x 1015 vg, or at least about 1 x 1016 vg.
19. A cell comprising a target DNA edited by the gene editing pair or composition of any one of claims 10-14.
20. A cell edited by the method of any one of claims 15-18, preferably wherein the cell:2020289591 30 Jun 2026(a) is a prokaryotic cell;(b) is a eukaryotic cell, preferably wherein the eukaryotic cell:(i) is selected from the group consisting of a plant cell, a fungal cell, a protist cell, a mammalian cell, a reptile cell, an insect cell, an avian cell, a fish cell, a parasite cell, an arthropod cell, a cell of an invertebrate, a cell of a vertebrate, a rodent cell, a mouse cell, a rat cell, a primate cell, and a non-human primate; or(ii) is a human cell.
21. A polynucleotide encoding the chimeric CasX variant of any one of claims 1-9.
22. A vector comprising:(a) the polynucleotide of claim 21; or(b) encoding the chimeric CasX variant of any one of claims 1-9,preferably wherein the vector:(i) is an Adeno-Associated Viral (AAV) vector, preferablyAAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV-Rh74, or AAVRh10;(ii) is a lentiviral vector.;(iii) is a virus-like particle (VLP);(iv) is a non-viral particle.
23. A cell comprising the polynucleotide of claim 21, or the vector of claim 22.
24. A composition, comprising the chimeric CasX variant of any one of claims 1-9,preferably:(a) wherein the chimeric CasX variant and the gNA are associated together in a ribonuclear protein complex (RNP);(b) further comprising a donor template nucleic acid wherein the donor template comprises a nucleotide sequence having homology to a target DNA; and / or(c) further comprising a buffer, a nuclease inhibitor, a protease inhibitor, a liposome, a therapeutic agent, a label, a label visualization reagent, or any combination of the foregoing.2020289591 30 Jun 202625. A kit, comprising the chimeric CasX variant of any one of claims 1-9 and a container, preferably further comprising a donor template nucleic acid wherein the donor template comprises a nucleotide sequence having homology to a target sequence of a target DNA, preferably further comprising a buffer, a nuclease inhibitor, a protease inhibitor, a liposome, a therapeutic agent, a label, a label visualization reagent, or any combination of the foregoing.
26. A chimeric CasX variant containing at least two domains from difference sources,wherein the CasX variant comprises any one of the sequences SEQ ID NOS: 333-336 listed in Table 3, or a sequence with at least 90% sequence identity thereto.
27. A gene editing pair, or composition, comprising the gene editing pair or composition of any one of claims 10-14, or a vector of claim 22, for use:(a) as a medicament; or(b) in a method of treatment, wherein the method comprises editing or modifying a target DNA; optionally wherein the editing occurs in a subject having a mutation in an allele of a gene wherein the mutation causes a disease or disorder in the subject, preferably wherein the editing changes the mutation to a wild type allele of the gene or knocks down or knocks out an allele of a gene causing a disease or disorder in the subject.
Citation Information
Patent Citations
WO2018064371A1