Engineered casx system
Modified CasX proteins and gNAs with optimized domains and scaffold regions significantly enhance gene editing efficiency, addressing the need for improved Class 2 CRISPR/Cas systems in therapeutic and research applications.
Patent Information
- Application Number
- JP2025147107
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2020-05-27
- Filing Date
- 2025-09-04
- Publication Date
- 2025-12-16
AI Technical Summary
There is a need for additional Class 2 CRISPR/Cas systems that are optimized for various therapeutic, diagnostic, and research applications, offering improvements over existing systems.
Development of variant CasX nuclease proteins and guide nucleic acids (gNAs) with specific modifications to enhance binding and cleavage efficiency, including domains such as non-target strand binding, target strand loading, helical domains, oligonucleotide binding, and RuvC DNA cleavage, along with optimized scaffold regions.
The modified CasX proteins and gNAs exhibit improved gene editing capabilities, achieving up to 81-fold enhancement in editing efficiency compared to reference systems.
Smart Images

Figure 2025183286000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application Nos. 62,858,750, filed June 7, 2019, 62 / 944,892, filed December 6, 2019, and 63 / 030,838, filed May 27, 2020, the contents of each of which are incorporated herein by reference in their entirety.
[0002] Incorporation by reference of sequence listing This application contains a Sequence Listing that has been submitted in ASCII format via EFS-WEB and is incorporated herein by reference in its entirety. The ASCII copy, created on June 5, 2020, is named SCRB_011_03WO_SeqList_25 and is 3.63 MB in size. [Background technology]
[0003] CRISPR-Cas systems confer adaptive immunity to bacteria and archaea against phages and viruses. Intensive research over the past decade has clarified the biochemistry of these systems. CRISPR-Cas systems consist of a Cas protein, responsible for acquiring, targeting, and cleaving foreign DNA or RNA, and a CRISPR array containing tandem repeats flanked by short spacer sequences that guide the Cas protein to its target. Class 2 CRISPR-Cas is a streamlined version in which a single RNA-bound Cas protein is responsible for binding and cleaving the targeted sequence. The programmable nature of these minimal systems has facilitated their use as a versatile technology that is revolutionizing the field of genome engineering. Summary of the Invention [Problem to be solved by the invention]
[0004] To date, only a few widely used Class 2 CRISPR / Cas systems have been discovered. Thus, there is a need in the art for additional Class 2 CRISPR / Cas systems (e.g., combinations of Cas proteins and guide RNAs) that are optimized for use in a variety of therapeutic, diagnostic, and research applications and / or offer improvements over previous generation systems. [Means for solving the problem]
[0005] In some aspects, the disclosure provides a variant of a reference CasX nuclease protein, wherein the CasX variant is capable of forming a complex with a guide nucleic acid (NA), wherein the complex is capable of binding to a target DNA, wherein the target DNA comprises a non-target strand and a target strand, and wherein the CasX variant comprises at least one modification to a domain of the reference CasX, and wherein the variant exhibits one or more improved characteristics compared to the reference CasX protein. The domains of the reference CasX protein include: (a) a non-target strand binding (NTSB) domain that binds to the non-target strand of DNA, the NTSB domain comprising a four-stranded beta sheet; (b) a target strand loading (TSL) domain that positions the target DNA at the cleavage site of the CasX variant, the TSL domain comprising three positively charged amino acids, which bind to the target strand of DNA; (c) a helical I domain that interacts with both the target DNA and the spacer region of the guide NA, the helical I domain comprising one or more alpha helices; (d) a helical II domain that interacts with both the target DNA and the scaffold stem of the guide NA; (e) an oligonucleotide binding domain (OBD) that binds to the triplex region of the guide NA; and (f) a RuvC DNA cleavage domain.
[0006] In some aspects, the present disclosure provides a variant of a reference guide nucleic acid (gNA) capable of binding to a CasX protein, wherein the reference guide nucleic acid comprises at least one modification in a region compared to the reference guide nucleic acid sequence, and the variant exhibits one or more improved characteristics compared to the reference guide RNA. The scaffold region of the gNA comprises (a) an extended stem-loop, (b) a scaffold stem-loop, (c) a triplex, and (d) a pseudoknot. In some cases, the scaffold stem of the variant gNA further comprises a bubble. In other cases, the scaffold of the variant gNA further comprises a triplex loop region. In other cases, the scaffold of the variant gNA further comprises a 5' unstructured region.
[0007] In some aspects, the present disclosure provides a gene editing pair comprising a CasX protein and a gNA of any of the embodiments described herein.
[0008] In some aspects, the present disclosure provides the polynucleotide and vector that encodes the CasX protein, gNA and gene editing pair described herein.In some embodiments, the vector is a viral vector, such as adeno-associated virus (AAV) vector or lentivirus vector.In other embodiments, the vector is a non-viral particle, such as virus-like particle or nanoparticle.
[0009] In some aspects, the present disclosure provides cells comprising the polynucleotides, vectors, CasX proteins, gNAs, and gene editing pairs described herein. In other aspects, the present disclosure provides cells comprising target DNA edited by embodiments of the editing methods described herein.
[0010] In some aspects, the present disclosure provides kits comprising the polynucleotides, vectors, CasX proteins, gNAs and gene editing pairs described herein.
[0011] In some aspects, the present disclosure provides methods of editing target DNA, comprising contacting the target DNA with one or more of the gene editing pairs described herein, wherein the contacting results in editing of the target DNA.
[0012] In other aspects, the present disclosure provides methods of treating a subject in need thereof comprising administering a gene editing pair of any of the embodiments described herein or a vector comprising or encoding the gene editing pair.
[0013] In another aspect, provided herein is a gene editing pair, a composition comprising a gene editing pair, or a vector comprising or encoding a gene editing pair, for use as a pharmaceutical.
[0014] In another aspect, provided herein is a gene editing pair, a composition comprising a gene editing pair, or a vector comprising or encoding a gene editing pair, for use in a method of treatment, the method comprising editing or modifying target DNA, optionally wherein the editing is performed in a subject who has a mutation in an allele of a gene, the mutation causing a disease or disorder in the subject, and preferably wherein the editing changes the mutation to a wild-type allele of the gene, or knocks down or knocks out the allele of the gene that causes the disease or disorder in the subject. [Brief explanation of the drawings]
[0015] The novel features of the invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention are utilized, and the accompanying drawings of which:
[0016] [Figure 1]
[0023] Figure 1 shows an exemplary method for generating CasX protein and guide RNA variants of the present disclosure using Deep Mutational Evolution (DME). In some exemplary embodiments, DME constructs and tests nearly all possible mutations, insertions, and deletions in biomolecules and their combinations / multiples, providing a nearly comprehensive and unbiased assessment of the biomolecular fitness landscape and pathways in sequence space toward a desired outcome. As described herein, DME can be applied to both CasX proteins and guide RNAs.
[0017] [Figure 2]
[0023] Figure 1 shows an exemplary method for assaying the efficacy of a reference CasX protein or single guide RNA (sgRNA), or their variants, and an exemplary fluorescence-activated cell sorting (FACS) plot. A reporter (e.g., a GFP reporter) linked to a gRNA target sequence complementary to the gRNA spacer is incorporated into a reporter cell line. Cells are transformed or transfected with a CasX protein and / or sgRNA variant bearing a spacer motif of the sgRNA that is complementary to and targets the reporter gRNA target sequence. The ability of the CasX:sgRNA ribonucleoprotein complex to cleave the target sequence is assayed by FACS. Cells that lose reporter expression indicate the occurrence of CasX:sgRNA ribonucleoprotein complex-mediated cleavage and indel formation.
[0018] [Figure 3-1]3A and 3B are heat maps showing exemplary DME mutagenesis results for a reference sgRNA encoded by SEQ ID NO:5, as described in Example 3. Figure 3A shows the effects of single base pair (single base) substitutions, double base pair (double base) substitutions, single base pair insertions, single base pair deletions, and single base pair deletions, as well as single base pair substitutions, at each position in the reference sgRNA shown at the top. Figure 3B shows the effects of double base pair insertions and single base pair insertions, as well as single base pair substitutions, at each position in the improved reference sgRNA. The reference sgRNA sequence of SEQ ID NO:5 is shown at the top of Figure 3A and at the bottom of Figure 3B. In Figures 3A and 3B, the Log2-fold enrichment of variants in the DME library compared to the reference sgRNA after selection is shown in grayscale. Enrichment is a proxy for activity; greater enrichment indicates more active molecules. The results indicate regions of the reference sgRNA that should not be mutated and important regions that should be targeted for mutagenesis. [Figure 3-2] Same as above. [Figure 3-3] Same as above. [Figure 3-4] Same as above.
[0019] [Figure 4-1]
[0023] Figure 1 shows the results of an exemplary DME experiment using a reference sgRNA, as described in Example 3. An improved reference sgRNA (sgRNA) having the sequence of SEQ ID NO: 5 is shown at the top, and the Log2-fold enrichment of variants in the DME library compared to the reference sgRNA after selection is shown in grayscale. Enrichment is a proxy for activity; greater enrichment indicates more active molecules. The heatmap shows an exemplary DME experiment showing four replicates of a library in which every base pair in the reference sgRNA has been replaced with every possible alternative base pair. [Figure 4-2] Same as above.
[0020] [Figure 4-3]A series of eight plots comparing biological replicates of different DME libraries. The Log2-fold enrichment of individual variants relative to the reference sgRNA sequence for pairs of DME replicates is plotted against each other. Plots are shown for single deletion, single insertion, and single substitution DME experiments, as well as the wild-type control, demonstrating a significant amount of concordance for each replicate.
[0021] [Figure 4-4] 1 is a heatmap of an exemplary DME experiment showing four replicates of a library in which all positions within the reference sgRNA received a single base pair insertion. The DME experiment was performed as described in Example 3 using the reference sgRNA of SEQ ID NO: 5 (top). The Log2-fold enrichment of variants in the DME library compared to the reference sgRNA after selection is shown in grayscale. [Figure 4-5] Same as above.
[0022] [Figure 5-1]Figure 5A is a series of plots showing that sgNA variants can improve gene editing by more than two-fold in an EGFP disruption assay, as described in Examples 2 and 3. Editing was measured by indel formation and GFP disruption in HEK293 cells with a GFP reporter. Figure 5A shows the fold change in editing efficiency of the CasX sgRNA reference of SEQ ID NO: 4 and variants of that reference having the sequence of SEQ ID NO: 5 across 10 targets. Averaged across 10 targets, the editing efficiency of sgRNA SEQ ID NO: 5 was improved by 176% compared to SEQ ID NO: 4. Figure 5B shows that further improvement of the sgRNA scaffold of SEQ ID NO: 5 is possible by replacing the extended stem-loop sequence with an additional sequence to generate a scaffold whose sequence is shown in Table 2. The fold change in editing efficiency is shown on the Y-axis. Figure 5C is a plot showing the fold improvement of sgNA variants (including the variant having SEQ ID NO: 17) generated by DME mutations normalized to SEQ ID NO: 5 as the CasX reference sgRNA. Figure 5D is a plot showing the fold improvement of sgNA variants of the sequences shown in Table 2 generated by appending ribozyme sequences to reference sgRNA sequences, normalized to SEQ ID NO: 5 as the CasX reference sgRNA. Figure 5E is a plot showing the fold improvement, normalized to the reference sgRNA of SEQ ID NO: 5, of variants created by combining (overlapping) scaffold stem mutations that show improved cleavage, DME mutations that show improved cleavage, and using ribozyme appendages that show improved cleavage. The resulting sgNA variants provide more than a two-fold improvement in cleavage compared to SEQ ID NO: 5 in this assay. EGFP editing assays were performed using spacer target sequences of E6 and E7. [Figure 5-2] Same as above. [Figure 5-3] Same as above. [Figure 5-4] Same as above. [Figure 5-5] Same as above.
[0023] [Figure 6]1 shows the hepatitis delta virus (HDV) genomic ribozyme used in exemplary gNA variants (SEQ ID NOs: 18-22).
[0024] [Figure 7-1] 7A-7D show the effects of single amino acid substitutions, insertions, and deletions at each amino acid position in the reference CasX protein of SEQ ID NO:2, as described in Example 4. Data were generated by a DME assay performed at 37°C. The Y-axis indicates each possible substitution or insertion (from top to bottom: R, H, K, D, E, S, T, N, Q, C, G, P, A, I, L, M, F, W, Y, or V; boxes indicate amino acid identity to the reference protein), and the X-axis indicates the amino acid position in the reference CasX protein. Log2-fold enrichment of CasX variant proteins compared to the reference CasX protein of SEQ ID NO:2 in the enriched DME library is shown. As used herein, "enrichment" is a proxy for activity; greater enrichment indicates more active molecules. (*) indicates the active site. Figures 7A-7D show the effects of single amino acid substitutions. Figures 7E-7H show the effects of single amino acid insertions. Figure 7I shows the effects of single amino acid deletions. [Figure 7-2] Same as above. [Figure 7-3] Same as above. [Figure 7-4] Same as above. [Figure 7-5] Same as above. [Figure 7-6] Same as above. [Figure 7-7] Same as above. [Figure 7-8] Same as above. [Figure 7-9] Same as above.
[0025] [Figure 8-1]Figure 8 is a series of heat maps showing the effects of single amino acid substitutions, insertions, and deletions at each amino acid position in the reference CasX protein of SEQ ID NO:2, as described in Example 4. Data were generated by a DME assay performed at 45°C. Figure 8A shows the effect of a single amino acid substitution. Figure 8B shows the effect of a single amino acid insertion. Figure 8C shows the effect of a single amino acid deletion. For all Figures 8A-8C, the Y-axis indicates each possible substitution or insertion (from top to bottom: R, H, K, D, E, S, T, N, Q, C, G, P, A, I, L, M, F, W, Y, or V; boxes indicate amino acid identity to the reference protein), and the X-axis indicates the amino acid position in the reference CasX protein. The Log2-fold enrichment of the CasX variant proteins relative to the reference CasX protein of SEQ ID NO:2 in the enriched DME library is shown in grayscale, with greater enrichment indicating more active molecules. (*) indicates the active site. Running this assay at 45°C enriched for different variants compared to running the same assay at 37°C (see Figures 7A-7I), thereby indicating that amino acid residues and changes are important for thermal stability and folding. [Figure 8-2] Same as above. [Figure 8-3] Same as above.
[0026] [Figure 9] Figure 2 shows a comprehensive mutational landscape survey of all single mutations in the reference CasX protein of SEQ ID NO: 2. The Y-axis is the fold enrichment of the CasX variant compared to the reference CasX protein for single substitutions (top), insertions (middle), or deletions (bottom). The X-axis is the amino acid position of the reference CasX protein. Key regions that yield improved CasX variants include the first helix region and the region in the RuvC domain adjacent to the target strand loading (TLS) domain.
[0027] [Figure 10-1]1 is a plot showing that the evaluated CasX variant proteins improved editing by more than three-fold compared to a reference CasX protein in an EGFP disruption assay, as described in Example 5. The CasX proteins were tested for their ability to cleave an EGFP reporter at two different target sites in human HEK293 cells, and the normalized improvement in genome editing at these sites relative to the basic reference CasX protein of SEQ ID NO:2 is shown. From left to right, the variants (indicated by amino acid substitution, insertion, or deletion at the given residue number) are Y789T, [P793], Y789D, T72S, I546V, E552A, A636D, F536S, A708K, Y797L, L792G, A739V, G791M, ^G661, A788W, K390R, A751S, E385A, ^P696, ^M773, G695H, ^AS793, ^AS795, C477R, C477K, C477R, C477K. 79A, C479L, I55F, K210R, C233S, D231N, Q338E, Q338R, L379R, K390R, L481Q, F495S, D600N, T886K, A739V, K460N, I199F, G492P, T153I, R591I, ^AS795, ^AS796, ^L889, E121D, S270W, E712Q, K942Q, E552K, K25Q, N47D, ^T696, L685I, N880D, Q10 2R, M734K, A724S, T704K, P224K, K25R, M29E, H152D, S219R, E475K, G226R, A377K, E480K, K416E, H164R, K767R, I7F, M29R , H435R, E385Q, E385K, I279F, D489S, D732N, A739T, W885R, E53K, A238T, P283Q, E292K, Q628E, R388Q, G791M, L792K, L79 2E, M779N, G27D, K955R, S867R, R693I, F189Y, V635M, F399L, E498K, E386S, V254G, P793S, K188E, QT945KI, T620P, T946P , TT949PP, N952T, K682E, K975R, L212P, E292R, I303K, C349E, E385P, E386N, D387K, L404K, E466H, C477Q, C477H, C479A,D659H, T806V, K808S, ^AS797, V959M, K975Q, W974G, A708Q, V711K, D733T, L742W, V747K, F755M, M771A, M771Q, W78 2Q, G791F, L792D, L792K, P793Q, P793G, Q804A, Y966N, Y723N, Y857R, S890R, S932M, L897M, R624G, S603G, N737S, L3 07K, I658V^PT688, ^SA794, S877R, N580T, V335G, T620S, W345G, T280S, L406P, A612D, A751S, E386R, V351M, K210N, D40A, E773G, H207L, T62A, T287P, T832A, A893S, ^V14, ^AG13, R11V, R12N, R13H, ^Y13, R12L, ^Q13, V15S, ^D17. ^ indicates an insertion, and [] indicates a deletion. [Figure 10-2] Same as above. [Figure 10-3] Same as above.
[0028] [Figure 11]1 is a plot showing that individual beneficial mutations can be combined (sometimes referred to as "stacked") for even greater improvement in gene editing activity. CasX proteins were tested for their ability to cleave at two different target sites in human HEK293 cells using E6 and E7 spacers targeting an EGFP reporter, as described in Example 5. From left to right, the variants are S794R+Y797L, K416E+A708K, A708K+[P793], [P793]+P793AS, Q367K+I425S, A708K+[P793]+A793V, Q338R+A339E, Q338R+A339K, S507G+G508R, L379R+A708K+[P793], C477K+A708K+[P793], L379R+ C477K+A708K+[P793], L379R+A708K+[P793]+A739V, C477K+A708K+[P793]+A739V, L379R+C477K+A708K+[P793]+A7 39V, L379R+A708K+[P793]+M779N, L379R+A708K+[P793]+M771N, L379R+A708K+[P793]+D489S, L379R+A708K+[P793] +A739T, L379R+A708K+[P793]+D732N, L379R+A708K+[P793]+G791M, L379R+A708K+[P793]+Y797L, L379R+C477K+A7 08K+[P793]+M779N, L379R+C477K+A708K+[P793]+M771N, L379R+C477K+A708K+[P793]+D489S, L379R+C477K+A708K+ [P793]+A739T, L379R+C477K+A708K+[P793]+D732N, L379R+C477K+A708K+[P793]+G791M, L379R+C477K+A708K+[P793]+Y797L, L379R+C477K+A708K+[P793]+T620P, A708K+[P793]+E386S, E386R+F399L+[P793] and R4581I+A739V. [ ] indicates the deleted amino acid residue at the specified position of SEQ ID NO:2.
[0029] [Figure 12-1] Figure 12A shows a pair of plots demonstrating that combining a CasX protein and an sgNA variant can improve activity by more than six-fold compared to a reference sgRNA and reference CasX protein pair. The sgNA:protein pairs were assayed for their ability to cleave a GFP reporter in HEK293 cells as described in Example 5. The Y-axis shows the percentage of cells in which GFP reporter expression was inhibited by CasX-mediated gene editing. Figure 12A shows CasX protein and an sgNA assayed using an E6 spacer targeted to GFP. Figure 12B shows CasX protein and an sgNA assayed using an E7 spacer targeted to GFP. iGFP stands for "inducible GFP." [Figure 12-2] Same as above.
[0030] [Figure 13-1]These results demonstrate that the generation and screening of DME libraries, as described in Examples 1 and 3, enabled the generation and identification of variants exhibiting 1- to 81-fold improvements in editing efficiency. Figure 13A shows RFP+ and GFP+ reporters in E. coli cells assayed for CRISPR-interfered suppression of GFP using a reference nuclease-inactive CasX protein and sgNA. Figure 13B shows the same reporter cells assayed for GFP suppression using nuclease-inactive CasX variants screened from the DME library. Figure 13C shows the improved editing efficiency of selected CasX protein and sgNA variants compared to a reference with five spacers targeting the endogenous B2M locus in HEK293 human cells. The Y-axis shows inhibition of B2M staining by HLA1 antibodies, indicating gene disruption due to CasX editing and indel formation. The improved CasX variants improved editing of this locus by up to 81-fold over the reference for guide spacer #43. Shown are CasX pairs with the reference sgRNA:protein pairs of SEQ ID NO:5 and SEQ ID NO:2 assayed with an sgRNA variant with a shortened stem-loop and T10C substitution, encoded by the sequence TACTGGCGCCTTTATCTCATTACTTTGAGAGCCATCACCAGCGACTATGTCGTATGGGTAAAGCGCTTACGGACTTCGGTCCGTAAGAAGCATCAAAG (SEQ ID NO:23), and the L379R+A708K+[P793] CasX variant protein of SEQ ID NO:2. The following spacer sequences were used: #9: GTGTAGTACAAGAGATAGAA (SEQ ID NO:24), #14: TGAAGCTGACAGCATTCGGG (SEQ ID NO:25), #20: tagATCGAGACATGTAAGCA (SEQ ID NO:26), #37: GGCCGAGATGTCTCGCTCCG (SEQ ID NO:27), and #43: AGGCCAGAAAGAGAGAGTAG (SEQ ID NO:28). [Figure 13-2] Same as above. [Figure 13-3] Same as above.
[0031] [Figure 14-1] A series of structural models of a prototype CasX protein showing the location of mutations in the disclosed CasX variant proteins, demonstrating improved activity. Figure 14A shows the deletion of P at 793 of SEQ ID NO:2, with a deletion in a loop that may affect folding. Figure 14B shows the substitution of alanine (A) with lysine (K) at position 708 of SEQ ID NO:2. This mutation surfaces a salt bridge to the gNA in addition to the gNA 5' end. Figure 14C shows the substitution of cysteine (C) with lysine (K) at position 477 of SEQ ID NO:2. This mutation surfaces the gNA. Approximately 14 bases of the salt bridge to the gNAbb (gNA phosphatase backbone) are affected. This mutation removes a surface-exposed cysteine. Figure 14D shows the substitution of leucine (L) with arginine (R) at position 379 of SEQ ID NO:2. A salt bridge to the target DNAbb (DNA phosphate backbone) for base pairs 22-23 is affected. Figure 14E shows one diagram of the combination of the P deletion at 793 with the A708K substitution. Figure 14F shows another diagram showing that the effects of individual mutations are additive and that single mutations can be combined (stacked) for even greater improvement. Arrows indicate the locations of mutations throughout Figures 14A-14F. [Figure 14-2] Same as above. [Figure 14-3] Same as above. [Figure 14-4] Same as above. [Figure 14-5] Same as above. [Figure 14-6] Same as above.
[0032] [Figure 15] 1 is a plot showing the identification of optimal Planctomycetes CasX PAMs and spacers for a gene of interest, as described in Example 6. The Y-axis shows the percentage of GFP-negative cells, which show cleavage of the GFP reporter. The X-axis shows various PAM sequences and spacers: ATC PAM, CTC PAM, and TTC PAM. GTC, TTT, and CTT PAMs were also tested and showed no activity.
[0033] [Figure 16] 1 is a plot showing that improved CasX variants generated by DME can edit both canonical and non-canonical PAMs more efficiently than a reference CasX protein, as described in Example 6. The Y-axis shows the average fold improvement in editing compared to a reference sgRNA:protein pair with two targets (SEQ ID NO:2, SEQ ID NO:5) (N=6). The protein variants from left to right in each set of bars were A708K+[P793]+A739V, L379R+A708K+[P793], C477K+A708K+[P793], L379R+C477K+A708K+[P793], L379R+A708K+[P793]+A739V, C477K+A708K+[P793]+A739V, and L379R+C477K+A708K+[P793]+A739V. The reference CasX and protein variants were assayed using the reference sgRNA scaffold of SEQ ID NO: 5 with DNA encoding the spacer sequences, from left to right: E6 (SEQ ID NO: 29) and TTC PAM, E7 (SEQ ID NO: 30) and TTC PAM, GFP8 (SEQ ID NO: 31) and TTC PAM, B1 (SEQ ID NO: 32) and CTC PAM, and A7 (SEQ ID NO: 33) and ATC PAM.
[0034] [Figure 17-1]
[0049] Figure 17A and Figure 17D are a series of plots showing that a reference CasX protein and a reference sgRNA scaffold pair are highly specific for their target sequences, as described in Example 7. In Figures 17A and 17D, Streptococcus pyogenes Cas9 (SpyCas9) was assayed using two different gNA spacers and 5' PAM sites (SEQ ID NOS: 34-65) and (SEQ ID NOS: 136-166) for its ability to edit templates with target sequences complementary to the spacer sequences (arrows) or with one, two, three, or four mutations in the target sequence relative to the spacer sequences. In Figures 17B and 17E, Staphylococcus aureus Cas9 (SauCas9) was assayed using two different gNA spacers and 5' PAM sites (SEQ ID NOS: 66-103) and (SEQ ID NOS: 167-204) for its ability to edit templates with target sequences complementary to the spacer sequences (arrows) or with one, two, three, or four mutations in the target sequence relative to the spacer sequences. In Figures 17C and 17F, the reference Plm CasX protein and sgNA scaffold pair were assayed for their ability to edit target sequences complementary to the spacer sequence (arrows), or templates with 1, 2, 3, or 4 mutations in the target sequence relative to the spacer sequence, using two different sgNA spacers and 3' PAM sites (SEQ ID NOS: 104-135) and (SEQ ID NOS: 205-236). In all of Figures 17A-17F, the X-axis indicates the percentage of cells in which gene editing at the target sequence occurs. [Figure 17-2] Same as above. [Figure 17-3] Same as above. [Figure 17-4] Same as above. [Figure 17-5] Same as above. [Figure 17-6] Same as above.
[0035] [Figure 18] 1 shows the scaffold stem-loop of an exemplary reference sgRNA of the present disclosure (SEQ ID NO: 237).
[0036] [Figure 19] 1 shows the extended stem-loop sequence of an exemplary reference sgRNA of the present disclosure (SEQ ID NO: 238).
[0037] [Figure 20-1]
[0023] Figure 20A is a pair of plots showing that a particular subset of changes discovered by DME of CasX, as described in Example 4, is likely to predict improved activity. The plots represent data from the experiments described in Figures 7 and 8. Figure 20A shows that changing amino acids within 10 angstroms (A) of the guide RNA to hydrophobic residues (A, V, I, L, M, F, Y, W) results in a protein with significantly less activity. In contrast, Figure 20B shows that changing residues within 10 A of the RNA to positively charged amino acids (R, H, K) can improve activity. [Figure 20-2] Same as above.
[0038] [Figure 21-1] An alignment of two reference CasX protein sequences (SEQ ID NO: 1, top; SEQ ID NO: 2, bottom) is shown, with domains annotated. [Figure 21-2] Same as above.
[0039] [Figure 22] The domain organization of the reference CasX protein of SEQ ID NO: 1 is shown. The domains have the following coordinates: non-target strand binding (NTSB) domain: amino acids 101-191, helical I domain: amino acids 57-100 and 192-332, helical II domain: 333-509, oligonucleotide binding domain (OBD): amino acids 1-56 and 510-660, RuvC DNA cleavage domain (RuvC): amino acids 551-824 and 935-986, target strand loading (TSL) domain: amino acids 825-934. Note that the helical I, OBD, and RuvC domains are non-contiguous.
[0040] [Figure 23]An alignment of two CasX reference sgRNA scaffolds, SEQ ID NO: 5 (top) and SEQ ID NO: 4 (bottom) is shown.
[0041] [Figure 24] 1 shows an SDS-PAGE gel of StX2 (see CasX in SEQ ID NO: 2) purified fractions visualized by colloidal Coomassie staining, as described in Example 8. Lanes from left to right: Pellet: insoluble portion after cell lysis; Lysate: soluble portion after cell lysis; Flow-through: proteins that did not bind to the heparin column; Wash: proteins eluted from the column in wash buffer; Elute: proteins eluted from the heparin column with elution buffer; Flow-through: proteins that did not bind to the StrepTactin column; Elute: proteins eluted from the StrepTactin column with elution buffer; Inject: concentrated proteins injected onto an s200 gel filtration column; Frozen: pooled fractions from the s200 elution that were concentrated and frozen.
[0042] [Figure 25] 1 shows a chromatogram from a size exclusion chromatography assay of StX2, as described in Example 8.
[0043] [Figure 26] Shown is an SDS-PAGE gel of StX2 purified fractions visualized by colloidal Coomassie staining, as described in Example 8. From right to left: input sample, molecular weight marker; lanes 3-9: samples from the indicated elution volumes.
[0044] [Figure 27] 1 shows a chromatogram from a size-exclusion chromatography assay of CasX 119 using Superdex 200 16 / 600 pg gel filtration, as described in Example 8. The peak at 67.47 mL corresponds to the apparent molecular weight of CasX variant 119 and contained the majority of the CasX variant 119 protein.
[0045] [Figure 28]Figure 1 shows an SDS-PAGE gel of purified CasX 119 fractions visualized by colloidal Coomassie staining, as described in Example 8. Samples from the indicated fractions were separated by SDS-PAGE and stained with colloidal Coomassie. From right to left, Inject: Protein sample injected onto the gel filtration column, molecular weight marker; Lanes 3-10: Samples from the indicated elution volumes.
[0046] [Figure 29] An SDS-PAGE gel of a purified sample of CasX 438 visualized on a Bio-Rad Stain-Free™ gel is shown. Lanes from left to right are: Pellet: insoluble portion after cell lysis; Lysate: soluble portion after cell lysis; Flow-through: protein that did not bind to the heparin column; Elute: protein eluted from the heparin column with elution buffer; Flow-through: protein that did not bind to the StrepTactin column; Elute: protein eluted from the StrepTactin column with elution buffer; Inject: concentrated protein injected onto the s200 gel filtration column; Pool: pooled CasX-containing fractions; Final: pooled fractions from the concentrated and frozen s200 elution.
[0047] [Figure 30] 1 shows a chromatogram from a size-exclusion chromatography assay of CasX 438 using Superdex 200 16 / 600 pg gel filtration, as described in Example 8. The peak at 69.13 mL corresponds to the apparent molecular weight of CasX variant 438 and contained the majority of the CasX variant 438 protein.
[0048] [Figure 31] Figure 1 shows an SDS-PAGE gel of purified CasX 438 fractions visualized by colloidal Coomassie staining, as described in Example 8. Samples from the indicated fractions were separated by SDS-PAGE and stained with colloidal Coomassie. From right to left, Inject: Protein sample injected onto the gel filtration column, molecular weight marker; Lanes 3-10: Samples from the indicated elution volumes.
[0049] [Figure 32] Shown is an SDS-PAGE gel of a purified sample of CasX 457 visualized on a Bio-Rad Stain-Free™ gel. Lanes from left to right are: Pellet: insoluble portion after cell lysis; Lysate: soluble portion after cell lysis; Flow-through: protein that did not bind to the heparin column; Wash, Elute: protein eluted from the heparin column with elution buffer; Flow-through: protein that did not bind to the StrepTactin column; Elute: protein eluted from the StrepTactin column with elution buffer; Inject: concentrated protein injected onto an s200 gel filtration column; Final: pooled fractions from the s200 elution that were concentrated and frozen.
[0050] [Figure 33] 1 shows a chromatogram from a size-exclusion chromatography assay of CasX 457 using Superdex 200 16 / 600 pg gel filtration as described in Example 8. The peak at 67.52 mL corresponds to the apparent molecular weight of CasX variant 457 and contained the majority of the CasX variant 457 protein.
[0051] [Figure 34] Figure 1 shows an SDS-PAGE gel of purified CasX 457 fractions visualized by colloidal Coomassie staining, as described in Example 8. Samples from the indicated fractions were separated by SDS-PAGE and stained with colloidal Coomassie. From right to left, Inject: Protein sample injected onto the gel filtration column, molecular weight marker; Lanes 3-10: Samples from the indicated elution volumes.
[0052] [Figure 35] FIG. 1 is a schematic diagram showing the organization of the components of the pSTX34 plasmid used to assemble the CasX constructs, as described in Example 9.
[0053] [Figure 36]FIG. 1 is a schematic diagram showing the steps for generating CasX 119 variants, as described in Example 9.
[0054] [Figure 37] 1 is a graph showing the results of an assay for quantifying the active fraction of RNPs formed by sgRNA174 and CasX variants 119 and 457, as described in Example 19. Equimolar amounts of RNP and target were co-incubated, and the amount of cleaved target was determined at the indicated time points. The mean and standard deviation of three independent replicates for each time point are shown. The biphasic match of the combined replicates is shown. "2" refers to the reference CasX protein of SEQ ID NO: 2.
[0055] [Figure 38] 1 is a graph of the results of an assay for quantifying the active fraction of RNPs formed by CasX2 and reference guide 2, modified sgRNA guides 32, 64, and 174, as described in Example 19. Equimolar amounts of RNP and target were co-incubated, and the amount of cleaved target was determined at the indicated time points. The mean and standard deviation of three independent replicates for each time point are shown. The biphasic matches of the combined replicates are shown. "2" refers to the reference gRNA of SEQ ID NO: 5, and the identification numbers of the modified sgRNAs are shown in Table 2.
[0056] [Figure 39] Graph showing the results of an assay for quantifying the cleavage rate of RNPs formed by sgRNA174 and CasX variants 119 and 457, as described in Example 19. Target DNA was incubated with a 20-fold excess of the indicated RNPs, and the amount of cleaved target was determined at the indicated time points. The average and standard deviation of three independent replicates for each time point are shown. The monophasic match of the combined replicates is shown.
[0057] [Figure 40]Graph of the results of an assay for quantifying the cleavage rate of RNPs formed by CasX2 and sgRNA guide variants 2, 32, 64, and 174, as described in Example 19. Target DNA was incubated with a 20-fold excess of the indicated RNPs, and the amount of cleaved target was determined at the indicated time points. The average and standard deviation of three independent replicates for each time point are shown. The monophasic match of the combined replicates is shown.
[0058] [Figure 41] 12 is a graph of the results of an assay for quantifying the initial velocity of RNPs formed by CasX2 and sgRNA guide variants 2, 32, 64, and 174, as described in Example 19. The first two time points from previous cleavage experiments were fitted to a linear model to determine the initial cleavage rate.
[0059] [Figure 42] FIG. 1 is a schematic diagram showing examples of CasX protein and scaffold DNA sequences for packaging into adeno-associated virus (AAV), as described in Example 20. The DNA segment between the AAV inverted terminal repeats (ITRs), consisting of the DNA encoding CasX and its promoter, and the DNA encoding the scaffold and its promoter, is packaged into the AAV capsid during AAV production.
[0060] [Figure 43] Figure 2 is a graph showing representative results of AAV titration by qPCR, as described in Example 20. During AAV purification, the flow-through (FT) and successive elution fractions (1-6) are collected and titrated by qPCR. Most of the virus, in this example approximately 1e14 viral genomes, is found in the second elution fraction.
[0061] [Figure 44]This figure shows the results of an AAV-mediated gene editing experiment in an SOD1-GFP reporter cell line, as described in Example 21. CasX constructs (CasX 119 and guide 64 containing SOD1-targeting spacer 2, ATGTTCATGAGTTTGGAGAT, SEQ ID NO: 239) and SauCas9 containing the SOD1-targeting spacer were packaged into AAV vectors and used to transduce SOD1-GFP reporter cells at various multiplicities of infection (MOI, number of viral genomes / cell). After 12 days, cells were assayed for GFP disruption by FACS. In this example, CasX and SauCas9 showed comparable levels of editing, with 1-2% of cells showing GFP disruption at the highest MOIs of 1e7 or 1e6.
[0062] [Figure 45] This figure shows the results of a second AAV-mediated gene editing experiment in the SOD1-GFP reporter cell line, as described in Example 21. The CasX construct 119.64 containing the SOD1-targeting spacer (2, ATGTTCATGAGTTTGGAGAT, SEQ ID NO: 239) and SauCas9 containing the SOD1-targeting spacer were packaged into AAV vectors and used to transduce SOD1-GFP reporter cells at various multiplicities of infection (MOI, number of viral genomes / cell). After 12 days, cells were assayed for GFP disruption by FACS. In this example, CasX and SauCas9 showed comparable levels of editing at the highest MOI, with approximately 2-4% of cells showing GFP disruption.
[0063] [Figure 46]
[0049] Figure 2 shows the results of an AAV-mediated gene editing experiment in neural progenitor cells (NPCs) derived from the G93A mouse model of ALS, as described in Example 21. A CasX construct (CasX 119 and guide 64, including SOD1 targeting spacer 2, ATGTTCATGAGTTTGGAGAT, SEQ ID NO: 239) was packaged into an AAV vector and used to transduce G93A NPCs at various multiplicities of infection (MOI, number of viral genomes / cell). After 12 days, cells were assayed for gene editing by T7E1 assay. The agarose gel image from the T7E1 assay shown here demonstrates successful editing of the SOD1 locus. The double arrow indicates two DNA bands resulting from successful editing in the cells.
[0064] [Figure 47]
[0023] Figure 1 shows the results of an editing assay of six target genes in HEK293T cells, as described in Example 23. Each dot represents the results using an individual spacer.
[0065] [Figure 48] FIG. 1 shows the results of editing assays of six target genes in HEK293T cells, with each bar representing the results obtained with an individual spacer, as described in Example 23.
[0066] [Figure 49]
[0023] Figure 1 shows the results of an editing assay of four target genes in HEK293T cells, as described in Example 23. Each dot represents the results using an individual spacer utilizing CTC(CTCN)PAM.
[0067] [Figure 50]
[0033] Figure 1 is a schematic diagram showing the steps of deep mutational evolution used to generate a library of genes encoding CasX variants, as described in Example 24. The pSTX1 backbone is minimal, consisting only of a high-copy number origin of replication and a KanR resistance gene, making it compatible with the recombinant E. coli strain EcNR2. pSTX2 is a BsmbI destination plasmid for Tc-inducible expression in E. coli.
[0068] [Figure 51] 1 is a dot plot graph showing the results of CRISPRi screening for mutations in libraries D1, D2, and D3, as described in Example 24. In the absence of CRISPRi, E. coli constitutively express both GFP and RFP and fluoresce strongly at both wavelengths, represented by dots in the upper right region of the plot. CasX proteins that result in CRISPRi of GFP can reduce the green fluorescence by more than 10-fold without changing the red fluorescence, and these cells are within sort gate 1 shown. The total percentage of cells showing CRISPRi is shown.
[0069] [Figure 52] Figure 2 shows a photograph of colonies grown in a ccdB assay, as described in Example 24. Assaying 10-fold dilutions in the presence of glucose or arabinose to induce expression of the ccdB toxin resulted in an approximately 1000-fold difference between functional and non-functional protein. When grown in liquid medium, the resolution was approximately 10,000-fold, as seen on the right.
[0070] [Figure 53]
[0033] Figure 2 shows a graph of HEK iGFP genome editing efficiency for CasX variants tested using sgRNA 2 (SEQ ID NO: 5) with appropriate spacers, as described in Example 24. Data are expressed as fold improvement over wild-type CasX protein (SEQ ID NO: 2) in a HEK iGFP editing assay. Single mutations are shown at the top, and groups of mutations are shown at the bottom of the graph. Error bars represent a combined internal measurement error (SD) and inter-experiment measurement error (SD across replicate experiments for variants tested more than once) across at least three assays.
[0071] [Figure 54] FIG. 10 is a scatter plot showing the results of an SOD1-GFP reporter assay for CasX variants using sgRNA scaffold 2, which utilizes two different spacers for GFP, as described in Example 24.
[0072] [Figure 55] FIG. 22 is a graph showing the results of a HEK293 iGFP genome editing assay assessing editing across four different PAM sequences comparing wild-type CasX (SEQ ID NO: 2) and CasX variant 119, both of which utilized sgRNA scaffold 1 (SEQ ID NO: 4) with spacers utilizing four different PAM sequences, as described in Example 24.
[0073] [Figure 56] 2 is a graph showing the results of the genome editing activity of CasX variant 119 and sgRNA 174 compared to wild-type CasX 2 and guide scaffold 1 in an iGFP lipofection assay utilizing two different spacers, as described in Example 24.
[0074] [Figure 57]22A-22C are graphs showing the results of the genome editing activity of CasX variant 119 and sgRNA 174 compared to wild-type CasX and guide in an iGFP lentiviral transduction assay using two different spacers, as described in Example 24.
[0075] [Figure 58] 10A-B are graphs showing genome editing results in a more stringent lentiviral assay to compare the editing activity of four CasX variants (119, 438, 488, and 491) and the optimized sgNA 174 and two different spacers, as described in Example 24. The results show the incremental improvement in editing efficiency achieved by further modifications and domain swaps introduced into the starting 119 variant.
[0076] [Figure 59-1] Figure 59 shows the results of NGS analysis of a library of sgRNAs, as described in Example 25. Figure 59A shows the distribution of substitutions, deletions, and insertions. Figure 59B is a scatter plot showing the high reproducibility of variant representation in two separate library pools after CRISPRi assays on unsorted naive cell populations. (Library pools D3 and D2 are two different forms of dCasX protein and represent replicate CRISPRi assays.) [Figure 59-2] Same as above.
[0077] [Figure 60-1]The structures of wild-type CasX and the RNA guide (SEQ ID NO: 4) are shown. Figure 60A shows the cryoEM structure of the Deltaproteobacteria CasX protein:sgRNA RNP complex (PDB id: 6YN2), containing two stem-loops, a pseudoknot, and a triplex. Figure 60B shows that the secondary structure of the sgRNA was identified from the structure shown in (A) using the tool RNAPDBee 2.0 (rnapdbee.cs.put.poznan.pl / , using the tool 3DNA / DSSR, and using the VARNA visualization tool). The RNA region is indicated. Residues that were not apparent in the PDB crystal structure file are indicated by standard letters (i.e., not boxed) and are not included in the residue numbering. [Figure 60-2] Same as above.
[0078] [Figure 61-1] A comparison between two guide RNA scaffolds is shown. Figure 61A provides a sequence alignment between single-guide scaffold 1 (SEQ ID NO: 4) and scaffold 2 (SEQ ID NO: 5). Figure 61B shows the predicted secondary structure of scaffold 1 (without the 5' ACAUCU base, which was absent in the cryoEM structure). Predictions were made using RNAfold (v2.1.7) (see Figures 60A-60B) using constraints derived from base pairs observed in the cryoEM structure. This constraint required that the base pairs observed in the cryoEM structure be formed, and that bases involved in triplex formation be unpaired. This structure has different base pairing from the lowest energy predicted structure at the 5' end (i.e., a pseudoknot and triplex loop). Figure 61C shows the predicted secondary structure of scaffold 2. Predictions for scaffold 1 were made using similar constraints based on sequence alignment. [Figure 61-2] Same as above. [Figure 61-3] Same as above.
[0079] [Figure 62]1 shows a graph comparing the GFP knockdown ability of scaffold 1 and scaffold 2 in a GFP-lipofection assay using four different spacers utilizing different PAM sequences, as described in Example 25. The results show greater editing conferred by the use of modified scaffold 2 compared to wild-type scaffold 1, with wild-type scaffold 1 showing no editing with spacers utilizing the GTC and CTC PAM sequences.
[0080] [Figure 63-1] Figure 63 shows graphs depicting the enrichment of single variants across scaffolds, highlighting the mutated regions, as described in Example 25. Figure 63A shows the substituted bases (A, T, G, or C, from top to bottom), Figure 63B shows the inserted bases (A, T, G, or C, from top to bottom), and Figure 63C shows the deletions (X-axis) at individual nucleotide positions across scaffold 2. Enrichment values were averaged across the three inactive CasX forms relative to the average WT value. Scaffolds with relative log2 enrichment greater than 0 are considered "enriched" because they are more represented in the sorted population compared to the naive population than the wild-type scaffold. Error bars represent the confidence interval across the three catalytically inactive CasX experiments. [Figure 63-2] Same as above. [Figure 63-3] Same as above.
[0081] [Figure 64] Scatter plot showing that enrichment values obtained across different dCasX variants are largely consistent, as described in Example 25. Libraries D2 and DDD have highly correlated enrichment scores, while D3 is more clearly distinct.
[0082] [Figure 65] 1 shows a bar graph of the cleavage activity of several scaffold variants in a more stringent lipofection assay at the SOD1-GFP locus, as described in Example 25.
[0083] [Figure 66] A bar graph of cleavage activity for several scaffold variants using two different spacers, 8.2 and 8.4 targeting the SOD1-GFP locus (and the non-targeting spacer NT), with low MOI lentiviral transduction using a p34 plasmid backbone, as described in Example 25.
[0084] [Figure 67] The diagram shows the secondary structure of a single guide 174 at the top and the linear structure at the bottom, with the lines connecting these segments held together by base pairing or other noncovalent interactions. The scaffold stem (white, unfilled) (and loop) and extended stem (gray, unfilled) (and loop) are adjacent from 5' to 3' in sequence. However, the pseudoknot and extended stem are formed from strands with intervening regions within the sequence. In the case of the single guide 174, a triplex is formed, including nucleotides 5'-CUUUG'-3' and 5'-CAAAG-3', which form a base-paired duplex, and nucleotide 5'-UUU-3', which combines with 5'-AAA-3' to form the triplex region.
[0085] [Figure 68-1] Figure 68A shows a comparison of the highly evolved single guide 174 with scaffolds 1 and 2, which serve as starting points for the DME procedure described in Example 25. Figure 68A shows a bar graph of cleavage activity for a direct comparison of the cleavage activity of guide scaffolds with five different spacers in a plasmid lipofection assay at the GFP locus in HEK-GFP cells. Figure 68B shows a sequence alignment between scaffold 2 and guide 174 (SEQ ID NO: 2238). Asterisks indicate point mutations, and dotted boxes indicate entire extended stem exchanges. [Figure 68-2] Same as above.
[0086] [Figure 69-1]69A and 69B show scattergrams of HEK-iGFP cleavage assays for scaffold sequences compared to the WT scaffold using two spacers, 4.76 (FIG. 69A) and 4.77 (FIG. 69B), as described in Example 25. [Figure 69-2] Same as above.
[0087] [Figure 70] Figure 1 shows a scatter plot comparing the normalized cleavage activity of several scaffolds compared to WT using two spacers (4.76 and 4.77), as described in Example 25. Error bars combine, in quadrature, the intra- and inter-experimental measurement error (SD) and the inter-experimental measurement error (SD across replicate experiments for variants tested more than once).
[0088] [Figure 71] Figure 63 shows a scatter plot comparing the cleavage activity of multiple scaffolds normalized to WT in a HEK-iGFP cleavage assay with the enrichment obtained from a comprehensive CRISPRi screen, as described in Example 25. Generally, highly enriched (>1.5) scaffold mutations have cleavage activity comparable to or greater than WT. Two variants have high cleavage activity at low enrichment scores (C18G and T17G); interestingly, these substitutions are located at the same positions as several highly enriched insertions (Figures 63A-63C). Labels indicate the mutations for the subsets compared.
[0089] [Figure 72] 1 shows the results of flow cytometry analysis of Cas-mediated editing at the RHO locus in APRE19 RHO-GFP cells 14 days after transfection with CasX variant constructs 438, 499, and 491, as described in Example 26. Dots represent results for individual samples, and light dashed lines represent upper and lower quartiles.
[0090] [Figure 73]Quantification of the cleavage rates of RNPs formed by sgRNA174 and CasX variants on targets with different PAMs is shown. Target DNA was incubated with a 20-fold excess of the indicated RNPs, and the amount of cleaved target was determined at the indicated time points. Monophasic replication of the combined RNPs is shown. DETAILED DESCRIPTION OF THE INVENTION
[0091] While exemplary embodiments have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions will now occur to those skilled in the art without departing from the invention claimed herein. It is understood that various alternatives to the embodiments described herein can be used in practicing the embodiments of the present disclosure. The following claims define the scope of the invention, and it is intended that methods and structures within the scope of these claims and their equivalents be covered thereby.
[0092] definition Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Although methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present embodiments, suitable methods and materials are described below. In case of conflict, the present specification, including definitions, will control. In addition, the materials, methods, and examples are illustrative only and are not intended to be limiting. Numerous variations, changes, and substitutions will now occur to those skilled in the art without departing from the invention.
[0093] The terms "polynucleotide" and "nucleic acid," used interchangeably herein, refer to polymeric forms of nucleotides of any length, either ribonucleotides or deoxyribonucleotides. Thus, the terms "polynucleotide" and "nucleic acid" encompass single-stranded DNA, double-stranded DNA, multi-stranded DNA, single-stranded RNA, double-stranded RNA, multi-stranded RNA, genomic DNA, cDNA, DNA-RNA hybrids, and polymers containing purine and pyrimidine bases or other natural, chemically or biochemically modified, non-natural, or derivatized nucleotide bases.
[0094] "Hybridizable" and "complementary" are used interchangeably and mean that a nucleic acid (e.g., RNA, DNA) contains a sequence of nucleotides that allows it to "anneal" or "hybridize" to another nucleic acid in a non-covalent manner, i.e., by forming Watson-Crick base pairs and / or G / U base pairs, in a sequence-specific, nonparallel manner (i.e., a nucleic acid that specifically binds to a complementary nucleic acid) under appropriate in vitro and / or in vivo conditions of temperature and solution ionic strength. It is understood that a polynucleotide sequence need not be 100% complementary to that of its target nucleic acid to specifically hybridize; it can have at least about 70%, at least about 80%, or at least about 90%, or at least about 95% sequence identity and still be able to hybridize to a target nucleic acid. Furthermore, a polynucleotide can hybridize across one or more segments (e.g., a loop or hairpin structure, a "bulge," a "bubble," etc.) such that intervening or adjacent segments are not involved in the hybridization event.
[0095] For purposes of this disclosure, a "gene" includes a DNA region that encodes a gene product (e.g., protein, RNA), as well as a DNA region that regulates the production of the gene product, regardless of whether such regulatory sequences are adjacent to the coding and / or transcribed sequence. Thus, a gene can include regulatory sequences, including, but not necessarily limited to, promoter sequences, terminators, translational regulatory sequences such as ribosome binding sites and internal ribosome entry sites, enhancers, silencers, insulators, boundary elements, origins of replication, matrix attachment sites, and locus control regions. A coding sequence encodes a gene product upon transcription or transcription and translation; coding sequences of this disclosure can include fragments and need not include a full-length open reading frame. A gene can include both the transcribed strand, e.g., the strand containing the coding sequence, and the complementary strand.
[0096] The term "downstream" refers to a nucleotide sequence located 3' relative to a reference nucleotide sequence. In certain embodiments, a downstream nucleotide sequence refers to a sequence following the start of transcription. For example, the translation start codon of a gene is located downstream of the transcription start site.
[0097] The term "upstream" refers to a nucleotide sequence located 5' relative to a reference nucleotide sequence. In certain embodiments, an upstream nucleotide sequence refers to a sequence located 5' of a coding region or the start of transcription. For example, most promoters are located upstream of the transcription start site.
[0098] The term "regulatory element" is used interchangeably herein with the term "regulatory sequence" and is intended to include promoters, enhancers, and other expression control elements (e.g., transcription termination signals such as polyadenylation signals and polyU sequences). Exemplary regulatory elements include, but are not limited to, transcription promoters such as CMV, CMV + intron A, SV40, RSV, HIV-Ltr, elongation factor 1 alpha (EF1α), MMLV-Ltr, internal ribosome entry sites (IRES), or P2A peptides that enable translation of multiple genes from a single transcript, metallothionein, transcriptional enhancer elements, transcription termination signals, polyadenylation sequences, sequences for optimizing translation initiation, and translation termination sequences. It will be understood that the selection of appropriate regulatory elements will depend on the encoded components (e.g., protein or RNA) to be expressed, or whether the nucleic acid contains multiple components that require different polymerases or are not intended to be expressed as a fusion protein.
[0099] The term "promoter" refers to a DNA sequence that contains an RNA polymerase binding site, a transcription initiation site, a TATA box, and / or a B recognition element and that supports or facilitates the transcription and expression of an associated transcribable polynucleotide sequence and / or gene (or transgene). Promoters can be produced synthetically or derived from known or naturally occurring promoter sequences or from another promoter sequence. A promoter can be proximal or distal to the gene to be transcribed. Promoters can also include chimeric promoters, which contain a combination of two or more heterologous sequences to confer specific properties. Promoters of the present disclosure can include variants of promoter sequences that are similar in composition but not identical to other promoter sequences known or provided herein. Promoters can be classified according to criteria related to the expression pattern of the associated coding or transcribable sequence or gene operably linked to the promoter, such as constitutive, developmental, tissue-specific, inducible, etc.
[0100] The term "enhancer" refers to a regulatory element DNA sequence that regulates the expression of an associated gene upon binding of specific proteins called transcription factors. Enhancers can be located in the introns of a gene or 5' or 3' of the coding sequence of a gene. Enhancers can be proximal to the gene (i.e., within tens or hundreds of base pairs (bp) of the promoter) or distal to the gene (i.e., thousands, hundreds of thousands, or even millions of bp away from the promoter). A single gene can be regulated by more than one enhancer, all of which are contemplated within the scope of this disclosure.
[0101] As used herein, "recombinant" means that a particular nucleic acid (DNA or RNA) is the product of various combinations of cloning, restriction, and / or ligation steps, resulting in a construct with structural coding or non-coding sequences distinguishable from endogenous nucleic acids found in natural systems. Generally, DNA sequences encoding structural coding sequences can be assembled from cDNA fragments and short oligonucleotide linkers or from a series of synthetic oligonucleotides to provide a recombinant transcription unit contained in a cell or a synthetic nucleic acid that can be expressed from a cell-free transcription and translation system. Such sequences can be provided in the form of an open reading frame uninterrupted by internal non-translated sequences or introns typically present in eukaryotic genes. Genomic DNA containing the relevant sequences can also be used to form recombinant genes or transcription units. Sequences of non-translated DNA may be present 5' or 3' from the open reading frame; such sequences do not interfere with the manipulation or expression of the coding region but can actually act to regulate production of the desired product by various mechanisms (see "enhancer" and "promoter" above).
[0102] The term "recombinant polynucleotide" or "recombinant nucleic acid" refers to something that does not occur in nature, e.g., something that is created by the artificial combination of two otherwise separated segments of sequence through human intervention. This artificial combination is often achieved by chemical synthesis means or by the artificial manipulation of isolated segments of nucleic acid, e.g., by genetic engineering techniques. This can usually be done to replace a codon with a redundant codon that encodes the same or a conservative amino acid, while introducing or removing a sequence recognition site. Alternatively, it can be performed by joining nucleic acid segments of desired functions together to produce a desired combination of functions. This artificial combination is often achieved by chemical synthesis means or by the artificial manipulation of isolated segments of nucleic acid, e.g., by genetic engineering techniques.
[0103] Similarly, the term "recombinant polypeptide" or "recombinant protein" refers to a non-naturally occurring polypeptide or protein, made, for example, by human intervention, by the artificial combination of two otherwise separated segments of amino acid sequence. Thus, for example, a protein comprising a heterologous amino acid sequence is recombinant.
[0104] As used herein, the term "contact" refers to establishing a physical bond between two or more entities.For example, contacting a target nucleic acid with a guide nucleic acid means that the target nucleic acid and the guide nucleic acid are made to share a physical bond, for example, if the sequences share sequence similarity, they can hybridize.
[0105] "Dissociation constant" or "K d " are used interchangeably and refer to the affinity between a ligand "L" and a protein "P", i.e., how tightly the ligand binds to a particular protein. This is expressed in the formula K d = [L][P] / [LP], where [P], [L], and [LP] represent the molar concentrations of the protein, ligand, and complex, respectively.
[0106] The present disclosure provides compositions and methods useful for editing a target nucleic acid sequence. As used herein, "editing" is used interchangeably with "modifying" and includes, but is not limited to, truncation, nicking, deletion, knock-in, knock-out, etc.
[0107] As used herein, " homology-directed repair " (HDR) refers to the form of DNA repair that occurs during the repair of double-strand breaks in cells.This process requires nucleotide sequence homology, and uses a donor template to repair or knock out target DNA, resulting in the transfer of genetic information from donor (such as donor template) to target.If the donor template is different from the target DNA sequence, and some or all of the sequence of the donor template is integrated into the target DNA at the correct genomic locus, homology-directed repair can cause the sequence change of the target nucleic acid sequence by insertion, deletion or mutation.
[0108] As used herein, "non-homologous end joining" (NHEJ) refers to the repair of double-stranded breaks in DNA by direct ligation of the broken ends to each other, without the need for a homologous template (as opposed to homology recombination repair, which requires a homologous sequence to guide the repair). NHEJ often results in indels, losses (deletions) or insertions of nucleotide sequences near the double-stranded break site.
[0109] As used herein, "microhomology-mediated end joining" (MMEJ) refers to a mutagenic DSB repair mechanism that does not require a homologous template (as opposed to homologous recombination repair, which requires a homologous sequence to guide the repair) and is always associated with a deletion adjacent to the break site. MMEJ often results in the loss (deletion) of nucleotide sequences near the double-strand break site.
[0110] A polynucleotide or polypeptide (or protein) has a certain percentage "sequence similarity" or "sequence identity" to another polynucleotide or polypeptide, meaning that when aligned, that percentage of bases or amino acids are the same and are in the same relative positions when the two sequences are compared. Sequence similarity (sometimes called percent similarity, percent identity, or homology) can be determined in several different ways. To determine sequence similarity, sequences can be aligned using methods and computer programs known in the art, including BLAST, available on the World Wide Web at ncbi.nlm.nih.gov / BLAST. The percent complementarity between specific stretches of nucleic acid sequence within a nucleic acid can be determined using any convenient method. Exemplary methods include the use of the BLAST program (Basic Local Alignment Search Tool) and the PowerBLAST program (Altschul et al., J. Mol. Biol., 1990, 215, 403-410; Zhang and Madden, Genome Res., 1997, 7, 649-656) or, for example, the Smith and Waterman algorithm (Adv. Appl. Math., 1981, 2, 482-489), using default settings, or the use of the Gap program (Wisconsin Sequence Analysis Package, Version 8 for Unix, Genetics Computer Group, University Research Park, Madison Wis.).
[0111] The terms "polypeptide" and "protein" are used interchangeably herein and refer to polymeric forms of amino acids of any length, which can include coded and non-coded amino acids, chemically or biochemically modified or derivatized amino acids, and polypeptides with modified peptide backbones. This term includes fusion proteins, including, but not limited to, fusion proteins with heterologous amino acid sequences.
[0112] A "vector" or "expression vector" is a replicon, such as a plasmid, phage, virus, or cosmid, to which another DNA segment, or "insert," can be attached, so as to bring about the replication or expression of the attached segment in a cell.
[0113] As used herein, the term "naturally-occurring" or "unmodified" or "wild-type" as applied to a nucleic acid, polypeptide, cell, or organism refers to a nucleic acid, polypeptide, cell, or organism that is found in nature.
[0114] As used herein, a "mutation" refers to an insertion, deletion, substitution, duplication, or inversion of one or more amino acids or nucleotides compared to a wild-type or reference amino acid sequence or a wild-type or reference nucleotide sequence.
[0115] As used herein, the term "isolated" is meant to refer to a polynucleotide, polypeptide, or cell that is present in an environment that is different from the environment that the polynucleotide, polypeptide, or cell naturally occurs in. An isolated genetically modified host cell may be present in a mixed population of genetically modified host cells.
[0116] As used herein, "host cell" means a eukaryotic cell, a prokaryotic cell, or a cell derived from a multicellular organism (e.g., a cell line) cultured as a unicellular entity, including the progeny of the original cell that has been used as a recipient of a nucleic acid (e.g., an expression vector) and that has been genetically modified by the nucleic acid. It is understood that the progeny of a single cell may not necessarily be completely identical in morphology or genomic or total DNA complement to the original parent due to natural, accidental, or deliberate mutation. A "recombinant host cell" (also called a "genetically modified host cell") is a host cell into which a heterologous nucleic acid, e.g., an expression vector, has been introduced.
[0117] The term "conservative amino acid substitution" refers to the interchangeability of amino acid residues with similar side chains in proteins. For example, the group of amino acids with aliphatic side chains consists of glycine, alanine, valine, leucine, and isoleucine; the group of amino acids with aliphatic-hydroxyl side chains consists of serine and threonine; the group of amino acids with amide-containing side chains consists of asparagine and glutamine; the group of amino acids with aromatic side chains consists of phenylalanine, tyrosine, and tryptophan; the group of amino acids with basic side chains consists of lysine, arginine, and histidine; and the group of amino acids with sulfur-containing side chains consists of cysteine and methionine. Exemplary conservative amino acid substitutions are valine-leucine-isoleucine, phenylalanine-tyrosine, lysine-arginine, alanine-valine, and asparagine-glutamine.
[0118] As used herein, "treatment" or "treating" are used interchangeably herein and refer to an approach to obtaining beneficial or desired results, including, but not limited to, therapeutic benefit and / or prophylactic benefit. Therapeutic benefit refers to the eradication or amelioration of the underlying disorder or disease being treated. Therapeutic benefit can also be achieved by the eradication or amelioration of one or more symptoms, or the improvement of one or more clinical parameters associated with the underlying disease, such that an improvement is observed in a subject, even though the subject may still suffer from the underlying disorder.
[0119] As used herein, the terms "therapeutically effective amount" and "therapeutically effective dose" refer to an amount of a composition, vector, cell, etc. that, when administered to a subject in one or multiple doses, can have any detectable beneficial effect on any symptom, aspect, measured parameter, or characteristic of a disease condition or state. Such an effect need not be absolutely beneficial. Such an effect may be temporary.
[0120] As used herein, "administering" refers to the method of giving a dosage of a composition of the present disclosure to a subject.
[0121] As used herein, a "subject" is a mammal, including, but not limited to, domestic animals, primates, non-human primates, humans, dogs, porcine (pigs), rabbits, mice, rats, and other rodents.
[0122] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference.
[0123] I. General Methods The practice of the present invention will employ, unless otherwise indicated, conventional techniques of immunology, biochemistry, chemistry, molecular biology, microbiology, cell biology, genomics, and recombinant DNA, which may be adapted from techniques such as those described in Molecular Cloning: A Laboratory Manual, 3rd Ed. (Sambrook et al., Harbor Laboratory Press 2001), Short Protocols in Molecular Biology, 4th Ed. (Ausubel et al. eds., John Wiley & Sons 1999), Protein Methods (Bollag et al., John Wiley & Sons 1996), Nonviral Vectors for Gene Therapy (Wagner et al. eds., Academic Press 1999), Viral Vectors (Kaplift & Loewy eds., Academic Press 1995), Immunology Methods Manual (I. Lefkovits ed., Academic Press 1997), and Cell and Tissue Culture: Laboratory Procedures in These can be found in standard textbooks such as Biotechnology (Doyle & Griffiths, John Wiley & Sons 1998), the disclosures of which are incorporated herein by reference.
[0124] When a range of values is provided, the endpoints are included, and intervening values, to the tenth of the unit of the lower limit, unless the context clearly dictates otherwise, are understood to be between the upper and lower limits of that range, and other stated and intervening values in that stated range are encompassed. The upper and lower limits of these smaller ranges can independently be included in the smaller ranges and are also included, subject to any specifically excluded limit in the stated range. When the stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included.
[0125] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. All publications mentioned herein are incorporated by reference to disclose and describe the methods and / or materials in connection with which the publications are cited.
[0126] It must be noted that as used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise.
[0127] It will be understood that certain features of the present disclosure that are, for clarity, described in the context of separate embodiments, may also be provided in combination in a single embodiment. In other cases, for brevity, various features of the present disclosure that are described in the context of a single embodiment may also be provided separately or in any suitable subcombination. All combinations of the embodiments related to the present disclosure are specifically embraced by the present disclosure and are intended to be disclosed herein as if each and every combination were individually and expressly disclosed. Furthermore, all subcombinations of the various embodiments and elements thereof are also specifically embraced by the present disclosure and are disclosed herein as if each and every such subcombination were individually and expressly disclosed herein.
[0128] II. CasX:gNA System In a first aspect, the present disclosure provides a CasX:gNA system comprising a CasX protein and one or more guide nucleic acids (gNAs) for use in modifying or editing target nucleic acids, including coding and non-coding regions. The terms CasX protein and CasX are used interchangeably herein, and the terms CasX variant protein and CasX variant are used interchangeably herein. The CasX protein and gNA of the CasX:gNA system, each independently provided herein, can be a reference CasX protein, a CasX variant protein, a reference gNA, a gNA variant, or any combination of a reference CasX protein, a reference gNA, a CasX variant protein, or a gNA variant. The gNA and CasX protein, the gNA variant and the CasX variant, or any combination thereof, can form a complex and associate through non-covalent interactions, referred to herein as a ribonucleoprotein (RNP) complex. In some embodiments, the use of a pre-complexed CasX:gNA confers advantages in the delivery of system components to cells or target nucleic acids for editing the target nucleic acid. In RNPs, the gNA can provide target specificity to the RNP complex by including a spacer sequence (targeting sequence) having a nucleotide sequence complementary to the sequence of the target nucleic acid. In RNPs, the CasX protein of the pre-complexed CasX:gNA provides site-specific activity and is guided to a target site within the target nucleic acid sequence to be modified by association with the gNA (and further stabilized at the target site). The CasX protein of the RNP complex provides the site-specific activity of the complex, such as binding, cleavage, or nicking of the target sequence by the CasX protein. Provided herein are compositions and cells comprising a CasX:gNA gene editing pair, including a reference CasX protein, a CasX variant protein, a reference gNA, a gNA variant, and any combination of CasX and gNA, as well as delivery modalities comprising the CasX:gNA.In other embodiments, the present disclosure provides vectors encoding or including a CasX:gNA pair, and optionally, donor templates for the production and / or delivery of CasX:gNA systems. Also provided herein are methods for making CasX proteins and gNAs, as well as methods for using CasX and gNAs, including gene editing and therapeutic methods. The CasX protein and gNA components of CasX:gNA and their characteristics, as well as delivery modes and methods of using the compositions, are described in more detail below.
[0129] The donor template of the CasX:gNA system is designed depending on whether it is being used to correct a mutation in a target gene, to insert a transgene at a different locus in the genome ("knock-in"), or to prevent expression of an aberrant gene product; for example, it contains one or more mutations that reduce expression of a gene product or render the protein nonfunctional ("knockdown" or "knockout"). In some embodiments, the donor template is a single-stranded DNA template or a single-stranded RNA template. In other embodiments, the donor template is a double-stranded DNA template. In some embodiments, the CasX:gNA system used to edit a target nucleic acid includes a donor template having all or at least a portion of the open reading frame of a gene in the target nucleic acid for insertion of a corrected wild-type sequence to correct the defective protein. In other cases, the donor template includes all or a portion of the wild-type gene for insertion at a different locus in the genome for expression of the gene product. In still other cases, a portion of a gene can be inserted upstream ('5) of a mutation in the target nucleic acid, with the donor template gene portion extending to the C-terminus of the gene, and its insertion into the target nucleic acid resulting in expression of the gene product. In other embodiments, the donor template can contain one or more mutations in the coding sequence compared to the normal wild-type sequence of the target gene utilized for insertion to knock out or knock down a defective target nucleic acid sequence (discussed in more detail below). In other embodiments, the donor template can contain regulatory elements, introns, or intron-exon junctions with sequences specifically designed to knock down or knock out a defective gene to allow expression of a functional gene product, or alternatively, to knock in a corrective sequence.In some embodiments, the donor polynucleotide comprises at least about 10, at least about 20, at least about 50, at least about 100, at least about 200, at least about 300, at least about 400, at least about 500, at least about 600, at least about 700, at least about 800, at least about 900, at least about 1000, at least about 10,000, at least about 15,000, at least about 25,000, at least about 50,000, at least about 100,000, or at least about 200,000 nucleotides. If there is a stretch of DNA sequence with a sufficient number of nucleotides with sufficient homology adjacent to the cleavage site of the target nucleic acid sequence targeted by CasX:gNA (i.e., 5' and 3' of the cleavage site) to support homology-directed repair (the flanking regions are "homology arms"), the use of such a donor template can result in its incorporation into the target nucleic acid by HDR. In other cases, the donor template can be inserted by non-homologous end joining (NHEJ, which does not require homologous arms) or microhomology-mediated end joining (MMEJ, which requires short homologous regions at the 5' and 3' ends). In some embodiments, the donor template contains homologous arms at the 5' and 3' ends, each of which has at least about 2, at least about 10, at least about 20, at least about 30, at least about 50, at least about 100, at least about 150, at least about 300, at least about 1000, at least about 1500 or more nucleotides that are homologous to sequences adjacent to the intended cleavage site of the target nucleic acid. In some embodiments, the CasX:gNA system utilizes two or more gNAs with targeting sequences complementary to overlapping or different regions of the target nucleic acid to excise the defective sequence by multiple double-strand breaks or by nicking at positions adjacent to the defective sequence, and the donor template is inserted by HDR to replace the excised sequence. As described above, gNAs are designed to contain targeting sequences that are 5' and 3' to the individual site or sequence to be excised. By appropriately selecting the targeting sequence of the gNA, defined regions of the target nucleic acid can be edited using the CasX:gNA system described herein.
[0130] III. Guide nucleic acid of the CasX:gNA system In other aspects, the present disclosure provides guide nucleic acids (gNAs) that can be utilized with the CasX:gNA system and have utility in editing target nucleic acids. The present disclosure provides specifically designed gNAs with targeting sequences (or "spacers") that are complementary to (and thus capable of hybridizing with) the target nucleic acid as components of the gene editing CasX:gNA system. In some embodiments, multiple gNAs (e.g., multiple gRNAs) are envisioned to be delivered by the CasX:gNA system for modification of different regions of a gene, including regulatory elements, exons, introns, or intron-exon junctions. In some embodiments, the targeting sequence of the gNA is complementary to a sequence containing one or more single nucleotide polymorphisms (SNPs) of the target nucleic acid. In other embodiments, the targeting sequence of the gNA is complementary to a sequence in an intergenic region. For example, if deletion of a protein-coding gene is desired, a pair of gNAs with targeting sequences to different or overlapping regions of the target nucleic acid sequence can be used to bind and cleave at two different sites within the gene, and the gene can then be edited by indel formation or homology-directed repair (HDR), which utilizes an inserted donor template to replace the deleted sequence and complete the edit.
[0131] a. Reference gNA and gNA variants In some embodiments, the gNA of the present disclosure comprises the sequence of a naturally occurring gNA ("reference gNA"). In other cases, the reference gNA of the present disclosure may be subjected to one or more mutagenesis methods, such as those described herein, which may include deep mutational evolution (DME), deep mutational scanning (DMS), error-prone PCR, cassette mutagenesis, random mutagenesis, staggered extension PCR, gene shuffling, or domain swapping, to generate one or more gNA variants with improved or altered properties compared to the reference gNA. gNA variants also include variants containing one or more exogenous sequences, for example, fused to either the 5' or 3' end or inserted internally. The activity of the reference gNA may be used as a benchmark against which the activity of the gNA variants is compared, thereby measuring the improvement in function or other characteristics of the gNA variants. In other embodiments, the reference gNA may be subjected to one or more deliberate, targeted mutations to produce gNA variants, such as rationally designed variants. As used herein, the terms gNA, gRNA, and gDNA include naturally occurring molecules (reference molecules), as well as sequence variants.
[0132] In some embodiments, the gNA is a deoxyribonucleic acid molecule ("gDNA"), in some embodiments, the gNA is a ribonucleic acid molecule ("gRNA"), and in other embodiments, the gNA is chimeric and comprises both DNA and RNA.
[0133] The gNAs of the present disclosure comprise two segments: a targeting sequence and a protein-binding segment (which constitute the scaffold discussed herein). The targeting segment of a gNA comprises a nucleotide sequence (interchangeably referred to herein as a guide sequence, spacer, targeting sequence, or targeting region) that is complementary to (and therefore hybridizes with) a specific sequence (target site) within a target nucleic acid sequence (e.g., a target ssRNA, a target ssDNA, the complementary strand of a double-stranded target DNA, etc.), as described in more detail below.
[0134] The targeting sequence of the gNA can bind to target nucleic acid sequences, including coding sequences, complements of coding sequences, non-coding sequences, and regulatory elements. The protein-binding segment (or "protein-binding sequence") interacts with (e.g., binds to) the CasX protein. The protein-binding segment is alternatively referred to herein as a "scaffold." In some embodiments, the targeting sequence and the scaffold each comprise complementary stretches of nucleotides that hybridize to each other to form a double-stranded duplex (e.g., a dsRNA duplex for a gRNA). Site-specific binding and / or cleavage of a target nucleic acid sequence (e.g., genomic DNA) by the CasX:gNA can occur at one or more positions in the target nucleic acid, determined by base pair complementarity between the targeting sequence of the gNA and the target nucleic acid sequence.
[0135] The gNA provides target specificity to the complex by having a nucleotide sequence complementary to the target sequence of the target nucleic acid. The CasX of the complex provides the site-specific activity of the complex, such as binding, cleaving, or nicking of the target sequence of the target nucleic acid by the CasX nuclease, as described below, and / or the activity provided by the fusion partner in the case of a CasX-containing fusion protein. In some embodiments, the present disclosure provides a gene editing pair of CasX and gNA of any of the embodiments described herein that can be combined together before use in gene editing and are thus "pre-complexed" as an RNP. The use of a pre-complexed RNP provides advantages in delivering system components to cells or target nucleic acid sequences for editing the target nucleic acid sequence. The CasX protein of the RNP provides site-specific activity that is directed to (e.g., stabilized at) the target site within the target nucleic acid sequence by association with a guide RNA containing a targeting sequence.
[0136] In some embodiments where the gRNA is a gRNA, the term "targeter" or "targeter RNA" is used herein to refer to the crRNA-like molecule (crRNA: "CRISPR RNA") of a CasX dual guide RNA (dgRNA). In a single guide RNA (sgRNA), the "activator" and "targeter" are linked together, e.g., by an intervening nucleotide). Thus, for example, a guide RNA (dgRNA or sgRNA) comprises a guide sequence and a duplex-forming segment of crRNA, sometimes referred to as a crRNA repeat. Because the targeter sequence of a guide sequence hybridizes with a specific target nucleic acid sequence, the targeter may be modified by the user to hybridize with a desired target nucleic acid sequence. In some embodiments, the sequence of the targeter may often be a non-naturally occurring sequence. The targeter and activator each have a duplex-forming segment, where the duplex-forming segment of the targeter and the duplex-forming segment of the activator are complementary to each other and hybridize to form a double-stranded duplex (a dsRNA duplex for a gRNA). In some embodiments, the targeter contains both the guide sequence of the CasX guide RNA and a stretch of nucleotides that form half of the dsRNA duplex of the protein-binding segment of the gNA. The corresponding tracrRNA-like molecule (the activator "trans-acting CRISPR RNA") also contains a duplex-forming stretch of nucleotides that form the other half of the dsRNA duplex of the protein-binding segment of the CasX guide RNA. In some cases, the activator contains one or more stem-loops that can interact with the CasX protein. Thus, the targeter and activator as a corresponding pair hybridize to form a CasX dual-guide NA, referred to herein as a "dual-guide NA," "dgNA," "dual-molecular guide NA," or "bi-molecular guide NA."
[0137] In some embodiments, the activator and targeter of a reference gNA are covalently linked to each other and comprise a single molecule, referred to herein as a "single-molecule guide NA," "one-molecule guide NA," "single-guide NA," "single-guide RNA," "single-molecule guide RNA," "one-molecule guide RNA," "single-guide DNA," "single-molecule DNA," or "single-molecule guide DNA," ("sgNA," "sgRNA," or "sgDNA"). In some embodiments, an sgNA comprises an "activator" or a "targeter," and thus may be an "activator RNA" and a "targeter RNA," respectively.
[0138] A reference gRNA of the present disclosure comprises four distinct regions or domains: an RNA triplex, a scaffold stem, an extended stem, and a targeting sequence (specific for the target nucleic acid. The RNA triplex, scaffold stem, and extended stem are collectively referred to as the "scaffold" of the reference gRNA, upon which further gNA variants are generated.
[0139] b. RNA triplex In some embodiments of the guide NAs provided herein, the gNA comprises an RNA triplex, which ends with AAAG after two intervening stem loops (a scaffold stem loop and an extended stem loop), UUU--N X The gRNA contains a (approximately 4-15)-UUU stem-loop (SEQ ID NO: 241) sequence, forming a pseudoknot that can extend beyond the triplex into a duplex pseudoknot. The triplex UU-UUU-AAA sequence forms a nexus between the targeting sequence, the scaffold stem, and the extended stem. In an exemplary gRNA, the UUU-loop-UUU region is encoded first, followed by the scaffold stem-loop, then the extended stem-loop connected by the tetraloop, and then AAAG, which terminates the triplex before the targeting sequence.
[0140] c. Scaffold stem loop In some embodiments of gNAs of the present disclosure, the triplex region is followed by a scaffold stem-loop. The scaffold stem-loop is the region of the gNA to which a CasX protein (such as a reference or CasX variant protein) binds. In some embodiments, the scaffold stem-loop is a fairly short and stable stem-loop, increasing the overall stability of the gNA. In some cases, the scaffold stem-loop does not tolerate much variation and requires some form of RNA bubble. In some embodiments, the scaffold stem is necessary for gNA function. While likely similar to the nexus stem of Cas9 in being a critical stem-loop, in some embodiments, the scaffold stem of a gNA possesses a necessary bulge (RNA bubble) that differs from many other stem-loops found in CRISPR / Cas systems. In some embodiments, the presence of this bulge is conserved across gNAs that interact with different CasX proteins. An exemplary sequence of a gNA scaffold stem-loop sequence comprises the sequence CCAGCGACUAUGUCGUAUGG (SEQ ID NO: 242). In other embodiments, the present disclosure provides gNA variants in which the scaffold stem-loop is replaced with an RNA stem-loop sequence from a heterologous RNA source having proximal 5' and 3' ends, such as, but not limited to, a stem-loop sequence selected from MS2, Qβ, U1 hairpin II, Uvsx, or PP7 stem-loop. In some cases, the heterologous RNA stem-loop of the gNA can bind to a protein, RNA structure, DNA sequence, or small molecule.
[0141] d. extended stem-loop In some embodiments of gNAs of the present disclosure, the scaffold stem-loop is followed by an extended stem-loop. In some embodiments, the extended stem comprises a synthetic tracr and crRNA fusion without the majority of the CasX protein bound. In some embodiments, the extended stem-loop can be highly flexible. In some embodiments, a single-guide gRNA is created using a GAAA tetraloop linker or a GAGAAA linker between the tracr and crRNA within the extended stem-loop. In some cases, the targeter and activator of the sgNA are linked to each other by an intervening nucleotide, and the linker can have a length of 3 to 20 nucleotides. In some embodiments of sgNAs of the present disclosure, the extended stem is a large 32-bp loop located outside the CasX protein of the ribonucleoprotein complex. An exemplary sequence of an extended stem-loop sequence of a sgNA comprises the sequence GCGCUUAUUUAUCGGAGAGAAAUCCGAUAAAUAAGAAGC (SEQ ID NO: 15). In some embodiments, the extended stem-loop comprises a GAGAAA spacing sequence. In some embodiments, the present disclosure provides gNA variants in which the extended stem-loop is replaced with an RNA stem-loop sequence from a heterologous RNA source having proximal 5' and 3' ends, such as, but not limited to, a stem-loop sequence selected from MS2, Qβ, U1 hairpin II, Uvsx, or PP7 stem-loop. In such cases, the heterologous RNA stem-loop increases the stability of the gNA. In other embodiments, the present disclosure provides gNA variants with an extended stem-loop region comprising at least 10, at least 100, at least 500, at least 1000, or at least 10,000 nucleotides.
[0142] e. targeting sequence In some embodiments of gNAs of the present disclosure, the extended stem-loop is followed by a region that forms part of a triplex, followed by a targeting sequence (or "spacer"). The targeting sequence can be designed to target the CasX ribonucleoprotein holocomplex to a specific region of the target nucleic acid sequence. Thus, the gNA targeting sequence of the gNAs of the present disclosure has a sequence that is complementary to, and can therefore hybridize with, a portion of a target nucleic acid within the nucleic acid of a eukaryotic cell (e.g., a eukaryotic chromosome, a chromosomal sequence, a eukaryotic RNA, etc.) as a component of an RNP, when any one of the PAM sequences TTC, ATC, GTC, or CTC is located one nucleotide 5' to the non-target strand sequence that is complementary to the target sequence.
[0143] In some embodiments, the present disclosure provides gNAs whose targeting sequence is complementary to a target nucleic acid sequence containing one or more mutations compared to the wild-type gene sequence, for the purpose of editing the mutation-containing sequence using the CasX:gNA system of the present disclosure. In some embodiments, the targeting sequence of the gNA is designed to be specific to an exon of the gene of the target nucleic acid. In other embodiments, the targeting sequence of the gNA is designed to be specific to an intron of the gene of the target nucleic acid. In other embodiments, the targeting sequence of the gNA is designed to be specific to an intron-exon junction of the gene of the target nucleic acid. In other embodiments, the targeting sequence of the gNA is designed to be specific to a regulatory element of the gene of the target nucleic acid. In some embodiments, the targeting sequence of the gNA is designed to be complementary to a sequence containing one or more single nucleotide polymorphisms (SNPs) in the gene of the target nucleic acid. SNPs in both coding and non-coding sequences are within the scope of the present disclosure. In other embodiments, the targeting sequence of the gNA is designed to be complementary to a sequence in an intergenic region of the gene of the target nucleic acid.
[0144] In some embodiments, the targeting sequence of the gNA is designed to be specific to the regulatory element that regulates the expression of the gene product of the target nucleic acid. Such regulatory elements include, but are not limited to, promoter regions, enhancer regions, intergenic regions, 5' untranslated regions (5'UTR), 3' untranslated regions (3'UTR), conserved elements, and regions containing cis-regulatory elements. The promoter region is intended to include nucleotides within 5 kb from the start of the coding sequence, or in the case of a gene enhancer element or conserved element, it may be thousands, hundreds of thousands, or even millions of bp away from the coding sequence of the gene of the target nucleic acid. In some of the above embodiments, the target is intended to knock out or knock down the coding gene of the target, so that the encoded protein containing the mutation is not expressed or is expressed at a lower level in the cell.
[0145] In some embodiments, the targeting sequence of the gNA comprises 14 to 35 contiguous nucleotides. In some embodiments, the targeting sequence comprises 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, or 35 contiguous nucleotides. In some embodiments, the targeting sequence of the gNA comprises 20 contiguous nucleotides. In some embodiments, the targeting sequence comprises 19 contiguous nucleotides. In some embodiments, the targeting sequence comprises 18 contiguous nucleotides. In some embodiments, the targeting sequence comprises 17 contiguous nucleotides. In some embodiments, the targeting sequence comprises 16 contiguous nucleotides. In some embodiments, the targeting sequence comprises 15 contiguous nucleotides. In some embodiments, the targeting sequence has 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, or 35 contiguous nucleotides, and the targeting sequence contains 0 to 5, 0 to 4, 0 to 3, or 0 to 2 mismatches to the target nucleic acid sequence, and can retain sufficient binding specificity so that an RNP containing a gNA containing the targeting sequence can form a complementary bond to the target nucleic acid.
[0146] In some embodiments, the CasX:gNA system includes a first gNA and further includes a second (and optionally a third, fourth, fifth, or more) gNA, where the second or additional gNA has a targeting sequence complementary to a different or overlapping portion of the target nucleic acid sequence compared to the targeting sequence of the first gNA, thereby targeting multiple points within the target nucleic acid, e.g., introducing multiple cuts into the target nucleic acid by CasX. In such cases, it will be understood that the second or additional gNA forms a complex with additional copies of the CasX protein. By selecting the targeting sequence of the gNA, a defined region of the target nucleic acid sequence surrounding a mutation can be modified or edited using the CasX:gNA system described herein, including facilitating the insertion of a donor template.
[0147] f. gNA scaffold Except for the targeting sequence region, the remaining region of the gNA is referred to herein as a scaffold.In some embodiments, the gNA scaffold is derived from a naturally occurring sequence, which is described below as a reference gNA.In other embodiments, the gNA scaffold is a variant of the reference gNA, which is introduced with mutations, insertions, deletions, or domain substitutions to give the gNA desirable properties.
[0148] In some embodiments, the reference gRNA comprises a sequence isolated or derived from Deltaproteobacteria. In some embodiments, the sequence is a CasX tracrRNA sequence. Exemplary CasX reference tracrRNA sequences isolated or derived from Deltaproteobacteria may include ACAUCUGGCGCGUUUAUUCCAUUACUUUGGAGCCAGUCCCAGCGACUAUGUCGUAUGGACGAAGCGCUUAUUUAUCGGAGA (SEQ ID NO: 6) and ACAUCUGGCGCGUUUAUUCCAUUACUUUGGAGCCAGUCCCAGCGACUAUGUCGUAUGGACGAAGCGCUUAUUUAUCGG (SEQ ID NO: 7). An exemplary crRNA sequence isolated or derived from Deltaproteobacteria may comprise the sequence CCGAUAAGUAAAACGCAUCAAAG (SEQ ID NO: 243). In some embodiments, the reference gNA comprises a sequence that is at least 60% identical, at least 65% identical, at least 70% identical, at least 75% identical, at least 80% identical, at least 81% identical, at least 82% identical, at least 83% identical, at least 84% identical, at least 85% identical, at least 86% identical, at least 86% identical, at least 87% identical, at least 88% identical, at least 89% identical, at least 89% identical, at least 90% identical, at least 91% identical, at least 92% identical, at least 93% identical, at least 94% identical, at least 95% identical, at least 96% identical, at least 97% identical, at least 98% identical, at least 99% identical, at least 99.5% identical, or 100% identical to a sequence isolated or derived from Deltaproteobacteria.
[0149] In some embodiments, the reference guide RNA comprises a sequence isolated or derived from Planctomycetes. In some embodiments, the sequence is a CasX tracrRNA sequence. Exemplary reference tracrRNA sequences isolated or derived from Planctomycetes include UACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUGGGUAAAGCGCUUAUUUAUCGGAGA (SEQ ID NO: 8) and An exemplary crRNA sequence isolated from or derived from Planctomycetes may include the sequence UCUCCGAUAAAUAAGAAGCAUCAAAG (SEQ ID NO: 244). In some embodiments, the reference gNA comprises a sequence that is at least 60% identical, at least 65% identical, at least 70% identical, at least 75% identical, at least 80% identical, at least 81% identical, at least 82% identical, at least 83% identical, at least 84% identical, at least 85% identical, at least 86% identical, at least 86% identical, at least 87% identical, at least 88% identical, at least 89% identical, at least 89% identical, at least 90% identical, at least 91% identical, at least 92% identical, at least 93% identical, at least 94% identical, at least 95% identical, at least 96% identical, at least 97% identical, at least 98% identical, at least 99% identical, at least 99.5% identical, or 100% identical to a sequence isolated or derived from Planctomycetes.
[0150] In some embodiments, the reference gNA comprises a sequence isolated or derived from Candidatus Sungbacteria. In some embodiments, the sequence is a CasX tracrRNA sequence. Exemplary CasX reference tracrRNA sequences isolated or derived from Candidatus Sungbacteria may include the sequences GUUUACACACUCCCUCUCAUAGGGU (SEQ ID NO: 10), GUUUACACACUCCCUCUCAUGAGGU (SEQ ID NO: 11), UUUUACAUACCCCCUCUCAUGGGAU (SEQ ID NO: 12), and GUUUACACACUCCCUCUCAUGGGGG (SEQ ID NO: 13). In some embodiments, the reference guide RNA comprises a sequence that is at least 60% identical, at least 65% identical, at least 70% identical, at least 75% identical, at least 80% identical, at least 81% identical, at least 82% identical, at least 83% identical, at least 84% identical, at least 85% identical, at least 86% identical, at least 86% identical, at least 87% identical, at least 88% identical, at least 89% identical, at least 89% identical, at least 90% identical, at least 91% identical, at least 92% identical, at least 93% identical, at least 94% identical, at least 95% identical, at least 96% identical, at least 97% identical, at least 98% identical, at least 99% identical, at least 99.5% identical, or 100% identical to a sequence isolated or derived from Candidatus Sungbacteria.
[0151] Table 1 provides sequences of reference gRNAs tracr, cr, and scaffold sequences. In some embodiments, the present disclosure provides gNA sequences having a scaffold comprising a sequence having at least one nucleotide modification compared to a reference gNA sequence having the sequence of any one of SEQ ID NOS: 4-16 in Table 1. In these embodiments in which the vector comprises DNA encoding a sequence for the gNA, or the gNA is a gDNA or a chimera of RNA and DNA, it will be understood that thymine (T) bases can be substituted for uracil (U) bases in any of the gNA sequence embodiments described herein. [Table 1]
[0152] g.gNA variant In another aspect, the present disclosure relates to guide nucleic acid variants (alternatively referred to herein as "gNA variants" or "gRNA variants") that contain one or more modifications compared to a reference gRNA scaffold. As used herein, "scaffold" refers to all portions of a gNA required for gNA function, excluding spacer sequences.
[0153] In some embodiments, the gNA variants comprise one or more nucleotide substitutions, insertions, deletions, or exchanged or substituted regions compared to the reference gRNA sequence of the present disclosure. In some embodiments, mutations can occur in any region of the reference gRNA scaffold to generate the gNA variant. In some embodiments, the scaffold of the gNA variant sequence has at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, or at least 70%, at least 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to the sequence of SEQ ID NO: 4 or SEQ ID NO: 5.
[0154] In some embodiments, the gNA variant comprises one or more nucleotide changes in one or more regions of the reference gRNA scaffold that improve the characteristics of the reference gRNA. Exemplary regions include an RNA triplex, a pseudoknot, a scaffold stem-loop, and an extended stem-loop. In some cases, the variant scaffold stem further comprises a bubble. In other cases, the variant scaffold further comprises a triplex loop region. In still other cases, the variant scaffold further comprises a 5' unstructured region. In some embodiments, the gNA variant scaffold comprises a scaffold stem-loop that has at least 60% sequence identity, at least 70% sequence identity, at least 80% sequence identity, at least 90% sequence identity, at least 95% sequence identity, or at least 99% sequence identity to SEQ ID NO: 14. In some embodiments, the gNA variant scaffold comprises a scaffold stem-loop that has at least 60% sequence identity to SEQ ID NO: 14. In other embodiments, the gNA variant comprises a scaffold stem-loop with the sequence CCAGCGACUAUGUCGUAGUGG (SEQ ID NO: 245). In other embodiments, the disclosure provides a gNA scaffold comprising a modified extended stem-loop, compared to SEQ ID NO: 5, with a C18G substitution, a G55 insertion, a U1 deletion, and the original 6 nt loop and 13 most proximal loop base pairs (32 nucleotides total) replaced by a Uvsx hairpin (4 nt loop and 5 proximal loop base pairs, 14 nucleotides total), and the distal loop bases of the extended stem converted into a fully base-paired stem adjacent to the new Uvsx hairpin by deletion of A99 and substitution of G65U. In the foregoing embodiment, the gNA scaffold comprises the sequence ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGUGGGUAAAGCUCCCUCUUCGGAGGGAGCAUCAAAG (SEQ ID NO: 2238).
[0155] All gNA variants in which the variant gNA has one or more improved characteristics or one or more new functions added compared to the reference gRNA described herein are considered to be within the scope of the present disclosure. A representative example of such a gNA variant is guide 174 (SEQ ID NO: 2238), the design of which is described in the Examples. In some embodiments, the gNA variant adds new functions to the RNP containing the gNA variant. In some embodiments, the gNA variant has an improved characteristic selected from improved stability, improved solubility, improved transcription of the gNA, improved resistance to nuclease activity, increased folding rate of the gNA, reduced by-product formation during folding, increased productive folding, improved binding affinity for the CasX protein, improved binding affinity for the target DNA when complexed with the CasX protein, improved gene editing when complexed with the CasX protein, improved editing specificity when complexed with the CasX protein, and an improved ability to utilize a broader spectrum of one or more PAM sequences, including ATC, CTC, GTC, or TTC, in editing the target DNA when complexed with the CasX protein, and any combination thereof. In some cases, one or more of the improved characteristics of the gNA variant is improved by at least about 1.1 to about 100,000 fold compared to the reference gNA of SEQ ID NO: 4 or SEQ ID NO: 5. In other cases, one or more improved characteristics of the gNA variant are improved by at least about 1.1, at least about 10, at least about 100, at least about 1000, at least about 10,000, at least about 100,000 or more times compared to the reference gNA of SEQ ID NO: 4 or SEQ ID NO: 5.In other cases, one or more of the improved features of the gNA variant may be about 1.1-100,00 fold, about 1.1-10,00 fold, about 1.1-1,000 fold, about 1.1-500 fold, about 1.1-100 fold, about 1.1-50 fold, about 1.1-20 fold, about 10-100,00 fold, about 10-10,00 fold, about 10-1,000 fold, about 10-500 fold, about 10-100 fold, about 10-50 fold, about 10-20 fold, about 2-70 fold, about 2-50 fold, about 2-30 fold, about 2-20 fold, about 2-10 fold, about 5-50 fold, about 5 to 30 times, about 5 to 10 times, about 100 to 100,00 times, about 100 to 10,00 times, about 100 to 1,000 times, about 100 to 500 times, about 500 to 100,00 times, about 500 to 10,000 times, about 500 to 1,000 times, about 500 to 750 times, about 1,000 to 100,00 times, about 10,000 to 100,00 times, about 20 to 500 times, about 20 to 250 times, about 20 to 200 times, about 20 to 100 times, about 20 to 50 times, about 50 to 10,000 times, about 50 to 1,000 times, about 50 to 500 times, about 50 to 200 times, or about 50 to 100 times improvement. In other cases, one or more improved characteristics of the gNA variant may be about 1.1-fold, 1.2-fold, 1.3-fold, 1.4-fold, 1.5-fold, 1.6-fold, 1.7-fold, 1.8-fold, 1.9-fold, 2-fold, 3-fold, 4-fold, 5-fold, 6-fold, 7-fold, 8-fold, 9-fold, 10-fold, 11-fold, 12-fold, 13-fold, 14-fold, 15-fold, 16-fold, 17-fold, 18-fold, 19-fold, 20-fold, 25-fold, 30-fold, 40-fold, 45-fold, 50-fold, 55-fold, 60-fold, 75-fold, 80-fold, 85-fold, 86-fold, 87-fold, 88-fold, 89-fold, 90-fold, 91-fold, 92-fold, 93-fold, 94-fold, 95-fold, 96-fold, 97-fold, 98-fold, 99-fold, 100-fold, 101-fold, 102-fold, 103-fold, 104-fold, 105-fold, 106-fold, 107-fold, 108-fold, 109-fold, 110-fold, 111-fold, 112-fold, 113-fold, 114-fold, 115-fold, 116-fold, 117-fold, 118-fold, 119-fold, 120-fold, 121-fold, 122-fold, 123-fold, 124-fold, 125-fold 0x, 80x, 90x, 100x, 110x, 120x, 130x, 140x, 150x, 160x, 170x, 180x, 190x, 200x, 210x, 220x, 230x, 240x, 250x, 260x, 270x, 280x, 290x, 300x, 310x, 320x, 330x, 340x, 350x, 360x, 370x, 380x, 390x, 400x, 425x, 450x, 475x, or 500x improvement.
[0156] In some embodiments, gNA variants can be generated by subjecting a reference gNA to one or more mutagenesis methods, such as those described herein below, which may include deep mutational evolution (DME), deep mutational scanning (DMS), error-prone PCR, cassette mutagenesis, random mutagenesis, staggered extension PCR, gene shuffling, or domain swapping, to generate the gNA variants of the present disclosure. The activity of the reference gNA can be used as a benchmark against which the activity of the gNA variants can be compared, thereby measuring the improvement in function of the gNA variant. In other embodiments, the reference gNA can be subjected to one or more deliberate, targeted mutations, substitutions, or domain swaps to generate gNA variants, such as rationally designed variants. Exemplary gNA variants generated by such methods are described in the Examples, and representative sequences of gNA scaffolds are shown in Table 2.
[0157] In some embodiments, the gNA variants comprise one or more modifications compared to a reference guide nucleic acid scaffold sequence, the one or more modifications being selected from at least one nucleotide substitution in a region of the reference gNA, at least one nucleotide deletion in a region of the reference gNA, at least one nucleotide insertion in a region of the reference gNA, a substitution of all or part of a region of the reference gNA, a deletion of all or part of a region of the reference gNA, or any combination of the foregoing. In some cases, the modification is a substitution of 1 to 15 consecutive or non-consecutive nucleotides of the reference gNA in one or more regions. In other cases, the modification is a deletion of 1 to 10 consecutive or non-consecutive nucleotides of the reference gNA in one or more regions. In other cases, the modification is an insertion of 1 to 10 consecutive or non-consecutive nucleotides of the reference gNA in one or more regions. In other cases, the modification is a replacement of the scaffold stem-loop or extended stem-loop with an RNA stem-loop sequence from a heterologous RNA source having proximal 5' and 3' ends. In some cases, the gNA variants of the present disclosure comprise two or more modifications in a single region compared to a reference gRNA. In other cases, the gNA variants of the present disclosure comprise modifications in two or more regions. In other cases, the gNA variants comprise any combination of the modifications described in this paragraph. In some embodiments, exemplary modifications of the gNAs of the present disclosure comprise the modifications in Table 24.
[0158] In some embodiments, because transcription from the U6 promoter is more efficient and consistent with respect to the start site when the +1 nucleotide is G, a 5' G is added to the gNA variant sequence for in vivo expression compared to the reference gRNA. In other embodiments, because T7 polymerase strongly prefers G at the +1 position and a purine at the +2 position, two 5' Gs are added to generate gNA variant sequences for in vitro transcription to increase production efficiency. In some cases, a 5' G base is added to the reference scaffold in Table 1. In other cases, a 5' G base is added to the variant scaffold in Table 2.
[0159] Table 2 provides exemplary gNA variant scaffold sequences of the present disclosure. In Table 2, (-) indicates a deletion at the specified position relative to the reference sequence of SEQ ID NO: 5, (+) indicates an insertion of a specified base at the indicated position relative to SEQ ID NO: 5, and (:) indicates a range of bases at the specified start:stop coordinates of the deletion or substitution relative to SEQ ID NO: 5, with multiple insertions, deletions, or substitutions separated by commas, e.g., A14C, T17G. In some embodiments, the gNA variant scaffold comprises any one of SEQ ID NOs: 2101-2280, sequences set forth in Table 2, or a sequence having at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity thereto. In those embodiments in which the vector comprises DNA encoding a sequence for a gNA, or the gNA is a chimera of gDNA or RNA and DNA, it will be understood that thymine (T) bases can be substituted for uracil (U) bases in any of the gNA sequence embodiments described herein. [Table 2-1] [Table 2-2] [Table 2-3] [Table 2-4] [Table 2-5] [Table 2-6] [Table 2-7] [Table 2-8] [Table 2-9] [Table 2-10] [Table 2-11] [Table 2-12] [Table 2-13] [Table 2-14]
[0160] In some embodiments, the gNA variant comprises a tracrRNA stem loop comprising the sequence -UUU-N4-25-UUU- (SEQ ID NO: 240). For example, the gNA variant comprises a scaffold stem loop or a substitution thereof flanked by two triplet U motifs that contribute to the triplex region. In some embodiments, the scaffold stem loop or a substitution thereof comprises at least 4 nucleotides, at least 5 nucleotides, at least 6 nucleotides, at least 7 nucleotides, at least 7 nucleotides, at least 8 nucleotides, at least 9 nucleotides, at least 10 nucleotides, at least 11 nucleotides, at least 12 nucleotides, at least 13 nucleotides, at least 14 nucleotides, at least 15 nucleotides, at least 16 nucleotides, at least 17 nucleotides, at least 18 nucleotides, at least 19 nucleotides, at least 20 nucleotides, at least 21 nucleotides, at least 22 nucleotides, at least 23 nucleotides, at least 24 nucleotides, or at least 25 nucleotides.
[0161] In some embodiments, the gNA variant comprises a crRNA sequence with -AAAG- at the 5' position relative to the spacer region. In some embodiments, the -AAAG- sequence is immediately 5' to the spacer region.
[0162] In some embodiments, at least one nucleotide modification to a reference gNA to produce a gNA variant comprises at least one nucleotide deletion in the CasX variant gNA compared to the reference gRNA. In some embodiments, the gNA variant comprises a deletion of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 consecutive or non-consecutive nucleotides compared to the reference gNA. In some embodiments, the at least one deletion comprises a deletion of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 or more consecutive nucleotides compared to the reference gNA. In some embodiments, the gNA variant comprises a deletion of 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 or more nucleotides compared to the reference gRNA, and the deletion is not in a contiguous nucleotide sequence. In these embodiments, where the gNA variant has two or more non-contiguous deletions compared to the reference gRNA, any length of the deletion and any combination of deletion lengths as described herein are contemplated within the scope of the present disclosure. For example, in some embodiments, the gNA variant may comprise a first deletion of one nucleotide and a second deletion of two nucleotides, and the two deletions are not contiguous. In some embodiments, the gNA variant comprises at least two deletions in different regions of the reference gRNA. In some embodiments, the gNA variant comprises at least two deletions in the same region of the reference gRNA. For example, the region can be an extended stem-loop, a scaffold stem-loop, a scaffold stem-bubble, a triplex loop, a pseudoknot, a triplex, or the 5' end of a gRNA variant. Any nucleotide deletion in the reference gRNA is contemplated as being within the scope of this disclosure.
[0163] In some embodiments, at least one nucleotide modification of the reference gRNA to generate a gNA variant comprises at least one nucleotide insertion. In some embodiments, the gNA variant comprises an insertion of 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 consecutive or non-consecutive nucleotides compared to the reference gRNA. In some embodiments, the insertion of at least one nucleotide comprises an insertion of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 or more consecutive nucleotides compared to the reference gRNA. In some embodiments, the gNA variant comprises two or more insertions compared to the reference gRNA, and the insertions are not contiguous. In these embodiments, where there are two or more non-contiguous insertions in the gNA variant compared to the reference gRNA, any length of the insertion and any combination of insertion lengths as described herein are contemplated within the scope of the present disclosure. For example, in some embodiments, a gNA variant may include a first insertion of one nucleotide and a second insertion of two nucleotides, where the two insertions are not contiguous. In some embodiments, a gNA variant includes at least two insertions in different regions of a reference gRNA. In some embodiments, a gNA variant includes at least two insertions in the same region of a reference gRNA. For example, the region may be an extended stem-loop, a scaffold stem-loop, a scaffold stem-bubble, a triplex loop, a pseudoknot, a triplex, or the 5' end of the gNA variant. Any insertion of A, G, C, U (or T in the corresponding DNA) or a combination thereof at any position in the reference gRNA is contemplated within the scope of the present disclosure.
[0164] In some embodiments, at least one nucleotide modification of a reference gRNA to generate a gNA variant comprises at least one nucleic acid substitution. In some embodiments, the gNA variant comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 or more consecutive or non-consecutive substituted nucleotides compared to the reference gRNA. In some embodiments, the gNA variant comprises 1 to 4 nucleotide substitutions compared to the reference gRNA. In some embodiments, at least one substitution comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 or more consecutive nucleotide substitutions compared to the reference gRNA. In some embodiments, the gNA variant comprises two or more substitutions compared to the reference gRNA, and the substitutions are not consecutive. In these embodiments in which there are two or more non-consecutive substitutions in a gNA variant compared to a reference gRNA, any length of the substituted nucleotides and any combination of lengths of the substituted nucleotides as described herein are contemplated within the scope of the present disclosure. For example, in some embodiments, a gNA variant may include a first substitution of one nucleotide and a second substitution of two nucleotides, where the two substitutions are not consecutive. In some embodiments, a gNA variant includes at least two substitutions in different regions of a reference gRNA. In some embodiments, a gNA variant includes at least two substitutions in the same region of a reference gRNA. For example, the region may be a triplex, an extended stem-loop, a scaffold stem-loop, a scaffold stem-bubble, a triplex loop, a pseudoknot, a triplex, or the 5' end of a gNA variant. Any substitution of A, G, C, U (or T in the corresponding DNA) or a combination thereof at any position in a reference gRNA is contemplated within the scope of the present disclosure.
[0165] Any of the substitutions, insertions and deletions described herein can be combined to generate the gNA variant of the present disclosure.For example, gNA variant can comprise at least one substitution and at least one deletion compared to reference gRNA, at least one substitution and at least one insertion compared to reference gRNA, at least one insertion and at least one deletion compared to reference gRNA, or at least one substitution, one insertion and one deletion compared to reference gRNA.
[0166] In some embodiments, the gNA variant comprises a scaffold region that is at least 20% identical, at least 30% identical, at least 40% identical, at least 50% identical, at least 60% identical, at least 65% identical, at least 70% identical, at least 75% identical, at least 80% identical, at least 85% identical, at least 90% identical, at least 91% identical, at least 92% identical, at least 93% identical, at least 94% identical, at least 95% identical, at least 96% identical, at least 97% identical, at least 98% identical, or at least 99% identical to any one of SEQ ID NOs: 4-16. In some embodiments, the gNA variant comprises a scaffold region that is at least 60% homologous (or identical) to any one of SEQ ID NOs: 4-16.
[0167] In some embodiments, the gNA variant comprises a tracr stem loop that is at least 60% identical, at least 65% identical, at least 70% identical, at least 75% identical, at least 80% identical, at least 85% identical, at least 90% identical, at least 91% identical, at least 92% identical, at least 93% identical, at least 94% identical, at least 95% identical, at least 96% identical, at least 97% identical, at least 98% identical, or at least 99% identical to SEQ ID NO: 14. In some embodiments, the gNA variant comprises a tracr stem loop that is at least 60% homologous (or identical) to SEQ ID NO: 14.
[0168] In some embodiments, the gNA variant comprises an extended stem loop that is at least 60% identical, at least 65% identical, at least 70% identical, at least 75% identical, at least 80% identical, at least 85% identical, at least 90% identical, at least 91% identical, at least 92% identical, at least 93% identical, at least 94% identical, at least 95% identical, at least 96% identical, at least 97% identical, at least 98% identical, or at least 99% identical to SEQ ID NO: 15. In some embodiments, the gNA variant comprises an extended stem loop that is at least 60% homologous (or identical) to SEQ ID NO: 15.
[0169] In some embodiments, the gNA variant comprises the sequence of any one of SEQ ID NOs: 412-3295. In some embodiments, the gNA variant comprises the sequence of any one of SEQ ID NOs: 2236, 2237, 2238, 2241, 2244, 2248, 2249, or 2259-2280. In some embodiments, the gNA variant comprises the sequence of any one of SEQ ID NOs: 2236, 2237, 2238, 2241, 2244, 2248, 2249, or 2259-2280.
[0170] In some embodiments, the gNA variant comprises an exogenous extended stem-loop with such differences from a reference gNA as described below: In some embodiments, the exogenous extended stem-loop has little or no identity to a reference stem-loop region disclosed herein (e.g., SEQ ID NO: 15). In some embodiments, the exogenous stem loop is at least 10 bp, at least 20 bp, at least 30 bp, at least 40 bp, at least 50 bp, at least 60 bp, at least 70 bp, at least 80 bp, at least 90 bp, at least 100 bp, at least 200 bp, at least 300 bp, at least 400 bp, at least 500 bp, at least 600 bp, at least 700 bp, at least 800 bp, at least 900 bp, at least 1,000 bp, at least 2,000 bp, at least 3,000 bp, at least 4,000 bp, at least 5,000 bp, at least 6,000 bp, at least 7,000 bp, at least 8,000 bp, at least 9,000 bp, at least 10,000 bp, at least 12,000 bp, at least 15,000 bp, or at least 20,000 bp. In some embodiments, the gNA variant comprises an extended stem-loop region comprising at least 10, at least 100, at least 500, at least 1000, or at least 10,000 nucleotides. In some embodiments, the heterologous stem-loop increases the stability of the gNA. In some embodiments, the heterologous RNA stem-loop can bind to a protein, an RNA structure, a DNA sequence, or a small molecule.In some embodiments, the exogenous stem-loop region is an RNA stem-loop or hairpin, e.g., a thermostable RNA, such as MS2 (ACAUGAGGAUUACCCAUGU; SEQ ID NO: 4278), Qβ (UGCAUGUCUAAGACAGCA; SEQ ID NO: 4279), U1 hairpin II (AAUCCAUUGCACUCCGGAUU; SEQ ID NO: 4280), Uvsx (CCUCUUCGGAGG; SEQ ID NO: 4281), PP7 (AGGAGUUUCUAUGGAAACCCU; SEQ ID NO: 4282), phage replication loop (AGGUGGGACGACCUCUCGGUCGUCCUAUCU; SEQ ID NO: 4283), kissing loop_a (UGCUCGCUCCGUUCGAGCA; SEQ ID NO: 4284), kissing loop The exogenous stem loop may comprise a loop_b1 (UGCUCGACGCGUCCUCGAGCA; SEQ ID NO: 4285), a kissing loop_b2 (UGCUCGUUUGCGGCUACGAGCA; SEQ ID NO: 4286), a G-quadriplex M3q (AGGGAGGGAGGGAGAGG; SEQ ID NO: 4287), a G-quadriplex telomere basket (GGUUAGGGUUAGGGUUAGG; SEQ ID NO: 4288), a sarcin-ricin loop (CUGCUCAGUACGAGAGGAACCGCAG; SEQ ID NO: 4289), or a pseudoknot (UACACUGGGAUCGCUGAAUUAGAGAUCGGCGUCCUUUCAUUCUAUAUACUUUGGAGUUUUAAAAUGUCUCUAAGUACA; SEQ ID NO: 4290). In some embodiments, the exogenous stem loop comprises an RNA scaffold. As used herein, "RNA scaffold" refers to a multidimensional RNA structure that can interact with, organize, or localize one or more proteins. In some embodiments, the RNA scaffold is synthetic or does not exist in nature. In some embodiments, the exogenous stem-loop comprises a long non-coding RNA (lncRNA). As used herein, lncRNA refers to a non-coding RNA that is longer than about 200 bp in length. In some embodiments, the 5' and 3' ends of the exogenous stem-loop form base pairs, i.e., interact to form a duplex RNA region.In some embodiments, the 5' and 3' ends of the exogenous stem-loop are base-paired, and one or more regions between the 5' and 3' ends of the exogenous stem-loop are not base-paired. In some embodiments, the at least one nucleotide modification comprises (a) a substitution of 1 to 15 consecutive or non-consecutive nucleotides of the gNA variant in one or more regions, (b) a deletion of 1 to 10 consecutive or non-consecutive nucleotides of the gNA variant in one or more regions, (c) an insertion of 1 to 10 consecutive or non-consecutive nucleotides of the gNA variant in one or more regions, (d) a substitution of the scaffold stem-loop or extended stem-loop with an RNA stem-loop sequence from a heterologous RNA source having proximal 5' and 3' ends, or any combination of (a)-(d).
[0171] In some embodiments, the gNA variant comprises the sequence or subsequence of any one of SEQ ID NOs: 412-3295 and the sequence of an exogenous stem-loop. In some embodiments, the gNA variant comprises the sequence or subsequence of any one of SEQ ID NOs: 2236, 2237, 2238, 2241, 2244, 2248, 2249, or 2259-2280 and the sequence of an exogenous stem-loop. In some embodiments, the gNA variant comprises the sequence or subsequence of any one of SEQ ID NOs: 2236, 2237, 2238, 2241, 2244, 2248, 2249, or 2259-2280 and the sequence of an exogenous stem-loop.
[0172] In some embodiments, the gNA variant comprises a scaffold stem loop having at least 60% identity to SEQ ID NO: 14. In some embodiments, the gNA variant comprises a scaffold stem loop having at least 60% identity, at least 70% identity, at least 80% identity, at least 90% identity, at least 95% identity, at least 98% identity, or at least 99% identity to SEQ ID NO: 14. In some embodiments, the gNA variant comprises a scaffold stem loop comprising SEQ ID NO: 14.
[0173] In some embodiments, the gNA variant comprises a scaffold stem-loop sequence of CCAGCGACUAUGUCGUAGUGG (SEQ ID NO: 245). In some embodiments, the gNA variant comprises a scaffold stem-loop sequence of CCAGCGACUAUGUCGUAGUGG (SEQ ID NO: 245) with at least 1, 2, 3, 4, or 5 mismatches thereto.
[0174] In some embodiments, the gNA variant comprises an extended stem-loop region comprising fewer than 32 nucleotides, fewer than 31 nucleotides, fewer than 30 nucleotides, fewer than 29 nucleotides, fewer than 28 nucleotides, fewer than 27 nucleotides, fewer than 26 nucleotides, fewer than 25 nucleotides, fewer than 24 nucleotides, fewer than 23 nucleotides, fewer than 22 nucleotides, fewer than 21 nucleotides, or fewer than 20 nucleotides. In some embodiments, the gNA variant comprises an extended stem-loop region comprising fewer than 32 nucleotides. In some embodiments, the gNA variant further comprises a thermostable stem-loop.
[0175] In some embodiments, the sgRNA variant comprises the sequence of SEQ ID NO:2104, 2106, SEQ ID NO:2163, SEQ ID NO:2107, SEQ ID NO:2164, SEQ ID NO:2165, SEQ ID NO:2166, SEQ ID NO:2103, SEQ ID NO:2167, SEQ ID NO:2105, SEQ ID NO:2108, SEQ ID NO:2112, SEQ ID NO:2160, SEQ ID NO:2170, SEQ ID NO:2114, SEQ ID NO:2171, SEQ ID NO:2112, SEQ ID NO:2173, SEQ ID NO:2102, SEQ ID NO:2174, SEQ ID NO:2175, SEQ ID NO:2109, SEQ ID NO:2176, SEQ ID NO:2238, SEQ ID NO:2239, SEQ ID NO:2240, or SEQ ID NO:2241.
[0176] In some embodiments, the gNA variant comprises one or more additional changes to the sequence of any one of SEQ ID NOs: 2201-2280. In some embodiments, the gNA variant comprises or has at least about 80%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to the sequence of any one of SEQ ID NOs: 2236, 2237, 2238, 2241, 2244, 2248, 2249, or 2259-2280. In some embodiments, the gNA variant comprises one or more additional changes to the sequence of any one of SEQ ID NOs: 2201-2280. In some embodiments, the gNA variant comprises the sequence of any one of SEQ ID NOs: 2236, 2237, 2238, 2241, 2244, 2248, 2249, or 2259-2280.
[0177] In some embodiments, the sgRNA variant comprises one or more additional changes to the sequence of SEQ ID NO:2104, SEQ ID NO:2163, SEQ ID NO:2107, SEQ ID NO:2164, SEQ ID NO:2165, SEQ ID NO:2166, SEQ ID NO:2103, SEQ ID NO:2167, SEQ ID NO:2105, SEQ ID NO:2108, SEQ ID NO:2112, SEQ ID NO:2160, SEQ ID NO:2170, SEQ ID NO:2114, SEQ ID NO:2171, SEQ ID NO:2112, SEQ ID NO:2173, SEQ ID NO:2102, SEQ ID NO:2174, SEQ ID NO:2175, SEQ ID NO:2109, SEQ ID NO:2176, SEQ ID NO:2238, SEQ ID NO:2239, SEQ ID NO:2240, or SEQ ID NO:2241.
[0178] In some embodiments of gNA variants of the present disclosure, the gNA variant comprises at least one modification, relative to the reference guide scaffold of SEQ ID NO: 5, selected from one or more of: (a) a C18G substitution in the triplex loop, (b) a G55 insertion in the stem bubble, (c) a U1 deletion, or (d) an extended stem-loop modification in which (i) the 6-nt loop and 13 loop-proximal base pairs are replaced by a Uvsx hairpin, and (ii) an A99 deletion and G65U substitution resulting in a perfectly base-paired loop-distal base. In such embodiments, the gNA variant comprises the sequence of any one of SEQ ID NOs: 2236, 2237, 2238, 2241, 2244, 2248, 2249, or 2259-2280.
[0179] In some embodiments, the gNA variant scaffold comprises the sequence of any one of SEQ ID NOs: 2201-2280 in Table 2. In some embodiments, the gNA scaffold consists of, or consists essentially of, the sequence of any one of SEQ ID NOs: 2201-2280. In some embodiments, the gNA variant sequence scaffold is at least about 60% identical, at least about 65% identical, at least about 70% identical, at least about 75% identical, at least about 80% identical, at least about 85% identical, at least about 90% identical, at least about 91% identical, at least about 92% identical, at least about 93% identical, at least about 94% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, or at least about 99% identical to any one of SEQ ID NOs: 2201-2280.
[0180] In some embodiments, the gNA variants, as described in more detail above, further comprise a spacer (or targeting sequence) region, comprising at least 14 to about 35 nucleotides, wherein the spacer is designed with a sequence complementary to the target DNA. In some embodiments, the gNA variants comprise a targeting sequence of at least 10 to 30 nucleotides that is complementary to the target DNA. In some embodiments, the targeting sequence has 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, or 35 nucleotides. In some embodiments, the gNA variants comprise a targeting sequence having 20 nucleotides. In some embodiments, the targeting sequence has 25 nucleotides. In some embodiments, the targeting sequence has 24 nucleotides. In some embodiments, the targeting sequence has 23 nucleotides. In some embodiments, the targeting sequence has 22 nucleotides. In some embodiments, the targeting sequence has 21 nucleotides. In some embodiments, the targeting sequence has 20 nucleotides. In some embodiments, the targeting sequence has 19 nucleotides. In some embodiments, the targeting sequence has 18 nucleotides. In some embodiments, the targeting sequence has 17 nucleotides. In some embodiments, the targeting sequence has 16 nucleotides. In some embodiments, the targeting sequence has 15 nucleotides. In some embodiments, the targeting sequence has 14 nucleotides.
[0181] In some embodiments, the scaffold of the gNA variant is a variant that includes one or more additional changes relative to the sequence of a reference gRNA that includes SEQ ID NO: 4 or SEQ ID NO: 5. In these embodiments where the reference gRNA scaffold is derived from SEQ ID NO: 4 or SEQ ID NO: 5, one or more improved or added features of the gNA variant are improved compared to the same feature of SEQ ID NO: 4 or SEQ ID NO: 5.
[0182] In some embodiments, the scaffold of the gNA variant is part of an RNP having a reference CasX protein comprising SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 3. In other embodiments, the scaffold of the gNA variant is part of an RNP having a CasX variant protein comprising any one of the sequences in Tables 3, 8, 9, 10, and 12, or a sequence having at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto. In the foregoing embodiments, the gNA further comprises a spacer sequence.
[0183] h. Chemically modified gNA In some embodiments, the present disclosure provides chemically modified gNAs. In some embodiments, the present disclosure provides chemically modified gNAs that have guide NA functionality and reduced susceptibility to nuclease cleavage. A gNA containing any nucleotide other than the four standard ribonucleotides A, C, G, and U, or a deoxynucleotide, is a chemically modified gNA. In some cases, the chemically modified gNA contains any backbone or internucleotide linkage other than a natural phosphodiester internucleotide linkage. In certain embodiments, the retained functionality includes the ability of the modified gNA to bind to a CasX of any of the embodiments described herein. In certain embodiments, the retained functionality includes the ability of the modified gNA to bind to a target nucleic acid sequence. In certain embodiments, the retained functionality includes the ability to target a CasX protein or a pre-complexed RNP to bind to a target nucleic acid sequence. In certain embodiments, the retained functionality includes the ability to nick a target polynucleotide by the CasX-gNA. In certain embodiments, the retained functionality includes the ability to cleave a target nucleic acid sequence by the CasX-gNA. In certain embodiments, the functionality that is retained is any other known function of a gNA in a recombinant system with a CasX chimeric protein of an embodiment of the present disclosure.
[0184] In some embodiments, the present disclosure provides a nucleotide sugar modification comprising 2'-OC 1-4 Alkyl, e.g., 2'-O-methyl (2'-OMe), 2'-deoxy (2'-H), 2'-OC 1-3 Alkyl-OC 1-3Chemically modified gNAs are provided in which alkyl, e.g., 2'-methoxyethyl ("2'-MOE"), 2'-fluoro ("2'-F"), 2'-amino ("2'-NH"), 2'-arabinosyl ("2'-arabino") nucleotides, 2'-F-arabinosyl ("2'-F-arabino") nucleotides, 2'-locked nucleic acid ("LNA") nucleotides, 2'-unlocked nucleic acid ("ULNA") nucleotides, L-form sugars ("L-sugars"), and 4'-thioribosyl nucleotides are incorporated into the gNA. In other embodiments, the internucleotide linkage modification incorporated into the guide RNA is selected from the group consisting of phosphorothioate "P(S)" (P(S)), phosphonocarboxylate (P(CH)), phosphonothioate (P(CH)). n COOR), e.g., phosphonoacetate "PACE" (P(CHCOO - )), thiophosphonocarboxylate ((S)P(CH2) n COOR), e.g., thiophosphonoacetate "thioPACE" ((S)P(CH2) n COO - )), alkyl phosphonates (P(C 1-3 alkyl), such as methylphosphonate-P(CH), boranophosphonate (P(BH)), and phosphorodithioate (P(S)).
[0185] In certain embodiments, the present disclosure provides methods for preparing nucleic acid base ("base") modifications, including but not limited to, 2-thiouracil ("2-thioU"), 2-thiocytosine ("2-thioC"), 4-thiouracil ("4-thioU"), 6-thioguanine ("6-thioG"), 2-aminoadenine ("2-aminoA"), 2-aminopurine, pseudouracil, hypoxanthine, 7-deazaguanine, 7-deaza-8-azaguanine, 7-deazaadenine, 7-deaza-8-azaadenine, 5-methylcytosine ("5-methylC"), 5-methyluracil ("5-methylU"), 5-hydroxymethylcytosine, 5-hydroxymethyluracil, 5,6-dehydrouracil , 5-propynylcytosine, 5-propynyluracil, 5-ethynylcytosine, 5-ethynyluracil, 5-allyluracil ("5-allylU"), 5-allylcytosine ("5-allylC"), 5-aminoallyluracil ("5-aminoallylU"), 5-aminoallyl-cytosine ("5-aminoallylC"), abasic nucleotides, Z bases, P bases, unstructured nucleic acids ("UNA"), isoguanine ("isoG"), isocytosine ("isoC"), 5-methyl-2-pyrimidine, x(A,G,C,T), and y(A,G,C,T) are incorporated into the gNA.
[0186] In other embodiments, the present disclosure provides that one or more isotopic modifications are 15 N, 13 C. 14 C, deuterium, 3 H, 32 P, 125 I, 131 Chemically modified gNAs are provided, including nucleotides containing an I atom, or other atoms or elements used as tracers, introduced into the nucleotide sugar, nucleobase, phosphodiester bond, and / or nucleotide phosphate.
[0187] In some embodiments, the "terminal" modification incorporated into the gNA is selected from the group consisting of PEG (polyethylene glycol), hydrocarbon linkers (including heteroatom (O, S, N)-substituted hydrocarbon spacers, halo-substituted hydrocarbon spacers, keto-, carboxyl-, amido-, thionyl-, carbamoyl-, and thionocarbamayl-containing hydrocarbon spacers), spermine linkers, dyes (e.g., fluorescein, rhodamine, cyanine), including fluorescent dyes attached to linkers such as 6-fluorescein-hexyl, quenchers (e.g., dabcyl, BHQ), and other labels (e.g., biotin, digoxigenin, acridine, streptavidin, avidin, peptides, and / or proteins). In some embodiments, the "terminal" modification includes conjugation (or ligation) of the gNA to another molecule, including deoxynucleotide and / or ribonucleotide oligonucleotides, peptides, proteins, sugars, oligosaccharides, steroids, lipids, folate, vitamins, and / or other molecules. In certain embodiments, the present disclosure provides chemically modified gNAs in which the "terminal" modification (described above) is incorporated as a phosphodiester bond and is located internally in the gNA sequence via a linker, such as, for example, a 2-(4-butylamidofluorescein)propane-1,3-diol bis(phosphodiester) linker, which can be incorporated anywhere between two nucleotides in the gNA.
[0188] In some embodiments, the present disclosure provides a method for the preparation of nucleotides containing amines, thiols (or sulfhydryls), hydroxyls, carboxyls, carbonyls, thionyls, thiocarbonyls, carbamoyls, thiocarbamoyls, phosphoryls, alkenes, alkynes, halogens, or fluorescent dyes, non-fluorescent labels, tags (e.g., 14 For C, for example, biotin, avidin, streptavidin, or 15 N, 13 C, deuterium, 3 H, 32 P, 125Chemically modified gNAs are provided having terminal modifications including terminal functional groups, such as functional terminal linkers, that can be subsequently conjugated to a desired moiety selected from the group consisting of: a moiety containing an isotopic label such as I, an oligonucleotide (including an aptamer, including deoxynucleotides and / or ribonucleotides), an amino acid, a peptide, a protein, a sugar, an oligosaccharide, a steroid, a lipid, a folate, and a vitamin. Conjugation can be accomplished using N-hydroxysuccinimide, isothiocyanate, DCC (or DCI), and / or other methods such as those described in "Bioconjugate Techniques" by Greg T. Hermanson, Publisher: Essevier Science, 3 rd ed. (2013), the contents of which are incorporated herein by reference in their entirety.
[0189] i. Complex formation with CasX protein In some embodiments, the gNA variant has an improved ability to form a complex with a CasX protein (such as a reference CasX or CasX variant protein) compared to a reference gRNA. In some embodiments, the gNA variant has an improved affinity for a CasX protein (such as a reference or variant protein) compared to a reference gRNA, thereby improving its ability to form a ribonucleoprotein (RNP) complex with the CasX protein, as described in the Examples. Improving ribonucleoprotein complex formation can, in some embodiments, improve the efficiency with which functional RNPs are assembled. In some embodiments, more than 90%, more than 93%, more than 95%, more than 96%, more than 97%, more than 98%, or more than 99% of RNPs comprising the gNA variant and spacer are competent for gene editing of the target nucleic acid.
[0190] An exemplary nucleotide change that can improve the ability of a gNA variant to form a complex with a CasX protein may, in some embodiments, include replacing the scaffold stem with a thermostable stem-loop. Without wishing to be bound by any theory, replacing the scaffold stem with a thermostable stem-loop can increase the overall binding stability of the gNA variant with the CasX protein. Alternatively, or in addition, removing a large portion of the stem-loop can alter the folding kinetics of the gNA variant, making it easier and faster to generate a functional, folded gNA and structurally assemble it, for example, by reducing the extent to which the gNA variant can "tangle" itself. In some embodiments, the selection of the scaffold stem-loop sequence can vary depending on the different spacers utilized in the gNA. In some embodiments, the scaffold sequence can be tailored to the spacer, and therefore the target sequence. Biochemical assays can be used to assess the binding affinity of the CasX protein to the gNA variant to form RNPs, including the assays in the examples. For example, one skilled in the art can measure the change in the amount of fluorescently labeled gNA bound to the immobilized CasX protein in response to increasing concentrations of additional unlabeled "cold competitor" gNA. Alternatively, or in addition, one can monitor or confirm how the fluorescent signal changes when different amounts of fluorescently labeled gNA are flowed over the immobilized CasX protein. Alternatively, the ability to form RNPs can be assessed using an in vitro cleavage assay against a defined target nucleic acid sequence.
[0191] j.gNA stability In some embodiments, gNA variants have improved stability compared to reference gRNA.Increased stability and efficient folding can in some embodiments increase the extent to which gNA variants persist in target cells, thereby increasing the chance of forming functional RNPs that can perform CasX functions such as gene editing.Increased stability of gNA variants can also in some embodiments allow for the same result of less gNA being delivered to cells, which can then reduce the chance of off-target effects during gene editing.
[0192] In other embodiments, the present disclosure provides gNAs in which the scaffold stem loop and / or extended stem loop are replaced with a hairpin loop or a thermostable RNA stem loop, resulting in gNAs with increased stability and capable of interacting with specific cellular proteins or RNAs, depending on the loop selection. In some embodiments, the replaced RNA loop is selected from MS2, Qβ, U1 hairpin II, Uvsx, PP7, phage replication loop, kissing loop_a, kissing loop_b1, kissing loop_b2, G-quadruplex M3q, G-quadruplex telomere basket, sarcin-ricin loop, and pseudoknot. The sequences of gNA variants containing such components are listed in Table 2.
[0193] The stability of guide NAs can be assessed in a variety of ways, including, for example, by assembling the guide in vitro, incubating it in a solution that mimics the intracellular environment for various periods of time, and then measuring functional activity using the in vitro cleavage assay described herein. Alternatively, or in addition, gNAs can be harvested from cells at various times after initial transfection / transduction of the gNA to determine the duration of persistence of the gNA variant relative to a reference gRNA.
[0194] k. Solubility In some embodiments, the gNA variant has improved solubility compared to a reference gRNA. In some embodiments, the gNA variant has improved solubility of the CasX protein:gNA RNP compared to a reference gRNA. In some embodiments, the solubility of the CasX protein:gNA RNP is improved by adding a ribozyme sequence to the 5' or 3' end of the gNA variant, for example, the 5' or 3' end of the reference sgRNA. Some ribozymes, such as the M1 ribozyme, can increase protein solubility through RNA-mediated protein folding.
[0195] The increased solubility of CasX RNPs containing the gNA variants described herein can be assessed by various means known to those of skill in the art, such as by taking densitometric readings on gels of the soluble fraction of lysed E. coli in which CasX and the gNA variant are expressed.
[0196] l. Resistance to nuclease activity In some embodiments, gNA variants have improved resistance to nuclease activity compared with reference gRNA.Without wishing to be bound by any theory, increasing resistance to nucleases, such as the nucleases found in cells, can, for example, increase the persistence of variant gNA in intracellular environment, thereby improving gene editing.
[0197] Many nucleases are processive and degrade RNA in a 3' to 5' manner. Thus, in some embodiments, adding a nuclease-resistant secondary structure to one or both ends of a gNA or changing the nucleotide structure of a sgNA can generate a gNA variant with increased resistance to nuclease activity. Resistance to nuclease activity can be assessed by various methods known to those skilled in the art. For example, an in vitro method for measuring resistance to nuclease activity can include, for example, contacting a reference gNA and a variant with one or more exemplary RNA nucleases and measuring degradation. Alternatively, or in addition, the persistence of a gNA variant in a cellular environment using the methods described herein can indicate the degree to which the gNA variant is nuclease-resistant.
[0198] m. binding affinity to target DNA In some embodiments, the gNA variant has improved affinity for target DNA compared to the reference gRNA. In certain embodiments, the ribonucleoprotein complex containing the gNA variant has improved affinity for target DNA compared to the affinity of the RNP containing the reference gRNA. In some embodiments, the improved affinity of the RNP for target DNA includes improved affinity for the target sequence, improved affinity for the PAM sequence, improved ability of the RNP to search for DNA for the target sequence, or any combination thereof. In some embodiments, the improved affinity for target DNA is the result of increased overall DNA binding affinity.
[0199] Without wishing to be bound by theory, nucleotide changes in the gNA variant that affect the function of the OBD of the CasX protein may increase the affinity of the CasX variant protein for binding to a protospacer adjacent motif (PAM) and its ability to bind to or utilize an increased spectrum of PAM sequences other than the standard TTC PAM recognized by the reference CasX protein of SEQ ID NO:2, including PAM sequences selected from the group consisting of TTC, ATC, GTC, and CTC, thereby increasing the affinity and diversity of the CasX variant protein for target DNA sequences and thereby increasing the number of target nucleic acid sequences it can edit and / or bind to compared to the reference CasX. As described in more detail below, increasing the number of target nucleic acid sequences that can be edited compared to the reference CasX refers to both the PAM and protospacer sequences and their orientation relative to the orientation of the non-target strand. This does not imply that the PAM sequence of the non-target strand, rather than the target strand, determines cleavage or is mechanistically involved in target recognition. For example, when referring to a TTC PAM, it may actually be the complementary GAA sequence required for target cleavage, or it may be some combination of nucleotides from both strands. In the CasX proteins disclosed herein, the PAM is located 5' of the protospacer, with at least a single nucleotide separating the PAM from the first nucleotide of the protospacer. Alternatively, or in addition, changes in the gNA that affect the function of the Helical I and / or Helical II domains that increase the affinity of the CasX variant protein for the target DNA strand can increase the affinity of the CasX RNP containing the variant gNA for the target DNA.
[0200] Addition or change of n.gNA function In some embodiments, gNA variants can comprise larger structural changes that change the topology of gNA variants relative to reference gRNA, thereby enabling different gNA functionality.For example, in some embodiments, gNA variants replace the endogenous stem-loop of the reference gRNA scaffold with a previously identified stable RNA structure, or a stem-loop that interacts with a protein or RNA binding partner to recruit additional moieties to CasX, or that has a binding partner to this RNA structure, recruiting CasX to a specific location, such as the inside of a virus capsid.In other situations, RNAs can be recruited to each other, such as kissing loops, thereby allowing two CasX proteins to co-localize for more efficient gene editing at target DNA sequences. Such RNA structures can include MS2, Qβ, U1 hairpin II, Uvsx, PP7, phage replication loop, kissing loop_a, kissing loop_b1, kissing loop_b2, G-quadriplex M3q, G-quadriplex telomeric basket, sarcin-ricin loop, or pseudoknot.
[0201] In some embodiments, the gNA variant comprises a terminal fusion partner. The term gNA variant includes variants containing exogenous sequences, such as terminal fusions or internal insertions. Exemplary terminal fusions may include fusion of a gRNA to a self-cleaving ribozyme or a protein-binding motif. As used herein, "ribozyme" refers to an RNA or a segment thereof that has one or more catalytic activities similar to those of a protein enzyme. Exemplary ribozyme catalytic activities may include, for example, RNA cleavage and / or ligation, DNA cleavage and / or ligation, or peptide bond formation. In some embodiments, such fusions can improve scaffold folding or recruit DNA repair mechanisms. For example, in some embodiments, gRNAs can be fused to hepatitis delta virus (HDV) antigenome ribozyme, HDV genomic ribozyme, hatchet ribozyme (from metagenomic data), env25 pistol ribozyme (represented by Aliistipes putredinis), HH15 minimal hammerhead ribozyme, tobacco ringspot virus (TRSV) ribozyme, WT virus hammerhead ribozyme (and rational variants), or twisted sister 1 or RBMX recruiting motifs. Hammerhead ribozymes are RNA motifs that catalyze reversible cleavage and ligation reactions at specific sites within RNA molecules. Hammerhead ribozymes include type I, type II, and type III hammerhead ribozymes. HDV, pistol, and hatchet ribozymes have self-cleaving activity. gNA variants containing one or more ribozymes may enable expanded gNA functionality compared to the gRNA reference. For example, in some embodiments, a gNA containing a self-cleaving ribozyme can be transcribed and processed into a mature gNA as part of a polycistronic transcription product. Such fusions can occur at either the 5' or 3' end of the gNA. In some embodiments, a gNA variant contains fusions at both the 5' and 3' ends, with each fusion being independent as described herein. In some embodiments, a gNA variant contains a phage replication loop or tetraloop.In some embodiments, the gNA comprises a hairpin loop capable of binding to a protein. For example, in some embodiments, the hairpin loop is an MS2, Qβ, U1 hairpin II, Uvsx, or PP7 hairpin loop.
[0202] In some embodiments, the gNA variant comprises one or more RNA aptamers. As used herein, "RNA aptamer" refers to an RNA molecule that binds to a target with high affinity and high specificity.
[0203] In some embodiments, the gNA variant comprises one or more riboswitches. As used herein, "riboswitch" refers to an RNA molecule that changes state upon binding to a small molecule.
[0204] In some embodiments, the gNA variant further comprises one or more protein-binding motifs. Adding protein-binding motifs to a reference gRNA or gNA variant of the present disclosure may, in some embodiments, allow the CasX RNP to associate with additional proteins, which can, for example, add the function of those proteins to the CasX RNP.
[0205] IV. CasX Proteins for Modifying Target Nucleic Acids As used herein, the term "CasX protein" refers to a family of proteins and encompasses all naturally occurring CasX proteins, proteins that share at least 50% identity with a naturally occurring CasX protein, and CasX variants that have one or more improved characteristics compared to a naturally occurring reference CasX protein. Exemplary improved features of embodiments of CasX variants include, but are not limited to, improved folding of the variant, improved binding affinity for gNA, improved binding affinity for target nucleic acids, improved ability to utilize a broader spectrum of PAM sequences in editing and / or binding to target DNA, improved unwinding of target DNA, increased editing activity, improved editing efficiency, improved editing specificity, an increased proportion of eukaryotic genomes that can be effectively edited, increased nuclease activity, increased target strand loading for double-stranded cleavage, reduced target strand loading for single-stranded nicking, reduced off-target cleavage, improved binding of non-target strands of DNA, improved protein stability, improved protein:gNA (RNP) complex stability, improved protein solubility, improved protein:gNA (RNP) complex solubility, improved protein yield, improved protein expression, and improved fusion characteristics, as described in more detail below. In the foregoing embodiments, one or more of the improved characteristics of the CasX variant is improved by at least about 1.1 to about 100,000-fold when assayed in an equivalent manner compared to a reference CasX protein of SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 3. In other embodiments, the improvement is at least about 1.1-fold, at least about 2-fold, at least about 5-fold, at least about 10-fold, at least about 50-fold, at least about 100-fold, at least about 500-fold, at least about 1000-fold, at least about 5000-fold, at least about 10,000-fold, or at least about 100,000-fold when assayed in an equivalent manner compared to a reference CasX protein of SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 3.
[0206] The term CasX variant includes variants that are fusion proteins, i.e., CasX is "fused" to a heterologous sequence. This includes CasX variants that contain a CasX variant sequence and an N-terminal, C-terminal, or internal fusion of CasX to a heterologous protein or domain thereof.
[0207] The CasX proteins of the present disclosure comprise at least one of the following domains: a non-target strand binding (SB) domain, a target strand loading (TSL) domain, a helical I domain, a helical II domain, an oligonucleotide binding domain (OBD), and a RuvC DNA cleavage domain (the latter of which may be modified or deleted in catalytically dead CasX variants), as described in more detail below. Furthermore, the CasX variant proteins of the present disclosure have an enhanced ability to effectively edit and / or bind target DNA utilizing a PAM sequence selected from TTC, ATC, GTC, or CTC, compared to a wild-type reference CasX protein. In the foregoing, the PAM sequence is located at least one nucleotide 5' to the non-target strand of a protospacer having identity to the targeting sequence of a gNA in an assay system, compared to the editing efficiency and / or binding of an RNP containing the reference CasX protein in an equivalent assay system.
[0208] In some cases, the CasX protein is a naturally occurring protein (e.g., naturally occurring in and isolated from prokaryotic cells). In other embodiments, the CasX protein is not a naturally occurring protein (e.g., the CasX protein is a CasX variant protein, a chimeric protein, etc.). Naturally occurring CasX proteins (referred to herein as "reference CasX proteins") function as endonucleases that catalyze double-strand cleavage at specific sequences in targeted double-stranded DNA (dsDNA). Sequence specificity is provided by the targeting sequence of the associated gNA that forms the complex, which hybridizes to the target sequence within the target nucleic acid.
[0209] In some embodiments, the CasX protein is capable of binding to and / or modifying (e.g., cleaving, nicking, methylating, demethylating, etc.) target nucleic acids and / or polypeptides associated with target nucleic acids (e.g., methylating or acetylating histone tails). In some embodiments, the CasX protein is catalytically dead (dCasX) but retains the ability to bind to target nucleic acids. Exemplary catalytically dead CasX proteins comprise one or more mutations in the active site of the RuvC domain of the CasX protein. In some embodiments, the catalytically dead CasX protein comprises substitutions at residues 672, 769, and / or 935 of SEQ ID NO: 1. In one embodiment, the catalytically dead CasX protein comprises substitutions at D672A, E769A, and / or D935A in a reference CasX protein of SEQ ID NO: 1. In other embodiments, the catalytically dead CasX protein comprises substitutions at amino acids 659, 756, and / or 922 in a reference CasX protein of SEQ ID NO: 2. In some embodiments, the catalytically inactive CasX protein comprises the D659A, E756A, and / or D922A substitutions in the reference CasX protein of SEQ ID NO: 2. In further embodiments, the catalytically inactive CasX protein comprises a deletion of all or part of the RuvC domain of the CasX protein. It will be understood that the same aforementioned substitutions may also be introduced into the CasX variants of the present disclosure, resulting in a dCasX variant. In one embodiment, all or part of the RuvC domain is deleted from the CasX variant, resulting in a dCasX variant. In some embodiments, catalytically inactive dCasX variant proteins can be used for base editing or epigenetic modification. The increased affinity for DNA allows catalytically inactive dCasX variant proteins to find their target nucleic acids faster, remain bound to target nucleic acids for longer periods, bind to target nucleic acids in a more stable manner, or a combination thereof, compared to catalytically active CasX, thereby improving the function of the catalytically inactive CasX variant protein.
[0210] a. non-target strand binding domain The reference CasX protein of the present disclosure contains a non-target strand-binding domain (NTSBD). The NTSBD is a domain not previously found in any Cas protein; for example, this domain is not present in Cas proteins such as Cas9, Cas12a / Cpf1, Cas13, Cas14, CASCADE, CSM, or CSY. Without being bound by theory or mechanism, the NTSBD in CasX may enable binding to the non-target DNA strand and aid in unwinding of both the non-target and target strands. The NTSBD is presumed to be involved in unwinding or capturing the non-target DNA strand in its unwound state. The NTSBD directly contacts the non-target strand in previously derived cryoEM model structures and may contain a non-canonical zinc finger domain. The NTSBD may also play a role in stabilizing DNA during unwinding, guide RNA invasion, and R-loop formation. In some embodiments, an exemplary NTSBD comprises amino acids 101-191 of SEQ ID NO:1 or amino acids 103-192 of SEQ ID NO:2. In some embodiments, the NTSBD of the reference CasX protein comprises a four-stranded beta sheet.
[0211] b. Target strand loading domain The reference CasX proteins of this disclosure contain a target strand loading (TSL) domain. The TSL domain is a domain not found in certain Cas proteins, such as Cas9, CASCADE, CSM, or CSY. Without wishing to be bound by theory or mechanism, it is believed that the TSL domain is involved in assisting the CasX protein in loading the target DNA strand into the RuvC active site. In some embodiments, the TSL acts to position or capture the target strand in a folded state that positions the cleavable phosphate of the target strand DNA backbone into the RuvC active site. The TSL comprises a cys4 (CXXC (SEQ ID NO: 246, CXXC (SEQ ID NO: 246) zinc finger / ribbon domain separated by the majority of the TSL. In some embodiments, an exemplary TSL comprises amino acids 825-934 of SEQ ID NO: 1 or amino acids 813-921 of SEQ ID NO: 2.
[0212] C helical I domain The reference CasX protein of the present disclosure contains a helical I domain. Certain Cas proteins other than CasX have domains that may be similarly named. However, in some embodiments, the helical I domain of a CasX protein contains one or more unique structural features or a unique sequence, or a combination thereof, compared to non-CasX proteins. For example, in some embodiments, the helical I domain of a CasX protein contains one or more unique secondary structures compared to domains of other Cas proteins that may have similar names. For example, in some embodiments, the helical I domain of a CasX protein contains one or more alpha helices of a structure and sequence that are unique in terms of their arrangement, number, and length compared to other CRISPR proteins. In certain embodiments, the helical I domain is involved in interactions with the spacer of the bound DNA and guide RNA. Without wishing to be bound by theory, it is believed that in some cases the helical I domain may contribute to the binding of a protospacer adjacent motif (PAM). In some embodiments, an exemplary helical I domain comprises amino acids 57-100 and 192-332 of SEQ ID NO: 1, or amino acids 59-102 and 193-333 of SEQ ID NO: 2. In some embodiments, the helical I domain of a reference CasX protein comprises one or more alpha helices.
[0213] d. Helical II domain The reference CasX protein of the present disclosure contains a helical II domain. Certain Cas proteins other than CasX have domains that may be similarly named. However, in some embodiments, the helical II domain of a CasX protein contains one or more unique structural features, or a unique sequence, or a combination thereof, compared to domains of other Cas proteins that may have similar names. For example, in some embodiments, the helical II domain contains one or more unique structural alpha-helical bundles that align along the target DNA:guide RNA channel. In some embodiments, in CasXs containing a helical II domain, the target strand and guide RNA interact with the helical II (and in some embodiments, the helical I domain) to allow the RuvC domain to access the target DNA. The helical II domain is involved in binding to the guide RNA scaffold stem-loop and the bound DNA. In some embodiments, an exemplary helical II domain contains amino acids 333-509 of SEQ ID NO:1 or amino acids 334-501 of SEQ ID NO:2.
[0214] e. oligonucleotide binding domain The reference CasX protein of the present disclosure contains an oligonucleotide binding domain (OBD). Certain Cas proteins other than CasX have domains that may be similarly named. However, in some embodiments, the OBD contains one or more unique functional features or a sequence unique to the CasX protein, or a combination thereof. For example, in some embodiments, the bridge helix (BH), helical I domain, helical II domain, and oligonucleotide binding domain (OBD) together participate in binding of the CasX protein to the guide RNA. Thus, for example, in some embodiments, the OBD is unique to the CasX protein in that it functionally interacts with the helical I domain, the helical II domain, or both, each of which may be unique to the CasX protein as described herein. Specifically, in CasX, the OBD primarily binds to the RNA triplex of the guide RNA scaffold. The OBD may also participate in binding to the protospacer adjacent motif (PAM). An exemplary OBD domain includes amino acids 1-56 and 510-660 of SEQ ID NO:1, or amino acids 1-58 and 502-647 of SEQ ID NO:2.
[0215] f. RuvC DNA cleavage domain The reference CasX protein of this disclosure contains a RuvC domain, which includes two partial RuvC domains (RuvC-I and RuvC-II). The RuvC domain is the ancestral domain of all types of CRISPR proteins. The RuvC domain is derived from TNPB (transposase B)-like transposases. Like other RuvC domains, the CasX RuvC domain possesses a DED catalytic triad involved in magnesium (Mg) ion coordination and DNA cleavage. In some embodiments, RuvC possesses a DED motif active site involved in cleaving both strands of DNA (one at a time, possibly the non-target strand at 11-14 nucleotides (nt) within the first targeted sequence, and then the next target strand at 2-4 nucleotides after the target sequence). In particular, the RuvC domain in CasX is unique in that it is also involved in binding the guide RNA scaffold stem-loop, which is critical for CasX function. An exemplary RuvC domain includes amino acids 661-824 and 935-986 of SEQ ID NO:1, or amino acids 648-812 and 922-978 of SEQ ID NO:2.
[0216] g. Reference CasX protein The present disclosure provides a reference CasX protein. In some embodiments, the reference CasX protein is a naturally occurring protein. For example, the reference CasX protein can be isolated from a naturally occurring prokaryote, such as a Deltaproteobacteria, Planctomycetes, or Candidatus Sungbacteria species. The reference CasX protein (sometimes referred to herein as a reference CasX polypeptide) is a type II CRISPR / Cas endonuclease belonging to the CasX (sometimes referred to as Cas12e) family of proteins that can interact with a guide NA to form a ribonucleoprotein (RNP). In some embodiments, an RNP complex containing the reference CasX protein can be targeted to a specific site within a target nucleic acid through base pairing between the targeting sequence (or spacer) of the gNA and a target sequence within the target nucleic acid. In some embodiments, the RNP containing the reference CasX protein can cleave the target DNA. In some embodiments, the RNP containing the reference CasX protein can nick the target DNA. In some embodiments, the RNP containing the reference CasX protein is capable of editing target DNA, e.g., in these embodiments, the reference CasX protein is capable of cleaving or nicking DNA, followed by non-homologous end joining (NHEJ), homologous recombination modification (HDR), homology-independent targeted integration (HITI), microhomology-mediated end joining (MMEJ), single-strand annealing (SSA), or base excision repair (BER). In some embodiments, the RNP containing the CasX protein is a catalytically dead (catalytically inactive or substantially lacking cleavage activity) CasX protein (dCasX), but retains the ability to bind target DNA, as described in more detail above.
[0217] In some embodiments, the reference CasX protein is isolated from or derived from Deltaproteobacteria. In some embodiments, the CasX protein comprises a sequence at least 50% identical, at least 60% identical, at least 65% identical, at least 70% identical, at least 75% identical, at least 80% identical, at least 81% identical, at least 82% identical, at least 83% identical, at least 84% identical, at least 85% identical, at least 86% identical, at least 86% identical, at least 87% identical, at least 88% identical, at least 89% identical, at least 89% identical, at least 90% identical, at least 91% identical, at least 92% identical, at least 93% identical, at least 94% identical, at least 95% identical, at least 96% identical, at least 97% identical, at least 98% identical, at least 99% identical, at least 99.5% identical, or 100% identical to the following sequence: 1 MEKRINKIRK KLSADNATKP VSRSGPMKTL LVRVMTDDLK KRLEKRRKKP EVMPQVISNN 61 AANNLRMLLD DYTKMKEAIL QVYWQEFKDD HVGLMCKFAQ PASKKIDQNK LKPEMDEKGN 121 LTTAGFACSQ CGQPLFVYKL EQVSEKGKAY TNYFGRCNVA EHEKLILLAQ LKPEKDSDEA 181 VTYSLGKFGQ RALDFYSIHV TKESTHPVKP LAQIAGNRYA SGPVGKALSD ACMGTIASFL 241 SKYQDIIIEH QKVVKGNQKR LESLRELAGK ENLEYPSVTL PPQPHTKEGV DAYNEVIARV 301 RMWVNLNLWQ KLKLSRDDAK PLLRLKGFPS FPVVERRENE VDWWNTINEV KKLIDAKRDM 361 GRVFWSGVTA EKRNTILEGY NYLPNENDHK KREGSLENPK KPAKRQFGDL LLYLEKKYAG 421 DWGKVFDEAW ERIDKKIAGL TSHIEREEAR NAEDAQSKAV LTDWLRAKAS FVLERLKEMD 481 EKEFYACEIQ LQKWYGDLRG NPFAVEAENR VVDISGFSIG SDGHSIQYRN LLAWKYLENG 541 KREFYLLMNY GKKGRIRFTD GTDIKKSGKW QGLLYGGGKA KVIDLTFDPD DEQLIILPLA 601 FGTRQGREFI WNDLLSLETG LIKLANGRVI EKTIYNKKIG RDEPALFVAL TFERREVVDP 661 SNIKPVNLIG VDRGENIPAV IALTDPEGCP LPEFKDSSGG PTDILRIGEG YKEKQRAIQA 721 AKEVEQRRAG GYSRKFASKS RNLADDMVRN SARDLFYHAV THDAVLVFEN LSRGFGRQGK 781 RTFMTERQYT KMEDWLTAKL AYEGLTSKTY LSKTLAQYTS KTCSNCGFTI TTADYDGMLV 841 RLKKTSDGWA TTLNNKELKA EGQITYYNRY KRQTVEKELS AELDRLSEES GNNDISKWTK 901 GRRDEALFLL KKRFSHRPVQ EQFVCLDCGH EVHADEQAAL NIARSWLFLN SNSTEFKSYK 961 SGKQPFVGAW QAFYKRRLKE VWKPNA (Sequence number 1).
[0218] In some embodiments, the reference CasX protein is isolated from or derived from a Planctomycetes. In some embodiments, the CasX protein comprises a sequence at least 50% identical, at least 60% identical, at least 65% identical, at least 70% identical, at least 75% identical, at least 80% identical, at least 81% identical, at least 82% identical, at least 83% identical, at least 84% identical, at least 85% identical, at least 86% identical, at least 86% identical, at least 87% identical, at least 88% identical, at least 89% identical, at least 89% identical, at least 90% identical, at least 91% identical, at least 92% identical, at least 93% identical, at least 94% identical, at least 95% identical, at least 96% identical, at least 97% identical, at least 98% identical, at least 99% identical, at least 99.5% identical, or 100% identical to the following sequence: 1 MQEIKRINKI RRRLVKDSNT KKAGKTGPMK TLLVRVMTPD LRERLENLRK KPENIPQPIS 61 NTSRANLNKL LTDYTEMKKA ILHVYWEEFQ KDPVGLMSRV AQPAPKNIDQ RKLIPVKDGN 121 ERLTSSGFAC SQCCQPLYVY KLEQVNDKGK PHTNYFGRCN VSEHERLILL SPHKPEANDE 181 LVTYSLGKFG QRALDFYSIH VTRESNHPVK PLEQIGGNSC ASGPVGKALS DACMGAVASF 241 LTKYQDIILE HQKVIKKNEK RLANLKDIAS ANGLAFPKIT LPPQPHTKEG IEAYNNVVAQ 301 IVIWVNLNLW QKLKIGRDEA KPLQRLKGFP SFPLVERQAN EVDWWDMVCN VKKLINEKKE 361 DGKVFWQNLA GYKRQEALLP YLSSEEDRKK GKKFARYQFG DLLLHLEKKH GEDWGKVYDE 421 AWERIDKKVE GLSKHIKLEE ERRSEDAQSK AALTDWLRAK ASFVIEGLKE ADKDEFCRCE 481 LKLQKWYGDL RGKPFAIEAE NSILDISGFS KQYNCAFIWQ KDGVKKLNLY LIINYFKGGK 541 LRFKKIKPEA FEANRFYTVI NKKSGEIVPM EVNFNFDDPN LIILPLAFGK RQGREFIWND 601 LLSLETGSLK LANGRVIEKT LYNRRTRQDE PALFVALTFE RREVLDSSNI KPMNLIGIDR 661 GENIPAVIAL TDPEGCPLSR FKDSLGNPTH ILRIGESYKE KQRTIQAAKE VEQRRAGGYS 721 RKYASKAKNL ADDMVRNTAR DLLYYAVTQD AMLIFENLSR GFGRQGKRTF MAERQYTRME 781 DWLTAKLAYE GLPSKTYLSK TLAQYTSKTC SNCGFTITSA DYDRVLEKLK KTATGWMTTI 841 NGKELKVEGQ ITYYNRYKRQ NVVKDLSVEL DRLSEESVNN DISSWTKGRS GEALSLLKKR 901 FSHRPVQEKF VCLNCGFETH ADEQAALNIA RSWLFLRSQE YKKYQTNKTT GNTDKRAFVE 961 TWQSFYRKKL KEVWKPAV (SEQ ID NO: 2).
[0219] In some embodiments, the CasX protein comprises the sequence of SEQ ID NO:2, or at least 60% similar thereto. In some embodiments, the CasX protein comprises the sequence of SEQ ID NO:2, or at least 80% similar thereto. In some embodiments, the CasX protein comprises the sequence of SEQ ID NO:2, or at least 90% similar thereto. In some embodiments, the CasX protein comprises the sequence of SEQ ID NO:2, or at least 95% similar thereto. In some embodiments, the CasX protein consists of the sequence of SEQ ID NO:2. In some embodiments, the CasX protein comprises or consists of a sequence having at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 20, at least 30, at least 40, or at least 50 mutations relative to the sequence of SEQ ID NO:2. These mutations can be insertions, deletions, amino acid substitutions, or any combination thereof.
[0220] In some embodiments, the reference CasX protein is isolated from or derived from Candidatus Sungbacteria. In some embodiments, the CasX protein comprises a sequence at least 50% identical, at least 60% identical, at least 65% identical, at least 70% identical, at least 75% identical, at least 80% identical, at least 81% identical, at least 82% identical, at least 83% identical, at least 84% identical, at least 85% identical, at least 86% identical, at least 86% identical, at least 87% identical, at least 88% identical, at least 89% identical, at least 89% identical, at least 90% identical, at least 91% identical, at least 92% identical, at least 93% identical, at least 94% identical, at least 95% identical, at least 96% identical, at least 97% identical, at least 98% identical, at least 99% identical, at least 99.5% identical, or 100% identical to the following sequence: 1 MDNANKPSTK SLVNTTRISD HFGVTPGQVT RVFSFGIIPT KRQYAIIERW FAAVEAARER 61 LYGMLYAHFQ ENPPAYLKEK FSYETFKGR PVNLGLRDID PTIMTSAVFT ALRHKAEGAM 121 AAFHTNHRRL FEEARKKMRE YAECLKANEA LLRGAADIDW DKIVNALRTR LNTCLAPEYD 181 AVIADFGALC AFRALIAETN ALKGAYNHAL NQMLPALVKV DEPEEAEESP RLRFFNGRIN 241 DLPKFPVAER ETPPDTETII RQLEDMARVI PDTAEILGYI HRIRHKAARR KPGSAVPLPQ 301 RVALYCAIRM ERNPEEDPST VAGHFLGEID RVCEKRRQGL VRTPFDSQIR ARYMDIISFR 361 ATLAHPDRWT EIQFLRSNAA SRRVRAETIS APFEGFSWTS NRTNPAPQYG MALAKDANAP 421 ADAPELCICL SPSSAAFSVR EKGGDLIYMR PTGGRRGKDN PGKEITWVPG SFDEYPASGV 481 ALKLRLYFGR SQARRMLTNK TWGLLSDNPR VFAANAELVG KKRNPQDRWK LFFHMVISGP 541 PPVEYLDFSS DVRSRARTVI GINRGEVNPL AYAVVSVEDG QVLEEGLLGK KEYIDQLIET 601 RRRISEYQSR EQTPPRDLRQ RVRHLQDTVL GSARAKIHSL IAFWKGILAI ERLDDQFHGR 661 EQKIIPKKTY LANKTGFMNA LSFSGAVRVD KKGNPWGGMI EIYPGGISRT CTQCGTVWLA 721 RRPKNPGHRD AMVVIPDIVD DAAATGFDNV DCDAGTVDYG ELFTLSREWV RLTPRYSRVM 781 RGTLGDLERA IRQGDDRKSR QMLELALEPQ PQWGQFFCHR CGFNGQSDVL AATNLARRAI 841 SLIRRLPDTD TPPTP (SEQ ID NO: 3).
[0221] In some embodiments, the CasX protein comprises the sequence of SEQ ID NO: 3, or at least 60% similar thereto. In some embodiments, the CasX protein comprises the sequence of SEQ ID NO: 3, or at least 80% similar thereto. In some embodiments, the CasX protein comprises the sequence of SEQ ID NO: 3, or at least 90% similar thereto. In some embodiments, the CasX protein comprises the sequence of SEQ ID NO: 3, or at least 95% similar thereto. In some embodiments, the CasX protein consists of the sequence of SEQ ID NO: 3. In some embodiments, the CasX protein comprises or consists of a sequence having at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 20, at least 30, at least 40, or at least 50 mutations relative to the sequence of SEQ ID NO: 3. These mutations may be insertions, deletions, amino acid substitutions, or any combination thereof.
[0222] h.CasX variant protein The present disclosure provides variants of a reference CasX protein (interchangeably referred to herein as "CasX variants" or "CasX variant proteins"), where the CasX variants comprise at least one modification in at least one domain compared to the reference CasX protein, including, but not limited to, the sequences of SEQ ID NOs: 1-3. In some embodiments, the CasX variants exhibit at least one improved characteristic compared to the reference CasX protein. All variants that improve one or more functions or characteristics of the CasX variant protein compared to the reference CasX protein described herein are contemplated within the scope of the present disclosure. In some embodiments, the modification is a mutation in one or more amino acids of the reference CasX. In other embodiments, the modification is a substitution of one or more domains of the reference CasX with one or more domains from a different CasX. In some embodiments, the insertion comprises the insertion of part or all of a domain from a different CasX protein. Mutations can occur in any one or more domains of the reference CasX protein, including, for example, partial or complete deletion of one or more domains, or substitution, deletion, or insertion of one or more amino acids in any domain of the reference CasX protein. CasX protein domains include the non-target strand binding (NTSB) domain, the target strand loading (TSL) domain, the helical I domain, the helical II domain, the oligonucleotide binding domain (OBD), and the RuvC DNA cleavage domain. Any change in the amino acid sequence of the reference CasX protein that results in improved characteristics of the CasX protein is considered a CasX variant protein of the present disclosure. For example, a CasX variant can include one or more amino acid substitutions, insertions, deletions, or swapped domains, or any combination thereof, compared to the reference CasX protein sequence.
[0223] In some embodiments, the CasX variant protein comprises at least one modification in at least each of two domains of a reference CasX protein comprising a sequence of SEQ ID NO: 1-3. In some embodiments, the CasX variant protein comprises at least one modification in at least two domains, at least three domains, at least four domains, or at least five domains of the reference CasX protein. In some embodiments, the CasX variant protein comprises two or more modifications in at least one domain of the reference CasX protein. In some embodiments, the CasX variant protein comprises at least two modifications in at least one domain of the reference CasX protein, at least three modifications in at least one domain of the reference CasX protein, or at least four modifications in at least one domain of the reference CasX protein. In some embodiments, the CasX variant comprises two or more modifications compared to the reference CasX protein, each modification occurring in a domain independently selected from the group consisting of the NTSBD, TSLD, helical I domain, helical II domain, OBD, and RuvC DNA cleavage domain.
[0224] In some embodiments, at least one modification of the CasX variant protein comprises a deletion of at least a portion of one domain of the reference CasX protein, in some embodiments, the deletion is in the NTSBD, TSLD, helical I domain, helical II domain, OBD, or RuvC DNA cleavage domain.
[0225] Suitable mutagenesis methods for generating the CasX variant proteins of the present disclosure can include, for example, deep mutational evolution (DME), deep mutational scanning (DMS), error-prone PCR, cassette mutagenesis, random mutagenesis, staggered extension PCR, gene shuffling, or domain swapping. Exemplary methods for generating CasX variants with improved characteristics are provided in the Examples below. In some embodiments, CasX variants are designed, for example, by selecting one or more desired mutations in a reference CasX. In certain embodiments, the activity of the reference CasX protein is used as a benchmark against which the activities of one or more CasX variants are compared, thereby measuring the improvement in the function of the CasX variants. Exemplary improvements to CasX variants include, but are not limited to, improved folding of the variant, improved binding affinity for gNA, improved binding affinity for target DNA, improved ability to utilize a broader spectrum of PAM sequences in editing or binding to target DNA, improved unwinding of target DNA, increased editing activity, improved editing efficiency, improved editing specificity, increased nuclease activity, increased target strand loading for double-stranded cleavage, decreased target strand loading for single-stranded nicking, decreased off-target cleavage, improved binding of non-target strands of DNA, improved protein stability, improved CasX:gNA(RNP) complex stability, improved protein solubility, improved CasX:gNA(RNP) complex solubility, improved protein yield, improved protein expression, and improved fusion characteristics, as described in more detail below.
[0226] In some embodiments of the CasX variants described herein, at least one modification comprises (a) a substitution of 1 to 100 consecutive or non-consecutive amino acids in the CasX variant, (b) a deletion of 1 to 100 consecutive or non-consecutive amino acids in the CasX variant, (c) an insertion of 1 to 100 consecutive or non-consecutive amino acids in CasX, or (d) any combination of (a)-(c). In some embodiments, at least one modification comprises (a) a substitution of 5 to 10 consecutive or non-consecutive amino acids in the CasX variant, (b) a deletion of 1 to 5 consecutive or non-consecutive amino acids in the CasX variant, (c) an insertion of 1 to 5 consecutive or non-consecutive amino acids in CasX, or (d) any combination of (a)-(c).
[0227] In some embodiments, the CasX variant protein comprises or consists of a sequence having at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 20, at least 30, at least 40, or at least 50 mutations relative to the sequence of SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 3. These mutations can be insertions, deletions, amino acid substitutions, or any combination thereof.
[0228] In some embodiments, the CasX variant protein comprises at least one amino acid substitution in at least one domain of the reference CasX protein. In some embodiments, the CasX variant protein comprises at least about 1-4 amino acid substitutions, 1-10 amino acid substitutions, 1-20 amino acid substitutions, 1-30 amino acid substitutions, 1-40 amino acid substitutions, 1-50 amino acid substitutions, 1-60 amino acid substitutions, 1-70 amino acid substitutions, 1-80 amino acid substitutions, 1-90 amino acid substitutions, 1-100 amino acid substitutions, 2-10 amino acid substitutions, 2-20 amino acid substitutions, 2-30 amino acid substitutions, 3-10 amino acid substitutions, 3-20 amino acid substitutions, 3-30 amino acid substitutions, 4-10 amino acid substitutions, 4-20 amino acid substitutions, 3-300 amino acid substitutions, 5-10 amino acid substitutions, 5-20 amino acid substitutions, 5-30 amino acid substitutions, 10-50 amino acid substitutions, or 20-50 amino acid substitutions compared to a reference CasX protein. In some embodiments, the CasX variant protein contains at least about 100 amino acid substitutions compared to a reference CasX protein. In some embodiments, the CasX variant protein contains 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid substitutions compared to a reference CasX protein. In some embodiments, the CasX variant protein contains 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid substitutions in a single domain compared to a reference CasX protein. In some embodiments, the amino acid substitutions are conservative substitutions. In other embodiments, the substitutions are non-conservative, for example, a polar amino acid is substituted for a non-polar amino acid, or vice versa.
[0229] In some embodiments, the CasX variant protein has, compared to a reference CasX protein, 1 amino acid substitution, 2 to 3 consecutive amino acid substitutions, 2 to 4 consecutive amino acid substitutions, 2 to 5 consecutive amino acid substitutions, 2 to 6 consecutive amino acid substitutions, 2 to 7 consecutive amino acid substitutions, 2 to 8 consecutive amino acid substitutions, 2 to 9 consecutive amino acid substitutions, 2 to 10 consecutive amino acid substitutions, 2 to 20 consecutive amino acid substitutions, 2 to 30 consecutive amino acid substitutions, 2 to 40 consecutive amino acid substitutions, 2 to 50 consecutive amino acid substitutions, 2 to 60 consecutive amino acid substitutions, The CasX variant protein may comprise 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 consecutive amino acid substitutions. In some embodiments, the CasX variant protein comprises at least about 100 consecutive amino acid substitutions. As used herein, "consecutive amino acids" refers to amino acids that are adjacent in the primary sequence of a polypeptide.
[0230] In some embodiments, the CasX variant protein contains two or more substitutions compared to the reference CasX protein, and the two or more substitutions do not occur in consecutive amino acids of the reference CasX sequence. For example, a first substitution can be in a first domain of the reference CasX protein, and a second substitution can be in a second domain of the reference CasX protein. In some embodiments, the CasX variant protein contains 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 non-consecutive substitutions compared to the reference CasX protein. In some embodiments, the CasX variant protein contains at least 20 non-consecutive substitutions compared to the reference CasX protein. Each non-consecutive substitution can be any length of amino acids described herein, e.g., 1 to 4 amino acids, 1 to 10 amino acids, etc. In some embodiments, the two or more substitutions compared to the reference CasX protein are not the same length, e.g., the first substitution is one amino acid and the second substitution is three amino acids. In some embodiments, two or more substitutions compared to the reference CasX protein are the same length, e.g., both substitutions are two consecutive amino acids in length.
[0231] In the substitutions described herein, any amino acid can be substituted with any other amino acid. The substitution can be conservative (e.g., a basic amino acid is substituted with another basic amino acid). The substitution can be non-conservative (e.g., a basic amino acid is substituted with an acidic amino acid, or vice versa). For example, a proline in a reference CasX protein can be substituted with any of arginine, histidine, lysine, aspartic acid, glutamic acid, serine, threonine, asparagine, glutamine, cysteine, glycine, alanine, isoleucine, leucine, methionine, phenylalanine, tryptophan, tyrosine, or valine to generate a CasX variant protein of the present disclosure.
[0232] In some embodiments, the CasX variant protein comprises at least one amino acid deletion relative to a reference CasX protein. In some embodiments, the CasX variant protein comprises a deletion of 1 to 4 amino acids, 1 to 10 amino acids, 1 to 20 amino acids, 1 to 30 amino acids, 1 to 40 amino acids, 1 to 50 amino acids, 1 to 60 amino acids, 1 to 70 amino acids, 1 to 80 amino acids, 1 to 90 amino acids, 1 to 100 amino acids, 2 to 10 amino acids, 2 to 20 amino acids, 2 to 30 amino acids, 3 to 10 amino acids, 3 to 20 amino acids, 3 to 30 amino acids, 4 to 10 amino acids, 4 to 20 amino acids, 3 to 300 amino acids, 5 to 10 amino acids, 5 to 20 amino acids, 5 to 30 amino acids, 10 to 50 amino acids, or 20 to 50 amino acids relative to a reference CasX protein. In some embodiments, the CasX variant comprises a deletion of at least about 100 consecutive amino acids compared to a reference CasX protein. In some embodiments, the CasX variant protein comprises a deletion of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, or 100 consecutive amino acids compared to a reference CasX protein. In some embodiments, the CasX variant protein comprises a deletion of 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 consecutive amino acids.
[0233] In some embodiments, the CasX variant protein comprises two or more deletions compared to a reference CasX protein, where the two or more deletions are not consecutive amino acids. For example, the first deletion can be in a first domain of the reference CasX protein, and the second deletion can be in a second domain of the reference CasX protein. In some embodiments, the CasX variant protein comprises 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 non-consecutive deletions compared to the reference CasX protein. In some embodiments, the CasX variant protein comprises at least 20 non-consecutive deletions compared to the reference CasX protein. Each non-consecutive deletion can be any length of amino acids described herein, e.g., 1 to 4 amino acids, 1 to 10 amino acids, etc.
[0234] In some embodiments, the CasX variant protein comprises at least one amino acid insertion, such as an insertion of 1 amino acid, 2 to 3 consecutive amino acids, 2 to 4 consecutive amino acids, 2 to 5 consecutive amino acids, 2 to 6 consecutive amino acids, 2 to 7 consecutive amino acids, 2 to 8 consecutive amino acids, 2 to 9 consecutive amino acids, 2 to 10 consecutive amino acids, 2 to 20 consecutive amino acids, 2 to 30 consecutive amino acids, 2 to 40 consecutive amino acids, 2 to 50 consecutive amino acids, or 2 to 60 consecutive amino acids, relative to a reference CasX protein. The CasX variant protein may comprise an insertion of 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 consecutive amino acids. In some embodiments, the CasX variant protein comprises an insertion of at least about 100 consecutive amino acids.
[0235] In some embodiments, the CasX variant protein contains two or more insertions compared to a reference CasX protein, and the two or more insertions are not consecutive amino acids in the sequence. For example, a first insertion can be in a first domain of the reference CasX protein, and a second insertion can be in a second domain of the reference CasX protein. In some embodiments, the CasX variant protein contains 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 non-contiguous insertions compared to the reference CasX protein. In some embodiments, the CasX variant protein contains at least 10 to about 20 or more non-contiguous insertions compared to the reference CasX protein. Each non-contiguous insertion can be any length of amino acids described herein, e.g., 1 to 4 amino acids, 1 to 10 amino acids, etc.
[0236] Any amino acid or combination of amino acids can be inserted as described herein, for example, proline, arginine, histidine, lysine, aspartic acid, glutamic acid, serine, threonine, asparagine, glutamine, cysteine, glycine, alanine, isoleucine, leucine, methionine, phenylalanine, tryptophan, tyrosine, or valine, or any combination thereof, can be inserted into a reference CasX protein of the present disclosure to generate a CasX variant protein.
[0237] Any permutation of the substitution, insertion, and deletion embodiments described herein can be combined to generate a CasX variant protein of the present disclosure. For example, a CasX variant protein can include at least one substitution and at least one deletion compared to a reference CasX protein sequence, at least one substitution and at least one insertion compared to a reference CasX protein sequence, at least one insertion and at least one deletion compared to a reference CasX protein sequence, or at least one substitution, one insertion, and one deletion compared to a reference CasX protein.
[0238] In some embodiments, the CasX variant protein has at least about 60% sequence similarity, at least 70% similarity, at least 80% similarity, at least 85% similarity, at least 86% similarity, at least 87% similarity, at least 88% similarity, at least 89% similarity, at least 90% similarity, at least 91% similarity, at least 92% similarity, at least 93% similarity, at least 94% similarity, at least 95% similarity, at least 96% similarity, at least 97% similarity, at least 98% similarity, at least 99% similarity, at least 99.5% similarity, at least 99.6% similarity, at least 99.7% similarity, at least 99.8% similarity, or at least 99.9% similarity to one of SEQ ID NO:1, SEQ ID NO:2, or SEQ ID NO:3.
[0239] In some embodiments, the CasX variant protein has at least about 60% sequence similarity to SEQ ID NO:2, or a portion thereof. In some embodiments, the CasX variant protein has a Y789T substitution in SEQ ID NO:2, a deletion of P793 in SEQ ID NO:2, a Y789D substitution in SEQ ID NO:2, a T72S substitution in SEQ ID NO:2, a I546V substitution in SEQ ID NO:2, a E552A substitution in SEQ ID NO:2, a A636D substitution in SEQ ID NO:2, a F536S substitution in SEQ ID NO:2, a A708K substitution in SEQ ID NO:2, a Y797L substitution in SEQ ID NO:2, a L792G substitution in SEQ ID NO:2, a A739V substitution in SEQ ID NO:2, a G791M substitution in SEQ ID NO:2, an insertion of an A at position 661 in SEQ ID NO:2, Substitution of A788W in SEQ ID NO:2, substitution of K390R in SEQ ID NO:2, substitution of A751S in SEQ ID NO:2, substitution of E385A in SEQ ID NO:2, insertion of P at position 696 in SEQ ID NO:2, insertion of M at position 773 in SEQ ID NO:2, substitution of G695H in SEQ ID NO:2, insertion of AS at position 793 in SEQ ID NO:2, insertion of AS at position 795 in SEQ ID NO:2, substitution of C477R in SEQ ID NO:2, substitution of C477K in SEQ ID NO:2, substitution of C479A in SEQ ID NO:2, substitution of C479L in SEQ ID NO:2, substitution of I55F in SEQ ID NO:2, substitution of K210R in SEQ ID NO:2, sequence Substitution of C233S in SEQ ID NO:2, substitution of D231N in SEQ ID NO:2, substitution of Q338E in SEQ ID NO:2, substitution of Q338R in SEQ ID NO:2, substitution of L379R in SEQ ID NO:2, substitution of K390R in SEQ ID NO:2, substitution of L481Q in SEQ ID NO:2, substitution of F495S in SEQ ID NO:2, substitution of D600N in SEQ ID NO:2, substitution of T886K in SEQ ID NO:2, substitution of A739V in SEQ ID NO:2, substitution of K460N in SEQ ID NO:2, substitution of I199F in SEQ ID NO:2, substitution of G492P in SEQ ID NO:2, substitution of T153I in SEQ ID NO:2, substitution of R591I in SEQ ID NO:2 Substitution, insertion of AS at position 795 of SEQ ID NO:2, insertion of AS at position 796 of SEQ ID NO:2, insertion of L at position 889 of SEQ ID NO:2, substitution of E121D of SEQ ID NO:2, substitution of S270W of SEQ ID NO:2, substitution of E712Q of SEQ ID NO:2, substitution of K942Q of SEQ ID NO:2, substitution of E552K of SEQ ID NO:2, substitution of K25Q of SEQ ID NO:2, substitution of N47D of SEQ ID NO:2, insertion of T at position 696 of SEQ ID NO:2, substitution of L685I of SEQ ID NO:2, substitution of N880D of SEQ ID NO:2, substitution of Q102R of SEQ ID NO:2, substitution of M734K of SEQ ID NO:2,Substitution of A724S in SEQ ID NO:2, substitution of T704K in SEQ ID NO:2, substitution of P224K in SEQ ID NO:2, substitution of K25R in SEQ ID NO:2, substitution of M29E in SEQ ID NO:2, substitution of H152D in SEQ ID NO:2, substitution of S219R in SEQ ID NO:2, substitution of E475K in SEQ ID NO:2, substitution of G226R in SEQ ID NO:2, substitution of A377K in SEQ ID NO:2, substitution of E480K in SEQ ID NO:2, substitution of K416E in SEQ ID NO:2, substitution of H164R in SEQ ID NO:2, substitution of K767R in SEQ ID NO:2, substitution of I7F in SEQ ID NO:2, substitution of M29R in SEQ ID NO:2, substitution of H435 in SEQ ID NO:2 Substitution of R, substitution of E385Q in SEQ ID NO:2, substitution of E385K in SEQ ID NO:2, substitution of I279F in SEQ ID NO:2, substitution of D489S in SEQ ID NO:2, substitution of D732N in SEQ ID NO:2, substitution of A739T in SEQ ID NO:2, substitution of W885R in SEQ ID NO:2, substitution of E53K in SEQ ID NO:2, substitution of A238T in SEQ ID NO:2, substitution of P283Q in SEQ ID NO:2, substitution of E292K in SEQ ID NO:2, substitution of Q628E in SEQ ID NO:2, substitution of R388Q in SEQ ID NO:2, substitution of G791M in SEQ ID NO:2, substitution of L792K in SEQ ID NO:2, substitution of L792E in SEQ ID NO:2, Substitution of M779N in SEQ ID NO:2, Substitution of G27D in SEQ ID NO:2, Substitution of K955R in SEQ ID NO:2, Substitution of S867R in SEQ ID NO:2, Substitution of R693I in SEQ ID NO:2, Substitution of F189Y in SEQ ID NO:2, Substitution of V635M in SEQ ID NO:2, Substitution of F399L in SEQ ID NO:2, Substitution of E498K in SEQ ID NO:2, Substitution of E386R in SEQ ID NO:2, Substitution of V254G in SEQ ID NO:2, Substitution of P793S in SEQ ID NO:2, Substitution of K188E in SEQ ID NO:2, Substitution of QT945KI in SEQ ID NO:2, Substitution of T620P in SEQ ID NO:2, Substitution of T946P in SEQ ID NO:2, Substitution of TT9 in SEQ ID NO:2 Substitution of 49PP, substitution of N952T in SEQ ID NO:2, substitution of K682E in SEQ ID NO:2, substitution of K975R in SEQ ID NO:2, substitution of L212P in SEQ ID NO:2, substitution of E292R in SEQ ID NO:2, substitution of I303K in SEQ ID NO:2, substitution of C349E in SEQ ID NO:2, substitution of E385P in SEQ ID NO:2, substitution of E386N in SEQ ID NO:2, substitution of D387K in SEQ ID NO:2, substitution of L404K in SEQ ID NO:2, substitution of E466H in SEQ ID NO:2, substitution of C477Q in SEQ ID NO:2, substitution of C477H in SEQ ID NO:2, substitution of C479A in SEQ ID NO:2, substitution of D659H in SEQ ID NO:2,Substitution of T806V in SEQ ID NO:2, substitution of K808S in SEQ ID NO:2, insertion of AS at position 797 in SEQ ID NO:2, substitution of V959M in SEQ ID NO:2, substitution of K975Q in SEQ ID NO:2, substitution of W974G in SEQ ID NO:2, substitution of A708Q in SEQ ID NO:2, substitution of V711K in SEQ ID NO:2, substitution of D733T in SEQ ID NO:2, substitution of L742W in SEQ ID NO:2, substitution of V747K in SEQ ID NO:2, substitution of F755M in SEQ ID NO:2, substitution of M771A in SEQ ID NO:2, substitution of M771Q in SEQ ID NO:2, substitution of W782Q in SEQ ID NO:2, substitution of G791F in SEQ ID NO:2 , substitution of L792D in SEQ ID NO:2, substitution of L792K in SEQ ID NO:2, substitution of P793Q in SEQ ID NO:2, substitution of P793G in SEQ ID NO:2, substitution of Q804A in SEQ ID NO:2, substitution of Y966N in SEQ ID NO:2, substitution of Y723N in SEQ ID NO:2, substitution of Y857R in SEQ ID NO:2, substitution of S890R in SEQ ID NO:2, substitution of S932M in SEQ ID NO:2, substitution of L897M in SEQ ID NO:2, substitution of R624G in SEQ ID NO:2, substitution of S603G in SEQ ID NO:2, substitution of N737S in SEQ ID NO:2, substitution of L307K in SEQ ID NO:2, substitution of I658V in SEQ ID NO:2 Insertion of PT at position 688 of sequence number 2, insertion of SA at position 794 of sequence number 2, substitution of S877R of sequence number 2, substitution of N580T of sequence number 2, substitution of V335G of sequence number 2, substitution of T620S of sequence number 2, substitution of W345G of sequence number 2, substitution of T280S of sequence number 2, substitution of L406P of sequence number 2, substitution of A612D of sequence number 2, substitution of A751S of sequence number 2, substitution of E386R of sequence number 2, substitution of V351M of sequence number 2, substitution of K210N of sequence number 2, substitution of D40A of sequence number 2, substitution of E773G of sequence number 2 substitution of H207L in SEQ ID NO:2, substitution of T62A in SEQ ID NO:2, substitution of T287P in SEQ ID NO:2, substitution of T832A in SEQ ID NO:2, substitution of A893S in SEQ ID NO:2, insertion of V at position 14 of SEQ ID NO:2, insertion of AG at position 13 of SEQ ID NO:2, substitution of R11V in SEQ ID NO:2, substitution of R12N in SEQ ID NO:2, substitution of R13H in SEQ ID NO:2, insertion of Y at position 13 of SEQ ID NO:2, substitution of R12L in SEQ ID NO:2, insertion of Q at position 13 of SEQ ID NO:2, substitution of V15S in SEQ ID NO:2, insertion of D at position 17 of SEQ ID NO:2, or combinations thereof.
[0240] In some embodiments, the CasX variant comprises at least one modification in the NTSB domain.
[0241] In some embodiments, the CasX variant comprises at least one modification in the TSL domain, hi some embodiments, the at least one modification in the TSL domain comprises an amino acid substitution of one or more of amino acids Y857, S890, or S932 of SEQ ID NO:2.
[0242] In some embodiments, the CasX variant comprises at least one modification in the helical I domain. In some embodiments, the at least one modification in the helical I domain comprises an amino acid substitution of one or more of amino acids S219, L249, E259, Q252, E292, L307, or D318 of SEQ ID NO:2.
[0243] In some embodiments, the CasX variant comprises at least one modification in the helical II domain, in which the at least one modification in the helical II domain comprises an amino acid substitution of one or more of amino acids D361, L379, E385, E386, D387, F399, L404, R458, C477, or D489 of SEQ ID NO:2.
[0244] In some embodiments, the CasX variant comprises at least one modification in the OBD domain, in which the at least one modification in the OBD comprises an amino acid substitution of one or more of amino acids F536, E552, T620, or I658 of SEQ ID NO:2.
[0245] In some embodiments, the CasX variant comprises at least one modification in the RuvC DNA cleavage domain, in which the at least one modification in the RuvC DNA cleavage domain comprises an amino acid substitution of one or more of amino acids K682, G695, A708, V711, D732, A739, D733, L742, V747, F755, M771, M779, W782, A788, G791, L792, P793, Y797, M799, Q804, S819, or Y857 of SEQ ID NO: 2, or a deletion of amino acid P793.
[0246] In some embodiments, the CasX variant comprises at least one modification compared to the reference CasX sequence of SEQ ID NO: 2, selected from one or more of: (a) an amino acid substitution of L379R; (b) an amino acid substitution of A708K; (c) an amino acid substitution of T620P; (d) an amino acid substitution of E385P; (e) an amino acid substitution of Y857R; (f) an amino acid substitution of I658V; (g) an amino acid substitution of F399L; (h) an amino acid substitution of Q252K; (i) an amino acid substitution of L404K; and (j) an amino acid deletion of P793.
[0247] In some embodiments, the CasX variant protein contains at least two amino acid changes relative to the reference CasX protein amino acid sequence. The at least two amino acid changes may be substitutions, insertions, or deletions in the reference CasX protein amino acid sequence, or any combination thereof. The substitutions, insertions, or deletions may be any substitutions, insertions, or deletions in the sequence of the reference CasX protein described herein. In some embodiments, the changes are contiguous, non-contiguous, or a combination of contiguous and non-contiguous amino acid changes relative to the reference CasX protein sequence. In some embodiments, the reference CasX protein is SEQ ID NO:2. In some embodiments, the CasX variant protein comprises at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 30, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, or at least 100 amino acid changes relative to a reference CasX protein. In some embodiments, the CasX variant protein is 1-50, 3-40, 5-30, 5-20, 5-15, 5-10, 10-50, 10-40, 10-30, 10-20, 15-50, 15-40, 15-30, 2-25, 2-24, 2-22, 2-23, 2-22, 2-21, 2-20, 2-19, 2-18, 2-17, 2-16, 2-15, 2-1 4, 2-12, 2-11, 2-10, 2-9, 2-8, 2-7, 2-6, 2-5, 2-4, 2-3, 3-25, 3-24, 3-22, 3-23, 3-22, 3-21, 3-20, 3-19, 3-18, 3-17, 3-16, 3-15, 3-14, 3-12, 3-11, 3-10, 3-9, 3-8, 3-7, 3-6, 3-5, 3-4, 4-25, 4-24, 4-22, 4-23, 4-22, 4-21,In some embodiments, the CasX variant protein comprises 4 to 20, 4 to 19, 4 to 18, 4 to 17, 4 to 16, 4 to 15, 4 to 14, 4 to 12, 4 to 11, 4 to 10, 4 to 9, 4 to 8, 4 to 7, 4 to 6, 4 to 5, 5 to 25, 5 to 24, 5 to 22, 5 to 23, 5 to 22, 5 to 21, 5 to 20, 5 to 19, 5 to 18, 5 to 17, 5 to 16, 5 to 15, 5 to 14, 5 to 12, 5 to 11, 5 to 10, 5 to 9, 5 to 8, 5 to 7, or 5 to 6 amino acid changes. In some embodiments, the CasX variant protein comprises 15 to 20 changes relative to the reference CasX protein sequence. In some embodiments, the CasX variant protein comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 amino acid changes relative to a reference CasX protein sequence. In some embodiments, the at least two amino acid changes relative to the sequence of the reference CasX variant protein are a substitution of Y789T in SEQ ID NO:2, a deletion of P793 in SEQ ID NO:2, a substitution of Y789D in SEQ ID NO:2, a substitution of T72S in SEQ ID NO:2, a substitution of I546V in SEQ ID NO:2, a substitution of E552A in SEQ ID NO:2, a substitution of A636D in SEQ ID NO:2, a substitution of F536S in SEQ ID NO:2, a substitution of A708K in SEQ ID NO:2, a substitution of Y797L in SEQ ID NO:2, a substitution of L792G in SEQ ID NO:2, a substitution of A739V in SEQ ID NO:2, a substitution of G791M in SEQ ID NO:2, an insertion of an A at position 661 in SEQ ID NO:2, a substitution of A788W ... substitution of K390R in SEQ ID NO:2, a substitution of A75 Substitution of E385A in SEQ ID NO:2, insertion of P at position 696 in SEQ ID NO:2, insertion of M at position 773 in SEQ ID NO:2, substitution of G695H in SEQ ID NO:2, insertion of AS at position 793 in SEQ ID NO:2, insertion of AS at position 795 in SEQ ID NO:2, substitution of C477R in SEQ ID NO:2, substitution of C477K in SEQ ID NO:2, substitution of C479A in SEQ ID NO:2, substitution of C479L in SEQ ID NO:2, substitution of I55F in SEQ ID NO:2, substitution of K210R in SEQ ID NO:2, substitution of C233S in SEQ ID NO:2, substitution of D231N in SEQ ID NO:2, substitution of Q338E in SEQ ID NO:2, substitution of Q338R in SEQ ID NO:2, substitution of L379R in SEQ ID NO:2, substitution of K390R in SEQ ID NO:2, substitution of L481Q in SEQ ID NO:2, substitution of F495S in SEQ ID NO:2,Substitution of D600N in SEQ ID NO:2, substitution of T886K in SEQ ID NO:2, substitution of A739V in SEQ ID NO:2, substitution of K460N in SEQ ID NO:2, substitution of I199F in SEQ ID NO:2, substitution of G492P in SEQ ID NO:2, substitution of T153I in SEQ ID NO:2, substitution of R591I in SEQ ID NO:2, insertion of AS at position 795 in SEQ ID NO:2, insertion of AS at position 796 in SEQ ID NO:2, insertion of L at position 889 in SEQ ID NO:2, substitution of E121D in SEQ ID NO:2, substitution of S270W in SEQ ID NO:2, substitution of E712Q in SEQ ID NO:2, substitution of K942Q in SEQ ID NO:2, substitution of E552K in SEQ ID NO:2, Substitution of K25Q in SEQ ID NO:2, substitution of N47D in SEQ ID NO:2, insertion of T at position 696 in SEQ ID NO:2, substitution of L685I in SEQ ID NO:2, substitution of N880D in SEQ ID NO:2, substitution of Q102R in SEQ ID NO:2, substitution of M734K in SEQ ID NO:2, substitution of A724S in SEQ ID NO:2, substitution of T704K in SEQ ID NO:2, substitution of P224K in SEQ ID NO:2, substitution of K25R in SEQ ID NO:2, substitution of M29E in SEQ ID NO:2, substitution of H152D in SEQ ID NO:2, substitution of S219R in SEQ ID NO:2, substitution of E475K in SEQ ID NO:2, substitution of G226R in SEQ ID NO:2, A377 in SEQ ID NO:2 K substitution, E480K substitution in SEQ ID NO:2, K416E substitution in SEQ ID NO:2, H164R substitution in SEQ ID NO:2, K767R substitution in SEQ ID NO:2, I7F substitution in SEQ ID NO:2, M29R substitution in SEQ ID NO:2, H435R substitution in SEQ ID NO:2, E385Q substitution in SEQ ID NO:2, E385K substitution in SEQ ID NO:2, I279F substitution in SEQ ID NO:2, D489S substitution in SEQ ID NO:2, D732N substitution in SEQ ID NO:2, A739T substitution in SEQ ID NO:2, W885R substitution in SEQ ID NO:2, E53K substitution in SEQ ID NO:2, A238T substitution in SEQ ID NO:2 Substitution of P283Q, substitution of E292K in SEQ ID NO:2, substitution of Q628E in SEQ ID NO:2, substitution of R388Q in SEQ ID NO:2, substitution of G791M in SEQ ID NO:2, substitution of L792K in SEQ ID NO:2, substitution of L792E in SEQ ID NO:2, substitution of M779N in SEQ ID NO:2, substitution of G27D in SEQ ID NO:2, substitution of K955R in SEQ ID NO:2, substitution of S867R in SEQ ID NO:2, substitution of R693I in SEQ ID NO:2, substitution of F189Y in SEQ ID NO:2, substitution of V635M in SEQ ID NO:2, substitution of F399L in SEQ ID NO:2, substitution of E498K in SEQ ID NO:2, substitution of E386R in SEQ ID NO:2,Substitution of V254G in SEQ ID NO:2, substitution of P793S in SEQ ID NO:2, substitution of K188E in SEQ ID NO:2, substitution of QT945KI in SEQ ID NO:2, substitution of T620P in SEQ ID NO:2, substitution of T946P in SEQ ID NO:2, substitution of TT949PP in SEQ ID NO:2, substitution of N952T in SEQ ID NO:2, substitution of K682E in SEQ ID NO:2, substitution of K975R in SEQ ID NO:2, substitution of L212P in SEQ ID NO:2, substitution of E292R in SEQ ID NO:2, substitution of I303K in SEQ ID NO:2, substitution of C349E in SEQ ID NO:2, substitution of E385P in SEQ ID NO:2, substitution of E386N in SEQ ID NO:2 Substitution of D387K in SEQ ID NO:2, substitution of L404K in SEQ ID NO:2, substitution of E466H in SEQ ID NO:2, substitution of C477Q in SEQ ID NO:2, substitution of C477H in SEQ ID NO:2, substitution of C479A in SEQ ID NO:2, substitution of D659H in SEQ ID NO:2, substitution of T806V in SEQ ID NO:2, substitution of K808S in SEQ ID NO:2, insertion of AS at position 797 in SEQ ID NO:2, substitution of V959M in SEQ ID NO:2, substitution of K975Q in SEQ ID NO:2, substitution of W974G in SEQ ID NO:2, substitution of A708Q in SEQ ID NO:2, substitution of V711K in SEQ ID NO:2, substitution of D733T in SEQ ID NO:2, substitution of L in SEQ ID NO:2 Substitution of 742W, substitution of V747K in SEQ ID NO:2, substitution of F755M in SEQ ID NO:2, substitution of M771A in SEQ ID NO:2, substitution of M771Q in SEQ ID NO:2, substitution of W782Q in SEQ ID NO:2, substitution of G791F in SEQ ID NO:2, substitution of L792D in SEQ ID NO:2, substitution of L792K in SEQ ID NO:2, substitution of P793Q in SEQ ID NO:2, substitution of P793G in SEQ ID NO:2, substitution of Q804A in SEQ ID NO:2, substitution of Y966N in SEQ ID NO:2, substitution of Y723N in SEQ ID NO:2, substitution of Y857R in SEQ ID NO:2, substitution of S890R in SEQ ID NO:2, substitution of S932M in SEQ ID NO:2 , substitution of L897M in SEQ ID NO:2, substitution of R624G in SEQ ID NO:2, substitution of S603G in SEQ ID NO:2, substitution of N737S in SEQ ID NO:2, substitution of L307K in SEQ ID NO:2, substitution of I658V in SEQ ID NO:2, insertion of PT at position 688 in SEQ ID NO:2, insertion of SA at position 794 in SEQ ID NO:2, substitution of S877R in SEQ ID NO:2, substitution of N580T in SEQ ID NO:2, substitution of V335G in SEQ ID NO:2, substitution of T620S in SEQ ID NO:2, substitution of W345G in SEQ ID NO:2, substitution of T280S in SEQ ID NO:2, substitution of L406P in SEQ ID NO:2, substitution of A612D in SEQ ID NO:2,A substitution of A751S in SEQ ID NO:2, a substitution of E386R in SEQ ID NO:2, a substitution of V351M in SEQ ID NO:2, a substitution of K210N in SEQ ID NO:2, a substitution of D40A in SEQ ID NO:2, a substitution of E773G in SEQ ID NO:2, a substitution of H207L in SEQ ID NO:2, a substitution of T62A in SEQ ID NO:2, a substitution of T287P in SEQ ID NO:2, a substitution of T832A in SEQ ID NO:2, a substitution of A893S in SEQ ID NO:2, an insertion of V at position 14 in SEQ ID NO:2, an insertion of AG at position 13 in SEQ ID NO:2, a substitution of R11V in SEQ ID NO:2, a substitution of R12N in SEQ ID NO:2, a substitution of R13H in SEQ ID NO:2, an insertion of Y at position 13 in SEQ ID NO:2, a substitution of R12L in SEQ ID NO:2, an insertion of Q at position 13 in SEQ ID NO:2, a substitution of V15S in SEQ ID NO:2, and an insertion of D at position 17 in SEQ ID NO:2. In some embodiments, at least two amino acid changes relative to the reference CasX protein are selected from the amino acid changes disclosed in the sequences of Table 3. In some embodiments, the CasX variant comprises any combination of the preceding embodiments in this paragraph.
[0248] In some embodiments, the CasX variant protein comprises two or more substitutions, insertions, and / or deletions of the reference CasX protein amino acid sequence. In some embodiments, the reference CasX protein comprises or consists essentially of SEQ ID NO:2. In some embodiments, the CasX variant protein comprises the S794R substitution and the Y797L substitution of SEQ ID NO:2. In some embodiments, the CasX variant protein comprises the K416E substitution and the A708K substitution of SEQ ID NO:2. In some embodiments, the CasX variant protein comprises the A708K substitution and the P793 deletion of SEQ ID NO:2. In some embodiments, the CasX variant protein comprises the P793 deletion and the insertion of an AS at position 795 of SEQ ID NO:2. In some embodiments, the CasX variant protein comprises the Q367K substitution and the I425S substitution of SEQ ID NO:2. In some embodiments, the CasX variant protein comprises an A708K substitution, a deletion of P at position 793, and an A793V substitution of SEQ ID NO:2. In some embodiments, the CasX variant protein comprises a Q338R substitution and an A339E substitution of SEQ ID NO:2. In some embodiments, the CasX variant protein comprises a Q338R substitution and an A339K substitution of SEQ ID NO:2. In some embodiments, the CasX variant protein comprises a S507G substitution and a G508R substitution of SEQ ID NO:2. In some embodiments, the CasX variant protein comprises an L379R substitution, an A708K substitution, and a deletion of P at position 793 of SEQ ID NO:2. In some embodiments, the CasX variant protein comprises a C477K substitution, an A708K substitution, and a deletion of P at position 793 of SEQ ID NO:2. In some embodiments, the CasX variant protein comprises an L379R substitution, a C477K substitution, an A708K substitution, and a deletion of P at position 793 of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises an L379R substitution, an A708K substitution, a deletion of P at position 793, and an A739V substitution of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises a C477K substitution, an A708K substitution, a deletion of P at position 793, and an A739V substitution of SEQ ID NO: 2.In some embodiments, the CasX variant protein comprises a substitution of L379R, a substitution of C477K, a substitution of A708K, a deletion of P at position 793, and a substitution of A739V of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises a substitution of L379R, a substitution of A708K, a deletion of P at position 793, and a substitution of M779N of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises a substitution of L379R, a substitution of A708K, a deletion of P at position 793, and a substitution of M771N of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises a substitution of L379R, a substitution of 708K, a deletion of P at position 793, and a substitution of D489S of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises a substitution of L379R, a substitution of A708K, a deletion of P at position 793, and a substitution of A739T of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises a substitution of L379R, a substitution of A708K, a deletion of P at position 793, and a substitution of D732N of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises a substitution of L379R, a substitution of A708K, a deletion of P at position 793, and a substitution of G791M of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises a substitution of L379R, a substitution of 708K, a deletion of P at position 793, and a substitution of Y797L of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises a substitution of L379R, a substitution of C477K, a substitution of A708K, a deletion of P at position 793, and a substitution of M779N of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises a substitution of L379R, a substitution of C477K, a substitution of A708K, a deletion of P at position 793, and a substitution of M771N of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises a substitution of L379R, a substitution of C477K, a substitution of A708K, a deletion of P at position 793, and a substitution of D489S of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises a substitution of L379R, a substitution of C477K, a substitution of A708K, a deletion of P at position 793, and a substitution of A739T of SEQ ID NO: 2.In some embodiments, the CasX variant protein comprises a substitution of L379R, a substitution of C477K, a substitution of A708K, a deletion of P at position 793, and a substitution of D732N of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises a substitution of L379R, a substitution of C477K, a substitution of A708K, a deletion of P at position 793, and a substitution of G791M of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises a substitution of L379R, a substitution of C477K, a substitution of A708K, a deletion of P at position 793, and a substitution of Y797L of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises a substitution of L379R, a substitution of C477K, a substitution of A708K, a deletion of P at position 793, and a substitution of T620P of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises an A708K substitution, a P deletion at position 793, and an E386S substitution of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises an E386R substitution, an F399L substitution, and a P deletion at position 793 of SEQ ID NO: 2. In some embodiments, the CasX variant protein comprises an R581I and A739V substitution of SEQ ID NO: 2. In some embodiments, the CasX variant comprises any combination of the preceding embodiments in this paragraph.
[0249] In some embodiments, the CasX variant protein comprises two or more substitutions, insertions, and / or deletions of the reference CasX protein amino acid sequence. In some embodiments, the CasX variant protein comprises an A708K substitution, a P deletion at position 793, and an A739V substitution of SEQ ID NO:2. In some embodiments, the CasX variant protein comprises an L379R substitution, an A708K substitution, and a P deletion at position 793 of SEQ ID NO:2. In some embodiments, the CasX variant protein comprises a C477K substitution, an A708K substitution, and a P deletion at position 793 of SEQ ID NO:2. In some embodiments, the CasX variant protein comprises an L379R substitution, a C477K substitution, an A708K substitution, and a P deletion at position 793 of SEQ ID NO:2. In some embodiments, the CasX variant protein comprises an L379R substitution, an A708K substitution, a deletion of P at position 793, and an A739V substitution of SEQ ID NO:2. In some embodiments, the CasX variant protein comprises a C477K substitution, an A708K substitution, a deletion of P at position 793, and an A739 substitution of SEQ ID NO:2. In some embodiments, the CasX variant protein comprises an L379R substitution, a C477K substitution, an A708K substitution, a deletion of P at position 793, and an A739V substitution of SEQ ID NO:2. In some embodiments, the CasX variant protein comprises an L379R substitution, a C477K substitution, an A708K substitution, a deletion of P at position 793, and a T620P substitution of SEQ ID NO:2. In some embodiments, the CasX variant protein comprises an M771A substitution of SEQ ID NO:2. In some embodiments, the CasX variant protein comprises an L379R substitution, an A708K substitution, a deletion of P at position 793, and a D732N substitution of SEQ ID NO: 2. In some embodiments, the CasX variant comprises any combination of the preceding embodiments in this paragraph.
[0250] In some embodiments, the CasX variant protein comprises a W782Q substitution of SEQ ID NO:2. In some embodiments, the CasX variant protein comprises a M771Q substitution of SEQ ID NO:2. In some embodiments, the CasX variant protein comprises a R458I substitution and an A739V substitution of SEQ ID NO:2. In some embodiments, the CasX variant protein comprises a L379R substitution, an A708K substitution, a P deletion at position 793, and an M771N substitution of SEQ ID NO:2. In some embodiments, the CasX variant protein comprises a L379R substitution, an A708K substitution, a P deletion at position 793, and an A739T substitution of SEQ ID NO:2. In some embodiments, the CasX variant protein comprises a L379R substitution, a C477K substitution, an A708K substitution, a P deletion at position 793, and a D489S substitution of SEQ ID NO:2. In some embodiments, the CasX variant protein comprises a substitution of L379R, a substitution of C477K, a substitution of A708K, a deletion of P at position 793, and a substitution of D732N of SEQ ID NO:2. In some embodiments, the CasX variant protein comprises a substitution of V711K of SEQ ID NO:2. In some embodiments, the CasX variant protein comprises a substitution of L379R, a substitution of C477K, a substitution of A708K, a deletion of P at position 793, and a substitution of Y797L of SEQ ID NO:2. In some embodiments, the CasX variant protein comprises a substitution of L379R, a substitution of A708K, and a deletion of P at position 793 of SEQ ID NO:2. In some embodiments, the CasX variant protein comprises a substitution of L379R, a substitution of C477K, a substitution of A708K, a deletion of P at position 793, and a substitution of M771N of SEQ ID NO:2. In some embodiments, the CasX variant protein comprises an A708K substitution, a P substitution at position 793, and an E386S substitution of SEQ ID NO:2. In some embodiments, the CasX variant protein comprises an L379R substitution, a C477K substitution, an A708K substitution, and a deletion of P at position 793 of SEQ ID NO:2. In some embodiments, the CasX variant protein comprises an L792D substitution of SEQ ID NO:2. In some embodiments, the CasX variant protein comprises a G791F substitution of SEQ ID NO:2.In some embodiments, the CasX variant protein comprises an A708K substitution, a P deletion at position 793, and an A739V substitution of SEQ ID NO:2. In some embodiments, the CasX variant protein comprises an L379R substitution, an A708K substitution, a P deletion at position 793, and an A739V substitution of SEQ ID NO:2. In some embodiments, the CasX variant protein comprises a C477K substitution, an A708K substitution, and a P substitution at position 793 of SEQ ID NO:2. In some embodiments, the CasX variant protein comprises an L249I substitution and an M771N substitution of SEQ ID NO:2. In some embodiments, the CasX variant protein comprises a V747K substitution of SEQ ID NO:2. In some embodiments, the CasX variant protein comprises an L379R substitution, a C477 substitution, an A708K substitution, a P deletion at position 793, and an M779N substitution of SEQ ID NO:2. In some embodiments, the CasX variant protein comprises a substitution of F755M. In some embodiments, the CasX variant comprises any combination of the preceding embodiments in this paragraph.
[0251] In some embodiments, the CasX variant protein comprises at least one modification compared to the reference CasX sequence of SEQ ID NO: 2, wherein the at least one modification is selected from one or more of an L379R amino acid substitution, an A708K amino acid substitution, a T620P amino acid substitution, an E385P amino acid substitution, a Y857R amino acid substitution, an I658V amino acid substitution, an F399L amino acid substitution, a Q252K amino acid substitution, an L404K amino acid substitution, and a [P793] amino acid deletion. In other embodiments, the CasX variant protein comprises any combination of the foregoing substitutions or deletions compared to the reference CasX sequence of SEQ ID NO: 2. In other embodiments, the CasX variant protein can further comprise, in addition to the foregoing substitutions or deletions, a substitution of the NTSB and / or helical lb domain from the reference CasX of SEQ ID NO: 1.
[0252] In some embodiments, the CasX variant comprises any one of SEQ ID NOs: 247-337, 3301-3493, 3498-3501, 3505-3520, 3540-3549, and 4412-4415. In some embodiments, the CasX variant comprises any one of SEQ ID NOs: 247-337, 3498-3501, 3505-3520, 3540-3549, and 4412-4415. In some embodiments, the CasX variant comprises any one of SEQ ID NOs: 3498-3501, 3505-3520, and 3540-3549.
[0253] In some embodiments, the CasX variant comprises one or a modification to any one of SEQ ID NOs: 247-337, 3301-3493, 3498-3501, 3505-3520, 3540-3549, and 4412-4415. In some embodiments, the CasX variant comprises one or a modification to any one of SEQ ID NOs: 247-337, 3498-3501, 3505-3520, 3540-3549, and 4412-4415. In some embodiments, the CasX variant comprises one or a modification to any one of SEQ ID NOs: 3498-3501, 3505-3520, and 3540-3549.
[0254] In some embodiments, the CasX variant protein comprises between 400 and 2000 amino acids, between 500 and 1500 amino acids, between 700 and 1200 amino acids, between 800 and 1100 amino acids, or between 900 and 1000 amino acids.
[0255] In some embodiments, the CasX variant protein comprises one or more modifications in a region of non-adjacent residues that form a channel through which gNA:target DNA complex formation occurs. In some embodiments, the CasX variant protein comprises one or more modifications comprising a region of non-adjacent residues that form an interface that binds to the gNA. For example, in some embodiments of a reference CasX protein, the helical I, helical II, and OBD domains all contact or are close to the gNA:target DNA complex, and one or more modifications to non-adjacent residues within any of these domains can improve the function of the CasX variant protein.
[0256] In some embodiments, the CasX variant protein contains one or more modifications in a region of non-adjacent residues that form a channel that binds to non-target strand DNA. For example, the CasX variant protein can contain one or more modifications to non-adjacent residues in the NTSBD. In some embodiments, the CasX variant protein contains one or more modifications to a region of non-adjacent residues that form an interface that binds to the PAM. For example, the CasX variant protein can contain one or more modifications to non-adjacent residues in the helical I domain or the OBD. In some embodiments, the CasX variant protein contains one or more modifications that include a region of non-adjacent, surface-exposed residues. As used herein, "surface-exposed residue" refers to an amino acid on the surface of the CasX protein, or an amino acid where at least a portion of the amino acid, such as part of the backbone or side chain, is on the surface of the protein. Surface-exposed residues of cellular proteins, such as CasX, that are exposed to the aqueous intracellular environment are often selected from positively charged hydrophilic amino acids, such as arginine, asparagine, aspartic acid, glutamine, glutamic acid, histidine, lysine, serine, and threonine. Thus, for example, in some embodiments of the variants provided herein, the region of surface-exposed residues contains one or more insertions, deletions, or substitutions compared to a reference CasX protein. In some embodiments, one or more positively charged residues are replaced with one or more other positively charged residues, or negatively charged residues, or uncharged residues, or any combination thereof. In some embodiments, one or more amino acid residues for substitution are in a nearby bound nucleic acid; for example, residues in the RuvC domain or helical I domain that contacts target DNA, or residues in the OBD or helical II domain that binds to gNA, may be replaced with one or more positively charged or polar amino acids.
[0257] In some embodiments, the CasX variant protein contains one or more modifications in a region of non-adjacent residues that form a core through hydrophobic packing in the domain of a reference CasX protein. Without wishing to be bound by any theory, the region that forms the core through hydrophobic packing is rich in hydrophobic amino acids, such as valine, isoleucine, leucine, methionine, phenylalanine, tryptophan, and cysteine. For example, in some reference CasX proteins, the RuvC domain contains a hydrophobic pocket adjacent to the active site. In some embodiments, 2 to 15 residues in the region are charged, polar, or base-stacking. Charged amino acids (sometimes referred to herein as residues) can include, for example, arginine, lysine, aspartic acid, and glutamic acid, and the side chains of these amino acids can form salt bridges when a cross-linking partner is also present (see Figure 14). Polar amino acids can include, for example, glutamine, asparagine, histidine, serine, threonine, tyrosine, and cysteine. Polar amino acids, in some embodiments, can form hydrogen bonds as proton donors or acceptors, depending on the identity of their side chains. As used herein, "base stacking" refers to the interaction of aromatic side chains of amino acid residues (such as tryptophan, tyrosine, phenylalanine, or histidine) with stacked nucleotide bases in nucleic acids. Any modification to a region of non-adjacent amino acids that are spatially adjacent to form a functional portion of a CasX variant protein is contemplated within the scope of this disclosure.
[0258] i. CasX variant proteins with domains from multiple source proteins In certain embodiments, the present disclosure provides chimeric CasX proteins comprising protein domains from two or more different CasX proteins, e.g., two or more naturally occurring CasX proteins or two or more CasX variant protein sequences described herein. As used herein, "chimeric CasX protein" refers to a CasX comprising at least two domains isolated or derived from different sources, such as two naturally occurring proteins, which in some embodiments may be isolated from different species. For example, in some embodiments, the chimeric CasX protein comprises a first domain from a first CasX protein and a second domain from a second, different CasX protein. In some embodiments, the first domain may be selected from the group consisting of NTSB, TSL, helical I, helical II, OBD, and RuvC domains. In some embodiments, the second domain is selected from the group consisting of NTSB, TSL, helical I, helical II, OBD, and RuvC domains, and the second domain is different from the first domain. For example, a chimeric CasX protein can comprise the NTSB, TSL, helical I, helical II, and OBD domains from the CasX protein of SEQ ID NO: 2 and the RuvC domain from the CasX protein of SEQ ID NO: 1, or vice versa. As a further example, a chimeric CasX protein can comprise the NTSB, TSL, helical II, OBD, and RuvC domains from the CasX protein of SEQ ID NO: 2 and the helical I domain from the CasX protein of SEQ ID NO: 1, or vice versa. Thus, in certain embodiments, a chimeric CasX protein can comprise the NTSB, TSL, helical II, OBD, and RuvC domains from a first CasX protein and the helical I domain from a second CasX protein. In some embodiments of the chimeric CasX protein, a domain of a first CasX protein is derived from the sequence of SEQ ID NO:1, SEQ ID NO:2, or SEQ ID NO:3, and a domain of a second CasX protein is derived from the sequence of SEQ ID NO:1, SEQ ID NO:2, or SEQ ID NO:3, and the first and second CasX proteins are not the same.In some embodiments, a domain of a first CasX protein comprises a sequence derived from SEQ ID NO: 1, and a domain of a second CasX protein comprises a sequence derived from SEQ ID NO: 2. In some embodiments, a domain of a first CasX protein comprises a sequence derived from SEQ ID NO: 1, and a domain of a second CasX protein comprises a sequence derived from SEQ ID NO: 3. In some embodiments, a domain of a first CasX protein comprises a sequence derived from SEQ ID NO: 2, and a domain of a second CasX protein comprises a sequence derived from SEQ ID NO: 3. In some embodiments, the CasX variant is selected from the group consisting of CasX variants having the sequence of SEQ ID NO:328, SEQ ID NO:3540, SEQ ID NO:4413, SEQ ID NO:4414, SEQ ID NO:4415, SEQ ID NO:329, SEQ ID NO:3541, SEQ ID NO:330, SEQ ID NO:3542, SEQ ID NO:331, SEQ ID NO:3543, SEQ ID NO:332, SEQ ID NO:3544, SEQ ID NO:333, SEQ ID NO:3545, SEQ ID NO:334, SEQ ID NO:3546, SEQ ID NO:335, SEQ ID NO:3547, SEQ ID NO:336 and SEQ ID NO:3548. In some embodiments, the CasX variant comprises one or more further modifications to any one of SEQ ID NO:328, SEQ ID NO:3540, SEQ ID NO:4413, SEQ ID NO:4414, SEQ ID NO:4415, SEQ ID NO:329, SEQ ID NO:3541, SEQ ID NO:330, SEQ ID NO:3542, SEQ ID NO:331, SEQ ID NO:3543, SEQ ID NO:332, SEQ ID NO:3544, SEQ ID NO:333, SEQ ID NO:3545, SEQ ID NO:334, SEQ ID NO:3546, SEQ ID NO:335, SEQ ID NO:3547, SEQ ID NO:336, or SEQ ID NO:3548. In some embodiments, the one or more further modifications comprise an insertion, substitution, or deletion as described herein.
[0259] In some embodiments, the CasX variant protein comprises at least one chimeric domain comprising a first portion from a first CasX protein and a second portion from a second, different CasX protein. As used herein, "chimeric domain" refers to a domain comprising at least two portions isolated or derived from different sources, such as portions of domains from two naturally occurring proteins or two reference CasX proteins. The at least one chimeric domain can be any of the NTSB, TSL, helical I, helical II, OBD, or RuvC domains, as described herein. In some embodiments, the first portion of the CasX domain comprises the sequence of SEQ ID NO:1, and the second portion of the CasX domain comprises the sequence of SEQ ID NO:2. In some embodiments, the first portion of the CasX domain comprises the sequence of SEQ ID NO:1, and the second portion of the CasX domain comprises the sequence of SEQ ID NO:3. In some embodiments, the first portion of the CasX domain comprises the sequence of SEQ ID NO:2, and the second portion of the CasX domain comprises the sequence of SEQ ID NO:3. In some embodiments, at least one chimeric domain comprises a chimeric RuvC domain. As an example of the foregoing, the chimeric RuvC domain comprises amino acids 661-824 of SEQ ID NO:1 and amino acids 922-978 of SEQ ID NO:2. As an alternative example of the foregoing, the chimeric RuvC domain comprises amino acids 648-812 of SEQ ID NO:2 and amino acids 935-986 of SEQ ID NO:1. In some embodiments, the CasX protein comprises a first domain from a first CasX protein and a second domain from a second CasX protein, and at least one chimeric domain comprises at least two portions isolated from different CasX proteins using the approach of the embodiments described in this paragraph. In the foregoing embodiments, chimeric CasX proteins having domains or portions of domains derived from SEQ ID NOs:1, 2, and 3 may further comprise an amino acid insertion, deletion, or substitution of any of the embodiments disclosed herein.
[0260] In some embodiments, the CasX variant protein comprises a sequence set forth in Table 3, 8, 9, 10, or 12. In other embodiments, the CasX variant protein comprises a sequence at least 60% identical, at least 65% identical, at least 70% identical, at least 75% identical, at least 80% identical, at least 81% identical, at least 82% identical, at least 83% identical, at least 84% identical, at least 85% identical, at least 86% identical, at least 86% identical, at least 87% identical, at least 88% identical, at least 89% identical, at least 89% identical, at least 90% identical, at least 91% identical, at least 92% identical, at least 93% identical, at least 94% identical, at least 95% identical, at least 96% identical, at least 97% identical, at least 98% identical, at least 99% identical, or at least 99.5% identical to a sequence set forth in Table 3, 8, 9, 10, or 12. In other embodiments, the CasX variant protein comprises a sequence shown in Table 3 and further comprises one or more NLSs disclosed herein at either the N-terminus, the C-terminus, or both. It will be understood that in some cases, the N-terminal methionine of a CasX variant in the table is removed from the expressed CasX variant during post-translational modification. [Table 3-1] [Table 3-2] [Table 3-3] [Table 3-4] [Table 3-5]
[0261] In some embodiments, the CasX variant protein has one or more improved characteristics compared to a reference CasX protein, e.g., the reference protein of SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 3. In some embodiments, the improved characteristics of the CasX variant are at least about 1.1 to about 100,000-fold improved compared to the reference protein. In some embodiments, the improved characteristics of the CasX variant are at least about 1.1 to about 10,000-fold improved, at least about 1.1 to about 1,000-fold improved, at least about 1.1 to about 500-fold improved, at least about 1.1 to about 400-fold improved, at least about 1.1 to about 300-fold improved, at least about 1.1 to about 200-fold improved, at least about 1.1 to about 100-fold improved, at least about 1.1 to about 50-fold improved, at least about 1.1 to about 40-fold improved, at least about 1.1 to about 30-fold improved, at least about 1.1 to about 20-fold improved, at least about 1.1 to about 10-fold improved, at least about 1.1 to about 9-fold improved, or at least about 1.1 to about 100-fold improved. The improved characteristic of the CasX variant may be about 1.1 to about 8-fold improved, at least about 1.1 to about 7-fold improved, at least about 1.1 to about 6-fold improved, at least about 1.1 to about 5-fold improved, at least about 1.1 to about 4-fold improved, at least about 1.1 to about 3-fold improved, at least about 1.1 to about 2-fold improved, at least about 1.1 to about 1.5-fold improved, at least about 1.5 to about 3-fold improved, at least about 1.5 to about 4-fold improved, at least about 1.5 to about 5-fold improved, at least about 1.5 to about 10-fold improved, at least about 5 to about 10-fold improved, at least about 10 to about 20-fold improved, at least 10 to about 30-fold improved, at least 10 to about 50-fold improved, or at least 10 to about 100-fold improved. In some embodiments, the improved characteristic of the CasX variant is at least about 10 to about 1000-fold improved compared to the reference CasX protein.
[0262] In some embodiments, one or more improved characteristics of the CasX variant protein are improved by at least about 5, at least about 10, at least about 20, at least about 30, at least about 40, at least about 50, at least about 60, at least about 70, at least about 80, at least about 90, at least about 100, at least about 250, at least about 500, or at least about 1000, at least about 5,000, at least about 10,000, or at least about 100,000-fold compared to a reference CasX protein. In some embodiments, the improved characteristics of the CasX variant protein compared to a reference CasX protein include at least about 1.1, at least about 1.2, at least about 1.3, at least about 1.4, at least about 1.5, at least about 1.6, at least about 1.7, at least about 1.8, at least about 1.9, at least about 2, at least about 2.1, at least about 2.2, at least about 2.3, at least about 2.4, at least about 2.5, at least about 2.6, at least about 2.7, at least about 2.8, at least about 2.9, at least about 3 , at least about 3.5, at least about 4, at least about 4.5, at least about 5, at least about 5.5, at least about 6, at least about 6.5, at least about 7.0, at least about 7.5, at least about 8, at least about 8.5, at least about 9, at least about 9.5, at least about 10, at least about 11, at least about 12, at least about 13, at least about 14, at least about 15, at least about 20, at least about 30, at least about 40, at least about 50, at least about 60, at least about 70, at least about 80, at least about 90, at least about 100, at least about 500, at least about 1,000, at least about 10,000, or at least about 100,000 times improved.In other cases, one or more improved characteristics of the CasX variant may be about 1.1-100,00-fold, about 1.1-10,00-fold, about 1.1-1,000-fold, about 1.1-500-fold, about 1.1-100-fold, about 1.1-50-fold, about 1.1-20-fold, about 10-100,00-fold, about 10-10,00-fold, about 10-1,000-fold, about 10-500-fold, about 10-100-fold, about 10-50-fold, about 10-20-fold, about 2-70-fold, about 2-50-fold, about 2-30-fold, about 2-20-fold, about 2-10-fold, about 5-50-fold, or about 5-60-fold, as compared to a reference CasX of SEQ ID NO:1, SEQ ID NO:2, or SEQ ID NO:3. fold, about 5 to 30 fold, about 5 to 10 fold, about 100 to 100,00 fold, about 100 to 10,00 fold, about 100 to 1,000 fold, about 100 to 500 fold, about 500 to 100,00 fold, about 500 to 10,000 fold, about 500 to 1,000 fold, about 500 to 750 fold, about 1,000 to 100,00 fold, about 10,000 to 100,00 fold, about 20 to 500 fold, about 20 to 250 fold, about 20 to 200 fold, about 20 to 100 fold, about 20 to 50 fold, about 50 to 10,000 fold, about 50 to 1,000 fold, about 50 to 500 fold, about 50 to 200 fold, or about 50 to 100 fold improvement. In other cases, one or more improved characteristics of the CasX variant may be about 1.1-fold, 1.2-fold, 1.3-fold, 1.4-fold, 1.5-fold, 1.6-fold, 1.7-fold, 1.8-fold, 1.9-fold, 2-fold, 3-fold, 4-fold, 5-fold, 6-fold, 7-fold, 8-fold, 9-fold, 10-fold, 11-fold, 12-fold, 13-fold, 14-fold, 15-fold, 16-fold, 17-fold, 18-fold, 19-fold, 20-fold, 25-fold, 30-fold, 40-fold, 45-fold, 50-fold, 55-fold, or greater than ... 60x, 70x, 80x, 90x, 100x, 110x, 120x, 130x, 140x, 150x, 160x, 170x, 180x, 190x, 200x, 210x, 220x, 230x, 240x, 250x, 260x, 270x, 280x, 290x, 300x, 310x, 320x, 330x, 340x, 350x, 360x, 370x, 380x, 390x, 400x, 425x, 450x, 475x, or 500x improvement.Exemplary features that may be improved in a CasX variant protein compared to the same feature in a reference CasX protein include, but are not limited to, improved folding of the variant, improved binding affinity for gNAs, improved binding affinity for target DNA, improved ability to utilize a broader spectrum of PAM sequences in editing and / or binding to target DNA, improved unwinding of target DNA, increased editing activity, improved editing efficiency, improved editing specificity, increased nuclease activity, increased target strand loading for double-stranded cleavage, decreased target strand loading for single-stranded nicking, decreased off-target cleavage, improved binding of non-target strands of DNA, improved protein stability, improved CasX:gNA RNA complex stability, improved protein solubility, improved CasX:gNA RNP complex solubility, improved protein yield, improved protein expression, and improved fusion characteristics. In some embodiments, the variant comprises at least one improved feature. In other embodiments, the variant comprises at least two improved features. In further embodiments, the variant comprises at least three improved features. In some embodiments, the variants include at least four improved features. In still other embodiments, the variants include at least five, at least six, at least seven, at least eight, at least nine, at least ten, at least eleven, at least twelve, at least thirteen, or more improved features. These improved features are described in more detail below.
[0263] j. Protein stability In some embodiments, the present disclosure provides CasX variant proteins with improved stability compared to a reference CasX protein. In some embodiments, improved stability of the CasX variant protein results in higher steady-state expression of the protein, thereby improving editing efficiency. In some embodiments, improved stability of the CasX variant protein results in a greater proportion of the CasX protein remaining folded in a functional configuration, improving editing efficiency or improving purifiability for manufacturing purposes. As used herein, "functional configuration" refers to a CasX protein in a configuration in which the protein can bind to gNA and target DNA. In embodiments in which the CasX variant does not have one or more mutations that render it catalytically inactive, the CasX variant can cleave, nick, or otherwise modify target DNA. For example, functional CasX variants can be used in gene editing in some embodiments, and the functional configuration refers to an "editing-competent" configuration. In some exemplary embodiments, including those in which the CasX variant protein results in a greater proportion of CasX proteins remaining folded in a functional configuration, lower concentrations of the CasX variant are required for applications such as gene editing compared to a reference CasX protein. Thus, in some embodiments, a CasX variant with improved stability has improved efficiency compared to a reference CasX in one or more gene editing contexts.
[0264] In some embodiments, the present disclosure provides CasX variant proteins with improved thermostability compared to reference CasX proteins. In some embodiments, the CasX variant proteins have improved thermostability over a specific temperature range. Without wishing to be bound by any theory, some reference CasX proteins naturally function in organisms that thrive in groundwater and sediments, and therefore some reference CasX proteins may have evolved to function optimally at low or high temperatures, which may be desirable for certain applications. For example, one application of CasX variant proteins is gene editing in mammalian cells, which is typically performed at about 37°C. In some embodiments, the CasX variant proteins described herein have improved thermostability compared to a reference CasX protein at temperatures of at least 16°C, at least 18°C, at least 20°C, at least 22°C, at least 24°C, at least 26°C, at least 28°C, at least 30°C, at least 32°C, at least 34°C, at least 35°C, at least 36°C, at least 37°C, at least 38°C, at least 39°C, at least 40°C, at least 41°C, at least 42°C, at least 44°C, at least 46°C, at least 48°C, at least 50°C, at least 52°C, or higher. In some embodiments, the CasX variant proteins have improved thermostability and functionality compared to a reference CasX protein, resulting in improved gene editing functionality, such as in mammalian gene editing applications, which may include human gene editing applications.
[0265] In some embodiments, the present disclosure provides CasX variant proteins that exhibit improved stability of the CasX variant protein:gNA RNP complex compared to a reference CasX protein:gNA complex, such that the RNP remains in a functional form. Improved stability may include increased thermal stability, resistance to proteolysis, improved pharmacokinetic properties, and stability across a variety of pH, salt, and isotonic conditions. Improved stability of the complex may, in some embodiments, result in improved editing efficiency.
[0266] In some embodiments, the present disclosure provides CasX variant proteins that have improved thermal stability of the CasX variant protein:gNA complex compared to a reference CasX protein:gNA complex. In some embodiments, the CasX variant protein has improved thermal stability compared to a reference CasX protein. In some embodiments, the CasX variant protein:gNA RNP complex has improved thermal stability compared to a complex containing a reference CasX protein at temperatures of at least 16°C, at least 18°C, at least 20°C, at least 22°C, at least 24°C, at least 26°C, at least 28°C, at least 30°C, at least 32°C, at least 34°C, at least 35°C, at least 36°C, at least 37°C, at least 38°C, at least 39°C, at least 40°C, at least 41°C, at least 42°C, at least 44°C, at least 46°C, at least 48°C, at least 50°C, at least 52°C, or higher. In some embodiments, the CasX variant protein improves the thermal stability of the CasX variant protein:gNA RNP complex compared to a reference CasX protein:gNA complex, resulting in improved functionality for gene editing applications, such as mammalian gene editing applications, which may include human gene editing applications.
[0267] In some embodiments, the improved stability and / or thermal stability of the CasX variant protein comprises faster folding kinetics of the CasX variant protein compared to a reference CasX protein, slower unfolding kinetics of the CasX variant protein compared to a reference CasX protein, greater free energy release upon folding of the CasX variant protein compared to a reference CasX protein, a higher temperature (Tm) at which 50% of the CasX variant protein is unfolded compared to a reference CasX protein, or any combination thereof. These characteristics can be improved over a wide range of values, for example, by at least 1.1-, at least 1.5-, at least 10-, at least 50-, at least 100-, at least 500-, at least 1,000-, at least 5,000-, or at least 10,000-fold compared to a reference CasX protein. In some embodiments, the improved thermal stability of the CasX variant protein comprises a higher Tm of the CasX variant protein compared to a reference CasX protein. In some embodiments, the Tm of the CasX variant protein is about 20°C to about 30°C, about 30°C to about 40°C, about 40°C to about 50°C, about 50°C to about 60°C, about 60°C to about 70°C, about 70°C to about 80°C, about 80°C to about 90°C, or about 90°C to about 100°C. Thermal stability is measured by the "melting temperature" (T), defined as the temperature at which half of the molecule is denatured. m) is determined. Methods for measuring protein stability characteristics such as Tm and the free energy of unfolding are known to those skilled in the art and can be measured using standard biochemical techniques in vitro. For example, Tm can be measured using differential scanning calorimetry, a thermal analysis technique that measures the difference in the amount of heat required to increase the temperature of a sample and a reference as a function of temperature (Chen et al. (2003) Pharm Res 20:1952-60, Ghirlando et al. (1999) Immunol Lett 68:47-52). Alternatively, or in addition, the Tm of a CasX variant protein can be measured using commercially available methods such as the ThermoFisher Protein Thermal Shift system. Alternatively, or in addition, circular dichroism can be used to measure folding and unfolding kinetics and Tm (Murray et al. (2002) J. Chromatogr Sci 40:343-9). Circular dichroism (CD) relies on the unequal absorption of left-handed and right-handed circularly polarized light by asymmetric molecules such as proteins. Certain protein structures, such as alpha helices and beta sheets, have characteristic CD spectra. Therefore, in some embodiments, CD can be used to determine the secondary structure of CasX variant proteins.
[0268] In some embodiments, the improved stability and / or thermostability of the CasX variant protein comprises improved folding kinetics of the CasX variant protein compared to a reference CasX protein, hi some embodiments, the folding kinetics of the CasX variant protein is improved by at least about 5, at least about 10, at least about 50, at least about 100, at least about 500, at least about 1,000, at least about 2,000, at least about 3,000, at least about 4,000, at least about 5,000, or at least about 10,000-fold improvement compared to the reference CasX protein. In some embodiments, the folding kinetics of the CasX variant protein is improved by at least about 1 kJ / mol, at least about 5 kJ / mol, at least about 10 kJ / mol, at least about 20 kJ / mol, at least about 30 kJ / mol, at least about 40 kJ / mol, at least about 50 kJ / mol, at least about 60 kJ / mol, at least about 70 kJ / mol, at least about 80 kJ / mol, at least about 90 kJ / mol, at least about 100 kJ / mol, at least about 150 kJ / mol, at least about 200 kJ / mol, at least about 250 kJ / mol, at least about 300 kJ / mol, at least about 350 kJ / mol, at least about 400 kJ / mol, at least about 450 kJ / mol, or at least about 500 kJ / mol compared to a reference CasX protein.
[0269] Exemplary amino acid changes that can increase the stability of a CasX variant protein relative to a reference CasX protein can include, but are not limited to, amino acid changes that increase the number of hydrogen bonds in the CasX variant protein, amino acid changes that increase the number of disulfide bridges in the CasX variant protein, amino acid changes that increase the number of salt bridges in the CasX variant protein, amino acid changes that strengthen interactions between portions of the CasX variant protein, amino acid changes that increase buried hydrophobic surface areas of the CasX variant protein, or any combination thereof.
[0270] k. Protein yield In some embodiments, the present disclosure provides CasX variant proteins with improved yields during expression and purification compared to a reference CasX protein. In some embodiments, the yield of the CasX variant protein purified from a bacterial or eukaryotic host cell is improved compared to the reference CasX protein. In some embodiments, the bacterial host cell is an Escherichia coli cell. In some embodiments, the eukaryotic cell is a yeast, plant (e.g., tobacco), insect (e.g., Spodoptera frugiperda sf9 cell), mouse, rat, hamster, guinea pig, non-human primate, or human cell. In some embodiments, the eukaryotic host cell is a mammalian cell, including but not limited to, HEK293 cells, HEK293T cells, HEK293-F cells, Lenti-X 293T cells, BHK cells, HepG2 cells, Saos-2 cells, HuH7 cells, A549 cells, NS0 cells, SP2 / 0 cells, YO myeloma cells, P3X63 mouse myeloma cells, PER cells, PER.C6 cells, hybridoma cells, VERO cells, NIH3T3 cells, COS, WI38 cells, MRC5 cells, HeLa, HT1080 cells, or CHO cells.
[0271] In some embodiments, improved yields of CasX variant proteins are achieved through codon optimization. Cells use 64 different codons, 61 of which encode the 20 standard amino acids and three of which function as stop codons. In some cases, a single amino acid is encoded by more than one codon. Different organisms tend to use different codons for the same naturally occurring amino acid. Therefore, the choice of codons in a protein and the choice of codons that match the organism in which the protein is expressed can, in some cases, significantly affect protein translation and therefore protein expression levels. In some embodiments, the CasX variant protein is encoded by a codon-optimized nucleic acid. In some embodiments, the nucleic acid encoding the CasX variant protein is codon-optimized for expression in bacterial cells, yeast cells, insect cells, plant cells, or mammalian cells. In some embodiments, the mammalian cells are mouse, rat, hamster, guinea pig, monkey, or human. In some embodiments, the CasX variant protein is encoded by a nucleic acid that is codon-optimized for expression in human cells. In some embodiments, the CasX variant protein is encoded by a nucleic acid that has been deleted for nucleotide sequences that reduce the translation rate in prokaryotes and eukaryotes. For example, a run of more than three consecutive thymine residues may reduce the translation rate in certain organisms, or an internal polyadenylation signal may reduce translation.
[0272] In some embodiments, improved solubility and stability, as described herein, results in improved yield of the CasX variant protein compared to a reference CasX protein.
[0273] Improved protein yield during expression and purification can be assessed by methods known in the art. For example, the amount of CasX variant protein can be determined by running the protein on an SDS-page gel and comparing the CasX variant protein to a control of known quantity or concentration to determine absolute protein levels. Alternatively, or in addition, purified CasX variant protein can be run on an SDS-page gel next to a reference CasX protein that has undergone the same purification process to determine the relative improvement in CasX variant protein yield. Alternatively, or in addition, protein levels can be measured using immunohistochemical methods such as Western blot or ELISA using antibodies against CasX, or by HPLC. For proteins in solution, concentration can be determined by measuring the protein's intrinsic UV absorbance or by methods that use a protein-dependent color change, such as the Lowry assay, Smith copper / bicinchoninic assay, or Bradford dye assay. Such methods can be used to calculate the yield of total protein (e.g., total soluble protein) obtained by expression under specific conditions. This can be compared, for example, to the protein yield of a reference CasX protein under similar expression conditions.
[0274] l. protein solubility In some embodiments, the CasX variant protein has improved solubility compared to a reference CasX protein. In some embodiments, the CasX variant protein has improved solubility of the CasX:gNA ribonucleoprotein complex variant compared to a ribonucleoprotein complex comprising the reference CasX protein.
[0275] In some embodiments, improved protein solubility results in higher protein yields from protein purification techniques, such as purification from E. coli. In some embodiments, improved solubility of the CasX variant protein may allow for more efficient activity within cells, as more soluble proteins are less likely to aggregate within cells. Protein aggregates may, in certain embodiments, be toxic or burdensome to cells, and, without wishing to be bound by any theory, increased solubility of the CasX variant protein may ameliorate this consequence of protein aggregation. Furthermore, improved solubility of the CasX variant protein may allow for enhanced formulations that allow for the delivery of higher effective amounts of functional protein, for example, in desired gene editing applications. In some embodiments, the improved solubility of the CasX variant protein compared to the reference CasX protein results in an improved yield of the CasX variant protein during purification that is at least about 5, at least about 10, at least about 20, at least about 30, at least about 40, at least about 50, at least about 60, at least about 70, at least about 80, at least about 90, at least about 100, at least about 250, at least about 500, or at least about 1000 times greater.In some embodiments, the improved solubility of the CasX variant protein compared to a reference CasX protein increases the activity of the CasX variant protein in a cell by at least about 1.1, at least about 1.2, at least about 1.3, at least about 1.4, at least about 1.5, at least about 1.6, at least about 1.7, at least about 1.8, at least about 1.9, at least about 2, at least about 2.1, at least about 2.2, at least about 2.3, at least about 2.4, at least about 2.5, at least about 2.6, at least about at least about 2.7, at least about 2.8, at least about 2.9, at least about 3, at least about 3.5, at least about 4, at least about 4.5, at least about 5, at least about 5.5, at least about 6, at least about 6.5, at least about 7.0, at least about 7.5, at least about 8, at least about 8.5, at least about 9, at least about 9.5, at least about 10, at least about 11, at least about 12, at least about 13, at least about 14, at least about 15-fold, or at least about 20-fold greater activity.
[0276] Methods for measuring the solubility of CasX proteins and their improvement in CasX variant proteins will be readily apparent to those skilled in the art. For example, in some embodiments, the solubility of CasX variant proteins can be measured by performing densitometric readings on a gel of the soluble fraction of lysed E. coli. Alternatively, or in addition, the improved solubility of CasX variant proteins can be measured by measuring the maintenance of soluble protein product throughout the entire protein purification process, including the methods described in the Examples. For example, soluble protein product can be measured at one or more steps of gel affinity purification, tag cleavage, cation exchange purification, or running the protein on a size-exclusion chromatography (SEC) column. In some embodiments, densitometry readings of all protein bands on a gel are taken after each step in the purification process. In some embodiments, CasX variant proteins with improved solubility may maintain a higher concentration during one or more steps in the protein purification process compared to a reference CasX protein, while insoluble protein variants may be lost during one or more steps due to buffer exchange, filtration steps, interactions with the purification column, etc.
[0277] In some embodiments, the improved solubility of the CasX variant protein results in a higher yield in terms of mg / L of protein during protein purification compared to a reference CasX protein.
[0278] In some embodiments, the improved solubility of the CasX variant protein allows for a greater amount of editing events compared to a less soluble protein, as assessed in an editing assay such as the EGFP disruption assay described herein.
[0279] Affinity for m.gNA In some embodiments, the CasX variant protein has improved affinity for gNA compared to a reference CasX protein, resulting in the formation of a ribonucleoprotein complex. The increased affinity of the CasX variant protein for gNA may result in, for example, a lower K for the generation of an RNP complex. d In some embodiments, the K of the CasX variant protein relative to the gNA can be increased, which in some cases can result in more stable ribonucleoprotein complex formation. d is increased by at least about 1.1, at least about 1.2, at least about 1.3, at least about 1.4, at least 1.5, at least about 1.6, at least about 1.7, at least about 1.8, at least about 1.9, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 15, at least about 20, at least about 25, at least about 30, at least about 35, at least about 40, at least about 45, at least about 50, at least about 60, at least about 70, at least about 80, at least about 90, or at least about 100-fold compared to a reference CasX protein. In some embodiments, the CasX variant has an increased binding affinity for gNA by about 1.1 to about 10-fold compared to a reference CasX protein of SEQ ID NO: 2.
[0280] In some embodiments, the increased affinity of the CasX variant protein for gNA results in increased stability of the ribonucleoprotein complex when delivered to mammalian cells, including in vivo delivery to a subject. This increased stability may not only affect the function and utility of the complex within a subject's cells, but may also result in improved pharmacokinetic properties in the blood when delivered to a subject. In some embodiments, the increased affinity of the CasX variant protein and the resulting increased stability of the ribonucleoprotein complex allow for lower doses of the CasX variant protein to be delivered to a subject or cell while still achieving the desired activity, such as in vivo or in vitro gene editing. The increased ability to form RNPs and maintain them in a stable form can be assessed using assays such as the in vitro cleavage assay described herein. In some embodiments, the CasX variants of the present disclosure, when complexed with RNPs, ultimately exhibit a 2-fold, at least 5-fold, or at least 10-fold higher K when compared to the RNPs of a reference CasX. 切断 Speed can be achieved
[0281] In some embodiments, the higher affinity (tighter binding) of the CasX variant protein for gNA allows for a greater amount of editing events when both the CasX variant protein and the gNA remain in the RNP complex. Increased editing events can be assessed using editing assays such as the EGFP disruption and in vitro cleavage assays described herein.
[0282] Without wishing to be bound by theory, in some embodiments, amino acid changes in the helical I domain can increase the binding affinity of the CasX variant protein with a gNA targeting sequence, while changes in the helical II domain can increase the binding affinity of the CasX variant protein with a gNA scaffold stem loop, and changes in the oligonucleotide binding domain (OBD) increase the binding affinity of the CasX variant protein with a gNA triplex.
[0283] Methods for measuring CasX protein binding affinity to gNAs include in vitro methods using purified CasX protein and gNAs. If the gNA or CasX protein is tagged with a fluorophore, the binding affinity to the reference CasX and variant proteins can be measured by fluorescence polarization. Alternatively, or in addition, binding affinity can be measured by biolayer interferometry, electrophoretic mobility shift assay (EMSA), or filter binding. Additional standard techniques for quantifying the absolute affinity of RNA-binding proteins, such as reference CasX and variant proteins of the present disclosure, to specific gNAs, such as reference gNAs and their variants, include, but are not limited to, isothermal calorimetry (ITC) and surface plasmon resonance (SPR), as well as the methods described in the Examples.
[0284] n. affinity for target nucleic acid In some embodiments, the CasX variant proteins have improved binding affinity for a target nucleic acid compared to the affinity of a reference CasX protein for the target nucleic acid. CasX variants with higher affinity for their target nucleic acid can, in some embodiments, cleave a target nucleic acid sequence more rapidly than a reference CasX protein that does not have increased affinity for the target nucleic acid.
[0285] In some embodiments, improved affinity for a target nucleic acid includes improved affinity for the target sequence or protospacer sequence of the target nucleic acid, improved affinity for the PAM sequence, improved ability to search DNA for the target sequence, or any combination thereof. Without wishing to be bound by theory, it is believed that CRISPR / Cas system proteins, such as CasX, can find their target sequence by one-dimensional diffusion along a DNA molecule. This process is believed to involve (1) binding of the ribonucleoprotein to the DNA molecule followed by (2) termination at the target sequence, either of which, in some embodiments, can be affected by improved affinity of the CasX protein for the target nucleic acid sequence, thereby improving the function of the CasX variant protein compared to a reference CasX protein.
[0286] In some embodiments, CasX variant proteins with improved target nucleic acid affinity have increased overall affinity for DNA. In some embodiments, CasX variant proteins with improved target nucleic acid affinity have increased affinity for or ability to utilize specific PAM sequences other than the standard TTC PAM recognized by the reference CasX protein of SEQ ID NO:2, including PAM sequences selected from the group consisting of TTC, ATC, GTC, and CTC, thereby increasing the amount of target DNA that can be edited compared to wild-type CasX nuclease. Without wishing to be bound by theory, these protein variants may be able to interact more strongly with DNA overall and may have an increased ability to access and edit sequences within target DNA due to their ability to utilize additional PAM sequences beyond those of the wild-type reference CasX, thereby enabling a more efficient search process for the CasX protein for the target sequence. Higher overall affinity for DNA may also, in some embodiments, increase the frequency with which the CasX protein can efficiently initiate and complete binding and unwinding steps, thereby facilitating target strand invasion, R-loop formation, and ultimately cleavage of the target nucleic acid sequence.
[0287] Without wishing to be bound by theory, amino acid changes in the NTSBD that increase the efficiency of unwinding or capturing non-target DNA strands in the unwound state may increase the affinity of the CasX variant protein for target DNA. Alternatively, or in addition, amino acid changes in the NTSBD that increase the ability of the NTSBD to stabilize DNA during unwinding may increase the affinity of the CasX variant protein for target DNA. Alternatively, or in addition, amino acid changes in the OBD may increase the affinity of the CasX variant protein to bind to the protospacer adjacent-adjacent motif (PAM), thereby increasing the affinity of the CasX variant protein for target nucleic acids. Alternatively, or in addition, amino acid changes in the helical I and / or II, RuvC, and TSL domains that increase the affinity of the CasX variant protein for target nucleic acid strands may increase the affinity of the CasX variant protein for target nucleic acids.
[0288] In some embodiments, the binding affinity of a CasX variant protein of the present disclosure for a target nucleic acid molecule is increased by at least about 1.1, at least about 1.2, at least about 1.3, at least about 1.4, at least about 1.5, at least about 1.6, at least about 1.7, at least about 1.8, at least about 1.9, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 15, at least about 20, at least about 25, at least about 30, at least about 35, at least about 40, at least about 45, at least about 50, at least about 60, at least about 70, at least about 80, at least about 90, or at least about 100-fold compared to a reference CasX protein. In some embodiments, the CasX variant protein has an increased binding affinity for a target nucleic acid by about 1.1 to about 100-fold compared to the reference protein of SEQ ID NO:1, SEQ ID NO:2, or SEQ ID NO:3.
[0289] In some embodiments, the CasX variant protein has improved binding affinity for the non-target strand of the target nucleic acid. As used herein, the term "non-target strand" refers to a strand of the DNA target nucleic acid sequence that does not form Watson-Crick base pairs with the targeting sequence of the gNA and is complementary to the target DNA strand. In some embodiments, the CasX variant protein has an increased binding affinity for the non-target strand of the target nucleic acid of about 1.1 to about 100-fold compared to the reference protein of SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 3.
[0290] Methods for measuring the affinity of CasX proteins (such as reference or variants) for target and / or non-target nucleic acid molecules can include electrophoretic mobility shift assays (EMSA), filter binding, isothermal calorimetry (ITC), and surface plasmon resonance (SPR), fluorescence polarization, and biolayer interferometry (BLI). Additional methods for measuring the affinity of CasX proteins for targets include in vitro biochemical assays that measure DNA cleavage events over time.
[0291] o. Improved specificity for the target site In some embodiments, the CasX variant protein has improved specificity for the target nucleic acid sequence compared to a reference CasX protein. As used herein, "specificity," sometimes referred to as "target specificity," refers to the degree to which a CRISPR / Cas system ribonucleoprotein complex cleaves off-target sequences that are similar, but not identical, to the target nucleic acid sequence; for example, a CasX variant RNP with higher specificity will exhibit reduced off-target cleavage of sequences compared to a reference CasX protein. The specificity of a CRISPR / Cas system protein and the reduction of potentially harmful off-target effects can be crucial for achieving an acceptable therapeutic index for use in mammalian subjects.
[0292] In some embodiments, the CasX variant protein has improved specificity for a target site within a target sequence complementary to the targeting sequence of a gNA. Without wishing to be bound by theory, amino acid changes in the helical I and II domains that increase the specificity of the CasX variant protein for a target nucleic acid strand may increase the specificity of the CasX variant protein for the entire target nucleic acid. In some embodiments, amino acid changes that increase the specificity of the CasX variant protein for a target nucleic acid may also result in a decrease in the affinity of the CasX variant protein for DNA.
[0293] Methods for testing the target specificity of a CasX protein (e.g., variant or reference) can include guide and circularization for in vitro reporting of cleavage effects by sequencing (CIRCLE-seq) or similar methods. Briefly, in CIRCLE-seq technology, genomic DNA is sheared and circularized by ligation of stem-loop adapters, which are nicked in the stem-loop region to expose a four-nucleotide palindromic overhang. This is followed by intramolecular ligation and degradation of the remaining linear DNA. The circular DNA molecule containing the CasX cleavage site is then linearized with CasX, adapter adapters are ligated to the exposed ends, and high-throughput sequencing is then performed to generate paired-end reads containing information about off-target sites. Additional assays that can be used to detect off-target events, and thus CasX protein specificity, include assays used to detect and quantify indels (insertions and deletions) formed at selected off-target sites, such as mismatch-detecting nuclease assays and next-generation sequencing (NGS). An exemplary mismatch-detecting assay involves PCR-amplified, denatured, and rehybridized genomic DNA from cells treated with CasX and sgNA to form heteroduplex DNA containing one wild-type strand and one strand with an indel. The mismatch is recognized and cleaved by a mismatch-detecting nuclease, such as Surveyor nuclease or T7 endonuclease I.
[0294] p. protospacer and PAM sequence In this specification, the protospacer is defined as the DNA sequence complementary to the targeting sequence of the guide RNA and the DNA complementary to that sequence, and are referred to as the target strand and the non-target strand, respectively. As used herein, the PAM is a nucleotide sequence adjacent to the protospacer, which, together with the targeting sequence of the gNA, serves to orient and position CasX for potential cleavage of the protospacer strand.
[0295] PAM sequences may be degenerate, and specific RNP constructs may have different preferred PAM sequences that support different cleavage efficiencies. Unless otherwise indicated, by convention, this disclosure refers to both the PAM and protospacer sequences and their orientation relative to the orientation of the non-target strand. This does not imply that the PAM sequence of the non-target strand, rather than the target strand, determines cleavage or is mechanistically involved in target recognition. For example, when referring to a TTC PAM, it may actually be the complementary GAA sequence required for target cleavage, or it may be some combination of nucleotides from both strands. In the CasX proteins disclosed herein, the PAM is located 5' of the protospacer, with a single nucleotide separating it from the first nucleotide of the protospacer. Thus, for the reference CasX, the TTC PAM has the following formula: NNTTCN(protospacer)NNNNNN...3' (SEQ ID NO: 3296), where "N" is any DNA nucleotide and "(protospacer)" is a DNA sequence having identity to the targeting sequence of the guide RNA. For CasX variants with extended PAM recognition, TTC, CTC, GTC, or ATC PAM should be understood to mean a sequence of the following formula: 5'-...NNTTCN(protospacer)NNNNNN...3' (SEQ ID NO: 3296), 5'-...NNCTCN(protospacer)NNNNNN...3' (SEQ ID NO: 3297), 5'-...NNGTCN(protospacer)NNNNNN...3' (SEQ ID NO: 3298), or 5'-...NNATCN(protospacer)NNNNNN...3' (SEQ ID NO: 3299). Alternatively, TC PAM should be understood to mean a sequence of the following formula: 5'-...NNNTCN(protospacer)NNNNNN...3' (SEQ ID NO: 3300).
[0296] In some embodiments, the CasX variants have improved editing of PAM sequences, and exhibit greater editing efficiency and / or target sequence binding in target DNA when any one of the PAM sequences TTC, ATC, GTC, or CTC is located one nucleotide 5' to the non-target strand of a protospacer that has identity to the targeting sequence of a gNA in a cellular assay system, compared to the editing efficiency and / or binding of an RNP containing a reference CasX protein in an equivalent assay system. In some embodiments, the PAM sequence is TTC. In some embodiments, the PAM sequence is ATC. In some embodiments, the PAM sequence is CTC. In some embodiments, the PAM sequence is GTC.
[0297] q. DNA unwinding In some embodiments, the CasX variant protein has an improved ability to unwind DNA compared to a reference CasX protein. It has previously been shown that insufficient dsDNA unwinding impairs or prevents the ability of the CRISPR / Cas system proteins AnaCas9 or Cas14s to cleave DNA. Therefore, without wishing to be bound by any theory, the increased DNA cleavage activity of some of the CasX variant proteins disclosed herein is likely due, at least in part, to their increased ability to find and unwind dsDNA at target sites. Methods for measuring the ability of a CasX protein (such as a variant or reference) to unwind DNA include, but are not limited to, in vitro assays that observe an increased rate of dsDNA targeting in fluorescence polarization or biolayer interferometry.
[0298] Without wishing to be bound by theory, it is believed that amino acid changes in the NTSB domain can produce CasX variant proteins with increased DNA unwinding characteristics. Alternatively, or in addition, amino acid changes in the OBD or helical domain regions that interact with the PAM can also produce CasX variant proteins with increased DNA unwinding characteristics.
[0299] r.Catalytic activity The ribonucleoprotein complex of the CasX:gNA system disclosed herein comprises a reference CasX protein or a CasX variant complexed with a gNA that binds to a target nucleic acid and, in some cases, cleaves the target nucleic acid. In some embodiments, the CasX variant protein has improved catalytic activity compared to the reference CasX protein. Without wishing to be bound by theory, it is believed that in some cases, cleavage of the target strand may be a limiting factor for Cas12-like molecules in producing dsDNA cleavage. In some embodiments, the CasX variant protein improves bending of the target strand of DNA and cleavage of this strand, resulting in improved overall efficiency of dsDNA cleavage by the CasX ribonucleoprotein complex.
[0300] In some embodiments, the CasX variant protein has increased nuclease activity compared to a reference CasX protein. Variants with increased nuclease activity can be generated, for example, by amino acid changes in the RuvC nuclease domain. In some embodiments, amino acid substitutions at amino acid residues 708-804 of the RuvC domain can result in increased editing efficiency, as seen in Figure 10. In some embodiments, the CasX variant comprises a nuclease domain with nickase activity. In the aforementioned embodiments, the CasX nickase of the gene editing pair creates a single-strand break within 10-18 nucleotides 3' of the PAM site on the non-target strand. In other embodiments, the CasX variant comprises a nuclease domain with double-strand break activity. In the aforementioned embodiments, the CasX of the gene editing pair creates a double-strand break within 18-26 nucleotides 5' of the PAM site on the target strand and 10-18 nucleotides 3' of the PAM site on the non-target strand. Nuclease activity can be assayed by a variety of methods, including those in the Examples. In some embodiments, the CasX variant has a K that is at least 2-fold, or at least 3-fold, or at least 4-fold, or at least 5-fold, or at least 6-fold, or at least 7-fold, or at least 8-fold, or at least 9-fold, or at least 10-fold greater than a reference or wild-type CasX. 切断 It has a constant.
[0301] In some embodiments, the CasX variant protein has increased target strand loading for double-strand breaks. Variants with increased target strand loading activity can be generated, for example, by amino acid changes in the TLS domain. Without wishing to be bound by theory, amino acid changes in the TLS domain can result in CasX variant proteins with improved catalytic activity. Alternatively, or in addition, amino acid changes around the binding channel for RNA:DNA duplexes can also improve the catalytic activity of CasX variant proteins.
[0302] In some embodiments, the CasX variant protein has increased incidental cleavage activity compared to a reference CasX protein. As used herein, "incidental cleavage activity" refers to the additional non-targeted cleavage of nucleic acids after recognition and cleavage of a target nucleic acid. In some embodiments, the CasX variant protein has decreased incidental cleavage activity compared to a reference CasX protein.
[0303] In some embodiments, including those applications in which cleavage of a target nucleic acid is not a desired outcome, improving the catalytic activity of a CasX variant protein comprises altering, reducing, or abolishing the catalytic activity of a CasX variant protein. In some embodiments, a ribonucleoprotein complex comprising a dCasX variant protein binds to a target nucleic acid and does not cleave the target nucleic acid.
[0304] In some embodiments, a CasX ribonucleoprotein complex containing a CasX variant protein binds to target DNA but generates a single-stranded nick in the target DNA. In some embodiments, particularly those in which the CasX protein is a nickase, the CasX variant protein has reduced target strand loading for single-strand nicking. Variants with reduced target strand loading can be generated, for example, by amino acid changes in the TSL domain.
[0305] Exemplary methods for characterizing the catalytic activity of a CasX protein can include, but are not limited to, in vitro cleavage assays, including those in the Examples below. In some embodiments, electrophoresis of DNA products on an agarose gel can examine the kinetics of strand scission.
[0306] s. affinity for target RNA In some embodiments, a ribonucleoprotein complex comprising a reference CasX protein or a variant thereof binds to a target RNA and cleaves the target nucleic acid. In some embodiments, a variant of a reference CasX protein increases the specificity of the CasX variant protein for the target RNA or increases the activity of the CasX variant protein for the target RNA compared to the reference CasX protein. For example, the CasX variant protein may exhibit increased binding affinity for the target RNA or increased cleavage of the target RNA compared to the reference CasX protein. In some embodiments, a ribonucleoprotein complex comprising a CasX variant protein binds to and / or cleaves the target RNA. In some embodiments, the CasX variant has at least about 2-fold to about 10-fold increased binding affinity for the target nucleic acid compared to the reference protein of SEQ ID NO:1, SEQ ID NO:2, or SEQ ID NO:3.
[0307] t.CasX fusion protein In some embodiments, the present disclosure provides a CasX protein comprising a heterologous protein fused to CasX. In some cases, the CasX is a reference CasX protein. In other cases, the CasX is a CasX variant of any of the embodiments described herein.
[0308] In some embodiments, the CasX variant protein is fused to one or more proteins or domains thereof with different activities of interest, resulting in a fusion protein. For example, in some embodiments, the CasX variant protein is fused to a protein (or domain thereof) that inhibits transcription, modifies a target nucleic acid, or modifies a polypeptide associated with a nucleic acid (e.g., a histone modifier).
[0309] In some embodiments, the CasX variant comprises any one of SEQ ID NOs: 247-337, 3301-3493, 3498-3501, 3505-3520, 3540-3549, and 4412-4415 fused to one or more proteins or domains thereof having an activity of interest. In some embodiments, the CasX variant comprises any one of SEQ ID NOs: 247-337, 3498-3501, 3505-3520, 3540-3549, and 4412-4415 fused to one or more proteins or domains thereof having an activity of interest. In some embodiments, the CasX variant comprises any one of SEQ ID NOs: 3498-3501, 3505-3520, and 3540-3549 fused to one or more proteins or domains thereof having an activity of interest.
[0310] In some embodiments, a heterologous polypeptide (or heterologous amino acid, such as a cysteine residue or an unnatural amino acid) can be inserted at one or more positions within a CasX protein to generate a CasX fusion protein. In other embodiments, a cysteine residue can be inserted at one or more positions within a CasX protein, followed by conjugation of a heterologous polypeptide, as described below. In some alternative embodiments, a heterologous polypeptide or heterologous amino acid can be added to the N- or C-terminus of a reference or CasX variant protein. In other embodiments, a heterologous polypeptide or heterologous amino acid can be inserted internally within the sequence of a CasX protein.
[0311] In some embodiments, the reference CasX or variant fusion protein retains RNA guide sequence-specific target nucleic acid binding and cleavage activity. In some cases, the reference CasX or variant fusion protein has (retains) 50% or more of the activity (e.g., cleavage and / or binding activity) of the corresponding reference CasX or variant protein without the heterologous protein inserted. In some cases, the reference CasX or variant fusion protein retains at least about 60%, or at least about 70%, or at least about 80%, or at least about 90%, or at least about 92%, or at least about 95%, or at least about 98%, or about 100% of the activity (e.g., cleavage and / or binding activity) of the corresponding CasX protein without the heterologous protein inserted.
[0312] In some cases, the reference CasX or CasX variant fusion protein retains (has) target nucleic acid binding activity relative to the activity of a CasX protein without the heterologous amino acid or polypeptide inserted. In some cases, the reference CasX or CasX variant fusion protein retains at least about 60%, or at least about 70%, or at least about 80%, or at least about 90%, or at least about 92%, or at least about 95%, or at least about 98%, or about 100% of the binding activity of the corresponding CasX protein without the heterologous protein inserted.
[0313] In some cases, the reference CasX or CasX variant fusion protein retains (has) target nucleic acid binding and / or cleavage activity compared to the activity of the parent CasX protein without the heterologous amino acid or polypeptide insertion. For example, in some cases, the reference CasX or CasX variant fusion protein retains (has) 50% or more of the binding and / or cleavage activity of the corresponding parent CasX protein (the CasX protein without the insertion). For example, in some cases, the reference CasX or CasX variant fusion protein retains (has) 60% or more (70% or more, 80% or more, 90% or more, 92% or more, 95% or more, 98% or more, or 100%) of the binding and / or cleavage activity of the corresponding parent CasX protein (the CasX protein without the insertion). Methods for measuring the cleavage and / or binding activity of CasX proteins and / or CasX fusion proteins are known to those skilled in the art, and any convenient method can be used.
[0314] A variety of heterologous polypeptides are suitable for inclusion in the reference CasX or CasX variant fusion proteins of the present disclosure. In some cases, the fusion partner can modulate transcription of target DNA (e.g., inhibit transcription, increase transcription). For example, in some cases, the fusion partner is a protein (or a domain from a protein) that inhibits transcription (e.g., a protein that functions by a transcription repressor, recruiting a transcription inhibitor protein, modifying target DNA such as methylation, recruiting a DNA modifier, regulating histones associated with target DNA, recruiting histone modifiers such as those that modify histone acetylation and / or methylation, etc.). In some cases, the fusion partner is a protein (or a domain from a protein) that increases transcription (e.g., a protein that functions by a transcription activator, recruiting a transcription activator protein, modifying target DNA such as demethylation, recruiting a DNA modifier, regulating histones associated with target DNA, recruiting histone modifiers such as those that modify histone acetylation and / or methylation, etc.).
[0315] In some cases, the fusion partner has an enzymatic activity that modifies a target nucleic acid, e.g., a nuclease activity, a methyltransferase activity, a demethylase activity, a DNA repair activity, a DNA damaging activity, a deaminating activity, a dismutase activity, an alkylating activity, a depurinating activity, an oxidizing activity, a pyrimidine dimer forming activity, an integrase activity, a transposase activity, a recombinase activity, a polymerase activity, a ligase activity, a helicase activity, a photolyase activity, or a glycosylase activity.
[0316] In some cases, the fusion partner has an enzymatic activity that modifies a polypeptide (e.g., a histone) associated with the target nucleic acid, e.g., a methyltransferase activity, demethylase activity, acetyltransferase activity, deacetylase activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitinating activity, adenylating activity, deadenylating activity, sumoylating activity, desumoylating activity, ribosylation activity, deribosylation activity, myristoylating activity, or demyristoylating activity. In some embodiments, the CasX variant comprises any one of SEQ ID NOs: 247-337, 3301-3493, 3498-3501, 3505-3520, 3540-3549, and 4412-4415, and a polypeptide having methyltransferase activity, demethylase activity, acetyltransferase activity, deacetylase activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitinating activity, adenylating activity, deadenylating activity, sumoylating activity, desumoylating activity, ribosylation activity, deribosylation activity, myristoylating activity, or demyristoylating activity. In some embodiments, the CasX variant comprises any one of SEQ ID NOs: 247-337, 3498-3501, 3505-3520, 3540-3549, and 4412-4415, and a polypeptide having methyltransferase activity, demethylase activity, acetyltransferase activity, deacetylase activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitinating activity, adenylating activity, deadenylating activity, sumoylating activity, desumoylating activity, ribosylation activity, deribosylation activity, myristoylating activity, or demyristoylating activity.In some embodiments, the CasX variant comprises any one of SEQ ID NOs: 3498-3501, 3505-3520, and 3540-3549, and a polypeptide having methyltransferase activity, demethylase activity, acetyltransferase activity, deacetylase activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitinating activity, adenylating activity, deadenylating activity, sumoylating activity, desumoylating activity, ribosylation activity, deribosylation activity, myristoylating activity, or demyristoylating activity.
[0317] Examples of proteins (or fragments thereof) that can be used as suitable fusion partners to the reference CasX or CasX variants to increase transcription include transcription activators such as VP16, VP64, VP48, VP160, p65 subdomains (e.g., from NFkB), and the activation domain of EDLL, and / or transcription activator-like (TAL) activation domains (e.g., for activity in plants); histone lysine methyltransferases such as SET domain-containing 1A, histone lysine methyltransferase (SET1A), SET domain-containing 1B, histone lysine methyltransferase (SET1B), lysine methyltransferase 2A (MLL1)-5, ASCL1 (ASH1), achaete-scute family bHLH transcription factor 1 (ASH1), SET and MYND domain-containing 2 provided (SMYD2), nuclear receptor-binding SET domain protein 1 (NSD1), and the like. histone lysine demethylases, such as lysine demethylase 3A (JHDM2a) / lysine-specific demethylase 3B (JHDM2b), lysine demethylase 6A (UTX), and lysine demethylase 6B (JMJD3); lysine acetyltransferase 2A (GCN5), lysine acetyltransferase 2B (PCAF), CREB-binding protein (CBP), E1A-binding protein p30 (p300), and TATA-box-binding protein-associated factor 1 (TATA-box). histone acetyltransferases such as AF1), lysine acetyltransferase 5 (TIP60 / PLIP), lysine acetyltransferase 6A (MOZ / MYST3), lysine acetyltransferase 6B (MORF / MYST4), SRC proto-oncogene, non-receptor tyrosine kinase (SRC1), nuclear receptor coactivator 3 (ACTR), MYB-binding protein 1a (P160), and clock circadian regulator (CLOCK);and DNA demethylases such as Ten-Eleven Translocation (TET) dioxygenase 1 (TET1CD), tet methylcytosine dioxygenase 1 (TET1), demeter (DME), demeter-like 1 (DML1), demeter-like 2 (DML2), and protein ROS1 (ROS1);
[0318] Examples of proteins (or fragments thereof) that can be used as suitable fusion partners with the reference CasX or CasX variants to reduce transcription include transcription repressors such as Kruppel-associated box (KRAB or SKD); KOX1 repression domain; Mad mSIN3-interacting domain (SID); ERF repressor domain (ERD), SRDX repression domain (e.g., for repression in plants), etc.; histone lysine methyltransferases such as PR / SET domain-containing protein (Pr-SET)7 / 8, lysine methyltransferase 5B (SUV4-20H1), PR / SET domain 2 (RIZ1); lysine demethylase 4A (JMJD2A / JHDM3A), lysine demethylase 4B (JMJD2B), lysine demethylase 4C (JMJD2C / GASC1), lysine demethylase 5B (JMJD2C / GASC1), etc. histone lysine demethylases such as lysine demethylase 4D (JMJD2D), lysine demethylase 5A (JARID1A / RBP2), lysine demethylase 5B (JARID1B / PLU-1), lysine demethylase 5C (JARID1C / SMCX), and lysine demethylase 5D (JARID1D / SMCY); histone lysine deacetylases such as histone deacetylase 1 (HDAC1), HDAC2, HDAC3, HDAC8, HDAC4, HDAC5, HDAC7, HDAC9, sirtuin 1 (SIRT1), SIRT2, and HDAC11; HhaI These include, but are not limited to, DNA methylases such as DNA m5c-methyltransferase (M.HhaI), DNA methyltransferase 1 (DNMT1), DNA methyltransferase 3a (DNMT3a), DNA methyltransferase 3b (DNMT3b), methyltransferase 1 (MET1), S-adenosyl-L-methionine-dependent methyltransferase superfamily protein (DRM3) (plants), DNA cytosine methyltransferase MET2a (ZMET2), chromomethylase 1 (CMT1), chromomethylase 2 (CMT2) (plants); and peripheral mobilization elements such as Lamin A and Lamin B.
[0319] In some cases, the fusion partner to the reference CasX or CasX variant has an enzymatic activity that modifies a target nucleic acid (e.g., ssRNA, dsRNA, ssDNA, dsDNA). Examples of enzymatic activities that can be provided by the fusion partner include nuclease activity, such as that provided by a restriction enzyme (e.g., FokI nuclease), methyltransferase activity, such as that provided by a methyltransferase (e.g., Hhal DNA m5c-methyltransferase (M.Hhal), DNA methyltransferase 1 (DNMT1), DNA methyltransferase 3a (DNMT3a), DNA methyltransferase 3b (DNMT3b), METI, DRM3 (plants), ZMET2, CMT1, CMT2 (plants), etc.); demethylase activity, such as that provided by a demethylase (e.g., ten-eleven translocation (TET) dioxygenase 1 (TET1 CD), TET1, DME, DML1, DML2, ROS1, etc.), DNA repair activity, DNA damage activity, deaminases (e.g., cytosine deaminase enzymes, e.g., rat apolipoprotein B These include, but are not limited to, deaminating activity such as that provided by mRNA editing enzymes, APOBEC proteins such as catalytic polypeptide 1 {APOBEC1}, dismutase activity, alkylating activity, depurinating activity, oxidizing activity, pyrimidine dimer-forming activity, integrase activity such as that provided by integrases and / or resolvases (e.g., Gin invertases such as hyperactive mutants of Gin invertase, GinH106Y; human immunodeficiency virus type 1 integrase (IN); Tn3 resolvase, etc.), transposase activity, recombinase activity such as that provided by recombinases (e.g., the catalytic domain of Gin recombinase), polymerase activity, ligase activity, helicase activity, photolyase activity, and glycosylase activity.
[0320] In some cases, a reference CasX or CasX variant protein of the present disclosure is fused to a polypeptide selected from a domain for increasing transcription (e.g., a VP16 domain, a VP64 domain), a domain for decreasing transcription (e.g., a KRAB domain from a Kox1 protein), a core catalytic domain of a histone acetyltransferase (e.g., histone acetyltransferase p300), a protein / domain that provides a detectable signal (e.g., a fluorescent protein such as GFP), a nuclease domain (e.g., a Fokl nuclease), and a base editor (described further below).
[0321] In some embodiments, the CasX variant comprises any one of SEQ ID NOs: 247-337, 3301-3493, 3498-3501, 3505-3520, 3540-3549, and 4412-4415 fused to a polypeptide selected from the group consisting of a domain for reducing transcription, a domain having enzymatic activity, a core catalytic domain of a histone acetyltransferase, a protein / domain that provides a detectable signal, a nuclease domain, and a base editor. In some embodiments, the CasX variant comprises any one of SEQ ID NOs: 247-337, 3498-3501, 3505-3520, 3540-3549, and 4412-4415 fused to a polypeptide selected from the group consisting of a domain for reducing transcription, a domain having enzymatic activity, a core catalytic domain of a histone acetyltransferase, a protein / domain that provides a detectable signal, a nuclease domain, and a base editor. In some embodiments, the CasX variant comprises any one of SEQ ID NOs: 3498-3501, 3505-3520, and 3540-3549 fused to a polypeptide selected from the group consisting of a domain for reducing transcription, a domain having enzymatic activity, a core catalytic domain of a histone acetyltransferase, a protein / domain that provides a detectable signal, a nuclease domain, and a base editor.
[0322] In some cases, a reference CasX protein or CasX variant of...
Claims
1. A variant of a reference CasX protein (CasX variant), a. the CasX variant comprises at least one modification in the reference CasX protein; b. A variant of a reference CasX protein, wherein said CasX variant exhibits at least one improved characteristic compared to said reference CasX protein.
2. 2. The CasX variant of claim 1, wherein the improved characteristic of the CasX variant is selected from the group consisting of improved folding of the CasX variant, improved binding affinity for a guide nucleic acid (gNA), improved binding affinity for a target DNA, improved ability to utilize a broader spectrum of one or more PAM sequences, including ATC, CTC, GTC, or TTC, in editing the target DNA, improved unwinding of the target DNA, increased editing activity, improved editing efficiency, improved editing specificity, increased nuclease activity, increased target strand loading for double-stranded cleavage, reduced target strand loading for single-stranded nicking, reduced off-target cleavage, improved binding of non-target DNA strands, improved protein stability, improved protein solubility, improved protein:gNA complex (RNP) stability, improved protein:gNA complex solubility, improved protein yield, improved protein expression, improved fusion characteristics, or a combination thereof.
3. The at least one modification is a. at least one amino acid substitution in a domain of said CasX variant; b. At least one amino acid deletion in a domain of the CasX variant; c. at least one amino acid insertion in the domain of the CasX variant; d. Replacing all or part of a domain from a different CasX; e. Deletion of all or part of a domain of said CasX variant, or f. Any combination of (a) to (e).
4. The CasX variant of any one of claims 1 to 3, wherein the reference CasX protein comprises the sequence of SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO:
3.
5. The at least one modification is a. a non-target strand binding (NTSB) domain; b. a target strand loading (TSL) domain; c. helical I domain, d. Helical II domain, e. an oligonucleotide binding domain (OBD), or f. The CasX variant of any one of claims 1 to 4, wherein the CasX variant is in a domain selected from the group consisting of the RuvC DNA cleavage domain and the RuvC DNA cleavage domain.
6. The CasX variant of claim 5, comprising at least one modification in the NTSB domain.
7. The CasX variant of claim 5, comprising at least one modification in the TSL domain.
8. 8. The CasX variant of claim 7, wherein the at least one modification in the TSL domain comprises an amino acid substitution of one or more of amino acids Y857, S890, or S932 of SEQ ID NO:
2.
9. The CasX variant of claim 5, comprising at least one modification in the helical I domain.
10. 10. The CasX variant of claim 9, wherein the at least one modification in the helical I domain comprises an amino acid substitution of one or more of amino acids S219, L249, E259, Q252, E292, L307, or D318 of SEQ ID NO:
2.
11. A CasX variant according to any one of claims 5 to 10, comprising at least one modification in the helical II domain.
12. 12. The CasX variant of claim 11, wherein the at least one modification in the helical II domain comprises an amino acid substitution of one or more of amino acids D361, L379, E385, E386, D387, F399, L404, R458, C477, or D489 of SEQ ID NO:
2.
13. The CasX variant of claim 5, comprising at least one modification in the OBD domain.
14. 14. The CasX variant of claim 13, wherein the at least one modification in the OBD comprises an amino acid substitution of one or more of amino acids F536, E552, T620, or I658 of SEQ ID NO:
2.
15. 6. The CasX variant of claim 5, comprising at least one modification in the RuvC DNA cleavage domain.
16. 16. The CasX variant of Claim 15, wherein the at least one modification in the RuvC DNA cleavage domain comprises an amino acid substitution of one or more of amino acids K682, G695, A708, V711, D732, A739, D733, L742, V747, F755, M771, M779, W782, A788, G791, L792, P793, Y797, M799, Q804, S819, or Y857 of SEQ ID NO: 2, or a deletion of amino acid P793.
17. The CasX variant of any one of claims 5 to 16, wherein the modification results in an increased ability to edit the target DNA.
18. The CasX variant of any one of claims 1 to 17, wherein the CasX variant is capable of forming a ribonucleoprotein complex (RNP) with a guide nucleic acid (gNA).
19. The at least one modification is a. a substitution of 1 to 100 consecutive or non-consecutive amino acids in the CasX variant; b. a deletion of 1 to 100 consecutive or non-consecutive amino acids in the CasX variant; c. an insertion of 1 to 100 consecutive or non-consecutive amino acids in said CasX; or d. Any combination of (a) to (c).
20. The at least one modification is a. a substitution of 5 to 10 consecutive or non-consecutive amino acids in the CasX variant; b. a deletion of 1 to 5 consecutive or non-consecutive amino acids in said CasX variant; c. an insertion of 1 to 5 consecutive or non-consecutive amino acids in said CasX; or 20. The CasX variant of claim 19, comprising any combination of (a) to (c).
21. The CasX variant of any one of claims 1 to 20, wherein the CasX variant comprises two or more modifications in one domain.
22. The CasX variant of any one of claims 1 to 21, wherein the CasX variant comprises modifications in two or more domains.
23. 21. The CasX variant of any one of claims 1 to 20, comprising at least one modification of a region of non-adjacent amino acid residues of the CasX variant that forms a channel through which gNA:target DNA complex formation with the CasX variant occurs.
24. 21. The CasX variant of any one of claims 1 to 20, comprising at least one modification of a region of non-adjacent amino acid residues of the CasX variant that forms a binding interface with the gNA.
25. The CasX variant of any one of claims 1 to 20, comprising at least one modification of a region of non-adjacent amino acid residues of the CasX variant that forms a channel that binds to the non-target strand DNA.
26. 21. The CasX variant of any one of claims 1 to 20, comprising at least one modification of a region of non-adjacent amino acid residues of the CasX variant that forms a binding interface with a protospacer adjacent motif (PAM) of the target DNA.
27. 21. The CasX variant of any one of claims 1 to 20, comprising at least one modification of a region of non-adjacent surface-exposed amino acid residues of the CasX variant.
28. A CasX variant according to any one of claims 1 to 20, comprising at least one modification of a region of non-adjacent amino acid residues that form a core through hydrophobic packing in the domain of the CasX variant.
29. The CasX variant of any one of claims 23 to 28, wherein the modification is one or more of a deletion, insertion, or substitution of one or more amino acids in the region.
30. The CasX variant according to any one of claims 23 to 28, wherein 2 to 15 amino acid residues in said region of said CasX variant are substituted with charged amino acids.
31. The CasX variant according to any one of claims 23 to 28, wherein 2 to 15 amino acid residues in the region of the CasX variant are substituted with polar amino acids.
32. The CasX variant of any one of claims 23 to 28, wherein 2 to 15 amino acid residues in the region of the CasX variant are substituted with amino acids that stack with DNA or RNA bases.
33. the at least one modification compared to the reference CasX sequence of SEQ ID NO: 2 is a. L379R amino acid substitution; b. A708K amino acid substitution; c. T620P amino acid substitution; d. E385P amino acid substitution; e. Y857R amino acid substitution; f. I658V amino acid substitution; g. F399L amino acid substitution; h. an amino acid substitution of Q252K; i. an amino acid substitution of L404K, and j. an amino acid deletion of P793.
34. 6. The CasX variant of any one of claims 1 to 5, wherein the CasX variant has a sequence selected from the group consisting of the sequences of Tables 3, 8, 9, 10, and 12, or a sequence having at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, or at least about 95%, or at least about 96%, or at least about 97%, or at least about 98%, or at least about 99% sequence identity thereto.
35. The CasX variant of any one of claims 1 to 5, comprising a sequence selected from the group consisting of SEQ ID NOs: 258-327, 3508-3520, and 4412-4415.
36. 6. The CasX variant of any one of claims 1 to 5, further comprising a substitution of the NTSB and / or helical 1b domain from a different CasX.
37. 37. The CasX variant of claim 36, wherein the substituted NTSB and / or the helical lb domain is derived from the reference CasX of SEQ ID NO:
1.
38. 38. The CasX variant of any one of claims 1 to 37, further comprising one or more nuclear localization signals (NLS).
39. The one or more NLSs may be PKKKRKV (SEQ ID NO: 352), KRPAATKKAGQAKKKK (SEQ ID NO: 353), PAAKRVKLD (SEQ ID NO: 354), RQRRNELKRSP (SEQ ID NO: 355), NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY (SEQ ID NO: 356), RMRIZFKNKGKDTAELRRRRRVEVSVELRKAKKDEQILKRRNV (SEQ ID NO: 357), VSRKRPRP (SEQ ID NO: 358), PPKKARED (SEQ ID NO: 359), PQPKK KPL (SEQ ID NO: 360), SALIKKKKKMAP (SEQ ID NO: 361), DRLRR (SEQ ID NO: 362), PKQKKRK (SEQ ID NO: 363), RKLKKKIKKL (SEQ ID NO: 364), REKKKFLKRR (SEQ ID NO: 365), KRKGDEVDGVDEVAKKKSKK (SEQ ID NO: 366), RKCLQAGMNLEARKTKK (SEQ ID NO: 367), PRPRKIPR (SEQ ID NO: 368), PPRKKRTVV (SEQ ID NO: 369), NLSKKKKRKREK (SEQ ID NO: 370), RRPSRPFRKP (SEQ ID NO: 371), No. 371), KRPRSPSS (SEQ ID NO: 372), KRGINDRNFWRGENERKTR (SEQ ID NO: 373), PRPPKMARYDN (SEQ ID NO: 374), KRSFSKAF (SEQ ID NO: 375), KLKIKRPVK (SEQ ID NO: 376), PKKKRKVPPPPAAKRVKLD (SEQ ID NO: 377), PKTRRRPRRRSQRKRPPT (SEQ ID NO: 378), SRRRKANPTKLSENAKKLAKEVEN (SEQ ID NO: 379), KTRRRPRRSQRKRPPT (SEQ ID NO: 380), RRKKRRPR 39. The CasX variant of claim 38, selected from the group of sequences consisting of RKKRR (SEQ ID NO:381), PKKKKSRKPKKKSRK (SEQ ID NO:382), HKKKHPDASVNFSEFSK (SEQ ID NO:383), QRPGPYDRPQRPGPYDRP (SEQ ID NO:384), LSPSLSPLLSPSLSPL (SEQ ID NO:385), RGKGGKGLGKGGAKRHRK (SEQ ID NO:386), PKRGRGRPKRGRGR (SEQ ID NO:387), and PKKKRKVPPPPKKKRKV (SEQ ID NO:389).
40. 39. The CasX variant of claim 38, comprising the sequence of any one of SEQ ID NOs: 3540-3549.
41. 40. The CasX variant of claim 38 or 39, wherein the one or more NLSs are located at or near the C-terminus of the CasX protein.
42. 40. The CasX variant of claim 38 or 39, wherein the one or more NLSs are located at or near the N-terminus of the CasX protein.
43. 40. The CasX variant of claim 38 or 39, comprising at least two NLSs, said at least two NLSs being located at or near the N-terminus and at or near the C-terminus of said CasX protein.
44. 44. The CasX variant of any one of claims 2 to 43, wherein one or more of the improved characteristics of the CasX variant are improved by at least about 1.1 to about 100 fold or more compared to the reference CasX protein of SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO:
3.
45. 44. The CasX variant of any one of claims 2 to 43, wherein one or more of the improved characteristics of the CasX variant are improved by at least about 1.1-fold, at least about 2-fold, at least about 10-fold, at least about 100-fold or more compared to the reference CasX protein of SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO:
3.
46. 46. The CasX variant of any one of claims 2 to 45, wherein the improved characteristics comprise editing efficiency, and wherein the CasX variant comprises a 1.1 to 100 fold improvement in editing efficiency compared to the reference CasX protein of SEQ ID NO:
2.
47. 47. The CasX variant of any one of claims 1 to 46, wherein the RNP comprising the CasX variant exhibits greater editing efficiency and / or binding of a target sequence in the target DNA when any one of the PAM sequences TTC, ATC, GTC, or CTC is located one nucleotide 5' to the non-target strand of the protospacer having identity to a targeting sequence of the gNA in a cellular assay system compared to the editing efficiency and / or binding of an RNP comprising a reference CasX protein in an equivalent assay system.
48. 48. The CasX variant of claim 47, wherein the PAM sequence is TTC.
49. 48. The CasX variant of claim 47, wherein the PAM sequence is ATC.
50. 48. The CasX variant of claim 47, wherein the PAM sequence is CTC.
51. 48. The CasX variant of claim 47, wherein the PAM sequence is GTC.
52. 48. The CasX variant of claim 47, wherein the improved editing efficiency and / or binding of the RNP comprising the CasX variant to the target DNA is improved by at least about 1.1 to about 100-fold compared to the RNP comprising the reference CasX.
53. 53. The CasX variant of any one of claims 1 to 52, wherein the CasX variant comprises 400 to 2000 amino acids.
54. The CasX variant of any one of claims 1 to 53, wherein the CasX variant protein comprises a nuclease domain having nickase activity.
55. The CasX variant of any one of claims 1 to 53, wherein the CasX variant protein comprises a nuclease domain having double-strand cleavage activity.
56. 54. The CasX variant of any one of claims 1 to 53, wherein the CasX protein is a catalytically inactive CasX (dCasX) protein, and the dCasX and the gNA retain the ability to bind to the target DNA.
57. The dCasX is a. D672, and / or E769, and / or D935 corresponding to the CasX protein of SEQ ID NO: 1, or b) The CasX variant of claim 56, comprising a mutation at residues D659, and / or E756, and / or D922, corresponding to the CasX protein of SEQ ID NO:
2.
58. 58. The CasX variant of claim 57, wherein the mutation is a substitution of the residue with alanine.
59. 59. The CasX variant of any one of claims 1 to 58, wherein the CasX variant comprises a first domain from a first CasX protein and a second domain from a second CasX protein that is different from the first CasX protein.
60. 60. The CasX variant of claim 59, wherein the first domain is selected from the group consisting of the NTSB, TSL, helical I, helical II, OBD, and RuvC domains.
61. 60. The CasX variant of claim 59, wherein the second domain is selected from the group consisting of the NTSB, TSL, helical I, helical II, OBD, and RuvC domains.
62. 62. The CasX variant of any one of claims 59 to 61, wherein the first and second domains are not the same domain.
63. 63. The CasX variant of any one of claims 59 to 62, wherein the first domain comprises a portion of a sequence selected from the group consisting of amino acids 1-56, 57-100, 101-191, 192-332, 333-509, 510-660, 661-824, 825-934, and 935-986 of SEQ ID NO:1, and the second domain comprises a portion of a sequence selected from the group consisting of amino acids 1-58, 59-102, 103-192, 193-333, 334-501, 502-647, 648-812, 813-921, and 922-978 of SEQ ID NO:
2.
64. 64. The CasX variant of any one of claims 1 to 63, wherein the CasX variant is selected from the group consisting of CasX variants SEQ ID NO:328, SEQ ID NO:3540, SEQ ID NO:4413, SEQ ID NO:4414, SEQ ID NO:4415, SEQ ID NO:329, SEQ ID NO:3541, SEQ ID NO:330, SEQ ID NO:3542, SEQ ID NO:331, SEQ ID NO:3543, SEQ ID NO:332, SEQ ID NO:3544, SEQ ID NO:333, SEQ ID NO:3545, SEQ ID NO:334, SEQ ID NO:3546, SEQ ID NO:335, SEQ ID NO:3547, SEQ ID NO:336 and SEQ ID NO:3548.
65. 59. The CasX variant of any one of claims 1 to 58, wherein the CasX variant comprises at least one chimeric domain comprising a first portion from a first CasX protein and a second portion from a second CasX protein that is different from the first CasX protein.
66. 66. The CasX variant of claim 65, wherein the at least one chimeric domain is selected from the group consisting of the NTSB, TSL, helical I, helical II, OBD, and RuvC domains.
67. 67. The CasX variant of claim 65 or 66, wherein the first CasX protein comprises the sequence of SEQ ID NO: 1 and the second CasX protein comprises the sequence of SEQ ID NO:
2.
68. 67. The CasX variant of claim 66, wherein the at least one chimeric domain comprises a chimeric RuvC domain.
69. 69. The CasX variant of claim 68, wherein the chimeric RuvC domain comprises amino acids 661-824 of SEQ ID NO:1 and amino acids 922-978 of SEQ ID NO:
2.
70. 69. The CasX variant of claim 68, wherein the chimeric RuvC domain comprises amino acids 648-812 of SEQ ID NO:2 and amino acids 935-986 of SEQ ID NO:
1.
71. The CasX variant of any one of claims 1 to 5, comprising a sequence selected from the group consisting of SEQ ID NOs: 247-337, 3301-3493, 3498-3501, 3505-3520, 3540-3549 and 4412-4415.
72. 6. The CasX variant of any one of claims 1 to 5, comprising a sequence selected from the group consisting of SEQ ID NOs: 247-337, 3498-3501, 3505-3520, 3540-3549 and 4412-4415.
73. 6. The CasX variant of any one of claims 1 to 5, comprising a sequence selected from the group consisting of SEQ ID NOs: 3498-3501, 3505-3520, and 3540-3549.
74. 74. The CasX variant of any one of claims 1 to 73, comprising a heterologous protein or domain thereof fused to the CasX.
75. 75. The CasX variant of Claim 74, wherein the heterologous protein or domain thereof is a base editor.
76. 76. The CasX variant of Claim 75, wherein the base editor is an adenosine deaminase, cytosine deaminase, or guanine oxidase.
77. A reference guide nucleic acid scaffold variant (gNA variant) capable of binding to a reference CasX protein or CasX variant, a. the gNA variant comprises at least one modification compared to the reference guide nucleic acid scaffold sequence; b. A variant of a reference guide nucleic acid scaffold, wherein the gNA variant exhibits one or more improved characteristics compared to the reference guide nucleic acid scaffold.
78. 78. The gNA variant of claim 77, wherein the one or more improved characteristics are selected from the group consisting of improved stability, improved solubility, improved transcription of the gNA, improved resistance to nuclease activity, an increased folding rate of the gNA, reduced by-product formation during folding, increased productive folding, improved binding affinity for a CasX protein, improved binding affinity for a target DNA when complexed with the CasX protein, improved gene editing when complexed with the CasX protein, improved editing specificity when complexed with the CasX protein, and an improved ability to utilize a broader spectrum of one or more PAM sequences, including ATC, CTC, GTC, or TTC, in the editing of a target DNA when complexed with the CasX protein.
79. 79. The gNA variant of claim 77 or 78, wherein the reference guide scaffold comprises a sequence selected from the group consisting of SEQ ID NOs: 4 to 16.
80. The at least one modification is a. at least one nucleotide substitution in the region of the gNA variant; b. at least one nucleotide deletion in the region of the gNA variant; c. at least one nucleotide insertion in the region of the gNA variant; d. Substitution of all or part of the region of the gNA variant; e. a deletion of all or part of the region of said gNA variant, or f. any combination of (a) to (e).
81. 81. The gNA variant of claim 80, wherein the region of the gNA variant is selected from the group consisting of an extended stem loop, a scaffold stem loop, a triplex, and a pseudoknot.
82. 82. The gNA variant of claim 81, wherein the scaffold stem further comprises a bubble.
83. 83. The gNA variant of claim 81 or 82, wherein the scaffold further comprises a triplex loop region.
84. The gNA variant of any one of claims 81 to 83, wherein the scaffold further comprises a 5' unstructured region.
85. The at least one modification is a. substitution of 1 to 15 consecutive or non-consecutive nucleotides of said gNA variant in one or more regions; b. a deletion of 1 to 10 consecutive or non-consecutive nucleotides of said gNA variant in one or more regions; c. insertion of 1 to 10 consecutive or non-consecutive nucleotides of said gNA variant in one or more regions; d. Replacement of the scaffold stem-loop or the extended stem-loop with an RNA stem-loop sequence from a heterologous RNA source having proximal 5' and 3' ends, or e. Any combination of (a) to (d).
86. 86. The gNA variant of any one of claims 77 to 85, comprising an extended stem-loop region comprising at least 10, at least 100, at least 500, at least 1000, or at least 10,000 nucleotides.
87. 86. The gNA variant of claim 85, wherein the heterologous RNA stem-loop sequence increases the stability of the gNA.
88. 88. The gNA variant of claim 87, wherein the heterologous RNA stem loop is capable of binding to a protein, an RNA structure, a DNA sequence, or a small molecule.
89. The gNA variant of claim 87 or 88, wherein the heterologous RNA stem-loop sequence is selected from MS2, Qβ, U1 hairpin II, Uvsx, or PP7 stem-loop.
90. said at least one modification compared to said reference guide scaffold of SEQ ID NO: 5: a. a C18G substitution in the triplex loop; b. a G55 insertion in the stem bubble; c. U1 deletion, d. A modification of the extended stem loop, i. The 6 nt loop and 13 loop-proximal base pairs are replaced by a Uvsx hairpin; ii. The gNA variant of any one of claims 85-89, wherein the modification is selected from one or more of the following: deletion of A99 and substitution of G64U, which results in a perfectly base-pairing loop distal base.
91. The gNA variant of any one of claims 77 to 90, wherein the gNA variant comprises two or more modifications in one region.
92. 92. The gNA variant of any one of claims 77 to 91, wherein the gNA variant comprises modifications in two or more regions.
93. The gNA variant of any one of claims 77 to 92, wherein the gNA variant further comprises a targeting sequence, the targeting sequence being complementary to the target DNA sequence.
94. 94. The gNA variant of claim 93, wherein the targeting sequence has 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34 or 35 nucleotides.
95. 95. The gNA variant of claim 93 or 94, wherein the targeting sequence has 20 nucleotides.
96. The gNA variant of any one of claims 93 to 95, wherein the gNA is a single guide gNA comprising the scaffold sequence linked to the targeting sequence.
97. 97. The gNA variant of any one of claims 77-96, wherein one or more of the improved features of the CasX variant are improved by at least about 1.1 to about 100 fold or more compared to the reference gNA of SEQ ID NO: 4 or SEQ ID NO:
5.
98. 97. The gNA variant of any one of claims 77 to 96, wherein one or more of the improved features of the gNA variant are improved by at least about 1.1 fold, at least about 2 fold, at least about 10 fold, or at least about 100 fold or more compared to the reference gNA of SEQ ID NO: 4 or SEQ ID NO:
5.
99. 99. The gNA variant of any one of claims 77 to 98, comprising a scaffold region having at least 60% sequence identity to SEQ ID NO: 4 or SEQ ID NO: 5, excluding the extended stem region.
100. 99. The gNA variant of any one of claims 77 to 98, comprising a scaffold stem loop having at least 60% sequence identity to SEQ ID NO:
14.
101. The gNA variant of claim 100, comprising a scaffold stem-loop sequence of CCAGCGACUAUGUCGUAGUGG (SEQ ID NO: 245).
102. 102. The gNA variant of any one of claims 77 to 101, wherein the scaffold of the gNA variant sequence has at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, or at least 70%, at least 80%, at least 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to the sequence of SEQ ID NO:4 or SEQ ID NO:
5.
103. The gNA variant of any one of claims 77 to 101, wherein the scaffold of the gNA variant sequence comprises a sequence selected from the group of sequences of SEQ ID NOs: 2101 to 2280, or a sequence having at least about 80%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% identity thereto.
104. The gNA variant of claim 103, wherein the scaffold of the gNA variant sequence consists of a sequence selected from the group of sequences of SEQ ID NOs: 2101 to 2280.
105. 105. The gNA variant of any one of claims 77 to 104, further comprising one or more ribozymes.
106. The gNA variant of claim 105, wherein the one or more ribozymes are independently fused to the termini of the gNA variant.
107. 107. The gNA variant of claim 105 or 106, wherein at least one of the one or more ribozymes is a hepatitis delta virus (HDV) ribozyme, a hammerhead ribozyme, a pistol ribozyme, a hatchet ribozyme, or a tobacco ringspot virus (TRSV) ribozyme.
108. A gNA variant according to any one of claims 77 to 107, further comprising a protein-binding motif.
109. The gNA variant of any one of claims 77 to 108, further comprising a thermostable stem loop.
110. The gNA variant of any one of claims 77 to 109, wherein the gNA is chemically modified.
111. 111. The gNA variant of any one of claims 77 to 110, wherein the gNA comprises a first region from a first gNA and a second region from a second gNA that is different from the first gNA.
112. The gNA variant of claim 111, wherein the first region is selected from the group consisting of a triplex region, a scaffold stem loop, and an extended stem loop.
113. The gNA variant of claim 111 or 112, wherein the second region is selected from the group consisting of a triplex region, a scaffold stem loop, and an extended stem loop.
114. A gNA variant according to any one of claims 111 to 113, wherein the first and second regions are not the same region.
115. A gNA variant according to any one of claims 111 to 113, wherein the first gNA comprises the sequence of SEQ ID NO: 4 and the second gNA comprises the sequence of SEQ ID NO:
5.
116. 116. A gNA variant according to any one of claims 77 to 115, comprising at least one chimeric region comprising a first portion from a first gNA and a second portion from a second gNA.
117. The gNA variant of claim 116, wherein the at least one chimeric region is selected from the group consisting of a triplex region, a scaffold stem loop, and an extended stem loop.
118. 78. The gNA variant of claim 77, comprising the sequence of any one of SEQ ID NOs: 2101-2280.
119. 78. The gNA variant of claim 77, comprising the sequence of any one of SEQ ID NOs: 2236, 2237, 2238, 2241, 2244, 2248, 2249, or 2259-2280.
120. A gene editing pair comprising a CasX protein and a first gNA.
121. 121. The gene editing pair of claim 120, wherein the CasX and the gNA can associate together in a ribonucleoprotein complex (RNP).
122. 121. The gene editing pair of claim 120, wherein the CasX and the gNA associate together in a ribonucleoprotein complex (RNP).
123. The first gNA is a. a gNA variant according to any one of claims 93 to 119, or b. The gene editing pair of any one of Claims 120-122, comprising a reference guide nucleic acid of SEQ ID NO: 4 or 5 and a targeting sequence, wherein the targeting sequence is complementary to the target DNA.
124. The CasX is a. a CasX variant according to any one of claims 1 to 76, or b. The gene editing pair of any one of claims 120 to 123, comprising a reference CasX protein of SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO:
3.
125. The first gNA is a. a gNA variant according to any one of claims 93 to 119, and b. A gene editing pair according to any one of claims 120 to 124, comprising a CasX variant according to any one of claims 1 to 76.
126. 126. The gene editing pair of Claim 125, wherein the gene editing pair of the CasX variant and the gNA variant has one or more improved characteristics compared to a gene editing pair comprising a reference CasX protein of SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 3, and a reference guide nucleic acid of SEQ ID NO: 4 or 5.
127. 127. The gene editing pair of claim 126, wherein the one or more improved characteristics comprise improved CasX:gNA (RNP) complex stability, improved binding affinity between the CasX and gNA, improved kinetics of RNP complex formation, a higher percentage of cleavage-competent RNPs, improved RNP binding affinity for target DNA, ability to utilize an increased spectrum of PAM sequences, improved unwinding of the target DNA, increased editing activity, improved editing efficiency, improved editing specificity, increased nuclease activity, increased target strand loading for double-stranded cleavage, reduced target strand loading for single-stranded nicking, reduced off-target cleavage, improved binding of the non-target strand of DNA, or improved resistance to nuclease activity.
128. The gene editing pair of Claims 126 or 127, wherein at least one or more of the improved characteristics is improved by at least about 1.1 to about 100 fold or more compared to the gene editing pair of the reference CasX protein and the reference guide nucleic acid.
129. The gene editing pair of Claims 126 or 127, wherein one or more of the improved features of the CasX variant are improved by at least about 1.1, at least about 2, at least about 10, or at least about 100-fold or more compared to the gene editing pair of the reference CasX protein and the reference guide nucleic acid.
130. The gene editing pair of claim 126 or 127, wherein the improved characteristics comprise a 4-9 fold increase in editing activity compared to a reference editing pair of SEQ ID NO: 2 and SEQ ID NO:
5.
131. 131. The gene editing pair of claim 130, comprising a CasX selected from any one of SEQ ID NO:270, SEQ ID NO:292, SEQ ID NO:311, SEQ ID NO:333, SEQ ID NO:336, SEQ ID NOs:3498-3501, SEQ ID NOs:3505-3520, and SEQ ID NOs:3540-3549, and a gNA selected from any one of SEQ ID NOs:2104, 2106, or 2238.
132. a. a second gene editing pair comprising the CasX variant of any one of claims 1 to 76 or the reference CasX protein of any one of SEQ ID NOs: 1 to 3; and b. A composition comprising the gene editing pair of any one of claims 120-131, further comprising a second gNA variant or second reference guide nucleic acid of any one of claims 77-119, wherein the second gNA variant or second reference guide nucleic acid has a targeting sequence that is complementary to a different or overlapping portion of the target DNA compared to the targeting sequence of the first gNA.
133. 133. The gene editing pair of any one of claims 120-132, wherein the RNPs of the CasX variant and the gNA variant have a higher percentage of cleavage-competent RNPs compared to the RNPs of a reference CasX protein and a reference guide nucleic acid.
134. The gene editing pair of any one of claims 120 to 133, wherein the RNP is capable of binding to and cleaving target DNA.
135. The gene editing pair of any one of claims 120 to 132, wherein the RNP is capable of binding to target DNA but is not capable of cleaving the target DNA.
136. The gene editing pair of any one of claims 120 to 132, wherein the RNP is capable of binding to target DNA and generating one or more single-stranded nicks in the target DNA.
137. 137. A method of editing target DNA, comprising contacting the target DNA with a gene editing pair of any one of claims 120-136, wherein the contacting results in editing or modification of the target DNA.
138. 138. The method of claim 137, comprising contacting the target DNA with multiple gNAs comprising targeting sequences complementary to different or overlapping regions of the target DNA.
139. 139. The method of Claim 137 or 138, wherein the contacting by the gene editing pair comprises binding the target DNA, resulting in introducing a mutation, insertion, or deletion into the target DNA.
140. 139. The method of Claim 137 or 138, wherein said contacting introduces one or more single-stranded breaks in said target DNA and said editing comprises introducing a mutation, insertion, or deletion in said target DNA.
141. 139. The method of Claim 137 or 138, wherein said contacting comprises introducing one or more double-stranded breaks in said target DNA, and said editing comprises introducing a mutation, insertion, or deletion in said target DNA.
142. 142. The method of Claim 140 or 141, further comprising contacting the target DNA with a nucleotide sequence of a donor template nucleic acid, wherein the donor template comprises a nucleotide sequence having homology to the target DNA.
143. 143. The method of claim 142, wherein the donor template comprises homologous arms at the 5' and 3' ends of the donor template.
144. 144. The method of claim 142 or 143, wherein the donor template is inserted into the target DNA at the break site by homology directed repair.
145. 144. The method of Claim 142 or 143, wherein the donor template is inserted into the target DNA at the cleavage site by non-homologous end joining (NHEJ) or microhomology end joining (MMEJ).
146. 145. The method of any one of claims 137 to 144, wherein the editing is performed in vitro outside a cell.
147. 145. The method of any one of claims 137 to 144, wherein the editing is performed in vitro inside a cell.
148. 145. The method of any one of claims 137 to 144, wherein the editing is performed in vivo inside a cell.
149. 149. The method of claim 147 or 148, wherein the cell is a eukaryotic cell.
150. 150. The method of claim 149, wherein the eukaryotic cell is selected from the group consisting of a plant cell, a fungal cell, a protist cell, a mammalian cell, a reptilian cell, an insect cell, an avian cell, a fish cell, a parasite cell, an arthropod cell, an invertebrate cell, a vertebrate cell, a rodent cell, a mouse cell, a rat cell, a primate cell, and a non-human primate cell.
151. 150. The method of claim 149, wherein the eukaryotic cell is a human cell.
152. 152. The method of claim 151, wherein the cell is an embryonic stem cell, an induced pluripotent stem cell, a germ cell, a fibroblast, an oligodendrocyte, a glial cell, a hematopoietic stem cell, a neuronal progenitor cell, a neuron, a muscle cell, a bone cell, a hepatocyte, a pancreatic cell, a retinal cell, a cancer cell, a T cell, a B cell, a NK cell, a fetal cardiomyocyte, a myofibroblast, a mesenchymal stem cell, an autologous expanded cardiomyocyte, an adipocyte, a totipotent cell, a multipotent cell, a blood stem cell, a myoblast, an adult stem cell, a bone marrow cell, a mesenchymal cell, a parenchymal cell, an epithelial cell, an endothelial cell, a mesothelial cell, a fibroblast, an osteoblast, a chondrocyte, an exogenous cell, an endogenous cell, a stem cell, a hematopoietic stem cell, a bone marrow-derived progenitor cell, a cardiomyocyte, a skeletal cell, a fetal cell, an undifferentiated cell, a multipotent progenitor cell, a unipotent progenitor cell, a monocyte, a cardiac myoblast, a skeletal myoblast, a macrophage, a capillary endothelial cell, a xenogeneic cell, an allogeneic cell, or a postnatal stem cell.
153. 153. The method of claim 151 or 152, wherein the cell is in a subject.
154. 154. The method of Claim 153, wherein the editing is performed in said subject having a mutation in an allele of a gene, said mutation causing a disease or disorder in said subject.
155. 155. The method of Claim 154, wherein said editing changes said mutation to a wild-type allele of said gene.
156. 155. The method of Claim 154, wherein said editing knocks down or knocks out an allele of a gene that causes a disease or disorder in said subject.
157. 152. The method of Claim 151, wherein the editing is performed in vitro inside the cells prior to introducing the cells into a subject.
158. 158. The method of claim 157, wherein the cells are autologous or allogeneic.
159. 152. The method of any one of claims 147-151, wherein greater editing of a target sequence in the target DNA is achieved in a cellular assay system comprising an RNP comprising the CasX variant when any one of the PAM sequences TTC, ATC, GTC, or CTC is located one nucleotide 5' to the non-target strand of a protospacer having identity to the targeting sequence of the gNA in a cellular assay system compared to the editing efficiency of an RNP comprising a reference CasX protein in a comparable assay system.
160. 160. The method of any one of Claims 149-159, wherein the method comprises contacting the eukaryotic cell with a vector encoding or comprising the CasX protein and the gNA, and optionally further comprising the donor template.
161. 161. The method of claim 160, wherein the vector is an adeno-associated virus (AAV) vector.
162. The method of claim 161, wherein the AAV is AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV-Rh74, or AAVRh10.
163. 161. The method of claim 160, wherein the vector is a lentiviral vector.
164. 161. The method of claim 160, wherein the vector is a non-viral particle.
165. 161. The method of claim 160, wherein the vector is a virus-like particle (VLP).
166. 165. The method of any one of claims 160 to 164, wherein the vector is administered to a subject in need thereof using a therapeutically effective dose.
167. 165. The method of claim 164, wherein the subject is selected from the group consisting of a mouse, a rat, a pig, and a non-human primate.
168. 167. The method of claim 166, wherein the subject is a human.
169. 169. The method of any one of claims 166-168, wherein the vector is administered at a dose of at least about 1 x 10 vector genome (vg), at least about 1 x 10 vg, at least about 1 x 10 vg, at least about 1 x 10 vg, at least about 1 x 10 vg, at least about 1 x 10 vg, at least about 1 x 10 vg, at least about 1 x 10 vg, or at least about 1 x 10 vg.
170. 170. The method of any one of claims 166-169, wherein the vector is administered by a route of administration selected from the group consisting of intraparenchymal, intravenous, intraarterial, intracerebroventricular, intracisternal, intrathecal, intracranial, and intraperitoneal routes.
171. 148. The method of claim 147, wherein the cell is a prokaryotic cell.
172. A cell comprising target DNA edited by a gene editing pair or composition described in any one of claims 120 to 136.
173. A cell edited by the method of any one of claims 137 to 165.
174. 174. The cell of claim 172 or 173, wherein the cell is a prokaryotic cell.
175. 174. The cell of claim 172 or 173, wherein the cell is a eukaryotic cell.
176. 176. The cell of claim 175, wherein the eukaryotic cell is selected from the group consisting of a plant cell, a fungal cell, a protist cell, a mammalian cell, a reptilian cell, an insect cell, an avian cell, a fish cell, a parasite cell, an arthropod cell, an invertebrate cell, a vertebrate cell, a rodent cell, a mouse cell, a rat cell, a primate cell, and a non-human primate cell.
177. The cell of claim 175, wherein the eukaryotic cell is a human cell.
178. A polynucleotide encoding a CasX variant according to any one of claims 1 to 76.
179. A polynucleotide encoding the gNA variant of any one of claims 77 to 119.
180. 180. A vector comprising the polynucleotide of claim 178 or 179.
181. 120. A vector comprising encoding a CasX variant of any one of claims 1 to 76 and a gNA variant of any one of claims 77 to 119.
182. 181. The vector of claim 180, wherein the vector is an adeno-associated virus (AAV) vector.
183. The vector of claim 182, wherein the AAV is AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV-Rh74, or AAVRh10.
184. The vector of claim 180, wherein the vector is a lentiviral vector.
185. The vector of claim 180, wherein the vector is a virus-like particle (VLP).
186. The vector of claim 180, wherein the vector is a non-viral particle.
187. A cell comprising the polynucleotide of claim 178 or the vector of any one of claims 180 to 186.
188. A composition comprising a CasX variant according to any one of claims 1 to 76.
189. a) a gNA variant according to any one of claims 77 to 119, or The composition of claim 188, further comprising: b) the reference guide scaffold and targeting sequence of SEQ ID NO: 4 or 5.
190. 190. The composition of claim 188 or 189, wherein the CasX protein and the gNA associate together in a ribonucleoprotein complex (RNP).
191. 191. The composition of any one of claims 188-190, further comprising a donor template nucleic acid, wherein the donor template comprises a nucleotide sequence having homology to a target DNA.
192. 192. The composition of any one of claims 188-191, further comprising a buffer, a nuclease inhibitor, a protease inhibitor, a liposome, a therapeutic agent, a label, a labeled visualization reagent, or any combination of the foregoing.
193. A composition comprising a gNA variant according to any one of claims 77 to 119.
194. 194. The composition of claim 193, further comprising the CasX variant of any one of claims 1 to 76, or the CasX protein of SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO:
3.
195. 195. The composition of claim 194, wherein the CasX protein and the gNA are associated together in a ribonucleoprotein complex (RNP).
196. 196. The composition of any one of claims 193-195, further comprising a donor template nucleic acid, wherein the donor template comprises a nucleotide sequence having homology to a target DNA.
197. 197. The composition of any one of claims 193-196, further comprising a buffer, a nuclease inhibitor, a protease inhibitor, a liposome, a therapeutic agent, a label, a labeled visualization reagent, or any combination of the foregoing.
198. A composition comprising a gene editing pair described in any one of claims 120 to 136.
199. 200. The composition of claim 198, further comprising a donor template nucleic acid, wherein the donor template comprises a nucleotide sequence having homology to the target DNA.
200. 200. The composition of claim 198 or 199, further comprising a buffer, a nuclease inhibitor, a protease inhibitor, a liposome, a therapeutic agent, a label, a labeled visualization reagent, or any combination of the foregoing.
201. A kit comprising a CasX variant according to any one of claims 1 to 76 and a container.
202. a) a gNA variant according to any one of claims 93 to 119, or 202. The kit of claim 201, further comprising: b) the reference guide RNA and targeting sequence of SEQ ID NO: 4 or 5.
203. 203. The kit of claim 201 or 202, further comprising a donor template nucleic acid, wherein the donor template comprises a nucleotide sequence having homology to a target sequence of the target DNA.
204. 204. The kit of any one of claims 201-203, further comprising a buffer, a nuclease inhibitor, a protease inhibitor, a liposome, a therapeutic agent, a label, a label visualization reagent, or any combination of the foregoing.
205. A kit comprising a gNA variant according to any one of claims 77 to 119.
206. 206. The kit of claim 205, further comprising a CasX variant of any one of claims 1 to 76, or the CasX protein of SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO:
3.
207. 207. The kit of claim 205 or 206, further comprising a donor template nucleic acid, wherein the donor template comprises a nucleotide sequence having homology to a target sequence of the target DNA.
208. 208. The kit of any one of claims 205-207, further comprising a buffer, a nuclease inhibitor, a protease inhibitor, a liposome, a therapeutic agent, a label, a label visualization reagent, or any combination of the foregoing.
209. A kit comprising a gene editing pair or composition described in any one of claims 120 to 136.
210. 210. The kit of claim 209, further comprising a donor template nucleic acid, wherein the donor template comprises a nucleotide sequence having homology to a target DNA.
211. 211. The kit of claim 209 or 210, further comprising a buffer, a nuclease inhibitor, a protease inhibitor, a liposome, a therapeutic agent, a label, a label visualization reagent, or any combination of the foregoing.
212. A CasX variant comprising any one of the sequences listed in Table 3.
213. A gNA variant comprising any one of the sequences listed in Table 2.
214. 214. The gNA variant of claim 213, further comprising a targeting sequence of at least 10-30 nucleotides complementary to the target DNA.
215. The gNA variant of claim 214, wherein the targeting sequence has 20 nucleotides.
216. The gNA variant of claim 214, wherein the targeting sequence has 19 nucleotides.
217. The gNA variant of claim 214, wherein the targeting sequence has 18 nucleotides.
218. The gNA variant of claim 214, wherein the targeting sequence has 17 nucleotides.
219. A CasX variant comprising substitutions L379R and A708K and deletion of P793 of SEQ ID NO:
2.
220. A gNA variant comprising the sequence ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGUGGGUAAAGCUCCCUCUUCGGAGGGAGCAUCAAAG (SEQ ID NO: 2238).
221. A gene editing pair or composition comprising a gene editing pair or composition of any one of claims 120 to 136, or a vector of any one of claims 180 to 186, for use as a pharmaceutical.
222. 187. A gene editing pair or composition comprising the gene editing pair or composition of any one of claims 120 to 136 or the vector of any one of claims 180 to 186 for use in a method of treatment, the method comprising editing or modifying target DNA, optionally said editing being carried out in a subject with a mutation in an allele of a gene, said mutation causing a disease or disorder in said subject, preferably said editing changing the mutation to a wild type allele of said gene, or knocking down or knocking out an allele of a gene that causes a disease or disorder in said subject.
Citation Information
Patent Citations
RNA-guided nucleic acid modifying enzymes and methods of use thereof
WO2018064371A1