Compositions and methods for the targeting of BCL11A

Class 2, Type V CRISPR systems, using CasX and gRNAs, address the challenge of modulating BCL11A gene expression to treat hemoglobinopathies by precisely editing the gene, improving therapeutic outcomes for conditions like sickle cell anemia and beta-thalassemia.

GB2616795BActive Publication Date: 2025-08-13SCRIBE THERAPEUTICS INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
GB2023009871
Authority / Receiving Office
GB · GB
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-12-03
Filing Date
2021-12-02
Publication Date
2025-08-13
Estimated Expiration
2041-12-02

AI Technical Summary

Technical Problem

Current technologies are inadequate in effectively modulating the expression of the BCL11A gene, which regulates the switch from fetal hemoglobin to adult hemoglobin, leading to conditions like sickle cell anemia and beta-thalassemia.

Method used

Utilization of Class 2, Type V CRISPR nuclease proteins, such as CasX, in conjunction with specifically designed guide nucleic acids (gRNAs) to target and edit the BCL11A gene, either knocking it down or knocking it out, thereby reducing its expression.

Benefits of technology

This approach provides a precise and effective means to modify the BCL11A gene, potentially treating hemoglobinopathies by altering hemoglobin production, offering therapeutic benefits for conditions like sickle cell anemia and beta-thalassemia.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000001_0000
    Figure 00000001_0000
  • Figure 00000002_0000
    Figure 00000002_0000
  • Figure 00000003_0000
    Figure 00000003_0000
Patent Text Reader

Abstract

Provided herein are systems comprising Class 2, Type V CRISPR polypeptides, guide nucleic acids (gNA), and optionally donor template nucleic acids useful in the modification of a BCL11A gene. The syst
Need to check novelty before this filing date? Find Prior Art

Description

CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to U.S. provisional patent application number 63 / 120,885, filed on December 3, 2020, the contents of which are incorporated by reference in their entirety herein. INCORPORATION BY REFERENCE OF SEQUENCE LISTING

[0002] This application contains a Sequence Listing which has been submitted in ASCII format via EFS-WEB and is hereby incorporated by reference in its entirety. Said ASCII copy, created on December 1, 2021 is named SCRB_030_01WO_SeqList_ST25.txt and is 8.78MB in size. 15 05 25 BACKGROUND

[0003] Fetal hemoglobin (also hemoglobin F, HbF, or u.2v2) is the main oxygen carrier protein in the human fetus. HbF has a different composition from the adult forms of hemoglobin, which allows it to bind oxygen more strongly than the adult form, allowing the developing fetus to retrieve oxygen from the mother's bloodstream. HbF is a tetramer of two adult a-globin polypeptides and two fetal P-like y-globin polypeptides. During gestation, the duplicated y-globin genes constitute the predominant genes transcribed in the P-globin cluster. After birth, y-globin is replaced by adult P-globin, a process referred to as the “fetal switch”, a process that involves expression of BCL11 A, a regulator of HbF silencing (Sankaran, V.G., et al. Human Fetal Hemoglobin Expression Is Regulated by the Developmental Stage-Specific Repressor BCL11 A. Science 322(5909): 1839-1842 (2008); Liu, N., et al. Direct Promoter Repression by BCL11A Controls the Fetal to Adult Hemoglobin Switch. Cell 173(2):430 (2018)). In healthy adults, the composition of hemoglobin is hemoglobin A (-97%), hemoglobin A2 (2.2 - 3.5%) and hemoglobin F (<1%) (Thomas, C and Lumb, A.B. Physiology of haemoglobin. Continuing Education in Anaesthesia Critical Care &Pain. 12(5): 251-256 (2012)).

[0004] Hemoglobinopathies are inherited single-gene disorders that, in most cases, are inherited as autosomal co-dominant traits. Common hemoglobinopathies include sickle-cell disease and a- and P-thalassemias. Hemoglobinopathies are most common in populations from Africa, the Mediterranean basin and Southeast Asia. Most hemoglobinopathies, including sickle cell anemia, are simply structural abnormalities in the globin proteins themselves. Sickle cell 15 05 25 anemia results from a point mutation in the P-globin structural gene, HBB, leading to the production of an abnormal hemoglobin (HbS), which results in a reduced oxygen-carrying capacity of the blood. Thalassemias, in contrast, usually result in underproduction of normal globin proteins, often through mutations in regulatory genes, leading to deficient or absent adult hemoglobin (HbA) In P-thalassemia, where P-globin is deficient, increased y-globin expression reduces the imbalance of the a- and P-globin chains that underlies the pathophysiology of anemia in this condition (Liu, N., et al. Direct Promoter Repression by BCL11A Controls the Fetal to Adult Hemoglobin Switch. Cell 173(2): 430 (2018)). Both sickle cell disease and thalassemia may cause anemia.

[0005] B-cell lymphoma / leukemia 11A (BCL11 A) is a protein that in humans is encoded by the BCL11A gene. During hematopoietic cell differentiation, this gene is down-regulated and has been found to play a role in the suppression of fetal hemoglobin production. BCL11A is a major repressor protein of hemoglobin F production, by binding to the gene coding for the y subunit at the promoter region (Sankaran VG, et al. Human fetal hemoglobin expression is regulated by the developmental stage-specific repressor BCL11 A. Science 322:1839 (2008)). As increased y-globin reduces the clinical severity of the P-hemoglobinopathies, sickle-cell disease, and P-thalassemia caused by mutation or decreased expression of P-globin, respectively, gene editing of BCL11A to increase expression of y-globin beyond the residual ~1% fetal hemoglobin has been proposed as an attractive therapeutic strategy in adults with hemoglobinopathies (Smith, E.C., et al. Strict in vivo specificity of the Bell la erythroid enhancer. Blood 128(19):2338 (2016)).

[0006] The advent of CRISPR / Cas systems and the programmable nature of these minimal systems has facilitated their use as a versatile technology for genomic manipulation and engineering. To date, the use of CRISPR / Cas systems for the treatment of hemoglobinopathies have been limited to the editing of cells ex vivo, followed by transplantation into subjects suffering from the underlying hemoglobinopathy. Thus, there is a need for compositions and methods to regulate BCL11A to reduce direct -y-globin gene promoter repression in vivo in subjects with these diseases. Provided herein are compositions and methods for targeting the BCL11A gene to the address this need. SUMMARY 15 05 25 In a first aspect of the invention, there is provided a system for modifying a polypyrimidine tract-binding protein 1 (BCL11 A) gene target nucleic acid sequence, the system comprising a CasX variant protein and a guide ribonucleic acid (gRNA) variant, wherein a. the gRNA variant comprises: (i) a scaffold stem loop sequence of SEQ ID NO: 25; and (ii) a targeting sequence complementary to a target nucleic acid sequence comprising a region within a polypyrimidine tract-binding protein 1 (BCL11 A) gene; b. the Cas X variant protein is a chimeric CasX variant protein comprising: (i) the NTSB domain of SEQ ID NO: 1, or a sequence with at least 90% sequence identity thereto; (ii) the helical lb domain of SEQ ID NO: 1, or a sequence with at least 90% sequence identity thereto; and (iii) the RuvC a and RuvC b domains of SEQ ID NO: 2, or sequences with at least 90% sequence identity thereto. In a second aspect of the invention, there is provided system for modifying a polypyrimidine tract-binding protein 1 (BCL11 A) gene target nucleic acid sequence, the system comprising a CasX variant protein and a guide ribonucleic acid (gRNA) variant, wherein the gRNA variant comprises: (i) a scaffold stem loop sequence of SEQ ID NO: 25; and (ii) a targeting sequence complementary to a target nucleic acid sequence comprising a region within a polypyrimidine tract-binding protein 1 (BCL11 A) gene, wherein the targeting sequence of the gRNA variant comprises a sequence selected from the group consisting of SEQ ID NOS: 22, 23, 2949, 2948, 15747, and 15748, wherein the system is capable of modifying the BCL11A gene, wherein the modifying comprises introducing an insertion, deletion, substitution, duplication, or inversion of one or more nucleotides in the BCL11A gene. Preferred embodiments of the invention in any of its various aspects are as described below or as defined in the sub claims. All embodiments and disclosures described below are not part of the present invention unless they are encompassed by the scope of the appended claims. 15 05 25 In addition, references to methods of treatment by therapy, surgery or in vivo diagnosis are to be interpreted as references to compounds, pharmaceutic compositions and medicaments of the present invention for use in those methods.

[0007] The present disclosure relates to compositions of modified Class 2, Type V CRISPR proteins and guide nucleic acids used to alter a target nucleic acid comprising a BCL11A gene in cells. The Class 2, Type V CRISPR proteins and guide nucleic acids are modified for passive entry into target cells. The Class 2, Type V CRISPR proteins and guide nucleic acids are useful in a variety of methods for target nucleic acid modification of BCL11 A-related diseases, which methods are also provided.

[0008] The present disclosure relates to CasX:guide nucleic acid systems (CasX:gRNA systems) and methods used to knock-down or knock-out a BCL11A gene in order to reduce or eliminate expression of the BCL11A gene product in subjects having a P-hemoglobinopathy-related disease.

[0009] The CasX:gRNA system gRNA may be a gRNA, or a chimera of RNA and DNA, and may be a single-molecule gRNA or a dual-molecule gRNA. The CasX:gRNA system gRNA may have a targeting sequence complementary to a target nucleic acid sequence comprising a region within the BCL11A gene. The targeting sequence of the gRNA may be selected from the group consisting of SEQ ID NOS: 272-2100 and 2286-26789 or a sequence having at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 85%, at least about 90%, or at least about 95% identity thereto. The gRNA can comprise a targeting sequence comprising 15 to 20 consecutive nucleotides. The targeting sequence of the gRNA may consist of 20 nucleotides. The targeting sequence may consist of 19 nucleotides. The targeting sequence may consist of 18 nucleotides. The targeting sequence may consist of 17 nucleotides. The targeting sequence may consist of 16 nucleotides. The targeting sequence may consist of 15 nucleotides. The targeting sequence of the gRNA may have a sequence selected from the group consisting of SEQ ID NOS: 272-2100 and 2286-26789. The targeting sequence of the gRNA may have a sequence selected from the group consisting of SEQ ID NOS: 272-2100 and 2286-26789, with a single nucleotide removed from the 3' end of the sequence. The targeting sequence may consist of 18 nucleotides, have a sequence selected from the group consisting of SEQ ID NOS: 272-2100 and 2286-26789, with two nucleotides removed from the 3’ end of the sequence. The targeting sequence may consist of 17 nucleotides, have a sequence selected from the group consisting of SEQ ID NOS: 272-2100 and 2286-26789, with three nucleotides 15 05 25 removed from the 3’ end of the sequence. The targeting sequence many consist of 16 nucleotides, have a sequence selected from the group consisting of SEQ ID NOS: 272-2100 and 2286-26789, with four nucleotides removed from the 3’ end of the sequence. The targeting sequence may consist of 15 nucleotides, have a sequence selected from the group consisting of SEQ ID NOS: 272-2100 and 2286-26789, with five nucleotides removed from the 3’ end of the sequence.

[0010] The gRNA may have a scaffold comprising a sequence selected from the group consisting of sequences SEQ ID NOS: 2238-2285, 26794-26839 and 27219-27265, or as set forth in Table 3, or a sequence having at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% sequence identity thereto. The gRNA may have a scaffold comprising a sequence selected from the group consisting of sequences SEQ ID NOS: 2238-2285, 26794-26839 and 27219-27265. The gRNA may have a scaffold comprising a sequence selected from the group consisting of sequences SEQ ID NOS: 2101-2285, 26794-26839 and 27219-27265.

[0011] The CasX:gRNA systems may comprise a CasX variant sequence having a sequence selected from the group consisting of SEQ ID NOS: 36-99, 101-148, 26908-27154, or a sequence having at least about 50%, at least about 60%, at least about 70%, at least about 80%, or at least about 90%, or at least about 95%, or at least about 96%, or at least about 97%, or at least about 98%, or at least about 99% sequence identity thereto. The CasX:gRNA systems may comprise a CasX variant sequence having a sequence selected from the group consisting of SEQ ID NOS: 36-99, 101-148, 26908-27154. The CasX:gRNA systems may comprise a CasX variant sequence having a sequence selected from the group consisting of SEQ ID NOS: 59, 72-99, 101-148, and 26908-27154, or as set forth in Table 4, or a sequence having at least about 50%, at least about 60%, at least about 70%, at least about 80%, or at least about 90%, or at least about 95%, or at least about 96%, or at least about 97%, or at least about 98%, or at least about 99% sequence identity thereto. The CasX:gRNA systems may comprise a CasX variant sequence having a sequence selected from the group consisting of SEQ ID NOS: 59, 72-99, 101-148, and 26908-27154. The CasX:gRNA systems may comprise a CasX variant sequence having a sequence selected from the group consisting of SEQ ID NOS: 132-148 and 26908-27154, or a sequence having at least about 50%, at least about 60%, at least about 70%, at least about 80%, or at least about 90%, or at least about 95%, or at least about 96%, or at least about 97%, or at 15 05 25 least about 98%, or at least about 99% sequence identity thereto. The CasX:gRNA systems may comprise a CasX variant sequence having a sequence selected from the group consisting of SEQ ID NOS: 132-148 and 26908-27154. A CasX variant may exhibit one or more improved characteristics relative to any one of the reference CasX proteins of SEQ ID NOS: 1-3. The CasX variant protein may have binding affinity for a protospacer adjacent motif (PAM) sequence selected from the group consisting of TTC, ATC, GTC, and CTC. The CasX variant protein may have binding affinity for the PAM sequence that is at least 1.5-fold greater compared to the binding affinity of any one of the reference CasX proteins of SEQ ID NOS: 1-3 for the PAM sequences selected from the group consisting of TTC, ATC, GTC, and CTC.

[0012] The CasX molecule and the gRNA molecule may be associated together in a ribonuclear protein complex (RNP). The RNP comprising the CasX variant and the gRNA variant may exhibit greater editing efficiency and / or binding of a target sequence in the target DNA when any one of the PAM sequences TTC, ATC, GTC, or CTC is located 1 nucleotide 5’ to the non-target strand sequence having identity with the targeting sequence of the gRNA in a cellular assay system compared to the editing efficiency and / or binding of an RNP comprising a reference CasX protein and a reference gRNA in a comparable assay system.

[0013] The CasX:gRNA system may further comprise a donor template comprising a nucleic acid comprising at least a portion of a BCL11A gene and having at least 1 to about 5 mutations relative to the wild-type sequence, wherein the BCL11A gene portion is selected from the group consisting of a BCL11A exon, a BCL11A intron, a BCL11A intron-exon junction, a BCL11A regulatory element, or combinations thereof, wherein the donor template is used to knock down or knock out the BCL11A gene. In some cases, the donor sequence is a single-stranded DNA template or a single stranded RNA template. In other cases, the donor template is a doublestranded DNA template.

[0014] The disclosure relates to nucleic acids encoding the CasX:gRNA systems of any of the embodiments described herein, as well as vectors comprising the nucleic acids. The vector may be selected from the group consisting of a retroviral vector, a lentiviral vector, an adenoviral vector, an adeno-associated viral (AAV) vector, a herpes simplex virus (HSV) vector, a plasmid, a minicircle, a nanoplasmid, and an RNA vector. The vector may be a CasX delivery particle (XDP) comprising an RNP of a CasX and gRNA as described herein and, optionally, a donor template nucleic acid and a targeting moiety such as a viral-derived glycoprotein. 15 05 25

[0015] The disclosure provides a method of modifying a BCL11A target nucleic acid sequence of a cells of a population, wherein said method comprises introducing into the cell: a) CasX:gRNA system as disclosed herein; b) the nucleic acid as disclosed herein; c) the vector as disclosed herein; d) the XDP as disclosed herein; or e) a combination of the foregoing. In the method, the modifying may comprise introducing an insertion, deletion, substitution, duplication, or inversion of one or more nucleotides in the target nucleic acid sequence as compared to the wild-type sequence. The target BCL11A gene includes the GATA1 erythroid-specific enhancer binding site (GATA1) as a regulatory element. The method of modifying may comprise modification of the GATA1 sequence, wherein the BCL11A gene is knocked down or knocked out by the modification. The method may further comprise contacting the target nucleic acid with a donor template nucleic acid as disclosed herein. In the method, the donor template may comprise a nucleic acid comprising at least a portion of a BCL11A gene but with one or more mutations for knocking out or knocking down the BCL11A gene. In some cases, the modifying of the target nucleic acid sequence occurs in vitro or ex vivo. In some cases, the modifying of the target nucleic acid sequence occurs in vivo. The cell may be a eukaryotic cell selected from the group consisting of a rodent cell, a mouse cell, a rat cell, a primate cell, and a non-human primate cell. The cell may be a human cell. The cell may be a selected from the group consisting of a hematopoietic stem cell (HSC), a hematopoietic progenitor cell (HPC), a CD34+ cell, a mesenchymal stem cell (MSC), induced pluripotent stem cell (iPSC), a common myeloid progenitor cell, a proerythroblast cell, and a erythroblast cell. The cell may be an autologous cell derived from a subject with a p-hemoglobinopathy-related disease. The cell may be allogenic, but of the same species as the subject to be treated.

[0016] The disclosure provides methods of modifying a target nucleic acid sequence of the BCL11A gene wherein the target cells of a population are contacted using vectors encoding the CasX protein and one or more gRNAs comprising a targeting sequence complementary to the BCL11A gene, and optionally further comprising a donor template. In some cases, the vector is an Adeno-Associated Viral (AAV) vector selected from AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV-Rh74, or AAVRhlO. In other cases, the vector is a lentiviral vector. The disclosure provides methods wherein the target cells are contacted using a vector, and wherein the vector is a CasX delivery particle (XDP) comprising an RNP of a CasX and gRNA as described herein and, optionally, a donor template nucleic acid. The vector may be administered to a subject at a therapeutically effective dose. The subject can be a mouse, 15 05 25 rat, pig, non-human primate, or a human. The dose can be administered by a route of administration selected from transplantation, local injection, systemic infusion, or combinations thereof.

[0017] The disclosure provides a method of treating a p-hemoglobinopathy-related disease in a subject in need thereof, comprising modifying a gene encoding BCL11A gene in a cell of the subject, the modifying comprising either contacting said cell with: a) CasX:gRNA system as disclosed herein; b) the nucleic acid as disclosed herein; c) the vector as disclosed herein; d) the XDP as disclosed herein; or e) a combination of the foregoing. The P-hemoglobinopathy-related disease may be sickle cell anemia or beta-thalassemia. The methods of treating a subject with a P-hemoglobinopathy-related disease may result in improvement in at least one clinically-relevant parameter. The methods of treating a subject with a p-hemoglobinopathy-related disease may result in improvement in at least two clinically-relevant parameters.

[0018] The disclosure provides use of the CasX:gRNA systems, nucleic acids, vectors or XDP described herein for treating a P-hemoglobinopathy-related disease in a subject in need thereof. The use may comprise modifying a gene encoding BCL11A gene in a cell of the subject, the modifying comprising either contacting said cell with: a) CasX gRNA system as disclosed herein; b) the nucleic acid as disclosed disclosed herein; c) the vector as disclosed herein; d) the XDP as disclosed herein; or e) a combination of the foregoing. INCORPORATION BY REFERENCE

[0019] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. The contents of U.S. provisional applications 63 / 121,196, filed on December 3, 2020, 63 / 162,346 filed on March 17, 2021, and 63 / 208,855, filed on June 9, 2021, which disclose CasX variants and gRNA variants, are hereby incorporated by reference in their entireties. The contents of international application publications WO 2020 / 247882, published December 10, 2020, WO 2020 / 247883, published December 10, 2020, and WO 2021 / 113772, published June 10, 2021 are hereby incorporated by reference in their entireties. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The novel features of the disclosure are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present disclosure will be 15 05 25 obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the disclosure are utilized, and the accompanying drawings of which:

[0021] FIG. lisa graph of the results of an assay for the quantification of active fractions of RNP formed by sgRNA 174 (SEQ ID NO: 2238) and the CasX variants 119 (SEQ ID NO: 59), 457 (SEQ ID NO: 101), 488 (SEQ ID NO: 123) and 491 (SEQ ID NO: 126), as described in Example 8. Equimolar amounts of RNP and target were co-incubated and the amount of cleaved target was determined at the indicated timepoints. Mean and standard deviation of three independent replicates are shown for each timepoint. The biphasic fit of the combined replicates is shown. “2” refers to the reference CasX protein of SEQ ID NO: 2.

[0022] FIG. 2 shows the quantification of active fractions of RNP formed by CasX2 (reference CasX protein of SEQ ID NO:2) and the modified sgRNAs, as described in Example 8. Equimolar amounts of RNP and target were co-incubated and the amount of cleaved target was determined at the indicated timepoints. Mean and standard deviation of three independent replicates are shown for each timepoint. The biphasic fit of the combined replicates is shown.

[0023] FIG. 3 shows the quantification of active fractions of RNP formed by CasX 491 and the modified sgRNAs under guide-limiting conditions, as described in Example 8. Equimolar amounts of RNP and target were co-incubated and the amount of cleaved target was determined at the indicated timepoints. The biphasic fit of the data is shown.

[0024] FIG. 4 shows the quantification of cleavage rates of RNP formed by sgRNAl 74 and the CasX variants, as described in Example 8. Target DNA was incubated with a 20-fold excess of the indicated RNP and the amount of cleaved target was determined at the indicated time points. Mean and standard deviation of three independent replicates are shown for each timepoint, except for 488 and 491 where a single replicate is shown. The monophasic fit of the combined replicates is shown.

[0025] FIG. 5 shows the quantification of cleavage rates of RNP formed by CasX2 and the indicated sgRNA variants, as described in Example 8. Target DNA was incubated with a 20-fold excess of the indicated RNP and the amount of cleaved target was determined at the indicated time points. Mean and standard deviation of three independent replicates are shown for each timepoint. The monophasic fit of the combined replicates is shown. 15 05 25

[0026] FIG. 6 shows the quantification of initial velocities of RNP formed by CasX2 and the sgRNA variants, as described in Example 8. The first two time-points of the previous cleavage experiment were fit with a linear model to determine the initial cleavage velocity.

[0027] FIG. 7 shows the quantification of cleavage rates of RNP formed by CasX491 and the sgRNA variants, as described in Example 8. Target DNA was incubated with a 20-fold excess of the indicated RNP at 10°C and the amount of cleaved target was determined at the indicated time points. The monophasic fit of the timepoints is shown.

[0028] FIG. 8 shows the quantification of competent fractions of RNP of CasX variant 515 (SEQ ID NO: 133) and 526 (SEQ ID NO: 143) complexed with gRNA variant 174 compared to RNP of reference CasX 2 complexed with gRNA 2 using equimolar amounts of indicated RNP and a complementary target, as described in Example 8. The biphasic fit for each time course or set of combined replicates is shown.

[0029] FIG. 9 shows the quantification of cleavage rates of RNP of CasX variant 515 and 526 complexed with gRNA variant 174 compared to RNP of reference CasX 2 complexed with gRNA 2 using with a 20-fold excess of the indicated RNP, as described in Example 8.

[0030] FIG. 10A shows the quantification of cleavage rates of CasX variants on TTC PAM, as described in Example 5. Target DNA substrates with identical spacers and the indicated PAM sequence were incubated with a 20-fold excess of the indicated RNP at 37°C and the amount of cleaved target was determined at the indicated time points. Monophasic fit of a single replicate is shown.

[0031] FIG. 10B shows the quantification of cleavage rates of CasX variants on CTC PAM, as described in Example 5. Target DNA substrates with identical spacers and the indicated PAM sequence were incubated with a 20-fold excess of the indicated RNP at 37°C and the amount of cleaved target was determined at the indicated time points. Monophasic fit of a single replicate is shown.

[0032] FIG. 10C shows the quantification of cleavage rates of CasX variants on GTC PAM, as described in Example 5. Target DNA substrates with identical spacers and the indicated PAM sequence were incubated with a 20-fold excess of the indicated RNP at 37°C and the amount of cleaved target was determined at the indicated time points. Monophasic fit of a single replicate is shown.

[0033] FIG. 10D shows the quantification of cleavage rates of CasX variants on ATC PAM, as described in Example 5. Target DNA substrates with identical spacers and the indicated PAM 15 05 25 sequence were incubated with a 20-fold excess of the indicated RNP at 37°C and the amount of cleaved target was determined at the indicated time points. Monophasic fit of a single replicate is shown.

[0034] FIG. 11A shows the quantification of cleavage rates of RNP of CasX variant 491 and guide 174 on NTC PAMs, as described in Example 5. Timepoints were taken over the course of 10 minutes and the fraction cleaved was graphed for each target and timepoint, but only the first two minutes of the time course are shown for clarity.

[0035] FIG. 1 IB shows the quantification of cleavage rates of RNP of CasX variant 491 and guide 174 on NTT PAMs, as described in Example 5. Timepoints were taken over the course of 10 minutes and the fraction cleaved was graphed for each target and timepoint.

[0036] FIG. 12A shows the quantification of cleavage by RNP formed by sgRNA174 and the CasX variants 515 using spacer lengths of 18, 19, or 20 nucleotides, as described in Example 9. Target DNA was incubated with a 20-fold excess of the indicated RNP and the amount of cleaved target was determined at the indicated time points. Mean and standard deviation of three independent replicates are shown for each timepoint. The monophasic fit of the combined replicates is shown.

[0037] FIG. 12B shows the quantification of cleavage by RNP formed by sgRNA174 and the CasX variant 526 using spacer lengths of 18, 19, or 20 nucleotides, as described in Example 9. Target DNA was incubated with a 20-fold excess of the indicated RNP and the amount of cleaved target was determined at the indicated time points. Mean and standard deviation of three independent replicates are shown for each timepoint. The monophasic fit of the combined replicates is shown.

[0038] FIG. 13 is a schematic showing an example of CasX protein and scaffold DNA sequence for packaging in adeno-associated virus (AAV). The DNA segment between the AAV inverted terminal repeats (ITRs), comprised of a CasX-encoding DNA and its promoter, and scaffold-encoding DNA and its promoter gets packaged within an AAV capsid during AAV production.

[0039] FIG. 14 shows the results of an editing assay comparing gRNA scaffolds 229-237 (see Table 3 for corresponding sequences and SEQ ID NOs) to scaffold 174 in mouse neural progenitor cells (mNPC) isolated from the Ai9-tdtomato transgenic mice. Cells were nucleofected with the indicated doses of p59 plasmids encoding CasX 491, the scaffold, and spacer 11.30 (5’ AAGGGGCUCCGCACCACGCC 3’, SEQ ID NO: 27197) targeting mRHO. 15 05 25 Editing at the mRHO locus was assessed 5 days post-transfection by NGS, and show that editing with constructs with scaffolds 230, 231, 234 and 235 demonstrated greater editing compared to constructs with scaffold 174 at both doses.

[0040] FIG. 15 shows the results of an editing assay comparing gRNA scaffolds 229-237 to scaffold 174 in mNPC cells. Cells were nucleofected with the indicated doses of p59 plasmids encoding CasX491, the scaffold, and spacer 12.7 (5’ CUGCAUUCUAGUUGUGGUUU 3’, SEQ ID NO: 27198) targeting repeat elements preventing expression of the tdTomato fluorescent protein. Editing was assessed 5 days post-transfection by FACS, to quantify the fraction of tdTomato positive cells. Cells nucleofected with scaffolds 231-235 displayed approximately 35% greater editing compared to constructs with scaffold 174 at the high dose, and approximately 25% greater editing at the low dose.

[0041] FIG. 16 shows the results of an editing assay comparing CasX nucleases 2, 119, 491, 515, 527, 528, 529, 530, and 531 (see Table 4 for corresponding sequences and SEQ ID NOs) in a custom HEK293 cell line, PASSV1.01. Cells were lipofected with 2 pg of p67 plasmid encoding the indicated CasX protein. After five days, cell genomic DNA was extracted. PCR amplification and Next-Generation Sequencing was performed to isolate and quantify the fraction of edited cells at custom designed on-target editing sites. For each sample, editing was evaluated at target sites (individual points) consisting of the following PAM sequences: 48 TTC, 14 ATC, 22 CTC, 11 GTC individual sites, and percent editing was normalized to a vehicle control. Cells lipofected with any nuclease displayed higher mean editing at TTC PAM target sites (horizontal bar) than that of the wild-type nuclease CasX 2, except CasX 528. The relative preference of any given nuclease for the four different PAM sequences is also represented by the violin plots. In particular, CasX nucleases 527, 528, and 529 exhibit substantially different PAM preferences than that of the wild-type nuclease CasX 2.

[0042] FIG. 17 shows the results of an editing assay comparing improved CasX nuclease 491 to improved nucleases 532 and 533 in a custom HEK293 cell line, PASS VI .01. Cells were lipofected, in duplicate, with 2 pg of p67 plasmid encoding the indicated CasX protein and a puromycin resistance gene, and grown under puromycin selection. After three days, cell genomic DNA was extracted. PCR amplification and Next-Generation Sequencing was performed to isolate and quantify the fraction of edited cells at custom designed on-target editing sites. For each sample, editing was evaluated at target sites consisting of the following PAM sequences: 48 TTC, 14 ATC, 22 CTC, 11 GTC individual sites, and fraction editing was 15 05 25 normalized to a vehicle control. Cells lipofected with CasX 532 or 533 displayed higher mean editing than Cas 491 at each of the PAM sequences, with the exception of CasX 533 at TTC PAM target sites. Error bars represent standard error of the mean for n = 2 biological samples.

[0043] FIG. 18 shows the results of editing of the BCL11A erythroid enhancer locus in HEK293T cells by CasX protein variant 438 with scaffold 174 compared to a Cas9 system, as described in Example 13.

[0044] FIG. 19 shows the results of editing at the GATA1 binding region of the BCL11A erythroid enhancer locus in K562 cells by CasX protein variant 491 with scaffold 174 compared to CasX protein variant 119 with scaffold 174, as described in Example 14.

[0045] FIG. 20 shows the results of editing at the GATA1 binding region of the BCL11A erythroid enhancer locus in K562 cells by CasX protein variant 491 with scaffold 174 delivered by various doses of XDP, as described in Example 14.

[0046] FIG. 21 shows the results of editing at the GATA1 binding region of the BCL11A erythroid enhancer locus in HSC cells by CasX protein variant 491 with scaffold 174 compared to CasX protein variant 119 with scaffold 174, as described in Example 15.

[0047] FIG. 22 shows the results of editing at the GATA1 binding region of the BCL11A erythroid enhancer locus in HSC cells by CasX protein variant 491 with scaffold 174 delivered by various doses of XDP, as described in Example 15.

[0048] FIG. 23 is a schematic showing the positioning of the spacer 21.1 (SEQ ID NO: 22) relative to the GATA1 binding site sequence in the target nucleic acid. Top strand: SEQ ID NO: 26790, bottom strand: SEQ ID NO: 26791. DETAILED DESCRIPTION

[0049] While exemplary embodiments have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions will now occur to those skilled in the art without departing from the inventions claimed herein. It should be understood that various alternatives to the embodiments described herein may be employed in practicing the embodiments of the disclosure. It is intended that the claims define the scope of the invention and that methods and structures within the scope of these claims and their equivalents be covered thereby.

[0050] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention 15 05 25 belongs. Although methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present embodiments, suitable methods and materials are described below. In case of conflict, the patent specification, including definitions, will control. In addition, the materials, methods, and examples are illustrative only and not intended to be limiting. Numerous variations, changes, and substitutions will now occur to those skilled in the art without departing from the invention. Definitions

[0051] The terms "polynucleotide" and "nucleic acid," used interchangeably herein, refer to a polymeric form of nucleotides of any length, either ribonucleotides or deoxyribonucleotides. Thus, terms "polynucleotide" and "nucleic acid" encompass single-stranded DNA; doublestranded DNA; multi-stranded DNA; single-stranded RNA; double-stranded RNA; multistranded RNA; genomic DNA; cDNA; DNA-RNA hybrids; and a polymer comprising purine and pyrimidine bases or other natural, chemically or biochemically modified, non-natural, or derivatized nucleotide bases.

[0052] "Hybridizable" or "complementary" are used interchangeably to mean that a nucleic acid (e.g., RNA, DNA) comprises a sequence of nucleotides that enables it to non-covalently bind, i.e., form Watson-Crick base pairs and / or G / U base pairs, "anneal", or "hybridize," to another nucleic acid in a sequence-specific, antiparallel, manner (i.e., a nucleic acid specifically binds to a complementary nucleic acid) under the appropriate in vitro and / or in vivo conditions of temperature and solution ionic strength. It is understood that the sequence of a polynucleotide need not be 100% complementary to that of its target nucleic acid to be specifically hybridizable; it can have at least about 70%, at least about 80%>, or at least about 90%>, or at least about 95% sequence identity and still hybridize to the target nucleic acid. Moreover, a polynucleotide may hybridize over one or more segments such that intervening or adjacent segments are not involved in the hybridization event (e.g., a loop structure or hairpin structure, a 'bulge', ‘bubble’ and the like).

[0053] A “gene,” for the purposes of the present disclosure, includes a DNA region encoding a gene product (e.g., a protein, RNA), as well as all DNA regions which regulate the production of the gene product, whether or not such regulatory element sequences are adjacent to coding and / or transcribed sequences. Accordingly, a gene may include regulatory sequences including, but not necessarily limited to, promoter sequences, terminators, translational regulatory sequences such as ribosome binding sites and internal ribosome entry sites, enhancers, silencers, 15 05 25 insulators, boundary elements, replication origins, matrix attachment sites and locus control regions. Coding sequences encode a gene product upon transcription or transcription and translation; the coding sequences of the disclosure may comprise fragments and need not contain a full-length open reading frame. A gene can include both the strand that is transcribed, e.g. the strand containing the coding sequence, as well as the complementary strand.

[0054] The term "downstream" refers to a nucleotide sequence that is located 3' to a reference nucleotide sequence. In certain embodiments, downstream nucleotide sequences relate to sequences that follow the starting point of transcription. For example, the translation initiation codon of a gene is located downstream of the start site of transcription.

[0055] The term "upstream" refers to a nucleotide sequence that is located 5' to a reference nucleotide sequence. In certain embodiments, upstream nucleotide sequences relate to sequences that are located on the 5' side of a coding region or starting point of transcription. For example, most promoters are located upstream of the start site of transcription.

[0056] The term “adjacent to” with respect to polynucleotide or amino acid sequences refers to sequences that are next to, or adjoining each other in a polynucleotide or polypeptide. The skilled artisan will appreciate that two sequences can be considered to be adjacent to each other and still encompass a limited amount of intervening sequence, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides or amino acids.

[0057] The term “regulatory element” is used interchangeably herein with the term “regulatory sequence,” and is intended to include promoters, enhancers, and other expression regulatory elements (e.g. transcription termination signals, such as polyadenylation signals and poly-U sequences). Exemplary regulatory elements include a transcription promoter such as, but not limited to, CMV, CMV+intron A, SV40, RSV, HIV-Ltr, elongation factor 1 alpha (EFla), MMLV-ltr, internal ribosome entry site (IRES) or P2A peptide to permit translation of multiple genes from a single transcript, metallothionein, a transcription enhancer element, a transcription termination signal, poly adenylation sequences, sequences for optimization of initiation of translation, and translation termination sequences. It will be understood that the choice of the appropriate regulatory element will depend on the encoded component to be expressed (e.g., protein or RNA) or whether the nucleic acid comprises multiple components that require different polymerases or are not intended to be expressed as a fusion protein.

[0058] The term "promoter" refers to a DNA sequence that contains an RNA polymerase binding site, transcription start site, TATA box, and / or B recognition element and assists or 15 05 25 promotes the transcription and expression of an associated transcribable polynucleotide sequence and / or gene (or transgene). A promoter can be synthetically produced or can be derived from a known or naturally occurring promoter sequence or another promoter sequence. A promoter can be proximal or distal to the gene to be transcribed. A promoter can also include a chimeric promoter comprising a combination of two or more heterologous sequences to confer certain properties. A promoter of the present disclosure can include variants of promoter sequences that are similar in composition, but not identical to, other promoter sequence(s) known or provided herein. A promoter can be classified according to criteria relating to the pattern of expression of an associated coding or transcribable sequence or gene operably linked to the promoter, such as constitutive, developmental, tissue specific, inducible, etc.

[0059] The term “enhancer” refers to regulatory element DNA sequences that, when bound by specific proteins called transcription factors, regulate the expression of an associated gene. Enhancers may be located in the intron of the gene, or 5’ or 3’ of the coding sequence of the gene. Enhancers may be proximal to the gene ( / . e., within a few tens or hundreds of base pairs (bp) of the promoter), or may be located distal to the gene (z.e., thousands of bp, hundreds of thousands of bp, or even millions of bp away from the promoter). A single gene may be regulated by more than one enhancer, all of which are envisaged as within the scope of the instant disclosure.

[0060] As used herein, a “post-transcriptional regulatory element (PRE),” such as a hepatitis PRE, refers to a DNA sequence that, when transcribed creates a tertiary structure capable of exhibiting post-transcriptional activity to enhance or promote expression of an associated gene operably linked thereto.

[0061] The term “GATA binding site” refers to a DNA binding site for the GATA family of transcription factors. GATA transcription factors typically recognize a target site conforming to the consensus sequence WGATAR (where W = A or T and R = A or G).

[0062] "Recombinant," as used herein, means that a particular nucleic acid (DNA or RNA) is the product of various combinations of cloning, restriction, and / or ligation steps resulting in a construct having a structural coding or non-coding sequence distinguishable from endogenous nucleic acids found in natural systems. Generally, DNA sequences encoding the structural coding sequence can be assembled from cDNA fragments and short oligonucleotide linkers, or from a series of synthetic oligonucleotides, to provide a synthetic nucleic acid which is capable of being expressed from a recombinant transcriptional unit contained in a cell or in a cell-free 15 05 25 transcription and translation system. Such sequences can be provided in the form of an open reading frame uninterrupted by internal non-translated sequences, or introns, which are typically present in eukaryotic genes. Genomic DNA comprising the relevant sequences can also be used in the formation of a recombinant gene or transcriptional unit. Sequences of non-translated DNA may be present 5’ or 3’ from the open reading frame, where such sequences do not interfere with manipulation or expression of the coding regions, and may indeed act to modulate production of a desired product by various mechanisms (see "enhancers” and “promoters", above).

[0063] The term "recombinant polynucleotide" or "recombinant nucleic acid" refers to one which is not naturally occurring, e.g., is made by the artificial combination of two otherwise separated segments of sequence through human intervention. This artificial combination is often accomplished by either chemical synthesis means, or by the artificial manipulation of isolated segments of nucleic acids, e.g., by genetic engineering techniques. Such can be done to replace a codon with a redundant codon encoding the same or a conservative amino acid, while typically introducing or removing a sequence recognition site. Alternatively, it is performed to join together nucleic acid segments of desired functions to generate a desired combination of functions. This artificial combination is often accomplished by either chemical synthesis means, or by the artificial manipulation of isolated segments of nucleic acids, e.g., by genetic engineering techniques.

[0064] Similarly, the term "recombinant polypeptide” or “recombinant protein” refers to a polypeptide or protein which is not naturally occurring, e.g., is made by the artificial combination of two otherwise separated segments of amino sequence through human intervention. Thus, e.g., a protein that comprises a heterologous amino acid sequence is recombinant.

[0065] As used herein, the term "contacting" means establishing a physical connection between two or more entities. For example, contacting a target nucleic acid sequence with a guide nucleic acid means that the target nucleic acid sequence and the guide nucleic acid are made to share a physical connection; e.g., can hybridize if the sequences share sequence similarity.

[0066] “Dissociation constant”, or “Ka”, are used interchangeably and mean the affinity between a ligand “L” and a protein “P”; i.e., how tightly a ligand binds to a particular protein. It can be calculated using the formula Ka=[L] [P] / [LP], where [P], [L] and [LP] represent molar concentrations of the protein, ligand and complex, respectively. 15 05 25

[0067] The disclosure provides compositions and methods useful for editing a target nucleic acid sequence. As used herein “editing” is used interchangeably with “modifying” and includes but is not limited to cleaving, nicking, deleting, knocking in, knocking out, and the like.

[0068] The term "knock-out" refers to the elimination of a gene or the expression of a gene. For example, a gene can be knocked out by either a deletion or an addition of a nucleotide sequence that leads to a disruption of the reading frame. As another example, a gene may be knocked out by replacing a part of the gene with an irrelevant sequence. The term "knock-down" as used herein refers to reduction in the expression of a gene or its gene product(s). As a result of a gene knock-down, the protein activity or function may be attenuated or the protein levels may be reduced or eliminated.

[0069] As used herein, "homology-directed repair" (HDR) refers to the form of DNA repair that takes place during repair of double-strand breaks in cells. This process requires nucleotide sequence homology, and uses a donor template to repair or knock-out a target DNA, and leads to the transfer of genetic information from the donor (e.g., such as the donor template) to the target. Homology-directed repair can result in an alteration of the sequence of the target nucleic acid sequence by insertion, deletion, or mutation if the donor template differs from the target DNA sequence and part or all of the sequence of the donor template is incorporated into the target DNA at the correct genomic locus.

[0070] As used herein, "non-homologous end joining" (NHEJ) refers to the repair of doublestrand breaks in DNA by direct ligation of the break ends to one another without the need for a homologous template (in contrast to homology-directed repair, which requires a homologous sequence to guide repair). NHEJ often results in indels; the loss (deletion) or insertion of nucleotide sequence near the site of the double- strand break.

[0071] As used herein “micro-homology mediated end joining” (MMEJ) refers to a mutagenic double strand break (DSB) repair mechanism, which always associates with deletions flanking the break sites without the need for a homologous template (in contrast to homology-directed repair, which requires a homologous sequence to guide repair). MMEJ often results in the loss (deletion) of nucleotide sequence near the site of the double- strand break.

[0072] A polynucleotide or polypeptide (or protein) has a certain percent "sequence similarity" or "sequence identity" to another polynucleotide or polypeptide, meaning that, when aligned, that percentage of bases or amino acids are the same, and in the same relative position, when comparing the two sequences. Sequence similarity (sometimes referred to as percent similarity, 15 05 25 percent identity, or homology) can be determined in a number of different manners. To determine sequence similarity, sequences can be aligned using the methods and computer programs that are known in the art, including BLAST, available over the world wide web at ncbi.nlm.nih.gov / BLAST. Percent complementarity between particular stretches of nucleic acid sequences within nucleic acids can be determined using any convenient method. Example methods include BLAST programs (basic local alignment search tools) and PowerBLAST programs (Altschul et al., J. Mol. Biol., 1990, 215, 403-410; Zhang and Madden, Genome Res., 1997, 7, 649-656) or by using the Gap program (Wisconsin Sequence Analysis Package, Version 8 for Unix, Genetics Computer Group, University Research Park, Madison Wis.), e.g., using default settings, which uses the algorithm of Smith and Waterman (Adv. Appl. Math., 1981, 2, 482-489).

[0073] The terms "polypeptide," and "protein" are used interchangeably herein, and refer to a polymeric form of amino acids of any length, which can include coded and non-coded amino acids, chemically or biochemically modified or derivatized amino acids, and polypeptides having modified peptide backbones. The term includes fusion proteins, including, but not limited to, fusion proteins with a heterologous amino acid sequence.

[0074] A "vector" or "expression vector" is a replicon, such as plasmid, phage, virus, viruslike particle, or cosmid, to which another DNA segment, i.e., an "insert", may be attached so as to bring about the replication or expression of the attached segment in a cell.

[0075] The term "naturally-occurring" or "unmodified" or "wild type" as used herein as applied to a nucleic acid, a polypeptide, a cell, or an organism, refers to a nucleic acid, polypeptide, cell, or organism that is found in nature.

[0076] As used herein, a "mutation" refers to an insertion, deletion, substitution, duplication, or inversion of one or more amino acids or nucleotides as compared to a wild-type or reference amino acid sequence or to a wild-type or reference nucleotide sequence.

[0077] As used herein the term "isolated" is meant to describe a polynucleotide, a polypeptide, or a cell that is in an environment different from that in which the polynucleotide, the polypeptide, or the cell naturally occurs. An isolated genetically modified host cell may be present in a mixed population of genetically modified host cells.

[0078] A "host cell," as used herein, denotes a eukaryotic cell, a prokaryotic cell, or a cell from a multicellular organism (e.g., a cell line) cultured as a unicellular entity, which eukaryotic or prokaryotic cells are used as recipients for a nucleic acid (e.g., an expression vector), and 15 05 25 include the progeny of the original cell which has been genetically modified by the nucleic acid. It is understood that the progeny of a single cell may not necessarily be completely identical in morphology or in genomic or total DNA complement as the original parent, due to natural, accidental, or deliberate mutation. A "recombinant host cell" (also referred to as a "genetically modified host cell") is a host cell into which has been introduced a heterologous nucleic acid, e.g., an expression vector.

[0079] The term "conservative amino acid substitution" refers to the interchangeability in proteins of amino acid residues having similar side chains. For example, a group of amino acids having aliphatic side chains consists of glycine, alanine, valine, leucine, and isoleucine; a group of amino acids having aliphatic-hydroxyl side chains consists of serine and threonine; a group of amino acids having amide-containing side chains consists of asparagine and glutamine; a group of amino acids having aromatic side chains consists of phenylalanine, tyrosine, and tryptophan; a group of amino acids having basic side chains consists of lysine, arginine, and histidine; and a group of amino acids having sulfur-containing side chains consists of cysteine and methionine. Exemplary conservative amino acid substitution groups are: valine-leucine-isoleucine, phenylalanine-tyrosine, lysine-arginine, alanine-valine, and asparagine-glutamine.

[0080] As used herein, "treatment" or "treating," are used interchangeably herein and refer to an approach for obtaining beneficial or desired results, including but not limited to a therapeutic benefit and / or a prophylactic benefit. By therapeutic benefit is meant eradication or amelioration of the underlying disorder or disease being treated. A therapeutic benefit can also be achieved with the eradication or amelioration of one or more of the symptoms or an improvement in one or more clinical parameters associated with the underlying disease such that an improvement is observed in the subject, notwithstanding that the subject may still be afflicted with the underlying disease.

[0081] The terms "therapeutically effective amount" and "therapeutically effective dose", as used herein, refer to an amount of a drug or a biologic, alone or as a part of a composition, that is capable of having any detectable, beneficial effect on any symptom, aspect, measured parameter or characteristics of a disease state or condition when administered in one or repeated doses to a subject such as a human or an experimental animal. Such effect need not be absolute to be beneficial.

[0082] As used herein, "administering" is meant as a method of giving a dosage of a composition of the disclosure to a subject. 15 05 25

[0083] As used herein, a "subject" is a mammal. Mammals include, but are not limited to, domesticated animals, primates, non-human primates, humans, dogs, porcine (pigs), rabbits, mice, rats and other rodents.

[0084] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. I. General Methods

[0085] The practice of the present invention employs, unless otherwise indicated, conventional techniques of immunology, biochemistry, chemistry, molecular biology, microbiology, cell biology, genomics and recombinant DNA, which can be found in such standard textbooks as Molecular Cloning: A Laboratory Manual, 3rd Ed. (Sambrook et al., Harbor Laboratory Press 2001); Short Protocols in Molecular Biology, 4th Ed. (Ausubel et al. eds., John Wiley &Sons 1999); Protein Methods (Bollag et al., John Wiley &Sons 1996); Nonviral Vectors for Gene Therapy (Wagner et al. eds., Academic Press 1999); Viral Vectors (Kaplift &Loewy eds., Academic Press 1995); Immunology Methods Manual (I. Lefkovits ed., Academic Press 1997); and Cell and Tissue Culture: Laboratory Procedures in Biotechnology (Doyle &Griffiths, John Wiley &Sons 1998), the disclosures of which are incorporated herein by reference.

[0086] Where a range of values is provided, it is understood that endpoints are included and that each intervening value, to the tenth of the unit of the lower limit unless the context clearly dictates otherwise, between the upper and lower limit of that range and any other stated or intervening value in that stated range, is encompassed. The upper and lower limits of these smaller ranges may independently be included in the smaller ranges, and are also encompassed, subject to any specifically excluded limit in the stated range. Where the stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included.

[0087] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. All publications mentioned herein are incorporated herein by reference to disclose and describe the methods and / or materials in connection with which the publications are cited.

[0088] It must be noted that as used herein and in the appended claims, the singular forms “a,” “an,” and “the” include plural referents unless the context clearly dictates otherwise. 15 05 25

[0089] It will be appreciated that certain features of the disclosure, which are, for clarity, described in the context of separate embodiments, may also be provided in combination in a single embodiment. In other cases, various features of the disclosure, which are, for brevity, described in the context of a single embodiment, may also be provided separately or in any suitable sub-combination. It is intended that all combinations of the embodiments pertaining to the disclosure are specifically embraced by the present disclosure and are disclosed herein just as if each and every combination was individually and explicitly disclosed. In addition, all subcombinations of the various embodiments and elements thereof are also specifically embraced by the present disclosure and are disclosed herein just as if each and every such sub-combination was individually and explicitly disclosed herein. II. Systems for Genetic Editing of BCL11A Genes

[0090] In a first aspect, the present disclosure provides systems comprising a Class 2, Type V CRISPR nuclease protein and one or more guide nucleic acids (gRNA) for use in modifying or editing a BCL11A gene in order to reduce or eliminate expression of the BCL11A gene product. Exemplary Class 2, Type V CRISPR nuclease protein and guide nucleic acid systems include the CasX:gRNA system. The CasX:gRNA systems are specifically designed to modify the BCL11A gene in eukaryotic cells. In some cases, the CasX:gRNA systems are designed to knock-down or knock-out the BCL11A gene. Generally, any portion of the BCL11A gene can be targeted using the programable compositions and methods provided herein. In some embodiments, the BCL11A gene to be modified is a wild-type sequence, and the portion to be modified is selected from the group consisting of a BCL11A intron, a BCL11A exon, a BCL11A intron-exon junction, a BCL11A regulatory element, and an intergenic region, or the modification is deletion or mutation of one or more exons.

[0091] As used herein, a “system,” such as the systems comprising a CRISPR nuclease protein and one or more gRNAs the disclosure, as well as nucleic acids encoding the CRISPR nuclease proteins and gRNA and vectors comprising the nucleic acids or CRISPR nuclease protein and one or more gRNAs the disclosure, is used interchangeably with term “composition.”

[0092] The human BCL11A gene (HGNC: 13221) encodes a protein (Q9H165) having the sequence MSRRKQGKPQHLSKREFSPEPLEAILTDDEPDHGPLGAPEGDHDLLTCGQCQMNFPLGDILIFIEHKRKQCNGSLCL EKAVDKPPSPSPIEMKKASNPVEVGIQVTPEDDDCLSTSSRGICPKQEHIADKLLHWRGLSSPRSAHGALIPTPGMS AEYAPQGICKDEPSSYTCTTCKQPFTSAWFLLQHAQNTHGLRIYLESEHGSPLTPRVGIPSGLGAECPSQPPLHGIH 15 05 25 IADNNPFNLLRIPGSVSREASGLAEGRFPPTPPLFSPPPRHHLDPHRIERLGAEEMALATHHPSAFDRVLRLNPMAM EPPAMDFSRRLRELAGNTSSPPLSPGRPSPMQRLLQPFQPGSKPPFLATPPLPPLQSAPPPSQPPVKSKSCEFCGKT FKFQSNLWHRRSHTGEKPYKCNLCDHACTQASKLKRHMKTHMHKSSPMTVKSDDGLSTASSPEPGTSDLVGSASSA LKSWAKFKSENDPNLIPENGDEEEEEDDEEEEEEEEEEEEELTESERVDYGFGLSLEAARHHENSSRGAWGVGDE SRALPDVMQGMVLSSMQHFSEAFHQVLGEKHKRGHLAEAEGHRDTCDEDSVAGESDRIDDGTVNGRGCSPGESASGG LSKKLLLGSPSSLSPFSKRIKLEKEFDLPPAAMPNTENVYSQWLAGYAASRQLKDPFLSFGDSRQSPFASSSEHSSE NGSLRFSTPPGELDGGISGRSGTGSGGSTPHISGPGPGRPSSKEGRRSDTCEYCGKVFKNCSNLTVHRRSHTGERPY KCELCNYACAQSSKLTRHMKTHGQVGKDVYKCEICKMPFSVYSTLEKHMKKWHSDRVLNNDIKTE (SEQ ID NO: 100). The BCL11A gene is defined as the sequence that spans chr2 60450520-60554467 (GRCh38 / hg38 Ensembl 100) of the human genome on chromosome 2.

[0093] In some embodiments, the disclosure provides systems specifically designed to modify the BCL11A gene in eukaryotic cells; either in vitro, ex vivo, or in vivo in a subject. Generally, any portion of the BCL11A target nucleic acid can be targeted using the programmable compositions and methods provided herein. In some embodiments, the CRISPR nuclease is a Class 2, Type V nuclease. Although members of Class 2 Type V CRISPR-Cas systems have differences, they share some common characteristics that distinguish them from the Cas9 systems. Firstly, the Type V nucleases possess an RNA-guided single effector containing a RuvC domain but no HNH domain, and they recognize T-rich PAM 5' upstream to the target region on the non-targeted strand, which is different from Cas9 systems which rely on G-rich PAM at 3' side of target sequences. Type V nucleases generate staggered double-stranded breaks distal to the PAM sequence, unlike Cas9, which generates a blunt end in the proximal site close to the PAM. In addition, Type V nucleases degrade ssDNA in trans when activated by target dsDNA or ssDNA binding in cis. In some embodiments, the disclosure provides Class 2, Type V nuclease selected from the group consisting of Casl2a, Casl2b, Casl2c, Casl2d (CasY), Casl2j. Casl2k, CasZ, and CasX. In some embodiments, the disclosure provides systems comprising one or more CasX proteins and one or more guide nucleic acids (gRNA) as a CasX:gRNA system. In other embodiments, the CasX:gRNA systems of the disclosure comprise one or more CasX proteins, one or more guide nucleic acids (gRNA) and one or more donor template nucleic acids comprising a nucleic acid encoding a portion of a BCL11A gene wherein the donor template nucleic acid comprises a deletion, an insertion, or a mutation of one or more nucleotides in comparison to a genomic nucleic acid sequence encoding the BCL11A protein. Each of these components and their use in the editing of the BCL11A gene is described herein, below. 15 05 25

[0094] In some embodiments, the disclosure provides gene editing pairs of a CasX and a gRNA of any of the embodiments described herein that are capable of being bound together prior to their use for gene editing and, thus, are “pre-complexed” as a ribonuclear protein complex (RNP) The use of a pre-complexed RNP confers advantages in the delivery of the system components to a cell or target nucleic acid sequence for editing of the target nucleic acid sequence.

[0095] In some embodiments, the functional RNP can be delivered ex vivo to a cell by electrophoresis or by chemical means. In other embodiments, the functional RNP can be delivered either ex vivo or in vivo by a vector in their functional form. In some embodiments, the RNP can be delivered in vivo to a subject using a CasX delivery particle (XDP). The gRNA can provide target specificity to the complex by including a targeting sequence (or “spacer”) having a nucleotide sequence that is complementary to a sequence of the target nucleic acid sequence while the CasX variant protein of the pre-complexed CasX:gRNA provides the sitespecific activity, such as cleavage or nicking of the target sequence, that is guided to a target site (e.g., stabilized at a target site) within a target nucleic acid sequence by virtue of its association with the gRNA.

[0096] The systems have utility in the treatment of a subject having a hemoglobinopathy disease, such as sickle cell anemia or P-thalassemia. Each of the components of the CasX:gRNA systems, their functions, and their use in the editing of the target nucleic acids in cells is described more fully, below. III. Guide Nucleic Acids of the Systems for Genetic Editing

[0097] In another aspect, the disclosure relates to specifically-designed guide ribonucleic acids (gRNA) comprising a targeting sequence complementary to (and are therefore able to hybridize with) a target nucleic acid sequence of a BCL11A gene that have utility, when complexed with a CRISPR nuclease, in genome editing of the BCL11A target nucleic acid in a cell. It is envisioned that in some embodiments, multiple gRNAs are delivered in the systems for the modification of a target nucleic acid. For example, a pair of gRNAs with targeting sequences to different or overlapping regions of the target nucleic acid sequence can be used, when each is complexed with a CRISPR nuclease, in order to bind and cleave at two different or overlapping sites within the gene, which is then edited by non-homologous end joining (NHEJ), homology- 15 05 25 directed repair (HDR), homology-independent targeted integration (HITI), micro-homology mediated end joining (MMEJ), single strand annealing (SSA) or base excision repair (BER).

[0098] In some embodiments, the disclosure provides gRNAs utilized in the CasX:gRNA systems that have utility in genome editing a BCL11A gene in a eukaryotic cell. In a particular embodiment, the gRNA of the systems are capable of forming a complex with a CasX nuclease. The present disclosure provides specifically-designed gRNAs wherein the targeting sequence (or spacer, described more fully, below) of the gRNA is complementary to (and are therefore able to hybridize with) target nucleic acid sequences when used as a component of the gene editing CasX:gRNA systems. SEQ ID NOs of representative, but non-limiting examples of targeting sequences to the BCL11A target nucleic acid that can be utilized in the gRNA of the embodiments are presented in Table 1, described more fully below. a. Reference gRNA and gRNA variants

[0099] As used herein, a “reference gRNA" refers to a CRISPR guide nucleic acid comprising a wild-type sequence of a naturally-occurring gRNA. In some embodiments, a reference gRNA of the disclosure may be subjected to one or more mutagenesis methods, such as the mutagenesis methods described herein, which may include Deep Mutational Evolution (DME), deep mutational scanning (DMS), error prone PCR, cassette mutagenesis, random mutagenesis, staggered extension PCR, gene shuffling, or domain swapping, in order to generate one or more gRNA variants with enhanced or varied properties relative to the reference gRNA. gRNA variants also include variants comprising one or more exogenous sequences, for example fused to either the 5’ or 3’ end, or inserted internally. The activity of reference gRNAs may be used as a benchmark against which the activity of gRNA variants are compared, thereby measuring improvements in function or other characteristics of the gRNA variants. In other embodiments, a reference gRNA may be subjected to one or more deliberate, specifically-targeted mutations in order to produce a gRNA variant, for example a rationally designed variant.

[0100] The gRNAs of the disclosure comprise two segments: a targeting sequence and a protein-binding segment. The targeting segment of a gRNA includes a nucleotide sequence (referred to interchangeably as a guide sequence, a spacer, a targeter, or a targeting sequence) that is complementary to (and therefore hybridizes with) a specific sequence (a target site) within the target nucleic acid sequence (e.g., a target ssRNA, a target ssDNA, a strand of a double stranded target DNA, etc.), described more fully below. The targeting sequence of a gRNA is capable of binding to a target nucleic acid sequence, including a coding sequence, a complement 15 05 25 of a coding sequence, a non-coding sequence, and to regulatory elements. The protein-binding segment (or “activator” or “protein-binding sequence”) interacts with (e.g., binds to) a CasX protein as a complex, forming an RNP (described more fully, below). The protein-binding segment is alternatively referred to herein as a “scaffold”, which is comprised of several regions, described more fully, below.

[0101] In the case of a dual guide RNA (dgRNA), the targeter and the activator portions each have a duplex-forming segment, where the duplex forming segment of the targeter and the duplex-forming segment of the activator have complementarity with one another and hybridize to one another to form a double stranded duplex (dsRNA duplex for a gRNA). When the gRNA is a gRNA, the term “targeter” or “targeter RNA” is used herein to refer to a crRNA-like molecule (crRNA: "CRISPR RNA") of a CasX dual guide RNA (and therefore of a CasX single guide RNA when the “activator" and the "targeter” are linked together, e.g., by intervening nucleotides). The crRNA has a 5' region that anneals with the tracrRNA followed by the nucleotides of the targeting sequence. Thus, for example, a guide RNA (dgRNA or sgRNA) comprises a guide sequence and a duplex-forming segment of a crRNA, which can also be referred to as a crRNA repeat. A corresponding tracrRNA-like molecule (activator) also comprises a duplex-forming stretch of nucleotides that forms the other half of the dsRNA duplex of the protein-binding segment of the guide RNA. Thus, a targeter and an activator, as a corresponding pair, hybridize to form a dual guide NA, referred to herein as a “dual guide NA”, a “dual-molecule gRNA”, a “dgRNA”, a “double-molecule guide NA”, or a “two-molecule guide NA”. Site-specific binding and / or cleavage of a target nucleic acid sequence (e.g., genomic DNA) by the CasX protein can occur at one or more locations (e.g., a sequence of a target nucleic acid) determined by base-pairing complementarity between the targeting sequence of the gRNA and the target nucleic acid sequence. Thus, for example, the gRNA of the disclosure have sequences complementarity to and therefore can hybridize with the target nucleic acid that is adjacent to a sequence complementary to a TC PAM motif or a PAM sequence, such as ATC, CTC, GTC, or TTC. Because the targeting sequence of a guide sequence hybridizes with a sequence of a target nucleic acid sequence, a targeter can be modified by a user to hybridize with a specific target nucleic acid sequence, so long as the location of the PAM sequence is considered. Thus, in some cases, the sequence of a targeter may be a non-naturally occurring sequence. In other cases, the sequence of a targeter may be a naturally-occurring sequence, derived from the gene to be edited. In other embodiments, the 15 05 25 activator and targeter of the gRNA are covalently linked to one another (rather than hybridizing to one another) and comprise a single molecule, referred to herein as a “single-molecule gRNA,” “one-molecule guide NA,” “single guide NA”, “single guide RNA”, a “single-molecule guide RNA,” a “one-molecule guide RNA”, or a “sgRNA”. In some embodiments, the sgRNA includes an “activator” or a “targeter” and thus can be an “activator-RNA” and a “targeter-RNA,” respectively. In some embodiments, the gRNA is a ribonucleic acid molecule (“gRNA”), and in other embodiments, the gRNA is a chimera, and comprises both DNA and RNA. As used herein, the term gRNA cover naturally-occurring molecules, as well as sequence variants.

[0102] Collectively, the assembled gRNAs of the disclosure comprise four distinct regions, or domains: the RNA triplex, the scaffold stem, the extended stem, and the targeting sequence that, in the embodiments of the disclosure is specific for a target nucleic acid and is located on the 3’end of the gRNA. The RNA triplex, the scaffold stem, and the extended stem, together, are referred to as the “scaffold” of the gRNA. b. RNA triplex

[0103] In some embodiments of the guide NAs provided herein (including reference sgRNAs), there is a RNA triplex, and the RNA triplex comprises the sequence of a UUU--nX(~4-l 5)-UUU (SEQ ID NO: 226) stem loop that ends with an AAAG after 2 intervening stem loops (the scaffold stem loop and the extended stem loop), forming a pseudoknot that may also extend past the triplex into a duplex pseudoknot. The UU-UUU-AAA sequence of the triplex forms as a nexus between the targeting sequence, scaffold stem, and extended stem. In exemplary CasX sgRNAs, the UUU-loop-UUU region is coded for first, then the scaffold stem loop, and then the extended stem loop, which is linked by the tetraloop, and then an AAAG closes off the triplex before becoming the targeting sequence. c. Scaffold Stem Loop

[0104] In some embodiments of sgRNAs of the disclosure, the triplex region is followed by the scaffold stem loop. The scaffold stem loop is a region of the gRNA that is bound by CasX protein (such as a reference or CasX variant protein). In some embodiments, the scaffold stem loop is a fairly short and stable stem loop. In some cases, the scaffold stem loop does not tolerate many changes, and requires some form of an RNA bubble. In some embodiments, the scaffold stem is necessary for CasX sgRNA function. While it is perhaps analogous to the nexus stem of Cas9 as being a critical stem loop, the scaffold stem of a CasX sgRNA, in some embodiments, has a necessary bulge (RNA bubble) that is different from many other stem loops found in 15 05 25 CRISPR / Cas systems. In some embodiments, the presence of this bulge is conserved across sgRNA that interact with different CasX proteins. An exemplary sequence of a scaffold stem loop sequence of a gRNA comprises the sequence CCAGCGACUAUGUCGUAUGG (SEQ ID NO: 20). d. Extended Stem Loop

[0105] In some embodiments of the CasX sgRNAs of the disclosure, the scaffold stem loop is followed by the extended stem loop. In some embodiments, the extended stem comprises a synthetic tracr and crRNA fusion that is largely unbound by the CasX protein. In some embodiments, the extended stem loop can be highly malleable. In some embodiments, a single guide gRNA is made with a GAAA tetraloop linker or a GAGAAA linker between the tracr and crRNA in the extended stem loop. In some cases, the targeter and activator of a CasX sgRNA are linked to one another by intervening nucleotides and the linker can have a length of from 3 to 20 nucleotides. In some embodiments of the CasX sgRNAs of the disclosure, the extended stem is a large 32-bp loop that sits outside of the CasX protein in the ribonucleoprotein complex. An exemplary sequence of an extended stem loop sequence of a sgRNA comprises the sequence GCGCUUAUUUAUCGGAGAGAAAUCCGAUAAAUAAGAAGC (SEQ ID NO: 21). In some embodiments, the extended stem loop comprises a GAGAAA spacer sequence. e. Targeting Sequence

[0106] In some embodiments of the gRNAs of the disclosure, the extended stem loop is followed by a region that forms part of the triplex, and then the targeting sequence (or “spacer”) at the 3’ end of the gRNA. The targeting sequence targets the CasX ribonucleoprotein holo complex to a specific region of the target nucleic acid sequence of the gene to be modified. Thus, for example, gRNA targeting sequences of the disclosure have sequences complementarity to, and therefore can hybridize to, a portion of the BCL11A gene in a nucleic acid in a eukaryotic cell (e g., a eukaryotic chromosome, chromosomal sequence, a eukaryotic RNA, etc.) as a component of the RNP when the TC PAM motif or any one of the PAM sequences TTC, ATC, GTC, or CTC is located 1 nucleotide 5’ to the non-target strand sequence complementary to the target sequence. The targeting sequence of a gRNA can be modified so that the gRNA can target a desired sequence of any desired target nucleic acid sequence, so long as the PAM sequence location is taken into consideration. In some embodiments, the gRNA scaffold is 5’ of the targeting sequence, with the targeting sequence on the 3’ end of the gRNA. In some 15 05 25 embodiments, the PAM motif sequence recognized by the nuclease of the RNP is TC. In other embodiments, the PAM sequence recognized by the nuclease of the RNP is NTC.

[0107] In some embodiments, the targeting sequence of the gRNA is specific for a portion of a gene encoding a BCL11A protein. In some embodiments, the targeting sequence of a gRNA is specific for a BCL11A exon. In some embodiments, the targeting sequence of a gRNA is specific for a BCL11A intron. In some embodiments, the targeting sequence of the gRNA is specific for a BCL11A intron-exon junction. In some embodiments, the targeting sequence of the gRNA has a sequence that hybridizes with a BCL11A regulatory element, a BCL11A coding region, a BCL11A non-coding region, or combinations thereof (e.g., the intersection of two regions). In some embodiments, the regulatory element comprises a GATA binding sequence. In some embodiments, the targeting sequence of the gRNA is complementary to a sequence comprising one or more single nucleotide polymorphisms (SNPs) of the BCL11A gene or its complement. SNPs that are within BCL11A coding sequence or within BCL11A non-coding sequence are both within the scope of the instant disclosure. In other embodiments, the targeting sequence of the gRNA is complementary to a sequence of an intergenic region of the BCL11A gene or a sequence complementary to an intergenic region of the BCL11A gene.

[0108] In some embodiments, the targeting sequence of a gRNA is designed to be specific for a regulatory element that regulates expression of the BCL11A gene product. Such regulatory elements include, but are not limited to promoter regions, enhancer regions, intergenic regions, 5' untranslated regions (5' UTR), 3' untranslated regions (3' UTR), conserved elements, and regions comprising cis-regulatory elements. The promoter region is intended to encompass nucleotides within 5 kb of the initiation point of the encoding sequence or, in the case of gene enhancer elements or conserved elements, can be thousands of base pairs (bp), hundreds of thousands of bp, or even millions of bp away from the encoding sequence of the gene of the target nucleic acid. In particular embodiments, the targeting sequence of the gRNA hybridizes with a sequence that is complementary to a BCL11A regulatory element. In one embodiment, the targeting sequence of the gRNA is UGGAGCCUGUGAUAAAAGCA (SEQ ID NO: 22), which hybridizes with the BCL11A GATA1 erythroid-specific enhancer binding site sequence, or has at least 90% or at least 95% sequence identity thereto (see FIG. 23). In another embodiment, the targeting sequence of the gRNA is UGCUUUUAUCACAGGCUCCA (SEQ ID NO: 23), or has at least 90% or at least 95% sequence identity thereto. In another embodiment, the targeting sequence of the gRNA is UGCUUUUAUCACAGGCUCCA (SEQ 15 05 25 ID NO: 23), or has at least 90% or at least 95% sequence identity thereto. In other embodiments, the targeting sequence of the gRNA is selected from the group consisting of CAGGCUCCAGGAAGGGUUUG (SEQ ID NO: 2949), GAGGCCAAACCCUUCCUGGA (SEQ ID NO: 2948), AGUGCAAGCUAACAGUUGCU (SEQ ID NO: 15747), and AUACAACUUUGAAGCUAGUC (SEQ ID NO: 15748).

[0109] In subjects that are maturing after birth, GATA1 binding enhances BCL11A expression which, in turn, represses hemoglobin F (HbF) expression, in favor of hemoglobin gamma. However, in subjects with certain hemoglobinopathies, repressing BCL11A expression has been demonstrated to permit HbF expression to resume, which can compensate for otherwise defective hemoglobin in the subject.

[0110] In some embodiments, the targeting sequence of the gRNA has between 14 and 35 consecutive nucleotides. In some embodiments, the targeting sequence has 14, 15, 16, 18, 18, 19, 20, 21, 22, 23 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34 or 35 consecutive nucleotides. In some embodiments, the targeting sequence consists of 20 consecutive nucleotides. In some embodiments, the targeting sequence consists of 19 consecutive nucleotides. In some embodiments, the targeting sequence consists of 18 consecutive nucleotides. In some embodiments, the targeting sequence consists of 17 consecutive nucleotides. In some embodiments, the targeting sequence consists of 16 consecutive nucleotides. In some embodiments, the targeting sequence consists of 15 consecutive nucleotides. In some embodiments, the targeting sequence has 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34 or 35 consecutive nucleotides and the targeting sequence can comprise 0 to 5, 0 to 4, 0 to 3, or 0 to 2 mismatches relative to the target nucleic acid sequence and retain sufficient binding specificity such that the RNP comprising the gRNA comprising the targeting sequence can form a complementary bond with respect to the target nucleic acid.

[0111] Representative, but non-limiting examples of targeting sequences to the target nucleic acid sequence contemplated for use in the gRNA of the disclosure are presented as SEQ ID NOS: 272-2100 and 2286-26789 (see Table 1). In some embodiments, the disclosure provides targeting sequences for an ATC PAM comprising a sequence that is at least 50% identical, at least 55% identical, at least 60% identical, at least 65% identical, at least 70% identical, at least 75% identical, at least 80% identical, at least 85% identical, at least 90% identical, at least 95% identical, or 100% identical to a sequence of SEQ ID NOs: 272-2100 or 2286-5625. In some embodiments, the disclosure provides targeting sequences for an ATC PAM comprising a 15 05 25 sequence of SEQ ID NOs: 272-2100 or 2286-5625. In some embodiments, the disclosure provides targeting sequences for an CTC PAM comprising a sequence that is at least 50% identical, at least 55% identical, at least 60% identical, at least 65% identical, at least 70% identical, at least 75% identical, at least 80% identical, at least 85% identical, at least 90% identical, at least 95% identical, or 100% identical to a sequence of SEQ ID NOs: 5626-13616. In some embodiments, the disclosure provides targeting sequences for an CTC PAM comprising a sequence of SEQ ID NOs: 5626-13616. In some embodiments, the disclosure provides targeting sequences for an GTC PAM comprising a sequence that is at least 50% identical, at least 55% identical, at least 60% identical, at least 65% identical, at least 70% identical, at least 75% identical, at least 80% identical, at least 85% identical, at least 90% identical, at least 95% identical, or 100% identical to a sequence of SEQ ID NOs: 13617-17903. In some embodiments, the disclosure provides targeting sequences for an GTC PAM comprising a sequence of SEQ ID NOs: 13617-17903. In some embodiments, the disclosure provides targeting sequences for an TTC PAM comprising a sequence that is at least 50% identical, at least 55% identical, at least 60% identical, at least 65% identical, at least 70% identical, at least 75% identical, at least 80% identical, at least 85% identical, at least 90% identical, at least 95% identical, or 100% identical to a sequence of SEQ ID NOs: 17904-26789. In some embodiments, the disclosure provides targeting sequences for an TTC PAM comprising a sequence of SEQ ID NOs: 17904-26789. In some embodiments, the targeting sequence contemplated for use in the gRNA of the disclosure of the gRNA comprises a sequence of SEQ ID NOs: 272-2100 or 2286-26789 with a single nucleotide removed from the 3' end of the sequence. In other embodiments, the targeting sequence of the gRNA comprises a sequence of SEQ ID NOs: 272-2100 or 2286-26789 with two nucleotides removed from the 3' end of the sequence. In other embodiments, the targeting sequence of the gRNA comprises a sequence of SEQ ID NOs: 272-2100 or 2286-26789 with three nucleotides removed from the 3' end of the sequence. In other embodiments, the targeting sequence of the gRNA comprises a sequence of SEQ ID NOs: 272-2100 or 2286-26789 with four nucleotides removed from the 3' end of the sequence. In other embodiments, the targeting sequence of the gRNA comprises a sequence of SEQ ID NOs: 272-2100 or 2286-26789 with five nucleotides removed from the 3' end of the sequence. In the foregoing embodiments of the paragraph, thymine (T) nucleotides can be substituted for one or more or all of the uracil (U) nucleotides in any of the targeting sequences such that the gRNA targeting sequence can be a gDNA or a gRNA, or a chimera of RNA and DNA, or in those cases where the encoding sequence for the spacer is incorporated into an expression vector. In some embodiments, a targeting sequence of SEQ ID NOs: 272-2100 or 2286-26789 has at least 1, 2, 3, 4, 5, or 6 or more thymine nucleotides substituted for uracil nucleotides. Table 1. SEQ ID NOs for gRNA Targeting Sequences for BCL11A Gene PAM Type SEQ ID NO ATC 272-2100, 2286-5625 CTC 5626-13616 GTC 13617-17903 Hrnr / "5 1 1 X—' 17904-26789 15 05 25

[0112] In some embodiments, the CasX:gRNA system comprises a first gRNA and further comprises a second (and optionally a third, fourth, fifth, or more) gRNA, wherein the second gRNA or additional gRNA has a targeting sequence complementary to a different or overlapping portion of the target nucleic acid sequence compared to the targeting sequence of the first gRNA such that multiple points in the target nucleic acid are targeted, and, for example, multiple breaks are introduced in the target nucleic acid by the CasX. It will be understood that in such cases, the second or additional gRNA is complexed with an additional copy of the CasX protein. By selection of the targeting sequences of the gRNA, defined regions of the target nucleic acid sequence bracketing a particular location within the target nucleic acid can be modified or edited using the CasX:gRNA systems described herein, including facilitating the insertion of a donor template comprising a mutation of the BCL11A gene. In a particular embodiment, a second gRNA can comprise a targeting sequence complementary to a sequence that is 5’ or 3’ and adjacent to the GATA1 binding site such that the GATA1 binding site is disrupted. f. gRNA scaffolds

[0113] With the exception of the targeting sequence domain, the remaining components of the gRNA are referred to herein as the scaffold. In some embodiments, the gRNA scaffolds are derived from naturally-occurring sequences, described below as reference gRNA. In other embodiments, the gRNA scaffolds are variants of reference gRNA wherein mutations, insertions, deletions or domain substitutions are introduced to confer desirable or improved properties on the gRNA. [0H4] The term “adjacent to” with respect to polynucleotide or amino acid sequences refers to sequences that are next to, or adjoining each other in a polynucleotide or polypeptide. The skilled artisan will appreciate that two sequences can be considered to be adjacent to each other and still encompass a limited amount of intervening sequence, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides or amino acids.

[0115] Table 2 provides the sequences of reference gRNA tracr and scaffold sequences. In some embodiments, the disclosure provides gRNA sequences wherein the gRNA has a scaffold comprising a sequence of SEQ ID NOs: 4-16 as set forth in Table 2, or a sequence having at least one nucleotide modification relative to a reference gRNA sequence having a sequence of any one of SEQ ID NOS: 4-16 of Table 2. It will be understood that in those embodiments wherein a vector comprises a DNA encoding sequence for a gRNA, or where a gRNA is a chimera of RNA and DNA, that thymine (T) bases can be substituted for the uracil (U) bases of any of the gRNA sequence embodiments described herein. Table 2. Reference gRNA tracr and scaffold sequences 15 05 25 SEQ ID NO. Nucleotide Sequence 4 ACAUCUGGCGCGUUUAUUCCAUUACUUUGGAGCCAGUCCCAGCGACUAUGUCGUAUGGACGAAG CGCUUAUUUAUCGGAGAGAAACCGAUAAGUAAAACGCAUCAAAG 5 UACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUGGGUAAAGC GCUUAUUUAUCGGAGAGAAAUCCGAUAAAUAAGAAGCAUCAAAG 6 ACAUCUGGCGCGUUUAUUCCAUUACUUUGGAGCCAGUCCCAGCGACUAUGUCGUAUGGACGAAG CGCUUAUUUAUCGGAGA 7 ACAUCUGGCGCGUUUAUUCCAUUACUUUGGAGCCAGUCCCAGCGACUAUGUCGUAUGGACGAAG CGCUUAUUUAUCGG 8 UACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUGGGUAAAGC GCUUAUUUAU C GGAGA 9 UACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUGGGUAAAGC GCUUAUUUAUCGG 10 GUUUACACACU C C CU CU CAUAGGGU 11 GUUUACACACU C C CU CU CAUGAGGU 12 UUUUACAUAC C C C CU CU CAUGGGAU 13 GUUUACACACUCCCUCUCAUGGGGG 14 CCAGCGACUAUGUCGUAUGG 15 GCGCUUAUUUAUCGGAGAGAAAUCCGAUAAAUAAGAAGC 16 GGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUGGGUAAAGCGCUU AUUUAUCGGA 15 05 25 g. gRNA Variants

[0116] In another aspect, the disclosure relates to guide nucleic acid variants (referred to herein alternatively as “gRNA variant” or “gRNA variant”), which comprise one or more modifications relative to a reference gRNA scaffold. As used herein, “scaffold” refers to all parts to the gRNA necessary for gRNA function with the exception of the targeting sequence.

[0117] In some embodiments, a gRNA variant comprises one or more nucleotide substitutions, insertions, deletions, or swapped or replaced regions relative to a reference gRNA sequence of the disclosure. In some embodiments, a mutation can occur in any region of a reference gRNA scaffold to produce a gRNA variant. In some embodiments, the scaffold of the gRNA variant sequence has at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, or at least 70%, at least 80%, at least 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to the sequence of SEQ ID NO: 4 or SEQ ID NO: 5.

[0118] In some embodiments, a gRNA variant comprises one or more nucleotide changes within one or more regions of the reference gRNA that improve a characteristic relative to the reference gRNA. Exemplary regions include the RNA triplex, the pseudoknot, the scaffold stem loop, and the extended stem loop. In some cases, the variant scaffold stem further comprises a bubble. In other cases, the variant scaffold further comprises a triplex loop region. In still other cases, the variant scaffold further comprises a 5' unstructured region. In one embodiment, the gRNA variant scaffold comprises a scaffold stem loop having at least 60% sequence identity to SEQ ID NO: 14. In another embodiment, the gRNA variant comprises a scaffold stem loop having the sequence of CCAGCGACUAUGUCGUAGUGG (SEQ ID NO: 25). In another embodiment, the disclosure provides a gRNA scaffold comprising, relative to SEQ ID NO:5, a C18G substitution, a G55 insertion, a UI deletion, and a modified extended stem loop in which the original 6 nt loop and 13 most-loop-proximal base pairs (32 nucleotides total) are replaced by a Uvsx hairpin (4 nt loop and 5 loop-proximal base pairs; 14 nucleotides total) and the loop-distal base of the extended stem was converted to a fully base-paired stem contiguous with the new Uvsx hairpin by deletion of the A99 and substitution of G64U. In the foregoing embodiment, the gRNA scaffold comprises the sequence ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGUGGGUAAAGC UCCCUCUUCGGAGGGAGCAUCAAAG (SEQ ID NO: 2238) . 15 05 25

[0119] All gRNA variants that have one or more improved functions or characteristics, or add one or more new functions when the variant gRNA is compared to a reference gRNA described herein, are envisaged as within the scope of the disclosure. A representative example of such a gRNA variant is guide 174 (SEQ ID NO: 2238), the design of which (and the rationale for the design) is described in the Examples. In some embodiments, the gRNA variant adds a new function to the RNP comprising the gRNA variant. In some embodiments, the gRNA variant has an improved characteristic selected from: improved stability; improved solubility; improved transcription of the gRNA; improved resistance to nuclease activity; increased folding rate of the gRNA; decreased side product formation during folding; increased productive folding; improved binding affinity to a CasX protein; improved binding affinity to a target DNA when complexed with a CasX protein; improved gene editing when complexed with a CasX protein; improved specificity of editing when complexed with a CasX protein; and improved ability to utilize a greater spectrum of one or more PAM sequences, including ATC, CTC, GTC, or TTC, in the editing of target DNA when complexed with a CasX protein, or any combination thereof. In some cases, the one or more of the improved characteristics of the gRNA variant is at least about 1.1 to about 100,000-fold improved relative to the reference gRNA of SEQ ID NO: 4 or SEQ ID NO: 5. In other cases, the one or more improved characteristics of the gRNA variant is at least about 1.1, at least about 10, at least about 100, at least about 1000, at least about 10,000, at least about 100,000-fold or more improved relative to the reference gRNA of SEQ ID NO: 4 or SEQ ID NO: 5. In other cases, the one or more of the improved characteristics of the gRNA variant is about 1.1 to 100,00-fold, about 1.1 to 10,00-fold, about 1.1 to 1,000-fold, about 1.1 to 500-fold, about 1.1 to 100-fold, about 1.1 to 50-fold, about 1.1 to 20-fold, about 10 to 100,00-fold, about 10 to 10,00-fold, about 10 to 1,000-fold, about 10 to 500-fold, about 10 to 100-fold, about 10 to 50-fold, about 10 to 20-fold, about 2 to 70-fold, about 2 to 50-fold, about 2 to 30-fold, about 2 to 20-fold, about 2 to 10-fold, about 5 to 50-fold, about 5 to 30-fold, about 5 to 10-fold, about 100 to 100,00-fold, about 100 to 10,00-fold, about 100 to 1,000-fold, about 100 to 500-fold, about 500 to 100,00-fold, about 500 to 10,00-fold, about 500 to 1,000-fold, about 500 to 750-fold, about 1,000 to 100,00-fold, about 10,000 to 100,00-fold, about 20 to 500-fold, about 20 to 250-fold, about 20 to 200-fold, about 20 to 100-fold, about 20 to 50-fold, about 50 to 10,000-fold, about 50 to 1,000-fold, about 50 to 500-fold, about 50 to 200-fold, or about 50 to 100-fold, improved relative to the reference gRNA of SEQ ID NO: 4 or SEQ ID NO: 5. In other cases, the one or more improved characteristics of the gRNA variant is about 1.1 -fold, 1.2-fold, 1.3-fold, 15 05 25 1.4-fold, 1.5-fold, 1.6-fold, 1.7-fold, 1.8-fold, 1.9-fold, 2-fold, 3-fold, 4-fold, 5-fold, 6-fold, 7-fold, 8-fold, 9-fold, 10-fold, 11-fold, 12-fold, 13-fold, 14-fold, 15-fold, 16-fold, 17-fold, 18-fold, 19-fold, 20-fold, 25-fold, 30-fold, 40-fold, 45-fold, 50-fold, 55-fold, 60-fold, 70-fold, 80-fold, 90-fold, 100-fold, 110-fold, 120-fold, 130-fold, 140-fold, 150-fold, 160-fold, 170-fold, 180-fold, 190-fold, 200-fold, 210-fold, 220-fold, 230-fold, 240-fold, 250-fold, 260-fold, 270-fold, 280-fold, 290-fold, 300-fold, 310-fold, 320-fold, 330-fold, 340-fold, 350-fold, 360-fold, 370-fold, 380-fold, 390-fold, 400-fold, 425-fold, 450-fold, 475-fold, or 500-fold improved relative to the reference gRNA of SEQ ID NO: 4 or SEQ ID NO: 5.

[0120] In some embodiments, a gRNA variant can be created by subjecting a reference gRNA to a one or more mutagenesis methods, such as the mutagenesis methods described herein, below, which may include Deep Mutational Evolution (DME), deep mutational scanning (DMS), error prone PCR, cassette mutagenesis, random mutagenesis, staggered extension PCR, gene shuffling, or domain swapping, in order to generate the gRNA variants of the disclosure. The activity of reference gRNAs may be used as a benchmark against which the activity of gRNA variants are compared, thereby measuring improvements in function of gRNA variants compared to the reference gRNA. In other embodiments, a reference gRNA may be subjected to one or more deliberate, targeted mutations, substitutions, or domain swaps in order to produce a gRNA variant, for example a rationally designed variant. Exemplary gRNA variants produced by such methods are described in the Examples and representative sequences of gRNA scaffolds are presented in Table 3.

[0121] In some embodiments, the gRNA variant comprises one or more modifications compared to a reference guide nucleic acid scaffold sequence, wherein the one or more modification is selected from: at least one nucleotide substitution in a region of the gRNA variant; at least one nucleotide deletion in a region of the gRNA variant; at least one nucleotide insertion in a region of the gRNA variant; a substitution of all or a portion of a region of the gRNA variant; a deletion of all or a portion of a region of the gRNA variant; or any combination of the foregoing. In some cases, the modification is a substitution of 1 to 15 consecutive or non-consecutive nucleotides in the gRNA variant in one or more regions. In other cases, the modification is a deletion of 1 to 10 consecutive or non-consecutive nucleotides in the gRNA variant in one or more regions. In other cases, the modification is an insertion of 1 to 10 consecutive or non-consecutive nucleotides in the gRNA variant in one or more regions. In other cases, the modification is a substitution of the scaffold stem loop or the extended stem loop with 15 05 25 an RNA stem loop sequence from a heterologous RNA source with proximal 5' and 3' ends. In some cases, a gRNA variant of the disclosure comprises two or more modifications in one region. In other cases, a gRNA variant of the disclosure comprises modifications in two or more regions. In other cases, a gRNA variant comprises any combination of the foregoing modifications described in this paragraph.

[0122] In some embodiments, a 5' G is added to a gRNA variant sequence for expression in vivo, as transcription from a U6 promoter is more efficient and more consistent with regard to the start site when the +1 nucleotide is a G. In other embodiments, two 5' Gs are added to a gRNA variant sequence for in vitro transcription to increase production efficiency, as T7 polymerase strongly prefers a G in the +1 position and a purine in the +2 position. In some cases, the 5’ G bases are added to the reference scaffolds of Table 2. In other cases, the 5’ G bases are added to the variant scaffolds SEQ ID NOS: 2238-2285, 26794-26839 and 27219-2726 of Table 3.

[0123] Table 3 provides exemplary gRNA variant scaffold sequences of the disclosure. In some embodiments, the gRNA variant scaffold comprises any one of the sequences listed in Table 3, SEQ ID NOS: 2238-2285, 26794-26839 and 27219-27265, or a sequence having at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% sequence identity thereto. In some embodiments, the gRNA variant scaffold comprises any one of SEQ ID NOS: 2238-2285, 26794-26839 and 27219-27265. In some embodiments, the gRNA variant scaffold comprises any one of SEQ ID NOS: 2281-2285, 26794-26839 and 27219-27265, or a sequence having at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% sequence identity thereto. In some embodiments, the gRNA variant scaffold comprises any one of SEQ ID NOS: 2281-2285, 26794-26839 and 27219-27265. It will be understood that in those embodiments wherein a vector comprises a DNA encoding sequence for a gRNA, or where a gRNA is a chimera of RNA and DNA, that thymine (T) bases can be substituted for the uracil (U) bases of any of the gRNA sequence embodiments described herein. 15 05 25 Table 3. Exemplary gRNA Scaffold Sequences SEQ ID NO: NAME NUCLEOTIDE SEQUENCE OR DESCRIPTION OF MODIFICATION 2101 ND phage replication stable 2102 ND Kissing loopbl 2103 ND Kissing loop a 2104 ND 32, uvsX hairpin 2105 ND PP7 2106 ND 64, trip mut, extended stem truncation 2107 ND hyperstable tetraloop 2108 ND C18G 2109 ND U17G 2110 ND CUUCGG loop 2111 ND MS2 2112 ND -1, A2G, -78, G77U 2113 ND QB 2114 ND 45,44 hairpin 2115 ND U1A 2116 ND A14C.U17G 2117 ND CUUCGG loop modified 2118 ND Kissing Ioop b2 2119 ND -76:78, -83:87 2120 ND -4 2121 ND extended stem truncation 2122 ND C55 2123 ND trip mut 2124 ND -76:78 2125 ND -1:5 2126 ND -83:87 2127 ND =+G28, A82U. -84. 2128 ND =+51U 2129 ND -1:4, +G5A, +G86, 2130 ND =+A94 2131 ND =+G72 2132 ND shorten front, CUUCGG loop modified, extend extended 2133 ND A14C 2134 ND -1:3, +G3 2135 ND =+C45. +U46 2136 ND CUUCGG loop modified, fun start 2137 ND -93:94 2138 ND =+U45 2139 ND -69, -94 2140 ND -94 2141 ND modified CUUCGG. minus U in 1st triplex 2142 ND -1:4, +C4, A14C, U17G. +G72, -76:78, -83:87 15 05 25 SEQ ID NO: NAME NUCLEOTIDE SEQUENCE OR DESCRIPTION OF MODIFICATION 2143 ND U1C, -73 2144 ND Scaffold uuCG, stem uuCG. Stem swap, t shorten 2145 ND Scaffold uuCG, stem uuCG. Stem swap 2146 ND =+G60 2147 ND no stem Scaffold uuCG 2148 ND no stem Scaffold uuCG, fun start 2149 ND Scaffold uuCG, stem uuCG, fun start 2150 ND Pseudoknots 2151 ND Scaffold uuCG. stem uuCG 2152 ND Scaffold uuCG, stem uuCG, no start 2153 ND Scaffold uuCG 2154 ND =+GCUC36 2155 ND G quadriplex telomere basket+ ends 2156 ND G quadriplex M3q 2157 ND G quadriplex telomere basket no ends 2158 ND 45.44 hairpin (old version) 2159 ND Sarcin-ricin loop 2160 ND uvsX. C18G 2161 ND truncated stem loop, C18G, trip mut (U10C) 2162 ND short phage rep, C18G 2163 ND phage rep loop, C18G 2164 ND =+G18, stacked onto 64 2165 ND truncated stem loop, C18G, -1 A2G 2166 ND phage rep loop, C18G, trip mut (U10C) 2167 ND short phage rep, C18G, trip mut (U10C) 2168 ND uvsX. trip mut (U10C) 2169 ND truncated stem loop 2170 ND =+A17, stacked onto 64 2171 ND 3' HDV genomic ribozyme 2172 ND phage rep loop, trip mut (U10C) 2173 ND -79:80 2174 ND short phage rep, trip mut (U10C) 2175 ND extra truncated stem loop 2176 ND imr1 giog1 U 1 / U, CloLr 2177 ND short phage rep 2178 ND uvsX, C18G, -1 A2G 2179 ND uvsX. C18G, trip mut (U10C), -1 A2G, HDV -99 G65U 2180 ND 3' HDV antigenomic ribozyme 2181 ND uvsX, C18G, trip mut (U10C), -1 A2G, HDV AA(98:99)C 2182 ND 3' HDV ribozyme (Lior Nissim. Timothy Lu) 2183 ND TAC(1:3)GA, stacked onto 64 2184 ND uvsX, -1 A2G 2185 ND truncated stem loop, C18G, trip mut (U10C), -1 A2G. HDV -99 G65U 2186 ND short phage rep, C18G, trip mut (U10C), -1 A2G, HDV -99 G65U 15 05 25 SEQ ID NO: NAME NUCLEOTIDE SEQUENCE OR DESCRIPTION OF MODIFICATION 2187 ND 3' sTRSV WT viral Hammerhead ribozyme 2188 ND short phage rep, C18G, -1 A2G 2189 ND short phage rep, C18G, trip mut (U10C), -1 A2G. 3' genomic HDV 2190 ND phage rep loop, C18G. trip mut (U10C), -1 A2G. HDV -99 G65U 2191 ND 3' HDV ribozyme (Owen Ryan, Jamie Cate) 2192 ND phage rep loop, C18G, -1 A2G 2193 ND 0.14 2194 ND -78, G77U 2195 ND ND 2196 ND short phage rep, -1 A2G 2197 ND truncated stem loop, C18G, trip mut (U10C), -1 A2G 2198 ND -1, A2G 2199 ND truncated stem loop, trip mut (U10C), -1 A2G 2200 ND uvsX, C18G, trip mut (U10C), -1 A2G 2201 ND phage rep loop, -1 A2G 2202 ND phage rep loop, trip mut (U10C), -1 A2G 2203 ND phage rep loop, C18G, trip mut (U10C), -1 A2G 2204 ND truncated stem loop, C18G 2205 ND uvsX, trip mut (U10C), -1 A2G 2206 ND truncated stem loop. -1 A2G 2207 ND short phage rep. trip mut (U10C), -1 A2G 2208 ND 5'HDV ribozyme (Owen Ryan. Jamie Cate) 2209 ND 5'HDV genomic ribozyme 2210 ND truncated stem loop, C18G, trip mut (U10C), -1 A2G, HDV AA(98:99)C 2211 ND 5'env25 pistol ribozyme (with an added CUUCGG loop) 2212 ND 5'HDV antigenomic ribozyme 2213 ND 3' Hammerhead ribozyme (Lior Nissim, Timothy Lu) guide scaffold scar 2214 ND =+A27, stacked onto 64 2215 ND 5'Hammerhead ribozyme (Lior Nissim, Timothy Lu) smaller scar 2216 ND phage rep loop, C18G, trip mut (U10C), -1 A2G, HDV AA(98:99)C 2217 ND -27, stacked onto 64 2218 ND 3' Hatchet 2219 ND 3' Hammerhead ribozyme (Lior Nissim, Timothy Lu) 2220 ND 5'Hatchet 2221 ND 5'HDV ribozyme (Lior Nissim, Timothy Lu) 2222 ND 5'Hammcrhcad ribozyme (Lior Nissim. Timothy Lu) 2223 ND 3' HH15 Minimal Hammerhead ribozyme 2224 ND 5' RBMX recruiting motif 2225 ND 3' Hammerhead ribozyme (Lior Nissim, Timothy Lu) smaller scar 2226 ND 3' env25 pistol ribozyme (with an added CUUCGG loop) 2227 ND 3' Env-9 Twister 2228 ND =+AUUAUCUCAUUACU25 2229 ND 5'Env-9 Twister 2230 ND 3' Twisted Sister 1 15 05 25 SEQ ID NO: NAME NUCLEOTIDE SEQUENCE OR DESCRIPTION OF MODIFICATION 2231 ND no stem 2232 ND 5'HH15 Minimal Hammerhead ribozyme 2233 ND 5'Hammerhead ribozyme (Lior Nissim, Timothy Lu) guide scaffold scar 2234 ND 5'Twisted Sister 1 2235 ND 5'sTRSV WT viral Hammerhead ribozy me 2236 ND 148, =+G55, stacked onto 64 2237 ND 158, 103+148(+G55) -99, G65U 2238 174 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGUGGGUAA AGOUCCCUCUUCGGAGGGAGCAUCAAAG 2239 175 ACUGGCGCCUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUGGGUAAA GCGCUUACGGACUUCGGUCCGUAAGAAGCAUCAAAG 2240 176 GCUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGUGGGUAA AGCUCCCUCUUCGGAGGGAGCAUCAAAG 2241 177 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUGGGUAAA GCUCCCUCUUCGGAGGGAGCAUCAAAG 2242 181 ACUGGCGCCUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUGGGUAAA GCGCUUACGGACUUCGGUCCGUAAGAAGCAUCAAAG 2243 182 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUGGGUAAA GCGCUUACGGACUUCGGUCCGUAAGAAGCAUCAAAG 2244 183 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGUGGGUAA AGCGCUUACGGACUUCGGUCCGUAAGAAGCAUCAAAG 2245 184 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUUGGGUAA AGCUCCCUCUUCGGAGGGAGCAUCAAAG 2246 185 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUUGGGUAA AGCGCUUACGGACUUCGGUCCGUAAGAAGCAUCAAAG 2247 186 ACUGGCGCCUUUAUCAUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUGGGUAA AGCGCUUACGGACUUCGGUCCGUAAGAAGCAUCAAAG 2248 187 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGUGGGUAA AGCGCCCUCUUCGGAGGGAAGCAUCAAAG 2249 188 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGUGGGUAA AGCUCACAUGAGGAUCACCCAUGUGAGCAUCAAAG 2250 189 ACUGGCACUUUUACCUGAUUACUUUGAGAGCCAACACCAGCGACUAUGUCGUAGUGGGUAA AGCUCCCUCUUCGGAGGGAGCAUCAAAG 2251 190 ACUGGCACUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGUGGGUAA AGCUCCCUCUUCGGAGGGAGCAUCAAAG 2252 191 ACUGGCCCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGUGGGUAA AGCUCCCUCUUCGGAGGGAGCAUCAAAG 2253 192 ACUGGCGCUUUUACCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGUGGGUAA AGCUCCCUCUUCGGAGGGAGCAUCAAAG 2254 193 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAACACCAGCGACUAUGUCGUAGUGGGUAA AGCUCCCUCUUCGGAGGGAGCAUCAAAG 2255 195 ACUGGCACCUUUACCUGAUUACUUUGAGAGCCAACACCAGCGACUAUGUCGUAUGGGUAAA GCGCUUACGGACUUCGGUCCGUAAGAAGCAUCAAAG 2256 196 ACUGGCACCUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUGGGUAAA GCGCUUACGGACUUCGGUCCGUAAGAAGCAUCAAAG 2257 197 ACUGGCCCCUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUGGGUAAA GCGCUUACGGACUUCGGUCCGUAAGAAGCAUCAAAG 2258 198 ACUGGCGCCUUUAUCUGAUUACUUUGAGAGCCAACACCAGCGACUAUGUCGUAUGGGUAAA GCGCUUACGGACUUCGGUCCGUAAGAAGCAUCAAAG 2259 199 GCUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGUGGGUAA AGCUCCCUCUUCGGAGGGAGCAUCAAAG 15 05 25 SEQ ID NO: NAME NUCLEOTIDE SEQUENCE OR DESCRIPTION OF MODIFICATION 2260 200 GACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGUGGGUA AAGCUCCCUCUUCGGAGGGAGCAUCAAAG 2261 201 ACUGGCGCCUUUAUCUGAUUACUUUGGAGAGCCAUCACCAGCGACUAUGUCGUAGUGGGUA AAGCUCCCUCUUCGGAGGGAGCAUCAAAG 2262 202 ACUGGCGCAUUUAUCUGAUUACUUUGUGAGCCAUCACCAGCGACUAUGUCGUAGUGGGUAA AGCU C C CU CUU C GGAGGGAGCAU CAAAG 2263 203 ACUGGCGCCUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGUGGGUAA AGCUCCCUCUUCGGAGGGAGCAUCAAAG 2264 204 ACUGGCGCUUUUAUCUGAUUACUUUGGAGAGCCAUCACCAGCGACUAUGUCGUAGUGGGUA AAGCUCCCUCUUCGGAGGGAGCAUCAAAG 2265 205 ACUGGCGCAUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGUGGGUAA AGCU C C CU CUU C GGAGGGAGCAU CAAAG 2266 206 ACUGGCGCUUUUAUCUGAUUACUUUGUGAGCCAUCACCAGCGACUAUGUCGUAGUGGGUAA AGCUCCCUCUUCGGAGGGAGCAUCAAAG 2267 207 ACUGGCGCUUUUAUUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGUGGGUA AAGCUCCCUCUUCGGAGGGAGCAUCAAAG 2268 208 ACGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGUGGGUAAA GCUCCCUCUUCGGAGGGAGCAUCAAAG 2269 209 ACUGGCGCUUUUAUAUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGUGGGUAA AGCUCCCUCUUCGGAGGGAGCAUCAAAG 2270 210 ACUGGCGCUUUUAUCUUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGUGGGUA AAGCUCCCUCUUCGGAGGGAGCAUCAAAG 2271 211 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAGCACCAGCGACUAUGUCGUAGUGGGUAA AGCUCCCUCUUCGGAGGGAGCAUCAAAG 2272 212 ACUGGCGCUGUUAUCUGAUUACUUCGAGAGCCAUCACCAGCGACUAUGUCGUAGUGGGUAA AGCUCCCUCUUCGGAGGGAGCAUCGAAG 2273 213 ACUGGCGCUCUUAUCUGAUUACUUCGAGAGCCAUCACCAGCGACUAUGUCGUAGUGGGUAA AGCUCCCUCUUCGGAGGGAGCAUCGAAG 2274 214 ACUGGCGCUUGUAUCUGAUUACUCUGAGAGCCAUCACCAGCGACUAUGUCGUAGUGGGUAA AGCUCCCUCUU C GGAGGGAGCAU CAGAG 2275 215 ACUGGCGCUUCUAUCUGAUUACUCUGAGAGCCAUCACCAGCGACUAUGUCGUAGUGGGUAA AGCUCCCUCUUCGGAGGGAGCAUCAGAG 2276 216 ACUGGCGCUUUGAUCUGAUUACCUUGAGAGCCAUCACCAGCGACUAUGUCGUAGUGGGUAA AGCUCCCUCUUCGGAGGGAGCAUCAAGG 2277 217 ACUGGCGCUUUCAUCUGAUUACCUUGAGAGCCAUCACCAGCGACUAUGUCGUAGUGGGUAA AGCUCCCUCUUCGGAGGGAGCAUCAAGG 2278 218 ACUGGCGCUGUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGUGGGUAA AGCUCCCUCUUCGGAGGGAGCAUCAAAG 2279 219 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGUGGGUAA AGCUCCCUCUUCGGAGGGAGCAUCGAAG 2280 220 ACUGGCGCUUUUAUCUGAUUACUUCGAGAGCCAUCACCAGCGACUAUGUCGUAGUGGGUAA AGCUCCCUCUUCGGAGGGAGCAUCAAAG 2281 221 ACUGGCACUUCUAUCUGAUUACUCUGAGAGCCAUCACCAGCGACUAUGUCGUAUGGGUAAA GCCGCUUACGGACUUCGGUCCGUAAGAGGCAUCAGAG 2282 222 ACUGGCACUUCUAUCUGAUUACUCUGAGAGCCAUCACCAGCGACUAUGUCGUAGUGGGUAA AGCUCCCUCUU C GGAGGGAGCAU CAGAG 2283 223 ACUGGCACCUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUGGGUAAA GCCGCUUACGGACUUCGGUCCGUAAGAGGCAUCAAAG 2284 224 ACUGGCACUUGUAUCUGAUUACUCUGAGAGCCAUCACCAGCGACUAUGUCGUAUGGGUAAA GCCGCUUACGGACUUCGGUCCGUAAGAGGCAUCAGAG 2285 225 ACUGGCACUUGUAUCUGAUUACUCUGAGAGCCAUCACCAGCGACUAUGUCGUAGUGGGUAA AGCUCCCUCUUCGGAGGGAGCAUCAGAG 15 05 25 SEQ ID NO: NAME NUCLEOTIDE SEQUENCE OR DESCRIPTION OF MODIFICATION 27219 226 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGUGGGUAA AGCUGCACUAUGGGCGCAGCGUCAAUGACGCUGACGGUACAGGCCAGACAAUUAUUGUCUG GUAUAGUGCAGCAUCAAAG 26794 229 ACUGGCACUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUGGGUAAA GCGCUUACGGACUUCGGUCCGUAAGAAGCAUCAAAG 26795 230 ACUGGCACUUCUAUCUGAUUACUCUGAGAGCCAUCACCAGCGACUAUGUCGUAUGGGUAAA GCGCUUACGGACUUCGGUCCGUAAGAAGCAUCAGAG 26796 231 ACUGGCGCUUCUAUCUGAUUACUCUGAGAGCCAUCACCAGCGACUAUGUCGUAUGGGUAAA GCCGCUUACGGACUUCGGUCCGUAAGAGGCAUCAGAG 26797 232 ACUGGCACUUCUAUCUGAUUACUCUGAGCGCCAUCACCAGCGACUAUGUCGUAUGGGUAAA GCCGCUUACGGACUUCGGUCCGUAAGAGGCAUCAGAG 26798 233 ACUGGCGCUUCUAUCUGAUUACUCUGAGCGCCAUCACCAGCGACUAUGUCGUAUGGGUAAA GCCGCUUACGGACUUCGGUCCGUAAGAGGCAUCAGAG 26799 234 ACUGGCGCUUCUAUCUGAUUACUCUGAGCGCCAUCACCAGCGACUAUGUCGUAUGGGUAAA GCGCCUUACGGACUUCGGUCCGUAAGGAGCAUCAGAG 26800 235 ACUGGCGCUUCUAUCUGAUUACUCUGAGCGCCAUCACCAGCGACUAUGUCGUAGUGGGUAA AGCCGCUUACGGACUUCGGUCCGUAAGAGGCAUCAGAG 26801 236 ACGGGACUUUCUAUCUGAUUACUCUGAAGUCCCUCACCAGCGACUAUGUCGUAUGGGUAAA GCCGCUUACGGACUUCGGUCCGUAAGAGGCAUCAGAG 26802 237 ACCUGUAGUUCUAUCUGAUUACUCUGACUACAGUCACCAGCGACUAUGUCGUAUGGGUAAA GCCGCUUACGGACUUCGGUCCGUAAGAGGCAUCAGAG 26803 238 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGUGGGUAA AGCUGCACGGUGGGCGCAGCUUCGGCUGACGGUACACCGUGCAGCAUCAAAG 26804 239 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGUGGGUAA AGCUGCACGGUGGGCGCAGCUUCGGCUGACGGUACACCGGUGGGCGCAGCUUCGGCUGACG GUACACCGUGCAGCAUCAAAG 26805 240 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGUGGGUAA AGCUGCACGGUGGGCGCAGCUUCGGCUGACGGUACACCGGUGGGCGCAGCUUCGGCUGACG GUACACCGGUGGGCGCAGCUUCGGCUGACGGUACACCGUGCAGCAUCAAAG 26806 241 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGUGGGUAA AGCUGCACGGUGGGCGCAGCUUCGGCUGACGGUACACCGGUGGGCGCAGCUUCGGCUGACG GUACACCGGUGGGCGCAGCUUCGGCUGACGGUACACCGGUGGGCGCAGCUUCGGCUGACGG UACACCGUGCAGCAUCAAAG 26807 242 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGUGGGUAA AGCUGCACGGUGGGCGCAGCUUCGGCUGACGGUACACCGGUGGGCGCAGCUUCGGCUGACG GUACACCGGUGGGCGCAGCUUCGGCUGACGGUACACCGGUGGGCGCAGCUUCGGCUGACGG UACACCGGUGGGCGCAGCUUCGGCUGACGGUACACCGUGCAGCAUCAAAG 26808 243 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGUGGGUAA AGCUGCACCUAGCGGAGGCUAGGUGCAGCAUCAAAG 26809 244 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGUGGGUAA AGCUGCACCUCGGCUUGCUGAAGCGCGCACGGCAAGAGGCGAGGUGCAGCAUCAAAG 26810 245 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGUGGGUAA AGCUGCACCUCUCUCGACGCAGGACUCGGCUUGCUGAAGCGCGCACGG»AGAGGCGAGGG GCb-GCGACU GLrUGAGUACGCCAAAAAUUUUGACOA^CGGAGGCUAGAAGGAGAGAGGU gca. GCAUCAAAG 26811 246 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGUGGGUAA AGCUGCACGGUGCCCGUCUGUUGUGUCGAGAGACGCCAAAAAUUUUGACUAGCGGAGGCUA GAAGGAGAGAGAUGGGUGCCGUGCAGCAUCAAAG 26812 247 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGUGGGUAA AGCUGCACAUGGAGAGGAGAUGUGCAGCAUCAAAG 26813 248 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGUGGGUAA AGCUGCACAUGGAGAUGUGCAGCAUCAAAG 26814 249 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGUGGGUAA AGCUUGGGCGCAGCGUCAAUGACGCUGACGGUACAAGCAUCAAAG 15 05 25 SEQ ID NO: NAME NUCLEOTIDE SEQUENCE OR DESCRIPTION OF MODIFICATION 26815 250 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGUGGGUAA AGCUGCACUAUGGGCGCAGCGUCAAUGACGCUGACGGUACAGGCCACAUGAGGAUCACCCA UGUGGUAUAGUGCAGCAUCAAAG 26816 251 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGUGGGUAA AGCUGCACUAUGGGCGCAGCUCAUGAGGAUCACCCAUGAGCUGACGGUACAGGCCACAUGA GGAU CAC C CAU GU G GUAUAGU G CAG CAU CAAAG 26817 252 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGUGGGUAA AGCUGCACUAUGGGCGCAGCGUCAAUGACGCUGACGGUACAGGCCACAUGGCAGUCGUAAC GACGCGGGUGGUAUAGUGCAGCAUCAAAG 26818 253 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGUGGGUAA AGCUGCACUAUGGGCGCAGCAAACAUGGCAGUCCUAAGGACGCGGGUUUUGCUGACGGUAC AGGCCACAUGGCAGUCGUAACGACGCGGGUGGUAUAGUGCAGCAUCAAAG 26819 254 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGUGGGUAA AGCUGCACUAUGGGCGCAGACAUGGCAGUCGUAACGACGCGGGUCUGACGGUACAGGCCAC AUGAGGAUCACCCAUGUGGUAUAGUGCAGCAUCAAAG 26820 255 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGUGGGUAA AGCUGCACUAAGGAGUUUAUAUGGAAACCCUUAGUGCAGCAUCAAAG 26821 256 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGUGGGUAA AGCUCAGGAAGCACUAUGGGCGCAGCGUCAAUGACGCUGACGGUACAGGCCAGACAAUUAU UGUCUGGUAUAGUGCAGCAGCAGAACAAUUUGCUGAGGGCUAUUGAGGCGCAACAGCAUCU GUUGCAACUCACAGUCUGGGGCAUCAAGCAGCUCCAGGCAAGAAUCCUGAGCAUCAAAG 26822 257 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGUGGGUAA AGCUGCACGCCCUGAAGAAGGGCGUGCAGCAUCAAAG 26823 258 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGUGGGUAA AGCUGCACGGCUCGUGUAGCUCAUUAGCUCCGAGCCGUGCAGCAUCAAAG 26824 259 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGUGGGUAA AGCUGCACCCGUGUGCAUCCGCAGUGUCGGAUCCACGGGUGCAGCAUCAAAG 26825 260 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGUGGGUAA AGCUGCACGGAAUCCAUUGCACUCCGGAUUUCACUAGGUGCAGCAUCAAAG 26826 261 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGUGGGUAA AGCUGCACAUGCAUGUCUAAGACAGCAUGUGCAGCAUCAAAG 26827 262 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGUGGGUAA AGCU GCACAAAACAUAAGGAAAAC CUAU GUU GU GCAGCAU CAAAG 26828 263 ACUGGCGCUUCUAUCUGAUUACUCUGAGCGCCAUCACCAGCGACUAUGUCGUAGUGGGUAA AGCCGCUUACGGACUAUGGGCGCAGCGUCAAUGACGCUGACGGUACAGGCCAGACAAUUAU UGUCUGGUAUAGUCCGUAAGAGGCAUCAGAG 26829 264 ACUGGCGCUUCUAUCUGAUUACUCUGAGCGCCAUCACCAGCGACUAUGUCGUAGUGGGUAA AGCCGCUUACGGGUGGGCGCAGCGUCAAUGACGCUGACGGUACAGGCCAGACAAUUAUUGU CU GGUAC C C GUAAGAGGCAU CAGAG 26830 265 ACUGGCGCUUCUAUCUGAUUACUCUGAGCGCCAUCACCAGCGACUAUGUCGUAGUGGGUAA AGCCGCUUACGGUAUGGGCGCAGCGUCAAUGACGCUGACGGUACAGGCCACAUGAGGAUCA CCCAUGUGGUAUACCGUAAGAGGCAUCAGAG 26831 266 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGUGGGUAA AGCUCCCUAUGGGCGCAGCGUCAAUGACGCUGACGGUACAGGCCACAUGAGGAUCACCCAU GUGGUAUAGGGAGCAUCAAAG 26832 267 ACUGGCGCUUCUAUCUGAUUACUCUGAGCGCCAUCACCAGCGACUAUGUCGUAGUGGGUAA AGCCGCUUACGGUAUGGGCGCAGCUCAUGAGGAUCACCCAUGAGCUGACGGUACAGGCCAC AU GAGGAU GAG C CAUGUGGUAUAC C GUAAGAGGCAUCAGAG 26833 268 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGUGGGUAA AGCUCCCUAUGGGCGCAGCUCAUGAGGAUCACCCAUGAGCUGACGGUACAGGCCACAUGAG GAU CAC C CAU GU G GUAUAGG GAG CAU CAAAG 26834 269 ACUGGCGCUUCUAUCUGAUUACUCUGAGCGCCAUCACCAGCGACUAUGUCGUAGUGGGUAA AGCCGCUUACGGUAUGGGCGCAGCGUCAAUGACGCUGACGGUACAGGCCACAUGGCAGUCG UAACGACGCGGGUGGUAUACCGUAAGAGGCAUCAGAG 15 05 25 SEQ ID NO: NAME NUCLEOTIDE SEQUENCE OR DESCRIPTION OF MODIFICATION 26835 270 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGUGGGUAA AGCUCCCUAUGGGCGCAGCGUCAAUGACGCUGACGGUACAGGCCACAUGGCAGUCGUAACG ACGCGGGUGGUAUAGGGAGCAUCAAAG 26836 271 ACUGGCGCUUCUAUCUGAUUACUCUGAGCGCCAUCACCAGCGACUAUGUCGUAGUGGGUAA AGCCGCUUACGGUAUGGGCGCAGCAAACAUGGCAGUCCUAAGGACGCGGGUUUUGCUGACG GUACAGGCCACAUGGCAGUCGUAACGACGCGGGUGGUAUACCGUAAGAGGCAUCAGAG 26837 272 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGUGGGUAA AGCUCCCUAUGGGCGCAGCAAACAUGGCAGUCCUAAGGACGCGGGUUUUGCUGACGGUACA GGCCACAUGGCAGUCGUAACGACGCGGGUGGUAUAGGGAGCAUCAAAG 26838 273 ACUGGCGCUUCUAUCUGAUUACUCUGAGCGCCAUCACCAGCGACUAUGUCGUAGUGGGUAA AGCCGCUUACGGUAUGGGCGCAGACAUGGCAGUCGUAACGACGCGGGUCUGACGGUACAGG CCACAUGAGGAUCACCCAUGUGGUAUACCGUAAGAGGCAUCAGAG 26839 274 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGUGGGUAA AGCUCCCUAUGGGCGCAGACAUGGCAGUCGUAACGACGCGGGUCUGACGGUACAGGCCACA UGAGGAUCACCCAUGUGGUAUAGGGAGCAUCAAAG 27220 275 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCG UAGUGGGUAAAGCUGCACUAUGGGCGCAGCACCUGAGGAUCACCCAGGUGC UGACGGUACAGGCCACCUGAGGAUCACCCAGGUGGUAUAGUGCAGCAUCAA AG 27221 276 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCG UAGUGGGUAAAGCUGCACUAUGGGCGCAGCGCAUGAGGAUCACCCAUGCGC UGACGGUACAGGCCGCAUGAGGAUCACCCAUGCGGUAUAGUGCAGCAUCAA AG 27222 277 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCG UAGUGGGUAAAGCUGCACUAUGGGCGCAGCGCCUGAGGAUCACCCAGGCGC UGACGGUACAGGCCGCCUGAGGAUCACCCAGGCGGUAUAGUGCAGCAUCAA AG 27223 278 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCG UAGUGGGUAAAGCUGCACUAUGGGCGCAGCGCCUGAGCAUCAGCCAGGCGC UGACGGUACAGGCCGCCUGAGCAUCAGCCAGGCGGUAUAGUGCAGCAUCAA AG 27224 279 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCG UAGUGGGUAAAGCUGCACUAUGGGCGCAGCACAUGAGCAUCAGCCAUGUGC UGACGGUACAGGCCACAUGAGCAUCAGCCAUGUGGUAUAGUGCAGCAUCAA AG 27225 280 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCG UAGUGGGUAAAGCUGCACUAUGGGCGCAGCACAUGAGUAUCAACCAUGUGC UGACGGUACAGGCCACAUGAGUAUCAACCAUGUGGUAUAGUGCAGCAUCAA AG 27226 281 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCG UAGUGGGUAAAGCUGCACUAUGGGCGCAGCACAUGAGAAUCAGCCAUGUGC UGACGGUACAGGCCACAUGAGAAUCAGCCAUGUGGUAUAGUGCAGCAUCAA AG 27227 282 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCG UAGUGGGUAAAGCUGCACUAUGGGCGCAGCCCUUGAGGAUCACCCAUGUGC UGACGGUACAGGCCCCUUGAGGAUCACCCAUGUGGUAUAGUGCAGCAUCAA AG 15 05 25 SEQ ID NO: NAME NUCLEOTIDE SEQUENCE OR DESCRIPTION OF MODIFICATION 27228 283 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCG UAGUGGGUAAAGCUGCACUAUGGGCGCAGCACUUGAGGAUCACCCAUGUGC UGACGGUACAGGCCACUUGAGGAUCACCCAUGUGGUAUAGUGCAGCAUCAA AG 27229 284 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCG UAGUGGGUAAAGCUGCACUAUGGGCGCAGCACCUGAGGAUCACCCAUGUGC UGACGGUACAGGCCACCUGAGGAUCACCCAUGUGGUAUAGUGCAGCAUCAA AG 27230 285 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGU GGGUAAAGCUGCACUAUGGGCGCAGCACAUGAGGAUCACCUAUGUGCUGACGGUA CAGGCCACAUGAGGAUCACCUAUGUGGUAUAGUGCAGCAUCAAAG 27231 286 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGU GGGUAAAGCUGCACUAUGGGCGCAGCACAUUAGGAUCACCAAUGUGCUGACGGUA CAGGCCACAUUAGGAUCACCAAUGUGGUAUAGUGCAGCAUCAAAG 27232 287 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGU GGGUAAAGCUGCACUAUGGGCGCAGCACAUUAGGAUCACCGAUGUGCUGACGGUA CAGGCCACAUUAGGAUCACCGAUGUGGUAUAGUGCAGCAUCAAAG 27233 288 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGU GGGUAAAGCUGCACUAUGGGCGCAGCACAUUAGGAUCACCUAUGUGCUGACGGUA CAGGCCACAUUAGGAUCACCUAUGUGGUAUAGUGCAGCAUCAAAG 27234 289 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGU GGGUAAAGCUGCACUAUGGGCGCAGCACAUGAGGAUUACCCAUGUGCUGACGGUA CAGGCCACAUGAGGAUUACCCAUGUGGUAUAGUGCAGCAUCAAAG 27235 290 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGU GGGUAAAGCUGCACUAUGGGCGCAGCACAUGAGGAUAACCCAUGUGCUGACGGUA CAGGCCACAUGAGGAUAACCCAUGUGGUAUAGUGCAGCAUCAAAG 27236 291 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGU GGGUAAAGCUGCACUAUGGGCGCAGCACAUGAGGAUGACCCAUGUGCUGACGGUA CAGGCCACAUGAGGAUGACCCAUGUGGUAUAGUGCAGCAUCAAAG 27237 292 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGU GGGUAAAGCUGCACUAUGGGCGCAGCACAUGAGGACCACCCAUGUGCUGACGGUA CAGGCCACAUGAGGACCACCCAUGUGGUAUAGUGCAGCAUCAAAG 27238 293 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGU GGGUAAAGCUGCACUAUGGGCGCAGCAGAUGAGGAUCACCCAUGGGCUGACGGUA CAGGCCAGAUGAGGAUCACCCAUGGGGUAUAGUGCAGCAUCAAAG 27239 294 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGU GGGUAAAGCUGCACUAUGGGCGCAGCACAUGGGGAUCACCCAUGUGCUGACGGUA CAGGCCACAUGGGGAUCACCCAUGUGGUAUAGUGCAGCAUCAAAG 27240 295 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGU GGGUAAAGCUGCACUAUGGGCGCAGCACAUGAGGAUCACCCAUGUGCUGACGGUA CAGGCCACAUGAGGAUCACCCAUGUGGUAUAGUGCAGCAUCAAAG 27241 296 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGU GGGUAAAGCUCACCUGAGGAUCACCCAGGUGAGCAUCAAAG 27242 297 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGU GGGUAAAGCUCGCAUGAGGAUCACCCAUGCGAGCAUCAAAG 27243 298 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGU GGGUAAAGCUCGCCUGAGGAUCACCCAGGCGAGCAUCAAAG 27244 299 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGU GGGUAAAGCUCGCCUGAGCAUCAGCCAGGCGAGCAUCAAAG 15 05 25 SEQ ID NO: NAME NUCLEOTIDE SEQUENCE OR DESCRIPTION OF MODIFICATION 27245 300 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGU GGGUAAAGCUCACAUGAGCAUCAGCCAUGUGAGCAUCAAAG 27246 301 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGU GGGUAAAGCUCACAUGAGUAUCAACCAUGUGAGCAUCAAAG 27247 302 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGU GGGUAAAGCUCACAUGAGAAUCAGCCAUGUGAGCAUCAAAG 27248 303 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGU GGGUAAAGCUCCCUUGAGGAUCACCCAUGUGAGCAUCAAAG 27249 304 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGU GGGUAAAGCUCACUUGAGGAUCACCCAUGUGAGCAUCAAAG 27250 305 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGU GGGUAAAGCUCACCUGAGGAUCACCCAUGUGAGCAUCAAAG 27251 306 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGU GGGUAAAGCUCACAUGAGGAUCACCUAUGUGAGCAUCAAAG 27252 307 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGU GGGUAAAGCUCACAUUAGGAUCACCAAUGUGAGCAUCAAAG 27253 308 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGU GGGUAAAGCUCACAUUAGGAUCACCGAUGUGAGCAUCAAAG 27254 309 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGU GGGUAAAGCUCACAUUAGGAUCACCUAUGUGAGCAUCAAAG 27255 310 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGU GGGUAAAGCUCACAUGAGGAUUACCCAUGUGAGCAUCAAAG 27256 311 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGU GGGUAAAGCUCACAUGAGGAUAACCCAUGUGAGCAUCAAAG 27257 312 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGU GGGUAAAGCUCACAUGAGGAUGACCCAUGUGAGCAUCAAAG 27258 313 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGU GGGUAAAGCUCACAUGAGGACCACCCAUGUGAGCAUCAAAG 27259 314 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGU GGGUAAAGCUCAGAUGAGGAUCACCCAUGGGAGCAUCAAAG 27260 315 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGU GGGUAAAGCUCACAUGGGGAUCACCCAUGUGAGCAUCAAAG 27261 317 ACUGGCGCUUCUAUCUGAUUACUCUGAGCGCCAUCACCAGCGACUAUGUCGUAGUGGGUAA AGCUCACAUGAGGAUCACCCAUGUGAGCAUCAGAG 27262 318 ACUGGCGCUUCUAUCUGAUUACUCUGAGCGCCAUCACCAGCGACUAUGUCGUAGUGGGUAA AGCUGCACUAUGGGCGCAGCGUCAAUGACGCUGACGGUACAGGCCACAUGAGGAUCACCCA UGUGGUAUAGUGCAGCAUCAGAG 27263 319 ACUGGCGCUUCUAUCUGAUUACUCUGAGCGCCAUCACCAGCGACUAUGUCGUAGUGGGUAA AGCUGCACUAUGGGCGCAGCUCAUGAGGAUCACCCAUGAGCUGACGGUACAGGCCACAUGA GGAU CAC C CAU GU GGUAUAGU G CAG CAU CAGAG 27264 320 ACUGGCGCUUCUAUCUGAUUACUCUGAGCGCCAUCACCAGCGACUAUGUCGUAGUGGGUAA AGCUGCACUAUGGGCGCAGACAUGGCAGUCGUAACGACGCGGGUCUGACGGUACAGGCCAC AUGAGGAUCACCCAUGUGGUAUAGUGCAGCAUCAGAG 27265 321 ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGU GGGUAAAGCUGCACUAUGGGGCCACAUGAGGAUCACCCAUGUGGUGUACAGCGCA GCGUCAAUGACGCUGACGAUAGUGCAGCAUCAAAG

[0124] In some embodiments, a sgRNA variant comprises one or more additional changes to a sequence of SEQ ID NO:2238, SEQ ID NO:2239, SEQ ID NO:2240, SEQ ID NO:2241, SEQ ID 15 05 25 NO:2243, SEQ ID NO:2256, SEQ ID NO:2274, SEQ ID NO:2275, SEQ ID NO:2279, SEQ ID NO:2281, SEQ ID NO: 2285, SEQ ID NO: 26797, or SEQ ID NO: 26800 of Table 3.

[0125] In some embodiments of the gRNA variants of the disclosure, the gRNA variant comprises at least one modification, wherein the at least one modification compared to the reference guide scaffold of SEQ ID NO:5 is selected from one or more of: (a) a C18G substitution in the triplex loop; (b) a G55 insertion in the stem bubble; (c) a UI deletion; (d) a modification of the extended stem loop wherein (i) a 6 nt loop and 13 loop-proximal base pairs are replaced by a Uvsx hairpin; and (ii) a deletion of A99 and a substitution of G65U that results in a loop-distal base that is fully base-paired. In exemplary embodiments of the foregoing, the gRNA variant comprises the sequence of any one of SEQ ID NOS: 2238, 2241, 2244, 2248, 2249, 2256, 2259-2285, 26797 or 26800.

[0126] In some embodiments, a gRNA variant comprises an exogenous stem loop having a long non-coding RNA (IncRNA). As used herein, a IncRNA refers to a non-coding RNA that is longer than approximately 200 bp in length. In some embodiments, the 5' and 3' ends of the exogenous stem loop are base paired; i.e., interact to form a region of duplex RNA. In some embodiments, the 5' and 3' ends of the exogenous stem loop are base paired, and one or more regions between the 5' and 3' ends of the exogenous stem loop are not base paired.

[0127] In some embodiments, the disclosure provide gRNA variants with nucleotide modifications relative to reference gRNA having: (a) substitution of 1 to 15 consecutive or non-consecutive nucleotides in the gRNA variant in one or more regions; (b) a deletion of 1 to 10 consecutive or non-consecutive nucleotides in the gRNA variant in one or more regions; (c) an insertion of 1 to 10 consecutive or non-consecutive nucleotides in the gRNA variant in one or more regions; (d) a substitution of the scaffold stem loop or the extended stem loop with an RNA stem loop sequence from a heterologous RNA source with proximal 5' and 3' ends; or any combination of (a)-(d). Any of the substitutions, insertions and deletions described herein can be combined to generate a gNA variant of the disclosure. For example, a gNA variant can comprise at least one substitution and at least one deletion relative to a reference gRNA, at least one substitution and at least one insertion relative to a reference gRNA, at least one insertion and at least one deletion relative to a reference gRNA, or at least one substitution, one insertion and one deletion relative to a reference gRNA.

[0128] In some embodiments, a sgRNA variant of the disclosure comprises one or more additional changes to a previously generated variant, the previously generated variant itself 15 05 25 serving as the sequence to be modified. In some embodiments, a sgRNA variant comprises one or more additional changes to a sequence of SEQ ID NO: 2238, SEQ ID NO: 2239, SEQ ID NO: 2240, SEQ ID NO: 2241, SEQ ID NO:2241, SEQ ID NO:2274, SEQ ID NO:2275, SEQ ID NO: 2279, or SEQ ID NO: 2285, SEQ ID NO: 26797, or SEQ ID NO: 26800.

[0129] In exemplary embodiments, a gRNA variant comprises one or more modification relative to gRNA scaffold variant 174 (SEQ ID NO:2238), wherein the resulting gRNA variant exhibits a functional improvement compared to the parent 174, when assessed in an in vitro or in vivo assay under comparable conditions.

[0130] In exemplary embodiments, a gRNA variant comprises one or more modification relative to gRNA scaffold variant 175 (SEQ ID NO:2239), wherein the resulting gRNA variant exhibits a functional improvement compared to the parent 174, when assessed in an in vitro or in vivo assay under comparable conditions.

[0131] In exemplary embodiments, a gRNA variant comprises one or more modification relative to gRNA scaffold variant 215 (SEQ ID NO:2275), wherein the resulting gRNA variant exhibits a functional improvement compared to the parent 215, when assessed in an in vitro or in vivo assay under comparable conditions.

[0132] In exemplary embodiments, a gRNA variant comprises one or more modification relative to gRNA scaffold variant 221 (SEQ ID NO: 2281), wherein the resulting gRNA variant exhibits a functional improvement compared to the parent 221, when assessed in an in vitro or in vivo assay under comparable conditions.

[0133] In exemplary embodiments, a gRNA variant comprises one or more modification relative to gRNA scaffold variant 225 (SEQ ID NO: 2285), wherein the resulting gRNA variant exhibits a functional improvement compared to the parent 225, when assessed in an in vitro or in vivo assay under comparable conditions.

[0134] In exemplary embodiments, a gRNA variant comprises one or more modification relative to gRNA scaffold variant 235 (SEQ ID NO: 26800), wherein the resulting gRNA variant exhibits a functional improvement compared to the parent 225, when assessed in an in vitro or in vivo assay under comparable conditions.

[0135] In some embodiments, the gRNA variant comprises an exogenous extended stem loop, with such differences from a reference gRNA described as follows. In some embodiments, an exogenous extended stem loop has little or no identity to the reference stem loop regions disclosed herein (e.g., SEQ ID NO: 15). In some embodiments, an exogenous stem loop is at 15 05 25 least 10 bp, at least 20 bp, at least 30 bp, at least 40 bp, at least 50 bp, at least 60 bp, at least 70 bp, at least 80 bp, at least 90 bp, at least 100 bp, at least 200 bp, at least 300 bp, at least 400 bp, at least 500 bp, at least 600 bp, at least 700 bp, at least 800 bp, at least 900 bp, at least 1,000 bp, at least 2,000 bp, at least 3,000 bp, at least 4,000 bp, at least 5,000 bp, at least 6,000 bp, at least 7,000 bp, at least 8,000 bp, at least 9,000 bp, at least 10,000 bp, at least 12,000 bp, at least 15,000 bp or at least 20,000 bp. In some embodiments, the gRNA variant comprises an extended stem loop region comprising at least 10, at least 100, at least 500, at least 1000, or at least 10,000 nucleotides. In some embodiments, the heterologous stem loop increases the stability of the gRNA. In some embodiments, the heterologous RNA stem loop is capable of binding a protein, an RNA structure, a DNA sequence, or a small molecule. In some embodiments, an exogenous stem loop region replacing the stem loop comprises an RNA stem loop or hairpin in which the resulting gRNA has increased stability and, depending on the choice of loop, can interact with certain cellular proteins or RNA. Such exogenous extended stem loops can comprise, for example a thermostable RNA such as MS2 hairpin (ACAUGAGGAUCACCCAUGU (SEQ ID NO: 27)), Qp hairpin (UGCAUGUCUAAGACAGCA (SEQ ID NO: 28)), UI hairpin II (AAUCCAUUGCACUCCGGAUU (SEQ ID NO: 29)), Uvsx (CCUCUUCGGAGG (SEQ ID NO: 30)), PP7 hairpin (AGGAGUUUCUAUGGAAACCCU (SEQ ID NO: 31)), Phage replication loop (AGGUGGGACGACCUCUCGGUCGUCCUAUCU (SEQ ID NO: 32)), Kissing loop a (UGCUCGCUCCGUUCGAGCA (SEQ ID NO: 33)), Kissing loop bl (UGCUCGACGCGUCCUCGAGCA (SEQ ID NO: 34)), Kissing Ioop b2 (UGCUCGUUUGCGGCUACGAGCA (SEQ ID NO: 35)), G quadriplex M3q (AGGGAGGGAGGGAGAGG (SEQ ID NO: 149)), G quadriplex telomere basket (GGUUAGGGUUAGGGUUAGG (SEQ ID NO: 150)), Sarcin-ricin loop (CUGCUCAGUACGAGAGGAACCGCAG (SEQ ID NO: 151)) or Pseudoknots (UACACUGGGAUCGCUGAAUUAGAGAUCGGCGUCCUUUCAUUCUAUAUACUUUGG AGUUUUAAAAUGUCUCUAAGUACA (SEQ ID NO: 152)). In some embodiments, one of the foregoing hairpin sequences is incorporated into the stem loop to help traffic the incorporation of the gRNA (and an associated CasX in an RNP complex) into a budding XDP (described more fully, below).

[0136] In the embodiments of the gRNA variants, the gRNA variant further comprises a spacer (or targeting sequence) region located at the 3’ end of the gRNA, capable of hybridizing with a target nucleic acid specific to a DMPK sequence described more fully, supra, which 15 05 25 comprises at least 14 to about 35 nucleotides wherein the spacer is designed with a sequence that is complementary to a target DNA. In some embodiments, the encoded gRNA variant comprises a targeting sequence of at least 10 to 20 nucleotides complementary to a target DNA. In some embodiments, the targeting sequence has 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34 or 35 nucleotides. In some embodiments, the encoded gRNA variant comprises a targeting sequence having 20 nucleotides. In some embodiments, the targeting sequence has 25 nucleotides. In some embodiments, the targeting sequence has 24 nucleotides. In some embodiments, the targeting sequence has 23 nucleotides. In some embodiments, the targeting sequence has 22 nucleotides. In some embodiments, the targeting sequence has 21 nucleotides. In some embodiments, the targeting sequence has 20 nucleotides. In some embodiments, the targeting sequence has 19 nucleotides. In some embodiments, the targeting sequence has 18 nucleotides. In some embodiments, the targeting sequence has 17 nucleotides. In some embodiments, the targeting sequence has 16 nucleotides. In some embodiments, the targeting sequence has 15 nucleotides. In some embodiments, the targeting sequence has 14 nucleotides. h. Complex Formation with CasX Protein

[0137] In some embodiments, upon expression, the gRNA variant is complexed as an RNP with a CasX variant protein comprising any one of the sequences of Table 4 (SEQ ID NOS: 36-99, 101-148, and 26908-27154), or a sequence having at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto. In some embodiments, upon expression, the gRNA variant is complexed as an RNP with a CasX variant protein comprising any one of SEQ ID NOS: 59, 72-99, 101-148, or 26908-27154, or a sequence having at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto. In some embodiments, upon expression, the gRNA variant is complexed as an RNP with a CasX variant protein comprising any one of SEQ ID NOS 132-148, or 26908-27154 or a sequence having at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least 15 05 25 about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto.

[0138] In some embodiments, a gRNA variant has an improved ability to form a complex with a CasX protein (such as a reference CasX or a CasX variant protein) when compared to a reference gRNA. In some embodiments, a gRNA variant has an improved affinity for a CasX protein (such as a reference or variant protein) when compared to a reference gRNA, thereby improving its ability to form a ribonucleoprotein (RNP) complex with the CasX protein, as described in the Examples. Improving ribonucleoprotein complex formation may, in some embodiments, improve the efficiency with which functional RNPs are assembled. In some embodiments, greater than 90%, greater than 93%, greater than 95%, greater than 96%, greater than 97%, greater than 98% or greater than 99% of RNPs comprising a gRNA variant and its targeting sequence are competent for gene editing of a target nucleic acid.

[0139] Exemplary nucleotide changes that can improve the ability of gRNA variants to form a complex with CasX protein may, in some embodiments, include replacing the scaffold stem with a thermostable stem loop. Without wishing to be bound by any theory, replacing the scaffold stem with a thermostable stem loop could increase the overall binding stability of the gRNA variant with the CasX protein. Alternatively, or in addition, removing a large section of the stem loop could change the gRNA variant folding kinetics and make a functional folded gRNA easier and quicker to structurally-assemble, for example by lessening the degree to which the gRNA variant can get “tangled” in itself. In some embodiments, choice of scaffold stem loop sequence could change with different targeting sequences that are utilized for the gRNA. In some embodiments, scaffold sequence can be tailored to the targeting sequence and therefore the target sequence. Biochemical assays can be used to evaluate the binding affinity of CasX protein for the gRNA variant to form the RNP, including the assays of the Examples. For example, a person of ordinary skill can measure changes in the amount of a fluorescently tagged gRNA that is bound to an immobilized CasX protein, as a response to increasing concentrations of an additional unlabeled “cold competitor” gRNA. Alternatively, or in addition, fluorescence signal can be monitored to or seeing how it changes as different amounts of fluorescently labeled gRNA are flowed over immobilized CasX protein. Alternatively, the ability to form an RNP can be assessed using in vitro cleavage assays against a defined target nucleic acid sequence. 15 05 25 IV. Proteins for Modifying a Target Nucleic Acid

[0140] The present disclosure provides systems comprising a CRISPR nuclease that have utility in genome editing of eukaryotic cells. In some embodiments, the CRISPR nuclease employed in the genome editing systems is a Class 2 Type V nuclease. Although members of Class 2, Type V CRISPR-Cas systems have differences, they share some common characteristics that distinguish them from the Cas9 systems. Firstly, the Class 2, Type V nucleases possess a single RNA-guided RuvC domain-containing effector but no HNH domain, and they recognize T-rich PAM 5' upstream to the target region on the non-targeted strand, which is different from Cas9 systems which rely on G-rich PAM at 3' side of target sequences. Type V nucleases generate staggered double-stranded breaks distal to the PAM sequence, unlike Cas9, which generates a blunt end in the proximal site close to the PAM. In addition, Type V nucleases degrade ssDNA in trans when activated by target dsDNA or ssDNA binding in cis. In some embodiments, the Type V nucleases of the embodiments recognize a 5'-TC PAM motif and produce staggered ends cleaved solely by the RuvC domain. In some embodiments, the Type V nuclease is selected from the group consisting of Casl2a, Casl2b, Casl2c, Casl2d (CasY), Casl2j, Casl2k, CasZ and CasX. In some embodiments, the present disclosure provides systems comprising a CasX protein and one or more gRNA acids (CasX:gRNA system) that are specifically designed to modify a target nucleic acid sequence in eukaryotic cells.

[0141] The term “CasX protein”, as used herein, refers to a family of proteins, and encompasses all naturally occurring CasX proteins, proteins that share at least 50% identity to naturally occurring CasX proteins, as well as CasX variants possessing one or more improved characteristics relative to a naturally-occurring reference CasX protein.

[0142] CasX proteins of the disclosure comprise at least one of the following domains: a nontarget strand binding (NTSB) domain, a target strand loading (TSL) domain, a helical I domain, a helical II domain, an oligonucleotide binding domain (OBD), and a RuvC DNA cleavage domain.

[0143] In some embodiments, a CasX protein can bind and / or modify (e.g., nick, catalyze a double strand break, methylate, demethylate, etc.) a target nucleic acid at a specific sequence targeted by an associated gRNA, which hybridizes to a sequence within the target nucleic acid sequence. a. Reference CasX Proteins 15 05 25

[0144] The disclosure provides naturally-occurring CasX proteins (referred to herein as a "reference CasX protein"), which were subsequently modified to create the CasX variants of the disclosure. For example, reference CasX proteins can be isolated from naturally occurring prokaryotes, such as Deltaproteobacteria, Planctomycetes, or Candidatus Sungbacteria species. A reference CasX protein (interchangeably referred to herein as a reference CasX polypeptide) is a type II CRISPR / Cas endonuclease belonging to the CasX (interchangeably referred to as Casl2e) family of proteins that interacts with a guide RNA to form a ribonucleoprotein (RNP) complex.

[0145] In some cases, a reference CasX protein is isolated or derived from Dehaproieobacter having a sequence of: 1 MEKRINKIRK KLSADNATKP VSRSGPMKTL LVRVMTDDLK KRLEKRRKKP EVMPQVISNN 61 AANNLRMLLD DYTKMKEAIL QVYWQEFKDD HVGLMCKFAQ PASKKIDQNK LKPEMDEKGN 121 LTTAGFACSQ CGQPLFVYKL EQVSEKGKAY TNYFGRCNVA EHEKLILLAQ LKPEKDSDEA 181 VTYSLGKFGQ RALDFYSIHV TKESTHPVKP LAQIAGNRYA SGPVGKALSD ACMGTIASFL 241 SKYQDIIIEH QKWKGNQKR LESLRELAGK ENLEYPSVTL PPQPHTKEGV DAYNEVIARV 301 RMWVNLNLWQ KLKLSRDDAK PLLRLKGFPS FPWERRENE VDWWNTINEV KKLIDAKRDM 361 GRVFWSGVTA EKRNTILEGY NYLPNENDHK KREGSLENPK KPAKRQFGDL LLYLEKKYAG 421 DWGKVFDEAW ERIDKKIAGL TSHIEREEAR NAEDAQSKAV LTDWLRAKAS FVLERLKEMD 481 EKEFYACEIQ LQKWYGDLRG NPFAVEAENR WDISGFSIG SDGHSIQYRN LLAWKYLENG 541 KREFYLLMNY GKKGRIRFTD GTDIKKSGKW QGLLYGGGKA KVIDLTFDPD DEQLIILPLA 601 FGTRQGREFI WNDLLSLETG LIKLANGRVI EKTIYNKKIG RDEPALFVAL TFERREWDP 661 SNIKPVNLIG VDRGENIPAV IALTDPEGCP LPEFKDSSGG PTDILRIGEG YKEKQRAIQA 721 AKEVEQRRAG GYSRKFASKS RNLADDMVRN SARDLFYHAV THDAVLVFEN LSRGFGRQGK 781 RTFMTERQYT KMEDWLTAKL AYEGLTSKTY LSKTLAQYTS KTCSNCGFTI TTADYDGMLV 841 RLKKTSDGWA TTLNNKELKA EGQITYYNRY KRQTVEKELS AELDRLSEES GNNDISKWTK 901 GRRDEALFLL KKRFSHRPVQ EQFVCLDCGH EVHADEQAAL NIARSWLFLN SNSTEFKSYK 961 SGKQPFVGAW QAFYKRRLKE VWKPNA (SEQ ID NO: 1).

[0146] In some cases, a reference CasX protein is isolated or derived from Planctomycetes having a sequence of: 1 MQEIKRINKI RRRLVKDSNT KKAGKTGPMK TLLVRVMTPD LRERLENLRK KPENIPQPIS 61 NTSRANLNKL LTDYTEMKKA ILHVYWEEFQ KDPVGLMSRV AQPAPKNIDQ RKLIPVKDGN 121 ERLTSSGFAC SQCCQPLYVY KLEQVNDKGK PHTNYFGRCN VSEHERLILL SPHKPEANDE 181 LVTYSLGKFG QRALDFYSIH VTRESNHPVK PLEQIGGNSC ASGPVGKALS DACMGAVASF 241 LTKYQDIILE HQKVIKKNEK RLANLKDIAS ANGLAFPKIT LPPQPHTKEG IEAYNNWAQ 301 IVIWVNLNLW QKLKIGRDEA KPLQRLKGFP SFPLVERQAN EVDWWDMVCN VKKLINEKKE 361 DGKVFWQNLA GYKRQEALLP YLSSEEDRKK GKKFARYQFG DLLLHLEKKH GEDWGKVYDE 421 AWERIDKKVE GLSKHIKLEE ERRSEDAQSK AALTDWLRAK ASFVIEGLKE ADKDEFCRCE 481 LKLQKWYGDL RGKPFAIEAE NSILDISGFS KQYNCAFIWQ KDGVKKLNLY LIINYFKGGK 541 LRFKKIKPEA FEANRFYTVI NKKSGEIVPM EVNFNFDDPN LIILPLAFGK ROGREFIWND 601 LLSLETGSLK LANGRVIEKT LYNRRTRQDE PALFVALTFE RREVLDSSNI KPMNLIGIDR 661 GENIPAVIAL TDPEGCPLSR FKDSLGNPTH ILRIGESYKE KQRTIQAAKE VEQRRAGGYS 721 RKYASKAKNL ADDMVRNTAR DLLYYAVTQD AMLIFENLSR GFGRQGKRTF MAERQYTRME 781 DWLTAKLAYE GLPSKTYLSK TLAQYTSKTC SNCGFTITSA DYDRVLEKLK KTATGWMTTI 841 NGKELKVEGQ ITYYNRYKRQ NWKDLSVEL DRLSEESVNN DISSWTKGRS GEALSLLKKR 901 FSHRPVQEKF VCLNCGFETH ADEQAALNIA RSWLFLRSQE YKKYQTNKTT GNTDKRAFVE 961 TWQSFYRKKL KEVWKPAV (SEQ ID NO: 2).

[0147] In some cases, a reference CasX protein is isolated or derived from Candidatus 15 05 25 Sungbacteria having a sequence of 1 MDNANKPSTK SLVNTTRISD HFGVTPGQVT RVFSFGIIPT KRQYAIIERW FAAVEAARER 61 LYGMLYAHFQ ENPPAYLKEK FSYETFFKGR PVLNGLRDID PTIMTSAVFT ALRHKAEGAM 121 AAFHTNHRRL FEEARKKMRE YAECLKANEA LLRGAADIDW DKIVNALRTR LNTCLAPEYD 181 AVIADFGALC AFRALIAETN ALKGAYNHAL NQMLPALVKV DEPEEAEESP RLRFFNGRIN 241 DLPKFPVAER ETPPDTETII RQLEDMARVI PDTAEILGYI HRIRHKAARR KPGSAVPLPQ 301 RVALYCAIRM ERNPEEDPST VAGHFLGEID RVCEKRRQGL VRTPFDSQIR ARYMDIISFR 361 ATLAHPDRWT EIQFLRSNAA SRRVRAETIS APFEGFSWTS NRTNPAPQYG MALAKDANAP 421 ADAPELCICL SPSSAAFSVR EKGGDLIYMR PTGGRRGKDN PGKEITWVPG SFDEYPASGV 481 ALKLRLYFGR SQARRMLTNK TWGLLSDNPR VFAANAELVG KKRNPQDRWK LFFHMVISGP 541 PPVEYLDFSS DVRSRARTVI GINRGEVNPL AYAWSVEDG QVLEEGLLGK KEYIDQLIET 601 RRRISEYQSR EQTPPRDLRQ RVRHLQDTVL GSARAKIHSL IAFWKGILAI ERLDDQFHGR 661 EQKIIPKKTY LANKTGFMNA LSFSGAVRVD KKGNPWGGMI EIYPGGISRT CTQCGTVWLA 721 RRPKNPGHRD AMWIPDIVD DAAATGFDNV DCDAGTVDYG ELFTLSREWV RLTPRYSRVM 781 RGTLGDLERA IRQGDDRKSR QMLELALEPQ PQWGQFFCHR CGFNGQSDVL AATNLARRAI 841 SLIRRLPDTD TPPTP (SEQ ID NO: 3). b. CasX Variant Proteins

[0148] The present disclosure provides variants of a reference CasX protein (interchangeably referred to herein as “CasX variant” or “CasX variant protein”), wherein the CasX variants comprise at least one modification in at least one domain relative to the reference CasX protein, including but not limited to the sequences of SEQ ID NOS: 1-3.

[0149] The CasX variants of the disclosure have one or more improved characteristics compared to reference CasX proteins. Exemplary improved characteristics of the CasX variant embodiments include, but are not limited to improved folding of the variant, improved binding 15 05 25 affinity to the gRNA, improved binding affinity to the target nucleic acid, improved ability to utilize a greater spectrum of PAM sequences in the editing and / or binding of target DNA, improved unwinding of the target DNA, increased editing activity, improved editing efficiency, improved editing specificity, increased percentage of a eukaryotic genome that can be efficiently edited, increased activity of the nuclease, increased target strand loading for double strand cleavage, decreased target strand loading for single strand nicking, decreased off-target cleavage, improved binding of the non-target strand of DNA, improved protein stability, improved protein:gRNA (RNP) complex stability, improved protein solubility, improved protein:gRNA (RNP) complex solubility, improved protein yield, improved protein expression, and improved fusion characteristics, as described more fully, below. Exemplary improved characteristics are described in WO 2020 / 247882A1 and WO 2020 / 247883, incorporated by reference herein. In the foregoing embodiments, the one or more of the improved characteristics of the CasX variant is at least about 1.1 to about 100,000-fold improved relative to the reference CasX protein of SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 3, when assayed in a comparable fashion. In other embodiments, the improvement is at least about 1.1-fold, at least about 2-fold, at least about 5-fold, at least about 10-fold, at least about 50-fold, at least about 100-fold, at least about 500-fold, at least about 1000-fold, at least about 5000-fold, at least about 10,000-fold, or at least about 100,000-fold compared to the reference CasX protein of SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 3, when assayed in a comparable fashion. In other cases, the one or more improved characteristics of an RNP of the CasX variant and the gRNA variant are at least about 1.1, at least about 10, at least about 100, at least about 1000, at least about 10,000, at least about 100,000-fold or more improved relative to an RNP of the reference CasX protein of SEQ ID NO: 1, SEQ ID NO:2, or SEQ ID NO:3 and the gRNA of Table 2. In other cases, the one or more of the improved characteristics of an RNP of the CasX variant and the gRNA variant are about 1.1 to 100,00-fold, about 1.1 to 10,00-fold, about 1.1 to 1,000-fold, about 1.1 to 500-fold, about 1.1 to 100-fold, about 1.1 to 50-fold, about 1.1 to 20-fold, about 10 to 100,00-fold, about 10 to 10,00-fold, about 10 to 1,000-fold, about 10 to 500-fold, about 10 to 100-fold, about 10 to 50-fold, about 10 to 20-fold, about 2 to 70-fold, about 2 to 50-fold, about 2 to 30-fold, about 2 to 20-fold, about 2 to 10-fold, about 5 to 50-fold, about 5 to 30-fold, about 5 to 10-fold, about 100 to 100,00-fold, about 100 to 10,00-fold, about 100 to 1,000-fold, about 100 to 500-fold, about 500 to 100,00-fold, about 500 to 10,00-fold, about 500 to 1,000-fold, about 500 to 750-fold, about 1,000 to 100,00-fold, about 10,000 to 100,00-fold, about 20 to 500-fold, about 20 to 250- 15 05 25 fold, about 20 to 200-fold, about 20 to 100-fold, about 20 to 50-fold, about 50 to 10,000-fold, about 50 to 1,000-fold, about 50 to 500-fold, about 50 to 200-fold, or about 50 to 100-fold, improved relative to an RNP of the reference CasX protein of SEQ ID NO: 1, SEQ ID NO:2, or SEQ ID NO:3 and the gRNA of Table 2, when assayed in a comparable fashion. In other cases, the one or more improved characteristics of an RNP of the CasX variant and the gRNA variant areabout 1.1-fold, 1.2-fold, 1.3-fold, 1.4-fold, 1.5-fold, 1.6-fold, 1.7-fold, 1.8-fold, 1.9-fold, 2-fold, 3-fold, 4-fold, 5-fold, 6-fold, 7-fold, 8-fold, 9-fold, 10-fold, 11-fold, 12-fold, 13-fold, 14-fold, 15-fold, 16-fold, 17-fold, 18-fold, 19-fold, 20-fold, 25-fold, 30-fold, 40-fold, 45-fold, 50-fold, 55-fold, 60-fold, 70-fold, 80-fold, 90-fold, 100-fold, 110-fold, 120-fold, 130-fold, 140-fold, 150-fold, 160-fold, 170-fold, 180-fold, 190-fold, 200-fold, 210-fold, 220-fold, 230-fold, 240-fold, 250-fold, 260-fold, 270-fold, 280-fold, 290-fold, 300-fold, 310-fold, 320-fold, 330-fold, 340-fold, 350-fold, 360-fold, 370-fold, 380-fold, 390-fold, 400-fold, 425-fold, 450-fold, 475-fold, or 500-fold improved relative to an RNP of the reference CasX protein of SEQ ID NO: 1, SEQ ID NO:2, or SEQ ID NO:3 and the gRNA of Table 2, when assayed in a comparable fashion.

[0150] The term CasX variant is inclusive of variants that are fusion proteins; i.e. the CasX is “fused to” a heterologous sequence. This includes CasX variants comprising CasX variant sequences and N-terminal, C-terminal, or internal fusions of the CasX to a heterologous protein or domain thereof.

[0151] In some embodiments, the CasX variant comprises at least one modification in the NTSB domain. In some embodiments, the CasX variant comprises at least one modification in the TSL domain. In some embodiments, the CasX variant comprises at least one modification in the helical I domain. In some embodiments, the CasX variant comprises at least one modification in the helical II domain. In some embodiments, the CasX variant comprises at least one modification in the OBD domain. In some embodiments, the CasX variant comprises at least one modification in the RuvC DNA cleavage domain. In some embodiments, the at least one modification in the RuvC DNA cleavage domain comprises an amino acid substitution of one or more of amino acids K682, G695, A708, ¥711, D732, A739, D733, L742, V747, F755, M771, M779, W782, A788, G791, L792, P793, Y797, M799, Q804, S819, or Y857 or a deletion of amino acid P793 of SEQ ID NO:2.

[0152] In some embodiments, the CasX variant protein comprises at least one modification in at least 1 domain, in at least each of 2 domains, in at least each of 3 domains, in at least each of 4 15 05 25 domains or in at least each of 5 domains of the reference CasX protein, including the sequences of SEQ ID NOS: 1-3. In some embodiments, the CasX variant protein comprises two or more modifications in at least one domain of the reference CasX protein. In some embodiments, the CasX variant protein comprises at least two modifications in at least one domain of the reference CasX protein, at least three modifications in at least one domain of the reference CasX protein or at least four modifications in at least one domain of the reference CasX protein. In some embodiments, wherein the CasX variant comprises two or more modifications compared to a reference CasX protein, each modification is made in a domain independently selected from the group consisting of a NTSBD, TSLD, Helical I domain, Helical II domain, OBD, and RuvC DNA cleavage domain. In some embodiments, the at least one modification of the CasX variant protein comprises a deletion of at least a portion of one domain of the reference CasX protein of SEQ ID NOS: 1-3. In some embodiments, the deletion is in the NTSBD, TSLD, Helical I domain, Helical II domain, OBD, or RuvC DNA cleavage domain. In other embodiments, the disclosure provides CasX variants wherein the CasX variants comprise at least one modification relative to another CasX variant; e.g., CasX variant 515 is a variant of CasX variant 491. All variants that improve one or more functions or characteristics of the CasX variant protein when compared to a reference CasX protein (or the variant from which it was derived) described herein are envisaged as being within the scope of the disclosure.

[0153] In some embodiments, the modification of the CasX variant is a mutation in one or more amino acids of the reference CasX. In other embodiments, the modification is a substitution of one or more domains of the reference CasX with one or more domains from a different CasX. In some embodiments, insertion includes the insertion of a part or all of a domain from a different CasX protein. Mutations can occur in any one or more domains of the reference CasX protein, and may include, for example, deletion of part or all of one or more domains, or one or more amino acid substitutions, deletions, or insertions in any domain of the reference CasX protein. The domains of CasX proteins include the non-target strand binding (NTSB) domain, the target strand loading (TSL) domain, the helical I domain, the helical II domain, the oligonucleotide binding domain (OBD), and the RuvC DNA cleavage domain. Any change in amino acid sequence of a reference CasX protein that leads to an improved characteristic of the CasX protein is considered a CasX variant protein of the disclosure. For example, CasX variants can comprise one or more amino acid substitutions, insertions, 15 05 25 deletions, or swapped domains, or any combinations thereof, relative to a reference CasX protein sequence.

[0154] Suitable mutagenesis methods for generating CasX variant proteins of the disclosure may include, for example, Deep Mutational Evolution (DME), deep mutational scanning (DMS), error prone PCR, cassette mutagenesis, random mutagenesis, staggered extension PCR, gene shuffling, or domain swapping. In some embodiments, the CasX variants are designed, for example by selecting one or more desired mutations in a reference CasX. In certain embodiments, the activity of a reference CasX protein is used as a benchmark against which the activity of one or more CasX variants are compared, thereby measuring improvements in function of the CasX variants.

[0155] In some embodiments of the CasX variants described herein, the at least one modification comprises: (a) a substitution of 1 to 100 consecutive or non-consecutive amino acids in the CasX variant compared to a reference CasX of SEQ ID NO: 1, SEQ ID NO:2, SEQ ID NO:3, CasX variant 491 or CasX variant 515; (b) a deletion of 1 to 100 consecutive or non-consecutive amino acids in the CasX variant compared to a reference CasX or the variant from which it was derived; (c) an insertion of 1 to 100 consecutive or non-consecutive amino acids in the CasX compared to a reference CasX or the variant from which it was derived; or (d) any combination of (a)-(c). In some embodiments, the at least one modification comprises: (a) a substitution of 5-10 consecutive or non-consecutive amino acids in the CasX variant compared to a reference CasX of SEQ ID NO: 1, SEQ ID NO:2, SEQ ID NO:3, CasX 491 or CasX 515; (b) a deletion of 1-5 consecutive or non-consecutive amino acids in the CasX variant compared to a reference CasX or the variant from which it was derived; (c) an insertion of 1-5 consecutive or non-consecutive amino acids in the CasX compared to a reference CasX or the variant from which it was derived; or (d) any combination of (a)-(c).

[0156] In some embodiments, the CasX variant protein comprises or consists of a sequence that has at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at lease 80, at least 90, or at least 100 alterations relative to the sequence of SEQ ID NO: 1, SEQ ID NO:2, SEQ ID NO:3, CasX 491 (with reference to Table 4) or CasX 515 (with reference to Table 4). These alterations can be amino acid insertions, deletions, substitutions, or any combinations thereof. The alterations can be in one domain or in any domain or any combination of domains of the CasX variant. Any amino acid can be substituted for any other amino acid in the substitutions 15 05 25 described herein. The substitution can be a conservative substitution (e.g., a basic amino acid is substituted for another basic amino acid). The substitution can be a non-conservative substitution (e.g., a basic amino acid is substituted for an acidic amino acid or vice versa). For example, a proline in a reference CasX protein can be substituted for any of arginine, histidine, lysine, aspartic acid, glutamic acid, serine, threonine, asparagine, glutamine, cysteine, glycine, alanine, isoleucine, leucine, methionine, phenylalanine, tryptophan, tyrosine or valine to generate a CasX variant protein of the disclosure.

[0157] Any permutation of the substitution, insertion and deletion embodiments described herein can be combined to generate a CasX variant protein of the disclosure. For example, a CasX variant protein can comprise at least one substitution and at least one deletion relative to a reference CasX protein sequence, at least one substitution and at least one insertion relative to a reference CasX protein sequence, at least one insertion and at least one deletion relative to a reference CasX protein sequence, or at least one substitution, one insertion and one deletion relative to a reference CasX protein sequence.

[0158] In some embodiments, the CasX variant comprises at least one modification compared to the reference CasX sequence of SEQ ID NO:2 is selected from one or more of: (a) an amino acid substitution of L379R; (b) an amino acid substitution of A708K; (c) an amino acid substitution of T620P; (d) an amino acid substitution of E385P; (e) an amino acid substitution of Y857R; (f) an amino acid substitution of I658V; (g) an amino acid substitution of F399L; (h) an amino acid substitution of Q252K; (i) an amino acid substitution of L404K; and (j) an amino acid deletion of P793.

[0159] In some embodiments, the CasX variant protein comprises between 400 and 2000 amino acids, between 500 and 1500 amino acids, between 700 and 1200 amino acids, between 800 and 1100 amino acids, or between 900 and 1000 amino acids.

[0160] In some embodiments, a CasX variant protein comprises a sequence of SEQ ID NOS: 59, 72-99, 101-148, and 26908-27154 as set forth in Table 4. In some embodiments, a CasX variant protein consists of a sequence of SEQ ID NOS: 59, 72-99, 101-148, and 26908-27154 as set forth in Table 4. In other embodiments, a CasX variant protein comprises a sequence at least 60% identical, at least 65% identical, at least 70% identical, at least 75% identical, at least 80% identical, at least 81% identical, at least 82% identical, at least 83% identical, at least 84% identical, at least 85% identical, at least 86% identical, at least 86% identical, at least 87% identical, at least 88% identical, at least 89% identical, at least 89% identical, at least 90% 15 05 25 identical, at least 91% identical, at least 92% identical, at least 93% identical, at least 94% identical, at least 95% identical, at least 96% identical, at least 97% identical, at least 98% identical, at least 99% identical, at least 99.5% identical to a sequence of SEQ ID NOS: 59, 72-99, 101-148, or 26908-27154 as set forth in Table 4. In some embodiments, a CasX variant protein comprises or consists of a sequence of SEQ ID NOS: 536-99, 101-148, or 26908-27154. In other embodiments, a CasX variant protein comprises a sequence at least 60% identical, at least 65% identical, at least 70% identical, at least 75% identical, at least 80% identical, at least 81% identical, at least 82% identical, at least 83% identical, at least 84% identical, at least 85% identical, at least 86% identical, at least 86% identical, at least 87% identical, at least 88% identical, at least 89% identical, at least 89% identical, at least 90% identical, at least 91% identical, at least 92% identical, at least 93% identical, at least 94% identical, at least 95% identical, at least 96% identical, at least 97% identical, at least 98% identical, at least 99% identical, at least 99.5% identical to a sequence of SEQ ID NOS: 36-99, 101-148, or 26908-27154. In some embodiments, a CasX variant protein comprises or consists of a sequence of SEQ ID NOS: 132-148, or 26908-27154. In other embodiments, a CasX variant protein comprises a sequence at least 60% identical, at least 65% identical, at least 70% identical, at least 75% identical, at least 80% identical, at least 81% identical, at least 82% identical, at least 83% identical, at least 84% identical, at least 85% identical, at least 86% identical, at least 86% identical, at least 87% identical, at least 88% identical, at least 89% identical, at least 89% identical, at least 90% identical, at least 91% identical, at least 92% identical, at least 93% identical, at least 94% identical, at least 95% identical, at least 96% identical, at least 97% identical, at least 98% identical, at least 99% identical, at least 99.5% identical to a sequence of SEQ ID NOS: 132-148 or 26908-27154. Table 4: CasX Variant Sequences SEQ ID NO Variant Description of Variant 36 ND TSL, Helical I, Helical II, OBD and RuvC domains from SEQ ID NO: 2 and an NTSB domain from SEQ ID NO: 1 37 ND NTSB, Helical I, Helical II, OBD and RuvC domains from SEQ ID NO: 2 and a TSL domain from SEQ ID NO: 1. 38 ND TSL, Helical I, Helical IL OBD and RuvC domains from SEQ ID NO: 1 and an NTSB domain from SEQ ID NO: 2 15 05 25 SEQ ID NO Variant Description of Variant 39 ND NTSB, Helical I, Helical II, OBD and RuvC domains from SEQ ID NO: 1 and an TSL domain from SEQ ID NO: 2. 40 ND NTSB, TSL, Helical I, Helical II and OBD domains SEQ ID NO: 2 and an exogenous RuvC domain or a portion thereof from a second CasX protein. 41 ND ND 42 ND NTSB, TSL, Helical II, OBD and RuvC domains from SEQ ID NO: 2 and a Helical I domain from SEQ ID NO: 1 43 ND NTSB, TSL, Helical I, OBD and RuvC domains from SEQ ID NO: 2 and a Helical II domain from SEQ ID NO: 1 44 ND NTSB, TSL, Helical I, Helical II and RuvC domains from a first CasX protein and an exogenous OBD or a part thereof from a second CasX protein 45 ND ND 46 ND ND 47 ND substitution of L379R, a substitution of C477K, a substitution of A708K, a deletion of P at position 793 and a substitution of T620P of SEQ ID NO: 2 48 ND substitution of M771A of SEQ ID NO: 2. 49 ND substitution of L379R, a substitution of A708K, a deletion of P at position 793 and a substitution of D732N of SEQ ID NO: 2. 50 ND substitution ofW782Q of SEQ ID NO: 2. 51 ND substitution of M771Q of SEQ ID NO: 2 52 ND substitution of R458I and a substitution of A739V of SEQ ID NO: 2. 53 ND L379R, a substitution of A708K, a deletion of P at position 793 and a substitution of M771N of SEQ ID NO: 2 54 ND substitution of L379R, a substitution of A708K, a deletion of P at position 793 and a substitution of A739T of SEQ ID NO: 2 55 ND substitution of L379R, a substitution of C477K, a substitution of A708K, a deletion of P at position 793 and a substitution of D489S of SEQ ID NO: 2. 56 ND substitution of L379R, a substitution of C477K, a substitution of A708K, a deletion of P at position 793 and a substitution of D732N of SEQ ID NO: 2. 57 ND substitution of V711K of SEQ ID NO: 2. 58 ND substitution of L379R, a substitution of C477K, a substitution of A708K, a deletion of P at position 793 and a substitution of Y797L of SEQ ID NO: 2. 60 ND substitution of L379R, a substitution of C477K, a substitution of A708K, a deletion of 15 05 25 SEQ ID NO Variant Description of Variant P at position 793 and a substitution of M77 IN of SEQ ID NO: 2. 61 ND substitution of A708K, a deletion of P at position 793 and a substitution of E386S of SEQ ID NO: 2. 62 ND substitution of L379R, a substitution of C477K, a substitution of A708K and a deletion of P at position 793 of SEQ ID NO: 2. 63 ND substitution of L792D of SEQ ID NO: 2. 64 ND substitution of G791F of SEQ ID NO: 2. 65 ND substitution of A708K, a deletion of P at position 793 and a substitution of A739V of SEQ ID NO: 2. 66 ND substitution of L379R, a substitution of A708K, a deletion of P at position 793 and a substitution of A739V of SEQ ID NO: 2. 67 ND substitution of C477K, a substitution of A708K and a deletion of P at position 793 of SEQ ID NO: 2. 68 ND substitution of L249I and a substitution of M771N of SEQ ID NO: 2. 69 ND substitution of V747K of SEQ ID NO: 2. 70 ND substitution of L379R, a substitution of C477K, a substitution of A708K, a deletion of P at position 793 and a substitution of M779N of SEQ ID NO: 2. 71 ND L379R, F755M 59 119 ND 72 429 ND 73 430 ND 74 431 ND 75 432 ND 76 433 ND 77 434 ND 78 435 ND 79 436 ND 80 437 ND 81 438 ND 82 439 ND 15 05 25 SEQ ID NO Variant Description of Variant 83 440 ND 84 441 ND 85 442 ND 86 443 ND 87 444 ND 88 445 ND 89 446 ND 90 447 ND 91 448 ND 92 449 ND 93 450 ND 94 451 ND 95 452 ND 96 453 ND 97 454 ND 98 455 ND 99 456 ND 101 457 ND 102 458 ND 103 459 ND 104 460 ND 105 278 ND 106 279 ND 107 280 ND 108 285 ND 109 286 ND 15 05 25 SEQ ID NO Variant Description of Variant 110 287 ND 111 288 ND 112 290 ND 113 291 ND 114 293 ND 115 300 ND 116 492 ND 117 493 ND 118 387 ND 119 395 ND 120 485 ND 121 486 ND 122 487 ND 123 488 ND 124 489 ND 125 490 ND 126 491 ND 127 494 ND 128 328 ND 129 388 ND 130 389 ND 131 390 ND 132 514 ND 133 515 ND 134 516 ND 135 517 ND 15 05 25 SEQ ID NO Variant Description of Variant 136 518 ND 137 519 ND 138 520 ND 139 522 ND 140 523 ND 141 524 ND 142 525 ND 143 526 ND 144 527 ND 145 528 ND 146 529 ND 147 530 ND 148 531 ND 26908 532 ND 26909 533 ND 26910 534 ND 26911 535 ND 26912 536 ND 26913 537 ND 26914 538 ND 26915 539 ND 26916 540 ND 26917 541 ND 26918 542 ND 26919 543 ND 26920 544 ND 15 05 25 SEQ ID NO Variant Description of Variant 26921 545 ND 26922 546 ND 26923 547 ND 26924 548 ND 26925 550 ND 26926 551 ND 26927 552 ND 26928 553 ND 26929 554 ND 26930 555 ND 26931 556 ND 26932 557 ND 26933 558 ND 26934 559 ND 26935 560 ND 26936 561 ND 26937 562 ND 26938 563 ND 26939 564 ND 26940 565 ND 26941 566 ND 26942 567 ND 26943 568 ND 26944 569 ND 26945 570 ND 26946 571 ND 26947 572 ND 15 05 25 SEQ ID NO Variant Description of Variant 26948 573 ND 26949 574 ND 26950 575 ND 26951 576 ND 26952 577 ND 26953 578 ND 26954 579 ND 26955 580 ND 26956 581 ND 26957 582 ND 26958 583 ND 26959 584 ND 26960 585 ND 26961 586 ND 26962 587 ND 26963 588 ND 26964 589 ND 26965 590 ND 26966 591 ND 26967 592 ND 26968 593 ND 26969 594 ND 26970 595 ND 26971 596 ND 26972 597 ND 26973 598 ND 26974 599 ND 26975 600 ND 15 05 25 SEQ ID NO Variant Description of Variant 26976 601 ND 26977 602 ND 26978 603 ND 26979 604 ND 26980 605 ND 26981 606 ND 26982 607 ND 26983 608 ND 26984 609 ND 26985 610 ND 26986 611 ND 26987 612 ND 26988 613 ND 26989 614 ND 26990 615 ND 26991 616 ND 26992 617 ND 26993 618 ND 26994 619 ND 26995 620 ND 26996 621 ND 26997 622 ND 26998 623 ND 26999 624 ND 27000 625 ND 27001 626 ND 27002 627 ND 27003 628 ND 15 05 25 SEQ ID NO Variant Description of Variant 27004 629 ND 27005 630 ND 27006 631 ND 27007 632 ND 27008 633 ND 27009 634 ND 27010 635 ND 27011 636 ND 27012 637 ND 27013 638 ND 27014 639 ND 27015 640 ND 27016 641 ND 27017 642 ND 27018 643 ND 27019 644 ND 27020 645 ND 27021 646 ND 27022 647 ND 27023 648 ND 27024 649 ND 27025 650 ND 27026 651 ND 27027 652 ND 27028 653 ND 27029 654 ND 27030 655 ND 27031 656 ND 15 05 25 SEQ ID NO Variant Description of Variant 27032 657 ND 27033 658 ND 27034 659 ND 27035 660 ND 27036 661 ND 27037 662 ND 27038 663 ND 27039 664 ND 27040 665 ND 27041 666 ND 27042 667 ND 27043 668 ND 27044 669 ND 27154 670 ND 27045 671 ND 27046 672 ND 27047 673 ND 27048 674 ND 27049 675 ND 27050 676 ND 27051 677 ND 27052 678 ND 27053 679 ND 27054 680 ND 27055 681 ND 27056 682 ND 27057 683 ND 27058 684 ND 15 05 25 SEQ ID NO Variant Description of Variant 27059 685 ND 27060 686 ND 27061 687 ND 27062 688 ND 27063 689 ND 27064 690 ND 27065 691 ND 27066 692 ND 27067 693 ND 27068 694 ND 27069 701 ND 27070 702 ND 27071 703 ND 27072 704 ND 27073 705 ND 27074 706 ND 27075 707 ND 27076 708 ND 27077 709 ND 27078 710 ND 27079 711 ND 27080 712 ND 27081 713 ND 27082 714 ND 27083 715 ND 27084 716 ND 27085 717 ND 27086 718 ND 15 05 25 SEQ ID NO Variant Description of Variant 27087 719 ND 27088 720 ND 27089 721 ND 27090 722 ND 27091 723 ND 27092 724 ND 27093 725 ND 27094 726 ND 27095 727 ND 27096 728 ND 27097 729 ND 27098 730 ND 27099 731 ND 27100 732 ND 27101 733 ND 27102 734 ND 27103 735 ND 27104 736 ND 27105 737 ND 27106 738 ND 27107 739 ND 27108 740 ND 27109 741 ND 27110 742 ND 27111 743 ND 27112 744 ND 27113 745 ND 27114 746 ND 15 05 25 SEQ ID NO Variant Description of Variant 27115 747 ND 27116 748 ND 27117 749 ND 27118 750 ND 27119 751 ND 27120 752 ND 27121 753 ND 27122 754 ND 27123 755 ND 27124 756 ND 27125 757 ND 27126 758 ND 27127 759 ND 27128 760 ND 27129 761 ND 27130 762 ND 27131 763 ND 27132 764 ND 27133 765 ND 27134 766 ND 27135 767 ND 27136 768 ND 27137 769 ND 27138 770 ND 27139 777 ND 27140 778 ND 27141 779 ND 27142 780 ND SEQ ID NO Variant Description of Variant 27143 781 ND 27144 782 ND 27145 783 ND 27146 784 ND 27147 785 ND 27148 786 ND 27149 787 ND 27150 788 ND 27151 789 ND 27152 790 ND 27153 791 ND 15 05 25 c. CasX Variant Proteins with Domains from Multiple Source Proteins

[0161] In certain embodiments, the disclosure provides a chimeric CasX protein comprising protein domains from two or more different CasX proteins, such as two or more naturally occurring CasX proteins, or two or more CasX variant protein sequences as described herein. As used herein, a “chimeric CasX protein” refers to a CasX containing at least two domains isolated or derived from different sources, such as two naturally occurring proteins, which may, in some embodiments, be isolated from different species. For example, in some embodiments, a chimeric CasX protein comprises a first domain from a first CasX protein and a second domain from a second, different CasX protein. In some embodiments, the first domain can be selected from the group consisting of the NTSB, TSL, helical I, helical II, OBD and RuvC domains. In some embodiments, the second domain is selected from the group consisting of the NTSB, TSL, helical I, helical II, OBD and RuvC domains with the second domain being different from the foregoing first domain. In the case of split or non-contiguous domains such as helical I, RuvC and OBD, a portion of the non-contiguous domain can be replaced with the corresponding portion from any other source. For example, the helical I-I domain (sometimes referred to as helical La) in SEQ ID NO: 2 can be replaced with the corresponding helical I-I sequence from SEQ ID NO: 1, and the like. Domain sequences from reference CasX proteins, and their coordinates, are shown in Table 5. Representative examples of chimeric CasX proteins include the variants of CasX 472-483, 485-491 and 515, the sequences of which are set forth in Table 4. Table 5. Domain coordinates in Reference CasX proteins Domain Name Coordinates in SEQ ID NO: 1 Coordinates in SI IQ ID NO: 2 OBD a 1-55 1-57 helical I a 56-99 58-101 NTSB 100-190 102-191 helical I b 191-331 192-332 helical II 332-508 333-500 OBD b 509-659 501-646 RuvC a 660-823 647-810 TSL 824-933 811-920 RuvC b 934-986 921-978 *OBD a and b, helical I a and b, and RuvC a and b are also referred to herein as OBDI and II, helical I-I and I-II, and RuvC I and II. 15 05 25 d. Protein Affinity for the gRNA

[0162] In some embodiments, a CasX variant protein has improved affinity for the gRNA relative to a reference CasX protein, leading to the formation of the ribonucleoprotein complex (RNP). Increased affinity of the CasX variant protein for the gRNA may, for example, result in a lower Kdfor the generation of a RNP complex, which can, in some cases, result in a more stable ribonucleoprotein complex formation. In some embodiments, increased affinity of the CasX variant protein for the gRNA results in increased stability of the ribonucleoprotein complex when delivered to human cells. This increased stability can affect the function and utility of the complex in the cells of a subject, as well as result in improved pharmacokinetic properties in blood, when delivered to a subject. In some embodiments, increased affinity of the CasX variant protein, and the resulting increased stability of the ribonucleoprotein complex, allows for a lower dose of the CasX variant protein to be delivered to the subject or cells while still having the desired activity, for example in vivo or in vitro gene editing. In some embodiments, a higher affinity (tighter binding) of a CasX variant protein to a gRNA allows for a greater amount of editing events when both the CasX variant protein and the gRNA remain in an RNP complex. Increased editing events can be assessed using editing assays such as the tdTom editing assays described herein. In some embodiments, the Kd of a CasX variant protein for a gRNA is increased relative to a reference CasX protein by a factor of at least about 1.1, at least about 1.2, 15 05 25 at least about 1.3, at least about 1.4, at least about 1.5, at least about 1.6, at least about 1.7, at least about 1.8, at least about 1.9, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 15, at least about 20, at least about 25, at least about 30, at least about 35, at least about 40, at least about 45, at least about 50, at least about 60, at least about 70, at least about 80, at least about 90, or at least about 100. In some embodiments, the CasX variant has about 1.1 to about 10-fold increased binding affinity to the gRNA compared to the reference CasX protein of SEQ ID NO: 2.

[0163] In some embodiments, increased affinity of the CasX variant protein for the gRNA results in increased stability of the ribonucleoprotein complex when delivered to mammalian cells, including in vivo delivery to a subject. This increased stability can affect the function and utility of the complex in the cells of a subject, as well as result in improved pharmacokinetic properties in blood, when delivered to a subject. In some embodiments, increased affinity of the CasX variant protein, and the resulting increased stability of the ribonucleoprotein complex, allows for a lower dose of the CasX variant protein to be delivered to the subject or cells while still having the desired activity; for example in vivo or in vitro gene editing. The increased ability to form RNP and keep them in stable form can be assessed using assays such as the in vitro cleavage assays described in the Examples herein. In some embodiments, RNP comprising the CasX variants of the disclosure are able to achieve a kcieave rate when complexed as an RNP that is at last 2-fold, at least 5-fold, or at least 10-fold higher compared to RNP comprising a reference CasX of SEQ ID NOS: 1-3.

[0164] In some embodiments, a higher affinity (tighter binding) of a CasX variant protein to a gRNA allows for a greater amount of editing events when both the CasX variant protein and the gRNA remain in an RNP complex. Increased editing events can be assessed using editing assays such as the assays described herein.

[0165] Without wishing to be bound by theory, in some embodiments amino acid changes in the Helical I domain can increase the binding affinity of the CasX variant protein with the gRNA targeting sequence, while changes in the Helical II domain can increase the binding affinity of the CasX variant protein with the gRNA scaffold stem loop, and changes in the oligonucleotide binding domain (OBD) increase the binding affinity of the CasX variant protein with the gRNA triplex. 15 05 25

[0166] Methods of measuring CasX protein binding affinity for a gRNA include in vitro methods using purified CasX protein and gRNA. The binding affinity for reference CasX and variant proteins can be measured by fluorescence polarization if the gRNA or CasX protein is tagged with a fluorophore. Alternatively, or in addition, binding affinity can be measured by biolayer interferometry, electrophoretic mobility shift assays (EMSAs), or filter binding. Additional standard techniques to quantify absolute affinities of RNA binding proteins such as the reference CasX and variant proteins of the disclosure for specific gRNAs such as reference gRNAs and variants thereof include, but are not limited to, isothermal calorimetry (ITC), and surface plasmon resonance (SPR), as well as the methods of the Examples. e. Affinity for Target Nucleic Acid

[0167] In some embodiments, a CasX variant protein has improved binding affinity for a target nucleic acid sequence relative to the affinity of a reference CasX protein for a target nucleic acid sequence. In some embodiments, affinity of a CasX variant protein of the disclosure for a target nucleic acid molecule is increased relative to a reference CasX protein by a factor of at least about 1.1, at least about 1.2, at least about 1.3, at least about 1.4, at least about 1.5, at least about 1.6, at least about 1.7, at least about 1.8, at least about 1.9, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 15, at least about 20, at least about 25, at least about 30, at least about 35, at least about 40, at least about 45, at least about 50, at least about 60, at least about 70, at least about 80, at least about 90, or at least about 100.

[0168] CasX variants with higher affinity for their target nucleic acid may, in some embodiments, cleave the target nucleic acid sequence more rapidly than a reference CasX protein that does not have increased affinity for the target nucleic acid. In some embodiments, the improved affinity for the target nucleic acid sequence comprises improved affinity for the target nucleic acid sequence, improved binding affinity to a wider spectrum of PAM sequences, an improved ability to search DNA for the target nucleic acid sequence, or any combinations thereof, resulting in an increased ability to modify the target nucleic acid. In some embodiments, a CasX variant protein with improved target nucleic acid affinity has increased affinity for specific PAM sequences other than the canonical TTC PAM recognized by the reference CasX protein of SEQ ID NO: 2, including binding affinity for PAM sequences selected from the group consisting of TTC, ATC, GTC, and CTC. A higher overall affinity for DNA also, in some embodiments, can increase the frequency at which a CasX protein can effectively start and finish 15 05 25 a binding and unwinding step, thereby facilitating target strand invasion and R-loop formation, and ultimately the cleavage of a target nucleic acid sequence.

[0169] In some embodiments, a CasX variant protein has improved binding affinity for the non-target strand of the target nucleic acid. As used herein, the term “non-target strand” refers to the strand of the DNA target nucleic acid sequence that does not form Watson and Crick base pairs with the targeting sequence in the gRNA and is complementary to the target DNA strand. In some embodiments, the CasX variant protein has about 1.1 to about 100-fold increased binding affinity to the non-target stand of the target nucleic acid compared to the reference protein of SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 3.

[0170] Methods of measuring CasX variant protein affinity for a target nucleic acid molecule may include electrophoretic mobility shift assays (EMSAs), filter binding, isothermal calorimetry (ITC), and surface plasmon resonance (SPR), fluorescence polarization and biolayer interferometry (BLI). Further methods of measuring CasX protein affinity for a target include in vitro biochemical assays that measure DNA cleavage events over time; e.g., determination of the kcieave rate, as described in the Examples. f Improved Specificity for a Target Site

[0171] In some embodiments, a CasX variant protein has improved specificity for a target nucleic acid sequence relative to a reference CasX protein. As used herein, “specificity,” interchangeably referred to as “target specificity,” refers to the degree to which a CRISPR / Cas system ribonucleoprotein complex cleaves off-target sequences that are similar, but not identical to the target nucleic acid sequence; e.g., a CasX variant RNP with a higher degree of specificity would exhibit reduced off-target cleavage of sequences relative to a reference CasX protein. The specificity, and the reduction of potentially deleterious off-target effects, of CRISPR / Cas system proteins can be vitally important in order to achieve an acceptable therapeutic index for use in mammalian subjects.

[0172] In some embodiments, a CasX variant protein has improved specificity for a target site within the target sequence that is complementary to the targeting sequence of the gRNA relative to a reference CasX protein of SEQ ID NOS: 1-3. Without wishing to be bound by theory, it is possible that amino acid changes in the helical I and II domains that increase the specificity of the CasX variant protein for the target nucleic acid strand can increase the specificity of the CasX variant protein for the target nucleic acid overall. In some embodiments, amino acid 15 05 25 changes that increase specificity of CasX variant proteins for target nucleic acid may also result in decreased affinity of CasX variant proteins for DNA.

[0173] Methods of testing CasX protein (such as variant or reference) target specificity may include guide and Circularization for In vitro Reporting of Cleavage Effects by Sequencing (CIRCLE-seq), or similar methods. In brief, in CIRCLE-seq techniques, genomic DNA is sheared and circularized by ligation of stem-loop adapters, which are nicked in the stem-loop regions to expose 4 nucleotide palindromic overhangs. This is followed by intramolecular ligation and degradation of remaining linear DNA. Circular DNA molecules containing a CasX cleavage site are subsequently linearized with CasX, and adapter adapters are ligated to the exposed ends followed by high-throughput sequencing to generate paired end reads that contain information about the off-target site. Additional assays that can be used to detect off-target events, and therefore CasX protein specificity include assays used to detect and quantify indels (insertions and deletions) formed at those selected off-target sites such as mismatch-detection nuclease assays and next generation sequencing (NGS). Exemplary mismatch-detection assays include nuclease assays, in which genomic DNA from cells treated with CasX and sgRNA is PCR amplified, denatured and rehybridized to form hetero-duplex DNA, containing one wildtype strand and one strand with an indel. Mismatches are recognized and cleaved by mismatch detection nucleases, such as Surveyor nuclease or T7 endonuclease I. g. Protospacer and PAM Sequences

[0174] Herein, the protospacer is defined as the DNA sequence complementary to the targeting sequence of the guide RNA and the DNA complementary to that sequence, referred to as the target strand and non-target strand, respectively. As used herein, the PAM is a nucleotide sequence located is located 1 nucleotide 5' of the sequence in the non-target strand that is complementary to the target nucleic acid sequence in the target strand of the target nucleic acid that, in conjunction with the targeting sequence of the gRNA, helps the orientation and positioning of the CasX for the potential cleavage of the protospacer strand(s). PAM sequences may be degenerate, and specific RNP constructs may have different preferred and tolerated PAM sequences that support different efficiencies of cleavage. Following convention, unless stated otherwise, the disclosure refers to both the PAM and the protospacer sequence and their directionality according to the orientation of the non-target strand. This does not imply that the PAM sequence of the non-target strand, rather than the target strand, is determinative of cleavage or mechanistically involved in target recognition. For example, when 15 05 25 reference is to a TTC PAM, it may in fact be the complementary GAA sequence that is required for target cleavage, or it may be some combination of nucleotides from both strands. In the case of the CasX proteins disclosed herein, the PAM is located 5’ of the protospacer with a single nucleotide separating the PAM from the first nucleotide of the protospacer. Thus, in the case of reference CasX, a TTC PAM should be understood to mean a sequence following the formula 5’-...NNTTCN(protospacer)NNNNNN...3’ where ‘N’ is any DNA nucleotide and ‘(protospacer)’ is a DNA sequence having identity with the targeting sequence of the guide RNA. In the case of a CasX variant with expanded PAM recognition, a TTC, CTC, GTC, or ATC PAM should be understood to mean a sequence following the formulae: 5’- .. ,NNTTCN(protospacer)NNNNNN.. .3’; 5 ’ -... NNCTCN(protospacer)NNNNNN... 3 ’; 5’-...NNGTCN(protospacer)NNNNNN...3’; or 5’-.. ,NNATCN(protospacer)NNNNNN.. .3’.

[0175] Alternatively, a TC PAM should be understood to mean a sequence following the formula: 5’-.. .NNNTCN(protospacer)NNNNNN. .3’Additionally, the CasX variant proteins of the disclosure have an enhanced ability to efficiently edit and / or bind target DNA, when complexed with a gRNA as an RNP, utilizing a PAM TC motif, including PAM sequences selected from TTC, ATC, GTC, or CTC, (in a 5’ to 3’ orientation), compared to an RNP of a reference CasX protein and reference gRNA. In the foregoing, the PAM sequence is located at least 1 nucleotide 5’ to the non-target strand of the protospacer having identity with the targeting sequence of the gRNA in an assay system compared to the editing efficiency and / or binding of an RNP comprising a reference CasX protein and reference gRNA in a comparable assay system. In one embodiment, an RNP of a CasX variant and gRNA variant exhibits greater editing efficiency and / or binding of a target sequence in the target DNA compared to an RNP comprising a reference CasX protein and a reference gRNA in a comparable assay system, wherein the PAM sequence of the target DNA is TTC. In another embodiment, an RNP of a CasX variant and gRNA variant exhibits greater editing efficiency and / or binding of a target sequence in the target DNA compared to an RNP comprising a reference CasX protein and a reference gRNA in a comparable assay system, wherein the PAM sequence of the target DNA is ATC. In another embodiment, an RNP of a CasX variant and gRNA variant exhibits greater editing efficiency and / or binding of a target sequence in the target DNA compared to an RNP comprising a reference CasX protein and a reference gRNA in a comparable assay system. 15 05 25 wherein the PAM sequence of the target DNA is CTC. In another embodiment, an RNP of a CasX variant and gRNA variant exhibits greater editing efficiency and / or binding of a target sequence in the target DNA compared to an RNP comprising a reference CasX protein and a reference gRNA in a comparable assay system, wherein the PAM sequence of the target DNA is GTC. In the foregoing embodiments, the increased editing efficiency and / or binding affinity for the one or more PAM sequences is at least 1.5-fold greater or more compared to the editing efficiency and / or binding affinity of an RNP of any one of the CasX proteins of SEQ ID NOS: 1-3 and the gRNA of Table 2 for the PAM sequences. Exemplary assays demonstrating the improved editing are described herein, in the Examples. h. Unwinding of DNA

[0176] In some embodiments, a CasX variant protein has improved ability of unwinding DNA relative to a reference CasX protein. Poor dsDNA unwinding has been shown previously to impair or prevent the ability of CRISPR / Cas system proteins AnaCas9 or Cas 14s to cleave DNA. Therefore, without wishing to be bound by any theory, it is likely that increased DNA cleavage activity by some CasX variant proteins of the disclosure is due, at least in part, to an increased ability to find and unwind the dsDNA at a target site.

[0177] Without wishing to be bound by theory, it is thought that amino acid changes in the NT SB domain may produce CasX variant proteins with increased DNA unwinding characteristics. Alternatively, or in addition, amino acid changes in the OBD or the helical domain regions that interact with the PAM may also produce CasX variant proteins with increased DNA unwinding characteristics.

[0178] Methods of measuring the ability of CasX proteins (such as variant or reference) to unwind DNA include, but are not limited to, in vitro assays that observe increased on rates of dsDNA targets in fluorescence polarization or biolayer interferometry. i. Catalytic Activity

[0179] The ribonucleoprotein complex of the CasX:gRNA systems disclosed herein comprise a CasX variant that bind to a target nucleic acid sequence and cleaves the target nucleic acid sequence. In some embodiments, a CasX variant protein has improved catalytic activity relative to a reference CasX protein. Without wishing to be bound by theory, it is thought that in some cases cleavage of the target strand can be a limiting factor for Casl2-like molecules in creating a dsDNA break. In some embodiments, CasX variant proteins improve bending of the target 15 05 25 strand of DNA and cleavage of this strand, resulting in an improvement in the overall efficiency of dsDNA cleavage by the CasX ribonucleoprotein complex.

[0180] In some embodiments, a CasX variant protein has increased nuclease activity compared to a reference CasX protein. Variants with increased nuclease activity can be generated, for example, through amino acid changes in the RuvC nuclease domain. In some embodiments, the CasX variant comprises a RuvC nuclease domain having nickase activity. In the foregoing, the CasX nickase of a CasX:gRNA system generates a single-stranded break within 10-18 nucleotides 3' of a PAM site in the non-target strand. In other embodiments, the CasX variant comprises a RuvC nuclease domain having double-stranded cleavage activity. In the foregoing, the CasX of the CasX:gRNA system generates a double-stranded break within 18-26 nucleotides 5' of a PAM site on the target strand and 10-18 nucleotides 3’ on the non-target strand. Nuclease activity can be assayed by a variety of methods, including those of the Examples. In some embodiments, a CasX variant has a kcieave constant that is at least 2-fold, or at least 3-fold, or at least 4-fold, or at least 5-fold, or at least 6-fold, or at least 7-fold, or at least 8-fold, or at least 9-fold, or at least 10-fold greater compared to a reference CasX.

[0181] In some embodiments, a CasX variant protein has the improved characteristic of forming RNP with gRNA that result in a higher percentage of cleavage-competent RNP compared to an RNP of a reference CasX protein of SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 3 and the gRNA, as described in the Examples. By cleavage competent, it is meant that the RNP that is formed has the ability to cleave the target nucleic acid. In some embodiments, the RNP of the CasX variant and the gRNA exhibit at least a 2-fold, or at least a 3-fold, or at least a 4-fold, or at least a 5-fold, or at least a 10-fold cleavage rate compared to an RNP of a reference CasX protein of SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 3 and the gRNA of Table 2. In the foregoing embodiment, the improved competency rate can be demonstrated in an in vitro assay, such as described in the Examples.

[0182] In some embodiments, a CasX variant protein has increased target strand loading for double strand cleavage compared to a reference CasX. Variants with increased target strand loading activity can be generated, for example, through amino acid changes in the TLS domain. Without wishing to be bound by theory, amino acid changes in the TSL domain may result in CasX variant proteins with improved catalytic activity. Alternatively, or in addition, amino acid changes around the binding channel for the RNA:DNA duplex may also improve catalytic activity of the CasX variant protein. In some embodiments, a CasX variant protein has increased 15 05 25 collateral cleavage activity compared to a reference CasX protein. As used herein, “collateral cleavage activity” refers to additional, non-targeted cleavage of nucleic acids following recognition and cleavage of a target nucleic acid sequence. In some embodiments, a CasX variant protein has decreased collateral cleavage activity compared to a reference CasX protein.

[0183] In some embodiments, for example those embodiments encompassing applications where cleavage of the target nucleic acid sequence is not a desired outcome, improving the catalytic activity of a CasX variant protein comprises altering, reducing, or abolishing the catalytic activity of the CasX variant protein. In some embodiments, a ribonucleoprotein complex comprising a dCasX variant protein binds to a target nucleic acid sequence and does not cleave the target nucleic acid.

[0184] In some embodiments, the CasX ribonucleoprotein complex comprising a CasX variant protein binds a target DNA but generates a single stranded nick in the target DNA. In some embodiments, particularly those embodiments wherein the CasX protein is a nickase, a CasX variant protein has decreased target strand loading for single strand nicking. Variants with decreased target strand loading may be generated, for example, through amino acid changes in the TSL domain.

[0185] Exemplary methods for characterizing the catalytic activity of CasX proteins may include, but are not limited to, in vitro cleavage assays, including those of the Examples, below. In some embodiments, electrophoresis of DNA products on agarose gels can interrogate the kinetics of strand cleavage. j. CasX Fusion Proteins

[0186] In some embodiments, the disclosure provides CasX proteins comprising a heterologous protein fused to the CasX. In some cases, the CasX is a reference CasX protein. In other cases, the CasX is a CasX variant of any of the embodiments described herein.

[0187] In some embodiments, the CasX variant protein comprises any one of SEQ ID NOS: 59, 72-99, 101-148, and 26908-27154 of the sequences of Table 4 fused to one or more proteins or domains thereof that has a different activity of interest, resulting in a fusion protein. In some embodiments, the CasX variant protein comprises any one of SEQ ID NOS: 36-99, 101-148, 26908-27154 fused to one or more proteins or domains thereof. In some embodiments, the CasX variant protein comprises any one of SEQ ID NOS: 132-148, 26908-2715 fused to one or more proteins or domains thereof. For example, in some embodiments, the CasX variant protein is 15 05 25 fused to a protein (or domain thereof) that inhibits transcription, modifies a target nucleic acid sequence, or modifies a polypeptide associated with a nucleic acid (e.g., histone modification).

[0188] In some embodiments, a heterologous polypeptide (or heterologous amino acid such as a cysteine residue or a non-natural amino acid) can be inserted at one or more positions within a CasX protein to generate a CasX fusion protein. In other embodiments, a cysteine residue can be inserted at one or more positions within a CasX protein followed by conjugation of a heterologous polypeptide described below. In some alternative embodiments, a heterologous polypeptide or heterologous amino acid can be added at the N- or C-terminus of the CasX variant protein. In other embodiments, a heterologous polypeptide or heterologous amino acid can be inserted internally within the sequence of the CasX protein.

[0189] In some embodiments, the CasX variant fusion protein retains RNA-guided sequence specific target nucleic acid binding and cleavage activity. In some cases, the CasX variant fusion protein has (retains) 50% or more of the activity (e.g., cleavage and / or binding activity) of the corresponding CasX variant protein that does not have the insertion of the heterologous protein. In some cases, the CasX variant fusion protein retains at least about 60%, or at least about 70% or more, at least about 80%, or at least about 90%, or at least about 92%, or at least about 95%, or at least about 98%, or at least about 100% of the activity (e.g., cleavage and / or binding activity) of the corresponding CasX protein that does not have the insertion of the heterologous protein.

[0190] In some cases, the CasX variant fusion protein retains (has) target nucleic acid binding activity relative to the activity of the CasX protein without the inserted heterologous amino acid or heterologous polypeptide. In some cases, the CasX variant fusion protein retains at least about 60%, or at least about 70% or more, at least about 80%, or at least about 90%, or at least about 92%, or at least about 95%, or at least about 98%, or at least about 100% of the binding activity of the corresponding CasX protein that does not have the insertion of the heterologous protein.

[0191] In some cases, the CasX variant fusion protein retains (has) target nucleic acid binding and / or cleavage activity relative to the activity of the parent CasX protein without the inserted heterologous amino acid or heterologous polypeptide. For example, in some cases, the CasX variant fusion protein has (retains) 50% or more of the binding and / or cleavage activity of the corresponding parent CasX protein (the CasX protein that does not have the insertion). For example, in some cases, the CasX variant fusion protein has (retains) 60% or more (70% or more, 80% or more, 90% or more, 92% or more, 95% or more, 98% or more, or 100%) of the 15 05 25 binding and / or cleavage activity of the corresponding CasX parent protein (the CasX protein that does not have the insertion). Methods of measuring cleaving and / or binding activity of a CasX protein and / or a CasX fusion protein will be known to one of ordinary skill in the art and any convenient method can be used.

[0192] A variety of heterologous polypeptides are suitable for inclusion in a reference CasX or CasX variant fusion protein of the disclosure. In some cases, the fusion partner can modulate transcription (e.g., inhibit transcription, increase transcription) of a target DNA. For example, in some cases the fusion partner is a protein (or a domain from a protein) that inhibits transcription (e.g., a transcriptional repressor, a protein that functions via recruitment of transcription inhibitor proteins, modification of target DNA such as methylation, recruitment of a DNA modifier, modulation of histones associated with target DNA, recruitment of a histone modifier such as those that modify acetylation and / or methylation of histones, and the like).

[0193] In some cases the fusion partner is a protein (or a domain from a protein) that increases transcription (e.g., a transcription activator, a protein that acts via recruitment of transcription activator proteins, modification of target DNA such as demethylation, recruitment of a DNA modifier, modulation of histones associated with target DNA, recruitment of a histone modifier such as those that modify acetylation and / or methylation of histones, and the like).In some cases, a fusion partner has enzymatic activity that modifies a target nucleic acid sequence; e.g., nuclease activity, methyltransferase activity, demethylase activity, DNA repair activity, DNA damage activity, deamination activity, dismutase activity, alkylation activity, depurination activity, oxidation activity, pyrimidine dimer forming activity, integrase activity, transposase activity, recombinase activity, polymerase activity, ligase activity, helicase activity, photolyase activity or glycosylase activity. In some embodiments, a CasX variant comprises any one of SEQ ID NOS: 36-99, 101-148, or 26908-27154, or any one of SEQ ID NOS: 59, 72-99, 101-148, or 26908-27154, or any one of SEQ ID NOS 132-148, or 26908-27154, and a polypeptide with methyltransferase activity, demethylase activity, acetyltransferase activity, deacetylase activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitinating activity, adenylation activity, deadenylation activity, SUMOylating activity, deSUMOylating activity, ribosylation activity, deribosylation activity, myristoylation activity or demyristoylation activity.

[0194] In some embodiments, a CasX variant comprises any one of SEQ ID NOS: 36-99, 101-148, and 26908-27154, or any one of SEQ ID NOS: 59, 72-99, 101-148, or 26908-27154, or any one of SEQ ID NOS 132-148, or 26908-27154, and a fusion partner having enzymatic activity 15 05 25 that modifies a polypeptide (e.g., a histone) associated with a target nucleic acid (e.g., methyltransferase activity, demethylase activity, acetyltransferase activity, deacetylase activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitinating activity, adenylation activity, deadenylation activity, SUMOylating activity, deSUMOylating activity, ribosylation activity, deribosylation activity, myristoylation activity or demyristoylation activity). Examples of proteins (or fragments thereof) that can be used as a fusion partner to increase transcription include but are not limited to: transcriptional activators such as VP16, VP64, VP48, VP160, p65 subdomain (e.g., from NFkB), and activation domain of EDLL and / or TAL activation domain (e.g, for activity in plants); histone lysine methyltransferases such as SET1A, SET1B, MLLI to 5, ASH1, SYMD2, NSD1, and the like; histone lysine demethylases such as JHDM2a / b, UTX, JMJD3, and the like; histone acetyltransferases such as GCN5, PCAF, CBP, p300, TAF1, TIP60 / PLIP, M0Z / MYST3, MORF / MYST4, SRC1, ACTR, P160, CLOCK, and the like; and DNA demethylases such as Ten-Eleven Translocation (TET) dioxygenase 1 (TET1CD), TET1, DME, DML1, DML2, ROS1, and the like.

[0195] Examples of proteins (or fragments thereof) that can be used as a fusion partner to decrease transcription include but are not limited to: transcriptional repressors such as the Kruppel associated box (KRAB or SKD); KOX1 repression domain; the Mad mSIN3 interaction domain (SID); the ERF repressor domain (ERD), the SRDX repression domain (e.g., for repression in plants), and the like; histone lysine methyltransferases such as Pr-SET7 / 8, SUV4-20H1, RIZ1, and the like; histone lysine demethylases such as JMJD2A / JHDM3A, JMJD2B, JMJD2C / GASC I, JMJD2D, JARID1A / RBP2, JARID1B / PLU-I, JARID 1C / SMCX, JARID1D / SMCY, and the like; histone lysine deacetylases such as HDAC1, HDAC2, HDAC3, HDAC8, HDAC4, HDAC5, HDAC7, HDAC9, SIRT1, SIRT2, HDAC11, and the like; DNA methylases such as Hhal DNA m5c-methyltransferase (M.Hhal), DNA methyltransferase 1 (DNMT1), DNA methyltransferase 3a (DNMT3a), DNA methyltransferase 3b (DNMT3b), METI, DRM3 (plants), ZMET2, CMT1, CMT2 (plants), and the like; and periphery recruitment elements such as Lamin A, Lamin B, and the like.

[0196] In some cases, the fusion partner to a CasX variant has enzymatic activity that modifies the target nucleic acid sequence (e.g., ssRNA, dsRNA, ssDNA, dsDNA). Examples of enzymatic activity that can be provided by the fusion partner include but are not limited to: nuclease activity such as that provided by a restriction enzyme (e.g., FokI nuclease), methyltransferase activity such as that provided by a methyltransferase (e.g., Hhal DNA m5c-methyltransferase 15 05 25 (M.Hhal), DNA methyltransferase 1 (DNMT1), DNA methyltransferase 3a (DNMT3a), DNA methyltransferase 3b (DNMT3b), METI, DRM3 (plants), ZMET2, CMT1, CMT2 (plants), and the like); demethylase activity such as that provided by a demethylase (e.g., Ten-Eleven Translocation (TET) dioxygenase 1 (TET 1 CD), TET1, DME, DML1, DML2, ROS1, and the like), DNA repair activity, DNA damage activity, deamination activity such as that provided by a deaminase (e.g., a cytosine deaminase enzyme, e.g., an APOBEC protein such as rat APOBEC1), dismutase activity, alkylation activity, depurination activity, oxidation activity, pyrimidine dimer forming activity, integrase activity such as that provided by an integrase and / or resolvase (e.g., Gin invertase such as the hyperactive mutant of the Gin invertase, GinH106Y; human immunodeficiency virus type 1 integrase (IN); Tn3 resolvase; and the like), transposase activity, recombinase activity such as that provided by a recombinase (e.g., catalytic domain of Gin recombinase), polymerase activity, ligase activity, helicase activity, photolyase activity, and glycosylase activity).

[0197] In some cases, a CasX variant protein of the present disclosure is fused to a polypeptide selected from a domain for increasing transcription (e.g., a VP 16 domain, a VP64 domain), a domain for decreasing transcription (e.g., a KRAB domain, e.g., from the Koxl protein), a core catalytic domain of a histone acetyltransferase (e.g., histone acetyltransferase p300), a protein / domain that provides a detectable signal (e.g., a fluorescent protein such as GFP), a nuclease domain (e.g., a Fokl nuclease), or a base editor (e.g., cytidine deaminase such as APOBEC 1).

[0198] In some embodiments, a CasX variant comprises any one of SEQ ID NOS: 36-99, 101-148, or 26908-27154, or any one of SEQ ID NOS: 59, 72-99, 101-148, or 26908-27154, or any one of SEQ ID NOS 132-148, or 26908-27154, and a fusion partner having enzymatic activity that modifies a protein associated with the target nucleic acid (e.g., ssRNA, dsRNA, ssDNA, dsDNA) (e.g., a histone, an RNA binding protein, a DNA binding protein, and the like). Examples of enzymatic activity (that modifies a protein associated with a target nucleic acid) that can be provided by the fusion partner include but are not limited to: methyltransferase activity such as that provided by a histone methyltransferase (HMT) (e.g., suppressor of variegation 3-9 homolog 1 (SUV39H1, also known as KMT1 A), euchromatic histone lysine methyltransferase 2 (G9A, also known as KMT1C and EHMT2), SUV39H2, ESET / SETDB 1, and the like, SET1A, SET1B, MLL1 to 5, ASH1, SYMD2, NSD1, DOT1L, Pr-SET7 / 8, SUV4-20H1, EZH2, R1ZI), demethylase activity such as that provided by a histone demethylase (e.g., 15 05 25 Lysine Demethylase 1A (KDM1A also known as LSD1), JHDM2a / b, JMJD2A / JHDM3A, JMJD2B, JMJD2C / GASC1, JMJD2D, JARID1A / RBP2, JARID1B / PLU-1, JARID1C / SMCX, JARID1D / SMCY, UTX, JMJD3, and the like), acetyltransferase activity such as that provided by a histone acetylase transferase (e.g., catalytic core / fragment of the human acetyltransferase p300, GCN5, PCAF, CBP, TAF1, TIP60 / PLIP, MOZ / MYST3, MORF / MYST4, HB01 / MYST2, HMOF / MYST1, SRC1, ACTR, P160, CLOCK, and the like), deacetylase activity such as that provided by a histone deacetylase (e.g., HDAC1, HDAC2, HDAC3, HDAC8, HDAC4, HDAC5, HDAC7, HDAC9, SIRT1, SIRT2, HDAC11, and the like), kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitinating activity, adenylation activity, deadenylation activity, SUMOylating activity, deSUMOylating activity, ribosylation activity, deribosylation activity, myristoylation activity, and demyristoylation activity.

[0199] Additional examples of suitable fusion partners for a CasX variant are (i) a dihydrofolate reductase (DHFR) destabilization domain (e.g., to generate a chemically controllable subject RNA-guided polypeptide or a conditionally active RNA-guided polypeptide), and (ii) a chloroplast transit peptide. In some embodiments, a CasX variant comprises any one of SEQ ID NOS: 36-99, 101-148, or 26908-27154, or any one of SEQ ID NOS: 59, 72-99, 101-148, or 26908-27154, or any one of SEQ ID NOS 132-148, or 26908-27154, or a sequence of Table 4, and a chloroplast transit peptide including, but are not limited to: MASMISSSAVTTVSRASRGQSAAMAPFGGLKSMTGFPVRKVNTDITSITSNGGR VKCMQVWPPIGKKKFETLSYLPPLTRDSRA (SEQ ID NO: 154); MASMISSSAVTTVSRASRGQSAAMAPFGGLKSMTGFPVRKVNTDITSITSNGGRVKS (SEQ ID NO: 155); MAS SMLS S ATM VASPAQ ATMVAPFNGLKS S A AFP ATRKANNDITSITSNGGRVNCMQ V WPPIEKKKFETLSYLPDLTDSGGRVNC (SEQ ID NO: 156); MAQVSRICNGVQNPSLISNLSKSSQRKSPLSVSLKTQQHPRAYPISSSWGLKKSGMTLIG SELRPLKVMSSVSTAC (SEQ ID NO: 157); MAQVSRICNGVWNPSLISNLSKSSQRKSPLSVSLKTQQHPRAYPISSSWGLKKSGMTLIG SELRPLKVMSSVSTAC (SEQ ID NO: 158); M AQINNM AQGIQTLNPNSNFHKPQ VPK S S SFLVFGSKKLKNS ANSMLVLKKDSIFMQLF CSFRISASVATAC (SEQ ID NO: 159); MAALVTSQLATSGTVLSVTDRFRRPGFQGLRPRNPADAALGMRTVGASAAPKQSRKPH RFDRRCLSMVV (SEQ ID NO: 160); 15 05 25 MAALTTSQLATSATGFGIADRSAPSSLLRHGFQGLKPRSPAGGDATSLSVTTSARATPKQ QRSVQRGSRRFPSVVVC (SEQ ID NO: 161); MASSVLSSAAVATRSNVAQANMVAPFTGLKSAASFPVSRKQNLDITSIASNGGRVQC (SEQ ID NO: 162); MESLAATSVFAPSRVAVPAARALVRAGTVVPTRRTSSTSGTSGVKCSAAVTPQASPVIS RSAAAA (SEQ ID NO: 163); and MGAAATSMQSLKFSNRLVPPSRRLSPVPNNVTCNNLPKSAAPVRTVKCCASSWNSTING AAATTNGASAASS (SEQ ID NO: 164).

[0200] In some cases, a CasX variant protein of the present disclosure can include an endosomal escape peptide. In some cases, an endosomal escape polypeptide comprises the amino acid sequence GLFXALLXLLXSLWXLLLXA (SEQ ID NO: 165), wherein each X is independently selected from lysine, histidine, and arginine. In some cases, an endosomal escape polypeptide comprises the amino acid sequence GLFHALLHLLHSLWHLLLHA (SEQ ID NO: 166), or HHHHHHHHH (SEQ ID NO: 167).

[0201] Non-limiting examples of fusion partners for use with CasX variant proteins when targeting ssRNA target nucleic acid sequences include (but are not limited to): splicing factors (e.g., RS domains); protein translation components (e.g., translation initiation, elongation, and / or release factors; e.g, eIF4G); RNA methylases; RNA editing enzymes (e.g., RNA deaminases, e.g., adenosine deaminase acting on RNA (ADAR), including A to I and / or C to U editing enzymes); helicases; RNA-binding proteins; and the like. It is understood that a heterologous polypeptide can include the entire protein or in some cases can include a fragment of the protein (e.g., a functional domain).

[0202] In some embodiments, a CasX variant comprises any one of SEQ ID NOS: 36-99, 101-148, or 26908-27154, or any one of SEQ ID NOS: 59, 72-99, 101-148, or 26908-27154, or any one of SEQ ID NOS 132-148, or 26908-27154 and a fusion partner of any domain capable of interacting with ssRNA (which, for the purposes of this disclosure, includes intramolecular and / or intermolecular secondary structures, e.g., double-stranded RNA duplexes such as hairpins, stem-loops, etc.), whether transiently or irreversibly, directly or indirectly, including but not limited to an effector domain selected from the group comprising; endonucleases (for example RNase III, the CRR22 DYW domain, Dicer, and PIN (PilT N-terminus) domains from proteins such as SMG5 and SMG6); proteins and protein domains responsible for stimulating RNA cleavage (for example CPSF, CstF, CFIm and CFIIm); exonucleases (for example XRN-1 15 05 25 or Exonuclease T); deadenylases (for example HNT3); proteins and protein domains responsible for nonsense mediated RNA decay (for example UPF1, UPF2, UPF3, UPF3b, RNP SI, Y14, DEK, REF2, and SRml60); proteins and protein domains responsible for stabilizing RNA (for example PABP); proteins and protein domains responsible for repressing translation (for example Ago2 and Ago4); proteins and protein domains responsible for stimulating translation (for example Staufen); proteins and protein domains responsible for (e.g., capable of) modulating translation (e.g., translation factors such as initiation factors, elongation factors, release factors, etc., e.g., eIF4G); proteins and protein domains responsible for polyadenylation of RNA (for example PAPI, GLD-2, and Star- PAP); proteins and protein domains responsible for polyuridinylation of RNA (for example CI DI and terminal uridylate transferase); proteins and protein domains responsible for RNA localization (for example from IMP1, ZBP1, She2p, She3p, and Bicaudal-D); proteins and protein domains responsible for nuclear retention of RNA (for example Rrp6); proteins and protein domains responsible for nuclear export of RNA (for example TAP, NXF1, THO, TREX, REF, and Aly); proteins and protein domains responsible for repression of RNA splicing (for example PTB, Sam68, and hnRNP Al); proteins and protein domains responsible for stimulation of RNA splicing (for example serine / arginine-rich (SR) domains); proteins and protein domains responsible for reducing the efficiency of transcription (for example FUS (TLS)); and proteins and protein domains responsible for stimulating transcription (for example CDK7 and HIV Tat). Alternatively, the effector domain may be selected from the group comprising endonucleases; proteins and protein domains capable of stimulating RNA cleavage; exonucleases; deadenylases; proteins and protein domains having nonsense mediated RNA decay activity; proteins and protein domains capable of stabilizing RNA; proteins and protein domains capable of repressing translation; proteins and protein domains capable of stimulating translation; proteins and protein domains capable of modulating translation (e.g., translation factors such as initiation factors, elongation factors, release factors, etc., e.g., eIF4G); proteins and protein domains capable of polyadenylation of RNA; proteins and protein domains capable of polyuridinylation of RNA; proteins and protein domains having RNA localization activity; proteins and protein domains capable of nuclear retention of RNA; proteins and protein domains having RNA nuclear export activity; proteins and protein domains capable of repression of RNA splicing; proteins and protein domains capable of stimulation of RNA splicing; proteins and protein domains capable of reducing the efficiency of transcription; and proteins and protein domains capable of stimulating transcription. Another suitable 15 05 25 heterologous polypeptide is a PUF RNA-binding domain, which is described in more detail in WO2012068627, which is hereby incorporated by reference in its entirety.

[0203] Some RNA splicing factors that can be used (in whole or as fragments thereof) as a fusion partner with a CasX variant have modular organization, with separate sequence-specific RNA binding modules and splicing effector domains. For example, members of the serine / arginine-rich (SR) protein family contain N-terminal RNA recognition motifs (RRMs) that bind to exonic splicing enhancers (ESEs) in pre-mRNAs and C-terminal RS domains that promote exon inclusion. As another example, the hnRNP protein hnRNP Al binds to exonic splicing silencers (ESSs) through its RRM domains and inhibits exon inclusion through a C-terminal glycine-rich domain. Some splicing factors can regulate alternative use of splice site (ss) by binding to regulatory sequences between the two alternative sites. For example, ASF / SF2 can recognize ESEs and promote the use of intron proximal sites, whereas hnRNP Al can bind to ESSs and shift splicing towards the use of intron distal sites. One application for such factors is to generate ESFs that modulate alternative splicing of endogenous genes, particularly disease associated genes. For example, Bcl-x pre-mRNA produces two splicing isoforms with two alternative 5’ splice sites to encode proteins of opposite functions. The long splicing isoform Bcl-xL is a potent apoptosis inhibitor expressed in long-lived post mitotic cells and is up-regulated in many cancer cells, protecting cells against apoptotic signals. The short isoform Bcl-xS is a pro-apoptotic isoform and expressed at high levels in cells with a high turnover rate (e.g., developing lymphocytes). The ratio of the two Bcl-x splicing isoforms is regulated by multiple cc -elements that are located in either the core exon region or the exon extension region (i.e., between the two alternative 5’ splice sites). For more examples, see WO2010075303, which is hereby incorporated by reference in its entirety.

[0204] Further suitable fusion partners for use with a CasX variant include, but are not limited to proteins (or fragments thereof) that are boundary elements (e.g., CTCF), proteins and fragments thereof that provide periphery recruitment (e.g., Lamin A, Lamin B, etc.), and protein docking elements (e.g., FKBP / FRB, Pill / Abyl, etc.).

[0205] In some cases, a heterologous polypeptide (a fusion partner) for use with a CasX variant provides for subcellular localization, i.e., the heterologous polypeptide contains a subcellular localization sequence (e.g., a nuclear localization signal (NLS) for targeting to the nucleus, a sequence to keep the fusion protein out of the nucleus, e.g., a nuclear export sequence (NES), a sequence to keep the fusion protein retained in the cytoplasm, a mitochondrial 15 05 25 localization signal for targeting to the mitochondria, a chloroplast localization signal for targeting to a chloroplast, an ER retention signal, and the like). In some embodiments, a subject RNA-guided polypeptide or a conditionally active RNA-guided polypeptide and / or subject CasX fusion protein does not include a NLS so that the protein is not targeted to the nucleus (which can be advantageous, e.g., when the target nucleic acid sequence is an RNA that is present in the cytosol). In some embodiments, a fusion partner can provide a tag (i.e., the heterologous polypeptide is a detectable label) for ease of tracking and / or purification (e.g., a fluorescent protein, e.g., green fluorescent protein (GFP), yellow fluorescent protein (YFP), red fluorescent protein (RFP), cyan fluorescent protein (CFP), mCherry, tdTomato, and the like; a histidine tag, e.g., a 6XHis tag; a hemagglutinin (HA) tag; a FLAG tag; a Myc tag; and the like).

[0206] In some cases, non-limiting examples of NLSs suitable for use with a CasX variant include sequences having at least about 80%, at least about 90%, or at least about 95% identity or are identical to sequences derived from: the NLS of the SV40 virus large T-antigen, having the amino acid sequence PKKKRKV (SEQ ID NO: 168); the NLS from nucleoplasmin (e.g., the nucleoplasmin bipartite NLS with the sequence KRPAATKKAGQAKKKK (SEQ ID NO: 169); the c-myc NLS having the amino acid sequence PAAKRVKLD (SEQ ID NO: 170) or RQRRNELKRSP (SEQ ID NO: 171); the hRNPAl M9 NLS having the sequence NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY (SEQ ID NO: 172); the sequence RMRIZFKNKGKDTAELRRRRVEVSVELRKAKKDEQILKRRNV (SEQ ID NO: 173) of the IBB domain from importin-alpha; the sequences VSRKRPRP (SEQ ID NO: 174) and PPKKARED (SEQ ID NO: 175) of the myoma T protein; the sequence PQPKKKPL (SEQ ID NO: 176) of human p53; the sequence SALIKKKKKMAP (SEQ ID NO: 177) of mouse c-abl IV; the sequences DRLRR (SEQ ID NO: 178) and PKQKKRK (SEQ ID NO: 179) of the influenza virus NS1; the sequence RKLKKKIKKL (SEQ ID NO: 180) of the Hepatitis virus delta antigen; the sequence REKKKFLKRR (SEQ ID NO: 181) of the mouse Mxl protein; the sequence KRKGDEVDGVDEVAKKKSKK (SEQ ID NO: 182) of the human poly(ADP-ribose) polymerase; the sequence RKCLQAGMNLEARKTKK (SEQ ID NO: 183) of the steroid hormone receptors (human) glucocorticoid; the sequence PRPRKIPR (SEQ ID NO: 184) of Borna disease virus P protein (BDV-P1); the sequence PPRKKRTVV (SEQ ID NO: 185) of hepatitis C virus nonstructural protein (HCV-NS5A);the sequence NLSKKKKRKREK (SEQ ID NO: 186) of LEF1; the sequence RRPSRPFRKP (SEQ ID NO: 187) of ORF57 simirae; the sequence KRPRSPSS (SEQ ID NO: 188) of EBV LANA; the sequence 15 05 25 KRGINDRNFWRGENERKTR (SEQ ID NO: 189) of Influenza A protein; the sequence PRPPKMARYDN (SEQ ID NO: 190) of human RNA helicase A (RHA); the sequence KRSFSKAF (SEQ ID NO: 191) of nucleolar RNA helicase II; the sequence KLKIKRPVK (SEQ ID NO: 192) of TUS-protein; the sequence PKKKRKVPPPPAAKRVKLD (SEQ ID NO: 193) associated with importin-alpha; the sequence PKTRRRPRRSQRKRPPT (SEQ ID NO:26792) from the Rex protein in HTLV-1; the sequence SRRRKANPTKLSENAKKLAKEVEN (SEQ ID NO: 194) from the EGL-13 protein of Caenorhabditis elegans; and the sequences KTRRRPRRSQRKRPPT (SEQ ID NO: 195), RRKKRRPRRKKRR (SEQ ID NO: 196), PKKKSRKPKKKSRK (SEQ ID NO: 197), HKKKHPDASVNFSEFSK (SEQ ID NO: 198), QRPGPYDRPQRPGPYDRP (SEQ ID NO: 199), LSPSLSPLLSPSLSPL (SEQ ID NO: 200), RGKGGKGLGKGGAKRHRK (SEq NQ 2()1) pKRGRGRpKRGRGR (SEq ID N0. 202), PKKKRKVPPPPAAKRVKLD (SEQ ID NO: 203), PKKKRKVPPPPKKKRKV (SEQ ID NO: 204), PAKRARRGYKC (SEQ ID NO: 27199), KLGPRKATGRW (SEQ ID NO: 27200), PRRKREE (SEQ ID NO: 27201), PYRGRKE (SEQ ID NO: 27202), PLRKRPRR (SEQ ID NO: 27203), PLRKRPRRGSPLRKRPRR (SEQ ID NO: 27204), PAAKRVKLDGGKRTADGSEFESPKKKRKV (SEQ ID NO: 27205), PAAKRVKLDGGKRTADGSEFESPKKKRKVGIHGVPAA (SEQ ID NO: 27206), PAAKRVKLDGGKRTADGSEFESPKKKRKVAEAAAKEAAAKEAAAKA (SEQ ID NO: 207), PAAKRVKLDGGKRTADGSEFESPKKKRKVPG (SEQ ID NO: 27208), KRKGSPERGERKRHW (SEQ ID NO: 27209), KRTADSQHSTPPKTKRKVEFEPKKKRKV (SEQ ID NO: 27210), and PKKKRKVGGSKRTADSQHSTPPKTKRKVEFEPKKKRKV (SEQ ID NO: 27211). In some embodiments, the one or more NLS are linked to the CRISPR protein or to adjacent NLS with a linker peptide wherein the linker peptide is selected from the group consisting of RS, (G)n (SEQ ID NO: 27212), (GS)n (SEQ ID NO: 27213), (GSGGS)n (SEQ ID NO: 214), (GGSGGS)n (SEQ ID NO: 215), (GGGS)n (SEQ ID NO: 216), GGSG (SEQ ID NO: 217), GGSGG (SEQ ID NO: 218), GSGSG (SEQ ID NO: 219), GSGGG (SEQ ID NO: 220), GGGSG (SEQ ID NO: 221), GSSSG (SEQ ID NO: 222), GPGP (SEQ ID NO: 223), GGP, PPP, PPAPPA (SEQ ID NO: 224), PPPG (SEQ ID NO: 27214), PPPGPPP (SEQ ID NO: 225), PPP(GGGS)n (SEQ ID NO: 27215), (GGGS)nPPP (SEQ ID NO: 27216), AEAAAKEAAAKEAAAKA (SEQ ID NO: 27217), and TPPKTKRKVEFE (SEQ ID NO: 27218), where n is 1 to 5. In general, NLS (or multiple NLSs) are of sufficient strength to drive accumulation of a CasX variant fusion protein in the nucleus of a eukaryotic cell. Detection of 15 05 25 accumulation in the nucleus may be performed by any suitable technique. For example, a detectable marker may be fused to a CasX variant fusion protein such that location within a cell may be visualized. Cell nuclei may also be isolated from cells, the contents of which may then be analyzed by any suitable process for detecting protein, such as immunohistochemistry, Western blot, or enzyme activity assay. Accumulation in the nucleus may also be determined indirectly.

[0207] In general, NLS (or multiple NLSs) are of sufficient strength to drive accumulation of an expressed CasX variant fusion protein in the nucleus of a eukaryotic cell. Detection of accumulation in the nucleus may be performed by any suitable technique. For example, a detectable marker may be fused to a CasX variant fusion protein such that location within a cell may be visualized. Cell nuclei may also be isolated from cells, the contents of which may then be analyzed by any suitable process for detecting protein, such as immunohistochemistry, Western blot, or enzyme activity assay. Accumulation in the nucleus may also be determined indirectly.

[0208] In some cases, a CasX variant fusion protein includes a “Protein Transduction Domain” or PTD (also known as a CPP - cell penetrating peptide), which refers to a protein, polynucleotide, carbohydrate, or organic or inorganic compound that facilitates traversing a lipid bilayer, micelle, cell membrane, organelle membrane, or vesicle membrane. A PTD attached to another molecule, which can range from a small polar molecule to a large macromolecule and / or a nanoparticle, facilitates the molecule traversing a membrane, for example going from an extracellular space to an intracellular space, or from the cytosol to within an organelle. In some embodiments, a PTD is covalently linked to the amino terminus of a CasX variant fusion protein. In some embodiments, a PTD is covalently linked to the carboxyl terminus of a CasX variant fusion protein. In some cases, the PTD is inserted internally in the sequence of a CasX variant fusion protein at a suitable insertion site. In some cases, a CasX variant fusion protein includes (is conjugated to, is fused to) one or more PTDs (e.g., two or more, three or more, four or more PTDs). In some cases, a PTD includes one or more nuclear localization signals (NLS). Examples of PTDs include but are not limited to peptide transduction domain of HIV TAT comprising YGRKKRRQRRR (SEQ ID NO: 205), RKKRRQRR (SEQ ID NO: 206); YARAAARQARA (SEQ ID NO: 207); THRLPRRRRRR (SEQ ID NO: 208); and GGRRARRRRRR (SEQ ID NO: 209); a polyarginine sequence comprising a number of arginine residues sufficient to direct entry into a cell (e.g., 3, 4, 5, 6, 7, 8, 9, 10, or 10-50 arginine 15 05 25 residues (SEQ ID NO: 26793); a VP22 domain (Zender et al. (2002) Cancer Gene Ther. 9(6):489-96); an Drosophila Antennapedia protein transduction domain (Noguchi et al. (2003) Diabetes 52(7): 1732-1737); a truncated human calcitonin peptide (Trehin et al. (2004) Pharm. Research 21 :1248-1256); polylysine (Wender et al. (2000) Proc. Natl. Acad. Sci. USA 97: 13003-13008); RRQRRTSKLMKR (SEQ ID NO: 210); Transportan GWTLNSAGYLLGKINLKALAALAKKIL (SEQ ID NO: 211); KALAWEAI<LAI<ALAI<ALAI<HLAI<ALAKALI<CEA (SEQ ID NO: 212); and RQIKIWFQNRRMKWKK (SEQ ID NO: 213). In some embodiments, the PTD is an activatable CPP (ACPP) (Aguilera et al. (2009) Integr Biol (Camb) June; 1(5-6): 371-381). ACPPs comprise a polycationic CPP (e.g., Arg9 or “R9”) connected via a cleavable linker to a matching polyanion (e.g., Glu9 or “E9”), which reduces the net charge to nearly zero and thereby inhibits adhesion and uptake into cells. Upon cleavage of the linker, the polyanion is released, locally unmasking the polyarginine and its inherent adhesiveness, thus “activating” the ACPP to traverse the membrane.

[0209] In some embodiments, a CasX variant fusion protein for use in the systems can include a CasX protein that is linked to an internally inserted heterologous amino acid or heterologous polypeptide (a heterologous amino acid sequence) via a linker polypeptide (e.g., one or more linker polypeptides). In some embodiments, a CasX variant fusion protein can be linked at the C-terminal and / or N-terminal end to a heterologous polypeptide (fusion partner) via a linker polypeptide (e.g., one or more linker polypeptides). The linker polypeptide may have any of a variety of amino acid sequences. Proteins can be joined by a spacer peptide, generally of a flexible nature, although other chemical linkages are not excluded. Suitable linkers include polypeptides of between 4 amino acids and 40 amino acids in length, or between 4 amino acids and 25 amino acids in length. These linkers are generally produced by using synthetic, linkerencoding oligonucleotides to couple the proteins. Peptide linkers with a degree of flexibility can be used. The linking peptides may have virtually any amino acid sequence, bearing in mind that the preferred linkers will have a sequence that results in a generally flexible peptide. The use of small amino acids, such as glycine and alanine, are of use in creating a flexible peptide. The creation of such sequences is routine to those of skill in the art. A variety of different linkers are commercially available and are considered suitable for use. Exemplary linker polypeptides include peptides selected from the group consisting of RS, (G)n (SEQ ID NO: 27212), (GS)n (SEQ ID NO: 27213), (GSGGS)n (SEQ ID NO: 214), (GGSGGS)n (SEQ ID NO: 215), 15 05 25 (GGGS)n (SEQ ID NO: 216), where n is an integer of 1 to 5, GGSG (SEQ ID NO: 217), GGSGG (SEQ ID NO: 218), GSGSG (SEQ ID NO: 219), GSGGG (SEQ ID NO: 220), GGGSG (SEQ ID NO: 221), GSSSG (SEQ ID NO: 222), GPGP (SEQ ID NO: 223), GGP, PPP, PPAPPA (SEQ ID NO: 224), PPPG (SEQ ID NO: 27214), PPPGPPP (SEQ ID NO: 225), PPP(GGGS)n (SEQ ID NO: 27215), (GGGS)nPPP (SEQ ID NO: 27216), AEAAAKEAAAKEAAAKA (SEQ ID NO: 27217), and TPPKTKRKVEFE (SEQ ID NO: 27218), where n is 1 to 5. and the like. The ordinarily skilled artisan will recognize that design of a peptide conjugated to any elements described above can include linkers that are all or partially flexible, such that the linker can include a flexible linker as well as one or more portions that confer less flexible structure. V. Systems and Methods for Modification of BCL11A Genes

[0210] The CRISPR proteins, guide nucleic acids, and variants thereof provided herein are useful for various applications, including as therapeutics, diagnostics, and for research. In some embodiments, to effect the methods of the disclosure for gene editing, provided herein are programmable CasX:gRNA systems. The programmable nature of the systems provided herein allows for the precise targeting to achieve the desired modification at one or more regions of predetermined interest in the BCL11A gene target nucleic acid. A variety of strategies and methods can be employed to modify the target nucleic acid sequence in a cell using the systems provided herein. As used herein "modifying" includes, but is not limited to, cleaving, nicking, editing, deleting, knocking out, knocking down, mutating, correcting, exon-skipping and the like. Depending on the system components utilized, the editing event may be a cleavage event followed by introducing random insertions or deletions (indels) or other mutations (e.g., a substitution, duplication, or inversion of one or more nucleotides), for example by utilizing the imprecise non-homologous DNA end joining (NHEJ) repair pathway, which may generate, for example, a frame shift mutation. Alternatively, the editing event may be a cleavage event followed by homology-directed repair (HDR), homology-independent targeted integration (HITI), micro-homology mediated end joining (MMEJ), single strand annealing (SSA) or base excision repair (BER), resulting in modification of the target nucleic acid sequence.

[0211] In some embodiments of the method, the BCL11A gene to be modified comprises a sequence corresponding to a polynucleotide encoding all or a portion of the sequence of SEQ ID NO: 100 or comprises a polynucleotide sequence that spans all or a portion of chr2 60450520-60554467 (GRCh38 / hg38 Ensembl 100) of the human genome on chromosome 2. In other 15 05 25 embodiments of the method, the target nucleic acid sequence to be modified includes regions of the BCL11A gene encoding the BCL11A protein, a BCL11A regulatory element, a non-coding region of the BCL11A gene, or overlapping portions thereof. In a particular embodiment of the method, the target nucleic acid sequence to be modified comprises the GATA1 binding motif sequence or its complement.

[0212] In some embodiments, the disclosure provides methods of modifying a BCL11A target nucleic acid in a cell, the method comprising introducing into the cell a Class 2, Type V CRISPR system. In some embodiments of the methods, the cells to be modified are autologous with respect to a subject to be administered said cell(s). In other embodiments, the cells to be modified are allogeneic with respect to a subject to be administered said cell(s). Thus, the systems and methods described herein can be used to engineer a variety of cells in which mutations exist in the P-globin gene and are associated with disease, e.g., hemoglobinopathies, including sickle-cell disease and a- and P-thalassemias. This approach, therefore, can be used to modify cells for applications in a subject with a hemoglobinopathy-related disease such as, but not limited to sickle-cell disease and a- and P-thalassemias.

[0213] In some embodiments, the disclosure provides methods of modifying a BCL11A target nucleic acid in a cell, the method comprising introducing into the cell: i) a CasX:gRNA system comprising a CasX and a gRNA of any one of the embodiments described herein; ii) a CasX:gRNA system comprising a CasX, a gRNA, and a donor template of any one of the embodiments described herein; iii) a nucleic acid encoding the CasX and the gRNA, and optionally comprising the donor template ; iv) a vector comprising the nucleic acid of (iii), above; v) an XDP comprising the CasX:gRNA system of any one of the embodiments described herein; or vi) combinations of two or more of (i) to (v), wherein the target nucleic acid sequence of the cells is modified by the CasX protein and, optionally, the donor template. In some embodiments, the vector is an AAV vector. In some embodiments, the disclosure provides CasX:gRNA systems for use in the methods of modifying the BCL11A gene in a cell, wherein the system comprises a CasX variant selected from the group consisting of SEQ ID NOS: 36-99, 101-148, and 26908-27154, or a CasX variant selected from the group consisting of SEQ ID NOS: 59, 72-99, 101-148, and 26908-27154, or a CasX variant selected from the group consisting of SEQ ID NOS 132-148, and 26908-27154, or a variant sequence at least 60% identical, at least 70% identical, at least 80% identical, at least 81% identical, at least 82% identical, at least 83% identical, at least 84% identical, at least 85% identical, at least 86% 15 05 25 identical, at least 86% identical, at least 87% identical, at least 88% identical, at least 89% identical, at least 89% identical, at least 90% identical, at least 91% identical, at least 92% identical, at least 93% identical, at least 94% identical, at least 95% identical, at least 96% identical, at least 97% identical, at least 98% identical, at least 99% identical, or at least 99.5% identical thereto, the gRNA scaffold comprises a sequence selected from the group consisting of SEQ ID NOS: 2101-2285, 26794-26839 and 27219-27265 as set forth in Table 3 or from the group consisting of SEQ ID NOS: 2281-2285, 26794-26839 and 27219-27265, or a sequence at least 65% identical, at least 70% identical, at least 75% identical, at least 80% identical, at least 81% identical, at least 82% identical, at least 83% identical, at least 84% identical, at least 85% identical, at least 86% identical, at least 86% identical, at least 87% identical, at least 88% identical, at least 89% identical, at least 89% identical, at least 90% identical, at least 91% identical, at least 92% identical, at least 93% identical, at least 94% identical, at least 95% identical, at least 96% identical, at least 97% identical, at least 98% identical, at least 99% identical, at least 99.5% identical thereto, and the gRNA comprises a targeting sequence selected from the group consisting of SEQ ID NOS: 272-2100 or 2286-26789, or a sequence at least 65% identical, at least 70% identical, at least 75% identical, at least 80% identical, at least 85% identical, at least 90% identical, or at least 95% identical thereto and having between 15 and 20 nucleotides. In particular embodiments, the targeting sequence of the gRNA is complementary to, and therefore is capable of hybridizing with, a sequence within the GATA1 binding motif sequence or that is 5’ or 3’ to the GATA1 binding motif sequence. In one embodiment, the targeting sequence of the gRNA is UGGAGCCUGUGAUAAAAGCA (SEQ ID NO: 22), which hybridizes with the BCL11A GATA1 erythroid-specific enhancer binding site sequence, or is a sequence having at least 90% or at least 95% sequence identity thereto. In another embodiment, the targeting sequence of the gRNA is UGCUUUUAUCACAGGCUCCA (SEQ ID NO: 23), which hybridizes with a sequence that is complementary to the reverse complement of the BCL11A GATA1 erythroid-specific enhancer binding site sequence, or is a sequence having at least 90% or at least 95% sequence identity thereto. In another particular embodiment, the targeting sequence of the gRNA is complementary to, and therefore is capable of hybridizing with a sequence within the promoter of the BCL11A gene. In one embodiment of the method, the CasX and gRNA are associated together in a ribonuclear protein complex (RNP). In some embodiments of the method of modifying a BCL11A target nucleic acid sequence in a cell, the modification comprises introducing a single-stranded break in the target nucleic acid sequence. 15 05 25 In other embodiments of the method, the modification comprises introducing a double-stranded break in the target nucleic acid sequence. In some embodiments of the method, the modifying comprises introducing an insertion, deletion, substitution, duplication, or inversion of one or more nucleotides in the target nucleic acid sequence. As described herein, a CasX variant introducing double-stranded cleavage of the target nucleic acid generates a double-stranded break within 18-26 nucleotides 5’ of a PAM site on the target strand and 10-18 nucleotides 3’ on the non-target strand. Thus, in some embodiments, the resulting modification by the method can result in random insertions or deletions (indels), or a substitution, duplication, or inversion of one or more nucleotides in those region by non-homologous DNA end joining (NHEJ) repair mechanisms.

[0214] In other embodiments of the method of modifying a BCL11A target nucleic acid sequence in a cell, the method comprises contacting the target nucleic acid sequence with a CasX:gRNA system with a first and a second, or a plurality of gRNAs targeted to different or overlapping portions of the BCL11A gene (e.g., wherein the targeting sequence of the second gRNA is complementary to a sequence that is 5’ or 3’ to the GATA1 binding site) wherein the CasX protein introduces multiple breaks in the target nucleic acid that result in a permanent indel or mutation in the target nucleic acid, as described herein, or an excision of the GATA1 binding motif sequence with a corresponding modulation of expression or alteration in the function of the BCL11A gene product, thereby creating an edited cell. In some cases of the foregoing, the plurality of the gRNAs target locations 5’ and 3’ relative to the GATA1 binding motif sequence of the BCL11A gene such that some or all of the GATA1 binding motif sequence is excised from the target gene between the dual cut sites targeted by the two gRNA. It will be understood that the foregoing embodiments of the method can also be effected by use of encoding nucleic acids, vectors comprising the encoding acids, or XDP comprising the CasX:gRNA system components.

[0215] In some embodiments, the methods of the disclosure provide CasX protein and gRNA pairs that generate site-specific double strand breaks (DSBs) or single strand breaks (SSBs) (e.g., when the CasX protein is a nickase that can cleave only one strand of a target nucleic acid) within 18-24 nucleotides 3' of a PAM site, which can then be repaired either by non-homologous end joining (NHEJ), homology-directed repair (HDR), homology-independent targeted integration (HITI), micro-homology mediated end joining (MMEJ), single strand annealing (SSA) or base excision repair (BER), wherein the modification of the BCL11A gene comprises 15 05 25 introducing an insertion, a deletion, an inversion, or a duplication mutation of one or more nucleotides as compared to the wild-type sequence, with a corresponding modulation of expression or alteration in the function of the BCL11A gene product, thereby creating an edited cell.

[0216] In some cases, the CasX:gRNA system for use in the methods of modifying the BCL11A gene further comprises a donor template nucleic acid of any of the embodiments disclosed herein, wherein the donor template can be inserted by the homology-directed repair (HDR) or homology-independent targeted integration (HITI) repair mechanisms of the host cell. Thus, in some cases, the methods provided herein include contacting the BCL11A gene with a donor template by introducing the donor template (either in vitro inside a cell or in vivo inside a cell), wherein the donor template, a portion of the donor template, a copy of the donor template, or a portion of a copy of the donor template integrates into the BCL11A gene to replace a portion of the BCL11A gene. The donor template can be a short single-stranded or doublestranded oligonucleotide, or a long single-stranded or double-stranded oligonucleotide. In some embodiments, the donor template comprises at least a portion of the BCL11A gene, wherein the BCL11A gene portion is selected from the group consisting of a BCL11A exon, a BCL11A intron, a BCL11A intron-exon junction, a BCL11A regulatory element, or a combination thereof. In some embodiments, the disclosure provides donor templates for use in targeting, or disrupting, the transcriptional activator GATA1 binding site in the BCL11A target sequence wherein the donor template includes sequences that are nonhomologous to regions of DNA within or near GATA1 site in the BCL11A gene, flanked by two regions of homology (“homologous arms”) to the 5’ and 3’ sides of the break site(s) such that the repair mechanisms between the target DNA region and the two flanking sequences results in insertion of the donor template at the target region to facilitate insertion by HDR. The donor template may contain one or more single base changes, insertions, deletions, inversions or rearrangements with respect to the genomic sequence, provided that there is sufficient homology with the target nucleic acid sequence to support its integration into the target nucleic acid, which can result in a frame-shift or other mutation such that the BCL11A protein is not expressed (a knock-out) or is expressed at a lower level (a knock-down). The exogenous donor template inserted by HITI can be any length, for example, a relatively short sequence of between 10 and 50 nucleotides in length, or a longer sequence of about 50-1000 nucleotides in length. The lack of homology can be, for example, having no more than 20-50% sequence identity and / or lacking in specific hybridization 15 05 25 at low stringency. In other cases, the lack of homology can further include a criterion of having no more than 5, 6, 7, 8, or 9 bp identity. In some embodiments, the donor template polynucleotide comprises at least about 10, at least about 50, at least about 100, or at least about 200, or at least about 300, or at least about 400, or at least about 500, or at least about 600, or at least about 700, or at least about 800, or at least about 900, or at least about 1000, or at least about 10,000, or at least about 15,000 nucleotides. In other embodiments, the donor template comprises at least about 10 to about 15,000 nucleotides, or at least about 100 to about 10,000 nucleotides, or at least about 400 to about 8,000 nucleotides, or at least about 600 to about 5000 nucleotides, or at least about 1000 to about 2000 nucleotides. The donor template sequence may comprise certain sequence differences as compared to the genomic sequence, e.g., restriction sites, nucleotide polymorphisms, selectable markers (e.g., drug resistance genes, fluorescent proteins, enzymes etc.), etc., which may be used to assess for successful insertion of the donor nucleic acid at the cleavage site or in some cases may be used for other purposes (e.g., to signify expression at the targeted genomic locus). Alternatively, these sequence differences may include flanking recombination sequences such as FLPs, loxP sequences, or the like, that can be activated at a later time for removal of the marker sequence.

[0217] In some embodiments of the methods of modifying a BCL11A target nucleic acid of a cell in vitro or ex vivo, to induce cleavage or any desired modification to a target nucleic acid, the gRNA and / or the CasX protein of the present disclosure and, optionally, the donor template sequence, whether they be introduced as nucleic acids or polypeptides, complexed RNP, vectors or XDP, are provided to the cells for about 30 minutes to about 24 hours, or at least about 1 hour, 1.5 hours, 2 hours, 2.5 hours, 3 hours, 3.5 hours 4 hours, 5 hours, 6 hours, 7 hours, 8 hours, 12 hours, 16 hours, 18 hours, 20 hours, or any other period from about 30 minutes to about 24 hours, which may be repeated with a frequency of about every day to about every 4 days, e.g., every 1.5 days, every 2 days, every 3 days, or any other frequency from about every day to about every four days. The agent(s) may be provided to the subject cells one or more times, e.g., one time, twice, three times, or more than three times, and the cells allowed to incubate with the agent(s) for some amount of time following each contacting event e.g., 30 minutes to about 24 hours. In the case of in vitro-based methods, after the incubation period with the CasX and gRNA (and optionally the donor template), the media is replaced with fresh media and the cells are cultured further. 15 05 25

[0218] In some embodiments of the methods of modifying a BCL11A target nucleic acid in a cell, the methods further comprises contacting the target nucleic acid sequence of the cell with: a) an additional CRISPR nuclease and a gRNA targeting a different or overlapping portion of the BCL11A target nucleic acid compared to the first gRNA; b) a polynucleotide encoding the additional CRISPR nuclease and the gRNA of (a); c) a vector comprising the polynucleotide of (b); or d) a XDP comprising the additional CRISPR nuclease and the gRNA of (a), wherein the contacting results in modification of the BCL11A target nucleic acid at a different location in the sequence compared to the first gRNA. In some cases, the additional CRISPR nuclease is a CasX protein having a sequence different from the CasX protein of any of the preceding claims. In other cases, the additional CRISPR nuclease is not a CasX protein and is selected from the group consisting of Cas9, Casl2a, Casl2b, Casl2c, Casl2d (CasY), Casl2j, Casl2k, Casl3a, Casl3b, Cas 13c, Cas 13d, CasY, Cas 14, Cpfl, C2cl, Csn2, Cas Phi, and sequence variants thereof.

[0219] In those cases where the modification results in a knock-down of the BCL11A gene, expression of the BCL11A protein is reduced by at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, or at least about 90% in comparison to cells that have not been modified. In other cases, wherein the modification results in a knock-out of the BCL11A gene, the target nucleic acid of the cells of the population is modified such that expression of the BCL11A protein cannot be detected. Expression of a BCL11A protein can be measured by flow cytometry, ELISA, cell-based assays, Western blot, qRT-PCR, or other methods know in the art, or as described in the Examples.

[0220] In some embodiments, the disclosure provides methods of modifying a BCL11A target nucleic acid in a population of cells in vivo in a subject. In some embodiments, the modifying of the target nucleic acid sequence is carried out ex vivo in a eukaryotic cell, wherein the eukaryotic cell is selected from the group consisting of a hematopoietic stem cell (HSC), a hematopoietic progenitor cell (HPC), a CD34+ cell, a mesenchymal stem cell (MSC), induced pluripotent stem cell (iPSC), a common myeloid progenitor cell, a proerythroblast cell, and an erythroblast cell. In the foregoing embodiment, a population of the modified cells can be utilized in a method of treatment in a subject, wherein the modified cells are administered to the subject in need thereof, and wherein the subject is selected from the group consisting of mouse, rat, pig, non-human primate, and human. In some cases, the ex vivo cell is autologous and is isolated from the subject’s bone marrow or peripheral blood. In other cases, the ex vivo cell is allogeneic and is 15 05 25 isolated from a different subject’s bone marrow or peripheral blood. In the methods of treatment, the modified cell can be administered to the subject by a route of administration selected from intraparenchymal, intravenous, intra-arterial, intramuscular, subcuticular, intraarticular, intracardiac, intrapericardial, intravitreal, sub-capsular, or by subcutaneous injection and can be implanted into the subject by transplantation, local injection, systemic infusion, or combinations thereof. In the foregoing embodiment, the method results in the persistence of the modified cell or its progeny for at least about 1 month, at least about 2 months, at least about 3 months, at least about 4 months, at least about 6 months, at least about 7 months, at least about 8 months, at least about 9 months, at least about 10 months, at least about 11 months, at least about 12 months, at least about 18 months, at least about 2 years, at least about 3 years, at least about 4 years, or at least about 5 years.

[0221] In some embodiments of the methods of modifying a target nucleic acid sequence, modifying the BCL11A gene comprises binding of the CasX:gRNA complex to the target nucleic acid sequence and is introduced into the cells as an RNP. In some embodiments, the CasX is a catalytically inactive CasX (dCasX) protein that retains the ability to bind to the gRNA and the target nucleic acid sequence. For example, the target nucleic acid sequence comprises a BCL11A sequence comprising a sequence complementary to the GATA1 binding motif sequence, and binding of the dCasX:gRNA complex to the target sequence interferes with or represses transcription of the BCL11A allele. In some embodiments, the dCasX comprises a mutation at residues D672, E769, and / or D935 corresponding to the CasX protein of SEQ ID NO: 1 or D659, E756 and / or D922 corresponding to the CasX protein of SEQ ID NO: 2. In some embodiments of the foregoing, the mutation in the CasX variant protein is a substitution of alanine or glycine for the residue and can be utilized for any of the variants described herein.

[0222] Introducing recombinant expression vectors comprising the components or the nucleic acids encoding the components of the system embodiments into a target cell can be carried out in vivo, in vitro or ex vivo. In some embodiments of the method, vectors may be provided directly to a target host cell. Methods of introducing a nucleic acid (e.g., a nucleic acid comprising a donor polynucleotide sequence, one or more nucleic acids (DNA or RNA) encoding a CasX protein and / or gRNA, or a vector comprising same) into a cell are known in the art, and any convenient method can be used to introduce a nucleic acid (e.g., an expression construct) into a cell. Suitable methods include e.g., viral infection, transfection, lipofection, electroporation, calcium phosphate precipitation, polyethyleneimine (PEI)-mediated transfection, DEAE-dextran 15 05 25 mediated transfection, liposome-mediated transfection, particle gun technology, nucleofection, electroporation, direct addition by cell penetrating CasX proteins that are fused to or recruit donor DNA, cell squeezing, calcium phosphate precipitation, direct microinjection, nanoparticle-mediated nucleic acid delivery, and the like. Nucleic acids may be introduced into the cells using well-developed commercially-available transfection techniques such as use of TransMessenger® reagents from Qiagen, Stemfect™ RNA Transfection Kit from Stemgent, and TransIT®-mRNA Transfection Kit from Mirus Bio LLC, Lonza nucleofection, Maxagen electroporation and the like. Introducing recombinant expression vectors comprising sequences encoding the CasX:gRNA systems (and, optionally, the donor sequences) of the disclosure into cells under in vitro conditions can occur in any suitable culture media and under any suitable culture conditions that promote the survival of the cells. For example, cells may be contacted with vectors comprising the subject nucleic acids (e g., recombinant expression vectors having the donor template sequence and nucleic acid encoding the CasX and gRNA) such that the vectors are taken up by the cells. Vectors used for providing the nucleic acids encoding gRNAs and / or CasX proteins to a target host cell can include suitable promoters for driving the expression, that is, transcriptional activation of the nucleic acid of interest. In some cases, the encoding nucleic acid of interest will be operably linked to a promoter. This may include ubiquitously acting promoters, for example, the CMV-beta-actin promoter, or inducible promoters, such as promoters that are active in particular cell populations or that respond to the presence of drugs such as tetracycline or kanamycin. By transcriptional activation, it is intended that transcription will be increased above basal levels in the target host cell comprising the vector by at least about 10-fold, by at least about 100-fold, more usually by at least about 1000-fold. In addition, vectors used for providing a nucleic acid encoding a gRNA and / or a CasX protein to a cell may include nucleic acid sequences that encode for selectable markers in the target cells, so as to identify cells that have taken up the CasX protein and / or the gRNA.

[0223] For viral vector delivery, cells can be contacted with viral particles comprising the subject viral expression vectors and the nucleic acid encoding the CasX and gRNA and, optionally, the donor template. In some embodiments, the vector is an Adeno-Associated Viral (AAV) vector, wherein the AAV is selected from AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, AAV 44.9, AAV-Rh74, or AAVRhlO. In other cases, the AAV is selected from AAV1, AAV2, AAV5, AAV6, AAV7, AAV8, and AAV9, which are efficient for muscle transduction (Gruntman AM, et al. Gene transfer in 15 05 25 skeletal and cardiac muscle using recombinant adeno-associated virus. Curr Protoc Microbiol. 14(14D):3 (2013). Embodiments of AAV vectors are described more fully, below. In other embodiments, the vector is a lentiviral vector. Retroviruses, for example, lentiviruses, may be suitable for use in methods of the present disclosure. Commonly used retroviral vectors are ’’defective", e.g., are unable to produce viral proteins required for productive infection. Rather, replication of the vector requires growth in a packaging cell line. To generate viral particles comprising nucleic acids of interest, the retroviral nucleic acids comprising the nucleic acid are packaged into viral capsids by a packaging cell line. Different packaging cell lines provide a different envelope protein (ecotropic, amphotropic or xenotropic) to be incorporated into the capsid, and this envelope protein determines the specificity or tropism of the viral particle for the cells (ecotropic for murine and rat; amphotropic for most mammalian cell types including human, dog and mouse; and xenotropic for most mammalian cell types except murine cells). The appropriate packaging cell line may be used to ensure that the cells are targeted by the packaged viral particles. Methods of introducing subject vector expression vectors into packaging cell lines, and of collecting the viral particles that are generated by the packaging lines, are well known in the art, including U.S. Pat. No. 5,173,414; Tratschin et al., Mol. Cell. Biol. 5:3251-3260 (1985); Tratschin, et al., Mol. Cell. Biol. 4:2072-2081 (1984); Hermonat &Muzyczka, PNAS 81:6466-6470 (1984); and Samulski et al., J. Virol. 63:03822-3828 (1989). Nucleic acids can also be introduced by direct micro-injection (e.g., injection of RNA).

[0224] In other embodiments of the methods of modifying a BCL11A gene, the method utilizes CasX delivery particles (XDP) for the targeted delivery of RNPs to the cells of the subject. XDP are particles that closely resemble viruses, but do not contain viral genetic material and are therefore non-infectious. In some embodiments, the XDP comprise a CasX and gRNA complexed as an RNP and, optionally, a donor template comprising all or a portion of the BCL11A gene to either knock-down or knock-out the BCL11A gene or a portion of the gene by insertion via HDR or HITI mechanisms. Embodiments of XDPs are described more fully, below. VI. Polynucleotides and Vectors

[0225] In another aspect, the present disclosure relates to polynucleotides encoding the Class2, Type V nucleases and gRNA that have utility in the editing of the BCL11A gene. In some embodiments, the disclosure provides polynucleotides encoding the CasX proteins and the polynucleotides of the gRNAs of any of the CasX:gRNA system embodiments described herein. 15 05 25 In additional embodiments, the disclosure provides donor template polynucleotides encoding portions or all of a BCL11A gene. In some cases, the donor template comprises a mutation or a heterologous sequence for knocking down or knocking out the BCL11A gene upon its insertion in the target nucleic acid. In yet further embodiments, the disclosure provides vectors comprising polynucleotides encoding the CasX proteins and the CasX gRNAs described herein, as well as the donor templates of the embodiments.

[0226] In some embodiments, the disclosure provides a polynucleotide sequence encoding the CasX variants of any of the embodiments described herein, including the CasX protein variants of SEQ ID NOS: 59, 72-99, 101-148, and 26908-27154 as described in Table 4 or sequences having at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to a sequence of SEQ ID NOS: 59, 72-99, 101-148, and 26908-27154 of Table 4. In some embodiments, the disclosure provides a polynucleotide sequence encoding the CasX variants of any of the embodiments described herein, including the CasX protein variants of SEQ ID NOS: 36-99, 101-148, and 26908-27154 as described in Table 4 or sequences having at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to a sequence of SEQ ID NOS: 36-99, 101-148, and 26908-27154 of Table 4. In some embodiments, the disclosure provides an isolated polynucleotide sequence encoding a gRNA sequence of any of the embodiments described herein, including the sequences of SEQ ID NOS: 4-16, 2238-2285, 26794-26839 or 27219-27265 of Tables 2 and 3, together with the targeting sequences of SEQ ID NOS: 272-2100 or 2286-26789. In some embodiments, the disclosure provides an isolated polynucleotide sequence encoding a gRNA sequence of any of the embodiments described herein, including the sequences of SEQ ID NOS: 2101-2285, 26794-26839 and 27219-27265, together with the targeting sequences of SEQ ID NOS: 272-2100 or 2286-26789. In some embodiments, the disclosure provides an isolated polynucleotide sequence encoding a gRNA sequence of any of the embodiments described herein, including the sequences of SEQ ID NOS: 2281-2285, 26794-26839 and 27219-27265, together with the targeting sequences of SEQ ID NOS: 272-2100 or 2286-26789. In some embodiments, the sequences encoding the CasX protein are codon optimized for expression in a eukaryotic cell. 15 05 25

[0227] In some embodiments, the disclosure provides a polynucleotide encoding a gRNA scaffold sequence of SEQ ID NOS: 4-16, 2238-2285, 26794-26839 or 27219-27265, or as set forth in Table 2 or Table 3, or a sequence having at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% sequence identity thereto. In other embodiments, the disclosure provides a targeting sequence polynucleotide of Table 1, or a sequence having at least about 65%, at least about 75%, at least about 85%, or at least about 95% identity to a sequence of SEQ ID NOS: 272-2100 or 2286-26789. In some embodiments, the targeting sequence polynucleotide is, in turn, linked to the 3’ end of the gRNA scaffold sequence; either as a sgRNA or a dgRNA. In other embodiments, the disclosure provides gRNAs comprising targeting sequence polynucleotides having one or more single nucleotide polymorphisms (SNP) relative to a sequence of SEQ ID NOS: 272-2100 or 2286-26789.

[0228] In other embodiments, the disclosure provides an isolated polynucleotide sequence encoding a gRNA comprising a targeting sequence that is complementary to, and therefore is capable of hybridizing with, the BCL11A gene. In some embodiments, the polynucleotide sequence encodes a gRNA comprising a targeting sequence that hybridizes with a BCL11A exon. In other embodiments, the polynucleotide sequence encodes a gRNA comprising a targeting sequence that hybridizes with a BCL11A intron. In other embodiments, the polynucleotide sequence encodes a gRNA comprising a targeting sequence that hybridizes with a BCL11A intron-exon junction. In other embodiments, the polynucleotide sequence encodes a gRNA comprising a targeting sequence that hybridizes with an intergenic region of the BCL11A gene. In other embodiments, the polynucleotide sequence encodes a gRNA comprising a targeting sequence that hybridizes with a BCL11A regulatory element. In some cases, the BCL11A regulatory element is a BCL11A promoter or enhancer. In some cases, the BCL11A regulatory element is located 5’ of the BCL11A transcription start site, 3’ of the BCL11A transcription start, or in a BCL11A intron. In other embodiments, the polynucleotide sequence encodes a gRNA comprising a targeting sequence that hybridizes with a sequence located 5’ to the GATA1 binding motif sequence. In other embodiments, the polynucleotide sequence encodes a gRNA comprising a targeting sequence that hybridizes with a sequence overlapping the GATA1 binding motif sequence. In a particular embodiment of the foregoing, the polynucleotide sequence encodes a gRNA comprising a targeting sequence having SEQ ID NO: 15 05 25 22. In some cases, the BCL11A regulatory element is in an intron of the BCL11A gene. In other cases, the BCL11A regulatory element comprises the 5' UTR of the BCL11A gene. In still other cases, the BCL11A regulatory element comprises the 3' UTR of the BCL11A gene.

[0229] In other embodiments, the disclosure provides donor template nucleic acids, wherein the donor template comprises a nucleotide sequence having homology to a BCL11A target nucleic acid sequence. In some embodiments, the BCL11A donor template is intended for gene editing in conjunction with the CasX:gRNA system and comprises at least a portion of a BCL11A gene. In other embodiments, the BCL11A donor sequence comprises a sequence that encodes at least a portion of a BCL11A exon. In other embodiments, the BCL11A donor template has a sequence that encodes at least a portion of a BCL11A intron. In other embodiments, the BCL11A donor template has a sequence that encodes at least a portion of a BCL11A intron-exon junction. In other embodiments, the BCL11A donor template has a sequence that encodes at least a portion of an intergenic region of the BCL11A gene. In other embodiments, the BCL11A donor template has a sequence that encodes at least a portion of a BCL11A regulatory element. In some cases, the BCL11A donor template is a wild-type sequence that encodes at least a portion of SEQ ID NO: 100. In other cases, the BCL11A donor template sequence comprises one or more mutations relative to a wild-type BCL11A gene. In a particular embodiment, the donor template has a sequence that encodes a portion or all of the GATA1 binding motif sequence but with at least 1 to 5 mutations relative to the wild-type sequence. In the foregoing embodiments, the donor template is at least 10 nucleotides, at least 100 nucleotides, at least 200 nucleotides, at least 300 nucleotides, at least 400 nucleotides, at least 500 nucleotides, at least 600 nucleotides, at least 700 nucleotides, at least 800 nucleotides, at least 900 nucleotides, at least 1,000 nucleotides, at least 2,000 nucleotides, at least 3,000 nucleotides, at least 4,000 nucleotides, at least 5,000 nucleotides, at least 6,000 nucleotides, at least 7,000 nucleotides, at least 8,000 nucleotides, at least 9,000 nucleotides, at least 10,000 nucleotides, at least 12,000 nucleotides, or at least 15,000 nucleotides. In some embodiments, the donor template comprises at least about 10 to about 15,000 nucleotides. In some embodiments, the donor template is a single-stranded DNA template. In other embodiments, the donor template is a single stranded RNA template. In other embodiments, the donor template is a double-stranded DNA template. In some embodiments, the donor template can be provided as naked nucleic acid in the systems to edit the BCL11A gene and does not need to be incorporated 15 05 25 into a vector. In other embodiments, the donor template can be incorporated into a vector to facilitate its delivery to a cell; e.g., in a viral vector.

[0230] In other aspects, the disclosure relates to methods to produce polynucleotide sequences encoding the CasX variants, or the gRNA of any of the embodiments described herein, including homologous variants thereof, as well as methods to express the proteins expressed or RNA transcribed by the polynucleotide sequences. In general, the methods include producing a polynucleotide sequence coding for the CasX variants, or the gRNA of any of the embodiments described herein and incorporating the encoding gene into an expression vector appropriate for a host cell. Standard recombinant techniques in molecular biology can be used to make the polynucleotides and expression vectors of the present disclosure. For production of the encoded reference CasX, the CasX variants, or the gRNA of any of the embodiments described herein, the methods include transforming an appropriate host cell with an expression vector comprising the encoding polynucleotide, and culturing the host cell under conditions causing or permitting the resulting reference CasX, the CasX variants, or the gRNA of any of the embodiments described herein to be expressed or transcribed in the transformed host cell, thereby producing the CasX variants, or the gRNA, which are recovered by methods described herein or by standard purification methods known in the art or as described in the Examples.

[0231] In accordance with the disclosure, nucleic acid sequences that encode the CasX variants, or the gRNA of any of the embodiments described herein (or their complement) are used to generate recombinant DNA molecules that direct the expression in appropriate host cells. Several cloning strategies are suitable for performing the present disclosure, many of which are used to generate a construct that comprises a gene coding for a composition of the present disclosure, or its complement. In some embodiments, the cloning strategy is used to create a gene that encodes a construct that comprises nucleotides encoding the CasX variants, or the gRNA that is used to transform a host cell for expression of the composition.

[0232] In some approaches, a construct is first prepared containing the DNA sequence encoding a CasX variant or a gRNA. Exemplary methods for the preparation of such constructs are described in the Examples. The construct is then used to create an expression vector suitable for transforming a host cell, such as a prokaryotic or eukaryotic host cell for the expression and recovery of the protein construct, in the case of the CasX, or the gRNA. Where desired, the host cell is an A. coli. In other embodiments, the host cell is a eukaryotic cell. The eukaryotic host cell can be selected from Baby Hamster Kidney fibroblast (BHK) cells, human embryonic 15 05 25 kidney 293 (HEK293), human embryonic kidney 293T (HEK293T), NSO cells, SP2 / 0 cells, YO myeloma cells, P3X63 mouse myeloma cells, PER cells, PER.C6 cells, hybridoma cells, NIH3T3 cells, CV-1 (simian) in Origin with SV40 genetic material (COS), HeLa, Chinese hamster ovary (CHO), or yeast cells, or other eukaryotic cells known in the art suitable for the production of recombinant products. Exemplary methods for the creation of expression vectors, the transformation of host cells and the expression and recovery of the CasX variants or the gRNA are described in the Examples.

[0233] The gene encoding the CasX variant, or the gRNA construct can be made in one or more steps, either fully synthetically or by synthesis combined with enzymatic processes, such as restriction enzyme-mediated cloning, PCR and overlap extension, including methods more fully described in the Examples. The methods disclosed herein can be used, for example, to ligate sequences of polynucleotides encoding the various components (e.g., CasX and gRNA) genes of a desired sequence. Genes encoding polypeptide compositions are assembled from oligonucleotides using standard techniques of gene synthesis.

[0234] In some embodiments, the nucleotide sequence encoding a CasX protein is codon optimized for the intended host cell. This type of optimization can entail a mutation of an encoding nucleotide sequence to mimic the codon preferences of the intended host organism or cell while encoding the same CasX protein. Thus, the codons can be changed, but the encoded protein or gRNA remains unchanged. For example, if the intended target cell of the CasX protein was a human cell, a human codon-optimized CasX-encoding nucleotide sequence could be used. As another non-limiting example, if the intended host cell were a mouse cell, then a mouse codon-optimized CasX-encoding nucleotide sequence could be generated. The gene design can be performed using algorithms that optimize codon usage and amino acid composition appropriate for the host cell utilized in the production of the reference CasX or the CasX variants. In one method of the disclosure, a library of polynucleotides encoding the components of the constructs is created and then assembled, as described above. The resulting genes are then assembled and the resulting genes used to transform a host cell and produce and recover the CasX variants, or the gRNA compositions for evaluation of its properties, as described herein.

[0235] The disclosure provides for the use of plasmid expression vectors containing replication and control sequences that are compatible with and recognized by the host cell and are operably linked to the gene encoding the polypeptide for controlled expression of the 15 05 25 polypeptide or transcription of the RNA. Such vector sequences are well known for a variety of bacteria, yeast, and viruses. Useful expression vectors that can be used include, for example, segments of chromosomal, non-chromosomal and synthetic DNA sequences. "Expression vector" refers to a DNA construct containing a DNA sequence that is operably linked to a suitable control sequence capable of effecting the expression of the DNA encoding the polypeptide in a suitable host. The requirements are that the vectors are replicable and viable in the host cell of choice. Low- or high-copy number vectors may be used as desired. The control sequences of the vector include a promoter to effect transcription, an optional operator sequence to control such transcription, a sequence encoding suitable mRNA ribosome binding sites, and sequences that control termination of transcription and translation. In some embodiments, a nucleotide sequence encoding a gRNA is operably linked to a control element, e.g., a transcriptional control element, such as a promoter. In some embodiments, a nucleotide sequence encoding a CasX protein is operably linked to a control element, e.g., a transcriptional control element, such as a promoter. In other cases, the nucleotide encoding the CasX and gRNA are linked and are operably linked to a single control element. The promoter may be any DNA sequence, which shows transcriptional activity in the host cell of choice and may be derived from genes encoding proteins either homologous or heterologous to the host cell. Exemplary regulatory elements include a transcription promoter, a transcription enhancer element, a transcription termination signal, internal ribosome entry site (IRES) or P2A peptide to permit translation of multiple genes from a single transcript, polyadenylation sequences to promote downstream transcriptional termination, sequences for optimization of initiation of translation, and translation termination sequences. In some cases, the promoter is a constitutively active promoter. In some cases, the promoter is a regulatable promoter. In some cases, the promoter is an inducible promoter. In some cases, the promoter is a tissue-specific promoter. In some cases, the promoter is a cell type-specific promoter. In some cases, the transcriptional control element (e.g., the promoter) is functional in a targeted cell type or targeted cell population. For example, in some cases, the transcriptional control element can be functional in eukaryotic cells, e.g., packaging cells for viral or XDP vectors, hematopoietic stem cells (HSC), hematopoietic progenitor cells (HPC), CD34+ cells, mesenchymal stem cells (MSC), embryonic stem (ES) cells, induced pluripotent stem cells (iPSC), common myeloid progenitor cells, proerythroblast cells, and erythroblast cells. 15 05 25

[0236] Non-limiting examples of pol II promoters include, but are not limited to EF-1 alpha, EF-1 alpha core promoter, Jens Tomoe (JeT), promoters from cytomegalovirus (CMV), CMV immediate early (CMVIE), CMV enhancer, herpes simplex virus (HSV) thymidine kinase, early and late simian virus 40 (SV40), the SV40 enhancer, long terminal repeats (LTRs) from retrovirus, mouse metallothionein-I, adenovirus major late promoter (Ad MLP), CMV promoter full-length promoter, the minimal CMV promoter, the chicken (E<-actin promoter (CBA), CBA hybrid (CBh), chicken CE<-actin promoter with cytomegalovirus enhancer (CB7), chicken beta-Actin promoter and rabbit beta-Globin splice acceptor site fusion (CAG), the rous sarcoma virus (RSV) promoter, the HIV-Ltr promoter, the hPGK promoter, the HSV TK promoter, a 7SK promoter, the Mini-TK promoter, the human synapsin I (SYN) promoter which confers neuronspecific expression, beta-actin promoter, super core promoter 1 (SCP1), the Mecp2 promoter for selective expression in neurons, the minimal IL-2 promoter, the Rous sarcoma virus enhancer / promoter (single), the spleen focus-forming virus long terminal repeat (LTR) promoter, the TBG promoter, promoter from the human thyroxine-binding globulin gene (Liver specific),, the PGK promoter, the human ubiquitin C promoter (UBC), the UCOE promoter (Promoter of HNRPA2B1-CBX3), the synthetic CAG promoter, the Histone H2 promoter, the Histone H3 promoter, the UI al small nuclear RNA promoter (226 nt), the UI al small nuclear RNA promoter (226 nt), the Ulb2 small nuclear RNA promoter (246 nt) 26, the GUSB promoter, the CBh promoter, rhodopsin (Rho) promoter, silencing-prone spleen focus forming virus (SFFV) promoter, a human Hl promoter (Hl), a POLI promoter, the TTR minimal enhancer / promoter, the b-kinesin promoter, mouse mammary tumor virus long terminal repeat (LTR) promoter, the human eukaryotic initiation factor 4A (EIF4A1) promoter, the ROSA26 promoter, the glyceraldehyde 3-phosphate dehydrogenase (GAPDH) promoter, tRNA promoters, and truncated versions and sequence variants of the foregoing. In a particular embodiment, the pol II promoter is EF-lalpha, wherein the promoter enhances transfection efficiency, the transgene transcription or expression of the CRISPR nuclease, the proportion of expression-positive clones and the copy number of the episomal vector in long-term culture.

[0237] Non-limiting examples of pol III promoters include, but are not limited to U6, mini U6, U6 truncated promoters,7SK, and Hl variants, BiHl (Bidrectional Hl promoter), BiU6, Bi7SK, BiHl (Bidirectional U6, 7SK, and Hl promoters), gorilla U6, rhesus U6, human 7SK, human Hl promoters, and sequence variants thereof. In the foregoing embodiment, the pol III promoter enhances the transcription of the gRNA. 15 05 25

[0238] Selection of the appropriate vector and promoter is well within the level of ordinary skill in the art, as it related to controlling expression, e.g., for modifying a BCL11A gene. The expression vector may also contain a ribosome binding site for translation initiation and a transcription terminator. The expression vector may also include appropriate sequences for amplifying expression. The expression vector may also include nucleotide sequences encoding protein tags (e.g., 6xHis tag, hemagglutinin tag, fluorescent protein, etc.) that can be fused to the CasX protein, thus resulting in a chimeric CasX protein that are used for purification or detection.

[0239] Recombinant expression vectors of the disclosure can also comprise elements that facilitate robust expression of CasX proteins and the gRNAs of the disclosure. For example, recombinant expression vectors can include one or more of a polyadenylation signal (poly(A)), an intronic sequence or a post-transcriptional regulatory element such as a woodchuck hepatitis post-transcriptional regulatory element (WPRE). Exemplary poly(A) sequences include hGH poly(A) signal (short), HSV TK poly(A) signal, synthetic polyadenylation signals, SV40 poly(A) signal, P-globin poly(A) signal and the like. A person of ordinary skill in the art will be able to select suitable elements to include in the recombinant expression vectors described herein.

[0240] In some embodiments, provided herein are one or more recombinant expression vectors comprising one or more of: (i) a nucleotide sequence of a donor template nucleic acid where the donor template comprises a nucleotide sequence having homology to a sequence of the target BCL11A locus of the target nucleic acid (e.g., a target genome); (ii) a nucleotide sequence that encodes a gRNA that hybridizes to a target sequence of the BCL11A locus of the targeted genome (e.g., configured as a single or dual guide RNA) operably linked to a promoter that is operable in a target cell such as a eukaryotic cell; and (iii) a nucleotide sequence encoding a CasX protein operably linked to a promoter that is operable in a target cell such as a eukaryotic cell. In some embodiments, the sequences encoding the donor template, the gRNA and the CasX protein are in different recombinant expression vectors, and in other embodiments one or more polynucleotide sequences (for the donor template, CasX, and the gRNA) are in the same recombinant expression vector. In other cases, the CasX and gRNA are delivered to the target cell as an RNP (e.g., by electroporation or chemical means) and the donor template is delivered by a vector. 15 05 25

[0241] The polynucleotide sequence(s) are inserted into the vector by a variety of procedures. In general, DNA is inserted into an appropriate restriction endonuclease site(s) using techniques known in the art. Vector components generally include, but are not limited to, one or more of a signal sequence, an origin of replication, one or more marker genes, an enhancer element, a promoter, and a transcription termination sequence. Construction of suitable vectors containing one or more of these components employs standard ligation techniques which are known to the skilled artisan. Such techniques are well known in the art and well described in the scientific and patent literature. Various vectors are publicly available. The vector may, for example, be in the form of a plasmid, cosmid, viral particle, or phage that may conveniently be subjected to recombinant DNA procedures, and the choice of vector will often depend on the host cell into which it is to be introduced. Thus, the vector may be an autonomously replicating vector, i.e., a vector, which exists as an extrachromosomal entity, the replication of which is independent of chromosomal replication, e.g., a plasmid. Alternatively, the vector may be one which, when introduced into a host cell, is integrated into the host cell genome and replicated together with the chromosome(s) into which it has been integrated. Once introduced into a suitable host cell, expression of the protein involved in antigen processing, antigen presentation, antigen recognition, and / or antigen response can be determined using any nucleic acid or protein assay known in the art. For example, the presence of transcribed mRNA of reference CasX or the CasX variants can be detected and / or quantified by conventional hybridization assays (e.g., Northern blot analysis), amplification procedures (e.g. RT-PCR), SAGE (U.S. Pat. No. 5,695,937), and array-based technologies (see e.g., U.S. Pat. Nos. 5,405,783, 5,412,087 and 5,445,934), using probes complementary to any region of the polynucleotide.

[0242] The polynucleotides and recombinant expression vectors can be delivered to the target host cells by a variety of methods. Such methods include, but are not limited to, viral infection, transfection, lipofection, electroporation, calcium phosphate precipitation, polyethyleneimine (PEI)-mediated transfection, DEAE-dextran mediated transfection, microinjection, liposome-mediated transfection, particle gun technology, nucleofection, direct addition by cell penetrating CasX proteins that are fused to or recruit donor DNA, cell squeezing, calcium phosphate precipitation, direct microinjection, nanoparticle-mediated nucleic acid delivery, and using the commercially available TransMessenger® reagents from Qiagen, StemfectTM RNA Transfection Kit from Stemgent, and TransIT®-mRNA Transfection Kit from Mirus Bio LLC, Lonza nucleofection, Maxagen electroporation and the like. 15 05 25

[0243] A recombinant expression vector sequence can be packaged into a virus or virus-like particle (also referred to herein as a “particle” or “virion”) for subsequent infection and transformation of a cell, ex vivo, in vitro or in vivo. Such particles or virions will typically include proteins that encapsidate or package the vector genome. Suitable expression vectors may include viral expression vectors based on vaccinia virus; poliovirus; adenovirus; a retroviral vector (e.g., Murine Leukemia Virus), spleen necrosis virus, and vectors derived from retroviruses such as Rous Sarcoma Virus, Harvey Sarcoma Virus, avian leukosis virus, a lentivirus, human immunodeficiency virus, myeloproliferative sarcoma virus, and mammary tumor virus; and the like. In some embodiments, a recombinant expression vector of the present disclosure is a recombinant adeno-associated virus (AAV) vector. In some embodiments, a recombinant expression vector of the present disclosure is a recombinant lentivirus vector. In some embodiments, a recombinant expression vector of the present disclosure is a recombinant retroviral vector.

[0244] In some embodiments, a recombinant expression vector of the present disclosure is a recombinant adeno-associated virus (AAV) vector. In some embodiments, a recombinant expression vector of the present disclosure is a recombinant lentivirus vector. In some embodiments, a recombinant expression vector of the present disclosure is a recombinant retroviral vector.

[0245] AAV is a small (20 nm), nonpathogenic virus that is useful in treating human diseases in situations that employ a viral vector for delivery to a cell such as a eukaryotic cell, either in vivo or ex vivo for cells to be prepared for administering to a subject. A construct is generated, for example a construct encoding any of the CasX proteins and / or CasX gRNA embodiments as described herein, and is flanked with AAV inverted terminal repeat (ITR) sequences, thereby enabling packaging of the AAV vector into an AAV viral particle.

[0246] An “AAV” vector may refer to the naturally occurring wild-type virus itself or derivatives thereof. The term covers all subtypes, serotypes and pseudotypes, and both naturally occurring and recombinant forms, except where required otherwise. As used herein, the term “serotype” refers to an AAV which is identified by and distinguished from other AAVs based on capsid protein reactivity with defined antisera, e.g., there are many known serotypes of primate AAVs. In some embodiments, the AAV vector is selected from AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, AAV 44.9, AAV-Rh74 (Rhesus macaque-derived AAV), and AAVRh 10, and modified capsids of these serotypes. For 15 05 25 example, serotype AAV-2 is used to refer to an AAV which contains capsid proteins encoded from the cap gene of AAV-2 and a genome containing 5' and 3' ITR sequences from the same AAV-2 serotype. Pseudotyped AAV refers to an AAV that contains capsid proteins from one serotype and a viral genome including 5'-3' ITRs of a second serotype. Pseudotyped rAAV would be expected to have cell surface binding properties of the capsid serotype and genetic properties consistent with the ITR serotype. Pseudotyped recombinant AAV (rAAV) are produced using standard techniques described in the art. As used herein, for example, rAAVl may be used to refer an AAV having both capsid proteins and 5'-3' ITRs from the same serotype or it may refer to an AAV having capsid proteins from serotype 1 and 5'-3' ITRs from a different AAV serotype, e.g., AAV serotype 2. For each example illustrated herein the description of the vector design and production describes the serotype of the capsid and 5'-3' ITR sequences.

[0247] An “AAV virus” or “AAV viral particle” refers to a viral particle composed of at least one AAV capsid protein (preferably by all of the capsid proteins of a wild-type AAV) and an encapsidated polynucleotide. If the particle additionally comprises a heterologous polynucleotide (i.e., a polynucleotide other than a wild-type AAV genome to be delivered to a mammalian cell), it is typically referred to as “rAAV”. An exemplary heterologous polynucleotide is a polynucleotide comprising a CasX protein and / or sgRNA and, optionally, a donor template of any of the embodiments described herein.

[0248] By “adeno-associated virus inverted terminal repeats” or “AAV ITRs” is meant the art recognized regions found at each end of the AAV genome which function together in cis as origins of DNA replication and as packaging signals for the virus. AAV ITRs, together with the AAV rep coding region, provide for the efficient excision and rescue from, and integration of a nucleotide sequence interposed between two flanking ITRs into a mammalian cell genome. The nucleotide sequences of AAV ITR regions are known. See, for example Kotin, R.M. (1994) Human Gene Therapy 5:793-801; Berns, K. I. “Parvoviridae and their Replication” in Fundamental Virology, 2nd Edition, (B N. Fields and D. M. Knipe, eds.). As used herein, an AAV ITR need not have the wild-type nucleotide sequence depicted, but may be altered, e.g., by the insertion, deletion or substitution of nucleotides. Additionally, the AAV ITR may be derived from any of several AAV serotypes, including without limitation, AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV-Rh74, and AAVRhlO, and modified capsids of these serotypes. Furthermore, 5' and 3' ITRs which flank a selected nucleotide sequence in an AAV vector need not necessarily be identical or derived from the 15 05 25 same AAV serotype or isolate, so long as they function as intended, i.e., to allow for excision and rescue of the sequence of interest from a host cell genome or vector, and to allow integration of the heterologous sequence into the recipient cell genome when AAV Rep gene products are present in the cell. Use of AAV serotypes for integration of heterologous sequences into a host cell is known in the art (see, e.g., WO2018195555A1 and US20180258424A1, incorporated by reference herein).

[0249] By “AAV rep coding region” is meant the region of the AAV genome which encodes the replication proteins Rep 78, Rep 68, Rep 52 and Rep 40. These Rep expression products have been shown to possess many functions, including recognition, binding and nicking of the AAV origin of DNA replication, DNA helicase activity and modulation of transcription from AAV (or other heterologous) promoters. The Rep expression products are collectively required for replicating the AAV genome. By “AAV cap coding region” is meant the region of the AAV genome which encodes the capsid proteins VP1, VP2, and VP3, or functional homologues thereof. These Cap expression products supply the packaging functions which are collectively required for packaging the viral genome.

[0250] In some embodiments, AAV capsids utilized for delivery of the encoding sequences for the CasX and gRNA, and, optionally, the DMPK donor template nucleotides to a host cell can be derived from any of several AAV serotypes, including without limitation, AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, AAV 44.9, AAV-Rh74 (Rhesus macaque-derived AAV), and AAVRh 10, and the AAV ITRs are derived from AAV serotype 2. In a particular embodiment, AAV1, AAV7, AAV6, AAV8, or AAV9 are utilized for delivery of the CasX, gRNA, and, optionally, donor template nucleotides, to a host muscle cell.

[0251] In order to produce rAAV viral particles, an AAV expression vector is introduced into a suitable host cell using known techniques, such as by transfection. Packaging cells are typically used to form virus particles; such cells include HEK293 cells (and other cells known in the art), which package adenovirus. A number of transfection techniques are generally known in the art; see, e.g., Sambrook et al. (1989) Molecular Cloning, a laboratory manual, Cold Spring Harbor Laboratories, New York. Particularly suitable transfection methods include calcium phosphate co-precipitation, direct microinjection into cultured cells, electroporation, liposome mediated gene transfer, lipid-mediated transduction, and nucleic acid delivery using high-velocity microprojectiles. 15 05 25

[0252] In an advantage of rAAV constructs of the present disclosure, the smaller size of the CRISPR Type V nucleases; e.g., the CasX of the embodiments, permits the inclusion of all the necessary editing and ancillary expression components into the transgene such that a single rAAV particle can deliver and transduce these components into a target cell in a form that results in the expression of the CRISPR nuclease and gRNA that are capable of effectively modifying the target nucleic acid of the target cell. A representative schematic of such a construct is presented in FIG. 13. This stands in marked contrast to other CRISPR systems, such as Cas9, where typically a two-particle system is employed to deliver the necessary editing components to a target cell. Thus, in some embodiments of the rAAV systems, the disclosure provides; i) a first plasmid comprising the ITRs, sequences encoding the CasX variant, sequences encoding one or more gRNA, a first promoter operably linked to the CasX and a second promoter operably linked to the gRNA, and, optionally, one or more enhancer elements; ii) a second plasmid comprising the rep and cap genes; and iii) a third plasmid comprising helper genes, wherein upon transfection of an appropriate packaging cell, the cell is capable of producing an rAAV having the ability to deliver to a target cell, in a single particle, sequences capable of expressing the CasX nuclease and gRNA having the ability to edit the target nucleic acid of the target cell. In some embodiments of the rAAV systems, the sequence encoding the CRISPR protein and the sequence encoding the at least first gRNA are less than about 3100, less than about 3090, less than about 3080, less than about 3070, less than about 3060, less than about 3050, or less than about 3040 nucleotides in length, such that the sequences encoding the first and second promoter and, optionally, one or more enhance elements can have at least about 1300, at least about 1350, at least about 1360, at least about 1370, at least about 1380, at least about 1390, at least about 1400, at least about 1500, at least about 1600 nucleotides, at least 1650, at least about 1700, at least about 1750, at least about 1800, at least about 1850, or at least about 1900 nucleotides in combined length. In some embodiments of the rAAV systems, the sequence encoding the first promoter and the at least one accessory element have greater than at least about 1300, at least about 1350, at least about 1360, at least about 1370, at least about 1380, at least about 1390, at least about 1400, at least about 1500, at least about 1600 nucleotides, at least 1650, at least about 1700, at least about 1750, at least about 1800, at least about 1850, or at least about 1900 nucleotides in combined length. In some embodiments of the rAAV systems, the sequence encoding the first and second promoters and the at least one accessory element have greater than at least about 1300, at least about 1350, at least about 1360, 15 05 25 at least about 1370, at least about 1380, at least about 1390, at least about 1400, at least about 1500, at least about 1600 nucleotides, at least 1650, at least about 1700, at least about 1750, at least about 1800, at least about 1850, or at least about 1900 nucleotides in combined length.

[0253] In some embodiments, host cells transfected with the above-described AAV expression vectors are rendered capable of providing AAV helper functions in order to replicate and encapsidate the nucleotide sequences flanked by the AAV ITRs to produce rAAV viral particles. AAV helper functions are generally AAV-derived coding sequences which can be expressed to provide AAV gene products that, in turn, function in trans for productive AAV replication. AAV helper functions are used herein to complement necessary AAV functions that are missing from the AAV expression vectors. Thus, AAV helper functions include one, or both of the major AAV ORFs (open reading frames), encoding the rep and cap coding regions, or functional homologues thereof. Accessory functions can be introduced into and then expressed in host cells using methods known to those of skill in the art. Commonly, accessory functions are provided by infection of the host cells with an unrelated helper virus. In some embodiments, accessory functions are provided using an accessory function vector. Depending on the host / vector system utilized, any of a number of suitable transcription and translation control elements, including constitutive and inducible promoters, transcription enhancer elements, transcription terminators, etc., may be used in the expression vector. In some embodiments, the disclosure provides host cells comprising the AAV vectors of the embodiments disclosed herein.

[0254] In other embodiments, suitable vectors may include virus-like particles (VLP). Viruslike particles (VLPs) are particles that closely resemble viruses, but do not contain viral genetic material and are therefore non-infectious. In some embodiments, VLPs comprise a polynucleotide encoding a transgene of interest, for example any of the CasX protein and / or a gRNA embodiments, and, optionally, donor template polynucleotides described herein, packaged with one or more viral structural proteins. In other embodiments, the disclosure provides XDPs produced in vitro that comprise a CasX:gRNA RNP complex and, optionally, a donor template. Combinations of structural proteins from different viruses can be used to create XDPs, including components from virus families including Parvoviridae (e.g., adeno-associated virus), Retroviridae (e.g., alpharetrovirus, a betaretrovirus, a gammaretrovirus, a deltaretrovirus, a epsilonretrovirus, or a lentivirus), Flaviviridae (e.g., Hepatitis C virus), Paramyxoviridae (e.g., Nipah) and bacteriophages (e.g., QP, AP205). In some embodiments, the disclosure provides XDP systems designed using components of retrovirus, including lenti viruses (such as HIV) and 15 05 25 alpharetrovirus, betaretrovirus, gammaretrovirus, deltaretrovirus, epsilonretrovirus, in which individual plasmids comprising polynucleotides encoding the various components are introduced into a packaging cell that, in turn, produce the XDP. In some embodiments, the disclosure provides XDP comprising one or more components of i) protease, ii) a protease cleavage site, iii) one or more components of a gag polyprotein selected from a matrix protein (MA), a nucleocapsid protein (NC), a capsid protein (CA), a pl peptide, a p6 peptide, a P2A peptide, a P2B peptide, a PIO peptide, a pl2 peptide, a PP21 / 24 peptide, a P12 / P3 / P8 peptide, and a P20 peptide; v) CasX; vi) gRNA, and vi) targeting glycoproteins or antibody fragments wherein the resulting XDP particle encapsidates a CasX:gRNA RNP. The polynucleotides encoding the Gag, CasX and gRNA can further comprise paired components designed to assist the trafficking of the components out of the nucleus of the host cell and into the budding XDP. Non-limiting examples of such trafficking components include hairpin RNA such as MS2 hairpin, PP7 hairpin, QP hairpin, and UI hairpin II that have binding affinity for MS2 coat protein, PP7 coat protein, QP coat protein, and UI A signal recognition particle, respectively. In other embodiments, the gRNA can comprise Rev response element (RRE) or portions thereof that have binding affinity to Rev, which can be linked to the Gag polyprotein. In other embodiments, the gRNA can comprise one or more RRE and one or more MS2 hairpin sequences. In other embodiments, the gRNA can comprise Rev response element (RRE) or portions thereof that have binding affinity to Rev, which can be linked to the Gag polyprotein. The RRE can be selected from the group consisting of Stem IIB of Rev response element (RRE), Stem II-V of RRE, Stem II of RRE, Rev-binding element (RBE) of Stem IIB, and full-length RRE. In the foregoing embodiment, the components include sequences of UGGGCGCAGCGUCAAUGACGCUGACGGUACA (Stem IIB; SEQ ID NO: 27266), GCACUAUGGGCGCAGCGUCAAUGACGCUGACGGUACAGGCCAGACAAUUAUUGU CUGGUAUAGUGC (Stem II; SEQ ID NO: 27267), GCUGACGGUACAGGC (RBE, SEQ ID NO: 27268), CAGGAAGCACUAUGGGCGCAGCGUCAAUGACGCUGACGGUACAGGCCAGACAAU UAUUGUCUGGUAUAGUGCAGCAGCAGAACAAUUUGCUGAGGGCUAUUGAGGCGC AACAGCAUCUGUUGCAACUCACAGUCUGGGGCAUCAAGCAGCUCCAGGCAAGAA UCCUG (Stem II-V; SEQ ID NO: 27269), and AUrUrAUrC UUUvjUUCCUU CjCjCjU U C U U vAjCtACtC ACrCAUtjAACrC AC U AU GCjCjCGC AUrC GTTCA ATTGACGCITGACGGTTAC AGGCC AGACA AI H J AUUGUCI1GGI IAU AGI TGCAGCA A_x A— / xA xA A_x A_J xA A—A_J A—z A—* A-J A A_J A—* x A A—x A A__J A__J A— / x A jL A x A xA A-9 A- / A A-Z A-Z A—A— / A-9 A—J A—* x A A_x A A-J A-Z A-J / x A A— / xA 15 05 25 GCAGAACAAUUUGCUGAGGGCUAUUGAGGCGCAACAGCAUCUGUUGCAACUCAC AGUCUGGGGCAUCAAGCAGCUCCAGGCAAGAAUCCUGGCUGUGGAAAGAUACCU AAAGGAUCAACAGCUCCU (full-length RRE; SEQ ID NO: 27270). In other embodiments, the gRNA can comprise one or more RRE and one or more MS2 hairpin sequences. In a particular embodiment, the gRNA comprises an MS2 hairpin variant that is optimized to increase the binding affinity to the MS2 coat protein, thereby enhancing the incorporation of the gRNA and associated CasX into the budding XDP. gRNA variants comprising MS2 hairpin variants include gRNA variants 275-315 and 317-320 (SEQ ID NOS: 2722-27264).

[0255] The targeting glycoproteins or antibody fragments on the surface that provides tropism of the XDP to the target cell, wherein upon administration and entry into the target cell, the RNP molecule is free to be transported into the nucleus of the cell. The envelope glycoprotein can be derived from any enveloped viruses known in the art to confer tropism to XDP, including but not limited to the group consisting of Argentine hemorrhagic fever virus, Australian bat virus, Autographa califomica multiple nucleopolyhedrovirus, Avian leukosis virus, baboon endogenous virus, Bolivian hemorrhagic fever virus, Borna disease virus, Breda virus, Bunyamwera virus, Chandipura virus, Chikungunya virus, Crimean-Congo hemorrhagic fever virus, Dengue fever virus, Duvenhage virus, Eastern equine encephalitis virus, Ebola hemorrhagic fever virus, Ebola Zaire virus, enteric adenovirus, Ephemerovirus, Epstein-Bar virus (EBV), European bat virus 1, European bat virus 2, Fug Synthetic gP Fusion, Gibbon ape leukemia virus, Hantavirus, Hendra virus, hepatitis A virus, hepatitis B virus, hepatitis C virus, hepatitis D virus, hepatitis E virus, hepatitis G Virus (GB virus C), herpes simplex virus type 1, herpes simplex virus type 2, human cytomegalovirus (HHV5), human foamy virus, human herpesvirus (HHV), human Herpesvirus 7, human herpesvirus type 6, human herpesvirus type 8, human immunodeficiency virus 1 (HIV-1), human metapneumovirus, human T-lymphotropic virus 1, influenza A, influenza B, influenza C virus, Japanese encephalitis virus, Kaposi’s sarcoma-associated herpesvirus (HHV8), Kaysanur Forest disease virus, La Crosse virus, Lagos bat virus, Lassa fever virus, lymphocytic choriomeningitis virus (LCMV), Machupo virus, Marburg hemorrhagic fever virus, measles virus, Middle eastern respiratory syndrome-related coronavirus, Mokola virus, Moloney murine leukemia virus, monkey pox, mouse mammary tumor virus, mumps virus, murine gammaherpesvirus, Newcastle disease virus, Nipah virus, Nipah virus, Norwalk virus, Omsk hemorrhagic fever virus, papilloma virus, parvovirus, pseudorabies virus, Quaranfil virus, rabies virus, RD114 Endogenous Feline Retrovirus, 15 05 25 respiratory syncytial virus (RSV), Rift Valley fever virus, Ross River virus, rRotavirus, Rous sarcoma virus, rubella virus, Sabia-associated hemorrhagic fever virus, SARS-associated coronavirus (SARS-CoV), Sendai virus, Tacaribe virus, Thogotovirus, tick-borne encephalitis causing virus, varicella zoster virus (HHV3), varicella zoster virus (HHV3), variola major virus, variola minor virus, Venezuelan equine encephalitis virus,...

Claims

1. A system for modifying a polypyrimidine tract-binding protein 1 (BCL11 A) gene targetnucleic acid sequence, the system comprising a CasX variant protein and a guide ribonucleic acid (gRNA) variant, whereina. the gRNA variant comprises:(i) a scaffold stem loop sequence of SEQ ID NO: 25; and(ii) a targeting sequence complementary to a target nucleic acid sequence comprising a region within a poly pyrimidine tract-binding protein 1 (BCL11 A) gene;b. the Cas X variant protein is a chimeric CasX variant protein comprising:(i) the NTSB domain of SEQ ID NO: 1, or a sequence with at least 90% sequence identity thereto;(ii) the helical lb domain of SEQ ID NO: 1, or a sequence with at least 90% sequence identity thereto; and(iii) the RuvC a and RuvC b domains of SEQ ID NO: 2, or sequences with at least 90% sequence identity thereto.wherein the system is capable of modifying the BCL11A gene, wherein the modifying comprises introducing an insertion, deletion, substitution, duplication, or inversion of one or more nucleotides in the BCL11A gene.

2. The system of claim 1, wherein the region within the BCL11A gene is selected from the group consisting of:a. a BCL11A intron;b. a BCL11A exon;c. a BCL11A intron-exon j unction;d. a BCL11A regulatory element; ande. an mtergenic region.

3. The system of claim 1 or claim 2, wherein the BCL11A gene comprises a wild-type sequence.

4. The system of any one of claims 1-3, wherein the gRNA variant is a single-molecule gRNA (sgRNA).

5. The system of any one of claims 1-4, wherein the targeting sequence of the gRNA variant comprises a sequence selected from the group consisting of SEQ ID NOS: 272-2100 and 2286-26789, or a sequence having at least about 90% sequence identity thereto.

6. The system of any one of claims 1-5, wherein the targeting sequence is linked to the 3’ end of the gRNA variant.

7. The system of any one of claims 1-6, wherein the targeting sequence has a singlenucleotide removed from the 3’ end of the sequence, or has two, three, four, or five nucleotides removed from the 3’ end of the sequence.

8. The system of any one of claims 1-7, wherein the targeting sequence of the gRNA is complementary to a sequence of a BCL11A exon.

9. The system of claim 8, wherein the targeting sequence of the gRNA is complementary to a sequence selected from the group consisting of a BCL11A exon 1 sequence, BCL11A exon 2 sequence, BCL11A exon 3 sequence, BCL11A exon 4 sequence, BCL11A exon 5 sequence, BCL11A exon 6 sequence, BCL11A exon 7 sequence, BCL11A exon 8 sequence, and a BCL11A exon 9 sequence.

10. The system of claim 9, wherein the targeting sequence of the gRNA is complementary to a sequence selected from the group consisting of a BCL11A exon 1 sequence, BCL11A exon 2 sequence, and a BCL11A exon 3 sequence.

11. The system of any one of claims 1-7, wherein the targeting sequence of the gRNA is complementary to a sequence of a BCL11A regulatory element.

12. The system of claim 11, wherein the targeting sequence of the gRNA is complementary to a sequence of a promoter of the BCL11A gene.

13. The system of claim 11, wherein the targeting sequence of the gRNA is complementary to a sequence of an enhancer regulatory element.

14. The system of claim 13, wherein the targeting sequence of the gRNA is complementary to a sequence that comprises a GATA1 erythroid-specific enhancer binding site (GATA1) of the BCL11A gene or to a sequence that is 5’ or 3' to the GATA1 binding site of the BCL11A gene.

15. The system of claim 14, wherein the targeting sequence of the gRNA variant comprises a sequence of UGGAGCCUGUGAUAAAAGCA (SEQ ID NO: 22), or a sequence having at least 90% sequence identity thereto.

16. The system of claim 14, wherein the targeting sequence of the gRNA comprises a sequence of UGCUUUUAUCACAGGCUCCA (SEQ ID NO: 23), or a sequence having at least 90% sequence identity thereto.

17. The system of claim 14, wherein the targeting sequence of the gRNA comprises a sequence of CAGGCUCCAGGAAGGGUUUG (SEQ ID NO: 2949), or a sequence having at least 90% sequence identity thereto.

18. The system of claim 14, wherein the targeting sequence of the gRNA comprises a sequence of GAGGCCAAACCCUUCCUGGA (SEQ ID NO: 2948), or a sequence having at least 90% sequence identity thereto.

19. The system of claim 14, wherein the targeting sequence of the gRNA comprises a sequence of AGUGCAAGCUAACAGUUGCU (SEQ ID NO: 15747), or a sequence having at least 90% sequence identity thereto.

20. The system of claim 14, wherein the targeting sequence of the gRNA comprises a sequence of AUACAACUUUGAAGCUAGUC (SEQ ID NO: 15748), or a sequence having at least 90% sequence identity thereto.

21. The system of any one of claims 1-20, wherein the gRNA variant has a sequence comprising the sequence of SEQ ID NO: 2238, or a sequence with at least 90% sequence identity thereto.

22. The system of any one of claims 1-20, wherein the gRNA variant has a sequence comprising the sequence of SEQ ID NO: 26800, or a sequence having at least 90% sequence identity thereto.

23. The system of any one of claims 1-22, wherein the gRNA variant is chemically modified.

24. The system of any one of claims 1-23, wherein the chimeric CasX variant proteincomprises a sequence with at least 95% sequence identity to SEQ ID NO: 126.

25. The system of claim 24, wherein the chimeric CasX variant protein comprises a sequence with at least 99% sequence identity to SEQ ID NO: 126.

26. The system of any one of claims 1-23, wherein the CasX variant protein comprises a sequence of SEQ ID NO: 123, 133, or 143, or sequences having at least about 95% sequence identity thereto.

27. The system of any one of claims 1-26, wherein the CasX variant protein further comprises one or more nuclear localization signals (NLS).

28. The system of claim 27, wherein the one or more NLS are selected from the group of sequences consisting of PKKKRKV (SEQ ID NO: 168), KRPAATKKAGQAKKKK (SEQ ID NO: 169), PAAKRVKLD (SEQ ID NO: 170), RQRRNELKRSP (SEQ ID NO: 171), NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY (SEQ ID NO: 172), RMRIZFKNKGKDTAELRRRRVEVSVELRKAKKDEQILKRRNV (SEQ ID NO: 173), VSRKRPRP (SEQ ID NO: 174), PPKKARED (SEQ ID NO: 175), PQPKKKPL (SEQ ID NO: 176), SALIKKKKKMAP (SEQ ID NO: 177), DRLRR (SEQ ID NO: 178), PKQKKRK (SEQ ID NO: 179), RKLKKKIKKL (SEQ ID NO: 180), REKKKFLKRR (SEQ ID NO: 181), KRKGDEVDGVDEVAKKKSKK (SEQ ID NO: 182), RKCLQAGMNLEARKTKK (SEQ ID NO: 183), PRPRKIPR (SEQ ID NO: 184), PPRKKRTVV (SEQ ID NO: 185), NLSKKKKRKREK (SEQ ID NO: 186), RRPSRPFRKP (SEQ ID NO: 187), KRPRSPSS (SEQ ID NO: 188), KRGINDRNFWRGENERKTR (SEQ ID NO: 189), PRPPKMARYDN (SEQ ID NO: 190), KRSFSKAF (SEQ ID NO: 191), KLKIKRPVK (SEQ ID NO: 192), PKKKRKVPPPPAAKRVKLD (SEQ ID NO: 193), PKTRRRPRRSQRKRPPT (SEQ ID NO:26792), SRRRKANPTKLSENAKKLAKEVEN (SEQ ID NO: 194), KTRRRPRRSQRKRPPT (SEQ ID NO: 195), RRKKRRPRRKKRR (SEQ ID NO: 196), PKKKSRKPKKKSRK (SEQ ID NO: 197), HKKKHPDASVNFSEFSK (SEQ ID NO: 198), QRPGPYDRPQRPGPYDRP (SEQ ID NO: 199), LSPSLSPLLSPSLSPL (SEQ ID NO: 200), RGKGGKGLGKGGAKRHRK (SEQ ID NO: 201), PKRGRGRPKRGRGR (SEQ ID NO: 202), PKKKRKVPPPPAAKRVKLD (SEQ ID NO: 203), PKKKRKVPPPPKKKRKV (SEQ ID NO: 204), PAKRARRGYKC (SEQ ID NO: 27199), KLGPRKATGRW (SEQ ID NO: 27200), PRRKREE (SEQ ID NO: 27201), PYRGRKE (SEQ ID NO: 27202), PLRKRPRR (SEQ ID NO: 27203), PLRKRPRRGSPLRKRPRR (SEQ ID NO: 27204), PAAKRVKLDGGKRTADGSEFESPKKKRKV (SEQ ID NO: 27205), PAAKRVKLDGGKRTADGSEFESPKKKRKVGIHGVPAA (SEQ ID NO: 27206), PAAKRVKLDGGKRTADGSEFESPKKKRKVAEAAAKEAAAKEAAAKA (SEQ ID NO: 207), PAAKRVKLDGGKRTADGSEFESPKKKRKVPG (SEQ ID NO: 27208), KRKGSPERGERKRHW (SEQ ID NO: 27209), KRTADSQHSTPPKTKRKVEFEPKKKRKV (SEQ ID NO: 27210), and PKKKRKVGGSKRTADSQHSTPPKTKRKVEFEPKKKRKV (SEQ ID NO: 27211), wherein the one or more NLS are linked to the CasX variant protein or to adjacent NLS with a linker peptide, wherein the linker peptide is selected from the group consisting of RS, (G)n (SEQ ID NO: 27212), (GS)n (SEQ ID NO: 27213), (GSGGS)n (SEQ IDNO: 214), (GGSGGS)n (SEQ ID NO: 215), (GGGS)n (SEQ ID NO: 216), GGSG (SEQ ID NO: 217), GGSGG (SEQ ID NO: 218), GSGSG (SEQ ID NO: 219), GSGGG (SEQ ID NO: 220), GGGSG (SEQ ID NO: 221), GSSSG (SEQ ID NO: 222), GPGP (SEQ ID NO: 223), GGP, PPP, PPAPPA (SEQ ID NO: 224), PPPG (SEQ ID NO: 27214), PPPGPPP (SEQ ID NO: 225), PPP(GGGS)n (SEQ ID NO: 27215), (GGGS)nPPP (SEQ ID NO: 27216), AEAAAKEAAAKEAAAKA (SEQ ID NO: 27217), and TPPKTKRKVEFE (SEQ ID NO: 27218), wherein n is 1 to 5.

29. The system of claim 27 or claim 28, wherein the one or more NLS are located at or near the C-terminus of the CasX variant protein, at or near the N-terminus of the CasX variant protein, or at or near the N-terminus and at or near the C-terminus of the CasX variant protein.

30. The system of any one of claims 1-29, wherein the CasX variant protein forms a ribonuclear protein complex (RNP) with the gRNA variant.

31. The system of claim 30, wherein the RNP exhibits greater editing efficiency and / or binding of a target nucleic acid sequence when any one of the PAM sequences TTC, ATC, GTC, or CTC is located 1 nucleotide 5’ to the non-target strand of a protospacer having identity with the targeting sequence of the gRNA in a cellular assay system compared to the editing efficiency and / or binding of an RNP comprising a reference CasX protein of SEQ ID NOS: 1-3 and a reference gRNA in a comparable assay system.

32. The system of claim 31, wherein the PAM sequence is TTC.

33. The system of claim 32, wherein the targeting sequence of the gRNA variant comprises asequence selected from the group consisting of SEQ ID NOS: 17904-26789.

34. The system of claim 33, wherein the PAM sequence is ATC.

35. The system of claim 34, wherein the targeting sequence of the gRNA variant comprises asequence selected from the group consisting of SEQ ID NOS: 272-2100 and 2286-5625, 36. The system of claim 35, wherein the PAM sequence is CTC.

37. The system of claim 36, wherein the targeting sequence of the gRNA variant comprises asequence selected from the group consisting of SEQ ID NOS: 5626-13616.

38. The system of claim 37, wherein the PAM sequence is GTC.

39. The system of claim 38, wherein the targeting sequence of the gRNA variant comprises asequence selected from the group consisting of SEQ ID NOS: 13617-17903.

40. The system of any one of claims 31-39, wherein the increased editing and / or binding affinity for the one or more PAM sequences is at least 1.5-fold greater compared to the binding affinity of any one of the reference CasX proteins of SEQ ID NOS: 1-3 for the PAM sequences.

41. The system of any one of claims 31-39, wherein the RNP has at least a 5%, at least a 10%, at least a 15%, or at least a 20% higher percentage of cleavage-competent RNP compared to an RNP of the reference CasX protein and the gRNA of SEQ ID NOs: 4-16.

42. One or more nucleic acids comprising or encoding the CasX variant protein and the gRNA variant of the system of any one of claims 1-41.

43. The one or more nucleic acids of claim 42, wherein the sequence that encodes the CasX variant protein is codon optimized for expression in a eukaryotic cell.

44. A vector comprising the one or more nucleic acids of claim 42 or claim 43, wherein the vector is selected from the group consisting of a retroviral vector, a lentiviral vector, an adenoviral vector, an adeno-associated viral (AAV) vector, a herpes simplex virus (HSV) vector, a plasmid, a minicircle, a nanoplasmid, a DNA vector, an RNA vector, a lipid nanoparticle, and a liposome.

45. The vector of claim 44, wherein the vector is an AAV selected from AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV-Rh74, and AAVRhlO.

46. A host cell comprising the vector of claim 44 or claim 45.

47. A method of modifying a BCL11A target nucleic acid sequence in a cell in vitro or ex vivo, the method comprising introducing into the cell:a. the system of any one of claims 1-41;b. the one or more nucleic acids of claim 42 or claim 43;c. the vector of claim 44 or claim 45; ord. combinations thereof;wherein the BCL11A target nucleic acid sequence of a cell targeted by the gRNA variant is modified by the CasX variant protein.

48. The method of claim 47, wherein the modifying comprises introducing an insertion, deletion, substitution, duplication, or inversion of one or more nucleotides in the BCL11A target nucleic acid sequence of the cell.

49. The method of claim 48, wherein a GATA1 binding site sequence of the target nucleic acid is modified.

50. The method of any one of claims 47-49, wherein the cells are eukaryotic.

51. The method of claim 50, wherein the eukaryotic cells are selected from the group consisting of rodent cells, mouse cells, rat cells, and non-human primate cells.

52. The method of claim 51, wherein the eukaryotic cells are human cells selected from the group consisting of a hematopoietic stem cell (HSC), a hematopoietic progenitor cell (HPC), a CD34+ cell, a mesenchymal stem cell (MSC), induced pluripotent stem cell (iPSC), a common myeloid progenitor cell, a proerythroblast cell, and an erythroblast cell.

53. The system of any one of claims 1-41, the nucleic acid of claim 42 or claim 43, or the vector of claim 44 or claim 45, for use in the treatment of a hemoglobinopathy,54. A system for modifying a polypyrimidine tract-binding protein 1 (BCL11 A) gene target nucleic acid sequence, the system comprising a CasX variant protein and a guide ribonucleic acid (gRNA) variant, wherein the gRNA variant comprises:(i) a scaffold stem loop sequence of SEQ ID NO: 25; and(ii) a targeting sequence complementary to a target nucleic acid sequence comprising a region within a polypyrimidine tract-binding protein 1 (BCL11 A) gene, wherein the targeting sequence of the gRNA variant comprises a sequence selected from the group consisting of SEQ ID NOS: 22, 23, 2949, 2948, 15747, and 15748, wherein the system is capable of modifying the BCL11A gene, wherein the modifying comprises introducing an insertion, deletion, substitution, duplication, or inversion of one or more nucleotides in the BCL11A gene.

Citation Information

Patent Citations

  • Materials and methods for treatment of hemoglobinopathies

    US20190201553A1

  • Systems and methods for one-shot guide RNA (ogRNA) targeting of endogenous and source DNA

    US9963719B1

  • Novel CAS12b enzymes and systems

    WO2020033601A1