Compositions and methods for targeting BCL11A
A modified CRISPR/CasX system targets the BCL11A gene in vivo to increase γ-globin expression, addressing the limitations of ex vivo editing and offering a therapeutic solution for hemoglobin disorders like sickle cell disease and β-thalassemia.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- SCRIBE THERAPEUTICS INC
- Filing Date
- 2021-12-02
- Publication Date
- 2026-05-12
AI Technical Summary
Current CRISPR/Cas systems for treating hemoglobin disorders are limited to ex vivo cell editing and there is a need for compositions and methods to modulate BCL11A gene expression in vivo to reduce its suppression of the γ-globin gene, which can alleviate conditions like sickle cell disease and β-thalassemia.
A modified Class 2, Type V CRISPR protein (CasX) and guide nucleic acid system (gRNA) are used to passively enter cells and target the BCL11A gene, allowing for its knockdown or knockout, potentially increasing γ-globin expression.
This approach enhances γ-globin expression in vivo, providing a therapeutic strategy to reduce the clinical severity of β-abnormal hemoglobin disorders such as sickle cell disease and β-thalassemia.
Smart Images

Figure 0007857026000038 
Figure 0007857026000039 
Figure 0007857026000040
Abstract
Description
[Technical Field]
[0001] Cross-references to related applications This application claims priority to U.S. Provisional Patent Application No. 63 / 1200233,885, filed on 3 December 2020, the contents of which are incorporated herein by reference in their entirety.
[0002] Reference to sequence listings This application includes an ASCII-formatted sequence listing submitted via EFS-Web, which is incorporated herein by reference in its entirety. The ASCII copy was created on December 1, 2021, and is named SCRB_030_01WO_SeqList_ST25.txt with a file size of 8.78 MB. [Background technology]
[0003] Fetal hemoglobin (also known as hemoglobin F, HbF, or α2γ2) is the primary oxygen carrier protein in the human fetus. HbF has a different composition from adult hemoglobin, allowing it to bind more strongly to oxygen and enabling the developing fetus to reabsorb oxygen from the mother's bloodstream. HbF is a tetramer of two adult α-globin polypeptides and two fetal β-like γ-globin polypeptides. During pregnancy, the replicated γ-globin gene constitutes the main gene transcribed in the β-globin cluster. After birth, γ-globin is replaced by adult β-globin. This process is called the "fetal switch" and involves the expression of BCL11A (a regulator of HbF silencing) (Sankaran, VG, et al. Human Fetal Hemoglobin Expression Is Regulated by the Developmental Stage-Specific Repressor BCL11A. Science 322(5909):1839-1842(2008), Liu, N., et al. Direct Promoter Repression by BCL11A Controls the Fetal to Adult Hemoglobin Switch. Cell 173(2):430(2018)). In healthy adults, hemoglobin is composed of hemoglobin A (approximately 97%), hemoglobin A2 (2.2–3.5%), and hemoglobin F (<1%) (Thomas, C and Lumb, AB Physiology of hemoglobin. Continuing Education in Anaesthesia Critical Care & Pain. 12(5):251-256 (2012)).
[0004] Hemoglobin disorders are, in most cases, hereditary monogenic disorders inherited as autosomal codominant traits. Common hemoglobin disorders include sickle cell anemia, as well as alpha-thalassemia and beta-thalassemia. Hemoglobin disorders are most common in people from Africa, the Mediterranean region, and Southeast Asia. Most hemoglobin disorders, including sickle cell anemia, are simply structural abnormalities of the globin protein itself. Sickle cell anemia arises from a point mutation in the β-globin structural gene HBB, resulting in the production of abnormal hemoglobin (HbS) and a decrease in the blood's oxygen-carrying capacity. Thalassemia, in contrast, usually results in insufficient production of normal globin protein, often leading to a deficiency or absence of adult hemoglobin (HbA) through mutations in regulatory genes. In β-thalassemia, β-globin is deficient, and increased γ-globin expression reduces the imbalance between α-globin and β-globin chains that underlies the pathophysiology of anemia in this condition (Liu, N., et al. Direct Promoter Repression by BCL11A Controls the Fetal to Adult Hemoglobin Switch. Cell 173(2):430 (2018)). Both sickle cell disease and thalassemia can cause anemia.
[0005] B-cell lymphoma / leukemia 11A (BCL11A) is a protein encoded by the BCL11A gene in humans. During hematopoietic cell differentiation, this gene is downregulated and has been found to play a role in suppressing fetal hemoglobin production. BCL11A is the major repressor protein for hemoglobin F production by binding to the promoter region of the gene encoding the γ subunit (Sankaran VG, et al. Human fetal hemoglobin expression is regulated by the developmental stage-specific repressor BCL11A. Science 322:1839 (2008)). Increased gamma globin reduces the clinical severity of β-abnormal hemoglobin disorders, sickle cell disease, and β-thalassemia, which are caused by mutations or decreased expression of β-globin. Therefore, gene editing of BCL11A to increase gamma globin expression beyond the remaining approximately 1% of fetal hemoglobin has been proposed as an attractive therapeutic strategy in adults with abnormal hemoglobin disorders (Smith, EC, et al. Strict in vivo specificity of the Bcl11a erythroid enhancer. Blood 128(19):2338(2016)).
[0006] The emergence of CRISPR / Cas systems and the programmable nature of these minimal systems has facilitated their use as versatile technologies for genomic manipulation and engineering. To date, the use of CRISPR / Cas systems for the treatment of hemoglobin disorders has been limited to ex vivo cell editing followed by transplantation into subjects suffering from the underlying hemoglobin disorder. Therefore, in subjects with these diseases, there is a need for compositions and methods to modulate BCL11A to reduce direct in vivo suppression of the γ-globin gene promoter. To address this need, compositions and methods for targeting the BCL11A gene are provided herein. [Overview of the project]
[0007] This disclosure relates to a composition of a modified class 2, type V CRISPR protein and a guide nucleic acid used to modify a target nucleic acid containing the BCL11A gene in cells. The class 2, type V CRISPR protein and guide nucleic acid are modified to passively enter target cells. The class 2, type V CRISPR protein and guide nucleic acid are useful in various methods for modifying the target nucleic acid of BCL11A-related diseases, and these methods are also provided.
[0008] In one embodiment, the disclosure relates to a CasX: guide nucleic acid system (CasX: gRNA system) and a method used to knock down or knock out the BCL11A gene for reducing or eliminating the expression of the BCL11A gene product in subjects with β-abnormal hemoglobinopathy-related diseases.
[0009] In some embodiments, the gRNA in the CasX:gRNA system is a gRNA or a chimera of RNA and DNA, and may be a single-molecule gRNA or a dimorphic-molecule gRNA. In other embodiments, the gRNA in the CasX:gRNA system has a targeting sequence complementary to a target nucleic acid sequence containing a region within the BCL11A gene. In some embodiments, the targeting sequence of the gRNA is selected from the group consisting of sequence numbers 272-2100 and 2286-26789, or sequences having at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 85%, at least about 90%, or at least about 95% identity thereto. The gRNA may contain a targeting sequence containing 15-20 consecutive nucleotides. In other embodiments, the targeting sequence of the gRNA consists of 20 nucleotides. In other embodiments, the targeting sequence consists of 19 nucleotides. In other embodiments, the targeting sequence consists of 18 nucleotides. In other embodiments, the targeting sequence consists of 17 nucleotides. In other embodiments, the targeting sequence consists of 16 nucleotides. In other embodiments, the targeting sequence consists of 15 nucleotides. In other embodiments, the targeting sequence of the gRNA has a sequence selected from the group consisting of SEQ ID NOs. 272-2100 and 2286-26789. In other embodiments, the targeting sequence of the gRNA has a sequence selected from the group consisting of SEQ ID NOs. 272-2100 and 2286-26789, with a single nucleotide removed from the 3' end of the sequence. In other embodiments, the targeting sequence consists of 18 nucleotides, has a sequence selected from the group consisting of SEQ ID NOs. 272-2100 and 2286-26789, with two nucleotides removed from the 3' end of the sequence. In other embodiments, the targeting sequence consists of 17 nucleotides, has a sequence selected from the group consisting of SEQ ID NOs. 272-2100 and 2286-26789, with three nucleotides removed from the 3' end of the sequence. In other embodiments, the targeting sequence consists of 16 nucleotides and is selected from the group consisting of SEQ ID NOs. 272-2100 and 2286-26789, with four nucleotides removed from the 3' end of the sequence.In other embodiments, the targeting sequence consists of 15 nucleotides and is selected from the group consisting of SEQ ID NOs. 272-2100 and 2286-26789, with 5 nucleotides removed from the 3' end of the sequence.
[0010] In some embodiments, the gRNA has a scaffold containing a sequence selected from the group consisting of sequence numbers 2238-2285, 26794-26839, and 27219-27265, or a sequence shown in Table 3, or a sequence having at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity thereto. In some embodiments, the gRNA has a scaffold containing a sequence selected from the group consisting of sequence numbers 2238-2285, 26794-26839, and 27219-27265. In some embodiments, the gRNA has a scaffold containing a sequence selected from the group consisting of sequence numbers 2101-2285, 26794-26839, and 27219-27265.
[0011] In some embodiments, the CasX:gRNA system includes a CasX variant sequence having a sequence selected from the group consisting of SEQ ID NOs: 36-99, 101-148, and 26908-27154, or a sequence having at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity thereto. In some embodiments, the CasX:gRNA system includes a CasX variant sequence having a sequence selected from the group consisting of SEQ ID NOs: 36-99, 101-148, and 26908-27154. In some embodiments, the CasX:gRNA system includes a CasX variant sequence having a sequence selected from the group consisting of SEQ ID NOs. 59, 72-99, 101-148, and 26908-27154, or a sequence shown in Table 4, or a sequence having at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity thereto. In some embodiments, the CasX:gRNA system includes a CasX variant sequence having a sequence selected from the group consisting of SEQ ID NOs. 59, 72-99, 101-148, and 26908-27154. In some embodiments, the CasX:gRNA system includes a CasX variant sequence having a sequence selected from the group consisting of SEQ ID NOs: 132-148 and 26908-27154, or a sequence having at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity thereto. In some embodiments, the CasX:gRNA system includes a CasX variant sequence having a sequence selected from the group consisting of SEQ ID NOs: 132-148 and 26908-27154.In these embodiments, the CasX variant exhibits one or more improved features compared to any one of the reference CasX proteins of SEQ ID NOs: 1–3. In some embodiments, the CasX variant protein has a binding affinity to a protospacer adjacent motif (PAM) sequence selected from the group consisting of TTC, ATC, GTC, and CTC. In some embodiments, the CasX variant protein has a binding affinity to PAM sequences that is at least 1.5 times greater than the binding affinity to any one of the reference CasX proteins of SEQ ID NOs: 1–3 to a PAM sequence selected from the group consisting of TTC, ATC, GTC, and CTC.
[0012] In other embodiments of the CasX:gRNA system, the CasX molecule and the gRNA molecule associate together in a ribonuclear protein (RNP). In certain embodiments, an RNP containing a CasX variant and a gRNA variant exhibits higher editing efficiency and / or binding of the target sequence on target DNA compared to an RNP containing a reference CasX protein and reference gRNA in a comparable assay system, when one of the PAM sequences TTC, ATC, GTC, or CTC is located one nucleotide 5' to the non-target strand sequence that is identical to the gRNA's targeting sequence in the cell assay system.
[0013] In some embodiments, the CasX:gRNA system further comprises a donor template containing at least a portion of the BCL11A gene and a nucleic acid having at least 1 to about 5 mutations compared to the wild-type sequence, wherein the BCL11A gene portion is selected from the group consisting of BCL11A exons, BCL11A introns, BCL11A intron-exon junctions, BCL11A regulatory elements, or combinations thereof, and the donor template is used to knock down or knock out the BCL11A gene. In some cases, the donor sequence is a single-stranded DNA template or a single-stranded RNA template. In other cases, the donor template is a double-stranded DNA template.
[0014] In other embodiments, this disclosure relates to nucleic acids encoding any of the CasX:gRNA systems described herein, as well as vectors comprising nucleic acids. In some embodiments, the vector is selected from the group consisting of retroviral vectors, lentiviral vectors, adenovirus vectors, adeno-associated viral (AAV) vectors, herpes simplex virus (HSV) vectors, plasmids, minicircles, nanoplasmides, and RNA vectors. In other embodiments, the vector is a CasX delivery particle (XDP) comprising an RNP of CasX and gRNA from any of the embodiments described herein, and optionally a donor template nucleic acid and a targeting moiety (e.g., a viral glycoprotein).
[0015] In other embodiments, the Disclosure provides a method for modifying a BCL11A target nucleic acid sequence in a cell population, the method comprising introducing into cells a) a CasX:gRNA system of any embodiment disclosed herein, b) a nucleic acid of any embodiment disclosed herein, c) a vector of any embodiment disclosed herein, d) an XDP of any embodiment disclosed herein, or e) a combination thereof. In some embodiments of the Method, the modification involves introducing one or more nucleotide insertions, deletions, substitutions, duplications, or inversions in the target nucleic acid sequence compared to the wild-type sequence. The target BCL11A gene includes a GATA1 erythrocyte-specific enhancer binding site (GATA1) as a regulatory element. In some embodiments, the modification method comprises modifying the GATA1 sequence, and the BCL11A gene is knocked down or knocked out by the modification. In some cases, the Method further comprises contacting the target nucleic acid with a donor template nucleic acid of any embodiment disclosed herein. In some embodiments of this method, the donor template comprises a nucleic acid containing at least a portion of the BCL11A gene but having one or more mutations for knocking out or knocking down the BCL11A gene. In some cases, the modification of the target nucleic acid sequence occurs in vitro or ex vivo. In some embodiments, the modification of the target nucleic acid sequence occurs in vivo. In some embodiments, the cells are eukaryotic cells selected from the group consisting of rodent cells, mouse cells, rat cells, primate cells, and non-human primate cells. In some embodiments, the cells are human cells. In some embodiments, the cells are selected from the group consisting of hematopoietic stem cells (HSCs), hematopoietic progenitor cells (HPCs), CD34+ stem cells, mesenchymal stem cells (MSCs), induced pluripotent stem cells (iPSCs), myeloid common progenitor cells, proerythroblasts, and erythroblasts. In some embodiments, the cells are autologous cells derived from a subject with a β-abnormal hemoglobinopathy-related disorder.In other embodiments, the cells are of the same species as the target being treated.
[0016] In other embodiments, the Disclosure provides a method for modifying a target nucleic acid sequence of the BCL11A gene, wherein a target cell population is contacted using a vector comprising the CasX protein and one or more gRNAs containing a targeting sequence complementary to the BCL11A gene, and optionally further comprising a donor template. In some cases, the vector is an adeno-associated virus (AAV) vector selected from AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV-Rh74, or AAVRh10. In other cases, the vector is a lentiviral vector. In other embodiments, the Disclosure provides a method for contacting target cells using a vector, wherein the vector is a CasX delivery particle (XDP) comprising the RNP of CasX and gRNA from any of the embodiments described herein, and optionally a donor template nucleic acid. In some embodiments of the Method, the vector is administered to the subject in a therapeutically effective dose. The target organisms may be mice, rats, pigs, non-human primates, or humans. The dose can be administered via a route of administration selected from transplantation, local injection, systemic infusion, or a combination thereof.
[0017] In other embodiments, the Disclosure provides a method for treating a subject in need of treatment for a β-hemoglobinopathy-related disorder, comprising modifying the gene encoding the BCL11A gene in cells of the subject, the modification comprising contacting the cells with any of the following: a) a CasX:gRNA system of any embodiment disclosed herein, b) a nucleic acid of any embodiment disclosed herein, c) a vector of any embodiment disclosed herein, d) an XDP of any embodiment disclosed herein, or e) a combination thereof. In some embodiments, the β-hemoglobinopathy-related disorder is sickle cell anemia or β-thalassemia. In some cases, a method for treating a subject with a β-hemoglobinopathy-related disorder results in an improvement in at least one clinically relevant parameter. In other cases, a method for treating a subject with a β-hemoglobinopathy-related disorder results in an improvement in at least two clinically relevant parameters.
[0018] In other embodiments, the Disclosure provides the use of the CasX:gRNA system, nucleic acid, vector, or XDP described herein for treating subjects requiring treatment for β-abnormal hemoglobinopathy-related disorders. In some embodiments, the use involves modifying a gene encoding the BCL11A gene in a cell of interest, and the modification involves contacting the cell with any of the following: a) a CasX:gRNA system of any embodiment disclosed herein, b) a nucleic acid of any embodiment disclosed herein, c) a vector of any embodiment disclosed herein, d) an XDP of any embodiment disclosed herein, or e) any combination thereof.
[0019] Reference All publications, patents, and patent applications referenced herein are incorporated herein by reference as if each individual publication, patent, or patent application were specifically and individually indicated to be incorporated by reference. The contents of U.S. Provisional Applications No. 63 / 121,196, filed 3 December 2020, No. 63 / 162,346, filed 17 March 2021, and No. 63 / 208,855, filed 9 June 2021, which disclose CasX variants and gRNA variants, are incorporated herein by reference in their entirety. The contents of International Patent Publication Nos. 2020 / 247882, 2020 / 247883, published on December 10, 2020, and 2021 / 113772, published on June 10, 2021, are incorporated herein by reference in their entirety. [Brief explanation of the drawing]
[0020] The novel features of this disclosure are described in detail in the attached claims. The features and advantages of this disclosure will be better understood by referring to the following detailed description of exemplary embodiments utilizing the principles of this disclosure, as well as the attached drawings. [Figure 1] This graph shows the results of an assay to quantify the active fraction of RNP formed by sgRNA174 (SEQ ID NO: 2238) and CasX variants 119 (SEQ ID NO: 59), 457 (SEQ ID NO: 101), 488 (SEQ ID NO: 123), and 491 (SEQ ID NO: 126), as described in Example 8. Equimolar amounts of RNP and target were co-incubated, and the amount of cleaved target was determined at the indicated time points. The mean and standard deviation of three independent replicates are shown for each time point. The two-phase fit of the combined replicas is shown. "2" refers to the reference CasX protein of SEQ ID NO: 2. [Figure 2]As described in Example 8, the quantification of the active fraction of RNP formed by CasX2 (reference CasX protein of SEQ ID NO: 2) and modified sgRNA is shown. Equimolar amounts of RNP and target were co-incubated, and the amount of cleaved target was determined at the indicated time points. The mean and standard deviation of three independent replicates are shown for each time point. The two-phase fit of the combined replicas is shown. [Figure 3] As described in Example 8, the quantification of the active fraction of RNP formed by CasX 491 and modified sgRNA under guide-restriction conditions is shown. Equimolar amounts of RNP and target were co-incubated, and the amount of cleaved target was determined at the indicated time. Two-phase fit of the data is shown. [Figure 4] The quantification of the cleavage rate of RNPs formed by sgRNA174 and CasX variants is shown, as described in Example 8. Target DNA was incubated with a 20-fold excess of the indicated RNPs, and the amount of target cleaved was determined at the indicated time points. The mean and standard deviation of three independent repeats are shown for each time point, except for 488 and 491, where single repeats are shown. Monophase fitting of combined repeats is shown. [Figure 5] As described in Example 8, the quantification of the cleavage rate of RNP formed by CasX2 and the indicated sgRNA variant is shown. Target DNA was incubated with a 20-fold excess of the indicated RNP, and the amount of target cleaved was determined at the indicated time points. The mean and standard deviation of three independent replicates are shown for each time point. A uniphase fit of the combined replicates is shown. [Figure 6] As described in Example 8, the initial velocity of RNPs formed by CasX2 and sgRNA variants is quantified. The initial cleavage velocity was determined by fitting the first two time points of the previous cleavage experiment to a linear model. [Figure 7] As described in Example 8, the cleavage rate of RNPs formed by CasX491 and sgRNA variants is quantified. Target DNA was incubated with a 20-fold excess of the indicated RNPs at 10°C, and the amount of target cleaved was determined at the indicated time points. Single-phase fit at the time points is shown. [Figure 8] As described in Example 8, the quantification of competent fractions of RNPs of CasX variants 515 (SEQ ID NO: 133) and 526 (SEQ ID NO: 143) complexed with gRNA variant 174 is shown, compared to the RNPs of reference CasX 2 complexed with gRNA 2, using equimolar amounts of indicated RNPs and complementary targets. Two-phase fit is shown for each time course or combined set of replications. [Figure 9] As described in Example 8, the cleavage rates of RNPs of CasX variants 515 and 526 complexed with gRNA variant 174 are quantified using a 20-fold excess of the indicated RNPs, compared to the RNPs of reference CasX 2 complexed with gRNA 2. [Figure 10A] As described in Example 5, the quantification of the cleavage rate of the CasX variant against TTC PAM is shown. Target DNA substrates having the same spacer and the indicated PAM sequence were incubated with a 20-fold excess of the indicated RNP at 37°C, and the amount of cleaved target was determined at the indicated time point. Single repeat monophase fitting is shown. [Figure 10B] As described in Example 5, the quantification of the cleavage rate of the CasX variant against CTC PAM is shown. Target DNA substrates having the same spacer and the indicated PAM sequence were incubated with a 20-fold excess of the indicated RNP at 37°C, and the amount of cleaved target was determined at the indicated time point. Single repeat monophase fitting is shown. [Figure 10C] As described in Example 5, the quantification of the cleavage rate of the CasX variant against GTC PAM is shown. Target DNA substrates having the same spacer and the indicated PAM sequence were incubated with a 20-fold excess of the indicated RNP at 37°C, and the amount of target cleaved was determined at the indicated time point. Single repeat monophase fitting is shown. [Figure 10D]As described in Example 5, the quantification of the cleavage rate of the CasX variant against ATC PAM is shown. Target DNA substrates having the same spacer and the indicated PAM sequence were incubated with a 20-fold excess of the indicated RNP at 37°C, and the amount of target cleaved was determined at the indicated time point. Single repeat monophase fitting is shown. [Figure 11A] As described in Example 5, the quantification of the cutting rate of RNPs of CasX variant 491 and guide 174 against NTC PAM is shown. Time points were acquired over a 10-minute period, and the cut fractions were graphed for each target and time point, but for clarity, only the first 2 minutes of the time course are shown. [Figure 11B] As described in Example 5, the quantification of the RNP cutting rate of CasX variant 491 and guide 174 against NTT PAM is shown. Time points were acquired over a 10-minute period, and the cut fractions were graphed for each target and time point. [Figure 12A] As described in Example 9, the quantification of cleavage by RNP formed by sgRNA174 and CasX variant 515 is shown using 18, 19, or 20 nucleotide-length spacers. Target DNA was incubated with a 20-fold excess of the indicated RNP, and the amount of cleaved target was determined at the indicated time points. The mean and standard deviation of three independent replicates are shown for each time point. A uniphase fit of combined replicates is shown. [Figure 12B] As described in Example 9, the quantification of cleavage by RNP formed by sgRNA174 and CasX variant 526 is shown using 18, 19, or 20 nucleotide spacers. Target DNA was incubated with a 20-fold excess of the indicated RNP, and the amount of cleaved target was determined at the indicated time points. The mean and standard deviation of three independent replicates are shown for each time point. A uniphase fit of combined replicates is shown. [Figure 13]This is a schematic diagram showing an example of the CasX protein and scaffold DNA sequence for packaging in adeno-associated virus (AAV). The DNA segments between AAV inverted terminal repeats (ITRs), consisting of the CasX-coding DNA and its promoter, as well as the scaffold-coding DNA and its promoter, are packaged within the AAV capsid during AAV production. [Figure 14] This paper presents the results of an editing assay comparing gRNA scaffolds 229–237 (see Table 3 for corresponding sequences and SEQ ID NOs) and scaffold 174 in mouse neural progenitor cells (mNPCs) isolated from Ai9-tdtomato transgenic mice. Cells were nucleofected with CasX 491, scaffolds, and the p59 plasmid encoding spacer 11.30 (5'AAGGGGCUCCGCACCACGCC3', SEQ ID NO 27197) that targets mRHO. Editing at the mRHO locus was evaluated by NGS 5 days after transfection, showing that editing with constructs containing scaffolds 230, 231, 234, and 235 demonstrated greater editing compared to constructs containing scaffold 174 at both doses. [Figure 15] The results of an editing assay comparing gRNA scaffolds 229–237 and scaffold 174 in mNPC cells are presented. Cells were nucleofected with CasX 491, scaffolds, and p59 plasmid encoding spacer 12.7 (5'CUGCAUUCUAGUUGUGGUUU 3', SEQ ID NO: 27198), targeting repeat elements that prevent the expression of tdTomato fluorescent protein, at indicated doses. Five days after transfection, editing was assessed by FACS to quantify the percentage of tdTomato-positive cells. Cells nucleofected with scaffolds 231–235 showed approximately 35% greater editing at high doses and approximately 25% greater editing at low doses compared to constructs with scaffold 174. [Figure 16]This report presents the results of an editing assay comparing CasX nucleases 2, 119, 491, 515, 527, 528, 529, 530, and 531 (see Table 4 for corresponding sequences and sequence numbers) in the custom HEK293 cell line PASS_V1.01. Cells were lipofected with 2 μg of p67 plasmid encoding the indicated CasX proteins. After 5 days, genomic DNA was extracted from the cells. PCR amplification and next-generation sequencing were performed to isolate and quantify the fraction of cells edited at custom-designed on-target editing sites. For each sample, editing was evaluated at target sites (individual points) consisting of the following PAM sequences: 48 TTCs, 14 ATCs, 22 CTCs, and 11 GTCs, and the percentage of editing (%) was normalized to the vehicle control. Cells lipofected with any nuclease showed higher mean editing at TTC PAM target sites (horizontal bars) than those of the wild-type nuclease CasX 2, except for CasX 528. The relative priority of any given nuclease for four different PAM sequences is also shown in a violin plot. In particular, CasX nucleases 527, 528, and 529 show substantially different PAM priorities than the wild-type nuclease CasX 2. [Figure 17]This report presents the results of an editing assay comparing improved CasX nucleases 491 with improved CasX nucleases 532 and 533 in the custom HEK293 cell line PASS_V1.01. Cells were lipofected in double denominations with 2 μg of p67 plasmid encoding the indicated CasX protein and puromycin resistance gene, and grown under puromycin selectivity. After 3 days, genomic DNA was extracted from the cells. PCR amplification and next-generation sequencing were performed to isolate and quantify the fraction of cells edited at custom-designed on-target editing sites. For each sample, editing was evaluated at target sites consisting of the following PAM sequences: individual sites of 48 TTCs, 14 ATCs, 22 CTCs, and 11 GTCs, and the percentage of editing was normalized to the vehicle control. Cells lipofected with CasX 532 or 533 showed higher mean editing than Cas 491 at each of the PAM sequences, except for CasX 533 at the TTC PAM target site. The error bars represent the standard error of the mean for a biological sample of n=2. [Figure 18] As described in Example 13, the results of editing the BCL11A erythrocyte enhancer locus in HEK293T cells using the CasX protein variant 438 with scaffold 174, compared to the Cas9 system, are shown. [Figure 19] As described in Example 14, the results of editing the GATA1-binding region of the BCL11A erythrocyte enhancer locus in K562 cells by the CasX protein variant 491 having scaffold 174, compared to the CasX protein variant 119 having scaffold 174, are shown. [Figure 20] As described in Example 14, we show the results of editing the GATA1-binding region of the BCL11A erythrocyte enhancer locus in K562 cells using a CasX protein variant 491 with scaffold 174 delivered by various doses of XDP. [Figure 21]As described in Example 15, the results of editing the GATA1-binding region of the BCL11A erythrocyte enhancer locus in HSC cells by the CasX protein variant 491 having scaffold 174, compared to the CasX protein variant 119 having scaffold 174, are shown. [Figure 22] As described in Example 15, we show the results of editing the GATA1-binding region of the BCL11A erythrocyte enhancer locus in HSC cells using a CasX protein variant 491 having scaffold 174 delivered by various doses of XDP. [Figure 23] This is a schematic diagram showing the placement of spacer 21.1 (sequence number 22) relative to the GATA1 binding site sequence in the target nucleic acid. Top strand: sequence number 26790, bottom strand: sequence number 26791. [Modes for carrying out the invention]
[0021] While exemplary embodiments have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided merely as examples. Those skilled in the art will recall numerous variations, modifications, and substitutions without departing from the claimed invention as described herein. It should be understood that various alternative forms to those described herein may be used when carrying out embodiments of this disclosure. The claims define the scope of the invention, and the methods and structures within these claims, as well as their equivalents, are intended to be encompassed thereby.
[0022] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those generally understood by those skilled in the art in which the present invention pertains. Any methods and materials similar to or equivalent to those described herein may be used in the implementation or testing of these embodiments, but preferred methods and materials are described below. In case of any conflict, the present specification, including the definitions, shall prevail. In addition, the materials, methods, and examples are illustrative and not intended to limit the scope of the invention. Those skilled in the art will be able to conceive of numerous variations, modifications, and substitutions without departing from the present invention.
[0023] definition The terms “polynucleotide” and “nucleic acid” are used interchangeably herein and refer to polymeric forms of nucleotides of any length, which are either ribonucleotides or deoxyribonucleotides. Accordingly, the terms “polynucleotide” and “nucleic acid” encompass polymers containing single-stranded DNA, double-stranded DNA, multi-stranded DNA, single-stranded RNA, double-stranded RNA, multi-stranded RNA, genomic DNA, cDNA, DNA-RNA hybrids, and purines and pyrimidine bases or other natural, chemically or biochemically modified, unnatural, or derivatized nucleotide bases.
[0024] The terms "hybridizable" and "complementary" are used interchangeably and mean that a nucleic acid (e.g., RNA, DNA) contains a sequence of nucleotides that allows it to non-covalently bond to another nucleic acid in a sequence-specific antiparallel manner (i.e., form Watson-Crick base pairs and / or G / U base pairs) under appropriate in vitro and / or in vivo conditions of temperature and ionic strength of solution, "annealing" or "hybridizing" (i.e., the nucleic acid specifically binding to a complementary nucleic acid). It is understood that the sequence of a polynucleotide does not need to be 100% complementary to the sequence of its target nucleic acid in order to be specifically hybridizable. It may have at least about 70%, at least about 80%, at least about 90%, or at least about 95% sequence identity and still be able to hybridize to the target nucleic acid. Furthermore, a polynucleotide can hybridize across one or more segments without intervening or adjacent segments being involved in the hybridization event (e.g., loop or hairpin structures, "bulges," "bubbles," etc.).
[0025] For the purposes of this disclosure, “gene” includes DNA regions that code for a gene product (e.g., protein, RNA), and all DNA regions that regulate the production of the gene product (whether such regulatory element sequences are adjacent to the coding sequence and / or transcription sequence). Thus, a gene may include, but is not limited to, promoter sequences, terminators, translation regulatory sequences (e.g., ribosome binding sites and internal ribosome entry sites), enhancers, silencers, insulators, boundary elements, origins of replication, matrix attachment sites, and locus regulatory regions. The coding sequence codes for a gene product during transcription or transcription and translation. The coding sequence in this disclosure may include fragments of an open reading frame, but does not need to include the full length. A gene may include both the transcribed strand (e.g., the strand containing the coding sequence) and a complementary strand.
[0026] The term "downstream" refers to a nucleotide sequence located 3' to the reference nucleotide sequence. In certain embodiments, the downstream nucleotide sequence relates to the sequence following the transcription start site. For example, the translation start codon of a gene is located downstream of the transcription start site.
[0027] The term "upstream" refers to a nucleotide sequence located 5' to the reference nucleotide sequence. In certain embodiments, the upstream nucleotide sequence refers to a sequence located 5' to the coding region or transcription start site. For example, most promoters are located upstream of the transcription start site.
[0028] With respect to a polynucleotide or amino acid sequence, the term “adjacent to” refers to sequences that are adjacent to or adjacent to each other in a polynucleotide or polypeptide. Those skilled in the art will understand that two sequences can be considered adjacent to each other and still contain a limited number of intervening sequences, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides or amino acids.
[0029] The term “regulatory element” is used herein interchangeably with the term “regulatory sequence” and is intended to include promoters, enhancers, and other expression regulatory elements (e.g., transcription termination signals such as polyadenylation signals and poly-U sequences). Exemplary regulatory elements include, but are not limited to, transcription promoters (e.g., CMV, CMV+intron A, SV40, RSV, HIV-Ltr, elongation factor 1 alpha (EF1α), MMLV-ltr, internal ribosome entry site (IRES) or P2A peptides (enabling translation of multiple genes from a single transcript)), metallothionein, transcription enhancer elements, transcription termination signals, polyadenylation sequences, sequences for optimizing translation initiation, and translation termination sequences. It is understood that the selection of an appropriate regulatory element depends on the encoded component being expressed (e.g., protein or RNA) or whether the nucleic acid contains multiple components requiring different polymerases or is not intended to be expressed as a fusion protein.
[0030] The term “promoter” refers to a DNA sequence that includes an RNA polymerase binding site, a transcription start site, a TATA box, and / or a B recognition element, and that assists or promotes the transcription and expression of the associated transcribed polynucleotide sequence and / or gene (or transgene). Promoters may be synthetically produced or derived from known or naturally occurring promoter sequences or other promoter sequences. Promoters may be proximal or distal to the gene being transcribed. Promoters may also include chimeric promoters that include a combination of two or more heterologous sequences to confer specific characteristics. Promoters in this disclosure may include variants of promoter sequences that are similar in composition but are not identical to known or other promoter sequences provided herein. Promoters may be classified according to criteria relating to the expression pattern of the associated coding sequence or transcribed sequence or gene operably linked to the promoter (e.g., constitutive, developmental, tissue-specific, inducible, etc.).
[0031] The term "enhancer" refers to a regulatory element DNA sequence that modulates the expression of a related gene when bound to a specific protein called a transcription factor. Enhancers may be located in the intron of a gene, or at the 5' or 3' of the gene's coding sequence. Enhancers may be located proximal to the gene (i.e., within tens or hundreds of base pairs (bp) of the promoter) or distal to the gene (i.e., thousands, hundreds of thousands, or even millions of bp away from the promoter). A single gene may be regulated by two or more enhancers, all of which are assumed to be within the scope of this disclosure.
[0032] As used herein, a “post-transcriptional regulatory element (PRE)” (e.g., a hepatitis PRE) refers to a DNA sequence that, when transcribed, generates a tertiary structure that may exhibit post-transcriptional activity, thereby enhancing or promoting the expression of an associated gene operably linked to it.
[0033] The term "GATA binding site" refers to the DNA binding site of the GATA family of transcription factors. GATA transcription factors typically recognize target sites that match the consensus sequence WGATAR (where W=A or T, and R=A or G).
[0034] As used herein, “recombinant” means that a particular nucleic acid (DNA or RNA) is the product of various combinations of cloning, restriction, and / or ligation steps, resulting in a construct having a structurally recognizable coding or non-coding sequence from endogenous nucleic acids found in natural systems. Generally, DNA sequences encoding structural coding sequences can be assembled from cDNA fragments and short oligonucleotide linkers, or from a series of synthetic oligonucleotides, providing synthetic nucleic acids that can be expressed from recombinant transcription units contained in cells or cell-free transcription and translation systems. Such sequences may typically be provided in the form of an open reading frame uninterrupted by internal non-coding sequences or introns present in eukaryotic genes. Genomic DNA containing the relevant sequences can also be used in the formation of recombinant genes or transcription units. The non-coding DNA sequence may be located at 5' or 3' of the open reading frame, where such sequence does not interfere with the manipulation or expression of the coding region, but can actually act to regulate the production of the desired product by various mechanisms (see “enhancer” and “promoter” above).
[0035] The terms “recombinant polynucleotides” or “recombinant nucleic acids” refer to those that do not exist naturally, for example, those created through human intervention by artificially combining two otherwise separated sequences. This artificial combination is often achieved by either chemical synthesis or artificial manipulation of isolated nucleic acid segments (e.g., genetic engineering techniques). This can typically be done by substituting codons with degenerate codons encoding the same or conserved amino acids, while introducing or removing sequence recognition sites. Alternatively, it can be done by ligating nucleic acid segments with desired functions together to produce a combination of desired functions. This artificial combination is often achieved by either chemical synthesis or artificial manipulation of isolated nucleic acid segments (e.g., genetic engineering techniques).
[0036] Similarly, the terms “recombinant polypeptide” or “recombinant protein” refer to polypeptides or proteins that do not exist in nature, for example, those created through human intervention by artificially combining two segments of an otherwise separated amino sequence. Therefore, for example, a protein containing a heterologous amino acid sequence is a recombinant.
[0037] As used herein, the term “in contact” means to establish a physical connection between two or more entities. For example, bringing a target nucleic acid sequence into contact with a guide nucleic acid means that the target nucleic acid sequence and the guide nucleic acid are made to share a physical connection (for example, they can hybridize if they share sequence similarity).
[0038] "Dissociation constant" or "K" d The term "L" is used interchangeably and signifies the affinity between ligand "L" and protein "P" (i.e., how strongly the ligand binds to a particular protein). It is represented by formula K dIt can be calculated using the formula =[L][P] / [LP], where [P], [L], and [LP] represent the molar concentrations of the protein, ligand, and complex, respectively.
[0039] This disclosure provides compositions and methods useful for editing target nucleic acid sequences. As used herein, “editing” is used interchangeably with “modification” and includes, but is not limited to, cleavage, nicking, deletion, knock-in, knock-out, and the like.
[0040] The term "knockout" refers to the elimination of a gene or its expression. For example, a gene may be knocked out by either the deletion or addition of a nucleotide sequence that results in the disruption of the leading frame. Alternatively, a gene may be knocked out by replacing a portion of the gene with an unrelated sequence. As used herein, the term "knockdown" refers to the reduction of the expression of a gene or its gene product. As a result of gene knockdown, protein activity or function may be reduced, or protein levels may be reduced or eliminated.
[0041] As used herein, “homology-directed repair” (HDR) refers to a form of DNA repair that occurs during the repair of double-strand breaks in cells. This process requires homology of nucleotide sequences and uses a donor template to repair or knock out target DNA, resulting in the transfer of genetic information from the donor (e.g., a donor template) to the target. Homologous recombination repair can result in alteration of the target nucleic acid sequence by insertion, deletion, or mutation if the donor template differs from the target DNA sequence and some or all of the donor template sequence is incorporated into the target DNA at the correct genomic locus.
[0042] As used herein, “non-homologous end joining” (NHEJ) refers to the repair of double-strand breaks in DNA by directly ligating the broken ends together without requiring a homologous template (as opposed to homologous recombination repair, which requires homologous sequences to induce repair). NHEJ often results in indels (loss (deletion) or insertion of nucleotide sequences near the double-strand break site).
[0043] As used herein, "microhomology-mediated end joining" (MMEJ) refers to a mutagenic double-strand break (DSB) repair mechanism that does not require a homologous template (as opposed to homologous recombination repair, which requires homologous sequences to induce repair) and is always associated with a deletion adjacent to the break site. MMEJs often result in the loss (deletion) of nucleotide sequences near the double-strand break site.
[0044] A polynucleotide or polypeptide (or protein) has a certain percentage of "sequence similarity" or "sequence identity" with another polynucleotide or polypeptide, which means the percentage of the same bases or amino acids at the same relative positions when the two sequences are compared (aligned). Sequence similarity (also called percentage similarity, percentage identity, or homology) can be determined in several different ways. To determine sequence similarity, sequences can be aligned using methods and computer programs known in the art, including BLAST, which is available on the World Wide Web at ncbi.nlm.nih.gov / BLAST. Percentage complementarity between specific stretches of nucleic acid sequences within a nucleic acid can be determined using any convenient method. Examples of methods include the BLAST program (a basic local sorting search tool) and the PowerBLAST program (Altschul et al., J.Mol.Biol., 1990, 215, 403-410; Zhang and Madden, Genome Res., 1997, 7, 649-656) (or by using the Gap program (Wisconsin Sequence Analysis Package, Version 8 for Unix, Genetics Computer Group, University Research Park, Madison Wis.), for example, using the default settings that employ the Smith and Waterman algorithm (Adv.Appl.Math., 1981, 2, 482-489)).
[0045] The terms “polypeptide” and “protein” are used interchangeably herein and refer to polymeric forms of amino acids of any length, which may include coding and non-coding amino acids, chemically or biochemically modified or derivatized amino acids, and polypeptides having a modified peptide backbone. The term also includes fusion proteins (including, but not limited to, fusion proteins having heterologous amino acid sequences).
[0046] A "vector" or "expression vector" is a replicon such as a plasmid, phage, virus, virus-like particle, or cosmid, to which another DNA segment (i.e., an "insert") can be attached, resulting in the replication or expression of the attached segment in a cell.
[0047] As used herein, the terms “naturally occurring,” “unmodified,” or “wild-type” refer to nucleic acids, polypeptides, cells, or organisms that are found in nature, when applied to nucleic acids, polypeptides, cells, or organisms.
[0048] As used herein, “mutation” means the insertion, deletion, substitution, duplication, or inversion of one or more amino acids or nucleotides compared to the wild-type or reference amino acid sequence, or compared to the wild-type or reference nucleotide sequence.
[0049] As used herein, the term “isolated” means a polynucleotide, polypeptide, or cell in an environment different from the environment in which the cell naturally exists. Isolated genetically modified host cells may be present in a mixed population of genetically modified host cells.
[0050] As used herein, “host cell” means a eukaryotic cell, a prokaryotic cell, or a cell derived from a multicellular organism cultured as a single cell entity (e.g., a cell line), where the eukaryotic or prokaryotic cell is used as a recipient of nucleic acid (e.g., an expression vector) and includes offspring of the original cell that has been genetically modified by the nucleic acid. It is understood that single-cell offspring may not necessarily be completely identical to the original parent in morphology or in genomic or total DNA complementation due to natural, accidental, or intentional mutations. A “recombinant host cell” (also referred to as a “genetically modified host cell”) is a host cell into which a different nucleic acid (e.g., an expression vector) has been introduced.
[0051] The term "conservative amino acid substitution" refers to the interchangeability of amino acid residues with similar side chains in proteins. For example, the group of amino acids with aliphatic side chains consists of glycine, alanine, valine, leucine, and isoleucine; the group of amino acids with aliphatic hydroxyl side chains consists of serine and threonine; the group of amino acids with amide-containing side chains consists of asparagine and glutamine; the group of amino acids with aromatic side chains consists of phenylalanine, tyrosine, and tryptophan; the group of amino acids with basic side chains consists of lysine, arginine, and histidine; and the group of amino acids with sulfur-containing side chains consists of cysteine and methionine. Exemplary conservative amino acid substitutions include valine-leucine-isoleucine, phenylalanine-tyrosine, lysine-arginine, alanine-valine, and asparagine-glutamine.
[0052] As used herein, “treatment” or “to treat” is used interchangeably herein and refers to an approach to obtain beneficial or desired outcomes, including but not limited to therapeutic and / or preventive benefits. Therapeutic benefits mean the eradication or improvement of the underlying disorder or disease being treated. Therapeutic benefits can also be achieved by the eradication or improvement of one or more symptoms associated with the underlying disease, or by improvement of one or more clinical parameters, such that improvement is observed in the subject, even though the subject may still be suffering from the underlying disease.
[0053] The terms “therapeutic dose” and “therapeutic dosage,” as used herein, refer to the amount of a drug or biological product, either alone or as part of a composition, that, when administered to a subject (e.g., a human or an experimental animal) in a single dose or repeated dose, can have any detectable beneficial effect on any symptom, aspect, measured parameter, or characteristic of a disease state or condition. Such effect does not need to be absolute to be beneficial.
[0054] As used herein, “administer” means a method of giving a subject a dose of the composition of this disclosure.
[0055] As used herein, “subject” means mammal. Mammals include, but are not limited to, livestock, primates, non-human primates, humans, dogs, pigs (porcine), rabbits, mice, rats, and other rodents.
[0056] All publications, patents, and patent applications referenced herein are incorporated herein by reference as if each individual publication, patent, or patent application were specifically and individually indicated to be invoked by reference.
[0057] I. General Methods The implementation of this invention will, unless otherwise indicated, utilize conventional techniques of immunology, biochemistry, chemistry, molecular biology, microbiology, cell biology, genomics, and recombinant DNA, as well as standard textbooks such as Molecular Cloning: A Laboratory Manual, 3rd Ed. (Sambrook et al., Harbor Laboratory Press 2001), Short Protocols in Molecular Biology, 4th Ed. (Ausubel et al. eds., John Wiley & Sons 1999), Protein Methods (Bollag et al., John Wiley & Sons 1996), Nonviral Vectors for Gene Therapy (Wagner et al. eds., Academic Press 1999), Viral Vectors (Kaplift & Loewy eds., Academic Press 1995), Immunology Methods Manual (I. Lefkovits ed., Academic Press 1997), and Cell and Tissue This can be found in Culture: Laboratory Procedures in Biotechnology (Doyle & Griffiths, John Wiley & Sons 1998) (their disclosures are incorporated herein by reference).
[0058] Where a range of values is provided, it is understood that the endpoint is included, and that each intermediary value between the upper and lower limits of that range encompasses any other listed or intermediary values within that range, up to one-tenth of the lower limit unit, unless the context otherwise explicitly indicates otherwise. The upper and lower limits of these smaller ranges may independently be included within smaller ranges and are also included according to any specifically excluded limits within the listed range. If a listed range includes one or both limits, it also includes ranges that exclude one or both of those included limits.
[0059] Unless otherwise defined, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art in which the present invention pertains. All publications referenced herein are incorporated herein by reference to disclose and describe the methods and / or materials by which the publications are cited.
[0060] Note that, as used herein and in the appended claims, the singular forms "a," "an," and "the" refer to multiple objects unless the context clearly indicates otherwise.
[0061] For clarity, it will be understood that certain features of the Disclosure described in the context of separate embodiments may be provided in combination in a single embodiment. In other cases, for brevity, various features of the Disclosure described in the context of a single embodiment may be provided separately or in any preferred partial combination. All combinations of embodiments relating to the Disclosure are specifically encompassed by the Disclosure and are intended to be disclosed herein as if every possible combination were disclosed individually and expressly. In addition, all partial combinations of various embodiments and their elements are also specifically encompassed by the Disclosure and are disclosed herein as if every possible partial combination were disclosed herein individually and expressly.
[0062] II. A system for gene editing of the BCL11A gene In a first aspect, the Disclosure provides a system comprising a class 2, type V CRISPR nuclease protein and one or more guide nucleic acids (gRNAs) for use in modifying or editing the BCL11A gene in order to reduce or eliminate the expression of the BCL11A gene product. An example of a class 2, type V CRISPR nuclease protein and guide nucleic acid system is the CasX:gRNA system, which is specifically designed to modify the BCL11A gene in eukaryotic cells. In some cases, the CasX:gRNA system is designed to knock down or knock out the BCL11A gene. Generally, any portion of the BCL11A gene can be targeted using programmable compositions and methods provided herein. In some embodiments, the BCL11A gene to be modified is a wild-type sequence, and the modified portion is selected from the group consisting of BCL11A introns, BCL11A exons, BCL11A intron-exon junctions, BCL11A regulatory elements, and intergenetic regions, or the modification is the deletion or mutation of one or more exons.
[0063] As used herein, “Systems,” including the CRISPR nuclease protein and one or more gRNAs of the Disclosure, the nucleic acids encoding the CRISPR nuclease protein and gRNAs of the Disclosure, and systems comprising the nucleic acids or vectors comprising the CRISPR nuclease protein and one or more gRNAs of the Disclosure, are used interchangeably with the term “Composition.”
[0064] The human BCL11A gene (HGNC:13221) encodes a protein (Q9H165) with the following sequence: (Sequence ID 100). The BCL11A gene is defined as the sequence spanning chr2 60450520-60554467 (GRCh38 / hg38 Ensembl 100) of the human genome on chromosome 2.
[0065] In some embodiments, this disclosure provides systems specifically designed to modify the BCL11A gene in eukaryotic cells, either in vitro, ex vivo, or in vivo. Generally, any portion of the BCL11A target nucleic acid can be targeted using programmable compositions and methods provided herein. In some embodiments, the CRISPR nuclease is a class 2, type V nuclease. While members of the class 2 type V CRISPR-Cas system have differences, they share several common features that distinguish them from the Cas9 system. First, type V nucleases have a single RNA-inducible effector containing a RuvC domain but lacking an HNH domain, and recognize a T-rich PAM 5' upstream of the target region on the non-target strand. This differs from the Cas9 system, which relies on a G-rich PAM at the 3' side of the target sequence. Unlike Cas9, which produces a blunt end proximal to the PAM, type V nucleases produce a double-strand break distal to the PAM sequence. In addition, when activated by target dsDNA or cis-bound ssDNA, the V-type nuclease degrades the ssDNA in trans. In some embodiments, the Disclosure provides a class 2, V-type nuclease selected from the group consisting of Cas12a, Cas12b, Cas12c, Cas12d (CasY), Cas12j, Cas12k, CasZ, and CasX. In some embodiments, the Disclosure provides a CasX:gRNA system comprising one or more CasX proteins and one or more guide nucleic acids (gRNAs). In other embodiments, the CasX:gRNA system of the Disclosure comprises one or more CasX proteins, one or more guide nucleic acids (gRNAs), and one or more donor template nucleic acids comprising nucleic acids encoding a portion of the BCL11A gene, wherein the donor template nucleic acid comprises one or more nucleotide deletions, insertions, or mutations compared to the genomic nucleic acid sequence encoding the BCL11A protein. Each of these components and their use in editing the BCL11A gene are described herein below.
[0066] In some embodiments, the Disclosure provides a gene editing pair of CasX and gRNA from any of the embodiments described herein, which can be bound together before use for gene editing and are therefore “pre-complexed” as a ribonucleoprotein complex (RNP). The use of pre-complexed RNPs provides advantages in the delivery of system components to cells or target nucleic acid sequences for editing of target nucleic acid sequences.
[0067] In some embodiments, functional RNPs can be delivered to cells ex vivo by electrophoresis or chemical means. In other embodiments, functional RNPs can be delivered either ex vivo or in vivo by vectors in their functional form. In some embodiments, RNPs can be delivered in vivo to a target using CasX delivery particles (XDPs). The gRNA can provide target specificity to the complex by including a targeting sequence (or "spacer") having a nucleotide sequence complementary to the sequence of the target nucleic acid sequence. On the other hand, the CasX variant protein of the pre-complexed CasX:gRNA provides site-specific activity such as cleavage or nicking of the target sequence, which is guided to a target site in the target nucleic acid sequence (e.g., stabilized at the target site) by its association with the gRNA.
[0068] The system is useful in the treatment of subjects with abnormal hemoglobin disorders (e.g., sickle cell anemia or β-thalassemia). The components of the CasX:gRNA system, their functions, and their use in editing target nucleic acids in cells are described in more detail below.
[0069] III. Guide to Systems for Gene Editing: Nucleic Acids In another embodiment, the disclosure relates to a specially designed guide ribonucleic acid (gRNA) which comprises a targeting sequence complementary to (and therefore capable of hybridizing with) the target nucleic acid sequence of the BCL11A gene, and which, when complexed with a CRISPR nuclease, is useful in genome editing of the BCL11A target nucleic acid in cells. In some embodiments, it is envisioned that multiple gRNAs are delivered in the system for modification of the target nucleic acid. For example, when each gRNA complexes with a CRISPR nuclease to bind and cleave at two different or overlapping sites within a gene, a pair of gRNAs having targeting sequences for different or overlapping regions of the target nucleic acid sequence can be used, and the gene is then edited by non-homologous end joining (NHEJ), homologous recombination repair (HDR), homology-independent targeted integration (HITI), micro-homology-mediated end joining (MMEJ), single-strand annealing (SSA), or base excision repair (BER).
[0070] In some embodiments, this disclosure provides gRNAs for use in a CasX:gRNA system, which are useful in genome editing of the BCL11A gene in eukaryotic cells. In certain embodiments, the gRNAs of the system can form complexes with CasX nucleases. This disclosure provides specially designed gRNAs, the targeting sequences (or spacers, described in more detail below) of the gRNAs, which are complementary to (and therefore hybridize with) the target nucleic acid sequence when used as components of a gene-editing CasX:gRNA system. Table 1 presents representative but non-limiting examples of sequence numbers for targeting sequences for the BCL11A target nucleic acid that can be used with the gRNAs of this embodiment, described in more detail below.
[0071] a. Reference gRNA and gRNA variants As used herein, “reference gRNA” refers to a CRISPR guide nucleic acid containing a wild-type sequence of a naturally occurring gRNA. In some embodiments, the reference gRNA of this disclosure may be subjected to one or more mutagenesis methods (e.g., deep mutational evolution (DME), deep mutational scanning (DMS), error-prone PCR, cassette mutagenesis, random mutagenesis, adherend extension PCR, gene shuffling, or domain swapping, as described herein) to generate one or more gRNA variants having enhanced or different properties compared to the reference gRNA. The gRNA variants also include variants containing one or more exogenous sequences, for example, fused to or inserted into either the 5' or 3' end. The activity of the reference gRNA may be used as a benchmark against which the activity of the gRNA variants is compared, thereby measuring improvements in the function or other characteristics of the gRNA variants. In other embodiments, the reference gRNA may be subjected to one or more intentionally targeted mutations to produce a gRNA variant (e.g., a rationally designed variant).
[0072] The gRNAs of this disclosure comprise two segments: a targeting sequence and a protein-binding segment. The targeting segment of the gRNA comprises a nucleotide sequence (replaced, referred to interchangeably, as a guide sequence, spacer, targeter, or targeting sequence) that is complementary to (and therefore hybridizes with) a specific sequence (target site) within a target nucleic acid sequence (e.g., a target ssRNA, target ssDNA, or strand of a double-stranded target DNA) as described in more detail below. The targeting sequence of the gRNA can bind to the target nucleic acid sequence (including coding sequences, complements to coding sequences, and non-coding sequences) and regulatory elements. The protein-binding segment (or “activator” or “protein-binding sequence”) interacts with (e.g., binds to) the CasX protein as a complex forming an RNP (as described in more detail below). Alternatively, the protein-binding segment is referred to herein as a “scaffold,” which consists of several regions (as described in more detail below).
[0073] In the case of dual guide RNA (dgRNA), the targeter and activator portions each have a double-stranding segment, and the double-stranding segments of the targeter and activator are complementary to each other and hybridize to form a double-stranded double helix (dsRNA double helix for gRNA). When gRNA is gRNA, the terms “targeter” or “targeter RNA” are used herein to refer to the crRNA-like molecule (crRNA: “CRISPR RNA”) of CasX dual guide RNA (and therefore, of CasX single guide RNA when the “activator” and “targeter” are linked together, for example, by intervening nucleotides). crRNA has a 5' region that anneals with tracrRNA, followed by nucleotides of the targeting sequence. Thus, for example, guide RNA (dgRNA or sgRNA) includes a guide sequence and the double-stranding segment of crRNA (which may also be called a crRNA repeat). The corresponding tracrRNA-like molecule (activator) also includes a double-stranding stretch of nucleotides that forms the other half of the dsRNA double helix of the protein-binding segment of the guide RNA. Thus, the targeter and activator hybridize as a corresponding pair to form a dual guide NA (hereinafter referred to herein as “dual guide NA”, “two-molecule gRNA”, “dgRNA”, “two-molecule guide NA”, or “two-molecule guide NA”). Site-specific binding and / or cleavage of the target nucleic acid sequence (e.g., genomic DNA) by the CasX protein can occur at one or more locations (e.g., the sequence of the target nucleic acid) determined by the complementarity of base pairing between the targeting sequence of the gRNA and the target nucleic acid sequence. Thus, for example, the gRNA of this disclosure has a sequence complementary to the target nucleic acid adjacent to a sequence complementary to the TC PAM motif (or PAM sequence such as ATC, CTC, GTC, or TTC), and is therefore capable of hybridizing with the target nucleic acid. Since the targeting sequence of the guide sequence hybridizes with the sequence of the target nucleic acid sequence, the targeter can be modified by the user to hybridize with a specific target nucleic acid sequence, as long as the position of the PAM sequence is taken into consideration.Therefore, in some cases, the targeter sequence may be a sequence that does not exist in nature. In other cases, the targeter sequence may be a sequence that exists in nature derived from the gene being edited. In other embodiments, the activator and targeter of the gRNA are covalently bonded to each other (rather than hybridizing with each other) and comprise a single molecule (referred herein to as “single-molecule gRNA”, “single-molecule guide NA”, “single-guide NA”, “single-guide RNA”, “single-molecule guide RNA”, or “sgRNA”). In some embodiments, the sgRNA comprises an “activator” or a “targeter” and may therefore be “activator RNA” and “targeter RNA,” respectively. In some embodiments, the gRNA is a ribonucleic acid molecule ("gRNA"), and in other embodiments, the gRNA is a chimera and comprises both DNA and RNA. As used herein, the term gRNA encompasses naturally occurring molecules as well as sequence variants.
[0074] In summary, the constructed gRNA of the present disclosure comprises four distinct regions or domains: an RNA triple helix, a scaffold stem, an elongation stem, and a targeting sequence (specific to the target nucleic acid in the embodiments of the present disclosure and located at the 3' end of the gRNA). The RNA triple helix, scaffold stem, and elongation stem together are referred to as the “scaffold” of the gRNA.
[0075] b.RNA triplex In some embodiments of the guide NAs (including reference sgRNAs) provided herein, an RNA triple helix is present, and the RNA triple helix contains a UUU-nX(approximately 4-15)-UUU(SEQ ID NO: 226) stem-loop sequence ending in AAAG after two intervening stem-loops (a scaffold stem-loop and an elongation stem-loop), forming a pseudoknot that may extend across the triple helix to a double-stranded pseudoknot. The UU-UUU-AAA sequence of the triple helix is formed as a nexus between the targeting sequence, the scaffold stem, and the elongation stem. In the exemplary CasX sgRNA, the UUU-loop-UUU region is encoded first, followed by the scaffold stem-loop, then the elongation stem-loop linked by a tetraloop, then AAAG closes the triple helix before the targeting sequence.
[0076] c. Scaffolding stem loop In some embodiments of the sgRNAs of this disclosure, a triple-stranded region is followed by a scaffolding stem-loop. The scaffolding stem-loop is the region of the gRNA to which the CasX protein (e.g., reference or CasX variant protein) binds. In some embodiments, the scaffolding stem-loop is a fairly short and stable stem-loop. In some cases, the scaffolding stem-loop does not tolerate much variation and requires some form of RNA bubble. In some embodiments, the scaffolding stem is required for the function of the CasX sgRNA. This is perhaps analogous to the nexus stem of Cas9, which is a critical stem-loop, but the scaffolding stem of the CasX sgRNA has a required bulge (RNA bubble) in some embodiments, which differs from many other stem-loops found in the CRISPR / Cas system. In some embodiments, the presence of this bulge is conserved across sgRNAs interacting with different CasX proteins. An exemplary sequence of the scaffolding stem-loop sequence of the gRNA includes the sequence CCAGCGACUAUGUCGUAUGG (SEQ ID NO: 20).
[0077] d. Elongated stem loop In some embodiments of the CasX sgRNA of this disclosure, an elongation stem loop follows a scaffolding stem loop. In some embodiments, the elongation stem mainly comprises a synthetic tracrRNA and crRNA fusion that does not have the CasX protein bound to it. In some embodiments, the elongation stem loop may be highly malleable. In some embodiments, a single guide gRNA is constructed using a GAAA tetraloop linker or GAGAAA linker between the tracrRNA and crRNA in the elongation stem loop. In some cases, the targeter and activator of the CasX sgRNA are linked to each other by intervening nucleotides, and the linker may have a length of 3 to 20 nucleotides. In some embodiments of the CasX sgRNA of this disclosure, the elongation stem is a large 32 bp loop located outside the CasX protein in the ribonucleoprotein complex. An exemplary sequence of the sgRNA elongation stem loop sequence includes the sequence GCGCUUAUUUAUCGGAGAGAAAUCCGAUAAAUAAGAAGC (SEQ ID NO: 21). In some embodiments, the elongated stem loop includes a GAGAAA spacer array.
[0078] e. Targeted sequences In some embodiments of the gRNAs of this disclosure, the elongated stem-loop is followed by a region that forms part of a triple helix, and then a targeting sequence (or "spacer") at the 3' end of the gRNA. The targeting sequence targets the CasX ribonucleoprotein holocomplex to a specific region of the target nucleic acid sequence of the gene to be modified. Thus, for example, the gRNA targeting sequence of this disclosure has, as a component of the RNP, a sequence complementary to (and therefore hybridizable to) a portion of the BCL11A gene in eukaryotic nucleic acids (e.g., eukaryotic chromosomes, chromosome sequences, eukaryotic RNA, etc.) when one of the TC PAM motif or PAM sequences TTC, ATC, GTC, or CTC is located 1 nucleotide 5' to the non-target strand sequence complementary to the target sequence. The gRNA targeting sequence can be modified so that the gRNA can target a desired sequence of any desired target nucleic acid sequence, insofar as the position of the PAM sequence is taken into consideration. In some embodiments, the gRNA scaffold is located at the 5' end of the target sequence, and the targeting sequence is located at the 3' end of the gRNA. In some embodiments, the PAM motif sequence recognized by the RNP nuclease is TC. In other embodiments, the PAM sequence recognized by the RNP nuclease is NTC.
[0079] In some embodiments, the gRNA targeting sequence is specific to a portion of the gene encoding the BCL11A protein. In some embodiments, the gRNA targeting sequence is specific to an exon of BCL11A. In some embodiments, the gRNA targeting sequence is specific to an intron of BCL11A. In some embodiments, the gRNA targeting sequence is specific to an intron-exon junction of BCL11A. In some embodiments, the gRNA targeting sequence has a sequence that hybridizes with a regulatory element of BCL11A, a coding region of BCL11A, a non-coding region of BCL11A, or a combination thereof (e.g., the intersection of two regions). In some embodiments, the regulatory element includes a GATA-binding sequence. In some embodiments, the gRNA targeting sequence is complementary to a sequence containing one or more single nucleotide polymorphisms (SNPs) of the BCL11A gene or its complement. SNPs within the coding sequence or non-coding sequence of BCL11A are both within the scope of this disclosure. In other embodiments, the gRNA targeting sequence is complementary to the intergenetic region of the BCL11A gene, or is a sequence complementary to the intergenetic region of the BCL11A gene.
[0080] In some embodiments, the gRNA targeting sequence is designed to be specific to regulatory elements that regulate the expression of the BCL11A gene product. Such regulatory elements include, but are not limited to, regions containing promoter regions, enhancer regions, intergeneric regions, 5' untranslated region (5'UTR), 3' untranslated region (3'UTR), conserved elements, and cis-regulatory elements. The promoter region is intended to contain nucleotides within 5 kb of the coding sequence start point, or, in the case of gene enhancer elements or conserved elements, may be thousands of base pairs (bp), hundreds of thousands of bp, or millions of bp away from the coding sequence of the target nucleic acid gene. In certain embodiments, the gRNA targeting sequence hybridizes with a sequence complementary to the regulatory element of BCL11A. In one embodiment, the targeting sequence of the gRNA is UGGAGCCUGUGAUAAAAGCA (SEQ ID NO: 22), which hybridizes with the BCL11A GATA1 erythroid-specific enhancer binding site sequence or has at least 90% or at least 95% sequence identity with it (see Figure 23). In another embodiment, the targeting sequence of the gRNA is UGCUUUUAUCACAGGCUCCA (SEQ ID NO: 23), or has at least 90% or at least 95% sequence identity with it. In yet another embodiment, the targeting sequence of the gRNA is UGCUUUUAUCACAGGCUCCA (SEQ ID NO: 23), or has at least 90% or at least 95% sequence identity with it. In other embodiments, the gRNA targeting sequence is selected from the group consisting of CAGGCUCCAGGAAGGGUUUG (SEQ ID NO: 2949), GAGGCCAAACCCUUCCUGGA (SEQ ID NO: 2948), AGUGCAAGCUAACAGUUGCU (SEQ ID NO: 15747), and AUACAACUUUGAAGCUAGUC (SEQ ID NO: 15748).
[0081] In mature subjects after birth, GATA1 binding enhances BCL11A expression, which in turn suppresses hemoglobin F (HbF) expression in favor of hemoglobin gamma. However, it has been demonstrated that suppressing BCL11A expression in subjects with certain hemoglobin disorders resumes HbF expression, thus compensating for otherwise defective hemoglobin.
[0082] In some embodiments, the gRNA targeting sequence has 14 to 35 consecutive nucleotides. In some embodiments, the targeting sequence has 14, 15, 16, 18, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, or 35 consecutive nucleotides. In some embodiments, the targeting sequence consists of 20 consecutive nucleotides. In some embodiments, the targeting sequence consists of 19 consecutive nucleotides. In some embodiments, the targeting sequence consists of 18 consecutive nucleotides. In some embodiments, the targeting sequence consists of 17 consecutive nucleotides. In some embodiments, the targeting sequence consists of 16 consecutive nucleotides. In some embodiments, the targeting sequence consists of 15 consecutive nucleotides. In some embodiments, the targeting sequence has 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, or 35 consecutive nucleotides, and the targeting sequence may contain 0-5, 0-4, 0-3, or 0-2 mismatches with respect to the target nucleic acid sequence, and can maintain sufficient binding specificity so that an RNP containing the gRNA containing the targeting sequence can form a complementary bond to the target nucleic acid.
[0083] Representative but non-limiting examples of targeting sequences for target nucleic acid sequences intended for use in gRNAs of this disclosure are presented as SEQ ID NOs: 272-2100 and 2286-26789 (see Table 1). In some embodiments, this disclosure provides targeting sequences for ATC PAMs that are at least 50% identical, at least 55% identical, at least 60% identical, at least 65% identical, at least 70% identical, at least 75% identical, at least 80% identical, at least 85% identical, at least 90% identical, at least 95% identical, or 100% identical to sequences of SEQ ID NOs: 272-2100 or 2286-5625. In some embodiments, the Disclosure provides targeted sequences for CTC PAMs that are at least 50% identical, at least 55% identical, at least 60% identical, at least 65% identical, at least 70% identical, at least 75% identical, at least 80% identical, at least 85% identical, at least 90% identical, at least 95% identical, or 100% identical to sequences of sequence numbers 5626 to 13616. In some embodiments, the Disclosure provides targeted sequences for CTC PAMs that are at least 50% identical, at least 55% identical, at least 60% identical, at least 65% identical, at least 70% identical, at least 75% identical, at least 80% identical, at least 85% identical, at least 90% identical, at least 95% identical, or 100% identical to sequences of sequence numbers 13617 to 17903. In some embodiments, the present disclosure provides targeted sequences of GTC PAMs, including sequences 13617-17903.In some embodiments, the Disclosure provides targeting sequences for TTC PAMs that are at least 50% identical, at least 55% identical, at least 60% identical, at least 65% identical, at least 70% identical, at least 75% identical, at least 80% identical, at least 85% identical, at least 90% identical, at least 95% identical, or 100% identical to sequences of sequence numbers 17904-26789. In some embodiments, the Disclosure provides targeting sequences for TTC PAMs that include sequences of sequence numbers 17904-26789. In some embodiments, the targeting sequences intended for use in gRNAs of the Disclosure include sequences of sequence numbers 272-2100 or 2286-26789 with a single nucleotide removed from the 3' end of the sequence. In other embodiments, the targeting sequences for gRNAs include sequences of sequence numbers 272-2100 or 2286-26789 with two nucleotides removed from the 3' end of the sequence. In other embodiments, the gRNA targeting sequence includes sequences of sequence numbers 272-2100 or 2286-26789 with three nucleotides removed from the 3' end of the sequence. In other embodiments, the gRNA targeting sequence includes sequences of sequence numbers 272-2100 or 2286-26789 with four nucleotides removed from the 3' end of the sequence. In other embodiments, the gRNA targeting sequence includes sequences of sequence numbers 272-2100 or 2286-26789 with five nucleotides removed from the 3' end of the sequence. In the embodiments described above, thymine (T) nucleotides may substitute one or more or all uracil (U) nucleotides in any of the targeting sequences so that the gRNA targeting sequence may be gDNA or gRNA, or a chimera of RNA and DNA, or when a spacer coding sequence is incorporated into an expression vector. In some embodiments, the targeting sequences of SEQ ID NOs. 272-2100 or 2286-26789 have at least one, two, three, four, five, or six or more thymine nucleotides that are substituted with uracil nucleotides.
[0084] [Table 1]
[0085] In some embodiments, the CasX:gRNA system comprises a first gRNA and further comprises a second (and optionally, a third, fourth, fifth, or more) gRNA, the second or additional gRNA having a targeting sequence complementary to a different or overlapping portion of the target nucleic acid sequence compared to the targeting sequence of the first gRNA, resulting in the targeting of multiple points in the target nucleic acid, for example, by CasX introducing multiple cleavage in the target nucleic acid. In such cases, it will be understood that the second or additional gRNA is complexed with an additional copy of the CasX protein. By selecting the targeting sequence of the gRNA, a defined region of the target nucleic acid sequence surrounding a specific location within the target nucleic acid can be modified or edited using the CasX:gRNA system described herein, including facilitating the insertion of a donor template containing a mutation in the BCL11A gene. In certain embodiments, the second gRNA may contain a targeting sequence complementary to a sequence adjacent to the GATA1 binding site on the 5' or 3' side, such that the GATA1 binding site is disrupted.
[0086] f.gRNA scaffold Apart from the targeting sequence domain, the remaining components of the gRNA are referred to herein as the scaffold. In some embodiments, the gRNA scaffold is derived from a naturally occurring sequence described below as the reference gRNA. In other embodiments, the gRNA scaffold is a variant of the reference gRNA, which is modified by mutation, insertion, deletion, or domain substitution to confer desirable or improved properties to the gRNA.
[0087] With respect to a polynucleotide or amino acid sequence, the term “adjacent to” refers to sequences that are adjacent to or adjacent to each other in a polynucleotide or polypeptide. Those skilled in the art will understand that two sequences can be considered adjacent to each other and still contain a limited number of intervening sequences, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides or amino acids.
[0088] Table 2 provides reference gRNA tracr sequences and scaffold sequences. In some embodiments, the disclosure provides gRNA sequences having a scaffold comprising sequences of sequence numbers 4-16 shown in Table 2, or a reference gRNA sequence having at least one nucleotide modification compared to any one of sequence numbers 4-16 in Table 2. In embodiments where the vector comprises DNA encoding the gRNA sequence or the gRNA is an RNA-DNA chimera, it will be understood that thymine (T) bases may be substituted with uracil (U) bases of any of the embodiments of the gRNA sequences described herein.
[0089] [Table 2]
[0090] g.gRNA variant In another aspect, the disclosure relates to guide nucleic acid variants (hereinafter substituted for “gRNA variant” or “gRNA variant”) that include one or more modifications compared to a reference gRNA scaffold. As used herein, “scaffold” means all portions of gRNA necessary for the function of the gRNA, excluding the targeting sequence.
[0091] In some embodiments, the gRNA variant comprises one or more nucleotide substitutions, insertions, deletions, or swapped or substituted regions compared to the reference gRNA sequence of this disclosure. In some embodiments, the mutation may occur in any region of the reference gRNA scaffold to generate the gRNA variant. In some embodiments, the scaffold of the gRNA variant sequence has at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, or at least 70%, at least 80%, at least 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity with the sequence of SEQ ID NO: 4 or SEQ ID NO: 5.
[0092] In some embodiments, the gRNA variant includes one or more nucleotide changes within one or more regions of the reference gRNA that improve the features compared to the reference gRNA. Exemplary regions include RNA triple helix, pseudoknots, scaffolding stem-loops, and elongation stem-loops. In some cases, the variant scaffolding stem further includes bubbles. In other cases, the variant scaffolding further includes triple-helix loop regions. In yet another case, the variant scaffolding further includes a 5' unstructured region. In one embodiment, the gRNA variant scaffolding includes a scaffolding stem-loop having at least 60% sequence identity with SEQ ID NO: 14. In another embodiment, the gRNA variant includes a scaffolding stem-loop having the sequence CCAGCGACUAUGUCGUAGUGG (SEQ ID NO: 25). In another embodiment, the present disclosure provides a modified elongated stem-loop with C18G substitution, G55 insertion, U1 deletion, and altered sequence number 5 (sequence number 52 nucleotides total) of sequence number 5 (the original 6nt loop and the most-loop-proximal 13 base pairs (total 32 nucleotides) to the loop are substituted with a Uvsx hairpin (4nt loop and the most-loop-proximal 5 base pairs (total 14 nucleotides) to the loop), and the loop-distal bases of the elongated stem are converted to a fully base-paired stem adjacent to the new Uvsx hairpin by A99 deletion and G64U substitution). In the aforementioned embodiment, the gRNA scaffold comprises the sequence ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGUGGGUAAAGCUCCCUCUUCGGAGGGAGCAUCAAAG).
[0093] All gRNA variants that have one or more improved functions or features, or that add one or more new functions when the variant gRNA is compared to the reference gRNA described herein, are assumed to be within the scope of this disclosure. A representative example of such a gRNA variant is Guide 174 (SEQ ID NO: 2238), whose design (and rationale for its design) is described in the examples. In some embodiments, the gRNA variant adds a new function to the RNP containing the gRNA variant. In some embodiments, the gRNA variant has improved features selected from: improved stability, improved solubility, improved gRNA transcription, improved resistance to nuclease activity, increased gRNA folding rate, reduced byproduct formation during folding, increased productive folding, improved binding affinity to CasX protein, improved binding affinity to target DNA when complexed with CasX protein, improved gene editing when complexed with CasX protein, improved editing specificity when complexed with CasX protein, and improved ability to utilize one or more broadly categorized PAM sequences, including ATC, CTC, GTC, or TTC, in editing target DNA when complexed with CasX protein, or any combination thereof. In some cases, one or more of the improved features of the gRNA variant are improved by at least about 1.1 to about 100,000 times compared to the reference gRNA of SEQ ID NO: 4 or SEQ ID NO: 5. In other cases, one or more improved features of the gRNA variant are improved by at least approximately 1.1 times, at least approximately 10 times, at least approximately 100 times, at least approximately 1,000 times, at least approximately 10,000 times, at least approximately 100,000 times, or more compared to the reference gRNA of SEQ ID NO: 4 or SEQ ID NO: 5.In other cases, one or more of the improved features of the gRNA variant are approximately 1.1–100,000 times, 1.1–10,000 times, 1.1–1,000 times, 1.1–500 times, 1.1–100 times, 1.1–50 times, 1.1–20 times, 10–100,000 times, 10–10,000 times, 10–1,000 times, 10–500 times, 10–100 times, 10–50 times, 10–20 times, 2–70 times, 2–50 times, 2–30 times, 2–20 times, 2–10 times, and 5–50 times compared to the reference gRNA of SEQ ID NO: 4 or 5. The improvements are approximately 5-30 times, 5-10 times, 100-100.00 times, 100-10.00 times, 100-1,000 times, 100-500 times, 500-100.00 times, 500-10.00 times, 500-1,000 times, 500-750 times, 1,000-100.00 times, 10,000-100.00 times, 20-500 times, 20-250 times, 20-200 times, 20-100 times, 20-50 times, 50-10,000 times, 50-1,000 times, 50-500 times, 50-200 times, or 50-100 times. In other cases, one or more improved features of the gRNA variant are approximately 1.1 times, 1.2 times, 1.3 times, 1.4 times, 1.5 times, 1.6 times, 1.7 times, 1.8 times, 1.9 times, 2 times, 3 times, 4 times, 5 times, 6 times, 7 times, 8 times, 9 times, 10 times, 11 times, 12 times, 13 times, 14 times, 15 times, 16 times, 17 times, 18 times, 19 times, 20 times, 25 times, 30 times, 40 times, 45 times, 50 times, 55 times, 60 times compared to the reference gRNA of SEQ ID NO: 4 or 5. It is improved by 70 times, 80 times, 90 times, 100 times, 110 times, 120 times, 130 times, 140 times, 150 times, 160 times, 170 times, 180 times, 190 times, 200 times, 210 times, 220 times, 230 times, 240 times, 250 times, 260 times, 270 times, 280 times, 290 times, 300 times, 310 times, 320 times, 330 times, 340 times, 350 times, 360 times, 370 times, 380 times, 390 times, 400 times, 425 times, 450 times, 475 times, or 500 times.
[0094] In some embodiments, gRNA variants can be produced by subjecting a reference gRNA to one or more mutagenesis methods, e.g., mutagenesis methods described herein below, which may include comprehensive mutation evolution (DME), comprehensive mutation scanning (DMS), error-prone PCR, cassette mutagenesis, random mutagenesis, adhesion extension PCR, gene shuffling, or domain swapping. The activity of the reference gRNA may be used as a benchmark against which the activity of the gRNA variant is compared, thereby measuring the improvement in the function of the gRNA variant compared to the reference gRNA. In other embodiments, the reference gRNA may be subjected to one or more intentionally targeted mutations, substitutions, or domain swaps to produce a gRNA variant (e.g., a rationally designed variant). Exemplary gRNA variants produced by such methods are described in the examples, and representative sequences of gRNA scaffolds are presented in Table 3.
[0095] In some embodiments, the gRNA variant comprises one or more modifications compared to a reference guide nucleic acid scaffold sequence, wherein one or more modifications are selected from: at least one nucleotide substitution in a region of the gRNA variant; at least one nucleotide deletion in a region of the gRNA variant; at least one nucleotide insertion in a region of the gRNA variant; substitution of all or part of a region of the gRNA variant; deletion of all or part of a region of the gRNA variant; or any combination thereof. In some cases, the modification is the substitution of 1 to 15 consecutive or non-consecutive nucleotides in one or more regions of the gRNA variant. In other cases, the modification is the deletion of 1 to 10 consecutive or non-consecutive nucleotides in one or more regions of the gRNA variant. In other cases, the modification is the insertion of 1 to 10 consecutive or non-consecutive nucleotides in one or more regions of the gRNA variant. In other cases, the modification is the substitution of an RNA stem-loop sequence from a heterologous RNA source having the proximal 5' and 3' ends of the scaffold stem-loop or elongation stem-loop. In some cases, the gRNA variants of this disclosure include two or more modifications in one region. In other cases, the gRNA variants of this disclosure include modifications in two or more regions. In other cases, the gRNA variants include any combination of the modifications described in this paragraph.
[0096] In some embodiments, a 5'G is added to the gRNA variant sequence for in vivo expression because transcription from the U6 promoter is more efficient and more consistent with respect to the start site when the +1 nucleotide is G. In other embodiments, two 5'Gs are added to the gRNA variant sequence to increase production efficiency in in vitro transcription because T7 polymerase strongly prefers the G at position +1 and the purine at position +2. In some cases, the 5'G base is added to the reference scaffold in Table 2. In other cases, the 5'G base is added to the variant scaffolds of sequence numbers 2238-2285, 26794-26839, and 27219-2726 in Table 3.
[0097] Table 3 provides exemplary gRNA variant scaffold sequences of the present disclosure. In some embodiments, the gRNA variant scaffold includes one of the sequences listed in SEQ ID NOs. 2238-2285, 26794-26839, and 27219-27265 in Table 3, or a sequence having at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, and at least about 99% sequence identity thereto. In some embodiments, the gRNA variant scaffold includes one of SEQ ID NOs. 2238-2285, 26794-26839, and 27219-27265. In some embodiments, the gRNA variant scaffold comprises one of sequence numbers 2281-2285, 26794-26839, and 27219-27265, or a sequence having at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity with it. In some embodiments, the gRNA variant scaffold comprises one of sequence numbers 2281-2285, 26794-26839, and 27219-27265. In embodiments where the vector comprises DNA encoding the gRNA sequence or the gRNA is a chimera of RNA and DNA, it will be understood that thymine (T) bases may be substituted with uracil (U) bases of any of the embodiments of the gRNA sequence described herein.
[0098] [Table 3-1]
[0099] [Table 3-2]
[0100] [Table 3-3]
[0101] Table 3-4
[0102] Table 3-5
[0103] Table 3-6
[0104] Table 3-7
[0105] Table 3-8
[0106] Table 3-9
[0107] Table 3-10
[0108] Table 3-11
[0109] In some embodiments, the sgRNA variant includes one or more additional changes to the sequence of SEQ ID NO: 2238, SEQ ID NO: 2239, SEQ ID NO: 2240, SEQ ID NO: 2241, SEQ ID NO: 2243, SEQ ID NO: 2256, SEQ ID NO: 2274, SEQ ID NO: 2275, SEQ ID NO: 2279, SEQ ID NO: 2281, SEQ ID NO: 2285, SEQ ID NO: 26797, or SEQ ID NO: 26800 in Table 3.
[0110] In some embodiments of the gRNA variants of this disclosure, the gRNA variant comprises at least one modification, the at least one modification compared to the reference guide scaffold of SEQ ID NO: 5, selected from one or more of the following: (a) C18G substitution in the triple-stranded loop, (b) G55 insertion in the stem bubble, (c) U1 deletion, (d) an extension stem-loop modification in which (i) a 6nt loop and 13 base pairs proximal to the loop are substituted by a Uvsx hairpin, and (ii) a deletion of A99 and substitution of G65U, resulting in complete base pairing of the bases distal to the loop. In the exemplary embodiments described above, the gRNA variant comprises any one of the sequences SEQ ID NOs: 2238, 2241, 2244, 2248, 2249, 2256, 2259-2285, 26797, or 26800.
[0111] In some embodiments, the gRNA variant comprises an exogenous stem-loop having a long non-coding RNA (lncRNA). As used herein, lncRNA refers to non-coding RNA longer than approximately 200 bp. In some embodiments, the 5' and 3' ends of the exogenous stem-loop are base-paired (i.e., interact to form a region of double-stranded RNA). In some embodiments, the 5' and 3' ends of the exogenous stem-loop are base-paired, and one or more regions between the 5' and 3' ends of the exogenous stem-loop are not base-paired.
[0112] In some embodiments, the Disclosure provides gRNA variants having nucleotide modifications compared to a reference gRNA, including (a) substitution of 1 to 15 consecutive or non-consecutive nucleotides in one or more regions of the gRNA variant, (b) deletion of 1 to 10 consecutive or non-consecutive nucleotides in one or more regions of the gRNA variant, (c) insertion of 1 to 10 consecutive or non-consecutive nucleotides in one or more regions of the gRNA variant, (d) substitution of a scaffold stem-loop or elongation stem-loop with an RNA stem-loop sequence from a heterologous RNA source having the proximal 5' and 3' ends, or any combination of (a) to (d). Any combination of substitutions, insertions, and deletions described herein can be used to generate the gRNA variants of the Disclosure. For example, a gNA variant may include at least one substitution and at least one deletion compared to the reference gRNA, at least one substitution and at least one insertion compared to the reference gRNA, at least one insertion and at least one deletion compared to the reference gRNA, or at least one substitution, one insertion, and one deletion compared to the reference gRNA.
[0113] In some embodiments, the sgRNA variants of the present disclosure include one or more additional modifications to a previously generated variant, where the previously generated variant itself functions as the modified sequence. In some embodiments, the sgRNA variants include one or more further modifications to the sequence of SEQ ID NO: 2238, SEQ ID NO: 2239, SEQ ID NO: 2240, SEQ ID NO: 2241, SEQ ID NO: 2241, SEQ ID NO: 2274, SEQ ID NO: 2275, SEQ ID NO: 2279, or SEQ ID NO: 2285, SEQ ID NO: 26797, or SEQ ID NO: 26800.
[0114] In exemplary embodiments, the gRNA variant comprises one or more modifications to the gRNA scaffold variant 174 (SEQ ID NO: 2238), and the resulting gRNA variant exhibits functional improvement compared to the parent 174 when evaluated in an in vitro or in vivo assay under equivalent conditions.
[0115] In exemplary embodiments, the gRNA variant comprises one or more modifications to the gRNA scaffold variant 175 (SEQ ID NO: 2239), and the resulting gRNA variant exhibits functional improvement compared to parent 174 when evaluated in an in vitro or in vivo assay under equivalent conditions.
[0116] In exemplary embodiments, the gRNA variant comprises one or more modifications to the gRNA scaffold variant 215 (SEQ ID NO: 2275), and the resulting gRNA variant exhibits functional improvement compared to the parent 215 when evaluated in an in vitro or in vivo assay under equivalent conditions.
[0117] In exemplary embodiments, the gRNA variant comprises one or more modifications to the gRNA scaffold variant 221 (SEQ ID NO: 2281), and the resulting gRNA variant exhibits functional improvement compared to the parent 221 when evaluated in an in vitro or in vivo assay under equivalent conditions.
[0118] In exemplary embodiments, the gRNA variant comprises one or more modifications to the gRNA scaffold variant 225 (SEQ ID NO: 2285), and the resulting gRNA variant exhibits functional improvement compared to parent 225 when evaluated in an in vitro or in vivo assay under equivalent conditions.
[0119] In exemplary embodiments, the gRNA variant comprises one or more modifications to the gRNA scaffold variant 235 (SEQ ID NO: 26800), and the resulting gRNA variant exhibits functional improvement compared to the parent 225 when evaluated in an in vitro or in vivo assay under equivalent conditions.
[0120] In some embodiments, the gRNA variant includes an exogenous elongated stem-loop, the difference from the reference gRNA being as follows. In some embodiments, the exogenous elongated stem-loop has little to no identity with the reference stem-loop region disclosed herein (e.g., SEQ ID NO: 15). In some embodiments, the exogenous stem-loop is at least 10 bp, at least 20 bp, at least 30 bp, at least 40 bp, at least 50 bp, at least 60 bp, at least 70 bp, at least 80 bp, at least 90 bp, at least 100 bp, at least 200 bp, at least 300 bp, at least 400 bp, at least 500 bp, at least 600 bp, at least 700 bp, at least 800 bp, at least 900 bp, at least 1,000 bp, at least 2,000 bp, at least 3,000 bp, at least 4,000 bp, at least 5,000 bp, at least 6,000 bp, at least 7,000 bp, at least 8,000 bp, at least 9,000 bp, at least 10,000 bp, at least 12,000 bp, at least 15,000 bp, or at least 20,000 bp. In some embodiments, the gRNA variant includes an elongated stem-loop region containing at least 10, at least 100, at least 500, at least 1000, or at least 10,000 nucleotides. In some embodiments, the heterologous stem-loop increases the stability of the gRNA. In some embodiments, the heterologous RNA stem-loop can bind to proteins, RNA structures, DNA sequences, or small molecules. In some embodiments, the exogenous stem-loop region substituting the stem-loop includes an RNA stem-loop or a hairpin, and the resulting gRNA has increased stability and, depending on the choice of loop, can interact with specific cellular proteins or RNAs.Such exogenous elongated stem-loops include, for example, thermostable RNA, MS2 hairpin (ACAUGAGGAUCACCCAUGU (SEQ ID NO: 27)), Qβ hairpin (UGCAUGUCUAAGACAGCA (SEQ ID NO: 28)), U1 hairpin II (AAUCCAUUGCACUCCGGAUU (SEQ ID NO: 29)), Uvsx (CCUCUUCGGAGG (SEQ ID NO: 30)), PP7 hairpin (AGGAGUUUCUAUGGAAACCCU (SEQ ID NO: 31)), phage replication loop (AGGUGGGACGACCUCUCGGUCGUCCUAUCU (SEQ ID NO: 32)), kissing loop_a (UGCUCGCUCCGUUCGAGCA (SEQ ID NO: 33)), and kissing loop_b. This may include 1 (UGCUCGACGCGUCCUCGAGCA (SEQ ID NO: 34)), kissing loop_b2 (UGCUCGUUUGCGGCUACGAGCA (SEQ ID NO: 35)), G triple-stranded M3q (AGGGAGGGAGGGAGAGG (SEQ ID NO: 149)), G quadruple-stranded telomere basket (GGUUAGGGUUAGGGUUAGG (SEQ ID NO: 150)), salsin-lysine loop (CUGCUCAGUACGAGAGGAACCGCAG (SEQ ID NO: 151)), or pseudoknot (UACACUGGGAUCGCUGAAUUAGAGAUCGGCGUCCUUUCAUUCUAUAUACUUUGGAGUUUUAAAAUGUCUCUAAGUACA (SEQ ID NO: 152)). In some embodiments, one of the aforementioned hairpin sequences is incorporated into the stem-loop to help transport the incorporation of gRNA (and associated CasX in the RNP complex) into budding XDP (described in more detail below).
[0121] In embodiments of the gRNA variant, the gRNA variant further comprises a spacer (or targeting sequence) (described in more detail above) located at the 3' end of the gRNA and capable of hybridizing with a target nucleic acid specific to the DMPK sequence, wherein the spacer is designed using a sequence comprising at least 14 to about 35 nucleotides and complementary to the target DNA. In some embodiments, the encoded gRNA variant comprises a targeting sequence of at least 10 to 20 nucleotides complementary to the target DNA. In some embodiments, the targeting sequence has 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, or 35 nucleotides. In some embodiments, the encoded gRNA variant comprises a targeting sequence having 20 nucleotides. In some embodiments, the targeting sequence has 25 nucleotides. In some embodiments, the targeting sequence has 24 nucleotides. In some embodiments, the targeting sequence has 23 nucleotides. In some embodiments, the targeting sequence has 22 nucleotides. In some embodiments, the targeting sequence has 21 nucleotides. In some embodiments, the targeting sequence has 20 nucleotides. In some embodiments, the targeting sequence has 19 nucleotides. In some embodiments, the targeting sequence has 18 nucleotides. In some embodiments, the targeting sequence has 17 nucleotides. In some embodiments, the targeting sequence has 16 nucleotides. In some embodiments, the targeting sequence has 15 nucleotides. In some embodiments, the targeting sequence has 14 nucleotides.
[0122] Formation of a complex with the h.CasX protein In some embodiments, the gRNA variant is complexed as an RNP with a CasX variant protein containing, at the time of expression, one of the sequences in Table 4 (SEQ ID NOs. 36-99, 101-148, and 26908-27154), or a sequence having at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity with that sequence. In some embodiments, at expression, the gRNA variant is complexed as an RNP with a CasX variant protein containing one of sequence numbers 59, 72-99, 101-148, or 26908-27154, or a sequence having at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity with it. In some embodiments, at expression, the gRNA variant is complexed as an RNP with a CasX variant protein containing one of sequence numbers 132-148 or 26908-27154, or a sequence having at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity with it.
[0123] In some embodiments, the gRNA variant has an improved ability to form a complex with the CasX protein (e.g., reference CasX or CasX variant protein) compared to the reference gRNA. In some embodiments, the gRNA variant has an improved affinity for the CasX protein (e.g., reference protein or variant protein) compared to the reference gRNA, thereby improving its ability to form a ribonucleoprotein (RNP) complex with the CasX protein, as described in the examples. Improving the formation of the ribonucleoprotein complex can, in some embodiments, improve the efficiency of constructing a functional RNP. In some embodiments, more than 90%, more than 93%, more than 95%, more than 96%, more than 97%, more than 98%, or more than 99% of the RNP containing the gRNA variant and its targeting sequence are eligible for gene editing of the target nucleic acid.
[0124] Exemplary nucleotide alterations that can improve the ability of gRNA variants to complex with the CasX protein may, in some embodiments, involve replacing the scaffold stem with a thermostable stem-loop. While we do not wish to be bound by any theory, replacing the scaffold stem with a thermostable stem-loop may increase the overall binding stability between the gRNA variant and the CasX protein. Alternatively, or additionally, by removing a large portion of the stem-loop, the folding dynamics of the gRNA variant may be altered, for example, by reducing the degree to which the gRNA variant can "entangle" itself, thereby allowing for more easily and rapidly forming of a functionally folded gRNA. In some embodiments, the selection of the scaffold stem-loop sequence may vary depending on the different targeting sequences utilized for the gRNA. In some embodiments, the scaffold sequence may be tailored to the targeting sequence, and therefore the target sequence. Biochemical assays, including the assays of the examples, can be used to evaluate the binding affinity of the CasX protein to gRNA variants for RNP formation. For example, those skilled in the art can measure the change in the amount of fluorescently tagged gRNA bound to the immobilized CasX protein in response to an increase in the concentration of additional unlabeled "cold competitor" gRNA. Alternatively, or additionally, the changes in the fluorescence signal can be monitored or observed as different amounts of fluorescently labeled gRNA flow along the immobilized CasX protein. Alternatively, the ability to form RNPs can be assessed using an in vitro cleavage assay against a defined target nucleic acid sequence.
[0125] IV. Proteins for modifying target nucleic acids This disclosure provides a system comprising CRISPR nucleases useful in genome editing of eukaryotic cells. In some embodiments, the CRISPR nucleases used in the genome editing system are class 2, type V nucleases. While there are differences among the members of the class 2, type V CRISPR-Cas system, they share several common features that distinguish them from the Cas9 system. First, class 2, type V nucleases have a single RNA-inducible RuvC domain-containing effector but lack an HNH domain and recognize T-rich PAMs 5' upstream of the target region on the non-target strand, which differs from the Cas9 system, which relies on G-rich PAMs 3' on the target sequence side. Unlike Cas9, which produces blunt ends proximal to the PAM, type V nucleases produce double-strand breaks distal to the PAM sequence. In addition, when activated by target dsDNA or ssDNA bound in cis, type V nucleases degrade ssDNA in trans. In some embodiments, the V-type nuclease of this embodiment recognizes a 5'-TC PAM motif and generates a sticky end that is cleaved only by the RuvC domain. In some embodiments, the V-type nuclease is selected from the group consisting of Cas12a, Cas12b, Cas12c, Cas12d (CasY), Cas12j, Cas12k, CasZ, and CasX. In some embodiments, the disclosure provides a system (CasX:gRNA system) comprising a CasX protein and one or more gRNA acids, the gRNA acids being specifically designed to modify a target nucleic acid sequence in eukaryotic cells.
[0126] As used herein, the term “CasX protein” refers to a family of proteins, encompassing all naturally occurring CasX proteins, proteins that share at least 50% identity with naturally occurring CasX proteins, and CasX variants that possess one or more improved characteristics compared to naturally occurring reference CasX proteins.
[0127] The CasX proteins of the present disclosure include at least one of the following domains: a non-target strand binding (NTSB) domain, a target strand loading (TSL) domain, a helix I domain, a helix II domain, an oligonucleotide binding domain (OBD), and a RuvC DNA cleavage domain.
[0128] In some embodiments, the CasX protein can bind to and / or modify (e.g., nick, catalyze double-strand cleavage, methylate, demethylate, etc.) a target nucleic acid at a specific sequence that is targeted by an associated gRNA that hybridizes to a sequence within the target nucleic acid sequence.
[0129] a. Reference CasX protein The present disclosure provides a naturally occurring CasX protein (referred to herein as the "reference CasX protein"), which is subsequently modified to generate CasX variants of the present disclosure. For example, the reference CasX protein can be isolated from naturally occurring prokaryotes such as Deltaproteobacteria, Planctomycetes, or Candidatus Sungbacteria species. The reference CasX protein (alternatively referred to herein as the reference CasX polypeptide) is a type II CRISPR / Cas endonuclease belonging to the CasX (alternatively referred to as Cas12e) family of proteins that interact with guide RNAs to form ribonucleoprotein (RNP) complexes.
[0130] In some cases, the reference CasX protein is isolated from or derived from Deltaproteobacter and has the following sequence:
[0131] [Table 4]
[0132] In some cases, the reference CasX protein is isolated from or derived from Plantomycetes and has the following sequence:
[0133] [Table 5]
[0134] In some cases, the reference CasX protein is isolated from or derived from Candidatus Sungbacteria and has the following sequence:
[0135] [Table 6]
[0136] b.CasX variant protein This disclosure provides a variant of the reference CasX protein (hereinafter interchangeably referred to as "CasX variant" or "CasX variant protein"), the CasX variant comprising at least one modification in at least one domain compared to the reference CasX protein, and comprising, but not limited to, the sequences of SEQ ID NOs: 1-3.
[0137] The CasX variants of this disclosure have one or more improved features compared to the reference CasX protein. Exemplary improved features of embodiments of the CasX variants include, but are not limited to, improved variant folding, improved binding affinity to gRNA, improved binding affinity to target nucleic acids, improved ability to utilize a wider range of PAM sequences in editing and / or binding of target DNA, improved target DNA unwinding, increased editing activity, improved editing efficiency, improved editing specificity, increased percentage of eukaryotic genomes that can be efficiently edited, increased nuclease activity, increased target strand loading for double-strand breaks, decreased target strand loading for single-strand nicking, decreased off-target breaks, improved binding of non-target strands of DNA, improved protein stability, improved stability of the protein:gRNA(RNP) complex, improved protein solubility, improved solubility of the protein:gRNA(RNP) complex, improved protein yield, improved protein expression, and improved fusion features (described in more detail below). Exemplary improved features are described in International Publication No. 2020 / 247882(A1) and International Publication No. 2020 / 247883 (incorporated herein by reference). In the embodiments described above, one or more of the improved features of the CasX variant are improved by at least about 1.1 to about 100,000 times compared to the reference CasX protein of SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 3 when assayed in a comparable manner. In other embodiments, the improvement is at least about 1.1 times, at least about 2 times, at least about 5 times, at least about 10 times, at least about 50 times, at least about 100 times, at least about 500 times, at least about 100 times, at least about 1000 times, at least about 5000 times, at least about 10,000 times, or at least about 100,000 times compared to the reference CasX protein of SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 3 when assayed in a comparable manner.In other cases, one or more improved features of the RNPs of CasX variants and gRNA variants are improved by at least about 1.1, at least about 10, at least about 100, at least about 1000, at least about 10,000, at least about 100,000, or more compared to the RNPs of the reference CasX protein of SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 3 and the gRNAs of Table 2. In other cases, one or more of the improved features of the RNPs of CasX variants and gRNA variants, when assayed in a comparable manner, are approximately 1.1–100,000, 1.1–10,000, 1.1–1,000, 1.1–500, 1.1–100, 1.1–50, 1.1–20, 10–100,000, 10–10,000, 10–1,000, 10–500, 10–100, 10–50, 10–20, 2–70, and 2–50, respectively. 2 to 30 times, 2 to 20 times, 2 to 10 times, 5 to 50 times, 5 to 30 times, 5 to 10 times, 100 to 100,00 times, 100 to 10,00 times , about 100 to 1,000 times, about 100 to 500 times, about 500 to 100,00 times, about 500 to 10,00 times, about 500 to 1,000 times, about 500 to 750 The improvement is approximately 2x, 1,000-100,000x, 10,000-100,000x, 20-500x, 20-250x, 20-200x, 20-100x, 20-50x, 50-10,000x, 50-1,000x, 50-500x, 50-200x, or 50-100x.In other cases, one or more improved features of the RNPs of CasX variants and gRNA variants, when assayed in a comparable manner, were approximately 1.1x, 1.2x, 1.3x, 1.4x, 1.5x, 1.6x, 1.7x, 1.8x, 1.9x, 2x, 3x, 4x, 5x, 6x, 7x, 8x, 9x, 10x, 11x, 12x, 13x, 14x, 15x, 16x, 17x, 18x, 19x, 20x, 16x, 17x, 18x, 19x, 20x, 16x, 17x, 18x, 19x, 16x, 19x, 16x, 19x, 16x, 19x, 16x, 17x, 189x, 16x, 19x, 16x, 19x, 16x, 19x, 16x, 19x, 16x, 19x, 16x, 19x, 16x, 19x, 16x, 19x, 16x, 19x, 16x, 19x, 16x, 19x, 16x, 19x, It is improved by 25 times, 30 times, 40 times, 45 times, 50 times, 55 times, 60 times, 70 times, 80 times, 90 times, 100 times, 110 times, 120 times, 130 times, 140 times, 150 times, 160 times, 170 times, 180 times, 190 times, 200 times, 210 times, 220 times, 230 times, 240 times, 250 times, 260 times, 270 times, 280 times, 290 times, 300 times, 310 times, 320 times, 330 times, 340 times, 350 times, 360 times, 370 times, 380 times, 390 times, 400 times, 425 times, 450 times, 475 times, or 500 times.
[0138] The term CasX variant includes variants that are fusion proteins; that is, CasX is "fused" to a heterologous sequence. This includes CasX variant sequences and CasX variants that include N-terminal, C-terminal, or internal fusion of CasX to a heterologous protein or its domain.
[0139] In some embodiments, the CasX variant includes at least one modification in the NTSB domain. In some embodiments, the CasX variant includes at least one modification in the TSL domain. In some embodiments, the CasX variant includes at least one modification in the helix I domain. In some embodiments, the CasX variant includes at least one modification in the helix II domain. In some embodiments, the CasX variant includes at least one modification in the OBD domain. In some embodiments, the CasX variant includes at least one modification in the RuvC DNA cleavage domain. In some embodiments, at least one modification in the RuvC DNA cleavage domain includes one or more amino acid substitutions from among the amino acids K682, G695, A708, V711, D732, A739, D733, L742, V747, F755, M771, M779, W782, A788, G791, L792, P793, Y797, M799, Q804, S819, or Y857 of SEQ ID NO: 2, or a deletion of amino acid P793.
[0140] In some embodiments, the CasX variant protein includes at least one modification in at least one domain of the reference CasX protein containing sequences 1-3, in each of at least two domains, in each of at least three domains, in each of at least four domains, or in each of at least five domains. In some embodiments, the CasX variant protein includes two or more modifications in at least one domain of the reference CasX protein. In some embodiments, the CasX variant protein includes at least two modifications in at least one domain of the reference CasX protein, at least three modifications in at least one domain of the reference CasX protein, or at least four modifications in at least one domain of the reference CasX protein. In some embodiments, the CasX variant includes two or more modifications compared to the reference CasX, and each modification is independently made in a domain selected from the group consisting of NTSBD, TSLD, helix I domain, helix II domain, OBD, and RuvC DNA cleavage domain. In some embodiments, at least one modification of a CasX variant protein includes a deletion of at least a portion of one domain of the reference CasX protein of SEQ ID NOs: 1–3. In some embodiments, the deletion is located in the NTSBD, TSLD, helix I domain, helix II domain, OBD, or RuvC DNA cleavage domain. In other embodiments, the disclosure provides a CasX variant that includes at least one modification compared to another CasX variant. For example, CasX variant 515 is a variant of CasX variant 491. Any variant that improves one or more functions or characteristics of a CasX variant protein compared to the reference CasX protein (or variants from which it is derived) described herein is assumed to be within the scope of the disclosure.
[0141] In some embodiments, a modification of a CasX variant is a mutation in one or more amino acids of the reference CasX. In other embodiments, a modification is the substitution of one or more domains of the reference CasX with one or more domains derived from a different CasX. In some embodiments, an insertion includes the insertion of some or all of a domain derived from a different CasX protein. Mutations may occur in any one or more domains of the reference CasX protein, and may include, for example, the deletion of some or all of one or more domains in any domain of the reference CasX protein, or the substitution, deletion, or insertion of one or more amino acids. The domains of the CasX protein include the non-target chain binding (NTSB) domain, the targeted chain loading (TSL) domain, the helix I domain, the helix II domain, the oligonucleotide binding domain (OBD), and the RuvC DNA cleavage domain. Any change in the amino acid sequence of the reference CasX protein that results in improved characteristics of the CasX protein is considered a CasX variant protein as described herein. For example, a CasX variant may contain one or more amino acid substitutions, insertions, deletions, or swapped domains, or any combination thereof, compared to a reference CasX protein sequence.
[0142] Preferred mutagenesis methods for generating the CasX variant proteins of this disclosure may include, for example, comprehensive mutational evolution (DME), comprehensive mutational scanning (DMS), error-prone PCR, cassette mutagenesis, random mutagenesis, adhesion extension PCR, gene shuffling, or domain swapping. In some embodiments, the CasX variant is designed, for example, by selecting one or more desired mutations in a reference CasX. In certain embodiments, the activity of a reference CasX protein is used as a benchmark against which the activity of one or more CasX variants is compared, thereby measuring the improvement in the function of the CasX variant.
[0143] In some embodiments of the CasX variants described herein, at least one modification includes: (a) substitution of 1 to 100 consecutive or non-consecutive amino acids in the CasX variant compared to reference CasX of SEQ ID NO: 1, SEQ ID NO: 2, SEQ ID NO: 3, CasX variant 491, or CasX variant 515; (b) deletion of 1 to 100 consecutive or non-consecutive amino acids in the CasX variant compared to reference CasX or a variant derived therefrom; (c) insertion of 1 to 100 consecutive or non-consecutive amino acids in CasX compared to reference CasX or a variant derived therefrom; or (d) any combination of (a) to (c). In some embodiments, at least one modification includes: (a) substitution of 5 to 10 consecutive or non-consecutive amino acids in a CasX variant compared to reference CasX of SEQ ID NO: 1, SEQ ID NO: 2, SEQ ID NO: 3, CasX 491, or CasX 515; (b) deletion of 1 to 5 consecutive or non-consecutive amino acids in a CasX variant compared to reference CasX or a variant derived therefrom; (c) insertion of 1 to 5 consecutive or non-consecutive amino acids in CasX compared to reference CasX or a variant derived therefrom; or (d) any combination of (a) to (c).
[0144] In some embodiments, the CasX variant protein comprises or consists of a sequence having at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, or at least 100 modifications compared to the sequence of SEQ ID NO: 1, SEQ ID NO: 2, SEQ ID NO: 3, CasX 491 (see Table 4), or CasX 515 (see Table 4). These modifications may be amino acid insertions, deletions, substitutions, or any combination thereof. The modifications may be in one domain of the CasX variant, or in any of the domains, or in any combination of the domains. In the substitutions described herein, any amino acid may be substituted for any other amino acid. Substitutions may be conserved substitutions (e.g., a basic amino acid is substituted for another basic amino acid). Substitutions may be non-conservative substitutions (e.g., a basic amino acid is substituted for an acidic amino acid, or vice versa). For example, the CasX variant protein of this disclosure can be generated by substituting proline in the reference CasX protein with any of the following: arginine, histidine, lysine, aspartic acid, glutamic acid, serine, threonine, asparagine, glutamine, cysteine, glycine, alanine, isoleucine, leucine, methionine, phenylalanine, tryptophan, tyrosine, or valine.
[0145] Any permutation of substitutions, insertions, and deletions of the embodiments described herein can be combined to generate the CasX variant proteins of this disclosure. For example, a CasX variant protein may include at least one substitution and at least one deletion compared to a reference CasX protein sequence, at least one substitution and at least one insertion compared to a reference CasX protein sequence, at least one insertion and at least one deletion compared to a reference CasX protein sequence, or at least one substitution, one insertion, and one deletion compared to a reference CasX protein sequence.
[0146] In some embodiments, the CasX variant includes at least one modification selected from one or more of the following, compared to the reference CasX sequence of SEQ ID NO: (a) L379R amino acid substitution, (b) A708K amino acid substitution, (c) T620P amino acid substitution, (d) E385P amino acid substitution, (e) Y857R amino acid substitution, (f) I658V amino acid substitution, (g) F399L amino acid substitution, (h) Q252K amino acid substitution, (i) L404K amino acid substitution, and (j) P793 amino acid deletion.
[0147] In some embodiments, the CasX variant protein comprises 400-2000 amino acids, 500-1500 amino acids, 700-1200 amino acids, 800-1100 amino acids, or 900-1000 amino acids.
[0148] In some embodiments, the CasX variant protein includes sequences 59, 72-99, 101-148, and 26908-27154 shown in Table 4. In other ways, the CasX variant protein contains sequences that are at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 86%, at least 87%, at least 88%, at least 89%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, and at least 99.5% identical to sequences of sequence numbers 59, 72-99, 101-148, or 26908-27154 shown in Table 4. In some embodiments, the CasX variant protein contains or consists of sequences 536-99, 101-148, or 26908-27154. In other words, CasX variant proteins contain sequences that are at least 60% identical, at least 65% identical, at least 70% identical, at least 75% identical, at least 80% identical, at least 81% identical, at least 82% identical, at least 83% identical, at least 84% identical, at least 85% identical, at least 86% identical, at least 86% identical, at least 87% identical, at least 88% identical, at least 89% identical, at least 89% identical, at least 90% identical, at least 91% identical, at least 92% identical, at least 93% identical, at least 94% identical, at least 95% identical, at least 96% identical, at least 97% identical, at least 98% identical, at least 99% identical, and at least 99.5% identical to sequences of sequence numbers 36-99, 101-148, or 26908-27154.In some embodiments, the CasX variant protein comprises or consists of the sequences of SEQ ID NOs: 132-148 or 26908-27154. In other embodiments, the CasX variant protein comprises a sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to the sequences of SEQ ID NOs: 132-148 or 26908-27154.
[0149]
Table 7-1
[0150]
Table 7-2
[0151]
Table 7-3
[0152]
Table 7-4
[0153]
Table 7-5
[0154]
Table 7-6
[0155] [Table 7-7]
[0156] c. CasX variant protein having domains derived from multiple source proteins In certain embodiments, the Disclosure provides a chimeric CasX protein comprising protein domains derived from two or more different CasX proteins (e.g., two or more naturally occurring CasX proteins or two or more CasX variant protein sequences described herein). As used herein, “chimeric CasX protein” means a CasX comprising at least two domains isolated or derived from different sources (e.g., two naturally occurring proteins), which in some embodiments may be isolated from different species. For example, in some embodiments, a chimeric CasX protein comprises a first domain derived from a first CasX protein and a second domain derived from a second different CasX protein. In some embodiments, the first domain may be selected from the group consisting of NTSB, TSL, helix I, helix II, OBD, and RuvC domains. In some embodiments, the second domain is selected from the group consisting of NTSB, TSL, helix I, helix II, OBD, and RuvC domains, and the second domain is different from the first domain described herein. In the case of split or discontinuous domains such as helix I, RuvC, and OBD, a portion of the discontinuous domain can be replaced with a corresponding portion from any other source. For example, the helix II domain in SEQ ID NO: 2 (sometimes referred to as helix Ia) can be replaced with the corresponding helix II sequence from SEQ ID NO: 1. Table 5 shows the domain sequences and their coordinates from the reference CasX protein. Representative examples of chimeric CasX proteins include variants of CasX 472-483, 485-491, and 515, whose sequences are shown in Table 4.
[0157] [Table 8] * OBDa and b, helix Ia and b, and RuvCa and b are also referred to herein as OBDI and II, helix II and I-II, and RuvCI and II.
[0158] Protein affinity for d.gRNA In some embodiments, the CasX variant protein has improved affinity for gRNA compared to the reference CasX protein, leading to the formation of a ribonucleoprotein complex (RNP). The increased affinity of the CasX variant protein for gRNA results in, for example, lower K for RNP complex formation. d This can result in the formation of a more stable ribonucleoprotein complex. In some embodiments, increased affinity of the CasX variant protein to gRNA results in increased stability of the ribonucleoprotein complex when delivered to human cells. This increased stability can affect the function and utility of the complex in the target cells and may result in improved pharmacokinetic properties in the blood when delivered to the target. In some embodiments, increased affinity of the CasX variant protein and the resulting increased stability of the ribonucleoprotein complex result in a lower dose of the CasX variant protein delivered to the target or cell while still having the desired activity, e.g., gene editing in vivo or in vitro. In some embodiments, higher affinity (stronger binding) of the CasX variant protein to gRNA allows for more editing events if both the CasX variant protein and gRNA remain in the RNP complex. The increase in editing events can be evaluated using an editing assay such as the tdTom editing assay described herein. In some embodiments, the K of the CasX variant protein to gRNA dThe binding affinity is increased by at least about 1.1, at least about 1.2, at least about 1.3, at least about 1.4, at least about 1.5, at least about 1.6, at least about 1.7, at least about 1.8, at least about 1.9, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 15, at least about 20, at least about 25, at least about 30, at least about 35, at least about 40, at least about 45, at least about 50, at least about 60, at least about 70, at least about 80, at least about 90, or at least about 100 times compared to the reference CasX protein. In some embodiments, the CasX variant has a binding affinity to gRNA that is increased by about 1.1 to about 10 times compared to the reference CasX protein of SEQ ID NO: 2.
[0159] In some embodiments, increased affinity of the CasX variant protein to gRNA results in increased stability of the ribonucleoprotein complex when delivered to mammalian cells (including in vivo delivery to the target). This increased stability may affect the function and utility of the complex in the target cells and may result in improved pharmacokinetic properties in the blood when delivered to the target. In some embodiments, increased affinity of the CasX variant protein and increased stability of the resulting ribonucleoprotein complex result in lower doses of the CasX variant protein delivered to the target or cells while still having the desired activity (e.g., in vivo or in vitro gene editing). The increased ability to form RNPs and maintain them in a stable form can be evaluated using assays (e.g., in vitro cleavage assays described in the examples herein). In some embodiments, RNPs containing the CasX variant of this disclosure, when complexed as RNPs, exhibit at least 2-fold, at least 5-fold, or at least 10-fold higher kJ compared to RNPs containing the reference CasX of SEQ ID NOs: 1-3. 切断 Speed can be achieved.
[0160] In some embodiments, a higher affinity (stronger binding) of the CasX variant protein to the gRNA allows for more editing events if both the CasX variant protein and gRNA remain in the RNP complex. The increase in editing events can be evaluated using editing assays such as those described herein.
[0161] While we do not wish to be bound by theory, in some embodiments, amino acid changes in the helix I domain can increase the binding affinity between the CasX variant protein and the gRNA targeting sequence, changes in the helix II domain can increase the binding affinity between the CasX variant protein and the gRNA scaffold stem-loop, and changes in the oligonucleotide-binding domain (OBD) can increase the binding affinity between the CasX variant protein and the gRNA triple helix.
[0162] Methods for measuring the binding affinity of CasX protein to gRNA include in vitro methods using purified CasX protein and gRNA. Binding affinity to reference CasX and variant proteins can be measured by fluorescence polarization if the gRNA or CasX protein is tagged with a fluorophore. Alternatively or additionally, binding affinity can be measured by biolayer interferometry, electrophoretic mobility shift assay (EMSA), or membrane binding. Further standard techniques for quantifying the absolute affinity of RNA-binding proteins, such as reference CasX and the variant proteins of this disclosure, to specific gRNAs, such as reference gRNA and its variants, include, but are not limited to, isothermal calorimetry (ITC) and surface plasmon resonance (SPR), as well as the methods described in the examples.
[0163] e. Affinity for target nucleic acids In some embodiments, the CasX variant protein has improved binding affinity to the target nucleic acid sequence compared to the affinity of the reference CasX protein to the target nucleic acid sequence. In some embodiments, the affinity of the CasX variant protein of this disclosure to the target nucleic acid molecule is increased by at least about 1.1, at least about 1.2, at least about 1.3, at least about 1.4, at least about 1.5, at least about 1.6, at least about 1.7, at least about 1.8, at least about 1.9, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 15, at least about 20, at least about 25, at least about 30, at least about 35, at least about 40, at least about 45, at least about 50, at least about 60, at least about 70, at least about 80, at least about 90, or at least about 100 times compared to the reference CasX protein.
[0164] CasX variants with higher affinity for those target nucleic acids can, in some embodiments, cleave target nucleic acid sequences more rapidly than reference CasX proteins without increased affinity for the target nucleic acids. In some embodiments, improved affinity for target nucleic acid sequences results in increased ability to modify target nucleic acids, including improved affinity for target nucleic acid sequences, improved binding affinity for a wider range of PAM sequences, improved ability to search for DNA for target nucleic acid sequences, or any combination thereof. In some embodiments, CasX variant proteins with improved target nucleic acid affinity have increased affinity for specific PAM sequences other than the canonical TTC PAM recognized by the reference CasX protein of SEQ ID NO: 2, including binding affinity for PAM sequences selected from the group consisting of TTC, ATC, GTC, and CTC. Higher overall affinity for DNA can also, in some embodiments, increase the frequency with which the CasX protein can effectively initiate and terminate the binding and unwinding steps, thereby facilitating target strand entry and R-loop formation, and ultimately cleavage of the target nucleic acid sequence.
[0165] In some embodiments, the CasX variant protein has improved binding affinity to the non-target strand of the target nucleic acid. As used herein, the term “non-target strand” refers to the strand of the DNA target nucleic acid sequence that does not form Watson and Crick base pairs with the targeting sequence in the gRNA and is complementary to the target DNA strand. In some embodiments, the CasX variant protein has approximately 1.1 to approximately 100-fold increased binding affinity to the non-target strand of the target nucleic acid compared to the reference protein of SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 3.
[0166] Methods for measuring the affinity of CasX variant proteins to target nucleic acid molecules may include electrophoretic mobility shift assays (EMSA), membrane binding, isothermal calorimetry (ITC), and surface plasmon resonance (SPR), fluorescence polarization, and biolayer interferometry (BLI). Further methods for measuring the affinity of CasX proteins to targets include in vitro biochemical assays that measure DNA cleavage events over time (e.g., as described in the examples, k 切断 (Determining the speed) is one example.
[0167] f. Improved specificity for target sites In some embodiments, CasX variant proteins exhibit improved specificity to target nucleic acid sequences compared to the reference CasX protein. As used herein, “specificity” (interchangeably referred to as “target specificity”) refers to the degree to which the CRISPR / Cas system ribonucleoprotein complex cleaves off-target sequences that are similar to, but not identical to, the target nucleic acid sequence. For example, a CasX variant RNP with a higher degree of specificity exhibits reduced off-target cleavage of the sequence compared to the reference CasX protein. The specificity of CRISPR / Cas system proteins, and the reduction of potentially harmful off-target effects, can be critical to achieving an acceptable therapeutic index for use in mammalian subjects.
[0168] In some embodiments, CasX variant proteins exhibit improved specificity to target sites within target sequences complementary to the targeting sequence of gRNA, compared to the reference CasX proteins of SEQ ID NOs: 1–3. While we do not wish to be bound by theory, amino acid changes in the helix I and II domains that increase the specificity of the CasX variant protein to the target nucleic acid strand may also increase the specificity of the CasX variant protein to the entire target nucleic acid. In some embodiments, amino acid changes that increase the specificity of the CasX variant protein to the target nucleic acid may also result in a decrease in the affinity of the CasX variant protein to DNA.
[0169] Methods for testing the target specificity of CasX proteins (e.g., variants or reference) may include circularization for in vitro reporting of cleavage effects by sequencing (CIRCLE-seq), or similar methods. Briefly, the CIRCLE-seq technique involves shearing genomic DNA, circularizing it by ligating a stem-loop adapter, nicking the stem-loop region, and exposing a 4-nucleotide palindromic overhang. This is followed by intramolecular ligation and degradation of the remaining linear DNA. Subsequently, the circular DNA molecule containing the CasX cleavage site is linearized with CasX, the adapter adapter is ligated to the exposed ends, and then high-throughput sequencing is performed to generate paired-end reads containing information about off-target sites. Further assays that can be used to detect off-target events (and therefore the specificity of the CasX protein) include assays used to detect and quantify indels (insertions and deletions) formed at selected off-target sites (e.g., mismatch detection nuclease assays and next-generation sequencing (NGS)). An example of a mismatch detection assay is a nuclease assay, which involves PCR amplification, denaturation, and re-hybridization of genomic DNA from cells treated with CasX and sgRNA to form heterodouble-stranded DNA containing one wild-type strand and one indel-containing strand. The mismatch is recognized and cleaved by a mismatch detection nuclease (e.g., Surveyor nuclease or T7 endonuclease I).
[0170] g. Protospacer and PAM array In this specification, a protospacer is defined as a DNA sequence complementary to the targeting sequence of the guide RNA and a DNA sequence complementary to that sequence, referred to as the target strand and non-target strand, respectively. As used herein, PAM is a nucleotide sequence located one nucleotide 5' to the sequence in the non-target strand that is complementary to the target nucleic acid sequence in the target strand of the target nucleic acid, and together with the targeting sequence of the gRNA, assists in the orientation and positioning of CasX for potential cleavage of the protospacer strand. The PAM sequence may be degenerate, and a particular RNP construct may have different preferred acceptable PAM sequences supporting different cleavage efficiencies. By convention, unless otherwise stated, this disclosure refers to both the PAM and protospacer sequences, as well as their orientation according to the orientation of the non-target strand. This does not mean that the PAM sequence on the non-target strand rather than the target strand is a determinant of cleavage or is mechanistically involved in target recognition. For example, when referring to the TTC PAM, it may actually be a complementary GAA sequence required for target cleavage, or any combination of nucleotides from both strands. In the case of the CasX protein disclosed herein, the PAM is located at the 5' end of the protospacer, and a single nucleotide separates the PAM from the first nucleotide of the protospacer. Therefore, in the case of reference CasX, TTC PAM should be understood to mean a sequence following the formula 5'-...NNTTCN(protospacer)NNNNNN...3' (wherein "N" is any DNA nucleotide and "(protospacer)" is a DNA sequence identical to the targeting sequence of the guide RNA). For CasX variants with extended PAM recognition, TTC, CTC, GTC, or ATC PAM should be understood to mean a sequence following the formula: 5'-...NNTTCN(protospacer)NNNNNN...3', 5'-...NNCTCN (protospacer)NNNNNN...3', 5'-...NNGTCN (protospacer)NNNNNN...3', or 5'-...NNATCN (protospacer)NNNNNN...3'.
[0171] Alternatively, TC PAM should be understood to mean a sequence following the formula: 5'-...NNNTCN(protospacer)NNNNNN...3'. Additionally, the CasX variant proteins of this disclosure, when complexed with gRNA as RNPs, have an enhanced ability to efficiently edit and / or bind to target DNA when utilizing a PAM TC motif containing a PAM sequence selected from TTC, ATC, GTC, or CTC (in 5' to 3' orientation), compared to the RNPs of the reference CasX protein and reference gRNA. In the foregoing, the PAM sequence is located at least one nucleotide 5' to the non-target strand of the protospacer, which is identical to the targeting sequence of gRNA in the assay system, compared to the editing efficiency and / or binding of RNPs containing the reference CasX protein and reference gRNA in a comparable assay system. In one embodiment, RNPs of CasX variants and gRNA variants exhibit higher editing efficiency and / or binding of the target sequence to the target DNA compared to RNPs containing a reference CasX protein and reference gRNA in a comparable assay system, where the PAM sequence of the target DNA is TTC. In another embodiment, RNPs of CasX variants and gRNA variants exhibit higher editing efficiency and / or binding of the target sequence to the target DNA compared to RNPs containing a reference CasX protein and reference gRNA in a comparable assay system, where the PAM sequence of the target DNA is ATC. In yet another embodiment, RNPs of CasX variants and gRNA variants exhibit higher editing efficiency and / or binding of the target sequence to the target DNA compared to RNPs containing a reference CasX protein and reference gRNA in a comparable assay system, where the PAM sequence of the target DNA is CTC. In another embodiment, RNPs of CasX variants and gRNA variants exhibit higher editing efficiency and / or binding of target sequences to target DNA compared with RNPs containing reference CasX protein and reference gRNA in a comparable assay system, where the PAM sequence of the target DNA is GTC.In the embodiments described above, the increased editing efficiency and / or binding affinity for one or more PAM sequences is at least 1.5 times greater than the editing efficiency and / or binding affinity of any one of the CasX proteins (SEQ ID NOs. 1-3) and the RNPs of the gRNAs listed in Table 2 for the PAM sequences. Exemplary assays demonstrating improved editing are described in the examples herein.
[0172] h. DNA rewind In some embodiments, CasX variant proteins have improved DNA unwinding ability compared to the reference CasX protein. It has been previously shown that insufficient dsDNA unwinding impairs or interferes with the ability of CRISPR / Cas system proteins AnaCas9 or Cas14 to cleave DNA. Therefore, while we do not wish to be bound by any theory, the increased DNA cleavage activity by some of the CasX variant proteins of this disclosure is likely at least in part to be due to an increased ability to locate and unwind dsDNA at the target site.
[0173] While we do not wish to be bound by theory, it is thought that amino acid changes in the NTSB domain can produce CasX variant proteins with increased DNA unwinding characteristics. Alternatively, or additionally, amino acid changes in helix domain regions interacting with OBD or PAM can also generate CasX variant proteins with increased DNA unwinding characteristics.
[0174] Methods for measuring the ability of a CasX protein (e.g., variant or reference) to unwind DNA include, but are not limited to, in vitro assays that observe an increase in the on-rate of a dsDNA target in fluorescence polarization or biolayer interferometry.
[0175] i.Catalytic activity The CasX:RNA system ribonucleoprotein complex g disclosed herein includes a CasX variant that binds to and cleaves a target nucleic acid sequence. In some embodiments, the CasX variant protein has improved catalytic activity compared to the reference CasX protein. While we do not wish to be bound by theory, it is conceivable that, in some cases, cleavage of the target strand may be a limiting factor for Cas12-like molecules in the generation of dsDNA cleavage. In some embodiments, the CasX variant protein improves the bending and cleavage of the target strand of DNA, resulting in an overall improvement in the efficiency of dsDNA cleavage by the CasX ribonucleoprotein complex.
[0176] In some embodiments, the CasX variant protein has increased nuclease activity compared to the reference CasX protein. Variants with increased nuclease activity can be generated, for example, through amino acid changes in the RuvC nuclease domain. In some embodiments, the CasX variant includes a RuvC nuclease domain having nickase activity. In the above, the CasX nickase in the CasX:gRNA system generates a single-strand break within 10-18 nucleotides on the 3' side of the PAM site on the non-target strand. In other embodiments, the CasX variant includes a RuvC nuclease domain having double-strand break activity. In the above, the CasX in the CasX:gRNA system generates double-strand breaks within 18-26 nucleotides on the 5' side of the PAM site on the target strand and within 10-18 nucleotides on the 3' side of the non-target strand. Nuclease activity can be assayed by various methods, including the method of the examples. In some embodiments, the CasX variant is at least 2 times, or at least 3 times, or at least 4 times, or at least 5 times, or at least 6 times, or at least 7 times, or at least 8 times, or at least 9 times, or at least 10 times larger than the reference CasX. 切断 It has a constant.
[0177] In some embodiments, the CasX variant protein has the improved characteristic of forming an RNP with gRNA, resulting in a higher percentage cleavage-competent RNP compared to the RNP of the reference CasX protein and gRNA of SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 3, as described in the Examples. Cleavage competent means that the formed RNP has the ability to cleave the target nucleic acid. In some embodiments, the RNP of the CasX variant and gRNA exhibits at least 2 times, at least 3 times, at least 4 times, at least 5 times, or at least 10 times the cleavage rate compared to the RNP of the reference CasX protein and gRNA of SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 3 and Table 2. In the embodiments described above, the improved competency rate can be demonstrated in an in vitro assay as described in the Examples.
[0178] In some embodiments, CasX variant proteins have increased target strand loading for double-strand breaks compared to reference CasX. Variants with increased target strand loading activity may be generated, for example, via amino acid alterations in the TLS domain. While we do not wish to be bound by theory, amino acid alterations in the TSL domain may result in CasX variant proteins with improved catalytic activity. Alternatively, or additionally, amino acid alterations around the RNA:DNA double-strand binding channel can also improve the catalytic activity of CasX variant proteins. In some embodiments, CasX variant proteins have increased incidental cleavage activity compared to reference CasX protein. As used herein, “incidental cleavage activity” means further untargeted cleavage of the nucleic acid following recognition and cleavage of the target nucleic acid sequence. In some embodiments, CasX variant proteins have decreased incidental cleavage activity compared to reference CasX protein.
[0179] In some embodiments, for example, embodiments encompassing applications where cleavage of the target nucleic acid sequence is not the desired result, improving the catalytic activity of the CasX variant protein includes altering, reducing, or eliminating the catalytic activity of the CasX variant protein. In some embodiments, a ribonucleoprotein complex containing the dCasX variant protein binds to the target nucleic acid sequence but does not cleave the target nucleic acid.
[0180] In some embodiments, the CasX ribonucleoprotein complex containing the CasX variant protein binds to target DNA but generates single-strand nicks in the target DNA. In some embodiments, particularly those in which the CasX protein is a nickase, the CasX variant protein has reduced target strand loading for single-strand nicking. Variants with reduced target strand loading can be generated, for example, through amino acid changes in the TSL domain.
[0181] Exemplary methods for characterizing the catalytic activity of the CasX protein include, but are not limited to, in vitro cleavage assays, including the methods described in the following examples. In some embodiments, electrophoresis of DNA products on an agarose gel can be used to investigate the dynamics of strand breaks.
[0182] j.CasX fusion protein In some embodiments, the Disclosure provides a CasX protein comprising a heterologous protein fused to CasX. In some cases, CasX is a reference CasX protein. In other cases, CasX is a CasX variant of any of the embodiments described herein.
[0183] In some embodiments, the CasX variant protein comprises one of the sequences SEQ ID NOs. 59, 72-99, 101-148, and 26908-27154 from Table 4, fused to one or more proteins or their domains having different activities of interest, resulting in a fusion protein. In some embodiments, the CasX variant protein comprises one of the sequences SEQ ID NOs. 36-99, 101-148, and 26908-27154, fused to one or more proteins or their domains. In some embodiments, the CasX variant protein comprises one of the sequences SEQ ID NOs. 132-148 and 26908-2715, fused to one or more proteins or their domains. For example, in some embodiments, the CasX variant protein is fused to a protein (or its domain) that inhibits transcription, modifies a target nucleic acid sequence, or modifies a nucleic acid-related polypeptide (e.g., histone modification).
[0184] In some embodiments, a heterologous polypeptide (or heterologous amino acid, such as a cysteine residue or a non-natural amino acid) can be inserted at one or more positions within the CasX protein to generate a CasX fusion protein. In other embodiments, a cysteine residue can be inserted at one or more positions within the CasX protein, followed by the conjugation of the heterologous polypeptide described below. In some alternative embodiments, the heterologous polypeptide or heterologous amino acid may be added to the N-terminus or C-terminus of the CasX variant protein. In other embodiments, the heterologous polypeptide or heterologous amino acid may be internally inserted into the sequence of the CasX protein.
[0185] In some embodiments, the CasX variant fusion protein retains RNA-induced sequence-specific target nucleic acid binding and cleavage activity. In some cases, the CasX variant fusion protein has (retains) 50% or more of the activity (e.g., cleavage activity and / or binding activity) of the corresponding CasX variant protein without heterologous protein insertion. In some cases, the CasX variant fusion protein retains at least about 60%, or at least about 70%, at least about 80%, at least about 90%, at least about 92%, at least about 95%, at least about 98%, or at least about 100% of the activity (e.g., cleavage activity and / or binding activity) of the corresponding CasX protein without heterologous protein insertion.
[0186] In some cases, CasX variant fusion proteins retain (have) target nucleic acid binding activity compared to the activity of CasX proteins without inserted heterologous amino acids or heterologous polypeptides. In some cases, CasX variant fusion proteins retain at least about 60%, or at least about 70%, at least about 80%, at least about 90%, at least about 92%, at least about 95%, at least about 98%, or at least about 100% of the binding activity of the corresponding CasX protein without heterologous protein insertion.
[0187] In some cases, CasX variant fusion proteins retain (have) target nucleic acid binding and / or cleavage activity compared to the activity of the parental CasX protein without the inserted heterologous amino acid or heterologous polypeptide. For example, in some cases, CasX variant fusion proteins have (retain) 50% or more of the binding and / or cleavage activity of the corresponding parental CasX protein (CasX protein without the insertion). For example, in some cases, CasX variant fusion proteins have (retain) 60% or more (70% or more, 80% or more, 90% or more, 92% or more, 95% or more, 98% or more, or 100%) of the binding and / or cleavage activity of the corresponding parental CasX protein (CasX protein without the insertion). Methods for measuring the cleavage and / or binding activity of CasX proteins and / or CasX fusion proteins are known to those skilled in the art, and any convenient method can be used.
[0188] Various heterologous polypeptides are suitable for inclusion in the reference CasX or CasX variant fusion protein of this disclosure. In some cases, the fusion partner can regulate the transcription of target DNA (e.g., inhibit transcription, increase transcription). For example, in some cases, the fusion partner is a transcription-repressing protein (or protein-derived domain) (e.g., a protein that functions via transcriptional repressors, the recruitment of transcription-repressing proteins, modification of target DNA such as methylation, recruitment of DNA modifiers, regulation of histones associated with target DNA, or the recruitment of histone modifiers such as those that modify histone acetylation and / or methylation).
[0189] In some cases, the fusion partner is a protein (or protein-derived domain) that increases transcription (e.g., a protein that acts via transcription activators, recruitment of transcription activator proteins, modification of target DNA such as demethylation, recruitment of DNA modifiers, regulation of histones associated with target DNA, or recruitment of histone modifiers such as those that modify histone acetylation and / or methylation). In some cases, the fusion partner has enzymatic activity that modifies the target nucleic acid sequence, such as nuclease activity, methyltransferase activity, demethylase activity, DNA repair activity, DNA damage activity, deamination activity, dismutase activity, alkylation activity, depurination activity, oxidation activity, pyrimidine dimerization activity, integrase activity, transposase activity, recombinase activity, polymerase activity, ligase activity, helicase activity, photolyase activity, or glycosylase activity. In some embodiments, the CasX variant comprises one of SEQ ID NOs: 36-99, 101-148, or 26908-27154, or one of SEQ ID NOs: 59, 72-99, 101-148, or 26908-27154, or one of SEQ ID NOs: 132-148 or 26908-27154, and a polypeptide having methyltransferase activity, demethylase activity, acetyltransferase activity, deacetylase activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitination activity, adenylation activity, deadenylation activity, SUMOylation activity, deSUMOylation activity, ribosylation activity, deribosylation activity, myristoylation activity, or demyristoylation activity.
[0190] In some embodiments, the CasX variant comprises one of SEQ ID NOs: 36-99, 101-148, and 26908-27154, or one of SEQ ID NOs: 59, 72-99, 101-148, or 26908-27154, or one of SEQ ID NOs: 132-148 or 26908-27154, and a fusion partner having enzymatic activity that modifies a polypeptide (e.g., histone) related to a target nucleic acid (e.g., methyltransferase activity, demethylase activity, acetyltransferase activity, deacetylase activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitination activity, adenylation activity, deadenylation activity, SUMOylation activity, deSUMOylation activity, ribosylation activity, deribosylation activity, myristoylation activity, or demyristoylation activity). Examples of proteins (or fragments thereof) that can be used as fusion partners to increase transcription include transcription activators (e.g., VP16, VP64, VP48, VP160, p65 subdomain (e.g., derived from NFκB), and the activation domain and / or TAL activation domain of EDLL (e.g., active in plants); histone lysine methyltransferases (e.g., SET1A, SET1B, MLL1-5, ASH1, SYMD2, NSD1, etc.); histone lysine demethylases (e.g., JHDM2a / b, UTX, JMJD3, etc.); histone acetyltransferases (e.g., GCN5, PCAF, CBP, p300, TAF1, TIP60 / PLIP, MOZ / MYST3, MORF / MYST4, SRC1, ACTR, P160, CLOCK, etc.); and DNA demethylases (e.g., Ten-Eleven) Examples of translocation (TET) dioxygenases include, but are not limited to, translocation (TET) dioxygenase 1 (TET1CD), TET1, DME, DML1, DML2, ROS1, etc.
[0191] Examples of proteins (or fragments thereof) that can be used as fusion partners to reduce transcription include, but are not limited to, the following: transcriptional repressors (e.g., Kruppel-associated boxes (KRAB or SKD)); KOX1 repressor domains; Mad mSIN3 interaction domains (SID); ERF repressor domains (ERF repressor Histone lysine methyltransferases (e.g., ERD domain), SRDX repression domain (e.g., repression in plants); histone lysine demethylases (e.g., JMJD2A / JHDM3A, JMJD2B, JMJD2C / GASC1, JMJD2D, JARID1A / RBP2, JARID1B / PLU-1, JARID1C / SMCX, JARID1D / SMCY, etc.); histone lysine deacetylases (e.g., HDAC1, HDAC2, HDAC3, HDAC8, HDAC4, HDAC5, HDAC7, HDAC9, SIRT1, SIRT2, HDAC11, etc.); DNA methylases (e.g., HhaI DNA m5c-methyltransferase (M.HhaI), DNA methyltransferase 1 (DNMT1), DNA methyltransferase 3a (DNMT3a), DNA methyltransferase 3b (DNMT3b), METI, DRM3 (plant), ZMET2, CMT1, CMT2 (plant), etc.; and peripheral recruiting elements (e.g., Lamin A, Lamin B, etc.).
[0192] In some cases, the fusion partner for a CasX variant possesses enzymatic activity that modifies the target nucleic acid sequence (e.g., ssRNA, dsRNA, ssDNA, dsDNA). Examples of enzymatic activity that may be provided by a fusion partner include nuclease activity (e.g., provided by restriction enzymes (e.g., FokI nuclease)), methyltransferase activity (e.g., provided by methyltransferases (e.g., HhaI DNA m5c-methyltransferase (M.Hhal), DNA methyltransferase 1 (DNMT1), DNA methyltransferase 3a (DNMT3a), DNA methyltransferase 3b (DNMT3b), METI, DRM3 (plant), ZMET2, CMT1, CMT2 (plant), etc.)), demethylase activity (e.g., demethylase (e.g., Ten-Eleven Translocation (TET) dioxygenase 1 (TET1) Examples of DNA repair activity, DNA damage activity, deamination activity (provided by, for example, CD, TET1, DME, DML1, DML2, ROS1, etc.), deaminase activity (provided by, for example, deaminase enzymes such as cytosine deaminase enzymes, such as APOBEC proteins such as rat APOBEC1), dismutase activity, alkylation activity, depurination activity, oxidative activity, pyrimidine dimerization activity, integrase activity (provided by, for example, integrase and / or resolverase enzymes such as, for example, Gin invertase enzymes such as GinH106Y, such as human immunodeficiency virus type 1 integrase (IN), Tn3 resolverase, etc.), transposase activity, recombinase activity (provided by, for example, recombinase enzymes such as, for example, the catalytic domain of Gin recombinase), polymerase activity, ligase activity, helicase activity, photolyase activity, and glycosylase activity.
[0193] In some cases, the CasX variant protein of this disclosure is fused to a polypeptide selected from domains for increasing transcription (e.g., VP16 domain, VP64 domain), domains for decreasing transcription (e.g., KRAB domain (e.g., derived from Kox1 protein)), core catalytic domains of histone acetyltransferases (e.g., histone acetyltransferase p300), proteins / domains that provide a detectable signal (e.g., fluorescent proteins such as GFP), nuclease domains (e.g., FokI nuclease), or base editors (e.g., cytidine deaminases such as APOBEC1).
[0194] In some embodiments, the CasX variant comprises one of SEQ ID NOs: 36-99, 101-148, or 26908-27154, or one of SEQ ID NOs: 59, 72-99, 101-148, or 26908-27154, or one of SEQ ID NOs: 132-148 or 26908-27154, and a fusion partner (e.g., histone, RNA-binding protein, DNA-binding protein, etc.) having enzymatic activity to modify proteins related to a target nucleic acid (e.g., ssRNA, dsRNA, ssDNA, dsDNA).Examples of enzyme activity that can be provided by fusion partners (modifying proteins related to target nucleic acids) include histone methyltransferase activity (e.g., histone methyltransferase (HMT) (e.g., suppressor of variegated 3-9 homolog 1 (SUV39H1, also known as KMT1A), true chromatin histone lysine methyltransferase 2 (G9A, also known as KMT1C and EHMT2), SUV39H2, ESET / SETDB) Demethylase activity (e.g., histone demethylases (e.g., lysine demethylase 1A (also known as KDM1A, LSD1), JHDM2a / b, JMJD2A / JHDM3A, JMJD2B, JMJD2C / GASC1, JMJD2D, JARID1A / RBP2, JARID1B / PLU-1, JARID1C / SMCX, JARID1D / SMCY, UTX, JMJD3, etc.)), acetyltransferase activity (e.g., histone acetyltransferases (e.g., catalytic core / fragment of human acetyltransferase p300, GCN 5. Activities provided by (PCAF, CBP, TAF1, TIP60 / PLIP, MOZ / MYST3, MORF / MYST4, HB01 / MYST2, HMOF / MYST1, SRC1, ACTR, P160, CLOCK, etc.), deacetylase activity (e.g., histone deacetylases provided by (e.g., HDAC1, HDAC2, HDAC3, HDAC8, HDAC4, HDAC5, HDAC7, HDAC9, SIRT1, SIRT2, HDAC11, etc.)), kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitination activity, adenylation activity, deadenylation activity, SUMOylation activity, deSUMOylation activity, ribosylation activity, deribosylation activity, myristoylation activity, and demyristoylation activity are examples of, but are not limited to, these activities.
[0195] Further examples of suitable fusion partners for CasX variants include (i) a dihydrofolate reductase (DHFR) destabilization domain (for example, to generate a chemically controllable RNA-inducible polypeptide of a target or a conditionally active RNA-inducible polypeptide), and (ii) a chloroplast-transfer peptide. In some embodiments, the CasX variant is one of SEQ ID NOs: 36-99, 101-148, or 26908-27154, or one of SEQ ID NOs: 59, 72-99, 101-148, or 26908-27154, or one of SEQ ID NOs: 132-148 or 26908-27154, or the sequence in Table 4 and a chloroplast-transfer peptide (MASMISSSAVTTVSRASRGQSAAMAPFGGLKSMTGFPVRKVNTDITSITSNGGR VKCMQVWPPIGKKKFETLSYLPPLTRDSRA (SEQ ID NO: 154); MASMISSSAVTTVSRASRGQSAAMAPFGGLKSMTGFPVRKVNTDITSITSNGGRVKS (SEQ ID NO: 155); MASSMLSSATMVASPAQATMVAPFNGLKSSAAFPATRKANNDITSITSNGGRVNCMQV WPPIEKKKFETLSYLPDLTDSGGRVNC(Sequence ID 156);MAQVSRICNGVQNPSLISNLSKSSQRKSPLSVSLKTQQHPRAYPISSSWGLKKSGMTLIG SELRPLKVMSSVSTAC(Sequence ID 157);MAQVSRICNGVWNPSLISNLSKSSQRKSPLSVSLKTQQHPRAYPISSSWGLKKSGMTLIG SELRPLKVMSSVSTAC(Sequence ID 158);MAQINNMAQGIQTLNPNSNFHKPQVPKSSSFLVFGSKKLKNSANSMLVLKKDSIFMQLF CSFRISASVATAC(Sequence ID 159);MAALVTSQLATSGTVLSVTDRFRRPGFQGLRPRNPADAALGMRTVGASAAPKQSRKPH RFDRRCLSMVV(Sequence ID 160);MAALTTSQLATSATGFGIADRSAPSSLLRHGFQGLKPRSPAGGDATSLSVTTSARATPKQ QRSVQRGSRRFPSVVVC (Sequence ID 161);MASSVLSSAAVATRSNVAQANMVAPFTGLKSAASFPVSRKQNLDITSIASNGGRVQC (Sequence ID 162);MESLAATSVFAPSRVAVPAARALVRAGTVVPTRRTSSTSGTSGVKCSAAVTPQASPVIS RSAAAA (Sequence ID 163); and MGAAATSMQSLKFSNRLVPPSRRLSPVPNNVTCNNLPKSAAPVRTVKCCASSWNSTINGAAATTNGASAASS (Sequence ID 164), among others.
[0196] In some cases, the CasX variant protein of this disclosure may include an endosomal escape peptide. In some cases, the endosomal escape polypeptide includes the amino acid sequence GLFXALLXLLXSLWXLLLXA (SEQ ID NO: 165), where each X is independently selected from lysine, histidine, and arginine. In some cases, the endosomal escape polypeptide includes the amino acid sequence GLFHALLHLLHSLWHLLLHA (SEQ ID NO: 166) or HHHHHHHHH (SEQ ID NO: 167).
[0197] When targeting ssRNA target nucleic acid sequences, non-exclusive examples of fusion partners for use with CasX variant proteins include (but are not limited to) splicing factors (e.g., RS domains); protein translation components (e.g., translation initiation, elongation, and / or termination factors; e.g., eIF4G); RNA methylases; RNA editing enzymes (e.g., RNA deaminases (e.g., adenosine deaminase acting on RNA, ADAR)), including editing enzymes A through I and / or C through U); helicases; and RNA-binding proteins. It is understood that heterologous polypeptides may include whole proteins or, in some cases, protein fragments (e.g., functional domains).
[0198] In some embodiments, the CasX variant includes one of SEQ ID NOs: 36-99, 101-148, or 26908-27154, or one of SEQ ID NOs: 59, 72-99, 101-148, or 26908-27154, or one of SEQ ID NOs: 132-148 or 26908-27154, and a fusion partner of any domain capable of interacting with ssRNA (for the purposes of this disclosure, this includes intramolecular and / or intermolecular secondary structures, e.g., double-stranded RNA double helix, e.g., hairpin, stem-loop, etc.). Examples of effector domains, whether transient or irreversible, direct or indirect, include, but are not limited to, effector domains selected from the group including: endonucleases (e.g., RNase III, CRR22 DYW domain, Dicer), and protein-derived PINs (PilT) such as SMG5 and SMG6. N-terminal domains); proteins and protein domains involved in stimulating RNA cleavage (e.g., CPSF, CstF, CFIm, and CFIIm); exonucleases (e.g., XRN-1 or exonuclease T); deadenylases (e.g., HNT3); proteins and protein domains involved in nonsense-mediated RNA degradation (e.g., UPF1, UPF2, UPF3, UPF3b, RNPSI, Y14, DEK, REF2, and SRm160); proteins and protein domains involved in RNA stabilization (e.g., PABP); proteins and proteins involved in translational repression Domains (e.g., Ago2 and Ago4); proteins and protein domains involved in translational stimulation (e.g., Staufen); proteins and protein domains involved in translational regulation (e.g., those capable of regulating translation) (e.g., translation factors (e.g., initiation factors, elongation factors, termination factors, etc., e.g., eIF4G)); proteins and protein domains involved in RNA polyadenylation (e.g., PAP1, GLD-2, and Star-PAP); proteins and protein domains involved in RNA polyuridinylation (e.g., CIDl and terminal uridylate transferases);Proteins and protein domains involved in RNA localization (e.g., IMP1, ZBP1, She2p, She3p, and Bicaudal-D); proteins and protein domains involved in nuclear retention of RNA (e.g., Rrp6); proteins and protein domains involved in nuclear export of RNA (e.g., TAP, NXF1, THO, TREX, REF, and Aly); proteins and protein domains involved in repression of RNA splicing (e.g., PTB, Sam68, and hnRNP Al); proteins and protein domains involved in stimulation of RNA splicing (e.g., serine / arginine-rich (SR) domains); proteins and protein domains involved in the reduction of transcription efficiency (e.g., FUS(TLS)); proteins and protein domains involved in transcription stimulation (e.g., CDK7 and HIV Tat). Alternatively, effector domains can be selected from the following group: endonucleases; proteins and protein domains capable of stimulating RNA cleavage; exonucleases; deadenylases; proteins and protein domains with nonsense-mediated RNA degradation activity; proteins and protein domains capable of stabilizing RNA; proteins and protein domains capable of repressing translation; proteins and protein domains capable of stimulating translation; proteins and protein domains capable of regulating translation (e.g., translation factors (e.g., initiation factors, elongation factors, termination factors, etc., e.g., eIF4G)); proteins and protein domains capable of polyadenylation of RNA; proteins and protein domains capable of polyuridinylation of RNA; proteins and protein domains with RNA localization activity; proteins and protein domains capable of retaining RNA in the nucleus; proteins and protein domains with RNA nuclear export activity; proteins and protein domains capable of repressing RNA splicing; proteins and protein domains capable of stimulating RNA splicing; proteins and protein domains capable of reducing transcription efficiency;Furthermore, proteins and protein domains capable of stimulating transcription. Another preferred heterologous polypeptide is a PUF RNA-binding domain, which is described in detail in International Publication No. 2012 / 068627 (the entire publication of which is incorporated herein by reference).
[0199] Some RNA splicing factors that can be used (whole or in fragments) as fusion partners with CasX variants have a modular configuration with separate sequence-specific RNA-binding modules and splicing effector domains. For example, members of the serine / arginine-rich (SR) protein family contain an N-terminal RNA recognition motif (RRM) that binds to an exonic splicing enhancer (ESE) in premRNA and a C-terminal RS domain that promotes exon inclusion. Another example is the hnRNP protein hnRNP Al, which binds to an exonic splicing silencer (ESS) via its RRM domain and inhibits exon inclusion via its C-terminal glycine-rich domain. Some splicing factors can modulate the selective use of splice sites (ss) by binding to regulatory sequences between two selective sites. For example, ASF / SF2 can recognize ESEs and promote the use of proximal intron sites, while hnRNPAl can bind to ESSs and shift splicing to the use of distal intron sites. One application of such factors is the creation of ESFs that regulate alternative splicing of endogenous genes (particularly disease-associated genes). For example, Bcl-x premRNA produces two splicing isoforms, each having two alternative 5' splice sites encoding proteins with opposite functions. The long splicing isoform, Bcl-xL, is a potent apoptosis inhibitor expressed in long-lived postmittal cells, upregulated in many cancer cells, and protects cells from apoptotic signaling. The short isoform, Bcl-xS, is a pro-apoptotic isoform and is expressed at high levels in cells with high metabolic rates (e.g., developing lymphocytes).The ratio of two Bcl-x splicing isoforms is regulated by multiple cc elements located either in the core exon region or the exon extension region (i.e., between two alternative 5' splice sites). For further examples, see International Publication 2010 / 075303, which is incorporated herein by reference in its entirety.
[0200] Further suitable fusion partners for use with CasX variants include, but are not limited to, boundary element proteins (e.g., CTCF) (or their fragments), peripheral recruitment-providing proteins and their fragments (e.g., lamin A, lamin B, etc.), and protein docking elements (e.g., FKBP / FRB, Pill / Abyl, etc.).
[0201] In some cases, a heterologous polypeptide (fusion partner) for use with a CasX variant provides intracellular localization. That is, the heterologous polypeptide includes intracellular localization sequences (e.g., a nuclear localization signal (NLS) for targeting the nucleus, a sequence for retaining the fusion protein outside the nucleus (e.g., a nuclear export sequence (NES)), a sequence for retaining the fusion protein in the cytoplasm, a mitochondrial localization signal for targeting mitochondria, a chloroplast localization signal for targeting chloroplasts, an ER retention signal, etc.). In some embodiments, the target RNA guide polypeptide or conditionally activated RNA guide polypeptide and / or the target CasX fusion protein does not contain an NLS so that the protein does not target the nucleus (this may be advantageous, for example, when the target nucleic acid sequence is RNA present in the cytosol). In some embodiments, the fusion partner can provide a tag for facilitating tracking and / or purification (i.e., a detectable label for heterologous polypeptides) (e.g., fluorescent proteins, e.g., green fluorescent protein (GFP), yellow fluorescent protein (YFP), red fluorescent protein (RFP), cyan fluorescent protein (CFP), mCherry, tdTomato, etc.; histidine tags, e.g., 6×His tag; hemagglutinin (HA) tag; FLAG tag; Myc tag, etc.).
[0202]
[0203] In general, NLS (or multiple NLSs) are potent enough to drive the accumulation of CasX variant fusion proteins expressed in the nuclei of eukaryotic cells. Detection of nuclear accumulation can be carried out by any suitable technique. For example, a detectable marker can be fused to the CasX variant fusion protein so that its intracellular location can be visualized. The cell nucleus can also be isolated from the cell, and its contents can then be analyzed by any suitable process for detecting the protein (e.g., immunohistochemistry, Western blotting, or enzyme activity assay). Nuclear accumulation can also be determined indirectly.
[0204] In some cases, the CasX variant fusion protein contains a "protein transduction domain" or PTD (also known as a CPP cell permeable peptide), which refers to a protein, polynucleotide, carbohydrate, or organic or inorganic compound that facilitates translocation of lipid bilayers, micelles, cell membranes, organelle membranes, or vesicle membranes. A PTD conjugated to another molecule (which can range from small polar molecules to large macromolecules and / or nanoparticles) facilitates the molecule's translocation across membranes, e.g., from the extracellular space to the intracellular space, or from the cytosol to the organelle. In some embodiments, the PTD is covalently bonded to the amino terminus of the CasX variant fusion protein. In some embodiments, the PTD is covalently bonded to the carboxyl terminus of the CasX variant fusion protein. In some cases, the PTD is inserted into the sequence of the CasX variant fusion protein at a preferred insertion site. In some cases, the CasX variant fusion protein contains (is conjugated to, is fused to) one or more PTDs (e.g., two or more, three or more, four or more PTDs). In some cases, PTDs include one or more nuclear localization signals (NLSs).Examples of PTDs include, but are not limited to, the following: peptide transduction domains of HIV TAT containing YGRKKRRQRRR (SEQ ID NO: 205) and RKKRRQRR (SEQ ID NO: 206); YARAAARQARA (SEQ ID NO: 207); THRLPRRRRRR (SEQ ID NO: 208); and GGRRARRRRRR (SEQ ID NO: 209); a sufficient number of arginine residues to direct cell entry (e.g., 3, 4, 5, 6, 7, 8, 9, 10, or polyarginine sequences containing 10-50 arginine residues (SEQ ID NO: 26793); VP22 domain (Zender et al. (2002) Cancer Gene Ther. 9(6):489-96); Drosophila Antennapedia protein transduction domain (Noguchi et al. (2003) Diabetes 52(7):1732-1737); truncated human calcitonin peptide (Trehin et al. al. (2004) Pharm. Research 21:1248-1256); polylysine (Wender et al. (2000) Proc. Natl. Acad. Sci. USA 97:13003-13008); RRQRRTSKLMKR (SEQ ID NO: 210); transportan GWTLNSAGYLLGKINLKALAALAKKIL (SEQ ID NO: 211); KALAWEAKLAKALAKALAKHLAKALAKALKCEA (SEQ ID NO: 212); and RQIKIWFQNRRMKWKK (SEQ ID NO: 213). In some embodiments, PTD is activatable CPP (ACPP) (Aguilera et al. (2009) Integr Biol(Camb)June;1(5-6):371-381). ACPP contains a polycationic CPP (e.g., Arg9 or "R9") linked to a matching polyanion (e.g., Glu9 or "E9") via a cleavage-competent linker, which reduces the net charge to nearly zero, thereby inhibiting adhesion to and uptake of cells. When the linker is cleaved, the polyanion is released, locally unmasking polyarginine and its inherent adhesiveness, and thus "activating" the ACPP to pass through the membrane.
[0205] In some embodiments, a CasX variant fusion protein for use in a system may include a CasX protein linked to heterologous amino acids or heterologous polypeptide (heterologous amino acid sequence) inserted internally via a linker polypeptide (e.g., one or more linker polypeptides). In some embodiments, the CasX variant fusion protein may be linked to a heterologous polypeptide (fusion partner) via a linker polypeptide (e.g., one or more linker polypeptides) at its C-terminus and / or N-terminus. The linker polypeptide may have any of a variety of amino acid sequences. Proteins may be linked by spacer peptides of a generally flexible nature, but other chemical bonds are not excluded. Preferred linkers include polypeptides of 4 to 40 amino acids in length, or 4 to 25 amino acids in length. These linkers are generally produced by coupling proteins using oligonucleotides encoding synthetic linkers. Peptide linkers with some degree of flexibility may be used. The linked peptide may have substantially any amino acid sequence, considering that the preferred linker will generally have a sequence that results in a flexible peptide. The use of small amino acids such as glycine and alanine is useful in the production of flexible peptides. The generation of such sequences is routine for those skilled in the art. Various different linkers are commercially available and are considered suitable for use.Examples of linker polypeptides include RS, (G)n (SEQ ID NO: 27212), (GS)n (SEQ ID NO: 27213), (GSGGS)n (SEQ ID NO: 214), (GGSGGS)n (SEQ ID NO: 215), (GGGS)n (SEQ ID NO: 216), (where n is an integer from 1 to 5)GGSG (SEQ ID NO: 217), GGSGG (SEQ ID NO: 218), GSGSG (SEQ ID NO: 219), GSGGG (SEQ ID NO: 220), GGGSG (SEQ ID NO: 221), GSSSG (SEQ ID NO: 216) Examples of peptides selected from the group consisting of 222), GPGP (SEQ ID NO: 223), GGP, PPP, PPAPPA (SEQ ID NO: 224), PPPG (SEQ ID NO: 27214), PPPGPPP (SEQ ID NO: 225), PPP(GGGS)n (SEQ ID NO: 27215), (GGGS)nPPP (SEQ ID NO: 27216), AEAAAKEAAAKEAAAKA (SEQ ID NO: 27217), and TPPKTKRKVEFE (SEQ ID NO: 27218) (wherein n is 1 to 5). Those skilled in the art will recognize that the design of peptides conjugated to any of the above elements may include all or partly flexible linkers so that the linker may include one or more parts that confer both flexible and less flexible structures.
[0206] System and method for modifying the V.BCL11A gene The CRISPR proteins, guide nucleic acids, and their variants provided herein are useful for a variety of applications, including therapeutic, diagnostic, and research purposes. In some embodiments, a programmable CasX:gRNA system is provided herein for carrying out the methods of this disclosure for gene editing. The programmable nature of the systems provided herein allows for precise targeting to achieve desired modifications in one or more regions of a given interest within the BCL11A gene target nucleic acid. Various strategies and methods can be used to modify the target nucleic acid sequence in cells using the systems provided herein. As used herein, “modification” includes, but is not limited to, cleavage, nicking, editing, deletion, knockout, knockdown, mutagenesis, repair, and exon skipping. Depending on the system components used, the editing event may be a cleavage event followed by the introduction of random insertions or deletions (indels) or other mutations (e.g., substitutions, duplications, or inversions of one or more nucleotides) by utilizing an inaccurate non-homologous DNA end-joining (NHEJ) repair pathway, which may produce, for example, frameshift mutations. Alternatively, the editing event may be a cleavage event followed by homologous recombination repair (HDR), homology-independent targeted integration (HITI), microhomology-mediated end joining (MMEJ), single-strand annealing (SSA), or base excision repair (BER), resulting in modification of the target nucleic acid sequence.
[0207] In some embodiments of this method, the modified BCL11A gene includes a sequence corresponding to a polynucleotide encoding all or part of the sequence of Sequence ID No. 100, or includes a polynucleotide sequence spanning all or part of chr2 60450520-60554467 (GRCh38 / hg38 Ensembl 100) of the human genome on chromosome 2. In other embodiments of this method, the modified target nucleic acid sequence includes a region of the BCL11A gene encoding the BCL11A protein, a BCL11A regulatory element, a non-coding region of the BCL11A gene, or overlapping portions thereof. In certain embodiments of this method, the modified target nucleic acid sequence includes a GATA1 binding motif sequence or its complement.
[0208] In some embodiments, this disclosure provides a method for modifying a BCL11A target nucleic acid in cells, and provides a method comprising introducing a class 2, type V CRISPR system into the cells. In some embodiments of this method, the modified cells are autologous to the subject to which the cells are administered. In other embodiments, the modified cells are homogeneous to the subject to which the cells are administered. Accordingly, the systems and methods described herein can be used to manipulate various cells associated with diseases in which mutations are present in the β-globin gene, such as sickle cell disease and hemoglobin disorders including α- and β-thalassemia. Thus, this approach can be used to modify cells for application in subjects having hemoglobin disorders such as sickle cell disease and α- and β-thalassemia, but is not limited to these.
[0209] In some embodiments, the Disclosure provides a method for modifying a BCL11A target nucleic acid in a cell, the method comprising introducing into a cell: i) a CasX:gRNA system comprising CasX and gRNA as described in any one of the embodiments described herein; ii) a CasX:gRNA system comprising CasX, gRNA, and a donor template as described in any one of the embodiments described herein; iii) a nucleic acid encoding CasX and gRNA and optionally comprising a donor template; iv) a vector comprising the nucleic acid of (iii) above; v) an XDP comprising the CasX:gRNA system as described in any one of the embodiments described herein; or vi) a combination of two or more of (i) to (v), wherein the target nucleic acid sequence of the cell is modified by the CasX protein and optionally by the donor template. In some embodiments, the vector is an AAV vector. In some embodiments, the present disclosure provides a CasX:gRNA system for use in a method for modifying the BCL11A gene in cells, the method using a CasX variant selected from the group consisting of SEQ ID NOs. 36-99, 101-148, and 26908-27154, or a CasX variant selected from the group consisting of SEQ ID NOs. 59, 72-99, 101-148, and 26908-27154, or a CasX variant selected from the group consisting of SEQ ID NOs. 132-148 and 26908-27154, or a CasX variant that is at least 60% identical, at least 70% identical, at least 80% identical, at least 81% identical, at least 82% identical, at least 83% identical, and at least All contain variant sequences that are 84% identical, at least 85% identical, at least 86% identical, at least 86% identical, at least 87% identical, at least 88% identical, at least 89% identical, at least 89% identical, at least 90% identical, at least 91% identical, at least 92% identical, at least 93% identical, at least 94% identical, at least 95% identical, at least 96% identical, at least 97% identical, at least 98% identical, at least 99% identical, or at least 99.5% identical, and the gRNA scaffold is a sequence selected from the group consisting of SEQ ID NOs. 2101-2285, 26794-26839, and 27219-27265 shown in Table 3, or SEQ ID NOs. 2281-2285,Sequences selected from the group consisting of 26794-26839 and 27219-27265, or sequences that are at least 65% identical, at least 70% identical, at least 75% identical, at least 80% identical, at least 81% identical, at least 82% identical, at least 83% identical, at least 84% identical, at least 85% identical, at least 86% identical, at least 86% identical, at least 87% identical, at least 88% identical, at least 89% identical, at least 89% identical, at least 90% identical, at least 91% identical, at least 92% identical, at least 9 The gRNA contains a sequence that is 3% identical, at least 94% identical, at least 95% identical, at least 96% identical, at least 97% identical, at least 98% identical, at least 99% identical, or at least 99.5% identical, and contains a targeting sequence selected from the group consisting of SEQ ID NOs. 272-2100 or 2286-26789, or a sequence that is at least 65% identical, at least 70% identical, at least 75% identical, at least 80% identical, at least 85% identical, at least 90% identical, or at least 95% identical, and has 15-20 nucleotides. In certain embodiments, the targeting sequence of the gRNA is complementary to a sequence in the GATA1 binding motif sequence, or a sequence located at the 5' or 3' end of the GATA1 binding motif sequence, and can therefore hybridize with it. In one embodiment, the targeting sequence of the gRNA is UGGAGCCUGUGAUAAAAGCA (SEQ ID NO: 22), which hybridizes with the BCL11A GATA1 erythrocyte-specific enhancer binding site sequence, or a sequence having at least 90% or at least 95% sequence identity with it. In another embodiment, the targeting sequence of the gRNA is UGCUUUUAUCACAGGCUCCA (SEQ ID NO: 23), which hybridizes with a sequence complementary to the reverse complement of the BCL11A GATA1 erythrocyte-specific enhancer binding site sequence, or a sequence having at least 90% or at least 95% sequence identity with it. In another specific embodiment, the targeting sequence of the gRNA is complementary to a sequence in the promoter of the BCL11A gene and can therefore hybridize with it. In one embodiment of the method, CasX and gRNA areThey associate together in the ribonucleoprotein complex (RNP). In some embodiments of the method for modifying the BCL11A target nucleic acid sequence in cells, the modification includes introducing a single-strand break in the target nucleic acid sequence. In other embodiments of the method, the modification includes introducing a double-strand break in the target nucleic acid sequence. In some embodiments of the method, the modification includes introducing one or more nucleotide insertions, deletions, substitutions, duplications, or inversions in the target nucleic acid sequence. As described herein, the CasX variant introducing a double-strand break in the target nucleic acid generates double-strand breaks within 18-26 nucleotides on the 5' end of the PAM site on the target strand and within 10-18 nucleotides on the 3' end of the non-target strand. Therefore, in some embodiments, the modifications obtained by the method may result in one or more random insertions or deletions (indels), substitutions, duplications, or inversions of nucleotides in those regions via the non-homologous DNA end-joining (NHEJ) repair mechanism.
[0210] In other embodiments of the method for modifying the BCL11A target nucleic acid sequence in cells, the method comprises contacting the target nucleic acid sequence with a CasX:gRNA system having first and second or more gRNAs (for example, the targeting sequence of the second gRNA is complementary to a sequence located on the 5' or 3' side of the GATA1 binding site) that target different or overlapping portions of the BCL11A gene, the CasX protein introducing multiple cleavages into the target nucleic acid, resulting in persistent indels or mutations or excision of the GATA1 binding motif sequence in the target nucleic acid, which involves the regulation or modification of the corresponding expression or function of the BCL11A gene product as described herein, thereby generating edited cells. In some of the aforementioned cases, the multiple gRNAs target the 5' and 3' positions of the GATA1 binding motif sequence of the BCL11A gene such that some or all of the GATA1 binding motif sequence is excised from the target gene between the double cleavage sites targeted by the two gRNAs. It will be understood that the aforementioned embodiments of this method can also be achieved by using a vector containing coding nucleic acid, coding acid, or XDP containing CasX:gRNA system components.
[0211] In some embodiments, the methods of the present disclosure provide a CasX protein and gRNA pair that generate a site-specific double-strand break (DSB) or single-strand break (SSB) within 18-24 nucleotides on the 3' side of the PAM site (for example, if the CasX protein is a nickase capable of cleaving only one strand of the target nucleic acid), which can then be repaired by any of the following: non-homologous end joining (NHEJ), homologous recombination repair (HDR), homology-independent targeted integration (HITI), microhomology-mediated end joining (MMEJ), single-strand annealing (SSA), or base excision repair (BER), thereby introducing a mutation of one or more nucleotides, insertions, deletions, inversions, or duplications compared to the wild-type sequence, with corresponding regulation of expression or modification of function of the BCL11A gene product, thereby generating edited cells.
[0212] In some cases, a CasX:gRNA system for use in a method for modifying the BCL11A gene further comprises a donor template nucleic acid of any embodiment disclosed herein, the donor template being inserted by a host cell homologous recombination repair (HDR) or homology-independent targeted insertion (HITI) repair mechanism. Accordingly, in some cases, a method provided herein includes contacting the BCL11A gene with a donor template by introducing the donor template (in vitro or in vivo into the cell), the donor template, a portion of the donor template, a copy of the donor template, or a portion of a copy of the donor template being incorporated into the BCL11A gene to replace a portion of the BCL11A gene. The donor template may be a short single-stranded or double-stranded oligonucleotide, or a long single-stranded or double-stranded oligonucleotide. In some embodiments, the donor template comprises at least a portion of the BCL11A gene, the portion of the BCL11A gene being selected from the group consisting of BCL11A exons, BCL11A introns, BCL11A intron-exon junctions, BCL11A regulatory elements, or combinations thereof. In some embodiments, the present disclosure provides a donor template for use in targeting or disrupting a transcription activator GATA1 binding site in a BCL11A target sequence, wherein the donor template comprises a sequence that is non-homologous to a region of DNA in or near the GATA1 site in the BCL11A gene, adjacent to two homology regions ("homologous arms") to the 5' and 3' sides of the cleavage site, such that a repair mechanism between the target DNA region and two adjacent sequences results in insertion of the donor template in the target region to facilitate HDR insertion. The donor template may contain one or more single nucleotide changes, insertions, deletions, inversions, or rearrangements with respect to the genome sequence, but must have sufficient homology to the target nucleic acid sequence to support its incorporation into the target nucleic acid, which can result in a frameshift or other mutation that causes the BCL11A protein to be not expressed (knockout) or to be expressed at a lower level (knockdown).The exogenous donor template inserted by HITI may be of any length, e.g., a relatively short sequence of 10–50 nucleotides, or a longer sequence of about 50–1000 nucleotides. Lack of homology may, for example, have less than 20–50% sequence identity and / or lack specific hybridization with low strictness. In other cases, lack of homology may further include the criterion of having less than 5, 6, 7, 8, or 9 bp identity. In some embodiments, the donor template polynucleotide contains at least about 10, at least about 50, at least about 100, or at least about 200, or at least about 300, or at least about 400, or at least about 500, or at least about 600, or at least about 700, or at least about 800, or at least about 900, or at least about 1000, or at least about 10,000, or at least about 15,000 nucleotides. In other embodiments, the donor template comprises at least about 10 to about 15,000 nucleotides, or at least about 100 to about 10,000 nucleotides, or at least about 400 to about 8,000 nucleotides, or at least about 600 to about 5,000 nucleotides, or at least about 1,000 to about 2,000 nucleotides. The donor template sequence may include specific sequence differences compared to the genome sequence, such as restriction sites, nucleotide polymorphisms, selection markers (e.g., drug resistance genes, fluorescent proteins, enzymes, etc.), which may be used to assess the success of donor nucleic acid insertion at cleavage sites, or, in some cases, for other purposes (e.g., to indicate expression at a target genomic locus). Alternatively, these sequence differences may include adjacent recombinant sequences (e.g., FLP sequences, loxP sequences, etc.) which may be activated later to remove marker sequences.
[0213] In some embodiments of methods for modifying BCL11A target nucleic acids in cells in vitro or ex vivo to induce cleavage or any desired modification of the target nucleic acid, the gRNA and / or CasX protein of the Disclosure, and optionally a donor template sequence, whether introduced as nucleic acid or polypeptide, complexed RNP, vector, or XDP, are provided to cells for about 30 minutes to about 24 hours, or any other period of at least about 1 hour, 1.5 hours, 2 hours, 2.5 hours, 3 hours, 3.5 hours, 4 hours, 5 hours, 6 hours, 7 hours, 8 hours, 12 hours, 16 hours, 18 hours, 20 hours, or any other period of about 30 minutes to about 24 hours, which may be repeated at a frequency of about daily to about every 4 days, for example, every 1.5 days, every 2 days, every 3 days, or any other frequency of about daily to about every 4 days. The drug may be delivered to the target cells one or more times, for example, once, twice, three times, or more than three times, and the cells may be incubated with the drug for a certain period of time after each contact event (e.g., 30 minutes to approximately 24 hours). In the in vitro-based method, after the incubation period with CasX and gRNA (and optionally donor template), the medium is replaced with fresh medium and the cells are further cultured.
[0214] In some embodiments of the method for modifying the BCL11A target nucleic acid in cells, the method further comprises contacting the target nucleic acid sequence in the cell with a) an additional CRISPR nuclease and gRNA that targets a different or overlapping portion of the BCL11A target nucleic acid compared to a first gRNA, b) a polynucleotide encoding the additional CRISPR nuclease and gRNA of (a), c) a vector comprising the polynucleotide of (b), or d) an XDP comprising the additional CRISPR nuclease and gRNA of (a), wherein the contact results in a modification of the BCL11A target nucleic acid at a different position in the sequence compared to the first gRNA. In some cases, the additional CRISPR nuclease is a CasX protein having a different sequence from the CasX protein of any of the above claims. In other cases, the additional CRISPR nuclease is selected from the group consisting of Cas9, Cas12a, Cas12b, Cas12c, Cas12d (CasY), Cas12j, Cas12k, Cas13a, Cas13b, Cas13c, Cas13d, CasY, Cas14, Cpfl, C2cl, Csn2, Cas Phi, and their sequence variants, rather than the CasX protein.
[0215] In cases where the modification results in knockdown of the BCL11A gene, BCL11A protein expression is reduced by at least approximately 10%, at least approximately 20%, at least approximately 30%, at least approximately 40%, at least approximately 50%, at least approximately 60%, at least approximately 70%, at least approximately 80%, or at least approximately 90% compared to unmodified cells. In other cases where the modification results in knockout of the BCL11A gene, the target nucleic acid of the cell population is modified so that BCL11A protein expression is undetectable. BCL11A protein expression can be measured by flow cytometry, ELISA, cell-based assays, Western blotting, qRT-PCR, or other methods known in the art, or as described in the examples.
[0216] In some embodiments, the disclosure provides a method for modifying the BCL11A target nucleic acid of a cell population in a subject in vivo. In some embodiments, the modification of the target nucleic acid sequence is performed ex vivo in eukaryotic cells, and the eukaryotic cells are selected from the group consisting of hematopoietic stem cells (HSCs), hematopoietic progenitor cells (HPCs), CD34+ cells, mesenchymal stem cells (MSCs), induced pluripotent stem cells (iPSCs), myeloid common progenitor cells, proerythroblasts, and erythroblasts. In the aforementioned embodiments, the modified cell population may be used in a method of treatment in a subject, and the modified cells are administered to a subject in need, and the subject is selected from the group consisting of mice, rats, pigs, non-human primates, and humans. In some cases, the ex vivo cells are autologous and isolated from the bone marrow or peripheral blood of the subject. In other cases, the ex vivo cells are homogeneous and isolated from the bone marrow or peripheral blood of different subjects. In the treatment method, the modified cells can be administered to the subject by a route of administration selected from intraparenchymal, intravenous, intraarterial, intramuscular, subepidermal, intraarticular, intracardiac, intrapericardial, intravitreous, or subcapsular, or by subcutaneous injection, and can be transplanted to the subject by transplantation, local injection, systemic injection, or a combination thereof. In the embodiments described above, the method results in the survival of the modified cells or their offspring for at least about 1 month, at least about 2 months, at least about 3 months, at least about 4 months, at least about 6 months, at least about 7 months, at least about 8 months, at least about 9 months, at least about 10 months, at least about 11 months, at least about 12 months, at least about 18 months, at least about 2 years, at least about 3 years, at least about 4 years, or at least about 5 years.
[0217] In some embodiments of the method for modifying a target nucleic acid sequence, modifying the BCL11A gene involves binding a CasX:gRNA complex to the target nucleic acid sequence and introducing it into cells as an RNP. In some embodiments, CasX is a catalytically inactive CasX (dCasX) protein that retains the ability to bind to the gRNA and the target nucleic acid sequence. For example, the target nucleic acid sequence includes a BCL11A sequence containing a sequence complementary to the GATA1 binding motif sequence, and the binding of the dCasX:gRNA complex to the target sequence interferes with or represses the transcription of the BCL11A allele. In some embodiments, dCasX includes mutations in residues D672, E769, and / or D935 corresponding to the CasX protein of SEQ ID NO: 1, or D659, E756, and / or D922 corresponding to the CasX protein of SEQ ID NO: 2. In some embodiments described above, the mutations in the CasX variant protein are alanine or glycine substitutions of residues and can be utilized in any of the variants described herein.
[0218] The introduction of a recombinant expression vector containing components of a system embodiment or nucleic acids encoding components into target cells can be performed in vivo, in vitro, or ex vivo. In some embodiments of the method, the vector may be delivered directly to the target host cells. Methods for introducing nucleic acids (e.g., nucleic acids containing a donor polynucleotide sequence, one or more nucleic acids (DNA or RNA) encoding a CasX protein and / or gRNA, or vectors containing the same) into cells are known in the art, and nucleic acids (e.g., expression constructs) can be introduced into cells using any convenient method. Preferred methods include, for example, viral infection, transfection, lipofection, electroporation, calcium phosphate precipitation, polyethyleneimine (PEI)-mediated transfection, DEAE-dextran-mediated transfection, liposome-mediated transfection, particle gun technology, nucleofection, electroporation, direct addition with a cell-permeable CasX protein fused to or recruiting donor DNA, cell compression, calcium phosphate precipitation, direct microinjection, nanoparticle-mediated nucleic acid delivery, and the like. Nucleic acids can be introduced into cells using well-developed, commercially available transfection techniques, such as the use of TransMessenger® reagents from Qiagen, Stemfect® RNA transfection kits from Stemgent, and TransIT®-mRNA transfection kits from Mirus Bio LLC, Lonza nucleofection, Maxagen electroporation, etc. This can be done in any suitable culture medium and under any suitable culture conditions that promote cell viability by introducing a recombinant expression vector containing a sequence encoding the CasX:gRNA system (and optionally a donor sequence) of this disclosure into cells under in vitro conditions. For example, cells can be brought into contact with a vector containing the target nucleic acid (e.g., a recombinant expression vector having a donor template sequence and nucleic acids encoding CasX and gRNA) so that the vector is taken up by the cells.A vector used to deliver nucleic acids encoding gRNA and / or CasX protein to target host cells may contain a suitable promoter to drive the expression (i.e., transcriptional activation) of the nucleic acid of interest. In some cases, the encoding nucleic acid of interest is operably ligated to the promoter. This may include a ubiquitous promoter (e.g., a CMV-β-actin promoter) or an inducible promoter (e.g., a promoter that is active in a specific cell population or that responds to the presence of a drug such as tetracycline or kanamycin). The transcriptional activation is intended to increase transcription above the basal level in the target host cells containing the vector by at least about 10-fold, at least about 100-fold, and more typically at least about 1000-fold. In addition, a vector used to deliver nucleic acids encoding gRNA and / or CasX protein to cells may contain nucleic acid sequences encoding selectable markers in the target cells to identify cells that have taken up the CasX protein and / or gRNA.
[0219] In the case of viral vector delivery, cells can be brought into contact with viral particles containing the target viral expression vector, as well as nucleic acids encoding CasX and gRNA, and optionally a donor template. In some embodiments, the vector is an adeno-associated virus (AAV) vector, and the AAV is selected from AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, AAV44,9, AAV-Rh74, or AAVRh10. In other cases, the AAV is selected from AAV1, AAV2, AAV5, AAV6, AAV7, AAV8, and AAV9, which are efficient for transduction into muscle (Gruntman AM, et al. Gene transfer in skeletal and cardiac muscle using recombinant adeno-associated virus. Curr Protoc Microbiol. 14(14D):3(2013)). Embodiments of AAV vectors are described in more detail below. In other embodiments, the vector is a lentiviral vector. Retroviruses (e.g., lentiviruses) may be suitable for use in the methods of the present disclosure. Commonly used retroviral vectors have "defects," for example, they cannot produce the viral proteins necessary for proliferative infection. Rather, vector replication requires proliferation in a packaging cell line. To generate viral particles containing the nucleic acid of interest, the retroviral nucleic acid containing the nucleic acid is packaged into a viral capsid by a packaging cell line. Different packaging cell lines provide different envelope proteins (ecotropic, amphotropic, or xenotropic) incorporated into the capsid, and this envelope protein determines the specificity or tropism of the viral particles to the cells (ecotropic for mouse and rat; amphotropic for most mammalian cell types, including human, canine, and mouse; xenotropic for most mammalian cell types, excluding mouse cells). A suitable packaging cell line may be used to ensure that cells are targeted by the packaged viral particles.Methods for introducing target vector expression vectors into packaging cell lines and for collecting viral particles generated by the packaging cell lines are well known in the art, including U.S. Patent No. 5,173,414, Tratschin et al., Mol. Cell. Biol. 5:3251-3260 (1985), Tratschin et al., Mol. Cell. Biol. 4:2072-2081 (1984), Hermonat & Muzyczka, PNAS 81:6466-6470 (1984), and Samulski et al., J. Virol. 63:03822-3828 (1989). Nucleic acids can also be introduced by direct microinjection (e.g., RNA injection).
[0220] In other embodiments of the method for modifying the BCL11A gene, the method utilizes CasX delivery particles (XDPs) for targeted delivery of RNPs to target cells. XDPs are particles that closely resemble viruses but do not contain viral genetic material and are therefore non-infectious. In some embodiments, the XDPs include a donor template containing CasX and gRNA complexed as RNPs, and optionally, all or part of the BCL11A gene for knockdown or knockout of the BCL11A gene or a portion thereof by insertion via an HDR or HITI mechanism. Embodiments of XDPs are described in more detail below.
[0221] VI. Polynucleotides and Vectors In another embodiment, the disclosure relates to polynucleotides encoding class 2, type V nucleases and gRNAs useful in editing the BCL11A gene. In some embodiments, the disclosure provides polynucleotides encoding the CasX protein and polynucleotides of the gRNAs of the CasX:gRNA systems of the embodiments described herein. In additional embodiments, the disclosure provides donor template polynucleotides encoding part or all of the BCL11A gene. In some cases, the donor template includes mutations or heterologous sequences for knocking down or knocking out the BCL11A gene upon insertion into a target nucleic acid. In yet another embodiment, the disclosure provides a vector comprising the polynucleotides encoding the CasX protein and CasX gRNA described herein, as well as the donor templates of the embodiments.
[0222] In some embodiments, the Disclosure provides polynucleotide sequences encoding any of the CasX variants of the embodiments described herein, including sequences having at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity with sequences 59, 72-99, 101-148, and 26908-27154 of Table 4. In some embodiments, the Disclosure provides polynucleotide sequences encoding any of the CasX variants of the embodiments described herein, including sequences having at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity with the sequences of sequence numbers 36–99, 101–148, and 26908–27154 in Table 4. In some embodiments, the Disclosure provides isolated polynucleotide sequences encoding gRNA sequences of any of the embodiments described herein, including the targeting sequences 272-2100 or 2286-26789, along with the sequences 4-16, 2238-2285, 26794-26839, or 27219-27265 in Tables 2 and 3. In some embodiments, the Disclosure provides isolated polynucleotide sequences encoding gRNA sequences of any of the embodiments described herein, including the targeting sequences 272-2100 or 2286-26789, along with the sequences 2101-2285, 26794-26839, and 27219-27265.In some embodiments, the Disclosure provides isolated polynucleotide sequences encoding any of the gRNA sequences described herein, including the sequences of SEQ ID NOs. 2281-2285, 26794-26839, and 27219-27265, along with the targeting sequences of SEQ ID NOs. 272-2100 or 2286-26789. In some embodiments, the sequences encoding the CasX protein are codon-optimized for expression in eukaryotic cells.
[0223] In some embodiments, the Disclosure provides polynucleotides encoding gRNA scaffold sequences of SEQ ID NOs. 4-16, 2238-2285, 26794-26839, or 27219-27265, or gRNA scaffold sequences shown in Table 2 or Table 3, or sequences having at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity thereto. In other embodiments, the Disclosure provides targeted sequence polynucleotides of Table 1, or sequences having at least about 65%, at least about 75%, at least about 85%, or at least about 95% identity with sequences of SEQ ID NOs. 272-2100 or 2286-26789. In some embodiments, the targeted sequence polynucleotide is then ligated (as either sgRNA or dgRNA) to the 3' end of the gRNA scaffold sequence. In other embodiments, the disclosure provides gRNAs comprising a targeted sequence polynucleotide having one or more single nucleotide polymorphisms (SNPs) compared to sequences 272-2100 or 2286-26789.
[0224] In other embodiments, the disclosure provides isolated polynucleotide sequences encoding gRNAs containing a targeting sequence that is complementary to and therefore capable of hybridizing with the BCL11A gene. In some embodiments, the polynucleotide sequence encodes a gRNA containing a targeting sequence that hybridizes with a BCL11A exon. In other embodiments, the polynucleotide sequence encodes a gRNA containing a targeting sequence that hybridizes with a BCL11A intron. In other embodiments, the polynucleotide sequence encodes a gRNA containing a targeting sequence that hybridizes with a BCL11A intron-exon junction. In other embodiments, the polynucleotide sequence encodes a gRNA containing a targeting sequence that hybridizes with an intergenetic region of the BCL11A gene. In other embodiments, the polynucleotide sequence encodes a gRNA containing a targeting sequence that hybridizes with a BCL11A regulatory element. In some cases, the BCL11A regulatory element is a BCL11A promoter or enhancer. In some cases, the BCL11A regulatory element is located at the 5' end of the BCL11A transcription start site, at the 3' end of the BCL11A transcription start, or within the BCL11A intron. In other embodiments, the polynucleotide sequence encodes a gRNA containing a targeting sequence that hybridizes with a sequence located at the 5' end of the GATA1 binding motif sequence. In other embodiments, the polynucleotide sequence encodes a gRNA containing a targeting sequence that hybridizes with a sequence that overlaps with the GATA1 binding motif sequence. In the particular embodiments described above, the polynucleotide sequence encodes a gRNA containing a targeting sequence having sequence number 22. In some cases, the BCL11A regulatory element is located within the intron of the BCL11A gene. In other cases, the BCL11A regulatory element is located within the 5' UTR of the BCL11A gene. In yet another case, the BCL11A regulatory element is located within the 3' UTR of the BCL11A gene.
[0225] In other embodiments, the disclosure provides a donor template nucleic acid, the donor template comprising a nucleotide sequence homologous to the BCL11A target nucleic acid sequence. In some embodiments, the BCL11A donor template comprises at least a portion of the BCL11A gene for the purpose of gene editing in conjunction with a CasX:gRNA system. In other embodiments, the BCL11A donor sequence comprises a sequence encoding at least a portion of the BCL11A exon. In other embodiments, the BCL11A donor template has a sequence encoding at least a portion of the BCL11A intron. In other embodiments, the BCL11A donor template has a sequence encoding at least a portion of the BCL11A intron-exon junction. In other embodiments, the BCL11A donor template has a sequence encoding at least a portion of the intergenetic region of the BCL11A gene. In other embodiments, the BCL11A donor template has a sequence encoding at least a portion of the BCL11A regulatory element. In some cases, the BCL11A donor template is a wild-type sequence encoding at least a portion of Sequence ID No. 100. In other cases, the BCL11A donor template sequence contains one or more mutations compared to the wild-type BCL11A gene. In certain embodiments, the donor template has a sequence that encodes part or all of the GATA1 binding motif sequence but has at least 1 to 5 mutations compared to the wild-type sequence. In the embodiments described above, the donor template is at least 10 nucleotides, at least 100 nucleotides, at least 200 nucleotides, at least 300 nucleotides, at least 400 nucleotides, at least 500 nucleotides, at least 600 nucleotides, at least 700 nucleotides, at least 800 nucleotides, at least 900 nucleotides, at least 1,000 nucleotides, at least 2,000 nucleotides, at least 3,000 nucleotides, at least 4,000 nucleotides, at least 5,000 nucleotides, at least 6,000 nucleotides, at least 7,000 nucleotides, at least 8,000 nucleotides, at least 9,000 nucleotides, at least 10,000 nucleotides, at least 12,000 nucleotides, or at least 15,000 nucleotides.In some embodiments, the donor template contains at least about 10 to about 15,000 nucleotides. In some embodiments, the donor template is a single-stranded DNA template. In other embodiments, the donor template is a single-stranded RNA template. In other embodiments, the donor template is a double-stranded DNA template. In some embodiments, the donor template can be provided as a naked nucleic acid in a system for editing the BCL11A gene and does not need to be incorporated into a vector. In other embodiments, the donor template can be incorporated into a vector (e.g., a viral vector) to facilitate its delivery to cells.
[0226] In other embodiments, the Disclosure relates to methods for producing polynucleotide sequences encoding any CasX variant or gRNA (including homologous variants thereof) of any of the embodiments described herein, and methods for expressing proteins or RNAs expressed or transcribed by the polynucleotide sequences. Generally, the Method comprises producing a polynucleotide sequence encoding any CasX variant or gRNA of any of the embodiments described herein, and incorporating the coding gene into an expression vector suitable for a host cell. The polynucleotides and expression vectors of the Disclosure can be prepared using standard recombination techniques in molecular biology. For the production of a coded reference CasX, CasX variant, or gRNA of any of the embodiments described herein, the Method comprises transforming suitable host cells with an expression vector containing the coding polynucleotide, and culturing the host cells under conditions that cause or enable the resulting reference CasX, CasX variant, or gRNA of any of the embodiments described herein to be expressed or transcribed in the transformed host cells, thereby producing the CasX variant or gRNA, which is recovered by the Method described herein, or by standard purification methods known in the Art, or as described in the Examples.
[0227] According to this disclosure, recombinant DNA molecules that direct expression in appropriate host cells are generated using nucleic acid sequences encoding a CasX variant or gRNA (or their complements) from any of the embodiments described herein. Several cloning strategies are suitable for carrying out this disclosure, many of which are used to generate constructs containing genes encoding the compositions of this disclosure or their complements. In some embodiments, a cloning strategy is used to generate a gene encoding a construct containing nucleotides encoding a CasX variant or gRNA, which is used to transform host cells for expression of the compositions.
[0228] In some approaches, a construct containing a DNA sequence encoding a CasX variant or gRNA is first prepared. Exemplary methods for preparing such constructs are described in the examples. The construct is then used to generate an expression vector suitable for transforming host cells (e.g., prokaryotic or eukaryotic host cells for the expression and recovery of the protein construct, in the case of CasX or gRNA). The host cell is E. coli, if desired. In other embodiments, the host cell is a eukaryotic cell. The eukaryotic host cells can be selected from baby hamster kidney fibroblasts (BHK) cells, human fetal kidney 293 (HEK293), human fetal kidney 293T (HEK293T), NS0 cells, SP2 / 0 cells, YO myeloma cells, P3X63 mouse myeloma cells, PER cells, PER.C6 cells, hybridoma cells, NIH3T3 cells, CV-1 (monkey) origin (COS) cells with SV40 genetic material, HeLa, Chinese hamster ovary (CHO), or yeast cells, or other eukaryotic cells known in the art that are suitable for the production of recombinant products. Exemplary methods for generating expression vectors, transforming host cells, and expressing and recovering CasX variants or gRNAs are described in the examples.
[0229] Genes encoding CasX variants or gRNA constructs can be prepared in one or more steps, either entirely synthetically or in combination with enzymatic processes such as restriction enzyme-mediated cloning, PCR, and duplication extension, including methods described in more detail in the examples. The methods disclosed herein can be used, for example, to ligate the sequences of polynucleotides encoding various components (e.g., CasX and gRNA) of a desired sequence. Genes encoding polypeptide compositions are constructed from oligonucleotides using standard gene synthesis techniques.
[0230] In some embodiments, the nucleotide sequence encoding the CasX protein is codon-optimized for the intended host cell. This type of optimization may involve mutations in the coding nucleotide sequence to mimic the codon preferences of the intended host organism or cell, while still encoding the same CasX protein. Thus, the codons can be altered, but the encoded protein or gRNA remains unchanged. For example, if the intended target cells for the CasX protein are human cells, a human codon-optimized CasX coding nucleotide sequence can be used. As another non-limiting example, if the intended host cells are mouse cells, a mouse codon-optimized CasX coding nucleotide sequence can be generated. Gene design can be carried out using an algorithm that optimizes the codon usage frequency and amino acid composition appropriate for the host cell to be used in the production of the reference CasX or CasX variant. In one method of this disclosure, a library of polynucleotides encoding the components of the construct is prepared and then constructed as described above. The resulting gene is then assembled, the resulting gene is used to transform host cells, and CasX variants or gRNA compositions are produced and recovered for characterization.
[0231] This disclosure provides the use of plasmid expression vectors comprising replication and regulatory sequences that are compatible with and recognized by host cells and operably ligated to a gene encoding a polypeptide for controlled polypeptide expression or RNA transcription. Such vector sequences for various bacteria, yeasts, and viruses are well known. Useful expression vectors that may be used include, for example, segments of chromosomal, non-chromosomal, and synthetic DNA sequences. “Expression vector” means a DNA construct comprising a DNA sequence operably ligated to a suitable regulatory sequence that can result in the expression of polypeptide-encoding DNA in a suitable host. A requirement is that the vector is replicable and viable in a selected host cell. Low copy number vectors or high copy number vectors may be used as desired. The regulatory sequence of the vector includes a promoter that results in transcription, an optional operator sequence that controls such transcription, a sequence encoding a suitable mRNA-ribosome binding site, and a sequence that controls the termination of transcription and translation. In some embodiments, a nucleotide sequence encoding gRNA is operably ligated to a regulatory element (e.g., a transcriptional regulatory element such as a promoter). In some embodiments, the nucleotide sequence encoding the CasX protein is operably ligated to a regulatory element (e.g., a transcriptional regulatory element such as a promoter). In other cases, the nucleotides encoding CasX and gRNA are ligated and operably ligated to a single regulatory element. The promoter may be any DNA sequence that exhibits transcriptional activity in a selected host cell and may be derived from a gene encoding either an homogeneous or heterogeneous protein to the host cell. Exemplary regulatory elements include transcriptional promoters, transcriptional enhancer elements, transcription termination signals, internal ribosome entry sites (IRES) or P2A peptides that enable translation of multiple genes from a single transcript, polyadenylation sequences that promote downstream transcription termination, sequences for optimizing translation initiation, and translation termination sequences. In some cases, the promoter is a constitutively active promoter. In some cases, the promoter is a regulated promoter.In some cases, the promoter is an inducible promoter. In some cases, the promoter is a tissue-specific promoter. In some cases, the promoter is a cell-type-specific promoter. In some cases, the transcriptional regulatory element (e.g., the promoter) is functional in a targeted cell type or targeted cell population. For example, in some cases, the transcriptional regulatory element may be functional in eukaryotic cells (e.g., packaging cells for viruses or XDP vectors, hematopoietic stem cells (HSCs), hematopoietic progenitor cells (HPCs), CD34+ cells, mesenchymal stem cells (MSCs), embryonic stem cells (ES), induced pluripotent stem cells (iPSCs), myeloid common progenitor cells, proerythroblasts, and erythroblasts).
[0232] Non-limiting examples of pol II promoters include EF-1α, EF-1α core promoter, Jens Tornoe (JeT), cytomegalovirus (CMV)-derived promoter, CMV immediate early (CMVIE), CMV enhancer, herpes simplex virus (HSV) thymidine kinase, early and late simian virus 40 (SV40), SV40 enhancer, retrovirus-derived long terminal repeat (LTR), mouse metallothionein-I, adenovirus major late promoter (Ad MLP), CMV promoter full-length promoter, minimal CMV promoter, and chicken.
[0233] [ka] Chickens with (CBA), CBA hybrid (CBh), and cytomegalovirus enhancer
[0234] [ka] (CB7), chicken β-actin promoter and rabbit β-globin splice acceptor site fusion (CAG), Rous sarcoma virus (RSV) promoter, HIV-Ltr promoter, hPGK promoter, HSV TK promoter, 7SK promoter, Mini-TK promoter, human synapsin I (SYN) promoter for neuron-specific expression, β-actin promoter, super core promoter 1 (SCP1), Mecp2 promoter for selective expression in neurons, minimal IL-2 promoter, Rous sarcoma virus enhancer / promoter (single), splenic lesion-forming virus long chain end repeat (LTR) promoter, TBG promoter, human thyroxine-binding globulin gene-derived promoter (liver-specific), PGK promoter, human ubiquitin C promoter (ubiquitin C, UBC), UCOE promoter (promoter for HNRPA2B1-CBX3), synthetic CAG promoter, histone H2 promoter, histone H3 promoter, U1a1 nuclear small RNA promoter (226nt), U1a1 nuclear small RNA promoter (226nt), U1b2 nuclear small RNA promoter (246nt), GUSB promoter, CBh promoter, rhodopsin (Rho) promoter, silencing-prone spleen focus forming virus (SFFV) promoter, human H1 promoter (H1), POL1 promoter, TTR minimal enhancer / promoter, β-kinesin promoter, mouse mammary cancer virus long-chain terminal repeat (LTR) promoter, human eukaryotic initiation factor 4A (EIF4A1) promoter, ROSA26 promoter, glyceraldehyde 3-phosphate dehydrogenase Examples include, but are not limited to, the dehydrogenase (GAPDH) promoter, the tRNA promoter, and the aforementioned shortened versions and sequence variants.In certain embodiments, the pol II promoter is EF-1α, and the promoter enhances transfection efficiency, transcription or expression of the CRISPR nuclease transgene, the percentage of expression-positive clones, and the copy number of the episomal vector in long-term culture.
[0235] Non-limiting examples of pol III promoters include, but are not limited to, U6, mini-U6, U6 truncated promoter, 7SK, and H1 variants, BiH1 (bidirectional H1 promoter), BiU6, Bi7SK, BiH1 (bidirectional U6, 7SK, and H1 promoter), gorilla U6, rhesus monkey U6, human 7SK, human H1 promoter, and their sequence variants. In the embodiments described above, the pol III promoter enhances gRNA transcription.
[0236] The selection of appropriate vectors and promoters is well within the realm of those skilled in the art, as it relates to the control of expression (e.g., for modifying the BCL11A gene). Expression vectors may also contain ribosome binding sites and transcription terminators for translation initiation. Expression vectors may also contain appropriate sequences for amplification of expression. Expression vectors may also contain nucleotide sequences encoding protein tags (e.g., 6×His tags, hemagglutinin tags, fluorescent proteins, etc.) that can be fused to the CasX protein, thereby obtaining chimeric CasX proteins for use in purification or detection.
[0237] The recombinant expression vectors of this disclosure may also contain elements that promote potent expression of the CasX protein and gRNA of this disclosure. For example, the recombinant expression vector may contain one or more polyadenylation signals (poly(A)), intron sequences, or post-transcriptional regulatory elements (e.g., woodchuck hepatitis post-transcriptional regulatory element, WPRE)). Exemplary poly(A) sequences include the hGH poly(A) signal (short chain), HSV TK poly(A) signal, synthetic polyadenylation signal, SV40 poly(A) signal, and β-globin poly(A) signal. Those skilled in the art will be able to select suitable elements to be included in the recombinant expression vectors described herein.
[0238] In some embodiments, one or more recombinant expression vectors comprising one or more of the following are provided herein: (i) a nucleotide sequence of a donor template nucleic acid, wherein the donor template comprises a nucleotide sequence homologous to the sequence of the target BCL11A locus of a target nucleic acid (e.g., a target genome); (ii) a nucleotide sequence encoding a gRNA that hybridizes to the target sequence of the BCL11A locus of the target genome (e.g., configured as a single or dual guide RNA) operably ligated to a promoter operable in a target cell such as a eukaryotic cell; and (iii) a nucleotide sequence encoding a CasX protein operably ligated to a promoter operable in a target cell such as a eukaryotic cell. In some embodiments, the donor template, gRNA, and the sequence encoding the CasX protein are in different recombinant expression vectors, and in other embodiments, one or more polynucleotide sequences (of the donor template, CasX, and gRNA) are in the same recombinant expression vector. In other cases, CasX and gRNA are delivered to target cells as RNPs (e.g., by electroporation or chemical means), and the donor template is delivered by a vector.
[0239] Polynucleotide sequences are inserted into vectors by various procedures. Generally, DNA is inserted into appropriate restriction endonuclease sites using techniques known in the art. Vector components generally include, but are not limited to, one or more of the following: signal sequences, origins of replication, one or more marker genes, enhancer elements, promoters, and transcription termination sequences. The construction of suitable vectors containing one or more of these components is done using standard ligation techniques known to those skilled in the art. Such techniques are well known in the art and are described in detail in the scientific and patent literature. Various vectors are publicly available. Vectors may be in the form of plasmids, cosmids, viral particles, or phages that can be conveniently used in recombinant DNA procedures, and the choice of vector often depends on the host cell into which the vector is introduced. Thus, the vector may be an autonomously replicating vector, i.e., a vector that exists as an extrachromosomal entity and whose replication is independent of chromosomal replication (e.g., a plasmid). Alternatively, when introduced into a host cell, the vector may be integrated into the host cell genome and replicate together with the chromosome into which it is integrated. Once introduced into suitable host cells, the expression of proteins involved in antigen processing, antigen presentation, antigen recognition, and / or antigen response can be determined using any nucleic acid assay or protein assay known in the art. For example, the presence of transcribed mRNA of reference CasX or a CasX variant can be detected and / or quantified by conventional hybridization assays (e.g., Northern blot analysis), amplification procedures (e.g., RT-PCR), SAGE (U.S. Patent No. 5,695,937), and array-based technologies (see, for example, U.S. Patents No. 5,405,783, 5,412,087, and 5,445,934) using probes complementary to any region of the polynucleotide.
[0240] Polynucleotides and recombinant expression vectors can be delivered to target host cells by a variety of methods. Such methods include, but are not limited to, viral infection, transfection, lipofection, electroporation, calcium phosphate precipitation, polyethyleneimine (PEI)-mediated transfection, DEAE-dextran-mediated transfection, microinjection, liposome-mediated transfection, particle gun technology, nucleofection, direct addition by a cell-permeable CasX protein fused to or recruiting donor DNA, cell compression, calcium phosphate precipitation, direct microinjection, nanoparticle-mediated nucleic acid delivery, and the use of commercially available TransMessenger® reagents from Qiagen, Stemfect™ RNA transfection kits from Stemgent, and TransIT®-mRNA transfection kits from Mirus Bio LLC, Lonza nucleofection, Maxagen electroporation, etc.
[0241] Recombinant expression vector sequences can be packaged into viruses or virus-like particles (also referred to herein as “particles” or “virions”) for subsequent ex vivo, in vitro, or in vivo infection and transformation of cells. Such particles or virions typically contain proteins that capsidize or package the vector genome. Suitable expression vectors may include viral expression vectors based on: vaccinia virus; poliovirus; adenovirus; retroviral vectors (e.g., mouse leukemia virus), splenic necrosis virus, and retroviruses (e.g., Rous sarcoma virus, Harvey sarcoma virus, avian leukemia virus, lentivirus, human immunodeficiency virus, myeloproliferative sarcoma virus, and mammary cancer virus). In some embodiments, the recombinant expression vector of this disclosure is a recombinant adeno-associated virus (AAV) vector. In some embodiments, the recombinant expression vector of this disclosure is a recombinant lentiviral vector. In some embodiments, the recombinant expression vector of this disclosure is a recombinant retroviral vector.
[0242] In some embodiments, the recombinant expression vector of the Disclosure is a recombinant adeno-associated virus (AAV) vector. In some embodiments, the recombinant expression vector of the Disclosure is a recombinant lentiviral vector. In some embodiments, the recombinant expression vector of the Disclosure is a recombinant retroviral vector.
[0243] AAV is a small (20 nm) nonpathogenic virus useful for treating human diseases in situations where a viral vector is used for delivery to cells such as eukaryotic cells, either in vivo or ex vivo for cells prepared for administration to a target. A construct (e.g., a construct encoding either the CasX protein and / or CasX gRNA of the embodiments described herein) is generated and flanked by the AAV inverted terminal repeat (ITR) sequence, thereby enabling the packaging of the AAV vector into AAV viral particles.
[0244] The “AAV” vector may refer to the naturally occurring wild-type virus itself or its derivatives. This term encompasses all subtypes, serotypes, and pseudotypes, as well as both naturally occurring and recombinant forms, unless otherwise required. As used herein, the term “serotype” refers to an AAV identified and distinguished from other AAVs based on the reactivity of its capsid protein with a given antiserum. For example, there are many known serotypes of primate AAVs. In some embodiments, the AAV vector is selected from AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, AAV44.9, AAV-Rh74 (rhesus macaque-derived AAV), and AAVRh10, as well as modified capsids of these serotypes. For example, serotype AAV-2 is used to refer to an AAV that contains a capsid protein encoded from the AAV-2 cap gene and a genome containing 5' and 3' ITR sequences derived from the same AAV-2 serotype. Pseudotype AAV refers to an AAV that contains a capsid protein derived from one serotype and a viral genome containing a 5'-3' ITR from a second serotype. Pseudotype rAAV is expected to have the cell surface binding properties of the capsid serotype and the genetic properties that match the ITR serotype. Pseudotype recombinant AAV (rAAV) is produced using standard techniques described in the art. As used herein, for example, rAAV1 may be used to refer to an AAV having both a capsid protein and a 5'-3' ITR derived from the same serotype, or it may refer to an AAV having a capsid protein derived from serotype 1 and a 5'-3' ITR derived from a different AAV serotype (e.g., AAV serotype 2). For each example illustrated herein, the vector design and production are described, along with the serotype of the capsid and the 5'-3'ITR sequence.
[0245] "AAV virus" or "AAV viral particle" refers to a viral particle composed of at least one AAV capsid protein (preferably by all of the capsid proteins of wild-type AAV) and a packaged polynucleotide. If the particle further contains a heterologous polynucleotide (i.e., a polynucleotide other than the wild-type AAV genome to be delivered to mammalian cells), it is typically referred to as "rAAV". Exemplary heterologous polynucleotides are polynucleotides containing any of the CasX proteins and / or sgRNAs of the embodiments described herein, and optionally a donor template.
[0246] "Adeno-associated virus inverted terminal repeat" or "AAV ITR" means the region recognized in the art found at each end of the AAV genome, which functions cis together as a DNA replication origin and as a viral packaging signal. AAV ITRs, together with the AAV rep coding region, provide for efficient excision and rescue from the mammalian cell genome and integration of the nucleotide sequence inserted between two adjacent ITRs into the mammalian cell genome. The nucleotide sequences of the AAV ITR regions are known. For example, Kotin, R.M. (1994) Human Gene Therapy 5:793-801; Berns, K.I. "Parvoviridae and their Replication" in Fundamental Virology, 2 ndSee Edition, (BNFields and DMKnipe, eds.). As used herein, AAV ITRs do not need to have the indicated wild-type nucleotide sequence, but can be modified, for example, by nucleotide insertion, deletion, or substitution. Additionally, AAV ITRs may be derived from any of several AAV serotypes (including, but not limited to, AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV-Rh74, and AAVRh10) and modified capsids of these serotypes. Furthermore, the 5' and 3' ITRs adjacent to the selected nucleotide sequence in the AAV vector do not necessarily need to be identical or derived from the same AAV serotype or isolate, as long as they function as intended (i.e., enabling the excision and rescue of the desired sequence from the host cell genome or vector, and enabling the incorporation of the heterologous sequence into the recipient cell genome if the AAV Rep gene product is present in the cell). The use of AAV serotypes for the incorporation of heterologous sequences into host cells is known in the art (see, for example, International Publication No. 2018 / 195555(A1) and U.S. Patent Application Publication No. Incorporated 2018 / 0258424(A1), which are incorporated herein by reference).
[0247] The "AAV rep coding region" refers to the region of the AAV genome that encodes the replication proteins Rep78, Rep68, Rep52, and Rep40. These Rep expression products have been shown to have many functions, including recognition, binding, and nicking of AAV DNA replication origins, DNA helicase activity, and regulation of transcription from AAV (or other heterologous) promoters. Collectively, Rep expression products are required for AAV genome replication. The "AAV cap coding region" refers to the region of the AAV genome that encodes the capsid proteins VP1, VP2, and VP3, or their functional homologs. Collectively, these Cap expression products provide the packaging functions required to package the viral genome.
[0248] In some embodiments, the AAV capsid used to deliver the coding sequence of CasX and gRNA, and optionally the DMPK donor template nucleotide, to host cells can be derived from any of several AAV serotypes, including but not limited to AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, AAV44.9, AAV-Rh74 (rhesus monkey-derived AAV), and AAVRh10, while AAV ITR is derived from AAV serotype 2. In certain embodiments, AAV1, AAV7, AAV6, AAV8, or AAV9 is used for the delivery of CasX, gRNA, and optionally the donor template nucleotide to host muscle cells.
[0249] To produce rAAV virus particles, AAV expression vectors are introduced into suitable host cells using known techniques (e.g., transfection). Packaging cells are typically used to form the virus particles. Such cells include HEK293 cells (and other cells known in the art) for packaging adenoviruses. Many transfection techniques are generally known in the art. See, for example, Sambrook et al. (1989) Molecular Cloning, a laboratory manual, Cold Spring Harbor Laboratories, New York. Particularly preferred transfection methods include calcium phosphate coprecipitation, direct microinjection into cultured cells, electroporation, liposome-mediated gene transfer, lipid-mediated transduction, and nucleic acid delivery using high-speed microparticle guns.
[0250] An advantage of the rAAV constructs of this disclosure is that smaller-sized CRISPR type V nucleases (e.g., CasX in this embodiment) allow these components to be incorporated into a transgene so that a single rAAV particle can deliver and transduce all the necessary editing and auxiliary expression components to target cells in a form that results in the expression of CRISPR nucleases and gRNAs capable of effectively modifying the target nucleic acids of the target cells. A typical schematic diagram of such a construct is shown in Figure 13. This is in stark contrast to other CRISPR systems such as Cas9, which typically use a two-particle system to deliver the necessary editing components to target cells. Accordingly, in some embodiments of the rAAV system, the Disclosure provides a first plasmid comprising i) an ITR, a sequence encoding a CasX variant, a sequence encoding one or more gRNAs, a first promoter operably linked to CasX, a second promoter operably linked to the gRNA, and optionally one or more enhancer elements; ii) a second plasmid comprising rep and cap genes; and iii) a third plasmid comprising helper genes, so that upon transfection of appropriate packaging cells, the cells can produce an rAAV having the ability to deliver a sequence capable of expressing CasX nuclease and a gRNA capable of editing the target nucleic acid of the target cell in a single particle to the target cell.In some embodiments of the rAAV system, the sequence encoding the CRISPR protein and the sequence encoding at least the first gRNA have a nucleotide length of less than about 3100, less than about 3090, less than about 3080, less than about 3070, less than about 3060, less than about 3050, or less than about 3040, thereby allowing the sequences encoding the first promoter and the second promoter and optionally one or more enhancer elements to together have a nucleotide length of at least about 1300, at least about 1350, at least about 1360, at least about 1370, at least about 1380, at least about 1390, at least about 1400, at least about 1500, at least about 1600 nucleotides, at least 1650, at least about 1700, at least about 1750, at least about 1800, at least about 1850, or at least about 1900 nucleotides. In some embodiments of the rAAV system, the sequences encoding the first promoter and at least one accessory element together have a nucleotide length of at least about 1300, at least about 1350, at least about 1360, at least about 1370, at least about 1380, at least about 1390, at least about 1400, at least about 1500, at least about 1600 nucleotides, at least 1650, at least about 1700, at least about 1750, at least about 1800, at least about 1850, or at least more than about 1900. In some embodiments of the rAAV system, the sequences encoding the first and second promoters and at least one accessory element together have a nucleotide length of at least about 1300, at least about 1350, at least about 1360, at least about 1370, at least about 1380, at least about 1390, at least about 1400, at least about 1500, at least about 1600 nucleotides, at least 1650, at least about 1700, at least about 1750, at least about 1800, at least about 1850, or at least more than about 1900.
[0251] In some embodiments, host cells transfected with the above-described AAV expression vector are made capable of providing AAV helper functions to replicate and capsidize nucleotide sequences adjacent to the AAV ITR to produce rAAV virus particles. AAV helper functions are generally AAV-derived coding sequences that, when expressed, provide the AAV gene product, which can then function in trans for productive AAV replication. AAV helper functions are used herein to complement necessary AAV functions lacking in the AAV expression vector. Therefore, AAV helper functions include one or both of the major AAV ORFs (Open Reading Frames) encoding the rep and cap coding regions, or their functional homologs. Accessory functions can be introduced into host cells using methods known to those skilled in the art and subsequently expressed in the host cells. Typically, accessory functions are provided by infection of host cells with an unrelated helper virus. In some embodiments, accessory functions are provided using accessory function vectors. Depending on the host / vector system used, any of several suitable transcription and translational regulatory elements, including constitutive promoters and inducible promoters, transcriptional enhancer elements, and transcriptional terminators, may be used in the expression vector. In some embodiments, this disclosure provides host cells containing the AAV vector of the embodiments disclosed herein.
[0252] In other embodiments, preferred vectors may include virus-like particles (VLPs). VLPs are particles that closely resemble viruses but do not contain viral genetic material and are therefore non-infectious. In some embodiments, the VLP includes a polynucleotide encoding the desired transgene (e.g., either the CasX protein and / or gRNA of the embodiments), packaged with one or more viral structural proteins, and optionally a donor template polynucleotide as described herein. In other embodiments, the disclosure provides an in vitro produced XDP comprising a CasX:gRNA RNP complex and optionally a donor template. XDPs can be constructed using combinations of structural proteins derived from different viruses. These include components derived from viridae, including parvoviridae (e.g., adeno-associated viruses), retroviridae (e.g., alpha-retrovirus, beta-retrovirus, gamma-retrovirus, delta-retrovirus, epsilon-retrovirus, or lentivirus), flaviviridae (e.g., hepatitis C virus), paramyxoviridae (e.g., Nipah), and bacteriophages (e.g., Qβ, AP205). In some embodiments, the disclosure provides an XDP system designed using components of retroviruses (including lentiviruses (such as HIV) and alpha-retroviruses, beta-retroviruses, gamma-retroviruses, delta-retroviruses, and epsilon-retroviruses), in which individual plasmids containing polynucleotides encoding various components are introduced into packaging cells, which then produce XDP.In some embodiments, the Disclosure provides an XDP comprising one or more components from i) a protease, ii) a protease cleavage site, iii) one or more components of a gag polyprotein selected from: matrix protein (MA), nucleocapsid protein (NC), capsid protein (CA), p1 peptide, p6 peptide, P2A peptide, P2B peptide, P10 peptide, p12 peptide, PP21 / 24 peptide, P12 / P3 / P8 peptide, and P20 peptide, v) CasX, vi) gRNA, and vi) a targeted glycoprotein or antibody fragment, wherein the resulting XDP particles capsidize the CasX:gRNA RNP. The polynucleotides encoding Gag, CasX, and gRNA may further comprise pairs of components designed to assist in the transport of components from the nucleus of a host cell to the budding XDP. Non-limiting examples of such transport components include MS2 hairpins, PP7 hairpins, Qβ hairpins, and U1 hairpin II and U1A signal recognition particles having binding affinity to MS2 coat protein, PP7 coat protein, and Qβ coat protein, respectively. In other embodiments, the gRNA may include a Rev response element (RRE) or a portion thereof having binding affinity to Rev, which can be ligated to a Gag polyprotein. In other embodiments, the gRNA may include one or more RREs and one or more MS2 hairpin sequences. In other embodiments, the gRNA may include a Rev response element (RRE) or a portion thereof having binding affinity to Rev, which can be ligated to a Gag polyprotein. The RRE may be selected from the group consisting of stem IIB of the Rev response element (RRE), stems II-V of the RRE, stem II of the RRE, Rev-binding element (RBE) of stem IIB, and full-length RREs.In the above-described embodiment, the components include UGGGCGCAGCGUCAAUGACGCUGACGGUACA (Stem IIB; SEQ ID NO: 27266), GCACUAUGGGCGCAGCGUCAAUGACGCUGACGGUACAGGCCAGACAAUUAUUGUCUGGUAUAGUGC (Stem II; SEQ ID NO: 27267), GCUGACGGUACAGGC (RBE, SEQ ID NO: 27268), CAGGAAGCACUAUGGGCGCAGCGUCAAUGACGCUGACGGUACAGGCCAGACAAUUAUUGUCUGGUAUAGUGCAGCAGCAGAACAAUUUGCUGAGGGCUAUUGAGGCGCAACAGCAUCUGUUGCAACUCACAGUC The sequence includes UGGGGCAUCAAGCAGCUCCAGGCAAGAAUCCUG (Stem II-V; SEQ ID NO: 27269) and AGGAGCUUUGUUCCUUGGGUUCUUGGGAGCAGCAGGAAGCACUAUGGGCGCAGCGUCAAUGACGCUGACGGUACAGGCCAGACAAUUAUUGUCUGGUAUAGUGCAGCAGCAGAACAAUUUGCUGAGGGCUAUUGAGGCGCAACAGCAUCUGUUGCAACUCACAGUCUGGGGCAUCAAGCAGCUCCAGGCAAGAAUCCUGGCUGUGGAAAGAUACCUAAAGGAUCAACAGCUCCU (Full-length RRE; SEQ ID NO: 27270). In other embodiments, the gRNA may include one or more RREs and one or more MS2 hairpin sequences. In certain embodiments, the gRNA may include an MS2 hairpin variant optimized to increase binding affinity to the MS2 coat protein, thereby enhancing the incorporation of the gRNA and associated CasX into budding XDP. gRNA variants containing the MS2 hairpin variant include gRNA variants 275-315 and 317-320 (SEQ ID NOs. 2722-27264).
[0253] The surface-mounted targeted glycoprotein or antibody fragment provides tropism of XDP to target cells, and upon administration to and entry into target cells, the RNP molecule is freely transported into the cell nucleus. The envelope glycoprotein may be derived from any of the following enveloped viruses known in the art to confer tropism to XDP: but not limited to, Argentine hemorrhagic fever virus, Australian bat virus, Autographa californica nuclear polyhedron disease virus, avian leukemia virus, baboon endogenous virus, Bolivian hemorrhagic fever virus, Borna disease virus, Breda virus, Bunyamwella virus, Chandipla virus, Chikungunya virus, Crimean-Congo hemorrhagic fever virus, dengue virus, Duvenhaig virus, Eastern equine encephalitis virus, Ebola hemorrhagic fever virus, Ebola-Zaire virus, enteric adenovirus, ephemeral virus, Epstein-Barr virus. Viru (EBV), European bat virus 1, European bat virus 2, Fug synthetic gP fusion, gibbon leukemia virus, hantavirus, hendra virus, hepatitis A virus, hepatitis B virus, hepatitis C virus, hepatitis D virus, hepatitis E virus, hepatitis G virus (GB virus C), herpes simplex virus type 1, herpes simplex virus type 2, human cytomegalovirus (HHV5), human foamy virus, human herpesvirus (HHV), human herpesvirus type 7, human herpesvirus type 6, human herpesvirus type 8, human immunodeficiency virus type 1 (HIV-1), human metapneumovirus, human T lymphotropic virus 1, influenza A virus, influenza B virus, influenza C virus, Japanese encephalitis virus, Kaposi's sarcoma-associated herpesvirus (Kaposi's Sarcoma-associated herpesvirus (HHV8), Kyasanur Forest disease virus, Lacrosse virus, Lagos bat virus, Lassa fever virus, lymphocytic choriomeningitis virusViruses (LCMV), Machupovirus, Marburg hemorrhagic fever virus, measles virus, Middle East Respiratory Syndrome-related coronavirus, Mocola virus, Moloney's mouse leukemia virus, monkeypox, mouse mammary cancer virus, mumps virus, mouse gamma herpesvirus, Newcastle disease virus, Nipah virus, Norwalk virus, Omsk hemorrhagic fever virus, papillomavirus, parvovirus, pseudorabies virus, qualanfil virus, rabies virus, RD114 feline endogenous retrovirus, respiratory syncytial virus (RSV), Rift Valley fever virus, Ross River virus, r rotavirus, Rous sarcoma virus, rubella virus, Sabia-related hemorrhagic fever virus, SARS-associated coronavirus (SARS-CoV), Sendai virus, Takaribe virus, Sogoto virus, tick-borne encephalitis virus, varicella zoster virus Viruses (HHV3), varicella-zoster virus (HHV3), varicella-zoster virus, varicella-zoster virus, venezuelanovirus, venezuelanovirus, vesicular stomatitis virus (VSV), VSV-G, vesiclovirus, West Nile virus, Western equine encephalitis virus, and Zika virus.
[0254] In other embodiments, the Disclosure provides the aforementioned XDP, further comprising one or more components of a pol polyprotein (e.g., a protease), and optionally, a second CasX or donor template. The Disclosure intends multiple configurations of the arrangement of the encoded component, including several replications of the encoded component. The foregoing offers advantages over other vectors in the art in that viral transduction into dividing and non-dividing cells is efficient and XDP delivers potent, short-lived RNPs that evade the target immune surveillance mechanisms that would otherwise detect the foreign protein. Non-limiting exemplary XDP systems are described in International Application PCT / US20 / 63488 and International Publication 2021 / 113772 (A1) (incorporated herein by reference). In some embodiments, the Disclosure provides host cells comprising a polynucleotide or vector encoding any of the embodiments of the aforementioned XDP.
[0255] When generating and recovering XDP containing CasX:gRNA RNP in any of the embodiments described herein, the XDP may be used in a method for editing target cells by administration of such XDP, as described in more detail below.
[0256] VII.Cells In another embodiment, a cell population containing the BCL11A gene, modified ex vivo by any embodiment of the system or method described herein, is provided herein. In some embodiments, such genetically modified cells may be administered to a subject for purposes such as gene therapy. For example, in a method for treating hemoglobin disorders such as sickle cell disease or β-thalassemia, administration results in increased γ-globin expression and increased fetal hemoglobin (HbF) in the subject. In other embodiments, the disclosure provides a composition of modified cells for use as a pharmaceutical in the treatment of hemoglobin disorders.
[0257] Cells suitable for ex vivo modification of the BCL11A gene by one or more guides targeting class 2, type V Cas nucleases and BCL11A target nucleic acids may be hematopoietic progenitor cells (HPCs), hematopoietic stem cells (HSCs), CD34+ cells, mesenchymal stem cells (MSCs), induced pluripotent stem cells (iPSCs), myeloid common progenitor cells, proerythroblasts, or erythroblasts. In some embodiments, the cell population to be modified is animal cells, e.g., derived from rodent, rat, mouse, rabbit, canine cells, or non-human primate cells (e.g., cynomolgus monkey cells). In some embodiments, the cells are human cells. In some cases, the cells to be modified are autologous to the subject to which the cells are administered. In other cases, the cells are homogeneous to the subject to which the cells are administered. In some cases, the ex vivo cells are isolated from the bone marrow or peripheral blood of the subject.
[0258] In some embodiments, the Disclosure provides a method for introducing into each cell population one of the following: i) a CasX:gRNA system comprising CasX and gRNA from any one embodiment described herein; ii) a CasX:gRNA system comprising CasX, gRNA, and a donor template from any one embodiment described herein; iii) a nucleic acid encoding CasX and gRNA and optionally comprising a donor template; iv) a vector comprising the nucleic acid of (iii) above, which may be an AAV from any one embodiment described herein; v) an XDP comprising a CasX:gRNA system from any one embodiment described herein; or vi) a combination of two or more of (i) to (v); and a cell population thereby modified, wherein the BCL11A target nucleic acid sequence of the cells targeted by the gRNA is modified by the CasX protein and optionally by the donor template. In some embodiments, the Method further comprises administering a second gRNA or a nucleic acid encoding a second gRNA, wherein the second gRNA has a targeting sequence complementary to a different or overlapping portion of the target nucleic acid sequence compared to the first gRNA. In some cases, CasX and gRNA are delivered to the cell population as RNP (an embodiment thereof is described above herein) and optionally as a donor template. In some embodiments, the disclosure provides a cell population modified by the aforementioned method, wherein at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% of the modified cells are modified so as not to express BCL11A protein at a detectable level. In other embodiments, the disclosure provides a cell population in which cells are modified so that BCL11A protein expression is reduced by at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% compared to unmodified cells. In yet another embodiment, the disclosure provides a cell population in which BCL11A protein expression is undetectable in the modified cell population.The effects of the modifications may be evaluated by Western blotting, flow cytometry, ELISA, cell-based assays, qRT-PCR, and electrochemiluminescence assays, and the sense transcripts may be analyzed by RNA fluorescence in situ hybridization (FISH) assays, or by other methods known in the art or other methods described in the examples.
[0259] In some embodiments, the Disclosure provides methods for modifying BCL11A target nucleic acid in a cell population by in vitro or ex vivo methods. The Method provides that cells can be obtained from the subject using any number of techniques known to those skilled in the art (e.g., by obtaining a bone marrow biopsy or peripheral blood sample). The desired cells may be separated from the rest of the sample, washed to remove fluids and debris, and optionally placed in a suitable buffer or medium for subsequent processing steps. The Method may comprise one or more of the following steps: i) introducing CasX:gRNA system components into cells for editing the target nucleic acid; ii) introducing nucleic acid or vector encoding the CasX:gRNA system components into cells; iii) growing the cells in a suitable medium under conditions suitable for their growth; and iv) cryopreserving the cells for subsequent administration to the subject. Thus, using the CasX:gRNA systems and methods described herein, various cells associated with abnormal hemoglobinopathy can be modified to produce a cell population in which BCL11A protein expression is reduced or eliminated and HbF is increased. Therefore, this approach can be used in therapeutic methods, particularly in subjects with abnormal hemoglobin disorders such as sickle cell anemia or β-thalassemia. In some cases, cells are brought into contact with CasX and gRNA, where the gRNA is a guide RNA (gRNA). In other cases, cells are brought into contact with CasX and gRNA, where the gRNA is a chimera containing DNA and RNA. In embodiments of any combination (combination of scaffold and targeting sequence which may constitute sgRNA or dgRNA) as described herein, each of the gRNA molecules may be provided as an RNP having CasX of the embodiments described herein for incorporation into the cells to be modified. In one embodiment, the target nucleic acid of a cell is modified by bringing the cell into contact with the CasX protein, a guide nucleic acid (gRNA) containing a targeting sequence complementary to the BCL11A target nucleic acid, and a donor template, such that the donor template is inserted into or replaces a portion of the target nucleic acid sequence of the cell so that the BCL11A protein is expressed at a level that is not expressed or is reduced.In other cases, the CasX and gRNA in the vector are delivered to a cell population (an embodiment of which is described above herein), and the target nucleic acid is modified to express the BCL11A protein at a level that is absent or reduced.
[0260] In some implementations, a cell population is contacted with a CasX variant containing the sequence in Table 4 or a sequence that is at least 65%, at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 86%, at least 87%, at least 88%, at least 89%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical, and the gRNA scaffold is contacted with the sequence in Table 3 or a CasX variant containing a sequence that is at least 65%, at least 70%, at least 75%, at least 80%, and The gRNA contains a sequence that is 81% identical, at least 82% identical, at least 83% identical, at least 84% identical, at least 85% identical, at least 86% identical, at least 86% identical, at least 87% identical, at least 88% identical, at least 89% identical, at least 89% identical, at least 90% identical, at least 91% identical, at least 92% identical, at least 93% identical, at least 94% identical, at least 95% identical, at least 96% identical, at least 97% identical, at least 98% identical, at least 99% identical, and at least 99.5% identical, and the gRNA contains a target sequence selected from the group consisting of SEQ ID NOs. 272-2100 and 2286-26789 in Table 1, or a sequence that is at least 65% identical, at least 70% identical, at least 75% identical, at least 80% identical, at least 85% identical, at least 90% identical, or at least 95% identical, and has 15-20 amino acids. In other cases, CasX and one or more gRNAs are introduced into a cell population as coding polynucleotides using a vector (such embodiments are described herein).Further methods for modifying cells using CasX:gRNA system components include viral infection, transfection, conjugation, protoplast fusion, particle gun technology, calcium phosphate precipitation, and direct microinjection. The choice of method generally depends on the cell type being transformed and the environment in which the transformation takes place (e.g., in vitro, ex vivo, or in vivo). A general discussion of these methods can be found in Ausubel, et al., Short Protocols in Molecular Biology, 3rd ed., Wiley & Sons, 1995.
[0261] During hybridization with target nucleic acids using CasX and gRNA, CasX introduces one or more single-strand or double-strand breaks into the BCL11A gene, resulting in modifications to the target nucleic acid, such as persistent indels (deletions or insertions) or other mutations (e.g., substitutions, duplications, or inversions), which, in conjunction with the host cell's repair mechanisms, lead to a corresponding reduction or elimination of the expression of the functional BCL11A protein, thereby creating a modified cell population. As described herein, CasX variants that introduce double-strand breaks into the target nucleic acid generate double-strand breaks within 18-26 nucleotides on the 5' end of the PAM site on the target strand and within 10-18 nucleotides on the 3' end of the non-target strand. Therefore, in some embodiments, the modifications obtained by this method may result in random insertions or deletions (indels), substitutions, duplications, or inversions of one or more nucleotides in those regions via the non-homologous DNA end-joining (NHEJ) repair mechanism.
[0262] In some embodiments of the method for modifying a cell population, the first gRNA contains a targeting sequence complementary to a sequence proximal to or within any one of the exons of the BCL11A gene. In one embodiment, the first gRNA contains a targeting sequence complementary to a sequence proximal to, within, or adjacent to any one of the regulatory elements of the BCL11A gene. In a particular embodiment, the first gRNA contains a targeting sequence complementary to a sequence within or adjacent to the 5' side of the GATA1 binding motif sequence of the BCL11A gene. In a particular embodiment, the targeting sequence is Sequence ID No. 22. As described above, disruption of the target nucleic acid sequence results in modification of the BCL11A gene such that the expression of functional BCL11A protein is reduced or eliminated in the modified cell population.
[0263] In some embodiments of this method, the target nucleic acid of a cell population is modified using multiple (e.g., two, three, four, or more) gRNAs that target different or overlapping portions of the BCL11A gene, and the CasX protein introduces multiple cleavages into the target nucleic acid sequence, resulting in persistent indels (deletions or insertions) or other mutations (e.g., substitutions, duplications, or inversions of one or more nucleotides), thereby reducing or eliminating the expression of functional BCL11A protein in the modified cell population.
[0264] In other embodiments, the Disclosure provides a modified cell population by contacting cells with a CasX protein, one or more gRNAs containing a targeting sequence, and a donor template, the donor template being inserted into a cleavage site introduced by a nuclease and replacing all or part of the target nucleic acid sequence of the BCL11A gene to be modified. In one embodiment described above, the donor template comprises at least a portion of the BCL11A exon, and one or more mutations and cell modification result in a gene modification, thereby reducing or eliminating the expression of functional BCL11A protein in the modified cell population. In another embodiment described above, the donor template comprises a sequence in or adjacent to the 5' side of the GATA1 binding motif sequence but having one or more mutations compared to the wild-type sequence, and cell modification results in a reduction or elimination of the expression of functional BCL11A protein in the modified cell population. In such cases, the substitution by the donor template is understood to be larger in the 5' and 3' directions than the location of the cleavage site introduced by the nuclease at a specific portion of the target nucleic acid to be substituted, and further includes homologous arms located on the 5' and 3' sides of the cleavage site introduced by the nuclease to facilitate insertion. In some cases, the donor template is a single-stranded DNA template or a single-stranded RNA template. In other cases, the donor template is a double-stranded DNA template. Insertion of the donor template in the target region can be mediated by homologous recombination repair (HDR, as described above) or homology-independent targeted integration (HITI). The exogenous sequence inserted by HITI can be of any length, e.g., a relatively short sequence of 10–50 nucleotides, or a longer sequence of about 50–1000 nucleotides. The donor template sequence may contain specific sequence differences compared to the genome sequence, such as restriction sites, nucleotide polymorphisms, barcodes, and selection markers (e.g., drug resistance genes, fluorescent proteins, enzymes, etc.), which can be used to assess the success of donor nucleic acid insertion at cleavage sites, or, in some cases, for other purposes (e.g., to indicate expression at a target genomic locus).Alternatively, these sequence differences may include adjacent recombinant sequences (e.g., FLP sequences, loxP sequences, etc.) that can be activated later to remove marker sequences.
[0265] In some embodiments of the method for modifying a cell population, the method further comprises contacting the target nucleic acid sequence of the BCL11A gene of the cell population with i) an additional CRISPR nuclease and gRNA that targets a different or overlapping portion of the BCL11A target nucleic acid compared to a first gRNA, ii) a polynucleotide encoding the additional CRISPR nuclease and gRNA of (i), iii) a vector comprising the polynucleotide of (ii), or iv) an XDP comprising the additional CRISPR nuclease and gRNA of (i), wherein the contact results in a modification of the BCL11A gene at a different position in the sequence compared to the sequence targeted by the first gRNA. In one embodiment described above, the additional CRISPR nuclease is a CasX protein having a different sequence from the CasX protein of the embodiment described above. In the aforementioned alternative embodiment, the additional CRISPR nuclease is selected from the group consisting of Cas9, Cas12a, Cas12b, Cas12c, Cas12d(CasY), Cas12j, Cask, Cas13a, Cas13b, Cas13c, Cas13d, Cas14, Cpfl, C2cl, Csn2, and their sequence variants, rather than being a CasX protein.
[0266] In other embodiments, the disclosure provides a method for in vivo modification of BCL11A target nucleic acid in a target cell population. In one embodiment of the in vivo modification method, the method includes administering the vector of the embodiment described herein to the subject at a therapeutically effective dose. In some embodiments, the vector is at least about 1 × 10⁻¹⁶ 5 Vector genome (vg / kg), at least approximately 1 × 10⁻⁶ 6 vg / kg, at least about 1 × 10⁻⁶ 7 vg / kg, at least about 1 × 10⁻⁶ 8 vg / kg, at least about 1 × 10⁻⁶ 9 vg / kg, at least about 1 × 10⁻⁶ 10vg / kg, at least about 1 × 10⁻⁶ 11 vg / kg, at least about 1 × 10⁻⁶ 12 vg / kg, at least about 1 × 10⁻⁶ 13 vg / kg, at least about 1 × 10⁻⁶ 14 vg / kg, at least about 1 × 10⁻⁶ 15 vg / kg, or at least about 1 × 10⁻⁶ 16 The subject is administered a dose of vg / kg. In other embodiments, the vector is at least about 1 × 10⁻¹⁶ 5 vg / kg ~ at least about 1 × 10⁻⁶ 16 vg / kg, or at least about 1 × 10⁻⁶ 6 vg / kg ~ approx. 1×10 15 vg / kg, or at least about 1 × 10⁻⁶ 7 vg / kg ~ approx. 1×10 14 vg / kg, or at least about 1 × 10⁻⁶ 8 vg / kg ~ approx. 1×10 14 The subject is administered a dose of vg / kg. In another embodiment of the in vivo modification method, the method comprises administering a therapeutically effective dose of XDP to a subject, wherein XDP comprises CasX and gRNA complexed in RNP, and optionally a donor template (as described in more detail above), and XDP has tropism to target cells and can deliver RNP for editing of the BCL11A gene as described herein. Embodiments of XDP used in the aforementioned editing method are described herein. In one embodiment, XDP comprises at least about 1 × 10⁻¹⁶ 5 particles / kg, at least about 1 × 10⁻⁶ 6 particles / kg, at least about 1 × 10⁻⁶ 7 particles / kg at least about 1 × 10 8 particles / kg, at least about 1 × 10⁻⁶ 9 particles / kg, at least about 1 × 10⁻⁶ 10 particles / kg, at least about 1 × 10⁻⁶ 11 particles / kg, at least about 1 × 10⁻⁶ 12 particles / kg, at least about 1 × 10⁻⁶ 13 particles / kg, at least about 1 × 10⁻⁶ 14 particles / kg, at least about 1 × 10⁻⁶ 15 particles / kg, at least about 1 × 10⁻⁶16 particles / kg, or at least about 1 × 10⁻⁶ 6 Particles / kg ~ approx. 1×10 15 particles / kg, or at least about 1 × 10⁻⁶ 7 Particles / kg ~ approx. 1×10 14 The substance is administered to the subject at a dose of particles / kg. In the embodiments described above in this paragraph, the vector or XDP is administered to the subject by a route of administration selected from intraparenchymal, intravenous, intraarterial, intraperitoneal, intracapsular, subcutaneous, intramuscular, intraabdominally, or a combination thereof, and the method of administration is injection, blood transfusion, or transplantation.
[0267] VIII. Treatment method In another aspect, the disclosure relates to a method of treating a subject in need of treatment for an abnormal hemoglobin disorder-related disease, including but not limited to sickle cell disease or β-thalassemia, wherein the suppression or elimination of BCL11A protein expression by modifying the BCL11A gene in target cells of the subject improves the signs, symptoms, or effects of the disease, even though the subject may still have an underlying disease.
[0268] Several therapeutic strategies have been used to design compositions for use in methods of treating subjects with abnormal hemoglobinosis-related disorders. In some embodiments, the method involves administering a therapeutically effective dose of a class 2, type V CRISPR nuclease and guide RNA disclosed herein to a subject with an abnormal hemoglobinosis (e.g., sickle cell anemia or β-thalassemia). In some embodiments, the treatment method involves administering to a subject a therapeutically effective dose of: i) a CasX:gRNA system comprising a first CasX protein and a first gRNA having a targeting sequence complementary to the target nucleic acid; ii) a CasX:gRNA system comprising a first CasX protein, a first gRNA having a targeting sequence complementary to the target nucleic acid, and a donor template; iii) a nucleic acid encoding the CasX:gRNA system of (i) or (ii); iv) a vector comprising the nucleic acid of (iii), which may be an AAV of any embodiment described herein; v) an XDP comprising the CasX:gRNA system of (i) or (ii); or vi) a combination of two or more of (i) to (v), wherein 1) the BCL11A gene in the target cells targeted by the first gRNA is modified (e.g., knocked down or knocked out) by the CasX protein and optionally the donor template; and 2) an increase in hemoglobin F (HbF) production occurs in the subject. In some embodiments, the therapeutic method further comprises administering a second gRNA or a nucleic acid encoding the second gRNA, wherein the second gRNA has a targeting sequence complementary to a different or overlapping portion of the target nucleic acid sequence compared to the first gRNA. In some cases, the cells targeted for modification are selected from the group consisting of hematopoietic stem cells (HSCs), hematopoietic progenitor cells (HPCs), CD34+ cells, mesenchymal stem cells (MSCs), induced pluripotent stem cells (iPSCs), myeloid common progenitor cells, proerythroblasts, and erythroblasts. In some embodiments, the subject to be treated is selected from the group consisting of rodents, mice, rats, and non-human primates. In another embodiment, the subject is human.
[0269] In some embodiments of the therapeutic method, the vector is an AAV vector encoding the CasX:gRNA system, and is at least about 1 × 10⁻⁶. 5 Vector genome (vg / kg), at least approximately 1 × 10⁻⁶ 6 vg / kg, at least about 1 × 10⁻⁶ 7 vg / kg, at least about 1 × 10⁻⁶ 8 vg / kg, at least about 1 × 10⁻⁶ 9 vg / kg, at least about 1 × 10⁻⁶ 10 vg / kg, at least about 1 × 10⁻⁶ 11 vg / kg, at least about 1 × 10⁻⁶ 12 vg / kg, at least about 1 × 10⁻⁶ 13 vg / kg, at least about 1 × 10⁻⁶ 14 vg / kg, at least about 1 × 10⁻⁶ 15 vg / kg, or at least about 1 × 10⁻⁶ 16 The subject is administered a dose of vg / kg. In other embodiments of this method, the AAV vector is at least about 1 × 10⁶ 5 vg / kg ~ approx. 1×10 16 vg / kg, at least about 1 × 10⁻⁶ 6 vg / kg ~ approx. 1×10 15 vg / kg, or at least about 1 × 10⁻⁶ 7 vg / kg ~ approx. 1×10 14 The subject is administered a dose of vg / kg. In other embodiments, the treatment method includes administering to the subject a therapeutically effective dose of XDP containing the CasX:gRNA system. In one embodiment, XDP is at least about 1 × 10⁻¹⁶ 5 particles / kg, at least about 1 × 10⁻⁶ 6 particles / kg, at least about 1 × 10⁻⁶ 7 particles / kg, at least about 1 × 10⁻⁶ 8 particles / kg, at least about 1 × 10⁻⁶ 9 particles / kg, at least about 1 × 10⁻⁶ 10 particles / kg, at least about 1 × 10⁻⁶ 11 particles / kg, at least about 1 × 10⁻⁶ 12 particles / kg, at least about 1 × 10⁻⁶ 13 particles / kg, at least about 1 × 10⁻⁶ 14 particles / kg, at least about 1 × 10⁻⁶ 15particles / kg, at least about 1 × 10⁻⁶ 16 It is administered to the subject at a dose of particles / kg. In another embodiment, XDP is at least about 1 × 10⁻⁶ 5 Particles / kg ~ approx. 1×10 16 particles / kg, or at least about 1 × 10⁻⁶ 6 Particles / kg ~ approx. 1×10 15 particles / kg, or at least about 1 × 10⁻⁶ 7 Particles / kg ~ approx. 1×10 14 The substance is administered to the subject at a dose of particles / kg. In the embodiments described above, the vector or XDP is administered to the subject by a route of administration selected from intraparenchymal, intravenous, intraarterial, intraperitoneal, intracapsular, subcutaneous, intramuscular, intraabdominally, or a combination thereof, and the method of administration is by injection, blood transfusion, or transplantation. Administration may be once, twice, or multiple times using a regimen schedule of weekly, bi-weekly, monthly, quarterly, or every six months.
[0270] In some embodiments, the therapeutic method involves administering to a subject a vector containing a polynucleotide encoding CasX and multiple gRNAs that target different or overlapping regions of the BCL11A gene, the administration resulting in contact of the target nucleic acid sequence with the expression product of the vector within the target cell, and the BCL11A gene being modified in the target cell. In other embodiments of the therapeutic method, the method involves administering to a subject a vector encoding a CasX protein and gRNA, further comprising a donor template, the administration resulting in modification of the target nucleic acid sequence in the target cell by cleavage by the CasX protein and insertion of the donor template into the target nucleic acid. In other embodiments, the method involves administering to a subject a first vector containing a polynucleotide encoding CasX and multiple gRNAs that target different or overlapping sequences of the BCL11A gene, and a second vector containing a donor template polynucleotide encoding at least part or all of the BCL11A gene, wherein the administration of the vectors results in contact of the target nucleic acid sequence in the subject's cells with the expression products of the CasX and gRNA vectors and the donor template, thereby modifying the BCL11A gene in the subject's cells as described herein. In some embodiments of the therapeutic method, the vector administered to the subject is an AAV vector as described herein. As stated above, the AAV vector is selected from AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV-Rh74, or AAVRh10. In some embodiments of the therapeutic method, the vector administered to the subject is an XDP as described herein, containing an RNP of the CasX:gRNA system.
[0271] In some embodiments of this method, modification involves introducing a single-strand break in the BCL11A gene of a cell population. In other embodiments, modification involves introducing a double-strand break in the BCL11A gene of a cell population. In some embodiments, modification involves introducing one or more mutations in the BCL11A target nucleic acid, such as one or more nucleotide insertions, deletions, substitutions, duplications, or inversions in the BCL11A gene, resulting in a reduction of BCL11A protein expression in the target cells by at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, or at least about 90% compared to unmodified cells. In some cases, the BCL11A gene in the target cells is modified so that at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% of the modified cells do not express BCL11A protein at a detectable level. In other cases of the treatment method, the modification results in an increase in the production of HbF in the subject's circulating blood by at least about 5%, at least about 10%, at least about 20%, at least about 30%, at least about 40%, or at least about 50% compared to the HbF level in the subject before treatment. In other embodiments, the method results in an HbF to hemoglobin S (HbS) ratio in the subject's circulating blood of at least 0.01:1.0, at least 0.025:1.0, at least 0.05:1.0, at least 0.075:1.0, at least 0.1:1.0, at least 0.2:1.0, at least 0.3:1.0, at least 0.4:1.0, at least 0.5:1:0, at least 0.75:1.0, at least 1.0:1.0, at least 1.25:1.0, at least 1.5:1.0, or at least 1.75:1.0. In other embodiments, the method results in an HbF level of at least about 5%, or at least about 10%, or at least about 20%, or at least about 30% of the total hemoglobin in the circulating blood of the subject. In the embodiments described above, the subject is selected from the group consisting of mice, rats, pigs, non-human primates, and humans.Methods for obtaining samples (e.g., body fluids or tissues) from treated subjects for analysis to determine the effectiveness of the treatment, and methods for preparing samples to enable analysis, are well known to those skilled in the art. Methods for analysis at the RNA and protein levels have been discussed above and are well known to those skilled in the art. The effectiveness of the treatment can also be evaluated by measuring biomarkers associated with the expression of target genes in the aforementioned body fluids, tissues, or organs collected from animals contacted with one or more compounds of the present invention, using routine clinical methods known in the art. Biomarkers for abnormal hemoglobin disorders include, but are not limited to, the percentage of sickle cells in circulating blood, BCL11A levels, BCL11A RNA, hemoglobin S levels, hemoglobin-gamma levels, and hemoglobin F levels.
[0272] In some embodiments, a method for treating abnormal hemoglobinopathy in a subject further comprises administering an effective therapeutic dose of an additional CRISPR nuclease or a polynucleotide encoding an additional CRISPR nuclease. In one embodiment, the additional CRISPR nuclease is a CasX protein having a different sequence from the first CasX. In another embodiment, the additional CRISPR nuclease is not the CasX protein (i.e., Cas9, Cas12a, Cas12b, Cas12c, Cas12d(CasY), Cas12j, Cas12k, Cas13a, Cas13b, Cas13c, Cas13d, Cas14, Cpfl, C2cl, Csn2) or a sequence variant thereof. In some embodiments, a method for treating abnormal hemoglobinopathy in a subject further comprises administering a chemotherapeutic agent.
[0273] In other embodiments, the Disclosure provides a method for treating a patient in need of treatment for a disease related to abnormal hemoglobinopathy by administering a therapeutically effective amount of a cell population modified in vitro or ex vivo by a CasX:gRNA system composition of an embodiment described herein, including: i) a CasX:gRNA system comprising a first CasX protein and a first gRNA having a targeting sequence complementary to a target nucleic acid; ii) a CasX:gRNA system comprising a first CasX protein, a first gRNA having a targeting sequence complementary to a target nucleic acid, and a donor template; iii) a nucleic acid encoding the CasX:gRNA system of (i) or (ii); iv) a vector comprising the nucleic acid of (iii), which may be an AAV of any embodiment described herein; v) an XDP comprising the CasX:gRNA system of (i) or (ii); or vi) a combination of two or more of (i) to (v). In one embodiment, the treatment method comprises i) isolating induced pluripotent stem cells (iPSCs) or hematopoietic stem cells (HSCs) from a subject; ii) modifying the BCL11A target nucleic acid of the iPSCs or HSCs by any method of the embodiments described herein; iii) differentiating the modified iPSCs or HSCs into hematopoietic progenitor cells; and iv) transplanting the hematopoietic progenitor cells into a subject with an abnormal hemoglobinopathy, the method resulting in an increase of at least about 5%, at least about 10%, at least about 20%, at least about 30%, at least about 40%, or at least about 50% in the subject's circulating blood HbF level compared to the subject's hemoglobin F (HbF) level before treatment. In some cases, the cells are autologous to the subject to which the cells are administered and are isolated from the subject's bone marrow or peripheral blood. In other cases, the cells are allogeneic to the subject to which the cells are administered and are isolated from the bone marrow or peripheral blood of a different subject. Modified cells can be transplanted into a target by transplantation, local injection, systemic injection, or a combination thereof.Methods for modifying cells for administration to a target are described herein, but in short, the modification involves contacting cells with i) a CasX:gRNA system comprising a first CasX protein and a first gRNA having a targeting sequence complementary to a target nucleic acid, ii) a CasX:gRNA system comprising a first CasX protein, a first gRNA having a targeting sequence complementary to a target nucleic acid, and a donor template, iii) a nucleic acid encoding the CasX:gRNA system of (i) or (ii), iv) a vector comprising the nucleic acid of (iii), which may be an AAV of any embodiment described herein, v) an XDP comprising the CasX:gRNA system of (i) or (ii), or vi) a combination of two or more of (i) to (v), such that the expression of BCL11A protein is reduced or the cells do not express BCL11A protein at a detectable level. In some embodiments, the method further comprises administering a second gRNA or a nucleic acid encoding a second gRNA, wherein the second gRNA has a targeting sequence complementary to a different or overlapping portion of the target nucleic acid sequence compared to the first gRNA. In some cases, CasX and gRNA as RNPs (embodiments thereof are described above herein), and optionally a donor template, are delivered to a cell population, and the target nucleic acid is modified to express BCL11A protein at a level where it is not expressed or is reduced. In other cases, CasX and gRNA in a vector are delivered to a cell population (embodiments thereof are described above herein), and the target nucleic acid is modified to express BCL11A protein at a level where it is not expressed or is reduced. In some embodiments, the cell population modified by the administration of the composition is eukaryotic cells selected from the group consisting of rodent cells, mouse cells, rat cells, and non-human primate cells. In some embodiments, the eukaryotic cells are human cells. In some embodiments, eukaryotic cells are selected from the group consisting of hematopoietic stem cells (HSCs), hematopoietic progenitor cells (HPCs), CD34+ cells, mesenchymal stem cells (MSCs), induced pluripotent stem cells (iPSCs), myeloid common progenitor cells, proerythroblasts, and erythroblasts.In some embodiments of the present method, the cells or their offspring administered to a subject survive in the subject for at least 1 month, 2 months, 3 months, 4 months, 5 months, 6 months, 7 months, 8 months, 9 months, 10 months, 11 months, 12 months, 13 months, 14 months, 15 months, 16 months, 17 months, 18 months, 19 months, 20 months, 21 months, 22 months, 23 months, 2 years, 3 years, 4 years, or 5 years after administration. In some embodiments, the therapeutic method of the present disclosure results in an increase in the level of circulating blood HbF of at least about 5%, at least about 10%, at least about 20%, at least about 30%, at least about 40%, or at least about 50% compared to the level of hemoglobin F (HbF) in the subject before treatment. In other embodiments, the method yields an HbF to hemoglobin S (HbS) ratio in a subject of at least 0.01:1.0, at least 0.025:1.0, at least 0.05:1.0, at least 0.075:1.0, at least 0.1:1.0, at least 0.2:1.0, at least 0.3:1.0, at least 0.4:1.0, at least 0.5:1:0, at least 0.75:1.0, at least 1.0:1.0, at least 1.25:1.0, at least 1.5:1.0, or at least 1.75:1.0. In other embodiments, the method yields an HbF level of at least about 5%, or at least about 10%, or at least about 20%, or at least about 30% of the total circulating hemoglobin in a subject.
[0274] In other embodiments, the Disclosure provides a method for increasing fetal hemoglobin (HbF) in a subject having a hemoglobin disorder by genome editing, comprising: i) administering an effective dose of the vector or XDP of the embodiments described herein to the subject, the vector or XDP delivering a CasX:gRNA system to the cells of the subject; ii) editing the BCL11A target nucleic acid of the cells of the subject with CasX targeted by the first gRNA; and iii) the editing introducing one or more nucleotide insertions, deletions, substitutions, duplications or inversions in the target nucleic acid sequence such that the expression of the BCL11A protein is reduced or eliminated, resulting in an increase of at least about 5%, at least about 10%, at least about 20%, at least about 30%, at least about 40%, or at least about 50% in the subject's circulating blood HbF level compared to the subject's hemoglobin F (HbF) level before treatment. In the foregoing, the cells are selected from the group consisting of hematopoietic stem cells (HSCs), hematopoietic progenitor cells (HPCs), CD34+ cells, mesenchymal stem cells (MSCs), induced pluripotent stem cells (iPSCs), myeloid common progenitor cells, proerythroblasts, and erythroblasts. In one embodiment of this method, the target nucleic acid of the cells is edited such that the expression of the BCL11A protein is reduced by at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, or at least about 90% compared to the target nucleic acid of unedited cells. In some cases, the subjects are selected from the group consisting of mice, rats, pigs, and non-human primates. In other cases, the subjects are humans.
[0275] In some embodiments of the method for treating abnormal hemoglobin disorders in subjects, the method results in improvement of at least one clinically relevant parameter selected from the group consisting of: incidence of end-organ disease, albuminuria, hypertension, debilitation, hypotonic urine, diastolic dysfunction, functional exercise capacity, acute coronary syndrome, pain events, pain severity, anemia, hemolysis, tissue hypoxia, organ dysfunction, abnormal hematocrit levels, childhood mortality, stroke incidence, hemoglobin levels compared to baseline, HbF levels, reduced incidence of pulmonary embolism, incidence of vascular occlusive crisis, hemoglobin S concentration in red blood cells, hospitalization rate, hepatic iron concentration, required transfusions, and quality of life score. In other embodiments of the method for treating abnormal hemoglobin disorders in subjects, the method results in improvements in at least two clinically relevant parameters selected from the group consisting of: incidence of end-organ disease, albuminuria, hypertension, debilitation, hypotonic urine, diastolic dysfunction, functional exercise capacity, acute coronary syndrome, pain events, pain severity, anemia, hemolysis, tissue hypoxia, organ dysfunction, abnormal hematocrit levels, childhood mortality, stroke incidence, hemoglobin levels compared to baseline, HbF levels, reduced incidence of pulmonary embolism, incidence of vascular occlusive crisis, hemoglobin S concentration in red blood cells, hospitalization rate, hepatic iron concentration, required transfusions, and quality of life scores.
[0276] In some embodiments, the therapeutic method involves administering liposomes or lipid nanoparticles containing CasX protein and gRNA to a target. In some embodiments, the liposomes or lipid nanoparticles further comprise a donor template of any of the embodiments described herein.
[0277] In some embodiments, the Disclosure provides a method for treating a subject having an abnormal hemoglobinopathy-related disorder, the method comprising administering to the subject, using a therapeutically effective dose, a CasX:gRNA composition or vector or XDP containing RNP of a CasX:gRNA composition, in a therapeutic regime comprising one or more consecutive doses. In some embodiments of the therapeutic regime, the therapeutically effective dose of the composition or vector is administered as a single dose. In other embodiments of the therapeutic regime, the therapeutically effective dose is administered to the subject in two or more doses over a period of at least two weeks, or at least one month, or at least two months, or at least three months, or at least four months, or at least five months, or at least six months. In some embodiments of the therapeutic regime, the effective dose is administered by a route selected from the group consisting of transplantation, local injection, systemic infusion, or a combination thereof.
[0278] In some embodiments, the treatment method further comprises administering a chemotherapeutic agent, which includes, but is not limited to, hydroxyurea, L-glutamine oral powder, voxelotol, and analgesics, and is effective in improving signs or symptoms associated with abnormal hemoglobinosis-related disorders.
[0279] In some embodiments, the Disclosure provides CasX:gRNA compositions, nucleic acids encoding CasX:gRNA compositions, nucleic acid-containing vectors, or XDPs containing RNPs of CasX:gRNA for use as pharmaceuticals for the treatment of abnormal hemoglobin disorders including sickle cell disease or β-thalassemia.
[0280] XIV. Kits and Compositions In other embodiments, kits are provided herein comprising a CasX protein, one or more gRNAs from any embodiment of the disclosure comprising a targeting sequence specific to the BCL11A gene, and a suitable container (e.g., a tube, vial, or plate). In some embodiments, the kit further comprises a buffer, a nuclease inhibitor, a protease inhibitor, liposomes, a therapeutic agent, a label, a label visualization reagent, or any combination thereof. In some embodiments, the kit further comprises a pharmaceutically acceptable carrier, diluent, or excipient. In some embodiments, the kit comprises a suitable control composition for gene modification application and instructions for use. In some embodiments, the kit comprises a vector, the vector comprising a sequence encoding the CasX protein of the disclosure, a gRNA of the disclosure, optionally a donor template, or a combination thereof.
[0281] In other embodiments of the Kits of this Disclosure, the Kit comprises a composition for the treatment of atypical hemoglobinopathy in a subject by modifying the BCL11A target nucleic acid in isolated cells of the subject, the modification comprising contacting the target nucleic acid sequence of the cell with i) a CasX:gRNA system, ii) a nucleic acid encoding a component of the CasX:gRNA system, iii) a nucleic acid-containing vector, iv) an XDP containing a CasX protein and a guide nucleic acid (gRNA), or v) any combination of (i) to (iv) as described in embodiments disclosed herein, i) the contact results in modification of the BCL11A target nucleic acid sequence by the CasX protein, ii) a decrease in the expression of the BCL11A protein, and iii) an increase in hemoglobin F (HbF) production at cell maturation. In some cases, the cells are induced pluripotent stem cells (iPSCs). In other cases, the cells are hematopoietic stem cells (HSCs). In one embodiment, the use of the composition results in a reduction in BCL11A protein expression by mature cells, by at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, or at least about 90% compared to the unmodified target nucleic acid. In another embodiment, BCL11A protein expression by mature cells is undetectable.
[0282] In some embodiments, the kit comprises multiple cells edited using the CasX:gRNA system described herein.
[0283] This specification describes numerous exemplary configurations, methods, parameters, etc. However, it should be recognized that such descriptions are not intended to limit the scope of this disclosure, but rather are provided as descriptions of exemplary embodiments.
[0284] Listed embodiments The present invention may be defined by referring to the exemplary embodiments listed below.
[0285] Set I Embodiment 1. A composition comprising a class 2 V-type CRISPR protein and a first guide nucleic acid (gNA), wherein the gNA comprises a targeting sequence complementary to the polypyrimidine tract-binding protein 1 (BCL11A) gene target nucleic acid sequence.
[0286] Embodiment 2. gNA is a. BCL11A intron, b. BCL11A exon, c. BCL11A intron-exon junction, d. BCL11A adjustment element, and The composition according to Embodiment 1, comprising a targeting sequence complementary to a target nucleic acid sequence selected from the group consisting of e. intergenetic regions.
[0287] Embodiment 3. The composition according to Embodiment 1, wherein the BCL11A gene contains a wild-type sequence.
[0288] Embodiment 4. The composition according to any one of Embodiments 1 to 3, wherein gNA is a guide RNA (gRNA).
[0289] Embodiment 5. The composition according to any one of Embodiments 1 to 3, wherein gNA is guide DNA (gDNA).
[0290] Embodiment 6. The composition according to any one of Embodiments 1 to 3, wherein gNA is a chimera containing DNA and RNA.
[0291] The composition according to any one of Embodiments 1 to 6, wherein the embodiment gNA is single-molecule gNA (sgnA).
[0292] Embodiment 8. The composition according to any one of Embodiments 1 to 6, wherein gNA is a dual-molecule gNA (dgNa).
[0293] Embodiment 9. The composition according to any one of Embodiments 1 to 8, wherein the targeting sequence of gNA includes a sequence selected from the group consisting of SEQ ID NOs 272 to 2100 and 2286 to 26789, or a sequence having at least about 65%, at least about 75%, at least about 85%, or at least about 95% identity thereto.
[0294] Embodiment 10. The composition according to any one of Embodiments 1 to 8, wherein the targeting sequence of gNA includes a sequence selected from the group consisting of SEQ ID NOs: 272 to 2100 and 2286 to 26789.
[0295] Embodiment 11. The composition according to any one of Embodiments 1 to 8, wherein the targeting sequence of gNA includes sequences of SEQ ID NOs. 272-2100 and 2286-26789, with a single nucleotide removed from the 3' end of the sequence.
[0296] Embodiment 12. The composition according to any one of Embodiments 1 to 8, wherein the targeting sequence of gNA includes sequences of SEQ ID NOs. 272-2100 and 2286-26789, with two nucleotides removed from the 3' end of the sequence.
[0297] Embodiment 13. The composition according to any one of Embodiments 1 to 8, wherein the targeting sequence of gNA includes sequences of SEQ ID NOs. 272-2100 and 2286-26789, with three nucleotides removed from the 3' end of the sequence.
[0298] Embodiment 14. The composition according to any one of Embodiments 1 to 8, wherein the targeting sequence of gNA includes sequences of SEQ ID NOs. 272-2100 and 2286-26789, with four nucleotides removed from the 3' end of the sequence.
[0299] Embodiment 15. The composition according to any one of Embodiments 1 to 8, wherein the targeting sequence of gNA includes sequences of sequence numbers 272-2100 and 2286-26789, with five nucleotides removed from the 3' end of the sequence.
[0300] Embodiment 16. The composition according to any one of Embodiments 1 to 15, wherein the targeting sequence of gNA is complementary to the sequence of the BCL11A exon.
[0301] Embodime...
Claims
1. A system comprising a chimeric CasX variant protein and a guide ribonucleic acid (gRNA) variant, wherein the gRNA variant is a) A scaffold containing the sequence of sequence number 2238 or a sequence having at least 90% sequence identity therewith, and b) Targeting sequences complementary to the target nucleic acid sequence containing the B-cell lymphoma / leukemia 11A (BCL11A) gene. A system that includes this.
2. The aforementioned gRNA, a) BCL11A intron, b) BCL11A exon, c) BCL11A intron-exon junction, d) BCL11A adjustment element, and e) Intergenetic regions, The system according to claim 1, comprising a targeting sequence complementary to a target nucleic acid sequence selected from the group consisting of the following.
3. The system according to claim 1, wherein the BCL11A gene includes a wild-type sequence.
4. The system according to claim 1, wherein the gRNA variant is a single molecule gRNA (sgRNA).
5. The system according to claim 1, wherein the targeting sequence has a single nucleotide removed from its 3' end, or two, three, four, or five nucleotides removed from its 3' end.
6. The system according to claim 1, wherein the targeting sequence of the gRNA variant is complementary to the sequence of the BCL11A regulatory element, and the regulatory element is selected from the group consisting of promoter and enhancer regulatory elements of the BCL11A gene.
7. The system according to claim 6, wherein the targeting sequence of the gRNA variant is complementary to the sequence of an enhancer regulatory element selected from the group consisting of the GATA1 erythrocyte-specific enhancer binding site (GATA1) of the BCL11A gene and a sequence located 5' to the GATA1 binding site of the BCL11A gene.
8. The system according to claim 7, wherein the targeting sequence of the gRNA variant includes a sequence that has at least 90%, at least 95%, or at least 100% sequence identity with a sequence selected from the group consisting of UGGAGCCUGUGAAUAAAAGCA (SEQ ID NO: 22), UGCUUUUAUCACAGGCUUCCA (SEQ ID NO: 23), CAGGCUUCCAGGGAAGGUUGUUG (SEQ ID NO: 2949), GAGGCCAAACCCUUCCUUGGA (SEQ ID NO: 2948), AGUGCCAAGCUAACAGUUGCU (SEQ ID NO: 15747), and AUACACUUUGAAGGCUAGUC (SEQ ID NO: 15748).
9. The system according to claim 1, wherein the targeting sequence is ligated to the 3' end of the scaffold of the gRNA variant.
10. The system according to claim 1, wherein the chimeric CasX variant protein comprises the sequence of SEQ ID NO: 126 or a sequence having at least 90% sequence identity therewith.
11. The chimeric CasX variant protein further comprises one or more nuclear localization signals (NLS), The one or more NLSs mentioned above, a) At or near the C-terminus of the chimeric CasX variant protein b) At or near the N-terminus of the chimeric CasX variant protein, c) The N-terminus or vicinity thereof and the C-terminus or vicinity thereof of the chimeric CasX variant protein to be located, The system according to claim 1.
12. The chimeric CasX variant protein can form a ribonucleoprotein complex (RNP) with a gRNA variant, and the RNPs of the chimeric CasX variant protein and the gRNA variant exhibit at least one improved feature compared to the RNP of a reference CasX protein of SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 3 and a reference gRNA containing one of the sequences of SEQ ID NOs: 4-16, wherein the one or more improved features include improved folding of the chimeric CasX variant protein, improved binding affinity to gRNA, improved binding affinity to target DNA, and, in editing target DNA, one or more protospacer adjacent motifs including ATC, CTC, GTC, or TTC. The system according to claim 1, selected from one or more of the following: improved ability to utilize (PAM) sequences, improved unwinding of the target DNA, increased editing activity, improved editing efficiency, improved editing specificity, increased nuclease activity, improved cleavage rate of target nucleic acid sequences, increased target strand loading for double-strand breaks, decreased target strand loading for single-strand nicking, decreased off-target breaks, improved binding of non-target DNA strands, improved protein stability, improved protein solubility, improved ribonucleoprotein complex (RNP) formation, higher percentage cleavage-competent RNPs, improved stability of protein:gRNA complexes (RNPs), improved solubility of protein:gRNA complexes, improved protein yield, improved protein expression, and improved fusion characteristics.
13. The system according to claim 12, wherein if any one of the PAM sequences TTC, ATC, GTC, or CTC is located one nucleotide 5' to the non-target strand of a protospacer having identity with the targeting sequence of the gRNA in the cell assay system, the RNP containing the chimeric CasX variant protein and the gRNA variant exhibits higher editing efficiency and / or binding of the target nucleic acid sequence compared to the editing efficiency and / or binding of the RNP containing the reference CasX protein and reference gRNA in the comparative assay system.
14. The PAM sequence is TTC, and the targeting sequence of the gRNA variant includes a sequence selected from the group consisting of SEQ ID NOs: 22, 23, 2949, 2948, 15747, and 15748. The system according to claim 13.
15. The system according to claim 13, wherein the increased binding affinity to one or more PAM sequences is at least 1.5 times greater than the binding affinity to any one of the reference CasX proteins of SEQ ID NOs. 1 to 3 for the PAM sequences, and / or the RNPs have a percentage of cleavage-competent RNPs that are at least 5%, at least 10%, at least 15%, or at least 20% higher than the RNPs of the reference CasX protein and the reference gRNAs of SEQ ID NOs. 4 to 16.
16. One or more nucleic acids comprising a sequence encoding a chimeric CasX variant protein and a gRNA variant of the system according to any one of claims 1 to 15.
17. A nucleic acid comprising a sequence encoding a gRNA variant according to any one of claims 1 to 9.
18. A vector comprising a gRNA variant according to any one of claims 1 to 9, a chimeric CasX variant protein according to any one of claims 11 to 15, or one or more nucleic acids encoding the chimeric CasX variant protein or encoding or containing the gRNA variant.
19. The vector according to claim 18, wherein the vector is selected from the group consisting of retroviral vectors, lentiviral vectors, adenovirus vectors, adeno-associated virus (AAV) vectors, herpes simplex virus (HSV) vectors, virus-like particles (VLPs), plasmids, minicircles, nanoplasmides, DNA vectors, RNA vectors, lipid nanoparticles, and liposomes.
20. The vector according to claim 18, wherein the vector is an AAV vector, and the AAV vector is selected from AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV-Rh74, and AAVRh10.
21. A method for modifying a BCL11A target nucleic acid sequence in a cell population in vitro or ex vivo, the method comprising introducing into the cell population one or more nucleic acids comprising or encoding the system and a chimeric CasX variant protein and gRNA variant of the system according to any one of claims 1 to 9 and 11 to 15, wherein the target nucleic acid sequence of the BCL11A gene in the cells targeted by the gRNA variant is modified by the chimeric CasX variant protein, the modification comprising introducing one or more nucleotide insertions, deletions, substitutions, duplications or inversions in the BCL11A gene of the cell population, and the cells are not human germ cells or embryos, in vitro or ex vivo method.
22. The method according to claim 21, wherein the GATA1 binding site sequence of the target nucleic acid is modified.
23. a) The expression of the BCL11A protein is reduced by at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, or at least 90% compared to cells in which the BCL11A gene has not been modified, or b) so that at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, or at least 90% of the modified cells do not express BCL11A protein at a detectable level. The method according to claim 21, wherein the BCL11A gene of the cell population is modified.
24. The method according to claim 21, wherein the cells are selected from the group consisting of hematopoietic stem cells (HSCs), hematopoietic progenitor cells (HPCs), CD34+ cells, mesenchymal stem cells (MSCs), myeloid common progenitor cells, proerythroblasts, and erythroblasts.
25. A system according to any one of claims 1 to 9 and 11 to 15, or one or more nucleic acids encoding a chimeric CasX variant protein of the system, or comprising or encoding a gRNA variant, for use in the treatment of abnormal hemoglobin disorders.
26. A system according to any one of claims 1 to 9 and 11 to 15, or one or more nucleic acids encoding a chimeric CasX variant protein of the system, or comprising or encoding a gRNA variant, for use in the manufacture of a pharmaceutical for the treatment of abnormal hemoglobin disorders.