Systems and methods for treatment of hemoglobinopathies

JP2025081309A5Pending Publication Date: 2025-10-30EDITAS MEDICINE INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025006855
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2018-11-29
Filing Date
2025-01-17
Publication Date
2025-10-30

AI Technical Summary

Technical Problem

Current treatments for sickle cell disease (SCD) and β-thalassemia are limited by the risks associated with hematopoietic stem cell transplantation and the need for lifelong transfusions and chelation therapy, which come with complications such as infection risks and iron overload.

Method used

The use of genome editing systems, specifically CRISPR-mediated methods, to modify the promoter region of γ-globin genes (HBG1 and HBG2) to increase the expression of fetal hemoglobin (HbF), thereby addressing the underlying cause of these disorders.

Benefits of technology

Increased expression of HbF can alleviate symptoms of SCD and β-thalassemia by reducing the polymerization of sickle hemoglobin and improving hemoglobin production, potentially leading to improved clinical outcomes and reduced treatment complications.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

To provide genome editing systems and methods for altering a target nucleic acid sequence, or modulating expression of a target nucleic acid sequence, and methods for the alteration of genes encoding hemoglobin subunits and / or treatment of hemoglobinopathies.SOLUTION: The present invention provides an RNP complex comprising a CRISPR from Prevotella and Francisella1 (Cpf1) RNA guided nuclease or a variant thereof and a gRNA, wherein the gRNA is capable of binding to a target site in a promoter of an HBG gene in a cell.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Reference to Related Applications This application claims the benefit of U.S. Provisional Patent Application No. 62 / 643,168, filed Mar. 14, 2018; U.S. Provisional Patent Application No. 62 / 767,488, filed Nov. 14, 2018; and U.S. Provisional Patent Application No. 62 / 773,073, filed Nov. 29, 2018, the entire contents of each of which are incorporated herein by reference.

[0002] Sequence Listing This application includes a Sequence Listing submitted in ASCII format via EFS-Web, the entire contents of which are incorporated herein by reference. The ASCII copy, created on Mar. 14, 2019, is named 2019-03-14_EM115PCT_Editas_8013WO02_Sequence_Listing.txt and is 811 KB in size.

[0003] This disclosure relates to genome editing systems and methods for modifying a target nucleic acid sequence or regulating the expression of a target nucleic acid sequence, and their use in modifying genes encoding hemoglobin subunits and / or treating abnormal hemoglobinopathies.

Background Art

[0004] Hemoglobin (Hb) transports oxygen in red blood cells (RBCs) from the lungs to tissues. From during prenatal development through immediately after birth, hemoglobin exists in the form of fetal hemoglobin (HbF), a tetrameric protein composed of two α-globin chains and two γ-globin chains. Through a process known as globin switching, most of the HbF is replaced by adult hemoglobin (HbA), a tetrameric protein in which the γ-globin chains of HbF have been replaced by β-globin chains. The average adult makes less than 1% HbF from total hemoglobin (Thein 2009). The α-hemoglobin gene is located on chromosome 16, while the β-hemoglobin gene (HBB), the Aγ-globin chain (HBG1, also known as γ-globin A), and the Gγ-globin chain (HBG2, also known as γ-globin G) are located on chromosome 11 within the globin gene cluster (also referred to as the globin locus).

[0005] Mutations in HBB can cause hemoglobin disorders (i.e., abnormal hemoglobins), including sickle cell disease (SCD) and β-thalassemia (β-Thal). In the United States, approximately 93,000 people are diagnosed with abnormal hemoglobins. Worldwide, 300,000 children are born each year with abnormal hemoglobins (Angastiniotis 1998). Because these medical conditions are associated with HBB mutations, these symptoms typically do not manifest until the globin switches from HbF to HbA.

[0006] SCD is the most common genetic blood disorder in the United States, affecting approximately 80,000 people (Brousseau 2010). SCD is most common in people of African ancestry, with a prevalence of 1 in 500. In Africa, the prevalence of SCD is 15 million (Aliyu 2008). SCD is also more common in people of Indian, Saudi Arabian, and Mediterranean descent. In Hispanic Americans, the prevalence of sickle cell disease is 1 in 1,000 (Lewis 2014).

[0007] SCD is caused by one homozygous mutation of c.17A>T (HbS mutation) in the HBB gene. The sickle mutation is a point mutation (GAG>GTG) on HBB, which results in the substitution of valine for glutamic acid at amino acid position 6 in exon 1. Valine at position 6 of the β-hemoglobin chain is hydrophobic and causes a conformational change in the β-globin protein when not oxygen-bound. This conformational change causes polymerization of the HbS protein in the absence of oxygen, leading to deformation of RBCs (i.e., sickling). Since SCD is autosomal recessive, only patients with two HbS alleles have this disease. Heterozygous subjects have the sickle cell trait and may suffer from anemia and / or painful episodes when placed in severe dehydration or oxygen-deficient states.

[0008] Sickled RBCs cause multiple symptoms including anemia, sickle cell crises, onset of vaso-occlusion, aplastic episodes, and acute chest syndrome. Sickled RBCs are less elastic than wild-type red blood cells and thus cannot easily pass through the capillary bed, causing occlusion and ischemia (i.e., vaso-occlusion). Vaso-occlusive disease occurs when sickled red blood cells impede blood flow in the capillary beds of organs, resulting in pain, ischemia, and necrosis. These episodes typically last 5 - 7 days. The spleen plays a role in removing dysfunctional RBCs and thus typically enlarges in infancy and is frequently subject to vaso-occlusive episodes. By the end of childhood, the spleens of SCD patients often infarct, leading to autosplenectomy. Hemolysis is a constant feature of SCD and results in anemia. Sickle cells survive in circulation for 10 - 20 days, while healthy red blood cells survive for 90 - 120 days. SCD subjects are transfused as needed to maintain an appropriate hemoglobin concentration. Frequent transfusions expose subjects to the risk of HIV, hepatitis B, and hepatitis C infections. Subjects may also have problems with acute chest attacks and infarcts in the extremities, end organs, and central nervous system.

[0009] Subjects with SCD have a reduced average lifespan. The prognosis of patients with SCD has been steadily improved by careful management throughout their lives of the onset and anemia. As of 2001, the average remaining lifespan of subjects with sickle cell disease is in the mid- to late 50s. Current SCD treatment methods involve hydration and pain management during the onset and transfusion to correct anemia as needed.

[0010] Thalassemias (e.g., β-Thal, δ-Thal, and β / δ-Thal) cause chronic anemia. β-Thal is estimated to affect approximately 1 in 100,000 people worldwide. Its prevalence is high in certain populations including those of European descent, where it is approximately 1 in 10,000. β-Thal major is a more severe form of the disease and is life-threatening unless treated with lifelong transfusion and chelation therapy. In the United States, there are approximately 3,000 subjects with β-Thal major. β-Thal intermedia does not require transfusion, but it can cause growth retardation and significant systemic abnormalities and often requires lifelong chelation therapy. HbA makes up most of the hemoglobin in adult RBCs, but approximately 3% of adult hemoglobin is in the form of HbA2, a variant of HbA in which two γ-globin chains are replaced by two δ (Δ)-globin chains. δ-Thal is associated with mutations in the Δ-globin gene (HBD) that cause loss of HBD expression. Co-inheritance of HBD mutations can mask the diagnosis of β-Thal (i.e., β / δ-Thal) by reducing HbA2 levels to the normal range (Bouva 2006). β / δ-Thal is usually caused by deletions in the HBB and HBD sequences in both alleles. In homozygous (δo / δo βo / βo) patients, HBG is expressed, resulting in the production of HbF only.

[0011] Similar to SCD, β-Thal is caused by mutations in the HBB gene. The most common HBB mutations that result in β-Thal are c.-136C>G, c.92+1G>A, c.92+6T>C, c.93-21G>A, c.118C>T, c.316-106C>G, c.25_26delAA, c.27_28insG, c.92+5G>C, c.118C>T, c.135delC, c.315+1G>A, c.-78A>G, c.52A>T, c.59A>G, c.92+5G>C, c.124_127delTTCT, c.316-197C>T, c.-78A>G, c.52A>T, c.124_127delTTCT, c.316-197C>T, c.-138C>T, c.-79A>G, c.92+5G>C, c.75T>A, c.316-2A>G and c.316-2A>C. These and other mutations associated with β-Thal cause mutations or deletions in the β-globin chain, which disrupts the ratio of normal Hbα-hemoglobin to β-hemoglobin. Excess α-globin chains precipitate in the erythroid precursors of the bone marrow.

[0012] In β-Thal major, both alleles of HBB contain nonsense, frameshift or splicing mutations that result in a complete lack of β-globin production (denoted as β0 / β0). β-Thal major results in a severe reduction in β-globin chains, leading to marked precipitation of α-globin chains in RBCs and high-severity anemia.

[0013] β-Thal intermedia results from mutations in the 5’ or 3’ untranslated regions of HBB, mutations in the promoter region or polyadenylation signal of HBB, or splicing mutations within the HBB gene. The patient's genotype is denoted as βo / β+ or β+ / β+. βo represents the absence of β-globin chain expression; β+ represents β-globin chains that are dysfunctional but present. Phenotypic expression varies among patients. Since there is some production of β-globin in β-Thal intermedia, it results in less α-globin chain precipitation in erythroid precursors and anemia of lower severity than β-Thal major. However, the proliferation of the erythroid lineage secondary to chronic anemia has more serious consequences.

[0014] Subjects with β-Thal major develop at 6 months to 2 years of age and suffer from growth retardation, fever, hepatosplenomegaly, and diarrhea. Appropriate treatment includes regular blood transfusions. Treatments for β-Thal major also include splenectomy and treatment with hydroxyurea. If the patient receives regular blood transfusions, they develop normally until the beginning of their 20s. At that point, the patient requires chelation therapy (in addition to continued blood transfusions) to prevent complications of iron overload. Iron overload can manifest as delayed growth or delayed sexual maturation. In adulthood, inadequate chelation therapy can result in cardiomyopathy, cardiac arrhythmias, hepatic fibrosis and / or cirrhosis, diabetes, thyroid and parathyroid abnormalities, thrombosis, and osteoporosis. Frequent blood transfusions also expose the subject to the risks of HIV, hepatitis B, and hepatitis C infections.

[0015] Subjects with β-Thal intermedia generally present between 2 and 6 years of age. Subjects generally do not require blood transfusions. However, bone abnormalities occur due to chronic hyperplasia of the erythroid lineage to compensate for chronic anemia. Subjects may have fractures of long bones due to osteoporosis. Extramedullary erythropoiesis is common and results in swelling of the spleen, liver, and lymph nodes. This can also cause spinal cord compression and neurological problems. Subjects are at risk of developing thrombotic events such as lower limb ulcers, stroke, pulmonary embolism, and deep vein thrombosis. Treatment methods for β-Thal intermedia include splenectomy, folic acid supplementation, hydroxyurea treatment, and radiotherapy for extramedullary tumors. Chelation therapy is used in subjects who develop iron overload.

[0016] In β-Thal patients, life expectancy is often reduced. Subjects with β-Thal major who do not receive blood transfusion therapy generally die in their 20s or 30s. Subjects with β-Thal major who receive regular blood transfusions and appropriate chelation therapy can survive beyond their 50s. Heart failure secondary to iron toxicity is the main cause of death in β-Thal major subjects due to iron toxicity. Summary of the Invention

Problems to be Solved by the Invention

[0017] Currently, various novel therapies for SCD and β-Thal are being developed. Currently, the delivery of anti-sickling HBB genes via gene therapy is being investigated in clinical trials. However, the long-term efficacy and safety of this approach are unknown. Transplantation of hematopoietic stem cells (HSCs) from HLA-matched allogeneic stem cell donors has been demonstrated to cure SCD and β-Thal, but this procedure is associated with risks including those related to ablation therapy necessary to prepare the subject for transplantation, increasing the risk of life-threatening opportunistic infections and the risk of graft-versus-host disease after transplantation. Furthermore, a compatible allogeneic donor is often not identifiable. Therefore, there is a need for improved methods for managing these and other abnormal hemoglobinopathies.

Means for Solving the Problems

[0018] As used herein, provided are genome editing systems, ribonucleoprotein (RNP) complexes, guide RNAs, Cpf1 proteins including modified Cpf1 proteins (Cpf1 variants), and CRISPR-mediated methods for modifying the promoter region of one or more γ-globin genes (e.g., HBG1, HBG2, or both HBG1 and HBG2) to increase the expression of fetal hemoglobin (HbF). In certain embodiments, the RNP complex can include a guide RNA (gRNA) complexed with a wild-type Cpf1 or a modified Cpf1 RNA-guided nuclease (modified Cpf1 protein). In certain embodiments, the gRNA can include a sequence set forth in Table 13, Table 18, or Table 19. In certain embodiments, the gRNA can include a gRNA targeting domain. In certain embodiments, the gRNA targeting domain can include a sequence selected from the group consisting of SEQ ID NOs: 1002, 1254, 1258, 1260, 1262, and 1264. In certain embodiments, the gRNA can include a gRNA sequence set forth in Table 19. In certain embodiments, the gRNA can include a sequence selected from the group consisting of SEQ ID NOs: 1022, 1023, and 1041-1105. In certain embodiments, the RNP complex can include an RNP complex set forth in Table 21. For example, the RNP complex can include a gRNA including the sequence set forth in SEQ ID NO: 1051 and a modified Cpf1 protein encoded by the sequence set forth in SEQ ID NO: 1097 (RNP32 in Table 21).

[0019] The inventors have discovered herein that delivery of an RNP complex comprising a gRNA complexed with a modified Cpf1 protein can result in increased editing of a target nucleic acid. In certain embodiments, the modified Cpf1 protein can contain one or more modifications. In certain embodiments, the one or more modifications can include, without limitation, one or more mutations in the wild-type Cpf1 amino acid sequence, one or more mutations in the wild-type Cpf1 nucleic acid sequence, one or more mutations in the nuclear localization signal (NLS), one or more mutations in a purification tag (e.g., His tag), or combinations thereof. In certain embodiments, the modified Cpf1 can be encoded by a sequence set forth in SEQ ID NOs: 1000, 1001, 1008-1018, 1032, 1035-1039, 1094-1097, 1107-1109 (Cpf1 polypeptide sequences) or SEQ ID NOs: 1019-1021, 1110-1117 (Cpf1 polynucleotide sequences). In certain embodiments, the gRNA can be a modified or unmodified gRNA. In certain embodiments, the gRNA can include a sequence set forth in Table 13, Table 18, or Table 19. In certain embodiments, the RNP complex can include an RNP complex set forth in Table 21. For example, the RNP complex can include a gRNA comprising the sequence set forth in SEQ ID NO: 1051 and a modified Cpf1 protein encoded by the sequence set forth in SEQ ID NO: 1097 (RNP32 of Table 21). In certain embodiments, an RNP complex comprising a modified Cpf1 protein can increase editing of a target nucleic acid. In certain embodiments, an RNP complex comprising a modified Cpf1 protein can increase editing and result in an increase in productive indels. In various embodiments, the increase in editing of the target nucleic acid can be evaluated by any means known to those of skill in the art, such as PCR amplification of the target nucleic acid followed by sequence analysis (e.g., Sanger sequencing, next-generation sequencing), but is not limited thereto.

[0020] The inventors have also discovered herein that delivery of an RNP complex comprising a modified gRNA complexed with an unmodified or modified Cpf1 protein can result in increased editing of a target nucleic acid. In certain embodiments, the modified gRNA can comprise one or more modifications including phosphorothioate bond modifications, phosphorodithioate (PS2) bond modifications, 2'-O-methyl modifications, one, or more, or a stretch of deoxyribonucleic acid (DNA) bases (also referred to herein as "DNA extension"), one, or more, or a stretch of ribonucleic acid (RNA) bases (also referred to herein as "RNA extension") or combinations thereof. In certain embodiments, the DNA extension can comprise the sequences set forth in Table 24. For example, in certain embodiments, the DNA extension can comprise the sequences set forth in SEQ ID NOs: 1235-1250. In certain embodiments, the RNA extension can comprise the sequences set forth in Table 24. For example, in certain embodiments, the RNA extension can comprise the sequences set forth in SEQ ID NOs: 1231-1234, 1251-1253. In certain embodiments, the gRNA can comprise the sequences set forth in Table 13, Table 18, or Table 19. In certain embodiments, the RNP complex can comprise the RNP complex set forth in Table 21. For example, the RNP complex can comprise a gRNA comprising the sequence set forth in SEQ ID NO: 1051 and a modified Cpf1 protein encoded by the sequence set forth in SEQ ID NO: 1097 (RNP32 of Table 21). In certain embodiments, the RNP complex comprising the modified gRNA can increase editing of the target nucleic acid. In certain embodiments, the RNP complex comprising the modified gRNA can increase editing resulting in an increase in productive indels.

[0021] In certain embodiments, an RNP complex comprising a modified gRNA and a modified Cpf1 protein can increase editing of a target nucleic acid. In certain embodiments, an RNP complex comprising a modified gRNA and a modified Cpf1 protein can increase editing resulting in an increase in productive indels.

[0022] The inventors have also discovered that co-delivery of an RNP complex (e.g., "gRNA-Cpf1-RNP") containing a gRNA complexed with a Cpf1 molecule and a "booster element" can result in an increase in the editing of a target nucleic acid. In certain embodiments, the RNP complex can comprise the RNP complex described in Table 21. For example, the RNP complex can comprise a gRNA comprising the sequence set forth in SEQ ID NO: 1051 and a modified Cpf1 protein encoded by the sequence set forth in SEQ ID NO: 1097 (RNP32 in Table 21). As used herein, the term "booster element" refers to an element that increases the editing of a target nucleic acid when co-delivered with an RNP complex (a "gRNA-nuclease-RNP") comprising a gRNA complexed with an RNA-guided nuclease, as compared to the editing of the target nucleic acid without the booster element. In certain embodiments, one or more booster elements can be co-delivered with the gRNA-nuclease-RNP complex to increase the editing of the target nucleic acid. In certain embodiments, co-delivery of the booster element can increase the editing and result in an increase in productive indels. In various embodiments, the increase in the editing of the target nucleic acid can be evaluated by any means known to those of skill in the art, such as PCR amplification of the target nucleic acid followed by sequence analysis (e.g., Sanger sequencing, next-generation sequencing), but is not limited thereto.

[0023] In certain embodiments, the gRNA-nuclease-RNP may comprise a gRNA-Cpf1-RNP. In certain embodiments, the Cpf1 molecule of the gRNA-Cpf1-RNP complex can be a wild-type Cpf1 or a modified Cpf1. In certain embodiments, the Cpf1 molecule of the gRNA-Cpf1-RNP can be encoded by the sequences set forth in SEQ ID NOs: 1000, 1001, 1008-1018, 1032, 1035-1039, 1094-1097, 1107-1109 (Cpf1 polypeptide sequences) or SEQ ID NOs: 1019-1021, 1110-1117 (Cpf1 polynucleotide sequences). In certain embodiments, the gRNA-Cpf1-RNP complex may comprise a gRNA comprising a targeting domain set forth in Table 13 or Table 18. In certain embodiments, the gRNA-Cpf1-RNP complex may comprise a gRNA comprising the sequences set forth in Table 19. In certain embodiments, the gRNA can be a modified or unmodified gRNA.

[0024] In certain embodiments, the booster element can comprise a dead RNP (“dead gRNA-nuclease-RNP”) comprising a dead gRNA molecule complexed to an RNA-guided nuclease molecule. In certain embodiments, the dead gRNA-nuclease-RNP can comprise a dead gRNA complexed to a wild-type (WT) Cas9 molecule (“dead gRNA-Cas9-RNP”), a dead gRNA complexed to a Cas9 nickase molecule (“dead gRNA-nickase-RNP”), or a dead gRNA complexed to an enzymatically inactive (ei) Cas9 molecule (“dead gRNA-eiCas9-RNP”). In certain embodiments, the dead gRNA-nuclease-RNP complex can have a decrease in nuclease activity or a loss of activity. In certain embodiments, the dead gRNA of the dead gRNA-nuclease-RNP complex can comprise any of the dead gRNAs described herein. For example, a dead gRNA that can comprise a targeting domain can be identical to or differ by 3 nucleotides or less from a dead gRNA targeting domain described in Table 10 or Table 15. In certain embodiments, the dead gRNA can comprise a targeting domain that comprises truncation of the gRNA targeting domain. In certain embodiments, the gRNA targeting domain that is truncated can be a gRNA targeting domain described in Table 2, Table 10, or Table 15. In certain embodiments, the dead gRNA can be a modified or unmodified dead gRNA. As shown herein, co-delivery of gRNA-Cpf1-RNP and dead gRNA-Cas9-RNP (i.e., an RNP comprising a dead gRNA complexed to WT Cas9) or co-delivery of gRNA-Cpf1-RNP and dead gRNA-nickase-RNP (i.e., an RNP comprising a dead gRNA complexed to a Cas9 nickase (i.e., Cas9 D10A nickase)) resulted in an increase in overall editing beyond levels observed following delivery of gRNA-Cpf1-RNP alone (see, e.g., Examples 15, 17, 18). The dead gRNA molecule can comprise a targeting domain that is complementary to a region proximal to or within a target region in a target nucleic acid (e.g., a CCAAT box target region, a 13 nt target region, a proximal HBG1 / 2 promoter target sequence, and / or a GATA1 binding motif in BCL11Ae).In certain embodiments, "proximal" may refer to a region within 10, 25, 50, 100, or 200 nucleotides of a target region (e.g., a CCAAT box target region, a 13 nt target region, a proximal HBG1 / 2 promoter target sequence, and / or a GATA1 binding motif in BCL11Ae). In certain embodiments, one or more booster elements may be co-delivered with the gRNA-Cpf1-RNP and may include one or more dead gRNA-nuclease-RNPs such as a dead gRNA-Cas9-RNP, a dead gRNA-nickase-RNP, a dead gRNA-eiCas9-RNP. In certain embodiments, co-delivery of the dead gRNA-nuclease-RNP does not modify the indel profile of the gRNA-Cpf1-RNP.

[0025] In certain embodiments, the booster element can comprise an RNP complex ("gRNA-nickase-RNP") comprising a gRNA molecule complexed to an RNA-guided nuclease nickase molecule. In certain embodiments, the RNA-guided nuclease nickase molecule can be a Cas9 nickase molecule such as, for example, Cas9 D10A nickase. In certain embodiments, the gRNA of the gRNA-nickase-RNP can comprise any of the gRNAs described herein. For example, the gRNA can comprise a gRNA targeting domain described in Table 2, Table 10, or Table 15. In certain embodiments, the gRNA can be a modified or unmodified gRNA. As shown herein, co-delivery of gRNA-Cpf1-RNP and gRNA-nickase-RNP complex (an RNP comprising a guide RNA complexed to a Cas9 D10A nickase molecule) resulted in an increase in overall editing beyond the levels observed following delivery of gRNA-Cpf1-RNP alone (see, e.g., Examples 15, 16). Further, co-delivery of the gRNA-nickase-RNP complex and the gRNA-Cpf1-RNP complex modified the directionality, length, and / or position of the indel profile of the gRNA-Cpf1-RNP. In certain embodiments, a booster enhancer can be used to provide a desired editing outcome such as, for example, an increase in the ratio of productive indels. In certain embodiments, co-delivery of the gRNA-nickase-RNP complex and the gRNA-Cpf1-RNP complex can modify the indel profile of the gRNA-Cpf1-RNP.

[0026] In certain embodiments, the booster element can comprise a single-stranded oligodeoxynucleotide (ssODN) or a double-stranded oligodeoxynucleotide (dsODN). In certain embodiments, the ssODN can be any ssODN disclosed herein. In certain embodiments, the ssODN can comprise a sequence described in Table 11. For example, in certain embodiments, the ssODN can comprise the sequence set forth in SEQ ID NO: 1040.

[0027] In one aspect, the present disclosure relates to an RNP complex comprising a CRISPR from a Prevotella and Francisella 1 (Cpf1) RNA-guided nuclease or a variant thereof and a gRNA capable of binding to a target site in the promoter of the intracellular HBG gene. In certain embodiments, the gRNA can be modified or unmodified. In certain embodiments, the gRNA can include one or more modifications including, but not limited to, phosphorothioate bond modification, phosphorodithioate (PS2) bond modification, 2'-O-methyl modification, DNA extension, RNA extension, or combinations thereof. In certain embodiments, the DNA extension can include the sequences described in Table 24. In certain embodiments, the RNA extension can include the sequences described in Table 24. In certain embodiments, the gRNA can include the sequences described in Table 13, Table 18, or Table 19. In certain embodiments, the RNP complex can include the RNP complex described in Table 21. For example, the RNP complex can include a gRNA comprising the sequence set forth in SEQ ID NO: 1051 and a Cpf1 variant protein encoded by the sequence set forth in SEQ ID NO: 1097 (RNP32 of Table 21). In certain embodiments, the Cpf1 variant protein can contain one or more modifications. In certain embodiments, the one or more modifications can include, without limitation, one or more mutations of the wild-type Cpf1 amino acid sequence, one or more of the wild-type Cpf1 nucleic acid sequences, one or more nuclear localization signals (NLSs), one or more mutations of a purification tag (e.g., His tag), or combinations thereof. In certain embodiments, the Cpf1 variant protein can be encoded by the sequences set forth in SEQ ID NOs: 1000, 1001, 1008-1018, 1032, 1035-39, 1094-1097, 1107-09 (Cpf1 polypeptide sequences) or SEQ ID NOs: 1019-1021, 1110-17 (Cpf1 polynucleotide sequences).

[0028] In one aspect, the present disclosure relates to a method of modifying a promoter of a HBG gene in a cell, the method comprising contacting the cell with an RNP complex disclosed herein. In certain embodiments, the modification may include indels within one or more regions described in Table 17. In certain embodiments, the modification may include indels within the CCAAT box target region of the promoter of the HBG gene. For example, in certain embodiments, the modification may include indels within Chr11(NC_000011.10):5,249,955~5,249,987 (Table 17, Region 6), Chr11(NC_000011.10):5,254,879~5,254,909 (Table 17, Region 16) or combinations thereof. In certain embodiments, the RNP complex may include a gRNA and a Cpf1 protein. In certain embodiments, the gRNA may include an RNA targeting domain described in Table 19. In certain embodiments, the gRNA targeting domain may include a sequence selected from the group consisting of SEQ ID NOs: 1002, 1254, 1258, 1260, 1262 and 1264. In certain embodiments, the gRNA may include a gRNA sequence described in Table 19. In certain embodiments, the gRNA may include a sequence selected from the group consisting of SEQ ID NOs: 1022, 1023, 1041~1105. In certain embodiments, the gRNA may be configured to provide editing events at Chr11:5249973, Chr11:5249977 (HBG1); Chr11:5250042, Chr11:5250046 (HBG1); Chr11:5250055, Chr11:5250059 (HBG1); Chr11:5250179, Chr11:5250183 (HBG1); Chr11:5254897, Chr11:5254901 (HBG2); Chr11:5254897, Chr11:5254901 (HBG2); Chr11:5254966, 5254970 (HBG2); Chr11:5254979, 5254983 (HBG2) (Table 22, Table 23). In certain embodiments, the cell may be further contacted with a booster element. In certain embodiments, the booster element may include a single-stranded oligodeoxynucleotide (ssODN) or a double-stranded oligodeoxynucleotide (dsODN).In certain embodiments, the ssODN can be any ssODN disclosed herein. In certain embodiments, the ssODN can include the sequences set forth in Table 11. For example, in certain embodiments, the ssODN can include the sequence set forth in SEQ ID NO: 1040.

[0029] In one aspect, the present disclosure relates to an isolated cell comprising a modification in the promoter of the HBG gene generated by delivery of an RNP complex into the cell. In certain embodiments, the RNP complex can comprise a gRNA and a Cpf1 protein. In certain embodiments, the gRNA can be modified or unmodified. In certain embodiments, the gRNA can comprise one or more modifications including, but not limited to, phosphorothioate bond modification, phosphorodithioate (PS2) bond modification, 2'-O-methyl modification, DNA extension, RNA extension, or combinations thereof. In certain embodiments, the DNA extension can comprise the sequences described in Table 24. In certain embodiments, the RNA extension can comprise the sequences described in Table 24. In certain embodiments, the gRNA can comprise the sequences described in Table 13, Table 18, or Table 19. In certain embodiments, the RNP complex can comprise the RNP complex described in Table 21. For example, the RNP complex can comprise a gRNA comprising the sequence set forth in SEQ ID NO: 1051 and a Cpf1 variant protein encoded by the sequence set forth in SEQ ID NO: 1097 (RNP32 in Table 21). In certain embodiments, the Cpf1 variant protein can contain one or more modifications. In certain embodiments, the one or more modifications can include, without limitation, one or more mutations of the wild-type Cpf1 amino acid sequence, one or more mutations of the wild-type Cpf1 nucleic acid sequence, one or more nuclear localization signals (NLSs), one or more mutations of a purification tag (e.g., His tag), or combinations thereof. In certain embodiments, the Cpf1 variant protein can be encoded by the sequences set forth in SEQ ID NO: 1000, 1001, 1008-1018, 1032, 1035-1039, 1094-1097, 1107-1109 (Cpf1 polypeptide sequences) or SEQ ID NO: 1019-1021, 1110-1117 (Cpf1 polynucleotide sequences). In certain embodiments, a booster element can be co-delivered with the RNP complex. In certain embodiments, the booster element can comprise a single-stranded oligodeoxynucleotide (ssODN) or a double-stranded oligodeoxynucleotide (dsODN). In certain embodiments, the ssODN can be any ssODN disclosed herein.In certain embodiments, the ssODN may comprise the sequences set forth in Table 11. For example, in certain embodiments, the ssODN may comprise the sequence set forth in SEQ ID NO: 1040.

[0030] In one aspect, the present disclosure relates to an in vitro method of increasing the level of fetal hemoglobin (HbF) in human cells by genome editing using an RNP complex comprising a gRNA and a Cpf1 RNA-guided nuclease or a variant thereof, which affects modifications in the promoter of the HBG gene, thereby increasing the expression of HbF. In certain embodiments, the gRNA can be modified or unmodified. In certain embodiments, the gRNA can include one or more modifications including phosphorothioate bond modification, phosphorodithioate (PS2) bond modification, 2'-O-methyl modification, DNA extension, RNA extension, or combinations thereof. In certain embodiments, the DNA extension can include the sequences described in Table 24. In certain embodiments, the RNA extension can include the sequences described in Table 24. In certain embodiments, the gRNA can include the sequences described in Table 13, Table 18, or Table 19. In certain embodiments, the RNP complex can include the RNP complex described in Table 21. For example, the RNP complex can include a gRNA comprising the sequence set forth in SEQ ID NO: 1051 and a Cpf1 variant protein encoded by the sequence set forth in SEQ ID NO: 1097 (RNP32 in Table 21). In certain embodiments, the Cpf1 variant protein can contain one or more modifications. In certain embodiments, the one or more modifications can include, without limitation, one or more mutations in the wild-type Cpf1 amino acid sequence, one or more mutations in the wild-type Cpf1 nucleic acid sequence, one or more nuclear localization signals (NLSs), one or more mutations in a purification tag (e.g., His tag), or combinations thereof. In certain embodiments, the Cpf1 variant protein can be encoded by the sequences set forth in SEQ ID NOs: 1000, 1001, 1008-1018, 1032, 1035-1039, 1094-1097, 1107-1109 (Cpf1 polypeptide sequences) or SEQ ID NOs: 1019-1021, 1110-1117 (Cpf1 polynucleotide sequences). In certain embodiments, the booster element can be co-delivered with the RNP complex. In certain embodiments, the booster element can include a single-stranded oligodeoxynucleotide (ssODN) or a double-stranded oligodeoxynucleotide (dsODN).In certain embodiments, the ssODN can be any ssODN disclosed herein. In certain embodiments, the ssODN can include the sequences set forth in Table 11. For example, in certain embodiments, the ssODN can include the sequence set forth in SEQ ID NO: 1040.

[0031] In one aspect, the present disclosure relates to a population of CD34+ or hematopoietic stem cells, wherein one or more cells of the population comprise a modification in the promoter of the HBG gene, the modification being generated by delivering an RNP complex comprising a gRNA and a Cpf1 RNA-guided nuclease or variant thereof to the population of CD34+ or hematopoietic stem cells. In certain embodiments, the gRNA can be modified or unmodified. In certain embodiments, the gRNA can comprise one or more modifications including, but not limited to, phosphorothioate bond modification, phosphorodithioate (PS2) bond modification, 2'-O-methyl modification, DNA extension, RNA extension, or combinations thereof. In certain embodiments, the DNA extension can comprise the sequences described in Table 24. In certain embodiments, the RNA extension can comprise the sequences described in Table 24. In certain embodiments, the gRNA can comprise the sequences described in Table 13, Table 18, or Table 19. In certain embodiments, the RNP complex can comprise the RNP complexes described in Table 21. For example, the RNP complex can comprise a gRNA comprising the sequence set forth in SEQ ID NO: 1051 and a Cpf1 variant protein encoded by the sequence set forth in SEQ ID NO: 1097 (RNP32 of Table 21). In certain embodiments, the Cpf1 variant protein can contain one or more modifications. In certain embodiments, the one or more modifications can include, without limitation, one or more mutations of the wild-type Cpf1 amino acid sequence, one or more mutations of the wild-type Cpf1 nucleic acid sequence, one or more nuclear localization signals (NLSs), one or more mutations of a purification tag (e.g., His tag), or combinations thereof. In certain embodiments, the Cpf1 variant protein can be encoded by the sequences set forth in SEQ ID NO: 1000, 1001, 1008-1018, 1032, 1035-1039, 1094-1097, 1107-1109 (Cpf1 polypeptide sequences) or SEQ ID NO: 1019-1021, 1110-1117 (Cpf1 polynucleotide sequences). In certain embodiments, the booster element can be co-delivered with the RNP complex.In certain embodiments, the booster element can comprise single-stranded oligodeoxynucleotide (ssODN) or double-stranded oligodeoxynucleotide (dsODN). In certain embodiments, the ssODN can be any ssODN disclosed herein. In certain embodiments, the ssODN can comprise the sequences set forth in Table 11. For example, in certain embodiments, the ssODN can comprise the sequence set forth in SEQ ID NO: 1040.

[0032] In one aspect, the present disclosure relates to a method of alleviating one or more symptoms of sickle cell disease in a subject in need thereof, the method comprising: a) isolating a population of CD34+ or hematopoietic stem cells from the subject; b) modifying the isolated population of cells ex vivo by delivering an RNP complex comprising a gRNA and a Cpf1 RNA-guided nuclease or variant thereof, thereby affecting modification of the promoter of the HBG gene in one or more cells of the population; and c) administering the modified population of cells to the subject, thereby alleviating one or more symptoms of sickle cell disease in the subject. In certain embodiments, the modification can comprise an indel within the CCAAT box target region of the promoter of the HBG gene. In certain embodiments, the RNP complex can be delivered using electroporation. In certain embodiments, at least about 5%, at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80% or at least about 90% of the cells in the population of cells comprise productive indels.

[0033] In one aspect, the present disclosure relates to a gRNA comprising a 5' end and a 3' end, and including DNA extension at the 5' end and 2'-O-methyl-3'-phosphorothioate modification at the 3' end, the gRNA comprising an RNA segment capable of hybridizing to a target site and an RNA segment capable of binding to a Cpf1 RNA-guided nuclease. In certain embodiments, the DNA extension may comprise the sequences set forth in SEQ ID NOs: 1235-1250. In certain embodiments, the gRNA may be modified or unmodified. In certain embodiments, the gRNA may comprise one or more modifications including phosphorothioate linkage modification, phosphorodithioate (PS2) linkage modification, 2'-O-methyl modification, DNA extension, RNA extension, or combinations thereof. In certain embodiments, the DNA extension may comprise the sequences set forth in Table 24. In certain embodiments, the RNA extension may comprise the sequences set forth in Table 24. In certain embodiments, the gRNA may comprise the sequences set forth in Table 13, Table 18, or Table 19.

[0034] In one aspect, the present disclosure relates to an RNP complex comprising a Cpf1 RNA-guided nuclease as disclosed herein and a gRNA as disclosed herein.

[0035] Also provided herein are genome editing systems, guide RNAs, and CRISPR-mediated methods for modifying one or more γ-globin genes (e.g., HBG1, HBG2, or HBG1 and HBG2), the erythroid-specific enhancer of the BCL11A gene (BCL11Ae), or combinations thereof, and increasing the expression of fetal hemoglobin (HbF). In certain embodiments, one or more gRNAs comprising the sequences set forth in Table 18 or Table 19 are used, and modifications can be introduced into the promoter region of the HBG gene. In certain embodiments, the genome editing systems, guide RNAs, and CRISPR-mediated methods can modify a 13 nucleotide (nt) target region (the "13nt target region") that is 5' of the transcription site of the HBG1, HBG2, or HBG1 and HBG2 genes. In certain embodiments, the genome editing systems, guide RNAs, and CRISPR-mediated methods can modify a CCAAT box target region (the "CCAAT box target region") that is 5' of the transcription site of the HBG1, HBG2, or HBG1 and HBG2 genes. In certain embodiments, the CCAAT box target region can be the region of or in the vicinity of the distal CCAAT box, and includes the nucleotides of the distal CCAAT box, 25 nucleotides upstream (5') and 25 nucleotides downstream (3') of the distal CCAAT box (i.e., HBG1 / 2 c.-86~-140). In certain embodiments, the CCAAT box target region can be the region of or in the vicinity of the distal CCAAT box, and includes the nucleotides of the distal CCAAT box, 5 nucleotides upstream (5') and 5 nucleotides downstream (3') of the distal CCAAT box (i.e., HBG1 / 2 c.-106~-120). In certain embodiments, the CCAAT box target region can include an 18nt target region, a 13nt target region, an 11nt target region, a 4nt target region, a 1nt target region, a -117G>A target region, or combinations thereof as disclosed herein. In certain embodiments, the modification can be an 18nt deletion, a 13nt deletion, an 11nt deletion, a 4nt deletion, a 1nt deletion, a substitution of G to A at c.-117 of the HBG1, HBG2, or HBG1 and HBG2 genes, or combinations thereof. In certain embodiments, the modification can be a non-naturally occurring modification or a naturally occurring modification.In certain embodiments, one or more gRNAs comprising the targeting domains set forth in SEQ ID NOs: 251-901 or 940-942 can be used to introduce modifications into a 13 nt target region. In certain embodiments, one or more gRNAs comprising the sequences set forth in SEQ ID NOs: 251-901, 940-942, 996, 997, 970, 971, 1002 or 1003 can be used to introduce modifications into a CCAAT box target region. In certain embodiments, the genome editing system, guide RNA and CRISPR-mediated method can modify the GATA1 binding motif in BCL11Ae (the “GATA1 binding motif in BCL11Ae”) in the +58 DNaseI hypersensitive site (DHS) region of intron 2 of the BCL11A gene. In certain embodiments, one or more gRNAs comprising the targeting domains set forth in SEQ ID NOs: 952-955 can be used to introduce modifications into the GATA1 binding motif of BCL11Ae. In certain embodiments, one or more gRNAs can be used to introduce a GATA1 binding motif modification into BCL11Ae, and one or more gRNAs can be used to introduce modifications into the 13 nt target regions of HBG1 and / or HBG2.

[0036] In certain embodiments of the present specification, the use of optional genomic editing system components such as template nucleic acids (oligonucleotide donor templates) is also provided. In certain embodiments, examples of template nucleic acids for use in CCAAT target region targeting may include, without limitation, template nucleic acids encoding modifications of the CCAAT box target region. In certain embodiments, the CCAAT box target region may include an 18 nt target region, an 11 nt target region, a 4 nt target region, a 1 nt target region, or combinations thereof. In certain embodiments, the template nucleic acid can be a single-stranded oligodeoxynucleotide (ssODN) or a double-stranded oligodeoxynucleotide (dsODN). In certain embodiments, exemplary full-length donor templates encoding modifications in the 5' and 3' homology arms and the CCAAT box target region are also presented below (e.g., SEQ ID NOs: 904-909, 974-995). In certain embodiments, the template nucleic acid can be a plus strand or a minus strand. In certain embodiments, the ssODN can include a 5' homology arm, a substitution sequence, and a 3' homology arm. In certain embodiments, the 5' homology arm can be, for example, about 25 to about 200 nucleotides or more in length, such as at least about 25, 50, 75, 100, 125, 150, 175, or 200 nucleotides; the substitution sequence can include a length of 0 nucleotides; the 3' homology arm can be, for example, about 25 to about 200 nucleotides or more in length, such as at least about 25, 50, 75, 100, 125, 150, 175, or 200 nucleotides. In certain embodiments, the ssODN can include one or more phosphorothioates.

[0037] In certain embodiments, a genomic editing system, guide RNA, and CRISPR-mediated method for modifying one or more γ-globin genes (e.g., HBG1, HBG2, or HBG1 and HBG2) can include an RNA-guided nuclease. In certain embodiments, the RNA-guided nuclease can be Cas9 or a modified Cas9. In certain embodiments, the RNA-guided nuclease can be Cpf1 or a modified Cpf1 as disclosed herein.

[0038] In one aspect, the present disclosure relates to a composition comprising a plurality of cells generated by the method disclosed above, wherein at least 20%, 30%, 40%, 50%, 60%, 70%, 80% or 90% of the cells comprise a modification in the sequence of the human HBG1 or HBG2 gene or the 13 nt target region of the plurality of cells generated by the method disclosed above, at least 20%, 30%, 40%, 50%, 60%, 70%, 80% or 90% of the cells comprise a modification in the sequence of the 13 nt target region of the human HBG1 or HBG2 gene, and at least 20%, 30%, 40%, 50%, 60%, 70%, 80% or 90% of the cells comprise a modification in the GATA1 binding motif sequence of BCL11Ae. In certain embodiments, at least a portion of the plurality of cells can be within the erythroid lineage. In certain embodiments, the plurality of cells can be characterized by an increased level of fetal hemoglobin expression compared to an unmodified plurality of cells. In certain embodiments, the level of fetal hemoglobin can be increased by at least 20%, 30%, 40%, 50%, 60%, 70%, 80% or 90%. In certain embodiments, the composition can further comprise a pharmaceutically acceptable carrier.

[0039] In one aspect, the present disclosure relates to a method of modifying a cell, the method comprising the steps of unwinding a chromatin segment within or proximal to a target region of a nucleic acid within the cell, and generating a double-strand break (DSB) within the target region of the nucleic acid, thereby modifying the target region. In certain embodiments, the step of unwinding the chromatin segment may include the step of contacting the chromatin segment with an RNA-guided helicase. In certain embodiments, the step of unwinding the chromatin does not include the step of recruiting an exogenous trans-acting factor to the chromatin segment. The RNA-guided helicase can be an RNA-guided nuclease, and the RNA-guided nuclease can complex with a dead-guide RNA (dgRNA) comprising a first targeting domain sequence of 15 nucleotides or less in length. In certain embodiments, the dgRNA can include 5' or 3' end modifications including, but not limited to, an anti-reverse cap analog (ARCA) at the 5' end of the RNA, a 3' polyA tail at the RNA end, or both. In certain embodiments, the RNA-guided helicase can be an enzymatically active RNA-guided nuclease or can be configured to lack nuclease activity. In certain embodiments, the targeting domain sequence of the dgRNA can be complementary to a sequence proximal to the target region. In some embodiments herein, "proximal" can mean within 10, 25, 50, 100, or 200 nucleotides of the target region. In certain embodiments, the step of unwinding the chromatin segment may not include the step of forming a single-stranded or double-stranded break in the nucleic acid within the chromatin segment. In certain embodiments, the step of generating a DSB within the target region may include the step of contacting the chromatin segment with an RNA-guided nuclease having nuclease activity. In certain embodiments, the RNA-guided nuclease having nuclease activity can complex with a gRNA comprising a targeting domain configured to overlap the target region. In certain embodiments, the RNA-guided nuclease having nuclease activity can be a Cpf1 molecule.

[0040] Another aspect of the disclosure is a method of inducing accessibility of a nucleic acid to a target region for editing within a cell, the method comprising contacting the cell with an RNA-guided helicase and a dgRNA, and unwinding DNA within or proximal to the target region with the RNA-guided helicase, thereby inducing accessibility to the target region for editing. In various cases, the RNA-guided helicase and the dgRNA can be configured to bind within or proximal to the target region. In certain embodiments, the dgRNA can be configured not to provide an RNA-guided nuclease cleavage event. In certain embodiments, the RNA-guided helicase and the dgRNA can complex to form a dead ribonucleoprotein (RNP) lacking cleavage activity. In certain embodiments, the dgRNA can include a targeting domain sequence that is 15 nucleotides or less in length. In certain embodiments, the RNA-guided helicase can be an RNA-guided nuclease. In certain embodiments, the RNA-guided nuclease and the dgRNA are not configured to recruit an exogenous trans-acting factor to the target region. In certain embodiments, the RNA-guided nuclease can be Cas9 or a Cas9 fusion protein. In certain embodiments, Cas9 can be enzymatically active Cas9 or enzymatically dead Cas9. In certain embodiments, Cas9 can be a nickase, such as, for example, Cas9 D10A. In certain embodiments, the step of unwinding the DNA does not include forming a single-stranded or double-stranded break in the DNA. In certain embodiments, an RNA-guided nuclease having nuclease activity can complex with a gRNA comprising a targeting domain configured to overlap the target region. In certain embodiments, the RNA-guided nuclease having nuclease activity can be a Cpf1 molecule.

[0041] In another aspect, the present disclosure is a method of increasing the rate of indel formation in a nucleic acid, comprising using an RNA-guided helicase configured to bind within or proximal to a target region of the nucleic acid to unwind double-stranded DNA within or proximal to the target region of the nucleic acid, and generating a DSB within the target region. In certain embodiments, the step of generating a DSB within the target region results in the formation of an indel in the target region. In certain embodiments, the DSB can be repaired in a manner that forms an indel in the target region. In certain embodiments, the rate of indel formation in a gene achieved using an RNA-guided helicase is increased compared to the rate of indel formation in a gene achieved without using an RNA-guided helicase. In certain embodiments, the RNA-guided helicase can form an RNP complex with a dgRNA configured to bind within or proximal to the target region. In certain embodiments, the dgRNA can include a targeting domain sequence that is 15 nucleotides or less in length. In certain embodiments, the RNA-guided helicase can be an RNA-guided nuclease. In certain embodiments, the RNA-guided nuclease can be Cas9 or a Cas9 fusion protein. In certain embodiments, Cas9 can be enzymatically active Cas9 or enzymatically dead Cas9. In certain embodiments, Cas9 can be a nickase, such as, for example, Cas9 D10A. In certain embodiments, the RNA-guided nuclease and the dgRNA are not configured to recruit an exogenous trans-acting factor to the target region. In certain embodiments, the step of unwinding the double-stranded DNA does not include the step of forming a single-stranded or double-stranded break in the DNA.

[0042] In yet another aspect, the present disclosure relates to a method of deleting a segment of a target nucleic acid in a cell, the method comprising contacting the cell with an RNA-guided helicase and generating a DSB within the target region, whereby a segment of the target nucleic acid is deleted. In certain embodiments, the DSB can be repaired in a manner that deletes a segment of the target nucleic acid. In certain embodiments, the RNA-guided helicase can be configured to bind within or proximal to the target region of the target nucleic acid and unwind double-stranded DNA (dsDNA) within or proximal to the target region. In certain embodiments, the RNA-guided helicase can form a ribonucleoprotein complex with a dgRNA configured to bind within or proximal to the target region. In certain embodiments, the dgRNA can include a targeting domain sequence that is 15 nucleotides or less in length. In certain embodiments, the RNA-guided helicase can be an RNA-guided nuclease. In certain embodiments, the RNA-guided nuclease can be Cas9 or a Cas9 fusion protein. In certain embodiments, Cas9 can be enzymatically active Cas9 or enzymatically dead Cas9. In certain embodiments, the RNA-guided nuclease and the dgRNA are not configured to recruit exogenous trans-acting factors to the target region. In certain embodiments, the target nucleic acid can be a promoter region of a gene, a coding region of a gene, a non-coding region of a gene, an intron of a gene, or an exon of a gene. In certain embodiments, the segment of the target nucleic acid can be at least about 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, or 100 base pairs in length.

[0043] The present disclosure also relates to a dead gRNA (dgRNA) molecule comprising a targeting domain that includes truncation of the gRNA targeting domain. In certain embodiments, the gRNA targeting domain to be truncated can be the gRNA targeting domain described in Table 2, Table 10, or Table 15. In certain embodiments, the gRNA targeting domain can be truncated from the 5' end of the gRNA targeting domain. In certain embodiments, the dgRNA can include a targeting domain sequence that is 15 nucleotides or less in length. In certain embodiments, the first targeting domain can be identical to the dgRNA targeting domain described in Table 10 or Table 15, or can differ by 3 nucleotides or less.

[0044] Another aspect of the present disclosure relates to a composition comprising at least one polynucleotide encoding a plurality of gRNAs and one RNA-induced helicase, wherein at least one gRNA can be a dgRNA configured not to provide an RNA-induced nuclease cleavage event. In certain embodiments, the dgRNA can include a targeting domain sequence that is 15 nucleotides or less in length. In certain embodiments, the RNA-induced helicase can be an RNA-induced nuclease. In certain embodiments, the RNA-induced nuclease can be Cas9 or a Cas9 fusion protein. In certain embodiments, Cas9 can be enzymatically active Cas9 or enzymatically dead Cas9. In certain embodiments, Cas9 can be a nickase, such as, for example, Cas9 D10A. In certain embodiments, the RNA-induced nuclease and the dgRNA are not configured to recruit exogenous trans-acting factors to the target region. In certain embodiments, the composition further comprises a second RNA-induced nuclease configured to provide a cleavage event. In certain embodiments, the composition further comprises a second gRNA configured to provide a cleavage event.

[0045] In another aspect, the present disclosure relates to a genome editing system comprising an RNA-guided nuclease and an RNA-guided helicase configured to bind to a target nucleic acid proximal to a target region of the target nucleic acid to induce a conformational change in the target region, thereby facilitating access of the RNA-guided nuclease to the target region and forming a cleavage site in the target region. The present disclosure also relates to a genome editing system comprising a dgRNA comprising a targeting domain sequence of 15 nucleotides or less in length, a first RNA-guided nuclease, and an RNA-guided helicase. In certain embodiments, the genome editing system further comprises a gRNA. In certain embodiments, the gRNA and the first RNA-guided nuclease can bind to a target region in a target nucleic acid. In certain embodiments, the first RNA-guided nuclease can be a Cpf1 molecule. In certain embodiments, the gRNA and the first RNA-guided nuclease can bind to a first PAM sequence in the target nucleic acid, wherein the first PAM sequence is outward-facing. In certain embodiments, the RNA-guided helicase can be a second RNA-guided nuclease. In certain embodiments, the second RNA-guided nuclease and the dgRNA are not configured to recruit an exogenous trans-acting factor to the target region. In certain embodiments, the dgRNA and the second RNA-guided nuclease bind within or proximal to the target region in the target nucleic acid. In certain embodiments, the first RNA-guided nuclease and the second RNA-guided nuclease can complex with the gRNA and the dgRNA, respectively, to form first and second ribonucleoprotein complexes.

[0046] In another aspect, the present disclosure relates to a genome editing system comprising a dgRNA comprising a targeting domain sequence of 15 nucleotides or less in length, a first RNA-guided nuclease, and an RNA-guided helicase. In certain embodiments, the gRNA and the first RNA-guided nuclease can bind to a target region in a target nucleic acid. In certain embodiments, the first RNA-guided nuclease can be a Cpf1 molecule. In certain embodiments, the gRNA and the first RNA-guided nuclease can bind to a first protospacer adjacent motif (PAM) sequence in the target nucleic acid. In certain embodiments, the first PAM sequence can be outward-facing. In certain embodiments, the RNA-guided helicase can be a second RNA-guided nuclease. In certain embodiments, the second RNA-guided nuclease and the dgRNA are not configured to recruit an exogenous trans-acting factor to the target region. In certain embodiments, the dgRNA and the second RNA-guided nuclease can bind within or proximal to a target region in the target nucleic acid. In certain embodiments, the first RNA-guided nuclease and the second RNA-guided nuclease can complex with the gRNA and the dgRNA, respectively, to form first and second ribonucleoprotein complexes. In certain embodiments, the dgRNA and the second RNA-guided nuclease can bind to a second PAM sequence in the target nucleic acid, where the second PAM sequence can be outward-facing.

[0047] In another aspect, the present disclosure relates to a genome editing system comprising a dgRNA, a first gRNA comprising a second targeting domain sequence having a length greater than 17 nucleotides, a first RNA-guided nuclease, and a second RNA-guided nuclease. In certain embodiments, the first RNA-guided nuclease and the dgRNA can be configured to bind within a first target region in a target nucleic acid. In certain embodiments, the second RNA-guided nuclease and the first gRNA are configured to bind to a second target region to generate a double-strand break (DSB) in the target nucleic acid, thereby generating an indel between the first target region and the second target region. In certain embodiments, the second RNA-guided nuclease can be a Cpf1 molecule. In certain embodiments, the dgRNA can comprise a first targeting domain sequence having a length of 15 nucleotides or less. In certain embodiments, the dgRNA has reduced or no RNA-guided nuclease cleavage activity. In certain embodiments, the dgRNA can be configured not to provide an RNA-guided nuclease cleavage event. In certain embodiments, the dgRNA and the first RNA-guided nuclease can bind to a first protospacer adjacent motif (PAM) sequence in the target nucleic acid. In certain embodiments, the first PAM sequence can be outward-facing. In certain embodiments, the first gRNA and the second RNA-guided nuclease can bind to a second PAM sequence in the target nucleic acid. In certain embodiments, the second PAM sequence can be outward-facing.

[0048] In another aspect, the present disclosure relates to a method of modifying a cell, the method comprising contacting the cell with a dgRNA, a first gRNA comprising a second targeting domain sequence that is more than 17 nucleotides in length, a first RNA-guided nuclease, and a second RNA-guided nuclease. In certain embodiments, the first RNA-guided nuclease and the dgRNA can be configured to bind within a first target region in a target nucleic acid. In certain embodiments, the second RNA-guided nuclease and the first gRNA can bind to a second target region to create a double-strand break (DSB) in the target nucleic acid, thereby generating an indel between the first target region and the second target region. In certain embodiments, the second RNA-guided nuclease can be a Cpf1 molecule. In certain embodiments, the dgRNA can comprise a first targeting domain sequence that is 15 nucleotides or less in length. In certain embodiments, the dgRNA has reduced or no RNA-guided nuclease cleavage activity. In certain embodiments, the dgRNA can be configured to provide an RNA-guided nuclease cleavage event. In certain embodiments, the dgRNA and the first RNA-guided nuclease can bind to a first protospacer adjacent motif (PAM) sequence in the target nucleic acid. In certain embodiments, the first PAM sequence can be outward-facing. In certain embodiments, the first gRNA and the second RNA-guided nuclease can bind to a second PAM sequence in the target nucleic acid. In certain embodiments, the second PAM sequence can be outward-facing.

[0049] The present disclosure also relates to a method of modifying a cell, the method comprising contacting the cell with any of the genome editing systems disclosed herein. In certain embodiments, the step of contacting the cell can comprise contacting the cell with a solution comprising a first and a second ribonucleoprotein complex. In certain embodiments, the step of contacting the cell with the solution further comprises electroporating the cell, thereby introducing the first and second ribonucleoprotein complexes into the cell.

[0050] In another aspect, the disclosure relates to cells modified using the methods disclosed herein. Also disclosed herein are cells comprising productive indels that result in HbF expression. In certain embodiments, the indels can be generated by contacting the cells with a dgRNA, a first gRNA comprising a second targeting domain sequence that is greater than 17 nucleotides in length, a first RNA-guided nuclease, and a second RNA-guided nuclease. In certain embodiments, the first RNA-guided nuclease and the dgRNA can be configured to bind within a first target region in the target nucleic acid. In certain embodiments, the second RNA-guided nuclease and the first gRNA can bind to a second target region to create a double-strand break (DSB) in the target nucleic acid, thereby generating an indel between the first target region and the second target region. In certain embodiments, the cells disclosed herein can be capable of differentiating into erythroblasts, erythrocytes, or precursors of erythrocytes or erythroblasts. In certain embodiments, the cells can be CD34+ cells.

[0051] A genome editing system or method comprising any of the above - described features may comprise a target nucleic acid comprising the human HBG1, HBG2 gene or a combination thereof. In certain embodiments, the target region may be the CCAAT box target region of the human HBG1, HBG2 gene or a combination thereof. In certain embodiments, the first targeting domain sequence may be complementary to a first sequence flanking the CCAAT box target region of the human HBG1, HBG2 gene or a combination thereof, wherein the first sequence optionally overlaps with the CCAAT box target region of the human HBG1, HBG2 gene or a combination thereof. In certain embodiments, the second targeting domain sequence may be complementary to a second sequence flanking the CCAAT box target region of the human HBG1, HBG2 gene or a combination thereof, wherein the second sequence optionally overlaps with the CCAAT box target region of the human HBG1, HBG2 gene or a combination thereof. In certain embodiments, the first targeting domain may comprise truncation of the gRNA targeting domain. In certain embodiments, the gRNA targeting domain may comprise a gRNA described in Table 2, Table 10 or Table 15, and the gRNA targeting domain is truncated from the 5'-end of the gRNA targeting domain. In certain embodiments, the first targeting domain is identical to or differs by 3 nucleotides or less from the dgRNA targeting domain described in Table 10 or Table 15. In certain embodiments, the second targeting domain differs by 3 nucleotides or less from the gRNA targeting domain described in Table 2, Table 10 or Table 15. In certain embodiments, the indel may modify the CCAAT box target region indel. In certain embodiments, the indel may be a productive indel that results in an increase in the level of fetal hemoglobin expression. In certain embodiments, the gRNA, dgRNA or both may be synthesized in vitro or chemically synthesized.

[0052] In certain embodiments, the cell may comprise at least one modified allele of the HBG locus generated by any of the methods for modifying the cells disclosed herein, wherein the modified allele of the HBG locus comprises a modification of the human HBG1 gene, HBG2 gene or a combination thereof.

[0053] In certain embodiments, a population of isolated cells can be modified by any of the methods for modifying the cells disclosed herein, where the population of cells can include an indel distribution that differs from a population of isolated cells or their progeny of the same cell type that have not been modified by the method.

[0054] In certain embodiments, a plurality of cells can be generated by any of the methods for modifying the cells disclosed herein, where at least 20%, 30%, 40%, 50%, 60%, 70%, 80% or 90% of the cells can include a modification of a sequence within the CCAAT box target region of the human HBG1 gene, HBG2 gene or a combination thereof.

[0055] In certain embodiments, the cells disclosed herein can be used for a medicament. In certain embodiments, the cells can be used in the treatment of β-hemoglobinopathies. In certain embodiments, the β-hemoglobinopathy can be selected from the group consisting of sickle cell disease and β-thalassemia.

[0056] In one aspect, the present disclosure relates to a composition comprising a plurality of cells generated by the methods disclosed above, wherein at least 20%, 30%, 40%, 50%, 60%, 70%, 80% or 90% of the cells comprise a modification in the sequence of the CCAAT box target region of the human HBG1 or HBG2 gene, or at least 20%, 30%, 40%, 50%, 60%, 70%, 80% or 90% of the cells of the plurality of cells generated by the methods disclosed above comprise a modification in the sequence of the CCAAT box target region of human HBG1 or HBG2. In certain embodiments, at least a portion of the plurality of cells can be within the erythroid lineage. In certain embodiments, the plurality of cells can be characterized by an increased level of fetal hemoglobin expression compared to an unmodified plurality of cells. In certain embodiments, the level of fetal hemoglobin can be increased by at least 20%, 30%, 40%, 50%, 60%, 70%, 80% or 90%. In certain embodiments, the composition can further comprise a pharmaceutically acceptable carrier.

[0057] In one aspect, the present disclosure relates to a population of cells modified by a genome editing system comprising the dgRNA described above, wherein the population of cells comprises a higher percentage of productive indels compared to a population of cells not modified by the genome editing system. The present disclosure also relates to a population of cells modified by a genome editing system comprising the dgRNA described above, wherein a higher percentage of the population of cells can differentiate into a population of cells of the erythroid lineage expressing HbF compared to a population of cells not modified by the genome editing system. In certain embodiments, the higher percentage can be at least about 15%, at least about 20%, at least about 25%, at least about 30% or at least about 40% higher. In certain embodiments, the cells can be hematopoietic stem cells. In certain embodiments, the cells can be capable of differentiating into erythroblasts, erythrocytes or precursors of erythrocytes or erythroblasts. In certain embodiments, the indels can be generated by a repair mechanism other than microhomology-mediated end joining (MMEJ) repair.

[0058] The present disclosure also relates to the use of any of the cells disclosed herein in the manufacture of a medicament for treating β-thalassemia in a subject.

[0059] In one aspect, the present disclosure relates to a method of treating β-thalassemia in a subject in need thereof, the method comprising administering to the subject a cell disclosed herein. In certain embodiments, the method of treating β-thalassemia in a subject in need thereof may comprise administering to the subject a population of modified hematopoietic cells, wherein one or more cells are modified according to a method of modifying a cell disclosed herein.

[0060] In one aspect, the present disclosure relates to a genome editing system comprising an RNA-guided nuclease and a first guide RNA, wherein the first guide RNA may comprise a first targeting domain complementary to a first sequence flanking the CCAAT box target region of the human HBG1, HBG2 gene or a combination thereof, and the first sequence optionally overlaps with the CCAAT box target region of the human HBG1, HBG2 gene or a combination thereof. In certain embodiments, the genome editing system may further comprise a template nucleic acid encoding a modification of the CCAAT box target region of the human HBG1, HBG2 gene or a combination thereof. In certain embodiments, the template nucleic acid may be a single-stranded oligodeoxynucleotide (ssODN) or a double-stranded oligodeoxynucleotide (dsODN). In certain embodiments, the ssODN may comprise a 5' homology arm, a substitution sequence, and a 3' homology arm. In certain embodiments, the homology arms may be symmetric in length. In certain embodiments, the homology arms may be asymmetric in length. In certain embodiments, the ssODN may comprise one or more phosphorothioate modifications. In certain embodiments, the one or more phosphorothioate modifications may be at the 5' end, 3' end, or a combination thereof. In certain embodiments, the ssODN may be a plus strand or a minus strand. In certain embodiments, the modification may be a non-naturally occurring modification. In certain embodiments, the modification may comprise a deletion of the CCAAT box target region. In certain embodiments, the deletion may comprise an 18nt deletion, an 11nt deletion, a 4nt deletion, a 1nt deletion, or a combination thereof. In certain embodiments, the CCAAT box target region may comprise an 18nt target region, an 11nt target region, a 4nt target region, a 1nt target region, or a combination thereof. In certain embodiments, the 5' homology arm may be, for example, about 25 to about 200 or more nucleotides in length, such as at least about 25, 50, 75, 100, 125, 150, 175, or 200 nucleotides in length; the substitution sequence may have a length of 0 nucleotides; and the 3' homology arm may be, for example, about 25 to about 200 or more nucleotides in length, such as at least about 25, 50, 75, 100, 125, 150, 175, or 200 nucleotides in length.In certain embodiments, the 5' homology arm can include about 50-100 bp of homology, such as 55-95, 60-90, 70-90, or 80-90 bp, etc., on the 5' side of the 18 nt target region, and the 3' homology arm can include about 50-100 bp of homology, such as 55-95, 60-90, 70-90, or 80-90 bp, etc., on the 3' side of the 18 nt target region. In certain embodiments, the ssODN can comprise, consist essentially of, or consist of SEQ ID NO: 974 or SEQ ID NO: 975. In certain embodiments, the 5' homology arm can include about 50-100 bp of homology, such as 55-95, 60-90, 70-90, or 80-90 bp, etc., on the 5' side of the 11 nt target region, and the 3' homology arm can include about 50-100 bp of homology, such as 55-95, 60-90, 70-90, or 80-90 bp, etc., on the 3' side of the 11 nt target region. In certain embodiments, the ssODN can comprise, consist essentially of, or consist of SEQ ID NO: 976 or SEQ ID NO: 978. In certain embodiments, the 5' homology arm can include about 50-100 bp of homology, such as 55-95, 60-90, 70-90, or 80-90 bp, etc., on the 5' side of the 4 nt target region, and the 3' homology arm can include about 50-100 bp of homology, such as 55-95, 60-90, 70-90, or 80-90 bp, etc., on the 3' side of the 4 nt target region. In certain embodiments, the ssODN can comprise, consist essentially of, or consist of a sequence selected from the group consisting of SEQ ID NO: 984, SEQ ID NO: 985, SEQ ID NO: 986, SEQ ID NO: 987, SEQ ID NO: 988, SEQ ID NO: 989, SEQ ID NO: 990, SEQ ID NO: 991, SEQ ID NO: 992, SEQ ID NO: 993, SEQ ID NO: 994, and SEQ ID NO: 995. In certain embodiments, the 5' homology arm can include about 50-100 bp of homology, such as 55-95, 60-90, 70-90, or 80-90 bp, etc., on the 5' side of the 1 nt target region, and the 3' homology arm can include about 50-100 bp of homology, such as 55-95, 60-90, 70-90, or 80-90 bp, etc., on the 3' side of the 1 nt target region. In certain embodiments, the homology arms can be symmetric in length.In certain embodiments, the ssODN can comprise, consist essentially of, or consist of SEQ ID NO: 982 or SEQ ID NO: 983. In certain embodiments, the modification can be a naturally occurring modification. In certain embodiments, the modification can include a deletion or mutation of the CCAAT box target region. In certain embodiments, the CCAAT box target region can include a 13nt target region, a -117G>A target region, or a combination thereof. In certain embodiments, the modification can include a 13nt deletion in the 13nt target region, a G to A substitution in the -117G>A target region, or a combination thereof. In certain embodiments, the 5' homology arm can include about 50-100 bp of homology, such as 55-95, 60-90, 70-90, or 80-90 bp, on the 5' side of the 13nt target region, and the 3' homology arm can include about 50-100 bp of homology, such as 55-95, 60-90, 70-90, or 80-90 bp, on the 3' side of the 13nt target region. In certain embodiments, the ssODN can comprise, consist essentially of, or consist of SEQ ID NO: 977 or SEQ ID NO: 979. In certain embodiments, the 5' homology arm can include about 50-100 bp of homology, such as 55-95, 60-90, 70-90, or 80-90 bp, on the 5' side of the 13nt target region, and the 3' homology arm can include about 50-100 bp of homology, such as 55-95, 60-90, 70-90, or 80-90 bp, on the 3' side of the 13nt target region. In certain embodiments, the ssODN can comprise, consist essentially of, or consist of SEQ ID NO: 980 or SEQ ID NO: 981. In certain embodiments, the RNA-guided nuclease can be S. pyogenes Cas9. In certain embodiments, the RNA-guided nuclease can be a Cpf1 variant as disclosed herein. In certain embodiments, the first targeting domain can differ by 3 nucleotides or less from the targeting domains listed in Table 7, Table 18, Table 19, or the gRNAs of Table 12, Table 19.In certain embodiments, the genome editing system may further comprise a second guide RNA, wherein the second guide RNA may comprise a second targeting domain that is complementary to a second sequence on the side of the CCAAT box target region of the human HBG1, HBG2 gene or a combination thereof, and the second sequence optionally overlaps with the CCAAT box target region of the human HBG1, HBG2 gene or a combination thereof. In certain embodiments, the RNA-guided nuclease may be a nickase and optionally lacks RuvC activity. In certain embodiments, the genome editing system may comprise a first and a second RNA-guided nuclease. In certain embodiments, the first and second RNA-guided nucleases may complex with the first and second guide RNAs, respectively, to form first and second ribonucleoprotein complexes. In certain embodiments, the genome editing system may further comprise a third guide RNA and optionally a fourth guide RNA, wherein the third and fourth guide RNAs may comprise third and fourth targeting domains that are complementary to third and fourth sequences on the opposite side of the position of the GATA1 binding motif in the BCL11A erythroid enhancer (BCL11Ae) of the human BCL11A gene, and one or both of the third and fourth sequences optionally overlap with the GATA1 binding motif in the BCL11Ae of the human BCL11A gene. In certain embodiments, the genome editing system may further comprise a nucleic acid template encoding a deletion of the GATA1 binding motif in BCL11Ae. In certain embodiments, the RNA-guided nuclease may be S. pyogenes Cas9. In certain embodiments, the RNA-guided nuclease may be a nickase and optionally lacks RuvC activity. In certain embodiments, the third targeting domain may be complementary to a sequence within 1000 nucleotides upstream of the GATA1 binding motif in BCL11Ae. In certain embodiments, the third targeting domain may be complementary to a sequence within 100 nucleotides upstream of the GATA1 binding motif in BCL11Ae. In certain embodiments, one of the third and fourth targeting domains may be complementary to a sequence within 100 nucleotides downstream of the GATA1 binding motif in BCL11Ae.In certain embodiments, the fourth targeting domain can be complementary to a sequence within 50 nucleotides downstream of the GATA1 binding motif in BCL11Ae. In certain embodiments, at least one of the third and fourth targeting domains can differ by 3 nucleotides or less from the targeting domains listed in Table 9. In certain embodiments, the genome editing system can include the first and second RNA-guided nucleases. In certain embodiments, the first and second RNA-guided nucleases can complex with the third and fourth guide RNAs, respectively, to form the third and fourth ribonucleoprotein complexes.

[0061] In one aspect, the present disclosure relates to a method of modifying a cell, the method comprising contacting the cell with a genome editing system. In certain embodiments, the step of contacting the cell with the genome editing system can include contacting the cell with a solution comprising the first and second ribonucleoprotein complexes. In certain embodiments, the step of contacting the cell with the solution can further include electroporating the cell, whereby the first and second ribonucleoprotein complexes are introduced into the cell. In certain embodiments, the method of modifying the cell can further include contacting the cell with the genome editing system, and the step of contacting the cell with the genome editing system can include contacting the cell with a solution comprising the first, second, third, and optionally fourth ribonucleoprotein complexes. In certain embodiments, the step of contacting the cell with the solution can further include electroporating the cell, whereby the first, second, third, and optionally fourth ribonucleoprotein complexes are introduced into the cell. In certain embodiments, the cell can be capable of differentiating into an erythroblast, an erythrocyte, or a precursor of an erythrocyte or erythroblast. In certain embodiments, the cell can be a CD34+ cell.

[0062] In one aspect, the present disclosure is a CRISPR-mediated method of modifying a cell, comprising introducing a first single-stranded DNA break (SSB) or double-stranded DNA break (DSB) into the genome of the cell at positions c.-106 to -120 of the human HBG1 or HBG2 gene; and optionally, introducing a second SSB or DSB into the genome of the cell at positions c.-106 to -120 of the human HBG1 or HBG2 gene, wherein the first and second SSBs or DSBs can be repaired by the cell in a manner that modifies the CCAAT box target region of the human HBG1 or HBG2 gene. In certain embodiments, the first and second SSBs or DSBs can be repaired by the cell in a manner that results in modification of the CCAAT box target region of the human HBG1 or HBG2 gene. In certain embodiments, the CRISPR-mediated method can further comprise a template nucleic acid encoding modification of the CCAAT box target region of the human HBG1, HBG2 gene or a combination thereof. In certain embodiments, the template nucleic acid can be a single-stranded oligodeoxynucleotide (ssODN). In certain embodiments, the ssODN can comprise a 5' homology arm, a substitution sequence, and a 3' homology arm. In certain embodiments, the ssODN can be a plus strand or a minus strand. In certain embodiments, the modification can be a non-naturally occurring modification. In certain embodiments, the first and second SSBs or DSBs can be repaired by the cell in a manner that results in the formation of at least one of an indel, a deletion, or an insertion in the CCAAT box target region of the human HBG1 or HBG2 gene. In certain embodiments, the CCAAT box target region can comprise an 18nt target region, an 11nt target region, a 4nt target region, a 1nt target region, or a combination thereof. In certain embodiments, the 5' homology arm can be, for example, about 25 to about 200 nucleotides or more, such as at least about 25, 50, 75, 100, 125, 150, 175, or 200 nucleotides in length; the substitution sequence can comprise a length of 0 nucleotides; and the 3' homology arm can be, for example, about 25 to about 200 nucleotides or more, such as at least about 25, 50, 75, 100, 125, 150, 175, or 200 nucleotides in length.In certain embodiments, the 5' homology arm may include homology of about 50-100 bp, such as 55-95, 60-90, 70-90, or 80-90 bp, on the 5' side of the 18 nt target region, 11 nt target region, 4 nt target region, or 1 nt target region, and the 3' homology arm may include homology of about 50-100 bp, such as 55-95, 60-90, 70-90, or 80-90 bp, on the 3' side of the 18 nt target region, 11 nt target region, 4 nt target region, or 1 nt target region. In certain embodiments, the ssODN may comprise, consist essentially of, or consist of a sequence selected from the group consisting of SEQ ID NO: 974, SEQ ID NO: 975, SEQ ID NO: 976, SEQ ID NO: 978, SEQ ID NO: 984, SEQ ID NO: 985, SEQ ID NO: 986, SEQ ID NO: 987, SEQ ID NO: 988, SEQ ID NO: 989, SEQ ID NO: 990, SEQ ID NO: 991, SEQ ID NO: 992, SEQ ID NO: 993, SEQ ID NO: 994, SEQ ID NO: 995, SEQ ID NO: 982, and SEQ ID NO: 983. In certain embodiments, the modification may be a non-naturally occurring modification. In certain embodiments, the first and second SSB or DSB may be repaired by the cell in a manner that results in the formation of at least one indel, deletion, or insertion in the CCAAT box target region of the human HBG1 or HBG2 gene. In certain embodiments, the CCAAT box target region may include the 13 nt target region, the -117G>A target region, or a combination thereof. In certain embodiments, the modification may include a 13 nt deletion in the 13 nt target region, a substitution of G to A in the -117G>A target region, or a combination thereof. In certain embodiments, the 5' homology arm may include homology of about 50-100 bp, such as 55-95, 60-90, 70-90, or 80-90 bp, on the 5' side of the 13 nt target region or the -117G>A target region, and the 3' homology arm may include homology of about 50-100 bp, such as 55-95, 60-90, 70-90, or 80-90 bp, on the 3' side of the 13 nt target region or the -117G>A target region. In certain embodiments, the ssODN may comprise, consist essentially of, or consist of a sequence selected from the group consisting of SEQ ID NO: 977 or SEQ ID NO: 979, SEQ ID NO: 980 or SEQ ID NO: 981.

[0063] In one aspect, the present disclosure relates to a composition that may include a plurality of cells generated by a method of modifying the cells disclosed herein, where at least 20%, 30%, 40%, 50%, 60%, 70%, 80% or 90% of the cells may include a modification of the sequence of the CCAAT box target region of the human HBG1 gene, HBG2 gene or a combination thereof. In certain embodiments, the modification may include an 18nt deletion, an 11nt deletion, a 4nt deletion, a 1nt deletion, a 13nt deletion, a G to A substitution at -117 of the human HBG1 gene, HBG2 gene or a combination thereof. In certain embodiments, at least a portion of the plurality of cells may be within the erythroid lineage. In certain embodiments, the plurality of cells may be characterized by an increased level of fetal hemoglobin expression compared to an unmodified plurality of cells. In certain embodiments, the level of fetal hemoglobin may be increased by at least 20%, 30%, 40%, 50%, 60%, 70%, 80% or 90%. In certain embodiments, the composition may further include a pharmaceutically acceptable carrier.

[0064] In one aspect, the present disclosure relates to a cell comprising a synthetic genotype generated by a method of modifying the cells disclosed herein, where the cell may include an 18nt deletion, an 11nt deletion, a 4nt deletion, a 1nt deletion, a 13nt deletion, a G to A substitution at -117 of the human HBG1 gene, HBG2 gene or a combination thereof.

[0065] In one aspect, the present disclosure relates to a cell comprising at least one allele of the HBG locus generated by a method of modifying the cells disclosed herein, where the cell may encode an 18nt deletion, an 11nt deletion, a 4nt deletion, a 1nt deletion, a 13nt deletion, a G to A substitution at -117 of the human HBG1 gene, HBG2 gene or a combination thereof.

[0066] In one aspect, the present disclosure relates to an AAV vector that may include a template nucleic acid encoding a non-naturally occurring modification of the CCAAT box target region of the human HBG1, HBG2 gene or a combination thereof. In certain embodiments, the template nucleic acid may be a single-stranded oligodeoxynucleotide (ssODN). In certain embodiments, the CCAAT box target region may include an 18nt target region, an 11nt target region, a 4nt target region, a 1nt target region, or a combination thereof. In certain embodiments, the ssODN may include a 5' homology arm, a substitution sequence, and a 3' homology arm. In certain embodiments, the 5' homology arm may be, for example, about 25 to about 200 or more nucleotides in length, such as at least about 25, 50, 75, 100, 125, 150, 175, or 200 nucleotides in length; the substitution sequence may include a length of 0 nucleotides; the 3' homology arm may be, for example, about 25 to about 200 or more nucleotides in length, such as at least about 25, 50, 75, 100, 125, 150, 175, or 200 nucleotides in length. In certain embodiments, the 5' homology arm may include homology of about 50 to 100 bp, such as 55 to 95, 60 to 90, 70 to 90, or 80 to 90 bp, on the 5' side of the 18nt target region, 11nt target region, 4nt target region, or 1nt target region, and the 3' homology arm may include homology of about 50 to 100 bp, such as 55 to 95, 60 to 90, 70 to 90, or 80 to 90 bp, on the 3' side of the 18nt target region, 11nt target region, 4nt target region, or 1nt target region. In certain embodiments, the ssODN may include, consist essentially of, or consist of a sequence selected from the group consisting of SEQ ID NOs: 974-976, SEQ ID NO: 978, SEQ ID NOs: 982-995.

[0067] In one aspect, the present disclosure relates to a nucleotide sequence comprising a template nucleic acid encoding a non-naturally occurring modification of the CCAAT box target region of the human HBG1, HBG2 gene or a combination thereof. In certain embodiments, the template nucleic acid can be a single-stranded oligodeoxynucleotide (ssODN) or a double-stranded oligodeoxynucleotide (dsODN) comprising the modification. In certain embodiments, the CCAAT box target region can comprise an 18 nt target region, an 11 nt target region, a 4 nt target region, a 1 nt target region, or a combination thereof. In certain embodiments, the ssODN can comprise a 5' homology arm, a substitution sequence, and a 3' homology arm. In certain embodiments, the 5' homology arm can be, for example, about 25 to about 200 or more nucleotides in length, such as at least about 25, 50, 75, 100, 125, 150, 175, or 200 nucleotides in length; the substitution sequence can comprise a length of 0 nucleotides; the 3' homology arm can be, for example, about 25 to about 200 or more nucleotides in length, such as at least about 25, 50, 75, 100, 125, 150, 175, or 200 nucleotides in length. In certain embodiments, the 5' homology arm can comprise a homology of about 50 to 100 bp, such as 55 to 95, 60 to 90, 70 to 90, or 80 to 90 bp, on the 5' side of the 18 nt target region, 11 nt target region, 4 nt target region, or 1 nt target region; the 3' homology arm can comprise a homology of about 50 to 100 bp, such as 55 to 95, 60 to 90, 70 to 90, or 80 to 90 bp, on the 3' side of the 18 nt target region, 11 nt target region, 4 nt target region, or 1 nt target region. In certain embodiments, the ssODN can comprise, consist essentially of, or consist of a sequence selected from the group consisting of SEQ ID NOs: 974-976, SEQ ID NO: 978, SEQ ID NOs: 982-995.

[0068] In one aspect, the present disclosure relates to a cell comprising a synthetic genotype, wherein the cell can comprise an 18 nt deletion, an 11 nt deletion, a 4 nt deletion, a 1 nt deletion, a 13 nt deletion, a G to A substitution at -117 of the human HBG1 gene, HBG2 gene, or a combination thereof.

[0069] In one aspect, the present disclosure relates to a composition comprising a population of cells generated by a method of modifying the cells disclosed herein, wherein the cells comprise a higher frequency of modification of the sequence of the CCAAT box target region of the human HBG1 gene, HBG2 gene, or a combination thereof as compared to a population of unmodified cells. In certain embodiments, the higher frequency is at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% higher. In certain embodiments, the modification comprises an 18nt deletion, an 11nt deletion, a 4nt deletion, a 1nt deletion, a 13nt deletion, a substitution of G to A at -117 in the human HBG1 gene, HBG2 gene, or a combination thereof. In certain embodiments, at least a portion of the population of cells is within the erythroid lineage.

[0070] This listing is intended to be exemplary and explanatory, and not comprehensive and limiting. Additional aspects and embodiments may be described in or apparent from the remainder of the present disclosure and the claims.

[0071] The accompanying drawings are intended to provide illustrative and schematic examples of particular aspects and embodiments of the present disclosure, and not to be comprehensive. The drawings are not intended to limit or be bound by any particular theory or model, and are not necessarily to scale. Without limiting the foregoing, nucleic acids and polypeptides may be depicted as linear sequences or as schematic two-dimensional or three-dimensional structures; these depictions are intended to be exemplary and not to limit or be bound by any particular model or theory regarding their structure.

Brief Description of the Drawings

[0072]

Figure 1

Figure 2A

Figure 2B

Figure 3A

Figure 3B

Figure 3C

Figure 4A

Figure 4B

Figure 4C

Figure 5A

Figure 5B

Figure 6A

Figure 6B

Figure 7A

Figure 7B

Figure 8A

Figure 8B

Figure 8C

Figure 8D

Figure 9

Figure 10A

Figure 10B

Figure 10C

Figure 10D

Figure 10E

Figure 10F

Figure 10G

Figure 10H

Figure 11A

Figure 11B

Figure 11C

Figure 11D

Figure 12

Figure 13A

Figure 13B

Figure 14A

Figure 14B

Figure 14C

Figure 14D

Figure 14E

Figure 15A

Figure 15B

Figure 15C

Figure 16

Figure 17A

Figure 17B

Figure 17C

Figure 17D

Figure 18

Figure 19A

Figure 19B

Figure 19C

Figure 19D

Figure 19E

Figure 19F

Figure 19G

Figure 20A

Figure 20B

Figure 21A

Figure 21B

Figure 21C

Figure 21D

Figure 22A

Figure 22B

Figure 22C

Figure 22D

Figure 23A

Figure 23B

Figure 24

Figure 25

Figure 26A

Figure 26B

Figure 26C

Figure 27

Figure 28A

Figure 28B

Figure 29

Figure 30

Figure 31A

Figure 31B

Figure 32A

Figure 32B

Figure 33

Figure 34

Figure 35A

Figure 35B

Figure 35C

Figure 36A

Figure 36B

Figure 37

Figure 38A

Figure 38B

Figure 39

Figure 40

Figure 41

Figure 42

Figure 43

Figure 44A

Figure 44B

Figure 45A

Figure 45B

Figure 46A

Figure 46B

Figure 46C

Figure 47

Figure 48

Figure 49

Figure 50

Figure 51

Figure 52

Figure 53

Figure 54

Figure 55A

Figure 55B

Figure 56A

Figure 56B

Figure 57

Figure 58

Figure 59

Figure 60

Figure 61

Figure 62-1

Figure 62-2

Figure 62-3

Figure 62-4

Figure 62-5

Figure 62-6

Figure 62-7

Figure 62-8

Figure 62-9

Figure 62-10

Figure 62-11

Figure 62-12

Figure 62-13

Mode for Carrying Out the Invention

[0073] Definitions and Abbreviations Unless otherwise specified, each of the following terms has the meaning associated with it in this section.

[0074] The indefinite articles "a" and "an" refer to at least one of the associated nouns and are used interchangeably with the terms "at least one" and "one or more". For example, "module" means at least one module or one or more modules.

[0075] The conjunctions "or" and "and / or" are used synonymously as non-exclusive disjunctive alternatives.

[0076] "Domain" is used to represent a segment of a protein or nucleic acid. Unless otherwise specified, a domain need not have any specific functional characteristics.

[0077] The term "exogenous trans-acting factor" interacts with an RNA-guided nuclease or gRNA by means of a modification such as (a) insertion or fusion of a peptide or nucleotide into the RNA-guided nuclease or gRNA, and (b) interacts with the target DNA to modify its helical structure, and refers to any peptide or nucleotide component of a genome editing system. Peptide or nucleotide insertions or fusions may include, without limitation, direct covalent bonds between the RNA-guided nuclease or gRNA and the exogenous trans-acting factor and / or non-covalent bonds mediated by insertion or fusion of RNA / protein interaction domains such as MS2 loops and protein / protein interaction domains such as DZ, Lim or SH1, 2 or 3 domains. Other specific RNA and amino acid interaction motifs will be well known to those skilled in the art. Trans-acting factors may generally include transcriptional activators.

[0078] The term "booster element" refers to an element that increases the editing of a target nucleic acid when co-delivered with a ribonucleoprotein (RNP) complex ("gRNA-nuclease-RNP") containing a gRNA complexed to an RNA-guided nuclease, as compared to editing of the target nucleic acid without a booster element. In certain embodiments, the co-delivery can be sequential or simultaneous. In certain embodiments, the booster element can be an RNP complex comprising a dead guide RNA complexed to a WT Cas9 protein, a Cas9 nickase protein (e.g., a Cas9 D10A protein), or an enzymatically inactive Cas9 (eiCas9) protein. In certain embodiments, the booster element can be an RNP complex comprising a guide RNA complexed to a Cas9 nickase protein (e.g., a Cas9 D10A protein) or an enzymatically inactive Cas9 (eiCas9) protein. In certain embodiments, the booster element can be single-stranded or double-stranded donor template DNA. In certain embodiments, one or more booster elements can be co-delivered with the gRNA-nuclease-RNP to increase editing of the target nucleic acid. In certain embodiments, the booster element can be co-delivered with an RNP comprising a gRNA complexed to a Cpf1 molecule ("gRNA-Cpf1-RNP") to increase editing of the target nucleic acid.

[0079] "Productive indel" refers to an indel (deletion and / or insertion) that results in HbF expression. In certain embodiments, a productive indel can induce HbF expression. In certain embodiments, a productive indel can result in an increase in the level of HbF expression.

[0080] "Indel" is an insertion and / or deletion in a nucleic acid sequence. An indel can be the product of the repair of a DNA double-strand break, such as a double-strand break formed by the genome editing system of the present disclosure. Indels are most commonly formed when the interruption is repaired by a "error-prone" repair pathway, such as the NHEJ pathway described below.

[0081] "Gene conversion" refers to the modification of a DNA sequence by the incorporation of an endogenous homologous sequence (e.g., a homologous sequence within a gene array). "Gene correction" refers to the modification of a DNA sequence by the incorporation of an exogenous homologous sequence such as exogenous single-stranded or double-stranded donor template DNA. Gene conversion and gene correction are products of the repair of DNA double-strand breaks by HDR pathways such as those described below.

[0082] Indels, gene conversions, gene corrections, and other genome editing results are typically evaluated by sequencing (most commonly by "next-generation" or "synthetic sequencing" methods, although Sanger sequencing can still be used), and quantified by the relative frequency of numerical changes in all sequence reads (e.g., ±1, ±2, or more bases). DNA samples for sequencing can be prepared by a variety of methods known in the art, including amplification of the target site by polymerase chain reaction (PCR), capture of DNA ends generated by double-strand breaks as in the GUIDEseq process described in Tsai 2016 (incorporated herein by reference), or by other means well known in the art. Genome editing results can also be evaluated by in situ hybridization methods such as the FiberComb® system commercialized by Genomic Vision (Bagneux, France) and any other suitable methods known in the art.

[0083] "Alt-HDR", "alternative homology-directed repair", or "alternative HDR" are used interchangeably to refer to the process of using homologous nucleic acids (e.g., endogenous homologous sequences such as sister chromatids or exogenous nucleic acids such as template nucleic acids) to repair DNA damage. Alt-HDR differs from standard HDR in that the process utilizes a different pathway and can be inhibited by the standard HDR mediators RAD51 and BRCA2. Alt-HDR is also distinguished by the involvement of single-stranded or nicked homologous nucleic acid templates, whereas standard HDR generally involves double-stranded homologous templates.

[0084] "Standard HDR", "standard homology-directed repair", or "cHDR" refers to the process of repairing DNA damage using homologous nucleic acids (e.g., endogenous homologous sequences such as sister chromatids or exogenous nucleic acids such as template nucleic acids). Standard HDR typically functions when there is significant resection at a double-strand break and at least one single-stranded portion of DNA is formed. In normal cells, cHDR typically involves a series of steps such as break recognition, break stabilization, resection, single-stranded DNA stabilization, formation of DNA crossover intermediates, resolution of crossover intermediates, and ligation. The process requires RAD51 and BRCA2, and the homologous nucleic acids are typically double-stranded.

[0085] Unless otherwise specified, the term "HDR" as used herein encompasses both standard HDR and alt-HDR.

[0086] "Non-homologous end joining" or "NHEJ" refers to ligation-mediated repair and / or non-template-mediated repair such as standard NHEJ (cNHEJ) and alternative NHEJ (altNHEJ), which in turn includes microhomology-mediated end joining (MMEJ), single-strand annealing (SSA), and synthesis-dependent microhomology-mediated end joining (SD-MMEJ).

[0087] "Substituted" or "substitution", when used with respect to the modification of a molecule (e.g., a nucleic acid or a protein), does not require a process limitation and simply indicates the presence of a substituted entity.

[0088] "Subject" means a human, a mouse, or a non-human primate. A human subject can be of any age (e.g., an infant, a child, a young adult, or an adult) and can be suffering from a disease or in need of genetic modification.

[0089] "To treat", "treating", and "treatment" mean treating a disease (e.g., a human subject) including, but not limited to, suppressing the disease, i.e., preventing or precluding its onset or progression; alleviating the disease, i.e., causing regression of the disease state; alleviating one or more symptoms of the disease; and curing the disease.

[0090] "To prevent", "preventing", and "prevention" refer to preventing a disease in a subject, including, but not limited to: (a) avoiding or eliminating the disease; (b) affecting a predisposition to the disease; or (c) preventing or delaying the onset of at least one symptom of the disease.

[0091] "Kit" refers to any collection of two or more components that together constitute a functional unit that can be used for a particular purpose. By way of illustration, and not limitation, one kit according to the present disclosure can include a guide RNA complexed with or capable of complexing with an RNA-guided nuclease, and a pharmaceutically acceptable carrier (e.g., suspended or suspendable therein). In certain embodiments, the kit can include a booster element. For example, the kit can be used to introduce a complex into such a cell or subject for the purpose of causing a desired genomic modification in the cell or subject. The components of the kit can be packaged together or they can be packaged separately. A kit according to the present disclosure optionally includes, for example, a Directions for Use (DFU) that describes the use of the kit according to the methods of the present disclosure. The DFU can be physically packaged with the kit or it can be provided to the kit user, for example, by electronic means.

[0092] The terms "polynucleotide", "nucleotide sequence", "nucleic acid", "nucleic acid molecule", "nucleic acid sequence" and "oligonucleotide" refer to a series of nucleotide bases (also referred to as "nucleotides") in DNA and RNA, and mean any strand of two or more nucleotides. Polynucleotides, nucleotide sequences, nucleic acids, etc. can be single-stranded or double-stranded chimeric mixtures or their derivatives or modified versions. They can be modified, for example, in the base moiety, sugar moiety or phosphate backbone to improve the stability of the molecule, its hybridization parameters, etc. Nucleotide sequences typically carry genetic information including, but not limited to, the information used by cellular machinery to produce proteins and enzymes. These terms include both double-stranded or single-stranded genomic DNA, RNA, any synthetic and genetically engineered polynucleotides, and both sense and antisense polynucleotides. These terms also include nucleic acids containing modified bases.

[0093] As shown in Table 1 below, conventional IUPAC notation is used in the nucleotide sequences presented herein (see also Cornish-Bowden A, Nucleic Acids Res. 1985 May 10; 13(9):3021-30, which is incorporated herein by reference). However, it should be noted that when a sequence can be encoded by either DNA or RNA, for example in a gRNA targeting domain, "T" indicates "thymine or uracil".

[0094]

Table 1

[0095] The terms "protein", "peptide" and "polypeptide" are used synonymously and refer to a continuous chain of amino acids linked together via peptide bonds. The term includes individual proteins, groups or complexes of proteins that bind together and fragments or portions, variants, derivatives and analogs of such proteins. Peptide sequences are presented herein using the conventional notation starting from the amino or N-terminal on the left and proceeding to the carboxyl or C-terminal on the right. Standard one-letter or three-letter abbreviations may be used.

[0096] The notations such as "CCAAT box target region" refer to the sequences located on the 5'-side of the transcription start site (TSS) of the HBG1 and / or HBG2 genes. The CCAAT box is a highly conserved motif within the promoter regions of the α-like and β-like globin genes. The region within or near the CCAAT box plays an important role in the regulation of globin genes. For example, the γ-globin distal CCAAT box is associated with the hereditary persistence of fetal hemoglobin. Several transcription factors have been reported to bind to the overlapping CCAAT box regions of the γ-globin promoter, such as NF-Y, COUP-TFII (NF-E3), CDP, GATA1 / NF-E1, and DRED (Martyn 2017). Without wishing to be bound by theory, the binding site of the transcriptional activator NF-Y is thought to overlap with the transcriptional repressor of the γ-globin promoter. For example, HPFH mutations present within the distal γ-globin promoter region, such as within or near the CCAAT box, can modify the competitive binding of these factors and thus contribute to increased γ-globin expression and elevated HbF levels. The genomic positions provided herein for HBG1 and HBG2 are based on the coordinates provided in NCBI reference sequence NC_000011 “Homo sapiens chromosome 11, GRCh38.p12 Primary Assembly,” (Version NC_000011.10). The distal CCAAT boxes of HBG1 and HBG2 are located at HBG1 and HBG2 c.-111~-115 (the genomic positions are Hg38 Chr11:5,249,968~Chr11:5,249,972 and Hg38 Chr11:5,254,892~Chr11:5,254,896, respectively). The HBG1 c.-111~-115 region is exemplified at positions 2823~2827 by SEQ ID NO: 902 (HBG1), and the HBG2 c.-111~-115 region is exemplified at positions 2747~2751 by SEQ ID NO: 903 (HBG2).In certain embodiments, the "CCAAT box target region" refers to the region of or near the distal CCAAT box, and includes the nucleotides of the distal CCAAT box, and 25 nucleotides upstream (5') and 25 nucleotides downstream (3') of the distal CCAAT box (i.e., HBG1 / 2 c.-86~-140) (the genomic positions are Hg38 Chr11:5249943~Hg38 Chr11:5249997 and Hg38 Chr11:5254867~Hg38 Chr11:5254921, respectively). The HBG1 c.-86~-140 region is exemplified at positions 2798~2852 in SEQ ID NO: 902 (HBG1), and the HBG2 c.-86~-140 region is exemplified at positions 2723~2776 in SEQ ID NO: 903 (HBG2). In other embodiments, the "CCAAT box target region" refers to the region of or near the distal CCAAT box, and includes the nucleotides of the distal CCAAT box, and 5 nucleotides upstream (5') and 5 nucleotides downstream (3') of the distal CCAAT box (i.e., HBG1 / 2 c.-106~-120 (the genomic positions are Hg38 Chr11:5249963~Hg38 Chr11:5249977 (HGB1) and Hg38 Chr11:5254887~Hg38 Chr11:5254901, respectively)). The HBG1 c.-106~-120 region is exemplified at positions 2818~2832 in SEQ ID NO: 902 (HBG1), and the HBG2 c.-106~-120 region is exemplified at positions 2742~2756 in SEQ ID NO: 903 (HBG2). The term "CCAAT box target site modification" refers to a modification (e.g., deletion, insertion, mutation) of one or more nucleotides in the CCAAT box target region. Exemplary examples of CCAAT box target region modifications include, without limitation, 1nt deletion, 4nt deletion, 11nt deletion, 13nt deletion, 18nt deletion, and -117G>A modification. In the usage of this specification, the terms "CCAAT box" and "CAAT box" may be used synonymously.

[0097] Expressions such as "c.-114~-102 region", "c.-102~-114 region", "-102:-114", and "13nt target region" respectively refer to the sequences on the 5' side of the transcription start site (TSS) of the HBG1 and / or HBG2 genes at genomic positions Hg38 Chr11:5,249,959~Hg38 Chr11:5,249,971 and Hg38 Chr11:5,254,883~Hg38 Chr11:5,254,895. The HBG1 c.-102~-114 region is exemplified by positions 2824~2836 in SEQ ID NO: 902 (HBG1), and the HBG2 c.-102~-114 region is exemplified by positions 2748~2760 in SEQ ID NO: 903 (HBG2). Terms such as "13nt deletion" refer to the deletion of the 13nt target region.

[0098] Expressions such as "c.-121~-104 region", "c.-104~-121 region", "-104:-121", and "18nt target region" respectively refer to the sequences on the 5' side of the transcription start site (TSS) of the HBG1 and / or HBG2 genes at genomic positions Hg38 Chr11:5,249,961~Hg38 Chr11:5,249,978 and Hg38 Chr11:5,254,885~Hg38 Chr11:5,254,902. The HBG1 c.-104~-121 region is exemplified by positions 2817~2834 in SEQ ID NO: 902 (HBG1), and the HBG2 c.-104~-121 region is exemplified by positions 2741~2758 in SEQ ID NO: 903 (HBG2). Terms such as "18nt deletion" refer to the deletion of the 18nt target region.

[0099] Expressions such as "c.-105~-115 region", "c.-115~-105 region", "-105:-115", and "11nt target region" respectively refer to the sequences on the 5' side of the transcription start site (TSS) of the HBG1 and / or HBG2 genes at genomic positions Hg38 Chr11:5,249,962~Hg38 Chr11:5,249,972 and Hg38 Chr11:5,254,886~Hg38 Chr11:5,254,896. The HBG1 c.-105~-115 region is exemplified by positions 2823~2833 in SEQ ID NO: 902 (HBG1), and the HBG2 c.-105~-115 region is exemplified by positions 2747~2757 in SEQ ID NO: 903 (HBG2). Terms such as "11nt deletion" refer to the deletion of the 11nt target region.

[0100] Expressions such as "c.-115~-112 region", "c.-112~-115 region", "-112:-115", and "4nt target region" respectively refer to the sequences on the 5' side of the transcription start site (TSS) of the HBG1 and / or HBG2 genes at genomic positions Hg38 Chr11:5,249,969~Hg38 Chr11:5,249,972 and Hg38 Chr11:5,254,893~Hg38 Chr11:5,254,896. The HBG1 c.-112~-115 region is exemplified by positions 2823~2826 in SEQ ID NO: 902, and the HBG2 c.-112~-115 region is exemplified by positions 2747~2750 in SEQ ID NO: 903 (HBG2). Terms such as "4nt deletion" refer to the deletion of the 4nt target region.

[0101] Expressions such as "c.-116 region", "HBG-116", and "1nt target region" respectively refer to the sequences on the 5' side of the transcription start site (TSS) of the HBG1 and / or HBG2 genes at genomic positions Hg38 Chr11:5,249,973 and Hg38 Chr11:5,254,897. The HBG1 c.-116 region is exemplified by position 2822 in SEQ ID NO: 902, and the HBG2 c.-116 region is exemplified by position 2746 in SEQ ID NO: 903 (HBG2). Terms such as "1nt deletion" refer to the deletion of the 1nt target region.

[0102] The notations such as "c.-117G>A region", "HBG-117G>A", and "-117G>A target region" respectively refer to the sequences on the 5' side of the transcription start site (TSS) of the HBG1 and / or HBG2 genes at genomic positions Hg38 Chr11:5,249,974 to Hg38 Chr11:5,249,974 and Hg38 Chr11:5,254,898 to Hg38 Chr11:5,254,898. The HBG1 c.-117G>A region is exemplified by the substitution of guanine (G) to adenine (A) at position 2821 in SEQ ID NO: 902, and the HBG2 c.-117G>A region is exemplified by the substitution of G to A at position 2745 in SEQ ID NO: 903 (HBG2). Terms such as "-117G>A modification" refer to the substitution of G to A in the -117G>A target region.

[0103] The term "proximal HBG1 / 2 promoter target sequence" refers to the region within 50, 100, 200, 300, 400, or 500 bp proximal to the HBG1 / 2 promoter sequence that includes the 13nt target region. Modification by the genome editing system according to the present disclosure facilitates (e.g., tends to cause, promote, or increase the likelihood of) upregulation of HbF production in erythroid progeny.

[0104] The term "GATA1 binding motif in BCL11Ae" refers to the sequence of the GATA1 binding motif in the erythroid-specific enhancer of BCL11A (BCL11Ae) that is in the +58 DNaseI hypersensitive site (DHS) region of intron 2 of the BCL11A gene. The genomic coordinates of the GATA1 binding motif in BCL11Ae are chr2:60,495,265 to 60,495,270. The +58 DHS site contains the 115 base pair (bp) sequence described in SEQ ID NO: 968. The +58 DHS site sequence including approximately 500 bp upstream and approximately 200 bp downstream is described in SEQ ID NO: 969.

[0105] When ranges are provided in this specification, the endpoints are included. Further, unless otherwise specified or unless otherwise apparent from the context and / or the understanding of those skilled in the art, values expressed as ranges are to be understood as capable of taking any specific value within the described range down to one tenth of the unit of the lower limit of the range, in different embodiments of the invention, unless an exception is explicitly stated in the context. Also, unless otherwise specified or unless otherwise apparent from the context and / or the understanding of those skilled in the art, values expressed as ranges are capable of taking any sub-range within the given range, and the endpoints of the sub-range are to be understood as being expressed with the same precision as one tenth of the unit of the lower limit of the range.

[0106] Summary Various embodiments of the present disclosure generally relate to a genome editing system configured to introduce a modification (e.g., a deletion or insertion or other mutation) into chromosomal DNA so as to enhance the transcription of the HBG1 and / or HBG2 genes encoding the Aγ and Gγ subunits of hemoglobin, respectively. In certain embodiments, an increase in the expression of one or more γ-globin genes (e.g., HBG1, HBG2) using the methods provided herein results in the preferential formation of HbF over HbA and / or an increase in HbF levels as a percentage of total hemoglobin. In certain embodiments, the present disclosure generally relates to the use of an RNP complex comprising a gRNA complexed to a Cpf1 molecule. In certain embodiments, the gRNA may be unmodified or modified, and the Cpf1 molecule may be a wild-type Cpf1 protein or a modified Cpf1 protein. In certain embodiments, the gRNA may comprise a sequence described in Table 13, Table 18, or Table 19. In certain embodiments, the modified Cpf1 may be encoded by a sequence described in SEQ ID NOs: 1000, 1001, 1008-1018, 1032, 1035-1039, 1094-1097, 1107-1109 (Cpf1 polypeptide sequences) or SEQ ID NOs: 1019-1021, 1110-1117 (Cpf1 polynucleotide sequences). In certain embodiments, the RNP complex may comprise an RNP complex described in Table 21. For example, the RNP complex may comprise a gRNA comprising the sequence described in SEQ ID NO: 1051 and a modified Cpf1 protein encoded by the sequence described in SEQ ID NO: 1097 (RNP32 of Table 21).

[0107] Patients with the inherited persistence of fetal hemoglobin (HPFH) condition have mutations in the γ-globin regulatory elements, which have previously been shown to be not suppressed around birth and result in lifelong fetal γ-globin expression (Martyn 2017). This results in an increase in the expression of fetal hemoglobin (HbF). HPFH mutations can be deletion or non-deletion (e.g., point mutations). Subjects with HPFH exhibit lifelong HbF expression, i.e., do not undergo or only partially undergo globin switching without symptoms of anemia.

[0108] HbF expression can be induced through point mutations in γ-globin regulatory elements associated with naturally occurring HPFH variants, such as, for example, HBG1 c.-114 C>T; c.-117 G>A; c.-158 C>T; c.-167 C>T; c.-170 G>A; c.-175 T>G; c.-175 T>C; c.-195 C>G; c.-196 C>T; c.-197 C>T; c.-198 T>C; c.-201 C>T; c.-202 C>T; c.-211 C>T, c.-251 T>C; or c.-499 T>A; or HBG2 c.-109 G>T; c.-110 A>C; c.-114 C>A; c.-114 C>T; c.-114 C>G; c.-157 C>T; c.-158 C>T; c.-167 C>T; c.-167 C>A; c.-175 T>C; c.-197 C>T; c.-200+C; c.-202 C>G; c.-211 C>T; c.-228 T>C; c.-255 C>G; c.-309 A>G; c.-369 C>G; or c.-567 T>G.

[0109] Natural occurring mutations in the distal CCAAT box motif, found within the promoter of the HBG1 and / or HBG2 gene (i.e., HBG1 / 2 c.-111~-115), have also been shown to result in the continued expression of γ-globin and the HPFH phenotype. Modification (mutation or deletion) of the CCAAT box is thought to interfere with the binding of one or more transcriptional repressors and lead to the continued expression of the γ-globin gene and increased HbF expression (Martyn 2017). For example, a naturally occurring 13-base pair del c.-114~-102 (“13nt deletion”) has been shown to be associated with increased HbF levels (Martyn 2017). The distal CCAAT box is likely to overlap with binding motifs within and around the CCAAT box of negative regulatory transcription factors that are expressed in adulthood and suppress HBG (Martyn 2017).

[0110] The gene editing strategies disclosed herein are to increase HbF expression by interfering with one or more nucleotides within and / or around the distal CCAAT box. In certain embodiments, the "CCAAT box target region" can be the region of or in the vicinity of the distal CCAAT box, and includes the nucleotides of the distal CCAAT box, 25 nucleotides upstream (5') and 25 nucleotides downstream (3') of the distal CCAAT box (i.e., HBG1 / 2 c.-86~-140). In other embodiments, the "CCAAT box target region" can be the region of or in the vicinity of the distal CCAAT box, and includes the nucleotides of the distal CCAAT box, 5 nucleotides upstream (5') and 5 nucleotides downstream (3') of the distal CCAAT box (i.e., HBG1 / 2 c.-106~-120). Without limitation, the present disclosure provides for non-naturally occurring, native modifications of the CCAAT box target region that induce HBG expression, including but not limited to HBG del c.-104to-121 ("18nt deletion"), HBG del c.-105to-115 ("11nt deletion"), HBG del c.-112to-115 ("4nt deletion"), and HBG del c.-116 ("1nt deletion"). In certain embodiments, using the genome editing systems disclosed herein, modifications can be introduced into the CCAAT box target region of HBG1 and / or HBG2. In certain embodiments, the genome editing system can include one or more DNA donor templates encoding modifications (such as deletions, insertions, or mutations) of the CCAAT box target region. In certain embodiments, the modification can be a non-naturally occurring modification or a naturally occurring modification. In certain embodiments, the donor template can encode a 1nt deletion, 4nt deletion, 11nt deletion, 13nt deletion, 18nt deletion, or c.-117 G>A modification. In certain embodiments, the genome editing system can include an RNA-guided nuclease, including but not limited to Cas9, modified Cas9, Cpf1, or modified Cpf1. In certain embodiments, the genome editing system can include an RNP comprising a gRNA and a Cpf1 molecule.In certain embodiments, the gRNA can be unmodified or modified, and the Cpf1 molecule can be a wild-type Cpf1 protein, a modified Cpf1 protein, or a combination thereof. In certain embodiments, the gRNA can comprise the sequences set forth in Table 13, Table 18, or Table 19. In certain embodiments, the modified Cpf1 can be encoded by the sequences set forth in SEQ ID NOs: 1000, 1001, 1008-1018, 1032, 1035-1039, 1094-1097, 1107-1109 (Cpf1 polypeptide sequences) or SEQ ID NOs: 1019-1021, 1110-1117 (Cpf1 polynucleotide sequences). In certain embodiments, the RNP complex can comprise the RNP complex set forth in Table 21. For example, the RNP complex can comprise a gRNA comprising the sequence set forth in SEQ ID NO: 1051 and a modified Cpf1 protein encoded by the sequence set forth in SEQ ID NO: 1097 (RNP32 of Table 21).

[0111] HbF expression can also be induced through targeted disruption of the erythroid-specific expression of BCL11A, a transcriptional repressor that encodes a repressor that silences HBG1 and HBG2 expression (Canvers 2015). Another gene editing strategy disclosed herein is to increase HbF expression by targeted disruption of the erythroid-specific enhancer of BCL11A (BCL11Ae) (also discussed in International Publication No. WO 2015 / 148860 by Friedland et al. (「Friedland」), published October 1, 2015, assigned to the assignee of the present invention, which is hereby incorporated by reference in its entirety). In certain embodiments, the region of BCL11Ae targeted for disruption can be the GATA1 binding motif in BCL11Ae. In certain embodiments, using the genome editing system disclosed herein, modifications can be introduced into the GATA1 binding motif in BCL11Ae, the CCAAT box target region, the 13nt target region of HBG1 and / or HBG2, or combinations thereof.

[0112] The genome editing system of the present disclosure includes an RNA-guided nuclease such as Cas9 or Cpf1, one or more gRNAs having a targeting domain complementary to a sequence within or near the target region, and optionally one or more DNA donor templates encoding a specific mutation (such as a deletion or insertion) within or near the target region, and / or without limitation, random oligonucleotides, small molecule agonists or antagonists or peptide agents of gene products involved in DNA repair or DNA damage response, and may include agents that enhance the efficiency of such mutations to occur.

[0113] Various approaches for introducing mutations into the CCAAT box target region, 13 nt target region, proximal HBG1 / 2 promoter target sequence and / or GATA1 binding motif in BCL11Ae can be used in embodiments of the present disclosure. In one approach, a single modification such as a double-strand break is generated within the CCAAT box target region, 13 nt target region, proximal HBG1 / 2 promoter target sequence and / or GATA1 binding motif in BCL11Ae and is repaired in a manner that disrupts the function of the region, for example, by incorporation of a donor template sequence encoding the formation of an indel or deletion of the region. In a second approach, two or more modifications are generated on both sides of the region, resulting in deletion of intervening sequences including the CCAAT box target region, 13 nt target region, GATA1 binding motif in BCL11Ae.

[0114] Treatment of abnormal hemoglobinopathies by gene therapy and / or genome editing is complicated by the fact that red blood cells or RBCs, the cells phenotypically affected by the disease, are enucleated and do not contain the genetic material encoding either the abnormal hemoglobin protein (Hb) subunits or either the Aγ or Gγ subunits that are targeted in the exemplary genome editing approaches described above. This complex situation is addressed, in certain embodiments of the present disclosure, by modification of cells that are capable of differentiating into red blood cells or otherwise giving rise to red blood cells. Cells within the erythroid lineage that are modified according to various embodiments of the present disclosure include, without limitation, hematopoietic stem and progenitor cells (HSC), erythroblasts (including basophilic, polychromatic and / or orthochromatic erythroblasts), proerythroblasts, polychromatic red blood cells or reticulocytes, embryonic stem (ES) cells and / or induced pluripotent stem (iPSC) cells. These cells can be modified in situ (e.g., within the tissue of a subject) or ex vivo. Implementations of genome editing systems for in situ and ex vivo cell modification are described below under the heading "Implementations of Genome Editing Systems: Delivery, Formulations and Routes of Administration."

[0115] In certain embodiments, the modification that results in induction of Aγ and / or Gγ expression is obtained through use of a genome editing system that includes an RNA-guided nuclease and at least one gRNA having a targeting domain complementary to a sequence within or proximal to the CCAAT box target region of HBG1 and / or HBG2 (e.g., within 10, 20, 30, 40 or 50, 100, 200, 300, 400 or 500 bases of the CCAAT box target region). As discussed in more detail below, the RNA-guided nuclease and the gRNA form a complex that can bind to and modify the CCAAT box target region or a region proximal thereto. Examples of suitable gRNAs and gRNA targeting domains directed to the CCAAT box target region of HBG1 and / or HBG2 or a region proximal thereto for use in the embodiments disclosed herein include, without limitation, those set forth in SEQ ID NOs: 251-901, 940-942, 970, 971, 996, 997, 1002 and 1004.

[0116] In certain embodiments, the modifications that result in the induction of Aγ and / or Gγ expression are obtained through the use of a genome editing system comprising an RNA-guided nuclease and at least one gRNA having a targeting domain complementary to a sequence within or proximal to the 13 nt target region of HBG1 and / or HBG2 (e.g., within 10, 20, 30, 40 or 50, 100, 200, 300, 400 or 500 bases of the 13 nt target region). As discussed in more detail below, the RNA-guided nuclease and the gRNA bind to the 13 nt target region or a region proximal thereto to form a complex that can make modifications. Examples of suitable gRNAs and gRNA targeting domains directed to the 13 nt target region of HBG1 and / or HBG2 or a region proximal thereto for use in the embodiments disclosed herein include, without limitation, those set forth in SEQ ID NOs: 251-901, 940-942, 970, 971, 996, 997, 1002 and 1004.

[0117] In certain embodiments, the modifications that result in the induction of HbF expression are obtained through the use of a genome editing system comprising an RNA-guided nuclease and at least one gRNA having a targeting domain complementary to a sequence within or proximal to the GATA1 binding motif in BCL11Ae (e.g., within 10, 20, 30, 40 or 50, 100, 200, 300, 400 or 500 bases of the GATA1 binding motif in BCL11Ae). In certain embodiments, the RNA-guided nuclease and the gRNA bind to the GATA1 binding motif in BCL11Ae to form a complex that can make modifications. Examples of suitable targeting domains directed to the GATA1 binding motif in BCL11Ae for use in the embodiments disclosed herein include, without limitation, those set forth in SEQ ID NOs: 952-955.

[0118] The genome editing system can be implemented in various ways as discussed in detail below. As an example, the genome editing system of the present disclosure can be implemented as a ribonucleoprotein complex or multiple complexes in which multiple gRNAs are used. This ribonucleoprotein complex can be introduced into target cells using methods known in the art, including electroporation, as described in International Publication No. WO 2016 / 182959, jointly assigned to Jennifer Gori and published on November 17, 2016, which is hereby incorporated by reference in its entirety.

[0119] The ribonucleoprotein complexes within these compositions can be introduced into target cells by methods known in the art, including, without limitation, electroporation (e.g., use of nucleofection (trademark), a technology commercialized by Lonza, Basel, Switzerland, or a similar technology commercialized by MaxCyte Inc., Gaithersburg, Maryland) and lipofection (e.g., use of Lipofectamine (trademark) reagents commercialized by Thermo Fisher Scientific, Waltham, Massachusetts). Alternatively or in addition, the ribonucleoprotein complexes can be formed within the target cells following introduction of nucleic acids encoding the RNA-guided nuclease and / or gRNA. These and other delivery modes are described in the following general terms and in Gori.

[0120] Cells modified in vitro according to the present disclosure can be manipulated prior to their delivery to a subject (e.g., expanded, passaged, frozen, differentiated, dedifferentiated, transduced with a transgene, etc.). Cells can be delivered in various ways, either to the subject from which they were obtained ("autologous" transplantation) or to a recipient that is immunologically different from the donor of the cells ("allogeneic" transplantation).

[0121] In some cases, autologous transplantation involves obtaining a plurality of cells from a subject that are circulating in the peripheral blood or circulating within the bone marrow or other tissues (e.g., spleen, skin, etc.), and manipulating these cells (e.g., by inducing to generate iPSCs, purifying cells expressing specific cell surface markers such as CD34, CD90, CD49f, and / or purifying cells that do not express surface markers characteristic of non-erythroid lineages such as CD10, CD14, CD38) to enrich cells within the erythroid lineage. The cells are optionally or additionally grown, transduced with a transgene, exposed to a cytokine or other peptide or small molecule agent, and / or frozen / thawed prior to transduction with a genome editing system that targets the CCAAT box target region, 13nt target region, proximal HBG1 / 2 promoter target sequence, and / or GATA1 binding motif in BCL11Ae. The genome editing system can be implemented or delivered to the cells in any suitable form, such as as a ribonucleoprotein complex, as isolated protein and nucleic acid components, and / or as a nucleic acid encoding the components of the genome editing system.

[0122] In certain embodiments, CD34+ hematopoietic stem and progenitor cells (HSPCs) edited using the genome editing methods disclosed herein can be used for the treatment of abnormal hemoglobinopathies in a subject in need thereof. In certain embodiments, the abnormal hemoglobinopathy can be severe sickle cell disease (SCD) or a thalassemia such as beta thalassemia, delta thalassemia, or beta / delta-thalassemia. In certain embodiments, an exemplary protocol for the treatment of an abnormal hemoglobinopathy can include harvesting CD34+ HSPCs from a subject in need thereof, ex vivo editing of autologous CD34+ HSPCs using the genome editing methods disclosed herein, followed by re-infusing the edited autologous CD34+ HSPCs into the subject. In certain embodiments, treatment with the edited autologous CD34+ HSPCs can result in an increase in HbF induction.

[0123] In certain embodiments, prior to collection of CD34+ HSPCs, the subject may discontinue treatment with hydroxyurea, if applicable, and may receive transfusions to maintain adequate hemoglobin (Hb) levels. In certain embodiments, plerixafor (e.g., 0.24 mg / kg) may be administered intravenously to the subject to mobilize CD34+ HSPCs from the bone marrow into the peripheral blood. In certain embodiments, the subject may undergo one or more leukapheresis cycles (e.g., with approximately one month between cycles, where one cycle is defined as two consecutive plerixafor-mobilized leukapheresis collections). In certain embodiments, the number of leukapheresis cycles performed on the subject is the number necessary to achieve a dose of edited autologous CD34+ HSPCs (e.g., ≥ 1.5 × 106 cells / kg) for re-infusion into the subject, along with a dose of unedited autologous CD34+ HSPCs / kg for backup storage (e.g., ≥ 2 × 106 cells / kg, ≥ 3 × 106 cells / kg, ≥ 4 × 106 cells / kg, ≥ 5 × 106 cells / kg, 2 × 106 cells / kg to 3 × 106 cells / kg, 3 × 106 cells / kg to 4 × 106 cells / kg, 4 × 106 cells / kg to 5 × 106 cells / kg). In certain embodiments, the CD34+ HSPCs collected from the subject may be edited using any of the genome editing methods discussed herein. In certain embodiments, any one or more of the gRNAs and one or more of the RNA-guided nucleases disclosed herein may be used in the genome editing method.

[0124] In certain embodiments, the treatment may include autologous stem cell transplantation. In certain embodiments, the subject may undergo myeloablative conditioning by busulfan conditioning (e.g., dose-adjusted based on first-dose pharmacokinetics analysis at a test dose of 1 mg / kg). In certain embodiments, the conditioning may be performed for 4 consecutive days. In certain embodiments, after a 3-day busulfan drug holiday, edited autologous CD34+ HSPCs (e.g., ≧2×106 cells / kg, ≧3×106 cells / kg, ≧4×106 cells / kg, ≧5×106 cells / kg, 2×106 cells / kg to 3×106 cells / kg, 3×106 cells / kg to 4×106 cells / kg, 4×106 cells / kg to 5×106 cells / kg) may be re-infused into the subject (e.g., into the peripheral blood). In certain embodiments, the edited autologous CD34+ HSPCs may be manufactured and cryopreserved for a particular subject. In certain embodiments, the subject may achieve neutrophil engraftment following a continuous myeloablative conditioning regimen and infusion of edited autologous CD34+ cells. Neutrophil engraftment may be defined as three consecutive measurements of ANC of ≧0.5×109 / L.

[0125] However implemented, the genome editing system may include, or be co-delivered with, one or more factors that improve cell viability during and after editing, including, without limitation, aryl hydrocarbon receptor antagonists such as StemRegenin-1 (SR1), UM171, LGC0006, α-naphthoflavone, and CH-223191, and / or innate immune response antagonists such as cyclosporine A, dexamethasone, resveratrol, MyD88 inhibitory peptide, RNAi agent-targeted Myd88, B18R recombinant protein, glucocorticoids, OxPAPC, TLR antagonists, rapamycin, BX795, and RLR shRNA. These and other factors that improve cell viability during and after editing are described under the heading "I. Optimization of Stem Cells" on pages 36-61 of Gori, which is incorporated herein by reference.

[0126] Following delivery of the genome editing system, the cells are optionally manipulated, for example, HSCs and / or cells within the erythroid lineage are enriched and / or the edited cells are expanded, frozen / thawed, or otherwise prepared for return to the subject. Next, the edited cells are returned to the subject, for example, into the circulatory system or into solid tissue such as bone marrow, by means of intravenous delivery or other means of delivery.

[0127] Functionally, modification of the CCAAT box target region, 13nt target region, proximal HBG1 / 2 promoter target sequence, and / or GATA1 binding motif in BCL11Ae using the compositions, methods, and genome editing systems of the present disclosure results in a significant induction of Aγ and / or Gγ subunit expression (referred to synonymously as HbF expression), such as at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50% or more induction of Aγ and / or Gγ subunit expression in hemoglobin-expressing cells compared to an unmodified control, for example. This induction of protein expression generally includes at least one allele containing a modification of a sequence, such as an indel, insertion, or deletion, within or near the GATA1 binding motif in the CCAAT box target region, 13nt target region, proximal HBG1 / 2 promoter target sequence, and / or BCL11Ae, in at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, etc. of the plurality of cells treated, such as a portion or all of the plurality of cells, and is the result of modification of the GATA1 binding motif in the CCAAT box target region, 13nt target region, proximal HBG1 / 2 promoter target sequence, and / or BCL11Ae (e.g., represented as a percentage of the entire genome containing indel mutations in the plurality of cells).

[0128] The functional effects of modifications caused or facilitated by the genome editing systems and methods of the present disclosure can be evaluated in any number of suitable ways. For example, the effect of a modification on the expression of fetal hemoglobin can be evaluated at the protein or mRNA level. The expression of HBG1 and HBG2 mRNA can be evaluated by digital droplet PCR (ddPCR) performed on cDNA samples obtained by reverse transcription of mRNA taken from treated or untreated samples. Primers for HBG1, HBG2, HBB, and / or HBA can be used individually or multiplexed using methods known in the art. For example, ddPCR analysis of samples can be performed using the QX200™ ddPCR system commercialized by BioRad (Hercules, CA) and the associated protocols published by Bio-Rad. Fetal hemoglobin protein can be evaluated, for example, by high performance liquid chromatography (HPLC) or fast protein liquid chromatography (FPLC) according to the methods discussed on pages 143-144 of Chang 2017 (incorporated herein by reference), using ion exchange columns and / or reverse phase columns that separate HbF, HbB, and HbA and / or Aγ and Gγ globin chains as known in the art.

[0129] The rate at which the CCAAT box target region (e.g., 18nt, 11nt, 4nt, 1nt, c.-117 G>A target region), 13nt target region, proximal HBG1 / 2 promoter target sequence and / or the GATA1 binding motif in BCL11Ae are modified in target cells can be modified by the use of optional genome editing system components such as an oligonucleotide donor template. The design of the donor template is described in the following general terms under the heading "Design of Donor Template". Donor templates for use in targeting the 13nt target region can include, without limitation, modifications (e.g., deletions) encoding HBG1 c.-114 to -102 (corresponding to nucleotides 2824 to 2836 of SEQ ID NO: 902), HBG1 c.-225 to -222 (corresponding to nucleotides 2716 to 2719 of SEQ ID NO: 902)) and / or HBG2 c.-114 to -102 (corresponding to nucleotides 2748 to 2760 of SEQ ID NO: 903). Exemplary full-length donor templates encoding deletions such as c.-114 to -102 and exemplary 5' and 3' homology arms are also presented below (SEQ ID NOs: 904 to 909). In certain embodiments, donor templates for use in targeting the 18nt target region can include, without limitation, donor templates encoding modifications (e.g., deletions) of HBG1 c.-104 to -121, HBG2 c.-104 to -121 or combinations thereof. Exemplary full-length donor templates encoding deletions such as c.-104 to -121 include SEQ ID NOs: 974 and 975. In certain embodiments, donor templates for use in targeting the 11nt target region can include, without limitation, donor templates encoding modifications (e.g., deletions) of HBG1 c.-105 to -115, HBG2 c.-105 to -115 or combinations thereof. Exemplary full-length donor templates encoding deletions such as c.-105 to -115 include SEQ ID NOs: 976 and 978.In certain embodiments, donor templates for use in targeting a 4nt target region may include, without limitation, donor templates encoding modifications (e.g., deletions) of HBG1 c.-112~-115, HBG2 c.-112~-115, or combinations thereof. Exemplary full-length donor templates encoding deletions such as c.-112~-115 include SEQ ID NOs: 984~995. In certain embodiments, donor templates for use in targeting a 1nt target region may include, without limitation, donor templates encoding modifications (e.g., deletions) of HBG1 c.-116, HBG2 c.-116, or combinations thereof. Exemplary full-length donor templates encoding deletions such as c.-116 include SEQ ID NOs: 982 and 983. In certain embodiments, donor templates for use in targeting the c.-117 G>A target region may include, without limitation, donor templates encoding modifications (e.g., deletions) of HBG1 c.-117 G>A, HBG2 c.-117 G>A, or combinations thereof. Exemplary full-length donor templates encoding deletions such as c.-117 G>A include SEQ ID NOs: 980 and 981. In certain embodiments, the donor template can be the plus strand or the minus strand.

[0130] The donor templates used herein can be non-specific templates that are non-homologous to regions of DNA within or near the target sequence. In certain embodiments, donor templates for use in targeting a 13nt target region can include, without limitation, non-target-specific templates that are non-homologous to regions of DNA within or near the 13nt target region. For example, a non-specific donor template for use in targeting a 13nt target region can be non-homologous to a DNA region within or near the 13nt target region and can include a donor template encoding a deletion of HBG1 c.-225~-222 (corresponding to nucleotides 2716~2719 of SEQ ID NO: 902). In certain embodiments, donor templates for use in targeting the GATA1 binding motif in BCL11Ae can include, without limitation, non-target-specific templates that are non-homologous to regions of DNA within or near the GATA1 binding motif target sequence in BCL11Ae. Other donor templates for use in targeting BCL11Ae can include, without limitation, donor templates that include modifications (e.g., deletions) of BCL11Ae including, but not limited to, the GATA1 motif in BCL11Ae.

[0131] The embodiments described herein can be used in all classes of vertebrates including, but not limited to, primates, mice, rats, rabbits, pigs, dogs, and cats.

[0132] This summary focuses on a few exemplary embodiments that depict the principles of genome editing systems and CRISPR-mediated methods for modifying cells. However, for clarity, the present disclosure includes modifications and variations that, although not explicitly recited above, will be apparent to those of ordinary skill in the art. With this in mind, the following disclosure is intended to more generally illustrate the operating principles of genome editing systems. The following should not be understood as limiting, but rather as illustrative of certain principles of genome editing systems and CRISPR-mediated methods using these systems, which, in combination with the present disclosure, will provide those of ordinary skill in the art with information regarding additional implementations and modifications within the scope thereof.

[0133] RNA-induced helicase, guide RNA, and dead guide RNA Various embodiments of the present disclosure generally relate to genome editing systems configured to modify the helical structure of nucleic acids to enhance genome editing of target regions in nucleic acids (e.g., CCAAT box target regions, 13 nt target regions, proximal HBG1 / 2 promoter target sequences, and / or GATA1 binding motifs in BCL11Ae), as well as methods and compositions thereof. Many embodiments relate to the observation that placing an event that modifies the helical structure of DNA within or adjacent to a target region in a nucleic acid can improve the activity of a genome editing system directed to such a target region. Without wishing to be bound by any theory, it is believed that modification of the helical structure (e.g., by unwinding) within or proximal to the DNA target region may induce or increase the accessibility of the genome editing system to the target region, resulting in an increase in the editing of the target region by the genome editing system.

[0134] CRISPR nucleases evolved primarily to defend bacteria against viral pathogens whose genomes are not naturally organized into chromatin. In contrast, eukaryotic genomes are organized into nucleosome units containing genomic DNA segments wrapped around histones. CRISPR nucleases from some bacterial families have been found to be inactive for eukaryotic DNA editing, suggesting that the ability to edit nucleosome-bound DNA may vary by enzyme (Ran 2015). Biochemical evidence indicates that Cas9 from S. pyogenes can efficiently cleave DNA at the ends of nucleosomes but has reduced activity when the target site is located near the center of the nucleosome dyad (Hinz 2016).

[0135] In many cell types, the target site of interest may be strongly bound by nucleosomes or may possess only adjacent PAMs of enzymes that are not efficiently edited in the presence of nucleosomes. In this case, the problematic nucleosome can be first displaced by using an adjacent target site that is closer to the nucleosome edge or is bound by an enzyme that is more effective at binding nucleosomal DNA. However, cleavage at these adjacent sites can be detrimental to the therapeutic strategy. Thus, the availability of programmable enzymes that bind but do not cleave these adjacent sites would enable more efficient functional editing.

[0136] Related strategies utilize the recruitment of exogenous trans-acting factors to facilitate nucleosome replacement. However, the systems and methods of the present disclosure are advantageous over this strategy because they do not require modification of the gRNA beyond truncation of the targeting domain, do not require recruitment of exogenous trans-acting factors, and do not require transcriptional activation to achieve increased editing rates.

[0137] Various approaches for nucleic acid unwinding and modification are employed in various embodiments of the present disclosure. One approach involves unwinding (or opening) a chromatin segment within or proximal to a target region of a nucleic acid in a cell (e.g., a CCAAT box target region, a 13 nt target region, a proximal HBG1 / 2 promoter target sequence, and / or the GATA1 binding motif of BCL11Ae), generating a double-strand break (DSB) within the target region of the nucleic acid, thereby modifying the target region. In certain embodiments, the DSB can be repaired in a manner that modifies the target region. Unwinding a chromatin segment using the methods provided herein can facilitate increased access of a catalytically active RNP (e.g., a catalytically active RNA-guided nuclease and gRNA) to chromatin, enabling more efficient editing of DNA. For example, using these methods, a target region in chromatin that is difficult to access by ribonucleoprotein (e.g., an RNA-guided nuclease complexed with gRNA) because the chromatin is occupied by nucleosomes such as closed chromatin can be edited. In certain embodiments, the unwinding of the chromatin segment occurs via RNA-guided helicase activity. In certain embodiments, the unwinding step does not require recruitment of an exogenous trans-acting factor to the chromatin segment. In certain embodiments, the step of unwinding the chromatin segment does not include forming a single-strand or double-strand break in the nucleic acid within the chromatin segment.

[0138] In certain embodiments of the above approaches and methods, modification of the DNA helix structure is achieved through the action of an "RNA-guided helicase", a term generally used to refer to a molecule, typically a peptide, that (a) interacts (e.g., complexes) with a gRNA and (b) binds to and unwinds a target site together with the gRNA. In certain embodiments, the RNA-guided helicase can include an RNA-guided nuclease configured to lack nuclease activity. However, the inventors have observed that even an RNA-guided nuclease with cleavage ability can be adapted for use as an RNA-guided helicase by complexing it with a dead gRNA having a truncated targeting domain of 15 or fewer nucleotides in length. The complex of the dead gRNA and the wild-type RNA-guided nuclease shows a decrease or loss of RNA cleavage activity but appears to maintain helicase activity. RNA-guided helicases and dead gRNAs are described in more detail below.

[0139] Regarding RNA-induced helicases, according to the present disclosure, the RNA-induced helicase can include any of the RNA-induced nucleases disclosed herein and under the heading "RNA-induced nucleases" hereinafter, including but not limited to Cas9 or Cpf1 RNA-induced nucleases. The helicase activity of these RNA-induced nucleases enables the unwinding of DNA and provides increased access of genomic editing system components (e.g., catalytically active RNA-induced nucleases and gRNAs, without limitation) to the desired target region to be edited (e.g., CCAAT box target region, 13nt target region, proximal HBG1 / 2 promoter target sequence and / or GATA1 binding motif in BCL11Ae). In certain embodiments, the RNA-induced nuclease can be a catalytically active RNA-induced nuclease having nuclease activity. In certain embodiments, the RNA-induced helicase can be configured to lack nuclease activity. For example, in certain embodiments, the RNA-induced helicase can be a catalytically inactive RNA-induced nuclease lacking nuclease activity, such as a catalytically inactive Cas9 molecule, which still provides helicase activity. In certain embodiments, the RNA-induced helicase can complex with a dead gRNA to form a dead RNP that cannot cleave nucleic acids. In other embodiments, the RNA-induced helicase can be a catalytically active RNA-induced nuclease that complexes with a dead gRNA to form a dead RNP that cannot cleave nucleic acids. In certain embodiments, the RNA-induced nuclease is not configured to recruit exogenous trans-acting factors to the desired target region to be edited (e.g., CCAAT box target region, 13nt target region, proximal HBG1 / 2 promoter target sequence and / or GATA1 binding motif in BCL11Ae).

[0140] With reference to the dead gRNAs, these include any of the dead gRNAs discussed herein and under the heading “Dead gRNA Molecules” below. A dead gRNA (also referred to herein as “dgRNA”) can be generated by truncating the 5’ end of the gRNA targeting domain sequence to yield a targeting domain sequence that is 15 nucleotides or less in length. In certain embodiments, the dgRNA can be generated by truncating the 5’ end of any one of the gRNA targeting domain sequences disclosed in Table 2 or Table 10 herein. Dead guide RNA molecules according to the present disclosure include dead guide RNA molecules having reduced, low or undetectable cleavage activity. The dead guide RNA targeting domain sequence can be 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides shorter compared to the targeting domain sequence of the active guide RNA. The dead gRNA molecule can include a targeting domain that is complementary to a region proximal to or within a target region (e.g., CCAAT box target region, 13nt target region, proximal HBG1 / 2 promoter target sequence and / or GATA1 binding motif in BCL11Ae) in the target nucleic acid. In certain embodiments, “proximal” can indicate a region within 10, 25, 50, 100 or 200 nucleotides of a target region (e.g., CCAAT box target region, 13nt target region, proximal HBG1 / 2 promoter target sequence and / or GATA1 binding motif in BCL11Ae). In certain embodiments, the dead gRNA includes a targeting domain that is complementary to the transcribed or non-transcribed strand of the DNA. In certain embodiments, the dead guide RNA is not configured to recruit exogenous trans-acting factors to a target region (e.g., CCAAT box target region, 13nt target region, proximal HBG1 / 2 promoter target sequence and / or GATA1 binding motif in BCL11Ae).

[0141] A method is also provided herein for increasing the rate of indel formation in a target nucleic acid by unwinding DNA within or proximal to a target region (e.g., a CCAAT box target region, a 13 nt target region, a proximal HBG1 / 2 promoter target sequence, and / or a GATA1 binding motif in BCL11Ae) using an RNA-guided helicase to generate a DSB within the target region and forming an indel within the target region through repair of the DSB. The step of unwinding the DNA using the RNA-guided helicase provides increased indel formation as compared to indel formation methods that do not use a helicase.

[0142] The present disclosure further encompasses a method of deleting a segment of a target nucleic acid in a cell, the method comprising contacting the cell with an RNA-guided helicase and generating a double-strand break (DSB) within a target region (e.g., a CCAAT box target region, a 13 nt target region, a proximal HBG1 / 2 promoter target sequence, and / or a GATA1 binding motif in BCL11Ae). In certain embodiments, the RNA-guided helicase binds within or proximal to the target region of the target nucleic acid and is configured to unwind double-stranded DNA (dsDNA) within or proximal to the target region. In certain embodiments, the target nucleic acid is a promoter region of a gene, a coding region of a gene, a non-coding region of a gene, an intron of a gene, or an exon of a gene. In certain embodiments, the segment of the target nucleic acid to be deleted can be at least about 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, or 100 base pairs in length. In certain embodiments, the DSB is repaired by the method of deleting a segment of the target nucleic acid.

[0143] A genome editing system configured to introduce modifications into a helical structure can be implemented in various ways as discussed in detail below. As an example, the genome editing system of the present disclosure can be implemented as a ribonucleoprotein complex or multiple complexes in which multiple gRNAs are used. In certain embodiments, the ribonucleoprotein complex of the genome editing system can be an RNA-induced helicase complexed with a dead guide RNA. The ribonucleoprotein complex can be introduced into target cells using methods known in the art, such as electroporation, as described in Gori. The genome editing system incorporating the RNA-induced helicase can be modified in any suitable manner, including, without limitation, the incorporation of one or more DNA donor templates encoding specific mutations (such as deletions or insertions) within or near the target region and / or, without limitation, enhancing the efficiency of such mutations occurring with agents such as random oligonucleotides, small molecule agonists or antagonists or peptide agents of gene products involved in DNA repair or DNA damage response. These modifications are described in more detail below under the heading "Genome Editing Strategies". For clarity, the present disclosure includes compositions comprising one or more gRNAs, dead gRNAs, RNA-induced helicases, RNA-induced nucleases, or combinations thereof.

[0144] While some of the above exemplary embodiments focus on DNA unwinding, it should be noted that other helical modifications are within the scope of the present disclosure. These include, without limitation, overwinding, underwinding, an increase or decrease in torsional strain (e.g., through topoisomerase activity) on the DNA strand within or adjacent to the target region, denaturation or strand separation, and / or other suitable modifications that result in chromatin structure modification. Each of these modifications can be catalyzed by RNA-induced activity or the recruitment of endogenous factors to the target region.

[0145] Also provided herein are genome editing systems and methods for modifying one or more indels (e.g., indel signatures) generated by an active guide. As the inventors have discovered herein, the pairing of a dead RNP (i.e., a dead guide RNA complexed with an RNA-guided nuclease) with an active RNP (i.e., an active guide RNA complexed with an RNA-guided nuclease) can result in a change in the directionality of indels (e.g., indel signatures) generated by the active RNP alone (without dRNP). As shown in the following examples, the use of a dead guide RNA can result in an increase in the frequency of larger deletions that extend from the active guide RNA cleavage site towards the dead guide RNA binding site. Thus, a dead guide RNA can be used to effectively "direct" deletion editing towards a desired target site. In certain embodiments, the use of a dead guide RNA in combination with an active guide RNA can increase the frequency of deletions not associated with microhomology.

[0146] The examples disclosed in the Examples section below are directed to the modification of a CCAAT box target region, but one of ordinary skill in the art will appreciate that the genome editing systems, methods, cells, and compositions described herein can be used to modify any other target region, e.g., to increase the frequency of deletions in a target region and to increase the frequency of deletions in a target region not associated with microhomology (e.g., not repaired via MMEJ), without limitation.

[0147] This overview focuses on a few exemplary embodiments that depict the principles of genome editing systems and CRISPR-mediated methods for modifying cells. However, for clarity, the present disclosure includes modifications and variations that, while not explicitly recited above, will be apparent to those of ordinary skill in the art. With this in mind, the following disclosure is intended to more generally illustrate the operating principles of genome editing systems. The following should not be understood as limiting, but rather as illustrative of the specific principles of genome editing systems and CRISPR-mediated methods using these systems, which, in combination with the present disclosure, will provide those of ordinary skill in the art with information regarding additional implementations and modifications within its scope.

[0148] Genome editing system The term "genome editing system" refers to any system having RNA-guided DNA editing activity. The genome editing systems of the present disclosure include at least two components, a guide RNA (gRNA) and an RNA-guided nuclease, adapted from naturally occurring CRISPR systems. These two components bind to a specific nucleic acid sequence to form a complex that can edit DNA in or around the nucleic acid sequence by generating, for example, one or more single-strand breaks (SSBs or nicks), double-strand breaks (DSBs), and / or point mutations.

[0149] In certain embodiments, the genome editing system of the present disclosure may include a helicase for unwinding DNA. In certain embodiments, the helicase may be an RNA-guided helicase. In certain embodiments, the RNA-guided helicase may be an RNA-guided nuclease as described herein, such as a Cas9 or Cpf1 molecule. In certain embodiments, the RNA-guided nuclease is not configured to recruit exogenous trans-acting factors to the target region. In certain embodiments, the RNA-guided nuclease may be configured to lack nuclease activity. In certain embodiments, the RNA-guided helicase may complex with a dead guide RNA as disclosed herein. For example, the dead guide RNA (dgRNA) may include a targeting domain sequence less than 15 nucleotides in length. In certain embodiments, the dead guide RNA is not configured to recruit exogenous trans-acting factors to the target region.

[0150] Naturally occurring CRISPR systems are evolutionarily organized into two classes and five types (Makarova 2011, incorporated herein by reference), and while the genome editing system of the present disclosure can accommodate components of any type or class of naturally occurring CRISPR system, the embodiments presented herein are generally adapted from class 2 and type II or V CRISPR systems. Class 2 systems, including type II and V, are characterized by a relatively large multi-domain RNA-guided nuclease protein (e.g., Cas9 or Cpf1) and one or more guide RNAs (e.g., crRNA, optionally tracrRNA), which form a ribonucleoprotein (RNP) complex that associates (targets) with a specific locus complementary to the target (or spacer) sequence of the crRNA and cleaves. The genome editing system according to the present disclosure similarly targets and edits cellular DNA sequences, but is significantly different from naturally occurring CRISPR systems. For example, the single molecule guide RNAs described herein do not exist in nature, and the guide RNAs and RNA-guided nucleases according to the present disclosure can incorporate any number of modifications that do not exist in nature.

[0151] Genome editing systems can be implemented in a variety of ways (e.g., can be administered or delivered to cells or a subject), and different implementations may be suitable for different applications. For example, in certain embodiments, a genome editing system is implemented as a protein / RNA complex (ribonucleoprotein or RNP), which can be included in a pharmaceutical composition optionally containing a pharmaceutically acceptable carrier and / or encapsulant such as, but not limited to, lipid or polymer microparticles or nanoparticles, micelles, or liposomes. In certain embodiments, a genome editing system is implemented as one or more nucleic acids encoding the RNA-guided nuclease and guide RNA components described above (optionally together with one or more additional components); in certain embodiments, a genome editing system is implemented as one or more vectors containing such nucleic acids, such as a viral vector such as an adeno-associated virus (see the following section under the heading "Implementations of Genome Editing Systems: Delivery, Formulation, and Routes of Administration"); in certain embodiments, a genome editing system is implemented as any combination of the foregoing. Additional or modified implementations that operate according to the principles described herein will be apparent to those skilled in the art and are within the scope of the present disclosure.

[0152] It should be noted that the genome editing system of the present disclosure can target or can target a single specific nucleotide sequence and can edit two or more specific nucleotide sequences in parallel by using two or more guide RNAs. The use of multiple gRNAs is referred to throughout the present disclosure as "multiplexing" and can be used to target multiple unrelated target sequences of interest or to form multiple SSBs or DSBs within a single target domain, optionally for making specific edits within such a target domain. For example, International Publication No. 2015 / 138510 by Maeder et al. (referred to herein as "Maeder"), which is incorporated herein by reference, describes a genome editing system for correcting a point mutation (from C.2991+1655A to G) in the human CEP290 gene that results in the generation of a potential splice site and thereby reduces or eliminates the function of the gene. The Maeder genome editing system utilizes two guide RNAs that target the sequences flanking (i.e., sandwiching) the point mutation to form DSBs located on the mutant flank. This then promotes the deletion of the intervening sequence containing the mutation, thereby removing the potential splice site and restoring normal gene function.

[0153] As another example, International Publication No. 2016 / 073990 by Cotta-Ramusino (referred to herein as "Cotta-Ramusino") which is incorporated herein by reference, describes a genome editing system that utilizes two gRNAs in combination with a Cas9 nickase (a Cas9 that creates a single-strand nick such as S. pyogenes D10A) in an arrangement referred to as a "dual nickase system". The dual nickase system of Cotta-Ramusino is configured to create two nicks in the reverse strand of the target sequence that are offset by one or more nucleotides, and the nicks combine to create a double-strand break with an overhang (5' in this case, although a 3' overhang is also possible). The overhang can then, in some circumstances, promote a homology-directed repair event. As another example, International Publication No. 2015 / 070083 by Palestrant et al. (incorporated herein by reference) describes a gRNA (referred to as a "dominant RNA") that targets the nucleotide sequence encoding Cas9, which can be included in a genome editing system that includes one or more additional gRNAs, for example, to enable transient expression of Cas9 that would otherwise be constitutively expressed in some virus-transduced cells. These multiplexing applications are intended to be illustrative rather than limiting, and one of ordinary skill in the art will understand that other multiplexing applications will generally be compatible with the genome editing systems described herein.

[0154] As disclosed herein, in certain embodiments, the genome editing system can include a plurality of gRNAs that can be used to introduce mutations into the GATA1 binding motif in BCL11Ae or the 13 nt target region of HBG1 and / or HBG2. In certain embodiments, the genome editing system disclosed herein can include a plurality of gRNAs that are used to introduce mutations into the GATA1 binding motif in BCL11Ae and the 13 nt target region of HBG1 and / or HBG2.

[0155] A genome editing system can, in some cases, form double-strand breaks that are repaired by cellular DNA double-strand break repair mechanisms such as NHEJ or HDR. These mechanisms have been described throughout the literature (see, for example, Davis & Maizels 2014 (describing Alt-HDR); Frit 2014 (describing Alt-NHEJ); Iyama & Wilson 2013 (describing generally the canonical HDR and NHEJ pathways)).

[0156] When a genome editing system functions by forming DSBs, such a system optionally includes one or more components that promote or facilitate a particular mode of double-strand break repair or a particular repair outcome. For example, Cotta-Ramusino has also described a genome editing system into which a single-stranded oligonucleotide “donor template” has been added; the donor template can be incorporated into the target region of cellular DNA cleaved by the genome editing system, resulting in a change in the target sequence.

[0157] In certain embodiments, the genome editing system modifies a target sequence or modifies the expression of a gene in or near the target sequence without causing a single-stranded or double-stranded break. For example, the genome editing system can include an RNA-guided nuclease fused to a functional domain that acts on DNA, thereby modifying the target sequence or its expression. As an example, the RNA-guided nuclease can bind (e.g., fuse) to a cytidine deaminase functional domain and function by generating a targeted C to A substitution. Exemplary nuclease / deaminase fusions are described in Komor 2016, which is incorporated herein by reference. Alternatively, the genome editing system can utilize a cleavage-inactivated (i.e., “dead”) nuclease such as dead Cas9 (dCas9), which forms a stable complex on one or more targeted regions of cellular DNA, thereby functioning by interfering with functions involving the target region, including but not limited to mRNA transcription, chromatin remodeling, and the like. In certain embodiments, the genome editing system can include an RNA-guided helicase that unwinds DNA within or proximal to the target sequence without causing a single-stranded or double-stranded break. For example, the genome editing system can include an RNA-guided helicase configured to bind within or near the target sequence, unwind the DNA, and induce accessibility to the target sequence. In certain embodiments, the RNA-guided helicase can complex with a dead guide RNA configured to lack cleavage activity and be able to unwind DNA without causing DNA cleavage.

[0158] guide RNA (gRNA) molecule The terms "guide RNA" and "gRNA" refer to any nucleic acid that facilitates specific binding (or "targeting") of an RNA-guided nuclease, such as Cas9 or Cpf1, to a target sequence, such as a genomic or episomal sequence in a cell. The gRNA can be single molecule (including a single RNA molecule or also referred to as chimeric) or modular (e.g., including two or more, typically two separate RNA molecules, such as crRNA and tracrRNA, that normally bind to each other by duplexing). gRNAs and their components have been described throughout the literature (see, e.g., Briner 2014; Cotta-Ramusino, which are incorporated by reference). Examples of modular and single molecule gRNAs that can be used according to embodiments herein include, without limitation, the sequences set forth in SEQ ID NOs: 29-31 and 38-51. Examples of gRNA proximal and tail domains that can be used according to embodiments herein include, without limitation, the sequences set forth in SEQ ID NOs: 32-37.

[0159] In bacteria and archaea, type II CRISPR systems generally include an RNA-guided nuclease protein such as Cas9; a CRISPR RNA (crRNA) containing a 5' region complementary to a foreign sequence; and a trans-activating crRNA (tracrRNA) containing a 5' region that is complementary to and forms a duplex with the 3' region of the crRNA. Without intending to be bound by any theory, this duplex is thought to facilitate the formation of the Cas9 / gRNA complex and be required for its activity. During adapting the type II CRISPR system for use in gene editing, in one non-limiting example, it was discovered that the crRNA (its 3' end) and tracrRNA (its 5' end) can be ligated into a single single molecule or chimeric guide RNA by a 4 nucleotide (e.g., GAAA) "tetraloop" or "linker" sequence that bridges the complementary regions. (Mali 2013; Jiang 2013; Jinek 2012; all incorporated herein by reference).

[0160] The guide RNA, whether single molecule or modular type, contains a "targeting domain" that is fully or partially complementary to a target domain in a target sequence such as a DNA sequence in the genome of a cell where editing is desired. The targeting domain has been called by various names in the literature including, but not limited to, "guide sequence" (see Hsu 2013, incorporated herein by reference), "complementary region" (Cotta-Ramusino), "spacer" (Briner 2014), and collectively "crRNA" (Jiang). Regardless of the names given to them, the targeting domain is typically 10 to 30 nucleotides in length, and in certain embodiments is 16 to 24 nucleotides in length (e.g., 16, 17, 18, 19, 20, 21, 22, 23, or 24 nucleotides in length), and is located at or near the 5' end in the case of Cas9 gRNA and at or near the 3' end in the case of Cpf1 gRNA.

[0161] In addition to the targeting domain, the gRNA typically (but not necessarily as will be discussed below) contains multiple domains that can affect the formation or activity of the gRNA / Cas9 complex. For example, as described above, the double-stranded structure formed by the first and second complementary domains of the gRNA (also referred to as the repeat:anti-repeat duplex) can interact with the recognition (REC) lobe of Cas9 to mediate the formation of the Cas9 / gRNA complex (Nishimasu 2014; Nishimasu 2015; both incorporated herein by reference). It should be noted that the first and / or second complementary domains may contain one or more polyA strands that can be recognized as termination signals by RNA polymerase. Thus, the sequences of the first and second complementary domains can be optionally modified, for example, through the use of A-G swaps or A-U swaps as described in Briner 2014, such that these regions are removed to facilitate complete in vitro transcription of the gRNA. These and other similar modifications to the first and second complementary domains are within the scope of the present disclosure.

[0162] Together with the first and second complementary domains, Cas9 gRNA typically contains two or more additional double-stranded regions that are involved in nuclease activity in vivo but not necessarily in vitro (Nishimasu 2015). The first stem-loop 1 near the 3’ portion of the second complementary domain is variously called the “proximal domain” (Cotta-Ramusino), the “stem-loop 1” (Nishimasu 2014 and 2015), and the “nexus” (Briner 2014). One or more additional stem-loop structures are generally present near the 3’ end of the gRNA, and the number varies by species. S. pyogenes gRNA typically contains two 3’ stem-loops (a total of four stem-loop structures including the repeat:anti-repeat duplex), while S. aureus and other species have only one (a total of three stem-loop structures). An account of the conserved stem-loop structures (and more generally the gRNA structure) grouped by species is provided in Briner 2014.

[0163] The foregoing description has focused on gRNA for use with Cas9, but it should be understood that there are other RNA-guided nucleases that utilize gRNAs that differ in some respects from those described thus far. For example, Cpf1 (CRISPR from “Prevotella and Francicella 1”) is a recently discovered RNA-guided nuclease that does not require a tracrRNA to function. (Zetsche 2015, incorporated herein by reference). The gRNA for use in a Cpf1 genome editing system generally contains a targeting domain and a complementary domain (alternatively referred to as the “handle”). It should also be noted that in the gRNA for use with Cpf1, the targeting domain is typically present at the 3’ end or near it rather than at the 5’ end as described above for Cas9 gRNA (the handle is at the 5’ end or near it of the Cpf1 gRNA). Exemplary targeting domains of Cpf1 gRNA are described in Tables 13 and 18.

[0164] However, one of ordinary skill in the art will understand that while there may be structural differences between gRNAs from different prokaryotic species or between the gRNAs of Cpf1 and Cas9, the principles by which the gRNAs act are generally consistent. Due to this consistency in operation, gRNAs can be defined in a broad sense by their targeting domain sequences, and one of ordinary skill in the art will understand that a given targeting domain sequence can be incorporated into any suitable gRNA, including single-molecule or chimeric gRNAs or gRNAs containing one or more chemical and / or sequence modifications (substitutions, additional nucleotides, truncations, etc.). Thus, for the sake of economy of presentation in the present disclosure, gRNAs can be described only from the perspective of their targeting domain sequences.

[0165] More generally, one of ordinary skill in the art will understand that some aspects of the present disclosure relate to systems, methods, and compositions that can be implemented using multiple RNA-guided nucleases. For this reason, unless otherwise specified, the term gRNA should be understood to encompass not only gRNAs that are compatible with specific species of Cas9 or Cpf1, but also any suitable gRNA that can be used with any RNA-guided nuclease. By way of example, in certain embodiments, the term gRNA can include gRNAs for use with or derived from or adapted from an RNA-guided nuclease present in a class 2 CRISPR system such as type II or type V or any RNA-guided nuclease in a CRISPR system.

[0166] gRNA Design Methods for target array selection and validation, as well as off-target analysis, have been previously described (see, e.g., Mali 2013; Hsu 2013; Fu 2014; Heigwer 2014; Bae 2014; Xiao 2014). Each of these references is hereby incorporated by reference herein. By way of non-limiting example, gRNA design may involve the use of software tools to optimize the selection of potential target sequences corresponding to the user's target sequence, e.g., to minimize overall off-target activity across the genome. Off-target activity is not limited to cleavage, although cleavage efficiency at each off-target sequence can be predicted, e.g., using an experimentally derived weighting scheme. These and other guide selection methods are described in detail by Maeder and Cotta-Ramusino.

[0167] For the selection of gRNA targeting domain sequences directed to the HBG1 / 2 target site (e.g., 13 nt target region), an in silico gRNA target domain identification tool was utilized and hits were stratified into four tiers. In S. pyogenes, the first tier targeting domains were selected based on (1) the distance upstream or downstream from either end of the target site (i.e., the HBG1 / 2 13 nt target region), specifically within 400 bp of either end of the target site, (2) a high level of orthogonality, and (3) the presence of a 5’G. The second tier targeting domains were selected based on (1) the distance upstream or downstream from either end of the target site (i.e., the HBG1 / 2 13 nt target region), specifically within 400 bp of either end of the target site, and (2) a high level of orthogonality. The third tier targeting domains were selected based on (1) the distance upstream or downstream from either end of the target site (i.e., the HBG1 / 2 13 nt target region), specifically within 400 bp of either end of the target site, and (2) the presence of a 5’G. The fourth tier targeting domains were selected based on the distance upstream or downstream from either end of the target site (i.e., the HBG1 / 2 13 nt target region), specifically within 400 bp of either end of the target site.

[0168] In S. aureus, the first layer targeting domain was selected based on (1) the upstream or downstream distance from either end of the target site (i.e., the HBG1 / 2 13nt target region), specifically within 400 bp of either end of the target site, (2) a high level of orthogonality, (3) the presence of 5’G, and (4) a PAM having the NNGRRT sequence (SEQ ID NO: 204). The second layer targeting domain was selected based on (1) the upstream or downstream distance from either end of the target site (i.e., the HBG1 / 2 13nt target), specifically within 400 bp of either end of the target site, (2) a high level of orthogonality, and (3) a PAM having the NNGRRT sequence (SEQ ID NO: 204). The third layer targeting domain was selected based on (1) the upstream or downstream distance from either end of the target site (i.e., the HBG1 / 2 13nt target region), specifically within 400 bp of either end of the target site, and (2) a PAM having the NNGRRT sequence (SEQ ID NO: 204). The fourth layer targeting domain was selected based on (1) the upstream or downstream distance from either end of the target site (i.e., the HBG1 / 2 13nt target), specifically within 400 bp of either end of the target site, and (2) a PAM having the NNGRRV sequence (SEQ ID NO: 205).

[0169] Table 2 below shows the targeting domains of gRNAs of S. pyogenes and S. aureus classified by (a) layer (first, second, third, or fourth) and (b) HBG1 or HBG2.

[0170] [Table 2]

[0171] Additional gRNA sequences designed to target modifications to the CCAAT box target region include, but are not limited to, the sequences set forth in SEQ ID NOs: 970 and 971. Targeting domain sequences of gRNAs designed to target disruption of the CCAAT box target region include, but are not limited to, SEQ ID NO: 1002. Targeting domain sequences of gRNAs designed to target disruption of the CCAAT box target region plus PAM (UUUG) include, but are not limited to, SEQ ID NO: 1004. In certain embodiments, a gRNA comprising the sequence set forth in SEQ ID NO: 1002 and / or 1004 can complex with a Cpf1 protein or a modified Cpf1 protein to generate a modification in the CCAAT box target region. In certain embodiments, a gRNA comprising any of the Cpf1 gRNAs set forth in Table 15, Table 18, or Table 19 can complex with a Cpf1 protein or a modified Cpf1 protein to form an RNP (“gRNA-Cpf1-RNP”) and generate a modification in the CCAAT box target region. In certain embodiments, the modified Cpf1 protein can be His-AsCpf1-nNLS (SEQ ID NO: 1000) or His-AsCpf1-sNLS-sNLS (SEQ ID NO: 1001). In certain embodiments, the Cpf1 molecule of the gRNA-Cpf1-RNP can be encoded by the sequence set forth in SEQ ID NOs: 1000, 1001, 1008-1018, 1032, 1035-39 (Cpf1 polypeptide sequence) or SEQ ID NOs: 1019-1021 (Cpf1 polynucleotide sequence).

[0172] The gRNA can be designed to target the erythroid-specific enhancer of BCL11A (BCL11Ae) to interfere with the expression of the transcriptional repressor BCL11A (described in Friedland, which is incorporated herein by reference). The gRNA is designed to target the GATA1 binding motif in the erythroid-specific enhancer of BCL11A that is in the +58 DHS region of intron 2 (i.e., the GATA1 binding motif in BCL11Ae), and the +58 DHS enhancer region contains the sequence set forth in SEQ ID NO: 968. Targeting domain sequences of the gRNA designed to target the disruption of the GATA1 binding motif in BCL11Ae include, but are not limited to, the sequences set forth in SEQ ID NOs: 952-955. Targeting domain sequences of the gRNA designed to target the disruption of the GATA1 binding motif in BCL11Ae plus PAM (NGG) include, but are not limited to, the sequences set forth in SEQ ID NOs: 960-963.

[0173] gRNA modification The activity, stability or other characteristics of the gRNA can be modified by incorporating specific modifications. As an example, a nucleic acid transiently expressed or delivered can be, for example, prone to degradation by cellular nucleases. Thus, the gRNAs described herein can include one or more modified nucleosides or nucleotides that confer stability against nucleases. Without wishing to be bound by theory, it is also contemplated that certain modified gRNAs described herein may exhibit a reduced innate immune response when introduced into cells. Those skilled in the art are aware of certain cellular responses commonly observed in cells such as mammalian cells in response to exogenous nucleic acids, particularly those of viral or bacterial origin. Such responses, which can include the induction of cytokine expression and release and cell death, can be reduced or completely eliminated by the modifications presented herein.

[0174] The specific exemplary modifications discussed in this section, but not limited to, can be included at any position in the gRNA sequence such as the 5'-end or near it (e.g., within 1 to 10, 1 to 5, or 1 to 2 nucleotides of the 5'-end) and / or the 3'-end or near it (e.g., within 1 to 10, 1 to 5, or 1 to 2 nucleotides of the 3'-end). Optionally, the modification is located within a functional motif such as the repeat:anti-repeat duplex of the Cas9 gRNA, the stem-loop structure of the Cas9 or Cpf1 gRNA, and / or the targeting domain of the gRNA.

[0175] As an example, the 5'-end of the gRNA can contain a eukaryotic mRNA cap structure or a G-cap analog as shown below (e.g., a G(5')ppp(5')G cap analog, an m7G(5')ppp(5')G cap analog, or a 3'-O-Me-m7G(5')ppp(5')G anti-cap analog (ARCA)).

Chemical formula

[0176] In a similar manner, the 5'-triphosphate group can be absent at the 5'-end of the gRNA. For example, in vitro transcribed gRNA can be phosphatase-treated (e.g., by using calf intestinal alkaline phosphatase) to remove the 5'-triphosphate group.

[0177] Another common modification involves adding a plurality of (e.g., 1 to 10, 10 to 20, or 25 to 200) adenine (A) residues, referred to as a polyA tract, to the 3'-end of the gRNA. The polyA sequence can be added to the gRNA during chemical synthesis following in vitro transcription using polyadenosine polymerase (e.g., E. coli poly(A) polymerase) or in vivo by a polyadenylation sequence as described by Maeder.

[0178] The modifications described herein can be combined in any suitable manner. For example, it should be noted that gRNAs transcribed from DNA vectors in vivo or transcribed in vitro can contain either or both of a 5' cap structure or a cap analog and a 3' polyA sequence.

[0179] The guide RNA can be modified at the 3'-terminal U ribose. For example, the two terminal hydroxyl groups of the U ribose can be oxidized to aldehyde groups, and the simultaneous ring-opening of the ribose ring results in the modified nucleoside shown below,

Chemical formula

[0180] The 3'-terminal U ribose can be modified with a 2'3'-cyclic phosphate ester as shown below,

Chemical formula

[0181] The guide RNA can contain 3'-nucleotides that can be stabilized against degradation, for example, by incorporating one or more of the modified nucleotides described herein. In certain embodiments, uridine can be replaced with a modified uridine such as, for example, 5-(2-amino)propyluridine and 5-bromouridine or any of the modified uridines described herein; adenosine and guanosine can be replaced with modified adenosines and guanosines modified, for example, at the 8-position such as 8-bromoguanosine or any of the modified adenosines or guanosines described herein.

[0182] In certain embodiments, sugar-modified ribonucleotides can be incorporated into the gRNA. For example, the 2’OH group can be substituted with a group selected from H, -OR, -R (wherein R can be, for example, alkyl, cycloalkyl, aryl, aralkyl, heteroaryl or sugar), halo, -SH, -SR (wherein R can be, for example, alkyl, cycloalkyl, aryl, aralkyl, heteroaryl or sugar), amino (wherein amino can be, for example, NH2; alkylamino, dialkylamino, heterocyclyl, arylamino, diarylamino, heteroarylamino, diheteroarylamino or amino acid); or cyano (-CN). In certain embodiments, the phosphate backbone can be modified, for example, with phosphorothioate (PhTx) groups as described herein. In certain embodiments, one or more of the nucleotides of the gRNA are each independently 2’-sugar modifications such as 2'-O-methyl, 2'-O-methoxyethyl or modifications or unmodified nucleotides including, but not limited to, 2’-fluoro modifications such as 2’-F or 2'-O-methyl, adenosine (A), 2’-F or 2'-O-methyl, cytidine (C), 2’-F or 2'-O-methyl, uridine (U), 2’-F or 2'-O-methyl, thymidine (T), 1,2’-F or 2'-O-methyl, guanosine (G), 2'-O-methoxyethyl-5-methyluridine (Teo), 2'-O-methoxyethyladenosine (Aeo), 2'-O-methoxyethyl-5-methylcytidine (m5Ce0) and any combination thereof.

[0183] The guide gRNA may also include "locked" nucleic acids (LNAs), where the 2'OH group can be attached, for example, by a C1-6 alkylene or C1-6 heteroalkylene bridge to the 4' carbon of the same ribose sugar. To provide such a bridge, any suitable moiety may be used including, but not limited to, methylene, propylene, ether or amino bridges; O-amino (wherein the amino can be, for example, NH2; alkylamino, dialkylamino, heterocyclyl, arylamino, diarylamino, heteroarylamino or diheteroarylamino, ethylenediamine or polyamino) and aminoalkoxy or O(CH2)n-amino (wherein the amino can be, for example, NH2; alkylamino, dialkylamino, heterocyclyl, arylamino, diarylamino, heteroarylamino or diheteroarylamino, ethylenediamine or polyamino).

[0184] In certain embodiments, the gRNA may include modified nucleotides that are polycyclic (e.g., tricyclic); and "unlocked" forms such as glycol nucleic acid (GNA) (e.g., R-GNA or S-GNA in which ribose is replaced by glycol units attached by phosphodiester bonds) or threose nucleic acid (TNA) in which ribose is replaced by α-L-threofuranosyl-(3'→2').

[0185] Generally, gRNA contains ribose, a 5-membered cyclic sugar with oxygen. Representative modified gRNAs include, but are not limited to, substitution of oxygen in ribose (e.g., by sulfur (S), selenium (Se), or an alkylene such as methylene or ethylene); addition of a double bond (e.g., substituting ribose with cyclopentenyl or cyclohexenyl); ribose ring contraction (e.g., forming a 4-membered ring of cyclobutane or oxetane); ribose ring expansion (e.g., forming a 6- or 7-membered ring with additional carbon or heteroatoms such as anhydrohexitol, altritol, mannitol, cyclohexanyl, cyclohexenyl, and morpholino which also has a phosphoramidate backbone). Most of the sugar analog modifications are located at the 2'-position, but other positions including the 4'-position are also amenable to modification. In certain embodiments, the gRNA contains a 4'-S, 4'-Se, or 4'-C-aminomethyl-2'-O-Me modification.

[0186] In certain embodiments, deazapurines such as, for example, 7-deaza-adenosine can be incorporated into the gRNA. In certain embodiments, O- and N-alkylated nucleotides such as, for example, N6-methyladenosine can be incorporated into the gRNA. In certain embodiments, one or more or all of the nucleotides in the gRNA are deoxyribonucleotides.

[0187] In certain embodiments, the gRNA can be a modified or unmodified gRNA as used herein. In certain embodiments, the gRNA can contain one or more modifications. In certain embodiments, the one or more modifications can include phosphorothioate bond modifications, phosphorodithioate (PS2) bond modifications, 2'-O-methyl modifications, or combinations thereof. In certain embodiments, the one or more modifications can be at the 5'-end of the gRNA, the 3'-end of the gRNA, or combinations thereof.

[0188] In certain embodiments, the gRNA modification can include one or more phosphorodithioate (PS2) bond modifications.

[0189] In some embodiments, the gRNA used herein comprises one or more or a series of deoxyribonucleic acid (DNA) bases, also referred to herein as "DNA extension". In some embodiments, the gRNA used herein comprises a DNA extension at the 5' end of the gRNA, the 3' end of the gRNA, or a combination thereof. In certain embodiments, the DNA extension can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99 or 100 DNA bases in length. For example, in certain embodiments, the DNA extension can be 1, 2, 3, 4, 5, 10, 15, 20 or 25 DNA bases in length. In certain embodiments, the DNA extension can comprise one or more DNA bases selected from adenine (A), guanine (G), cytosine (C) or thymine (T). In certain embodiments, the DNA extension comprises the same DNA base. For example, the DNA extension can comprise a series of adenine (A) bases. In certain embodiments, the DNA extension can comprise a series of thymine (T) bases. In certain embodiments, the DNA extension comprises a combination of different DNA bases. In certain embodiments, the DNA extension can comprise the sequences described in Table 24. For example, the DNA extension can comprise the sequences described in SEQ ID NOs: 1235 to 1250. In certain embodiments, the gRNA used herein comprises a DNA extension and one or more phosphorothioate bond modifications, one or more phosphorodithioate (PS2) bond modifications, one or more 2'-O-methyl modifications, or a combination thereof. In certain embodiments, one or more modifications can be at the 5' end of the gRNA, the 3' end of the gRNA, or a combination thereof.In certain embodiments, the gRNA comprising a DNA extension may comprise the sequences set forth in Table 19 that include a DNA extension. In certain embodiments, the gRNA comprising a DNA extension may comprise the sequence set forth in SEQ ID NO: 1051. In certain embodiments, the gRNA comprising a DNA extension may comprise a sequence selected from the group consisting of SEQ ID NOs: 1046-1060, 1067, 1068, 1074, 1075, 1078, 1081-1084, 1086-1087, 1089-1090, 1092-1093, 1098-1102, and 1106. Without wishing to be bound by theory, any DNA extension may be used herein, provided that it does not hybridize to the target nucleic acid targeted by the gRNA and shows an increase in editing at the target nucleic acid site as compared to a gRNA that does not comprise such a DNA extension.

[0190] In some embodiments, the gRNA used herein comprises one or more or a continuous stretch of ribonucleic acid (RNA) bases, also referred to herein as "RNA extension". In some embodiments, the gRNA used herein comprises an RNA extension at the 5' end of the gRNA, the 3' end of the gRNA, or a combination thereof. In certain embodiments, the RNA extension can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99 or 100 RNA bases in length. For example, in certain embodiments, the RNA extension can be 1, 2, 3, 4, 5, 10, 15, 20 or 25 RNA bases in length. In certain embodiments, the RNA extension can comprise one or more RNA bases selected from adenine (rA), guanine (rG), cytosine (rC) or uracil (rU), where "r" represents RNA, 2'-hydroxy. In certain embodiments, the RNA extension comprises the same RNA base. For example, the RNA extension can comprise a continuous stretch of adenine (rA) bases. In certain embodiments, the RNA extension comprises a combination of different RNA bases. In certain embodiments, the RNA extension can comprise the sequences described in Table 24. For example, the RNA extension can comprise the sequences described in 1231 - 1234, 1251 - 1253. In certain embodiments, the gRNA used herein comprises an RNA extension and one or more phosphorothioate bond modifications, one or more phosphorodithioate (PS2) bond modifications, one or more 2'-O-methyl modifications, or a combination thereof. In certain embodiments, one or more modifications can be at the 5' end of the gRNA, the 3' end of the gRNA, or a combination thereof.In certain embodiments, a gRNA comprising an RNA extension may comprise a sequence set forth in Table 19 that includes an RNA extension. A gRNA comprising an RNA extension at the 5' end of the gRNA may comprise a sequence selected from the group consisting of SEQ ID NOs: 1042-1045, 1103-1105. A gRNA comprising an RNA extension at the 3' end of the gRNA may comprise a sequence selected from the group consisting of SEQ ID NOs: 1070-1075, 1079, 1081, 1098-1100.

[0191] It is contemplated that a gRNA as used herein may also include an RNA extension and a DNA extension. In certain embodiments, either or both of the RNA extension and the DNA extension may be at the 5' end of the gRNA, at the 3' end of the gRNA, or a combination thereof. In certain embodiments, the RNA extension is at the 5' end of the gRNA and the DNA extension is at the 3' end of the gRNA. In certain embodiments, the RNA extension is at the 3' end of the gRNA and the DNA extension is at the 5' end of the gRNA.

[0192] In some embodiments, a gRNA comprising both a 3' end phosphorothioate modification and a 5' end DNA extension complexes with an RNA-guided nuclease, such as Cpf1, to form an RNP, which is then used to edit hematopoietic stem cells (HSCs) or CD34+ cells ex vivo (i.e., outside of the body of the subject from which such cells are derived) at the HBG locus.

[0193] In the usage herein, an example of a gRNA comprises the sequence set forth in SEQ ID NO: 1051.

[0194] Dead gRNA molecule Examples of dead guide RNA (dgRNA) molecules according to the present disclosure include dead guide RNA molecules having reduced, low, or undetectable cleavage activity. The dead guide RNA targeting domain sequence is 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides shorter in length compared to the targeting domain sequence of the active guide RNA. In certain embodiments, the dead guide RNA molecule may include a targeting domain having a length of 15 nucleotides or less, 14 nucleotides or less, 13 nucleotides or less, 12 nucleotides or less, or 11 nucleotides or less. In some embodiments, the dead guide RNAs are configured such that they do not provide an RNA-guided nuclease cleavage event. The dead guide RNA can be generated by removal of the 5' end of the gRNA targeting domain sequence, resulting in a truncated targeting domain sequence. For example, if a gRNA sequence configured to provide a cleavage event (i.e., having a length of 17 nucleotides or more) has a targeting domain sequence that is 20 nucleotides in length, a dead guide RNA can be generated by removing 5 nucleotides from the 5' end of the gRNA sequence. For example, the dgRNA used herein can be truncated from the 5' end of the gRNA sequence and include a targeting domain described in Table 2 or Table 10 having a length of 15 nucleotides or less. In certain embodiments, the dgRNA can be configured to bind (or associate) to a nucleic acid sequence within or proximal to a target region to be edited (e.g., a CCAAT box target region, a 13nt target region, a proximal HBG1 / 2 promoter target sequence, and / or a GATA1 binding motif in BCL11Ae). For example, any of the dgRNAs described in Table 10 can be used to bind to a nucleic acid sequence proximal to the 13nt target region or the CCAAT box target region. In certain embodiments, proximal can refer to a region within 10, 25, 50, 100, or 200 nucleotides of a target region (e.g., a CCAAT box target region, a 13nt target region, a proximal HBG1 / 2 promoter target sequence, and / or a GATA1 binding motif in BCL11Ae). In certain embodiments, the dead guide RNA is not configured to recruit an exogenous trans-acting factor to the target region.In certain embodiments, the dgRNA is configured such that when complexed with an RNA-guided nuclease, it does not provide a DNA cleavage event. One of ordinary skill in the art will understand that a dead guide RNA molecule can be designed to include a targeting domain that is complementary to a region proximal to or within the target region in the target nucleic acid. In certain embodiments, the dead guide RNA includes a targeting domain sequence that is complementary to the transcribed or non-transcribed strand of double-stranded DNA. The dgRNA herein can include modifications at the 5' and 3' ends of the gRNA, as described for the guide RNA in the "gRNA Modifications" section herein. For example, in certain embodiments, the dead guide RNA can include an anti-reverse cap analog (ARCA) at the 5' end of the RNA. In certain embodiments, the dgRNA can include a polyA tail at the 3' end.

[0195] In certain embodiments, the use of the dead guide RNA by the genome editing systems and methods disclosed herein can increase the overall editing level of the active guide RNA. In certain embodiments, the use of the dead guide RNA by the genome editing systems and methods disclosed herein can increase the frequency of deletions. In certain embodiments, the deletion can extend from the cleavage site of the active guide RNA towards the dead guide RNA binding site. In this way, the dead guide RNA can change the directionality of the active guide RNA and direct editing towards the desired target region.

[0196] As used herein, the terms "dead gRNA" and "shortened gRNA" are used synonymously.

[0197] RNA-guided nuclease RNA-guided nucleases according to the present disclosure include, but are not limited to, naturally occurring class 2 CRISPR nucleases such as Cas9 and Cpf1, and other nucleases derived from or obtained from them. Also, certain RNA-guided nucleases such as Cas9 have been shown to have helicase activity that allows them to unwind nucleic acids. In certain embodiments, the RNA-guided helicase according to the present disclosure can be any of the RNA-nucleases described herein and in the section entitled "RNA-guided nuclease" above. In certain embodiments, the RNA-guided nuclease is not configured to recruit exogenous trans-acting factors to the target region. In certain embodiments, the RNA-guided helicase can be an RNA-guided nuclease configured to lack nuclease activity. For example, in certain embodiments, the RNA-guided helicase can be a catalytically inactive RNA-guided nuclease that lacks nuclease activity but still maintains its helicase activity. In certain embodiments, the RNA-guided nuclease can be mutated such that its nuclease activity is abolished (e.g., dead Cas9), generating a catalytically inactive RNA-guided nuclease that cannot cleave nucleic acids but can still unwind DNA. In certain embodiments, it can be an RNA-guided helicase complexed with any of the dead guide RNAs as described herein. For example, a catalytically active RNA-guided helicase (e.g., Cas9 or Cpf1) can form an RNP complex with a dead guide RNA, resulting in a catalytically inactive dead RNP (dRNP). In certain embodiments, a catalytically inactive RNA-guided helicase (e.g., dead Cas9) and a dead guide RNA can form a dRNP. These dRNPs cannot provide a cleavage event but still maintain the helicase activity important for nucleic acid unwinding.

[0198] Functionally, an RNA-guided nuclease is defined as a nuclease that (a) interacts with (e.g., forms a complex with) a gRNA; and (b) together with the gRNA, binds to a target region of DNA that includes a sequence complementary to the targeting domain of the gRNA and optionally (ii) an additional sequence referred to as a "protospacer adjacent motif" or "PAM", described in more detail below, and optionally cleaves or modifies it. As illustrated by the examples below, even though there can be variation between individual RNA-guided nucleases that share the same PAM specificity or cleavage activity, RNA-guided nucleases can be defined in a broad sense by their PAM specificity and cleavage activity. One of ordinary skill in the art will understand that some aspects of the present disclosure relate to systems, methods, and compositions that can be implemented using any suitable RNA-guided nuclease having a particular PAM specificity and / or cleavage activity. For this reason, unless otherwise specified, the term "RNA-guided nuclease" should be understood generically and is not limited to any particular type (e.g., Cpf1 as contrasted with Cas9), species (e.g., S. aureus as contrasted with S. pyogenes), or variation (e.g., truncated or split as contrasted with full length; engineered PAM specificity as contrasted with native PAM specificity, etc.) of an RNA-guided nuclease.

[0199] Various RNA-guided nucleases may require different sequence relationships between the PAM and the protospacer. In general, Cas9 recognizes a PAM sequence that is 3' of the protospacer. In contrast, Cpf1 generally recognizes a PAM sequence that is 5' of the protospacer.

[0200] In addition to recognizing the specific orientation of the PAM and the protospacer, the RNA-guided nuclease can also recognize a specific PAM sequence. For example, S. aureus Cas9 recognizes a PAM sequence of NNGRRT or NNGRRV, where the N residue is adjacent to the 3' of the region recognized by the gRNA targeting domain. S. pyogenes Cas9 recognizes the NGG PAM sequence. Also, F. novicida Cpf1 recognizes the TTN PAM sequence. PAM sequences have been identified for various RNA-guided nucleases, and strategies for identifying novel PAM sequences are described by Shmakov 2015. It should also be noted that engineered RNA-guided nucleases can have a PAM specificity different from that of the reference molecule (e.g., in the case of an engineered RNA-guided nuclease, the reference molecule can be the naturally occurring variant from which the RNA-guided nuclease is derived or a naturally occurring variant having the highest amino acid sequence homology with the engineered RNA-guided nuclease). Examples of PAMs that can be used according to the embodiments herein include, without limitation, the sequences set forth in SEQ ID NOs: 199-205.

[0201] In addition to their PAM specificities, RNA-guided nucleases can be characterized by their DNA cleavage activity, and engineered variants have been generated that, while native RNA-guided nucleases typically form DSBs in the target nucleic acid, yield only SSBs or do not cleave at all (discussed in Ran & Hsu 2013, which is hereby incorporated by reference herein and above).

[0202] Cas9 Crystal structures have been determined for S. pyogenes Cas9 (Jinek 2014) and S. aureus Cas9 complexed with single molecule guide RNA and target DNA (Nishimasu 2014; Anders 2014 and Nishimasu 2015).

[0203] Naturally occurring Cas9 protein comprises two lobes, a recognition (REC) lobe and a nuclease (NUC) lobe, each of which contains specific structural and / or functional domains. The REC lobe contains an arginine-rich bridging helix (BH) domain and at least one REC domain (e.g., REC1 domain, optionally REC2 domain). The REC lobe does not share structural similarity with other known proteins, suggesting that it is a unique functional domain. Without wishing to be bound by any theory, mutational analysis suggests specific functional roles for the BH and REC domains. The BH domain appears to play a role in gRNA:DNA recognition, while the REC domain is thought to interact with the repeat:anti-repeat duplex of the gRNA and mediate the formation of the Cas9 / gRNA complex.

[0204] The NUC lobe contains an RuvC domain, an HNH domain, and a PAM interaction (PI) domain. The RuvC domain shares structural similarity with members of the retroviral integrase superfamily and cleaves the non-complementary (i.e., bottom) strand of the target nucleic acid. It can be formed from two or more split RuvC motifs (such as RuvCI, RuvCII, and RuvCIII in S. pyogenes and S. aureus). On the other hand, the HNH domain is structurally similar to the HNN endonuclease motif and cleaves the complementary (i.e., top) strand of the target nucleic acid. As its name suggests, the PI domain contributes to PAM specificity. Examples of polypeptide sequences encoding Cas9 RuvC-like and Cas9 HNH-like domains that can be used according to embodiments herein are set forth in SEQ ID NOs: 15-23, 52-123 (RuvC-like domains) and SEQ ID NOs: 24-28, 124-198 (HNH-like domains).

[0205] Certain functions of Cas9 are associated with the specific domains described above (although not necessarily fully determined thereby), while these and other functions can be mediated or affected by other Cas9 domains or multiple domains on any of the lobes. For example, as described in Nishimasu 2014, in S. pyogenes Cas9, the repeat:anti-repeat duplex of the gRNA enters the groove between the REC and NUC lobes, and the nucleotides in the duplex interact with the amino acids in the BH, PI, and REC domains. As is the case for some nucleotides in the second and third stem-loops (RuvC and PI domains), some nucleotides in the first stem-loop structure also interact with the amino acids in multiple domains (PI, BH, and REC1). Examples of polypeptide sequences encoding Cas9 molecules that can be used according to the embodiments herein are set forth in SEQ ID NOs: 1-2, 4-6, 12, and 14.

[0206] Cpf1 The crystal structure of Acidaminococcus species Cpf1 complexed with crRNA and a double-stranded (ds) DNA target such as the TTTN PAM sequence has been analyzed by Yamano 2016 (incorporated herein by reference). Similar to Cas9, Cpf1 has two lobes, a REC (recognition) lobe and a NUC (nuclease) lobe. The REC lobe contains the REC1 and REC2 domains, which have no similarity to any known protein structure. In contrast, the NUC lobe contains three RuvC domains (RuvC-I, -II, and -III) and one BH domain. However, in contrast to Cas9, the Cpf1 REC lobe lacks the HNH domain and contains a structurally unique PI domain, which is another domain lacking similarity to known protein structures, three wedge (WED) domains (WED-I, -II, and -III), and a nuclease (Nuc) domain.

[0207] Cas9 and Cpf1 share similarities in structure and function, while it should be understood that certain Cpf1 activities are mediated by structural domains that are not similar to any Cas9 domain. For example, cleavage of the complementary strand of the target DNA appears to be mediated by the HNH domain of Cas9 and the Nuc domain, which is different in sequence and space. Furthermore, the non-targeting portion (handle) of the Cpf1 gRNA adopts a pseudoknot structure rather than the stem-loop structure formed by the repeat:anti-repeat duplex in the Cas9 gRNA.

[0208] In certain embodiments, the Cpf1 protein can be a modified Cpf1 protein. In certain embodiments, the modified Cpf1 protein can include one or more modifications. In certain embodiments, the modifications can be, without limitation, one or more mutations in the Cpf1 nucleotide sequence or Cpf1 amino acid sequence, one or more additional sequences such as His tags or nuclear localization signals (NLS), or combinations thereof. In certain embodiments, the modified Cpf1 can also be referred to herein as a Cpf1 variant.

[0209] In certain embodiments, the Cpf1 protein can be derived from a Cpf1 protein selected from the group consisting of the Acidaminococcus sp. strain BV3L6 Cpf1 protein (AsCpf1), the Lachnospiraceae bacterium ND2006 Cpf1 protein (LbCpf1), and the Lachnospiraceae bacterium MA2020 (Lb2Cpf1). In certain embodiments, the Cpf1 protein can include a sequence selected from the group consisting of SEQ ID NOs: 1016 - 1018, each having a codon-optimized nucleic acid sequence of SEQ ID NOs: 1019 - 1021.

[0210] In certain embodiments, the modified Cpf1 protein may include a nuclear localization signal (NLS). For example, without limitation, NLS sequences useful in connection with the methods and compositions disclosed herein will include amino acid sequences that can facilitate protein import into the cell nucleus. NLS sequences useful in connection with the methods and compositions disclosed herein are known in the art. Examples of such NLS sequences include the nucleoplasmin NLS having the amino acid sequence: KRPAATKKAGQAKKKK (SEQ ID NO: 1006) and the simian virus 40 "SV40" NLS having the amino acid sequence PKKKRKV (SEQ ID NO: 1007).

[0211] In certain embodiments, the NLS sequence of the modified Cpf1 protein is located at or near the C-terminus of the Cpf1 protein sequence. For example, without limitation, the modified Cpf1 protein can be selected from His-AsCpf1-nNLS (SEQ ID NO: 1000); His-AsCpf1-sNLS (SEQ ID NO: 1008); and His-AsCpf1-sNLS-sNLS (SEQ ID NO: 1001), wherein "His" refers to a 6-histidine purification sequence, "AsCpf1" refers to the Acidaminococcus species Cpf1 protein sequence, "nNLS" refers to the nucleoplasmin NLS, and "sNLS" refers to the SV40 NLS. For example, additional permutations of the identity and C-terminal position of the NLS sequence, such as the addition of two or more nNLS sequences or a combination of nNLS and sNLS sequences (or other NLS sequences) and the addition of sequences with or without a purification sequence such as a 6-histidine sequence, are within the scope of the presently disclosed subject matter.

[0212] In certain embodiments, the NLS sequence of the modified Cpf1 protein can be located at or near the N-terminus of the Cpf1 protein sequence. For example, without limitation, the modified Cpf1 protein can be selected from His-sNLS-AsCpf1 (SEQ ID NO: 1009), His-sNLS-sNLS-AsCpf1 (SEQ ID NO: 1010), and sNLS-sNLS-AsCpf1 (SEQ ID NO: 1011). For example, additional permutations of the NLS sequence identity and N-terminal position, such as the addition of two or more nNLS sequences or a combination of nNLS and sNLS sequences (or other NLS sequences), and the addition of sequences with or without a purification sequence such as a 6-histidine sequence, are within the scope of the presently disclosed subject matter.

[0213] In certain embodiments, the modified Cpf1 protein can include NLS sequences located at or near both the N-terminus and C-terminus of the Cpf1 protein sequence. For example, without limitation, the modified Cpf1 protein can be selected from His-sNLS-AsCpf1-sNLS (SEQ ID NO: 1012) and His-sNLS-sNLS-AsCpf1-sNLS-sNLS (SEQ ID NO: 1013). For example, additional permutations of the NLS sequence identity and N-terminal / C-terminal position, such as the addition of two or more nNLS sequences or a combination of nNLS and sNLS sequences (or other NLS sequences) to either the N-terminal / C-terminal position, and the addition of sequences with or without a purification sequence such as a 6-histidine sequence, are within the scope of the presently disclosed subject matter.

[0214] In certain embodiments, the modified Cpf1 protein can include a modification (e.g., deletion or substitution) at one or more cysteine residues of the Cpf1 protein sequence. For example, but not limited to, the modified Cpf1 protein can include a modification at a position selected from the group consisting of C65, C205, C334, C379, C608, C674, C1025, and C1248. In certain embodiments, the modified Cpf1 protein can include a substitution of one or more cysteine residues with serine or alanine. In certain embodiments, the modified Cpf1 protein can include a modification selected from the group consisting of C65S, C205S, C334S, C379S, C608S, C674S, C1025S, and C1248S. In certain embodiments, the modified Cpf1 protein can include a modification selected from the group consisting of C65A, C205A, C334A, C379A, C608A, C674A, C1025A, and C1248A. In certain embodiments, the modified Cpf1 protein can include a modification at positions C334 and C674 or C334, C379, and C674. In certain embodiments, the modified Cpf1 protein can include the modification of C334S and C674S or C334S, C379S, and C674S. In certain embodiments, the modified Cpf1 protein can include the modification of C334A and C674A or C334A, C379A, and C674A. In certain embodiments, the modified Cpf1 protein can include both one or more cysteine residue modifications, such as His-AsCpf1-nNLS Cys-less (SEQ ID NO: 1014) or His-AsCpf1-nNLS Cys-low (SEQ ID NO: 1015), and the introduction of one or more NLS sequences. In various embodiments, a Cpf1 protein comprising a deletion or substitution of one or more cysteine residues exhibits a decrease in aggregation.

[0215] In certain embodiments, other modified Cpf1 proteins known in the art can be used with the methods and systems described herein. For example, in certain embodiments, the modified Cpf1 can be Cpf1 containing the mutations S542R / K548V / N552R (“Cpf1 RVR”). Cpf1 RVR has been shown to cleave target sites having a TATV PAM. In certain embodiments, the modified Cpf1 can be Cpf1 containing the mutations S542R / K607R (“Cpf1 RR”). Cpf1 RR has been shown to cleave target sites having a TYCV / CCCC PAM.

[0216] In some embodiments, Cpf1 variants are used herein, where the Cpf1 variants are 11, 12, 13, 14, 15, 16, 17, 34, 36, 39, 40, 43, 46, 47, 50, 54, 57, 58, 111, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 532, 533, 534, 535, 536, 537, 538, 539, 540, 541, 542, 543, 544, 545, 546, 547, 548, 549, 550, 551, 552, 553, 554, 555, 556, 565, 566, 567, 568, 569, 570, 571, 572, 573, 574, 575, 592, 593, 594, 595, 596, 597, 598, 599, 600, 601, 602, 603, 604, 605, 606, 607, 608, 609, 610, 611, 612, 613, 614, 615, 616, 617, 618, 619, 620, 626, 627, 628, 629, 630, 631, 632, 633, 634, 635, 636, 637, 638, 642, 643, 644, 645, 646, 647, 648, 649, 651, 652, 653, 654, 655, 656, 676, 679, 680, 682, 683, 684, 685, 686, 687, 688, 689, 690, 691, 692, 693, 707, 711, 714, 715, 716, 717, 718, 719, 720, 721, 722, 739, 765, 768, 769, 773, 777, 778, 779, 780, 781, 782, 783, 784, 785, 786, 870, 871, 872, 873, 874, 875, 876, 877, 878, 879, 880, 881, 882, 883, 884 or 1048 or one or more residues of AsCpf1 (Acidaminococcus sp. BV3L6) selected from the group consisting of corresponding positions of orthologs, homologs or variants of AsCpf1 contain mutations.

[0217] In certain embodiments, the Cpf1 variant can include any of the Cpf1 proteins described in International Publication No. WO 2017 / 184768 by Zhang et al. (the “’768 Publication”), which is incorporated herein by reference for all purposes in the context of this disclosure.

[0218] In certain embodiments, the modified Cpf1 protein (also referred to as a Cpf1 variant) used herein can be encoded by any of the sequences set forth in SEQ ID NOs: 1000, 1001, 1008-1018, 1032, 1035-39, 1094-1097, 1107-09 (Cpf1 polypeptide sequences) or SEQ ID NOs: 1019-1021, 1110-17 (Cpf1 polynucleotide sequences). Table 20 sets forth exemplary Cpf1 variant amino acid and nucleotide sequences. These sequences are described in FIG. 62, which details the positions of the 6-histidine sequence (underlined characters) and the NLS sequence (bold). For example, additions to either the N-terminus / C-terminus positions of two or more nNLS sequences or combinations of nNLS and sNLS sequences (or other NLS sequences), and permutations of the NLS sequence identity and N-terminus / C-terminus positions that include or do not include the addition of a purification sequence such as, for example, a 6-histidine sequence, are within the scope of the presently disclosed subject matter.

[0219] In certain embodiments, either the Cpf1 protein or the modified Cpf1 protein disclosed herein can complex with one or more gRNAs comprising the targeting domain set forth in SEQ ID NO: 1002 and / or 1004 to modify the CCAAT box target region. In certain embodiments, either the Cpf1 protein or the modified Cpf1 protein disclosed herein can complex with one or more gRNAs comprising the sequences set forth in Table 18 or Table 19. In certain embodiments, the modified Cpf1 protein can be His-AsCpf1-nNLS (SEQ ID NO: 1000) or His-AsCpf1-sNLS-sNLS (SEQ ID NO: 1001). In certain embodiments, the modified Cpf1 protein used herein can be encoded by any of the sequences set forth in SEQ ID NO: 1000, 1001, 1008-1018, 1032, 1035-1039, 1094-1097, 1107-1109 (Cpf1 polypeptide sequences) or SEQ ID NO: 1019-1021, 1110-1117 (Cpf1 polynucleotide sequences). In certain embodiments, the modified Cpf1 protein can comprise the sequence set forth in SEQ ID NO: 1097.

[0220] Modification of RNA-guided nucleases The RNA-guided nucleases described above can have activities and properties useful for various applications, and those skilled in the art will recognize that RNA-guided nucleases can also be optionally modified such that their cleavage activity, PAM specificity, or other structural or functional characteristics can be altered.

[0221] First, turning to modifications that alter cleavage activity, mutations that reduce or eliminate the activity of domains within the NUC lobe are described above. Exemplary mutations that can be generated in the RuvC domain, Cas9 HNH domain, or Cpf1 Nuc domain are described in Ran & Hsu 2013 and Yamano and Cotta-Ramusino. Generally, mutations that reduce or eliminate the activity of one of the two nuclease domains result in an RNA-guided nuclease with nickase activity, but it should be noted that the type of nickase activity varies depending on which domain is inactivated. As an example, inactivation of the RuvC domain of Cas9 results in a nickase that cleaves the complementary strand or the top strand, as shown below (where C indicates the site of cleavage).

[0222] On the other hand, inactivation of the Cas9 HNH domain results in a nickase that cleaves the bottom strand or the non-complementary strand.

[0223] Modifications of PAM specificity compared to the naturally occurring Cas9 reference molecule have been described by Kleinstiver et al. for both S. pyogenes (Kleinstiver 2015a) and S. aureus (Kleinstiver 2015b). Kleinstiver et al. have also described modifications that improve the targeting fidelity of Cas9 (Kleinstiver 2016). Each of these references is incorporated herein by reference.

[0224] RNA-guided nucleases are split into two or more parts, as described by Zetsche 2015 and Fine 2015 (both of which are incorporated herein by reference).

[0225] RNA-guided nucleases can be size-optimized or shortened in certain embodiments through one or more deletions that reduce the size of the nuclease while still maintaining, for example, gRNA binding, target and PAM recognition, and cleavage activity. In certain embodiments, the RNA-guided nuclease is optionally covalently or non-covalently bound, by a linker, to another polypeptide, nucleotide, or other structure. Exemplary binding nucleases and linkers are described by Guilinger 2014, which is hereby incorporated by reference for all purposes.

[0226] The RNA-guided nuclease optionally includes a label, including but not limited to a nuclear localization signal, to facilitate the movement of the RNA-guided nuclease protein into the nucleus. In certain embodiments, the RNA-guided nuclease can incorporate a C-terminal and / or N-terminal nuclear localization signal. Nuclear localization sequences are known in the art and are described in Maeder and elsewhere.

[0227] The foregoing list of modifications is intended to be exemplary in nature, and one of ordinary skill in the art will understand, in view of the present disclosure, that other modifications may be possible or desirable in a particular application. Thus, for the sake of brevity, the exemplary systems, methods, and compositions of the present disclosure are presented with reference to particular RNA-guided nucleases, but it should be understood that the RNA-guided nucleases used can be modified in a manner that does not change their operating principles. Such modifications are within the scope of the present disclosure.

[0228] Nucleic acid encoding an RNA-guided nuclease For example, nucleic acids encoding RNA-guided nucleases such as Cas9, Cpf1, or functional fragments thereof are provided herein. Examples of nucleic acid sequences encoding Cas9 molecules that can be used according to the embodiments herein are set forth in SEQ ID NOs: 3, 7-11, 13. Exemplary nucleic acids encoding RNA-guided nucleases have been described previously (see, e.g., Cong 2013; Wang 2013; Mali 2013; Jinek 2012).

[0229] Optionally, the nucleic acid encoding the RNA-guided nuclease can be a synthetic nucleic acid sequence. For example, the synthetic nucleic acid molecule can be chemically modified. In certain embodiments, the mRNA encoding the RNA-guided nuclease will have one or more (e.g., all) of the following characteristics: it can be capped; it can be polyadenylated; it can be substituted with 5-methylcytidine and / or pseudouridine.

[0230] The synthetic nucleic acid sequence can also be codon-optimized, e.g., at least one non-common or less common codon is replaced by a common codon. For example, the synthetic nucleic acid can induce the synthesis of an optimized messenger mRNA optimized for expression in a mammalian expression system, such as described herein. Examples of codon-optimized Cas9 coding sequences are presented in Cotta-Ramusino.

[0231] Additionally or alternatively, the nucleic acid encoding the RNA-guided nuclease can include a nuclear localization sequence (NLS). Nuclear localization sequences are known in the art.

[0232] Functional analysis of candidate molecules Candidate RNA-guided nucleases, gRNAs, and their complexes can be evaluated by standard methods known in the art. See, e.g., Cotta-Ramusino. The stability of the RNP complex can be evaluated by differential scanning fluorimetry as described below.

[0233] Differential scanning fluorimetry (DSF) The thermal stability of a ribonucleoprotein (RNP) complex containing a gRNA and an RNA-guided nuclease can be measured by DSF. The DSF technique measures the thermal stability of a protein that can increase under favorable conditions such as the addition of a binding RNA molecule such as a gRNA.

[0234] The DSF assay can be performed according to any suitable protocol, including but not limited to: (a) testing different conditions (e.g., different stoichiometric ratios of gRNA:RNA-guided nuclease protein, different buffers, etc.) to identify optimal conditions for RNP formation; and (b) testing modifications of the RNA-guided nuclease and / or gRNA (e.g., chemical modifications, sequence alterations, etc.) to identify modifications that improve RNP formation or stability. One readout of the DSF assay is the change in the melting temperature of the RNP complex; a relatively high change indicates that the RNP complex is more stable (and thus may have greater activity or more favorable formation, degradation kinetics, or other functional characteristics) compared to a standard RNP complex characterized by a lower change. When the DSF assay is developed as a screening tool, a threshold melting temperature change can be identified, such that the results are one or more RNPs with a melting temperature change above the threshold. For example, the threshold can be 5-10 °C (e.g., 5 °C, 6 °C, 7 °C, 8 °C, 9 °C, 10 °C) or higher, and the results can be one or more RNPs characterized by a melting temperature change above the threshold.

[0235] Two non-limiting examples of DSF assay conditions are as follows.

[0236] To determine the best solution for forming the RNP complex, dispense a fixed concentration in water (e.g., 2 μM) of Cas9 + 10×SYPRO Orange® (Life Technologies, catalog number S-6650) into a 384-well plate. Next, add equimolar amounts of gRNA diluted in solutions with various pHs and salts. Incubate for 10 minutes at room temperature, briefly centrifuge to remove any bubbles, and then perform a gradient from 20°C to 90°C in 1°C increments every 10 seconds using a Bio-Rad CFX384™ Real-Time System C1000 Touch™ Thermal Cycler with Bio-Rad CFX Manager software.

[0237] The second assay consists of mixing various concentrations of gRNA with a fixed concentration (e.g., 2 μM) of Cas9i in the optimal buffer from Assay 1 above and incubating in a 384-well plate (e.g., for 10 minutes at room temperature). Add an equal volume of optimal buffer + 10×SYPRO Orange® (Life Technologies, catalog number S-6650), and seal the plate with Microseal® B adhesive (MSB-1001). Briefly centrifuge to remove any bubbles, and then perform a gradient from 20°C to 90°C in 1°C increments every 10 seconds using a Bio-Rad CFX384™ Real-Time System C1000 Touch™ Thermal Cycler with Bio-Rad CFX Manager software.

[0238] Genome Editing Strategies Using the genome editing systems described above, in various embodiments of the present disclosure, edit (i.e., modify) within the target region of DNA obtained intracellularly or from cells. Various strategies for making specific edits are described herein, and these strategies are generally described by the desired repair outcome, the number and location of individual edits (e.g., SSB or DSB), and the target site of such edits.

[0239] Genome editing strategies involving the formation of SSBs or DSBs are characterized by repair outcomes including (a) deletion of all or part of a target region; (b) insertion or replacement within all or part of a target region; or (c) disruption of all or part of a target region. This grouping is not intended to limit or be bound by any particular theory or model and is provided only for the economy of presentation. One of ordinary skill in the art will understand that the listed outcomes are not mutually exclusive and that some repairs may result in other outcomes. The description of a particular editing strategy or method should not be understood as requiring a particular repair outcome unless otherwise specified.

[0240] Replacement of a target region generally involves replacement of all or part of a sequence present within the target region by a homologous sequence, through gene correction or gene conversion, which are two repair outcomes mediated, for example, by the HDR pathway. HDR is facilitated by the use of a donor template, which can be single-stranded or double-stranded, as described in more detail below. The single-stranded or double-stranded template can be exogenous, in which case they promote gene correction, or they can be endogenous (e.g., homologous sequences in the cell genome) and promote gene conversion. Exogenous templates can have, for example, asymmetric overhangs (i.e., the portion of the template complementary to the site of the DSB can be offset in the 3' or 5' direction rather than being centered within the template), as described by Richardson 2016 (incorporated herein by reference). When the template is single-stranded, it can correspond to either the complementary (top) or non-complementary (bottom) strand of the target region.

[0241] As described by Ran & Hsu and Cotta-Ramusino, gene conversion and gene correction are sometimes facilitated by forming one or more nicks within or around the target region. Optionally, a double nickase strategy is used to form two offset SSBs, which are then used to form a single DSB with an overhang (e.g., a 5' overhang).

[0242] Disruptions and / or deletions of all or part of the target array can be achieved by various repair outcomes. As an example, as described by Maeder for the LCA10 mutation, the array can be deleted by simultaneously generating two or more DSBs flanking the target region, which is then excised when the DSBs are repaired. As another example, the array can be disrupted by the formation of a double-strand break with single-stranded overhangs, followed by deletions generated by nucleotide end-processing of the overhangs prior to repair.

[0243] One particular subset of target array disruptions is mediated by the formation of indels in the target array, where the repair outcome is typically mediated by the NHEJ pathway (including Alt-NHEJ). NHEJ is referred to as an "error-prone" repair pathway due to its association with indel mutations. However, in some cases, the DSB is repaired by NHEJ without modification of the surrounding sequence (so-called "perfect" or "scarless" repair); this generally requires that both ends of the DSB be fully ligated. On the other hand, indels are thought to result from enzymatic processing of the free DNA ends before they are ligated, which adds and / or removes nucleotides to one or both strands of one or both free ends.

[0244] Because the enzymatic processing of free DSB ends can be inherently stochastic, indel mutations tend to vary, occur along a distribution, and can be affected by various factors including specific target sites, cell types used, and genome editing strategies employed. Nevertheless, it is possible to make limited generalizations about indel formation: Deletions formed by the repair of a single DSB are most commonly in the range of 1 - 50 bp, but can exceed 100 - 200 bp. Insertions formed by the repair of a single DSB tend to be shorter and often contain short duplications of the sequence directly surrounding the cleavage site. However, it is possible to obtain large insertions, and in these cases, the inserted sequences have often been traced to other regions of the genome or plasmid DNA present within the cell.

[0245] Indel mutations and genome editing systems configured to generate indels are useful for disrupting target sequences, for example, when the generation of a specific final sequence is not required and / or when frameshift mutations are tolerated. They can also be useful in situations where a specific sequence is preferred, as long as the desired specific sequence has a tendency to preferentially result from the repair of SSBs or DSBs at a given site. Indel mutations are also useful tools for evaluating or screening the activity of specific genome editing systems and their components. In these and other settings, indels can be characterized by (a) their relative and absolute frequencies in the genome of cells contacted with the genome editing system, and (b) the distribution of numerical differences such as ±1, ±2, ±3, etc. relative to the unedited sequence. As an example, in a lead discovery setting, multiple gRNAs can be screened to identify the gRNA that most efficiently promotes cleavage at the target site based on indel reads under controlled conditions. Guides that generate indels above a threshold frequency or that generate a specific distribution of indels can be selected for further research and development. The frequency and distribution of indels can also be useful as a readout for evaluating the implementation or formulation and delivery methods of different genome editing systems, for example, by keeping the gRNA constant and varying other specific reaction conditions or delivery methods.

[0246] Multiple strategies The genome editing system according to the present disclosure can also be used for multiplex gene editing that generates two or more DSBs at either the same locus or different loci. Any of the RNA-guided nucleases and gRNAs disclosed herein can be used in a genome editing system for multiplex gene editing. Editing strategies involving the formation of multiple DSBs or SSBs are described, for example, in Cotta-Ramusino.

[0247] As disclosed herein, multiple gRNAs can be used in a genome editing system to introduce modifications (e.g., deletions, insertions) into the 13nt target regions of HBG1 and / or HBG2. In certain embodiments, one or more gRNAs comprising the targeting domains described in SEQ ID NOs: 251-901, 940-942 can be used to introduce modifications into the 13nt target regions of HBG1 and / or HBG2. In other embodiments, multiple gRNAs can be used in a genome editing system to introduce modifications into the CCAAT box target region. In certain embodiments, one or more gRNAs comprising the sequences described in SEQ ID NOs: 970, 971, 996, 997 can be used to introduce modifications into the CCAAT box target region. In certain embodiments, one or more gRNAs comprising the targeting domains described in SEQ ID NOs: 1002, 1004 can be used to introduce modifications into the CCAAT box target region. In other embodiments, multiple gRNAs can be used in a genome editing system to introduce modifications into the GATA1 binding motif in BCL11Ae. In certain embodiments, one or more gRNAs comprising the targeting domains described in SEQ ID NOs: 952-955 can be used to introduce modifications into the GATA1 binding motif of BCL11Ae. Multiple gRNAs can also be used in a genome editing system to introduce modifications into the GATA1 binding motif, CCAAT box target region, 13nt target regions of HBG1 and / or HBG2, or combinations thereof in BCL11Ae. In certain embodiments, one or more gRNAs comprising the targeting domains described in SEQ ID NOs: 952-955 can be used to introduce modifications into the GATA1 binding motif in BCL11Ae, and one or more gRNAs comprising the targeting domains described in SEQ ID NOs: 251-901, 940-942 can be used to introduce modifications into the 13nt target regions of HBG1 and / or HBG2. In certain embodiments, one or more gRNAs comprising the targeting domains described in SEQ ID NOs: 952-955 can be used to introduce modifications into the GATA1 binding motif in BCL11Ae, and one or more gRNAs comprising the targeting domains described in SEQ ID NOs: 970, 971, 996, 997 can be used to introduce modifications into the CCAAT box target region.In certain embodiments, one or more gRNAs comprising the targeting domains set forth in SEQ ID NOs: 952-955 can be used to introduce modifications into the GATA1 binding motifs in BCL11Ae, and one or more gRNAs comprising the targeting domains set forth in SEQ ID NOs: 1002, 1004 can be used to introduce modifications into the CCAAT box target region.

[0248] In certain embodiments, multiple gRNAs and an RNA-guided nuclease are used in a genome editing system to introduce modifications (e.g., deletions, insertions) into the CCAAT box target region of HBG1 and / or HBG2. In certain embodiments, the RNA-guided nuclease can be Cas9, a modified Cas9 (e.g., D10A), Cpf1, or a modified Cpf1.

[0249] Donor Template Design Donor template design is described in detail, for example, in the literature such as Cotta-Ramusino. A DNA oligomer donor template (oligodeoxynucleotide or ODN), which can be single-stranded (ssODN) or double-stranded (dsODN), can be used to facilitate repair based on HDR of DSBs or to improve the overall editing rate, which is particularly useful for introducing modifications into a target DNA sequence, inserting a new sequence into the target sequence, or completely replacing the target sequence.

[0250] Regardless of whether it is single-stranded or double-stranded, the donor template generally includes regions homologous to the DNA region in or near the target sequence to be cleaved (e.g., flanking or adjacent). These homologous regions are referred to herein as "homology arms" and are schematically shown below. [5'Homology Arm]--[Replacement Sequence]--[3'Homology Arm]

[0251] The homology arms can have any suitable length (including zero nucleotides if only one homology arm is used), and the 3' and 5' homology arms can have the same length or different lengths. The choice of appropriate homology arm length can be influenced by various factors, such as the desire to avoid homology or microhomology with specific sequences, such as Alu repeats or other very common elements. For example, the 5' homology arm can be shortened to avoid sequence repeat elements. In other embodiments, the 3' homology arm can be shortened to avoid sequence repeat elements. In some embodiments, both the 5' and 3' homology arms can be shortened to avoid inclusion of specific sequence repeat elements. Additionally, some homology arm designs can improve editing efficiency or increase the frequency of desired repair outcomes. For example, Richardson 2016, which is incorporated herein by reference, found that the relative asymmetry of the 3' and 5' homology arms of a single-stranded donor template affects the repair rate and / or outcome.

[0252] The replacement sequences in the donor template are described elsewhere, including by Cotta-Ramusino. The replacement sequences can be of any suitable length (including zero nucleotides if the desired repair outcome is a deletion) and typically include one, two, three or more sequence modifications relative to the native sequence in the cell where editing is desired. One common sequence modification involves altering the native sequence to repair a mutation associated with a disease or condition for which treatment is desired. Another common sequence modification involves one or more sequences that are complementary to the PAM sequence of an RNA-guided nuclease or the targeting domain of a gRNA used to generate an SSB or DSB, or thus modifying the PAM sequence or targeting domain to reduce or eliminate repeated cleavage of the target site after the replacement sequence is incorporated into the target site.

[0253] When a linear ssODN is used, it can be configured to anneal to (i) the nicked strand of the target nucleic acid, (ii) the intact target nucleic acid strand, (iii) the plus strand of the target nucleic acid, and / or (iv) the minus strand of the target nucleic acid. The ssODN can have any suitable length, for example, about or at least 80 to 200 nucleotides or less (e.g., 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, or 200 nucleotides).

[0254] It should be noted that the template nucleic acid can also be a nucleic acid vector such as a viral genome or a circular double-stranded DNA such as a plasmid. The nucleic acid vector containing the donor template can contain other coding or non-coding elements. For example, the template nucleic acid can contain specific genomic backbone elements (e.g., inverted terminal repeats in the case of the AAV genome) and can be delivered as part of a viral genome that optionally contains additional sequences encoding gRNA and / or RNA-guided nucleases (e.g., in the AAV or lentiviral genome). In certain embodiments, the donor template can be adjacent to or flanked by a target site recognized by one or more gRNAs, facilitating the formation of free DSBs at one or both ends of the donor template, which can participate in the repair of corresponding SSBs or DSBs formed in cellular DNA using the same gRNA. Exemplary nucleic acid vectors suitable for use as donor templates are described in Cotta-Ramusino, which is incorporated by reference.

[0255] Regardless of the form used, the template nucleic acid can be designed to avoid unwanted sequences. In certain embodiments, one or both of the homology arms can be shortened, for example, to avoid overlap with specific sequence repeat elements such as Alu repeats, LINE elements, etc.

[0256] In certain embodiments, silent non-pathogenic SNPs can be included in the ssODN donor template to enable the identification of gene editing events.

[0257] In certain embodiments, the donor template can be a non-specific template that is non-homologous to the region of DNA within or near the target sequence to be cleaved. In certain embodiments, donor templates for use in targeting the GATA1 binding motif in BCL11Ae can include, without limitation, non-target-specific templates that are non-homologous to the region of DNA within or near the GATA1 binding motif in BCL11Ae. In certain embodiments, donor templates for use in targeting the 13nt target region can include, without limitation, non-target-specific templates that are non-homologous to the region of DNA within or near the 13nt target region.

[0258] A donor template or template nucleic acid, as used herein, refers to a nucleic acid sequence that can be used in combination with an RNA nuclease molecule and one or more gRNA molecules to modify (e.g., delete, disrupt, or modify) a target DNA sequence. In certain embodiments, the template nucleic acid results in a modification (e.g., deletion) to the CCAAT box target region of HBG1 and / or HBG2. In certain embodiments, the modification is a non-naturally occurring modification. In certain embodiments, non-naturally occurring modifications in the CCAAT box target region of HBG1 and / or HBG2 can include an 18nt target region, an 11nt target region, a 4nt target region, or a 1nt target region, or combinations thereof. In certain embodiments, the modification is a naturally occurring modification. In certain embodiments, naturally occurring modifications in the CCAAT box target region of HBG1 and / or HBG2 can include a 13nt target region, a c.-117G>A target region, or combinations thereof. In certain embodiments, the template nucleic acid is an ssODN. In certain embodiments, the ssODN is a plus strand or a minus strand.

[0259] For example, a template nucleic acid for introducing an 18-nt deletion into an 18-nt target region (HBG1 c.-104 to -121, HBG2 c.-104 to -121, or a combination thereof) may include a 5’ homology arm, a substitution sequence, and a 3’ homology arm, and the substitution sequence is 0 nucleotides or 0 bp. In certain embodiments, the 5’ homology arm can be, for example, about 25 to about 200 nucleotides or more in length, such as at least about 25, 50, 75, 100, 125, 150, 175, or 200 nucleotides. In certain embodiments, the 5’ homology arm includes homology of about 50 to 100 bp, such as 55 to 95, 60 to 90, 70 to 90, or 80 to 90 bp, on the 5’ side of the 18-nt target region. In certain embodiments, the 3’ homology arm can be, for example, about 25 to about 200 nucleotides or more in length, such as at least about 25, 50, 75, 100, 125, 150, 175, or 200 nucleotides. In certain embodiments, the 3’ homology arm includes homology of about 50 to 100 bp, such as 55 to 95, 60 to 90, 70 to 90, or 80 to 90 bp, on the 3’ side of the 18-nt target region. In certain embodiments, the 5’ and 3’ homology arms are symmetric in length. In certain embodiments, the 5’ and 3’ homology arms are asymmetric in length. In certain embodiments, the template nucleic acid is an ssODN. In certain embodiments, the ssODN is a plus strand. In certain embodiments, the ssODN is a minus strand. In certain embodiments, the ssODN comprises, consists essentially of, or consists of SEQ ID NO: 974 (OLI16409) or SEQ ID NO: 975 (OLI16410).

[0260] In certain embodiments, the template nucleic acid for introducing an 11-nt deletion into an 11-nt target region (HBG1 c.-105 to -115, HBG2 c.-105 to -115, or combinations thereof) may include a 5' homology arm, a substitution sequence, and a 3' homology arm, and the substitution sequence is 0 nucleotides or 0 bp. In certain embodiments, the 5' homology arm can be, for example, about 25 to about 200 nucleotides in length, such as at least about 25, 50, 75, 100, 125, 150, 175, or 200 nucleotides. In certain embodiments, the 5' homology arm includes homology of about 50 to 100 bp, such as 55 to 95, 60 to 90, 70 to 90, or 80 to 90 bp, on the 5' side of the 11-nt target region. In certain embodiments, the 3' homology arm can be, for example, about 25 to about 200 nucleotides in length, such as at least about 25, 50, 75, 100, 125, 150, 175, or 200 nucleotides. In certain embodiments, the 3' homology arm includes homology ...

Claims

1. 5' end and 3' end, a DNA extension comprising at its 5' end a sequence selected from the group consisting of SEQ ID NOs: 1236 to 1250; a targeting domain complementary to a target site in the promoter of the HBG gene; and an RNA segment capable of binding to Cpf1 RNA-guided nuclease or a Cpf1 variant; a single-molecule gRNA comprising: a nucleic acid encoding the Cpf1 RNA-guided nuclease or the Cpf1 variant; A genome editing system comprising:

2. The genome editing system of claim 1, wherein the gRNA comprises a 2'-O-methyl modification and a phosphorothioate modification at the 3' end.

3. The genome editing system of claim 1 or 2, wherein the gRNA comprises a 2'-fluoro modification.

4. The targeting domain is selected from the group consisting of SEQ ID NOs: 1002, 1004, 1139, 1141, 1143, 1145, 1147, 1149, 1151, 1153, 1155, 1157, 1159, 1161, 1163, 1165, 1167, 1169, 1171, 1173, 1175, 1177, 1179, 1181, 1183, 1185, 1187, 1189, 1191, 1193, 1195, 1197, 1199, 1201, 1203, 1205, 1207, 1209, 1211, 1213, 1215, 1217, 1219, 1221, 1223, 1225, 1227, 1229, 1254, 1256, 1258, 1260, 1262, or 1264. The genome editing system of any one of claims 1 to 3, comprising a sequence selected from the group consisting of:

5. 5. The genome editing system of claim 1, wherein the target site comprises nucleotides located at Chr11 (NC_000011.10) 5,249,904 to 5,249,927 (Table 17, Region 6); Chr11 (NC_000011.10) 5,254,879 to 5,254,909 (Table 17, Region 16); or a combination thereof.

6. The genome editing system of any one of claims 1 to 5, wherein the Cpf1 variant comprises one or more modifications selected from the group consisting of one or more mutations of a wild-type Cpf1 amino acid sequence, one or more nuclear localization signals, one or more purification tags, and combinations thereof.

7. The genome editing system of any one of claims 1 to 6, wherein the Cpf1 variant comprises a sequence selected from the group consisting of SEQ ID NOs: 1000, 1001, 1008-1015 and 1035-1039.

8. The genome editing system of any one of claims 1 to 7, which is used in a method for modifying a promoter of an HBG gene in a cell, comprising contacting the cell with the gRNA and the nucleic acid.

9. The genome editing system of claim 8, wherein the cells are CD34+ cells or hematopoietic stem cells.

10. 10. The genome editing system of claim 8 or 9, wherein the gRNA and the nucleic acid are delivered to the cell using electroporation or lipid nanoparticles.

11. The genome editing system of any one of claims 8 to 10, wherein the nucleic acid comprises messenger RNA.

12. A genome editing system described in any one of claims 8 to 11 for use in alleviating one or more symptoms of sickle cell disease in a subject in need thereof.

13. A composition comprising the genome editing system of any one of claims 8 to 12, further comprising a pharmaceutically acceptable carrier.

14. A genome editing system according to any one of claims 1 to 7, and Pharmaceutically acceptable carrier Kit including: