Systems and methods for the treatment of abnormal hemoglobin disorders

JP7920328B2Active Publication Date: 2026-09-14EDITAS MEDICINE INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2025006855
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2018-11-29
Filing Date
2025-01-17
Publication Date
2026-09-14
Estimated Expiration
2039-03-14

Smart Images

  • Figure 0007920328000049
    Figure 0007920328000049
  • Figure 0007920328000050
    Figure 0007920328000050
  • Figure 0007920328000051
    Figure 0007920328000051
Patent Text Reader

Abstract

To provide genome editing systems and methods for altering a target nucleic acid sequence, or modulating expression of a target nucleic acid sequence, and methods for the alteration of genes encoding hemoglobin subunits and / or treatment of hemoglobinopathies.SOLUTION: The present invention provides an RNP complex comprising a CRISPR from Prevotella and Francisella1 (Cpf1) RNA guided nuclease or a variant thereof and a gRNA, wherein the gRNA is capable of binding to a target site in a promoter of an HBG gene in a cell.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Reference to Related Application This application claims the benefit of U.S. Provisional Patent Application No. 62 / 643,168 filed on March 14, 2018, U.S. Provisional Patent Application No. 62 / 767,488 filed on November 14, 2018, and U.S. Provisional Patent Application No. 62 / 773,073 filed on November 29, 2018, the entire contents of which are incorporated herein by reference.

[0002] Sequence Listing This application contains a Sequence Listing submitted in ASCII format via EFS-Web, the entire content of which is incorporated herein by reference. The ASCII copy created on March 14, 2019 is named 2019-03-14_EM115PCT_Editas_8013WO02_Sequence_Listing.txt and is 811 KB in size.

[0003] The present disclosure relates to genome editing systems and methods for modifying a target nucleic acid sequence or modulating the expression of a target nucleic acid sequence, as well as uses thereof associated with the modification of genes encoding hemoglobin subunits and / or the treatment of hemoglobinopathies. [Background Art]

[0004] Hemoglobin (Hb) carries oxygen from red blood cells (RBCs) from the lungs to the tissues. During prenatal development and immediately after birth, hemoglobin exists in the form of fetal hemoglobin (HbF), a tetrameric protein composed of two α-globin chains and two γ-globin chains. Through a process known as globin switching, HbF is largely replaced by adult hemoglobin (HbA), a tetrameric protein in which the γ-globin chains of HbF are replaced by β-globin chains. The average adult produces less than 1% HbF from total hemoglobin (Thein 2009). The α-hemoglobin gene is located on chromosome 16, while the β-hemoglobin gene (HBB), Aγ-globin chain (HBG1, also known as γ-globin A), and Gγ-globin chain (HBG2, also known as γ-globin G) are located on chromosome 11 within the globin gene cluster (also called the globin locus).

[0005] Mutations in the hemoglobin barium (HBB) can cause hemoglobin disorders (i.e., abnormal hemoglobin disorders), including sickle cell disease (SCD) and beta-thalassemia (β-thal). In the United States, approximately 93,000 people are diagnosed with abnormal hemoglobin disorders. Globally, 300,000 children are born with abnormal hemoglobin disorders each year (Angastiniotis 1998). Because these conditions are associated with HBB mutations, these symptoms typically do not manifest until the globin changes from HbF to HbA.

[0006] Sickle cell disease (SCD) is the most common genetic blood disorder in the United States, affecting approximately 80,000 people (Brousseau 2010). SCD is most common among people of African descent, with a prevalence of 1 in 500. In Africa, the prevalence of SCD is 15 million (Aliyu 2008). SCD is also more common among people of Indian, Saudi Arabian, and Mediterranean descent. Among Hispanic Americans, the prevalence of sickle cell disease is 1 in 1,000 (Lewis 2014).

[0007] SCD is caused by a single homozygous mutation in the c.17A>T (HbS mutation) of the HBB gene. The sickle cell mutation is a point mutation (GAG>GTG) on the HBB, which results in the substitution of glutamic acid with valine at amino acid position 6 of exon 1. The valine at position 6 of the β-hemoglobin chain is hydrophobic and, when not bound to oxygen, causes a conformational change in the β-globin protein. This conformational change causes polymerization of the HbS protein in the absence of oxygen, resulting in deformation of the RBC (i.e., sickle cell formation). SCD is inherited in an autosomal recessive manner, so only patients with two HbS alleles will have this disease. Heterozygous individuals have the sickle cell phenotype and may suffer from anemia and / or painful onset when subjected to severe dehydration or oxygen deficiency.

[0008] Sickle cell red blood cells (RBCs) cause multiple symptoms, including anemia, sickle cell crisis, vascular occlusion, aplastic syndrome, and acute chest syndrome. Because sickle RBCs are less elastic than wild-type red blood cells, they cannot easily pass through the capillary bed, leading to occlusion and ischemia (i.e., vascular occlusion). Vascular occlusion occurs when sickle cells obstruct blood flow in the capillary bed of an organ, resulting in pain, ischemia, and necrosis. These episodes typically last 5–7 days. The spleen plays a role in removing dysfunctional RBCs and is therefore typically enlarged in infancy, frequently suffering from vascular occlusion. By the end of childhood, the spleen of SCD patients often becomes infarcted, leading to splenic ligation. Hemolysis is a constant feature of SCD and results in anemia. Sickle cells survive in circulation for 10–20 days, while healthy red blood cells survive for 90–120 days. SCD patients receive transfusions as needed to maintain adequate hemoglobin levels. Frequent blood transfusions expose patients to the risk of HIV, hepatitis B, and hepatitis C infection. Patients may also experience acute chest cramps and infarctions of the limbs, peripheral organs, and central nervous system.

[0009] Individuals with sickle cell disease (SCD) have a reduced life expectancy. The prognosis for patients with SCD has steadily improved with careful lifelong management of the disease and anemia. As of 2001, the average life expectancy for individuals with sickle cell disease was in their mid-to-late 50s. Current treatment for SCD involves hydration and pain management during the disease, as well as blood transfusions to correct anemia as needed.

[0010] Thalassemia (e.g., β-Thal, δ-Thal, and β / δ-Thal) causes chronic anemia. β-Thal is estimated to affect approximately 1 in 100,000 people worldwide. Its prevalence is higher in certain populations, including those of European descent, where it affects approximately 1 in 10,000 people. β-Thal major is a more severe form of the disease and is life-threatening unless treated with lifelong transfusions and chelation therapy. In the United States, there are approximately 3,000 people eligible for β-Thal major. β-Thal intermediate does not require transfusions, but it can cause growth retardation and significant systemic abnormalities, often requiring lifelong chelation therapy. HbA makes up the majority of hemoglobin in adult red blood cells, but about 3% of adult hemoglobin is in the form of HbA2, an HbA variant in which two gamma-globin chains are replaced by two delta(Δ)-globin chains. δ-thal is associated with mutations in the Δ-hemoglobin gene (HBD) that result in loss of HBD expression. Co-inheritance of HBD mutations can mask the diagnosis of β-thal (i.e., β / δ-thal) by lowering HbA2 levels to the normal range (Bouva 2006). β / δ-thal is usually caused by deletions of HBB and HBD sequences in both alleles. In homozygous (δo / δo βo / βo) patients, HBG is expressed, resulting in the production of HbF only.

[0011] Similar to SCD, β-thal is caused by mutations in the HBB gene. The most common HBB mutations resulting in β-thal are c.-136C>G, c.92+1G>A, c.92+6T>C, c.93-21G>A, c.118C>T, c.316-106C>G, c.25_26delAA, c.27_28insG, c.92+5G>C, c.118C>T, c.135delC, c.315+1G>A, c.-78A>G, c. These include 52A>T, c.59A>G, c.92+5G>C, c.124_127delTTCT, c.316-197C>T, c.-78A>G, c.52A>T, c.124_127delTTCT, c.316-197C>T, c.-138C>T, c.-79A>G, c.92+5G>C, c.75T>A, c.316-2A>G, and c.316-2A>C. These and other β-thal-related mutations cause mutations or deletions in the β-globin chain, which disrupts the normal Hbα-hemoglobin to β-hemoglobin ratio. The excess α-globin chain precipitates in the erythrocyte precursors of the bone marrow.

[0012] In the β-thal major, both alleles of the HBB contain nonsense, frameshift, or splicing mutations that result in a complete absence of β-globin production (denoted as β0 / β0). The β-thal major leads to a severe reduction in β-globin chains, resulting in marked precipitation of α-globin chains in the RBCs and high-severity anemia.

[0013] The β-thal intermediate type arises from mutations in the 5' or 3' untranslated region of the HBB, mutations in the HBB promoter region or polyadenylation signaling pathway, or splicing mutations within the HBB gene. A patient's genotype is denoted as βo / β+ or β+ / β+. βo represents the absence of β-globin chain expression; β+ represents a dysfunctional but present β-globin chain. Phenotypic expression varies among patients. Because the β-thal intermediate type still produces some β-globin, it results in less α-globin chain precipitation in erythrocyte precursors and less severe anemia than the β-thal major type. However, it has more significant consequences for erythroid proliferation secondary to chronic anemia.

[0014] Patients with β-thal major develop symptoms between 6 months and 2 years of age, suffering from growth retardation, fever, hepatosplenomegaly, and diarrhea. Appropriate treatment includes regular blood transfusions. Splenectomy and hydroxyurea therapy are also options for treating β-thal major. Patients who receive regular transfusions develop normally until their early 20s. At that point, patients require chelation therapy (in addition to continuous transfusions) to prevent complications of iron overload. Iron overload can manifest as delayed growth or delayed sexual maturation. In adulthood, inadequate chelation therapy can lead to cardiomyopathy, cardiac arrhythmias, hepatic fibrosis and / or cirrhosis, diabetes, thyroid and parathyroid abnormalities, thrombosis, and osteoporosis. Frequent transfusions also expose patients to the risk of HIV, hepatitis B, and hepatitis C infection.

[0015] Intermediate β-thal anemia typically affects children between 2 and 6 years of age. Generally, these children do not require blood transfusions. However, bone abnormalities occur due to chronic hypertrophy of the red blood cell lineage to compensate for chronic anemia. They may also experience fractures of long bones due to osteoporosis. Extramedullary erythrocyte formation is common, leading to swelling of the spleen, liver, and lymph nodes. This can also cause spinal cord compression and neurological problems. Children are prone to lower extremity ulcers and have an increased risk of thrombotic events, including stroke, pulmonary embolism, and deep vein thrombosis. Treatment options for intermediate β-thal anemia include splenectomy, folic acid supplementation, hydroxyurea therapy, and radiation therapy for extramedullary tumors. Chelation therapy is used for patients who develop iron overload.

[0016] Patients with β-thal syndrome often have a reduced life expectancy. Patients with β-thal majors who are not receiving blood transfusions generally die in their 20s or 30s. Patients with β-thal majors who receive regular blood transfusions and appropriate chelation therapy may survive beyond their 50s. Heart failure secondary to iron toxicity is a major cause of death in patients with β-thal majors due to iron toxicity. [Overview of the project] [Problems that the invention aims to solve]

[0017] Currently, various novel therapies for SCD and β-Thal are being developed. Currently, the delivery of anti-sickle hemoglobin genes via gene therapy is being investigated in clinical trials. However, the long-term efficacy and safety of this approach remain unclear. While hematopoietic stem cell (HSC) transplantation from HLA-matched allogeneic stem cell donors has been shown to cure SCD and β-Thal, this procedure carries risks, including those associated with the ablation therapy required to prepare the subject for transplantation, and increases the risk of life-threatening opportunistic infections and post-transplant graft-versus-host disease. Furthermore, finding a suitable allogeneic donor is often difficult. Therefore, there is a need for improved methods to manage these and other hemoglobin disorders. [Means for solving the problem]

[0018] This specification provides a genome editing system, a ribonucleoprotein (RNP) complex, a guide RNA, a modified Cpf1 protein (Cpf1 mutant), and a CRISPR-mediated method for modifying the promoter region of one or more γ-globin genes (e.g., HBG1, HBG2, or HBG1 and HBG2) to increase the expression of fetal hemoglobin (HbF). In certain embodiments, the RNP complex may comprise a guide RNA (gRNA) complexed with wild-type Cpf1 or modified Cpf1 RNA-induced nuclease (modified Cpf1 protein). In certain embodiments, the gRNA may comprise sequences listed in Table 13, Table 18, or Table 19. In certain embodiments, the gRNA may comprise a gRNA targeting domain. In certain embodiments, the gRNA targeting domain may comprise a sequence selected from the group consisting of SEQ ID NOs: 1002, 1254, 1258, 1260, 1262, and 1264. In certain embodiments, the gRNA may include gRNA sequences listed in Table 19. In certain embodiments, the gRNA may include sequences selected from the group consisting of SEQ ID NOs: 1022, 1023, and 1041-1105. In certain embodiments, the RNP complex may include RNP complexes listed in Table 21. For example, the RNP complex may include a gRNA containing the sequence described in SEQ ID NO: 1051 and a modified Cpf1 protein encoded by the sequence described in SEQ ID NO: 1097 (RNP32 in Table 21).

[0019] The inventors hereby discover that delivery of an RNP complex containing gRNA complexed with a modified Cpf1 protein can result in increased editing of the target nucleic acid. In certain embodiments, the modified Cpf1 protein may contain one or more modifications. In certain embodiments, one or more modifications may include, but are not limited to, one or more mutations in the wild-type Cpf1 amino acid sequence, one or more mutations in the wild-type Cpf1 nucleic acid sequence, one or more mutations in the nuclear localization signal (NLS), one or more mutations in the purification tag (e.g., His tag), or a combination thereof. In certain embodiments, the modified Cpf1 may be encoded by sequences described in SEQ ID NOs: 1000, 1001, 1008-1018, 1032, 1035-39, 1094-1097, 1107-09 (Cpf1 polypeptide sequences) or SEQ ID NOs: 1019-1021, 1110-17 (Cpf1 polynucleotide sequences). In certain embodiments, the gRNA may be modified or unmodified. In certain embodiments, the gRNA may contain sequences listed in Table 13, Table 18, or Table 19. In certain embodiments, the RNP complex may contain RNP complexes listed in Table 21. For example, the RNP complex may contain a gRNA containing the sequence listed in SEQ ID NO: 1051 and a modified Cpf1 protein encoded by the sequence listed in SEQ ID NO: 1097 (RNP32 in Table 21). In certain embodiments, the RNP complex containing the modified Cpf1 protein may increase the editing of the target nucleic acid. In certain embodiments, the RNP complex containing the modified Cpf1 protein may increase the editing, resulting in an increase in productive indels. In various embodiments, the increase in editing of the target nucleic acid may be evaluated by PCR amplification of the target nucleic acid and subsequent sequence analysis (e.g., Sanger sequencing, next-generation sequencing), but not limited to these, by any means known to those skilled in the art.

[0020] The inventors have also found herein that delivery of an RNP complex containing modified gRNA complexed with an unmodified or modified Cpf1 protein may result in increased editing of the target nucleic acid. In certain embodiments, the modified gRNA may include one or more modifications, including phosphorothioate binding modifications, phosphorodithioate (PS2) binding modifications, and 2'-O-methyl modifications; one or more or a sequence of deoxyribonucleic acid (DNA) bases (also referred herein as "DNA extensions"); one or more or a sequence of ribonucleic acid (RNA) bases (also referred herein as "RNA extensions"); or a combination thereof. In certain embodiments, the DNA extensions may include sequences listed in Table 24. For example, in certain embodiments, the DNA extensions may include sequences listed in SEQ ID NOs: 1235-1250. In certain embodiments, the RNA extensions may include sequences listed in Table 24. For example, in certain embodiments, the RNA extensions may include sequences listed in SEQ ID NOs: 1231-1234, 1251-1253. In certain embodiments, the gRNA may include sequences listed in Table 13, Table 18, or Table 19. In certain embodiments, the RNP complex may include RNP complexes listed in Table 21. For example, the RNP complex may include a gRNA containing the sequence listed in SEQ ID NO: 1051 and a modified Cpf1 protein encoded by the sequence listed in SEQ ID NO: 1097 (RNP32 in Table 21). In certain embodiments, the RNP complex containing the modified gRNA may increase the editing of the target nucleic acid. In certain embodiments, the RNP complex containing the modified gRNA may increase editing, resulting in an increase in productive indels.

[0021] In certain embodiments, an RNP complex comprising modified gRNA and modified Cpf1 protein may increase the editing of a target nucleic acid. In certain embodiments, an RNP complex comprising modified gRNA and modified Cpf1 protein may increase editing, resulting in an increase in productive indels.

[0022] The inventors have also discovered that the co-delivery of an RNP complex containing gRNA complexed with a Cpf1 molecule (e.g., “gRNA-Cpf1-RNP”) and a “booster element” may result in increased editing of the target nucleic acid. In certain embodiments, the RNP complex may include the RNP complexes listed in Table 21. For example, the RNP complex may include a gRNA containing the sequence described in SEQ ID NO: 1051 and a modified Cpf1 protein encoded by the sequence described in SEQ ID NO: 1097 (RNP32 in Table 21). As used herein, the term “booster element” refers to an element that, when co-delivered with an RNP complex containing gRNA complexed with an RNA-induced nuclease (“gRNA-nuclease-RNP”), increases the editing of the target nucleic acid compared to editing of the target nucleic acid without the booster element. In certain embodiments, one or more booster elements may be co-delivered with a gRNA-nuclease-RNP complex to increase the editing of the target nucleic acid. In certain embodiments, simultaneous delivery of booster elements can increase editing and result in an increase in productive indels. In various embodiments, the increase in editing of the target nucleic acid can be assessed by PCR amplification of the target nucleic acid and subsequent sequencing analysis (e.g., Sanger sequencing, next-generation sequencing), but not limited to these, and any other means known to those skilled in the art.

[0023] In certain embodiments, the gRNA-nuclease-RNP may include gRNA-Cpf1-RNP. In certain embodiments, the Cpf1 molecule of the gRNA-Cpf1-RNP complex may be wild-type Cpf1 or modified Cpf1. In certain embodiments, the Cpf1 molecule of gRNA-Cpf1-RNP may be encoded by sequences described in SEQ ID NOs: 1000, 1001, 1008-1018, 1032, 1035-39, 1094-1097, 1107-09 (Cpf1 polypeptide sequences) or SEQ ID NOs: 1019-1021, 1110-17 (Cpf1 polynucleotide sequences). In certain embodiments, the gRNA-Cpf1-RNP complex may include gRNAs containing targeting domains described in Table 13 or Table 18. In certain embodiments, the gRNA-Cpf1-RNP complex may include gRNA containing the sequences listed in Table 19. In certain embodiments, the gRNA may be modified or unmodified gRNA.

[0024] In certain embodiments, the booster element may include a dead gRNA-nuclease-RNP ("dead gRNA-nuclease-RNP") which contains a dead gRNA molecule complexed with an RNA-inducible nuclease molecule. In certain embodiments, the dead gRNA-nuclease-RNP may include a dead gRNA complexed with a wild-type (WT) Cas9 molecule ("dead gRNA-Cas9-RNP"), a dead gRNA complexed with a Cas9 nickase molecule ("dead gRNA-nickase-RNP"), or a dead gRNA complexed with an enzymatically inactive (ei)Cas9 molecule ("dead gRNA-eiCas9-RNP"). In certain embodiments, the dead gRNA-nuclease-RNP complex may have reduced or no nuclease activity. In certain embodiments, the dead gRNA of the dead gRNA-nuclease-RNP complex may include any of the dead gRNAs described herein. For example, a dead gRNA that may contain a targeting domain may be identical to or differ by three nucleotides or less from the dead gRNA targeting domains listed in Table 10 or Table 15. In certain embodiments, the dead gRNA may contain a targeting domain that includes a truncated gRNA targeting domain. In certain embodiments, the truncated gRNA targeting domain may be one of the gRNA targeting domains listed in Table 2, Table 10, or Table 15. In certain embodiments, the dead gRNA may be modified or unmodified dead gRNA. As described herein, co-delivery of gRNA-Cpf1-RNP and dead gRNA-Cas9-RNP (i.e., an RNP containing dead gRNA complexed with WT Cas9) or co-delivery of gRNA-Cpf1-RNP and dead gRNA-nickase-RNP (i.e., an RNP containing dead gRNA complexed with Cas9 nickase (i.e., Cas9 D10A nickase)) resulted in an increase in total editing beyond the levels observed following delivery of gRNA-Cpf1-RNP alone (see, e.g., Examples 15, 17, and 18). The dead gRNA molecule may contain a targeting domain complementary to a region proximal to or within a target region in the target nucleic acid (e.g., a CCAAT box target region, a 13nt target region, a proximal HBG1 / 2 promoter target sequence, and / or a GATA1 binding motif in BCL11Ae).In certain embodiments, “proximal” may refer to a region of 10, 25, 50, 100, or 200 nucleotides or less within the target region (e.g., the CCAAT box target region, the 13nt target region, the proximal HBG1 / 2 promoter target sequence, and / or the GATA1 binding motif in BCL11Ae). In certain embodiments, one or more booster elements may include one or more dead gRNA-nuclease-RNPs, such as dead gRNA-Cas9-RNP, dead gRNA-nickase-RNP, and dead gRNA-eiCas9-RNP, which are co-delivered with gRNA-Cpf1-RNP. In certain embodiments, co-delivery of dead gRNA-nuclease-RNPs does not alter the indel profile of gRNA-Cpf1-RNP.

[0025] In certain embodiments, the booster element may comprise an RNP complex ("gRNA-nickase-RNP") comprising a gRNA molecule complexed with an RNA-guided nuclease nickase molecule. In certain embodiments, the RNA-guided nuclease nickase molecule can be, for example, a Cas9 nickase molecule such as Cas9 D10A nickase. In certain embodiments, the gRNA of the gRNA-nickase-RNP may comprise any of the gRNAs described herein. For example, the gRNA may comprise a gRNA targeting domain set forth in Table 2, Table 10 or Table 15. In certain embodiments, the gRNA can be a modified or unmodified gRNA. As shown herein, co-delivery of a gRNA-Cpf1-RNP and a gRNA-nickase-RNP complex (an RNP comprising a guide RNA complexed to a Cas9 D10A nickase molecule) resulted in an increase in total editing over the level observed following delivery of gRNA-Cpf1-RNP alone (see, e.g., Examples 15 and 16). Furthermore, co-delivery of a gRNA-nickase-RNP complex and a gRNA-Cpf1-RNP complex altered the orientation, length and / or position of the indel profile of gRNA-Cpf1-RNP. In certain embodiments, a booster enhancer may be used to provide a desired editing outcome, for example, an increase in the proportion of productive indels. In certain embodiments, co-delivery of a gRNA-nickase-RNP complex and a gRNA-Cpf1-RNP complex can alter the indel profile of gRNA-Cpf1-RNP.

[0026] In certain embodiments, the booster element may comprise a single-stranded oligodeoxynucleotide (ssODN) or a double-stranded oligodeoxynucleotide (dsODN). In certain embodiments, the ssODN can be any ssODN disclosed herein. In certain embodiments, the ssODN may comprise a sequence set forth in Table 11. For example, in certain embodiments, the ssODN may comprise the sequence set forth in SEQ ID NO: 1040.

[0027] In one embodiment, the disclosure relates to an RNP complex comprising CRISPR from Prevotella and Francisella 1 (Cpf1) RNA-induced nucleases or mutants thereof, and a gRNA capable of binding to a target site in the promoter of the intracellular HBG gene. In certain embodiments, the gRNA may be modified or unmodified. In certain embodiments, the gRNA may include one or more modifications, including phosphorothioate binding modifications, phosphorodithioate (PS2) binding modifications, 2'-O-methyl modifications, DNA extensions, RNA extensions, or combinations thereof. In certain embodiments, the DNA extensions may include sequences listed in Table 24. In certain embodiments, the RNA extensions may include sequences listed in Table 24. In certain embodiments, the gRNA may include sequences listed in Table 13, Table 18, or Table 19. In certain embodiments, the RNP complex may include the RNP complexes listed in Table 21. For example, the RNP complex may include a gRNA containing the sequence described in SEQ ID NO: 1051 and a Cpf1 mutant protein encoded by the sequence described in SEQ ID NO: 1097 (RNP32 in Table 21). In certain embodiments, the Cpf1 mutant protein may contain one or more modifications. In certain embodiments, one or more modifications may include, but are not limited to, one or more mutations in the wild-type Cpf1 amino acid sequence, one or more wild-type Cpf1 nucleic acid sequences, one or more nuclear localization signals (NLS), one or more mutations in purification tags (e.g., His tags), or a combination thereof. In certain embodiments, the Cpf1 mutant protein may be encoded by sequences described in SEQ ID NOs: 1000, 1001, 1008-1018, 1032, 1035-39, 1094-1097, 1107-09 (Cpf1 polypeptide sequences) or SEQ ID NOs: 1019-1021, 1110-17 (Cpf1 polynucleotide sequences).

[0028] In one embodiment, the present disclosure relates to a method for modifying the promoter of the HBG gene in a cell, comprising the step of contacting the cell with the RNP complex disclosed herein. In certain embodiments, the modification may include indels in one or more regions listed in Table 17. In certain embodiments, the modification may include indels in the CCAAT box target region of the HBG gene promoter. For example, in certain embodiments, the modification may include indels in Chr11(NC_000011.10):5,249,955~5,249,987 (Table 17, region 6), Chr11(NC_000011.10):5,254,879~5,254,909 (Table 17, region 16), or a combination thereof. In certain embodiments, the RNP complex may include gRNA and Cpf1 protein. In certain embodiments, the gRNA may include RNA targeting domains listed in Table 19. In certain embodiments, the gRNA targeting domain may include a sequence selected from the group consisting of SEQ ID NOs: 1002, 1254, 1258, 1260, 1262, and 1264. In certain embodiments, the gRNA may include gRNA sequences listed in Table 19. In certain embodiments, the gRNA may include a sequence selected from the group consisting of SEQ ID NOs: 1022, 1023, and 1041-1105. In certain embodiments, the gRNA may include Chr11:5249973, Chr11:5249977(HBG1); Chr11:5250042, Chr11:5250046(HBG1); Chr11:5250055, Chr11:5250059(HBG1); Chr11:5250179, Chr11:5250183(HBG1); C It may be configured to provide editing events in hr11:5254897, Chr11:5254901(HBG2); Chr11:5254897, Chr11:5254901(HBG2); Chr11:5254966, 5254970(HBG2); Chr11:5254979, 5254983(HBG2) (Tables 22 and 23). In certain embodiments, the cell may be further in contact with a booster element. In certain embodiments, the booster element may include single-stranded oligodeoxynucleotides (ssODNs) or double-stranded oligodeoxynucleotides (dsODNs).In certain embodiments, the ssODN can be any ssODN disclosed herein. In certain embodiments, the ssODN can comprise the sequence set forth in Table 11. For example, in certain embodiments, the ssODN can comprise the sequence set forth in SEQ ID NO: 1040.

[0029] In one embodiment, the disclosure relates to an isolated cell comprising a modification in the promoter of the HBG gene, which is generated by delivery of an RNP complex to the cell. In certain embodiments, the RNP complex may comprise a gRNA and a Cpf1 protein. In certain embodiments, the gRNA may be modified or unmodified. In certain embodiments, the gRNA may comprise one or more modifications, including phosphorothioate binding modification, phosphorodithioate (PS2) binding modification, 2'-O-methyl modification, DNA extension, RNA extension, or a combination thereof. In certain embodiments, the DNA extension may comprise the sequences listed in Table 24. In certain embodiments, the RNA extension may comprise the sequences listed in Table 24. In certain embodiments, the gRNA may comprise the sequences listed in Table 13, Table 18, or Table 19. In certain embodiments, the RNP complex may comprise the RNP complexes listed in Table 21. For example, the RNP complex may include a gRNA containing the sequence described in SEQ ID NO: 1051 and a Cpf1 mutant protein encoded by the sequence described in SEQ ID NO: 1097 (RNP32 in Table 21). In certain embodiments, the Cpf1 mutant protein may contain one or more modifications. In certain embodiments, one or more modifications may include, but are not limited to, one or more mutations in the wild-type Cpf1 amino acid sequence, one or more mutations in the wild-type Cpf1 nucleic acid sequence, one or more nuclear localization signals (NLS), one or more mutations in purification tags (e.g., His tags), or a combination thereof. In certain embodiments, the Cpf1 mutant protein may be encoded by sequences described in SEQ ID NOs: 1000, 1001, 1008-1018, 1032, 1035-39, 1094-1097, 1107-09 (Cpf1 polypeptide sequences) or SEQ ID NOs: 1019-1021, 1110-17 (Cpf1 polynucleotide sequences). In certain embodiments, the booster element may be delivered co-delivered with the RNP complex. In certain embodiments, the booster element may comprise a single-stranded oligodeoxynucleotide (ssODN) or a double-stranded oligodeoxynucleotide (dsODN). In certain embodiments, the ssODN may be any ssODN disclosed herein.In certain embodiments, ssODN may include sequences listed in Table 11. For example, in certain embodiments, ssODN may include the sequence listed in Sequence ID No. 1040.

[0030] In one embodiment, the disclosure relates to an in vitro method for increasing the level of fetal hemoglobin (HbF) in human cells by genome editing using a gRNA and an RNP complex comprising a Cpf1 RNA-induced nuclease or a variant thereof, thereby influencing modifications in the promoter of the HBG gene and thereby increasing HbF expression. In certain embodiments, the gRNA may be modified or unmodified. In certain embodiments, the gRNA may include one or more modifications, including phosphorothioate binding modifications, phosphorodithioate (PS2) binding modifications, 2'-O-methyl modifications, DNA extensions, RNA extensions, or combinations thereof. In certain embodiments, the DNA extensions may include sequences listed in Table 24. In certain embodiments, the RNA extensions may include sequences listed in Table 24. In certain embodiments, the gRNA may include sequences listed in Table 13, Table 18, or Table 19. In certain embodiments, the RNP complex may include RNP complexes listed in Table 21. For example, the RNP complex may include a gRNA containing the sequence described in SEQ ID NO: 1051 and a Cpf1 mutant protein encoded by the sequence described in SEQ ID NO: 1097 (RNP32 in Table 21). In certain embodiments, the Cpf1 mutant protein may contain one or more modifications. In certain embodiments, one or more modifications may include, but are not limited to, one or more mutations in the wild-type Cpf1 amino acid sequence, one or more mutations in the wild-type Cpf1 nucleic acid sequence, one or more nuclear localization signals (NLS), one or more mutations in purification tags (e.g., His tags), or a combination thereof. In certain embodiments, the Cpf1 mutant protein may be encoded by sequences described in SEQ ID NOs: 1000, 1001, 1008-1018, 1032, 1035-39, 1094-1097, 1107-09 (Cpf1 polypeptide sequences) or SEQ ID NOs: 1019-1021, 1110-17 (Cpf1 polynucleotide sequences). In certain embodiments, the booster element may be delivered co-delivered with the RNP complex. In certain embodiments, the booster element may comprise single-stranded oligodeoxynucleotides (ssODNs) or double-stranded oligodeoxynucleotides (dsODNs).In certain embodiments, ssODN may be any ssODN disclosed herein. In certain embodiments, ssODN may include sequences listed in Table 11. For example, in certain embodiments, ssODN may include sequences listed in Sequence ID No. 1040.

[0031] In one embodiment, the disclosure relates to a population of CD34+ or hematopoietic stem cells, wherein one or more cells in the population include a modification in the promoter of the HBG gene, the modification being produced by delivering an RNP complex comprising gRNA and a Cpf1 RNA-induced nuclease or a variant thereof to the population of CD34+ or hematopoietic stem cells. In certain embodiments, the gRNA may be modified or unmodified. In certain embodiments, the gRNA may include one or more modifications, including phosphorothioate binding modification, phosphorodithioate (PS2) binding modification, 2'-O-methyl modification, DNA extension, RNA extension, or a combination thereof. In certain embodiments, the DNA extension may include sequences listed in Table 24. In certain embodiments, the RNA extension may include sequences listed in Table 24. In certain embodiments, the gRNA may include sequences listed in Table 13, Table 18, or Table 19. In certain embodiments, the RNP complex may include RNP complexes listed in Table 21. For example, the RNP complex may include a gRNA containing the sequence described in SEQ ID NO: 1051 and a Cpf1 mutant protein encoded by the sequence described in SEQ ID NO: 1097 (RNP32 in Table 21). In certain embodiments, the Cpf1 mutant protein may contain one or more modifications. In certain embodiments, one or more modifications may include, but are not limited to, one or more mutations in the wild-type Cpf1 amino acid sequence, one or more mutations in the wild-type Cpf1 nucleic acid sequence, one or more nuclear localization signals (NLS), one or more mutations in purification tags (e.g., His tags), or a combination thereof. In certain embodiments, the Cpf1 mutant protein may be encoded by sequences described in SEQ ID NOs: 1000, 1001, 1008-1018, 1032, 1035-39, 1094-1097, 1107-09 (Cpf1 polypeptide sequences) or SEQ ID NOs: 1019-1021, 1110-17 (Cpf1 polynucleotide sequences). In certain embodiments, the booster element may be delivered co-delivered with the RNP complex.In certain embodiments, the booster element may include a single-stranded oligodeoxynucleotide (ssODN) or a double-stranded oligodeoxynucleotide (dsODN). In certain embodiments, the ssODN may be any ssODN disclosed herein. In certain embodiments, the ssODN may include the sequences listed in Table 11. For example, in certain embodiments, the ssODN may include the sequence listed in Sequence ID No. 1040.

[0032] In one embodiment, the present disclosure relates to a method for alleviating one or more symptoms of sickle cell disease in a subject in need, comprising: a) isolating a population of CD34+ or hematopoietic stem cells from a subject; b) modifying the isolated population of cells in vitro by delivering an RNP complex comprising gRNA and Cpf1 RNA-induced nuclease or a variant thereof to the isolated population, thereby influencing a modification of the HBG gene promoter in one or more cells of the population; and c) administering the modified population of cells to a subject, thereby alleviating one or more symptoms of sickle cell disease in the subject. In certain embodiments, the modification may include an indel within the CCAAT box target region of the HBG gene promoter. In certain embodiments, the RNP complex may be delivered by electroporation. In certain embodiments, at least about 5%, at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, or at least about 90% of the cells in a population of cells contain productive indels.

[0033] In one embodiment, the disclosure relates to a gRNA comprising a 5' end and a 3' end, and comprising a DNA extension at the 5' end and a 2'-O-methyl-3'-phosphorothioate modification at the 3' end, comprising an RNA segment capable of hybridizing to a target site and an RNA segment capable of binding to a Cpf1 RNA-induced nuclease. In certain embodiments, the DNA extension may comprise the sequences described in SEQ ID NOs: 1235-1250. In certain embodiments, the gRNA may be modified or unmodified. In certain embodiments, the gRNA may comprise one or more modifications, including phosphorothioate binding modifications, phosphorodithioate (PS2) binding modifications, 2'-O-methyl modifications, DNA extensions, RNA extensions, or combinations thereof. In certain embodiments, the DNA extension may comprise the sequences described in Table 24. In certain embodiments, the RNA extension may comprise the sequences described in Table 24. In certain embodiments, the gRNA may comprise the sequences described in Table 13, Table 18, or Table 19.

[0034] In one embodiment, the present disclosure relates to an RNP complex comprising a Cpf1 RNA-induced nuclease as disclosed herein and a gRNA as disclosed herein.

[0035] This specification also provides genome editing systems, guide RNAs, and CRISPR-mediated methods for modifying one or more γ-globin genes (e.g., HBG1, HBG2, or HBG1 and HBG2), the erythrocyte-specific enhancer of the BCL11A gene (BCL11Ae), or a combination thereof, and for increasing the expression of fetal hemoglobin (HbF). In certain embodiments, one or more gRNAs containing sequences listed in Table 18 or Table 19 may be used to introduce modifications to the promoter region of the HBG gene. In certain embodiments, the genome editing systems, guide RNAs, and CRISPR-mediated methods may modify the 13-nucleotide (nt) target region ("13nt target region") located at the 5' end of the transcription site of the HBG1, HBG2, or HBG1 and HBG2 gene. In certain embodiments, the genome editing system, guide RNA, and CRISPR-mediated method may modify the CCAAT box target region ("CCAAT box target region") which is located at the 5' end of the transcription site of the HBG1, HBG2, or HBG1 and HBG2 genes. In certain embodiments, the CCAAT box target region may be the region of the distal CCAAT box or a region adjacent thereto, and include the nucleotides of the distal CCAAT box and 25 nucleotides upstream (5') and downstream (3') of the distal CCAAT box (i.e., HBG1 / 2 c.-86 to -140). In certain embodiments, the CCAAT box target region may be the region of the distal CCAAT box or a region adjacent thereto, and include the nucleotides of the distal CCAAT box and 5 nucleotides upstream (5') and downstream (3') of the distal CCAAT box (i.e., HBG1 / 2 c.-106 to -120). In certain embodiments, the CCAAT box target region may include an 18nt target region, a 13nt target region, an 11nt target region, a 4nt target region, a 1nt target region, a -117G>A target region, or a combination thereof, as disclosed herein. In certain embodiments, the modification may be an 18nt deletion, a 13nt deletion, an 11nt deletion, a 4nt deletion, a 1nt deletion, a G to A substitution at c.-117 in the HBG1, HBG2, or HBG1 and HBG2 genes, or a combination thereof. In certain embodiments, the modification may be a non-spontaneous modification or a spontaneous modification.In certain embodiments, modifications may be introduced to the 13nt target region using one or more gRNAs containing the targeting domains described in SEQ ID NOs. 251-901 or 940-942. In certain embodiments, modifications may be introduced to the CCAAT box target region using one or more gRNAs containing the sequences described in SEQ ID NOs. 251-901, 940-942, 996, 997, 970, 971, 1002, or 1003. In certain embodiments, a genome editing system, guide RNA, and CRISPR-mediated methods may modify the GATA1 binding motif in BCL11Ae located in the +58 DNaseI hypersensitive site (DHS) region of intron 2 of the BCL11A gene ("GATA1 binding motif in BCL11Ae"). In certain embodiments, modifications may be introduced to the GATA1 binding motif of BCL11Ae using one or more gRNAs containing the targeting domains described in SEQ ID NOs. 952-955. In certain embodiments, one or more gRNAs may be used to introduce a GATA1 binding motif modification to BCL11Ae, and one or more gRNAs may be used to introduce a modification to the 13nt target region of HBG1 and / or HBG2.

[0036] In certain embodiments of this specification, the use of optional genome editing system components, such as template nucleic acids (oligonucleotide donor templates), is also provided. In certain embodiments, template nucleic acids for use in CCAAT target region targeting may include, but are not limited to, template nucleic acids that encode modifications to the CCAAT box target region. In certain embodiments, the CCAAT box target region may include 18nt target regions, 11nt target regions, 4nt target regions, 1nt target regions, or combinations thereof. In certain embodiments, the template nucleic acid may be a single-stranded oligodeoxynucleotide (ssODN) or a double-stranded oligodeoxynucleotide (dsODN). In certain embodiments, exemplary full-length donor templates encoding modifications in the 5' and 3' homology arms and the CCAAT box target region are also presented below (e.g., SEQ ID NOs. 904-909, 974-995). In certain embodiments, the template nucleic acid may be a plus-strand or a minus-strand. In certain embodiments, ssODN may include a 5' homology arm, a substitution sequence, and a 3' homology arm. In certain embodiments, the 5' homology arm may be about 25 to about 200 nucleotides or more in length, for example, at least about 25, 50, 75, 100, 125, 150, 175, or 200 nucleotides; the substitution sequence may include a length of 0 nucleotides; and the 3' homology arm may be about 25 to about 200 nucleotides or more in length, for example, at least about 25, 50, 75, 100, 125, 150, 175, or 200 nucleotides. In certain embodiments, ssODN may include one or more phosphorothioates.

[0037] In certain embodiments, a genome editing system, guide RNA, and CRISPR-mediated method for modifying one or more γ-globin genes (e.g., HBG1, HBG2, or HBG1 and HBG2) may include an RNA-inducing nuclease. In certain embodiments, the RNA-inducing nuclease may be Cas9 or modified Cas9. In certain embodiments, the RNA-inducing nuclease may be Cpf1 or modified Cpf1 as disclosed herein.

[0038] In one embodiment, the present disclosure relates to a composition comprising a plurality of cells produced by the method disclosed above, wherein at least 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% of the cells include a modification to the sequence of the human HBG1 or HBG2 gene or the 13nt target region of the plurality of cells produced by the method disclosed above, and at least 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% of the cells include a modification to the sequence of the 13nt target region of the human HBG1 or HBG2 gene, and at least 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% of the cells include a modification to the GATA1 binding motif sequence of BCL11Ae. In certain embodiments, at least a portion of the plurality of cells may be within the erythrocyte lineage. In certain embodiments, the plurality of cells may be characterized by increased levels of fetal hemoglobin expression compared to an unmodified plurality of cells. In certain embodiments, the level of fetal hemoglobin can be increased by at least 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90%. In certain embodiments, the composition may further comprise a pharmaceutically acceptable carrier.

[0039] In one embodiment, the present disclosure relates to a method for modifying a cell, comprising the steps of unwinding a chromatin segment within or proximal to a target region of a nucleic acid within the cell, and inducing a double-strand break (DSB) within the target region of the nucleic acid, thereby modifying the target region. In certain embodiments, the step of unwinding the chromatin segment may include contacting the chromatin segment with an RNA-induced helicase. In certain embodiments, the step of unwinding chromatin does not include recruiting an exogenous trans-acting factor to the chromatin segment. The RNA-induced helicase may be an RNA-induced nuclease, which may complex with a death guide RNA (dgRNA) containing a first targeting domain sequence of 15 nucleotides or less in length. In certain embodiments, the dgRNA may include, but is not limited to, a 5' or 3' end modification, including an anti-reverse cap analog (ARCA) at the 5' end of the RNA, a 3' poly-A tail at the RNA end, or both. In certain embodiments, the RNA-induced helicase may be an enzymatically active RNA-induced nuclease or may be configured to lack nuclease activity. In certain embodiments, the targeting domain sequence of the dgRNA may be complementary to the sequence proximal to the target region. In some embodiments of this specification, “proximal” may mean within 10, 25, 50, 100, or 200 nucleotides of the target region. In certain embodiments, the step of unwinding the chromatin segment may not include the step of forming single-strand or double-strand breaks in the nucleic acids within the chromatin segment. In certain embodiments, the step of generating DSBs within the target region may include the step of contacting the chromatin segment with an RNA-induced nuclease having nuclease activity. In certain embodiments, the RNA-induced nuclease having nuclease activity may complex with a gRNA containing a targeting domain configured to overlap the target region. In certain embodiments, the RNA-induced nuclease having nuclease activity may be the Cpf1 molecule.

[0040] Another aspect of this disclosure includes a method for inducing access to a target region of a nucleic acid for editing within a cell, comprising the steps of contacting a cell with an RNA-inducible helicase and dgRNA, and unwinding DNA within or proximal to the target region with the RNA-inducible helicase to thereby induce access to the target region for editing. In various cases, the RNA-inducible helicase and dgRNA may be configured to bind within or proximal to the target region. In certain embodiments, the dgRNA may be configured not to provide an RNA-inducible nuclease cleavage event. In certain embodiments, the RNA-inducible helicase and dgRNA may complex to form a dead ribonucleoprotein (RNP) lacking cleavage activity. In certain embodiments, the dgRNA may contain a targeting domain sequence with a length of 15 nucleotides or less. In certain embodiments, the RNA-inducible helicase may be an RNA-inducible nuclease. In certain embodiments, the RNA-inducible nuclease and dgRNA are not configured to recruit an exogenous trans-acting factor to the target region. In certain embodiments, the RNA-induced nuclease may be Cas9 or a Cas9 fusion protein. In certain embodiments, Cas9 may be enzymatically active Cas9 or enzymatically killed Cas9. In certain embodiments, Cas9 may be a nickase such as Cas9 D10A. In certain embodiments, the DNA unwinding step does not involve the step of forming single-strand or double-strand breaks in the DNA. In certain embodiments, the RNA-induced nuclease having nuclease activity may complex with a gRNA containing a targeting domain configured to overlap the target region. In certain embodiments, the RNA-induced nuclease having nuclease activity may be a Cpf1 molecule.

[0041] In another embodiment, the disclosure relates to a method for increasing the rate of indel formation in a nucleic acid, comprising the steps of: unwinding double-stranded DNA in or near a target region of a nucleic acid using an RNA-induced helicase configured to bind in or near a target region of the nucleic acid; and generating a double-stranded branch (DSB) in the target region. In certain embodiments, the step of generating a DSB in the target region results in indel formation in the target region. In certain embodiments, the DSB may be repaired in a manner that forms an indel in the target region. In certain embodiments, the rate of indel formation in a gene achieved using an RNA-induced helicase is increased compared to the rate of indel formation in a gene achieved without the use of an RNA-induced helicase. In certain embodiments, the RNA-induced helicase may form an RNP complex with a dgRNA configured to bind in or near a target region. In certain embodiments, the dgRNA may contain a targeting domain sequence of 15 nucleotides or less in length. In certain embodiments, the RNA-induced helicase may be an RNA-induced nuclease. In certain embodiments, the RNA-induced nuclease may be Cas9 or a Cas9 fusion protein. In certain embodiments, Cas9 may be enzymatically active Cas9 or enzymatically killed Cas9. In certain embodiments, Cas9 may be a nickase such as Cas9 D10A. In certain embodiments, the RNA-induced nuclease and dgRNA are not configured to recruit an exogenous transacting factor to the target region. In certain embodiments, the step of unwinding the double-stranded DNA does not include the step of forming single-strand or double-strand breaks in the DNA.

[0042] In another embodiment, the disclosure relates to a method for deleting a segment of a target nucleic acid within a cell, comprising the steps of contacting the cell with an RNA-induced helicase and generating a double-stranded bridge (DSB) within a target region, thereby deleting the segment of the target nucleic acid. In certain embodiments, the DSB may be repaired in a manner that deletes the segment of the target nucleic acid. In certain embodiments, the RNA-induced helicase may be configured to bind to or near a target region of the target nucleic acid and to unwind double-stranded DNA (dsDNA) within or near the target region. In certain embodiments, the RNA-induced helicase may form a ribonucleoprotein complex with a dgRNA configured to bind to or near the target region. In certain embodiments, the dgRNA may include a targeting domain sequence of 15 nucleotides or less in length. In certain embodiments, the RNA-induced helicase may be an RNA-induced nuclease. In certain embodiments, the RNA-induced nuclease may be Cas9 or a Cas9 fusion protein. In certain embodiments, Cas9 may be enzymatically active Cas9 or enzymatically killed Cas9. In certain embodiments, the RNA-induced nuclease and dgRNA are not configured to recruit an exogenous transacting factor to the target region. In certain embodiments, the target nucleic acid may be the promoter region of a gene, the coding region of a gene, the non-coding region of a gene, the intron of a gene, or the exon of a gene. In certain embodiments, the segment of the target nucleic acid may have a length of at least about 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, or 100 base pairs.

[0043] This disclosure also relates to dead gRNA (dgRNA) molecules containing a targeting domain, including a truncated gRNA targeting domain. In certain embodiments, the truncated gRNA targeting domain may be one of the gRNA targeting domains listed in Table 2, Table 10, or Table 15. In certain embodiments, the gRNA targeting domain may be truncated from its 5' end. In certain embodiments, the dgRNA may contain a targeting domain sequence with a length of 15 nucleotides or less. In certain embodiments, the first targeting domain may be identical to or differ from the dgRNA targeting domains listed in Table 10 or Table 15 by 3 nucleotides or less.

[0044] Another aspect of the present disclosure relates to a composition comprising at least one polynucleotide encoding a plurality of gRNAs and one RNA-induced helicase, wherein the at least one gRNA may be a dgRNA configured not to provide an RNA-induced nuclease cleavage event. In certain embodiments, the dgRNA may comprise a targeting domain sequence having a length of 15 nucleotides or less. In certain embodiments, the RNA-induced helicase may be an RNA-induced nuclease. In certain embodiments, the RNA-induced nuclease may be Cas9 or a Cas9 fusion protein. In certain embodiments, Cas9 may be enzymatically active Cas9 or enzymatically killed Cas9. In certain embodiments, Cas9 may be a nickase such as Cas9 D10A. In certain embodiments, the RNA-induced nuclease and the dgRNA are not configured to recruit an exogenous transacting factor to the target region. In certain embodiments, the composition further comprises a second RNA-induced nuclease configured to provide a cleavage event. In certain embodiments, the composition further comprises a second gRNA configured to provide a cleavage event.

[0045] In another embodiment, the disclosure relates to a genome editing system comprising an RNA-induced nuclease and an RNA-induced helicase configured to bind to a target nucleic acid proximal to a target region of the target nucleic acid and induce a conformational change of the target region, thereby facilitating access to the target region by the RNA-induced nuclease and forming a disc in the target region. The disclosure also relates to a genome editing system comprising a dgRNA containing a targeting domain sequence of 15 nucleotides or less in length, a first RNA-induced nuclease, and an RNA-induced helicase. In certain embodiments, the genome editing system further comprises a gRNA. In certain embodiments, the gRNA and the first RNA-induced nuclease may bind to a target region in the target nucleic acid. In certain embodiments, the first RNA-induced nuclease may be a Cpf1 molecule. In certain embodiments, the gRNA and the first RNA-induced nuclease may bind to a first PAM sequence in the target nucleic acid, where the first PAM sequence is outward-facing. In certain embodiments, the RNA-induced helicase may be a second RNA-induced nuclease. In certain embodiments, the second RNA-induced nuclease and dgRNA are not configured to recruit an exogenous trans-acting factor to the target region. In certain embodiments, the dgRNA and the second RNA-induced nuclease bind to or near the target region in the target nucleic acid. In certain embodiments, the first RNA-induced nuclease and the second RNA-induced nuclease may complex with gRNA and dgRNA, respectively, to form first and second ribonucleoprotein complexes.

[0046] In another embodiment, the disclosure relates to a genome editing system comprising a dgRNA containing a targeting domain sequence of 15 nucleotides or less in length, a first RNA-inducible nuclease, and an RNA-inducible helicase. In certain embodiments, the gRNA and the first RNA-inducible nuclease may bind to a target region in a target nucleic acid. In certain embodiments, the first RNA-inducible nuclease may be a Cpf1 molecule. In certain embodiments, the gRNA and the first RNA-inducible nuclease may bind to a first protospacer-adjacent motif (PAM) sequence in a target nucleic acid. In certain embodiments, the first PAM sequence may be outward-facing. In certain embodiments, the RNA-inducible helicase may be a second RNA-inducible nuclease. In certain embodiments, the second RNA-inducible nuclease and the dgRNA are not configured to recruit an exogenous trans-acting factor to the target region. In certain embodiments, the dgRNA and the second RNA-inducible nuclease may bind to or proximal to a target region in a target nucleic acid. In certain embodiments, the first RNA-inducible nuclease and the second RNA-inducible nuclease may complex with gRNA and dgRNA, respectively, to form first and second ribonucleoprotein complexes. In certain embodiments, the dgRNA and the second RNA-inducible nuclease may bind to a second PAM sequence in the target nucleic acid, where the second PAM sequence may be outward-facing.

[0047] In another embodiment, the disclosure relates to a genome editing system comprising a dgRNA, a first gRNA containing a second targeting domain sequence longer than 17 nucleotides, a first RNA-inducible nuclease, and a second RNA-inducible nuclease. In certain embodiments, the first RNA-inducible nuclease and the dgRNA may be configured to bind to a first target region in a target nucleic acid. In certain embodiments, the second RNA-inducible nuclease and the first gRNA may be configured to bind to a second target region to produce a double-strand break (DSB) in the target nucleic acid, thereby generating an indel between the first and second target regions. In certain embodiments, the second RNA-inducible nuclease may be a Cpf1 molecule. In certain embodiments, the dgRNA may contain a first targeting domain sequence shorter than 15 nucleotides. In certain embodiments, the dgRNA may have reduced or no RNA-inducible nuclease cleavage activity. In certain embodiments, the dgRNA may be configured not to provide an RNA-induced nuclease cleavage event. In certain embodiments, the dgRNA and the first RNA-induced nuclease may bind to a first protospacer-adjacent motif (PAM) sequence in the target nucleic acid. In certain embodiments, the first PAM sequence may be outward-facing. In certain embodiments, the first gRNA and the second RNA-induced nuclease may bind to a second PAM sequence in the target nucleic acid. In certain embodiments, the second PAM sequence may be outward-facing.

[0048] In another embodiment, the disclosure relates to a method for modifying cells, comprising the step of contacting the cells with a dgRNA, a first gRNA containing a second targeting domain sequence longer than 17 nucleotides, a first RNA-inducible nuclease, and a second RNA-inducible nuclease. In certain embodiments, the first RNA-inducible nuclease and the dgRNA may be configured to bind to a first target region in a target nucleic acid. In certain embodiments, the second RNA-inducible nuclease and the first gRNA may bind to the second target region to produce a double-strand break (DSB) in the target nucleic acid, thereby generating an indel between the first and second target regions. In certain embodiments, the second RNA-inducible nuclease may be a Cpf1 molecule. In certain embodiments, the dgRNA may contain a first targeting domain sequence less than or equal to 15 nucleotides. In certain embodiments, the dgRNA has reduced or no RNA-inducible nuclease cleavage activity. In certain embodiments, the dgRNA may be configured to provide an RNA-induced nuclease cleavage event. In certain embodiments, the dgRNA and the first RNA-induced nuclease may bind to a first protospacer-adjacent motif (PAM) sequence in the target nucleic acid. In certain embodiments, the first PAM sequence may be outward-facing. In certain embodiments, the first gRNA and the second RNA-induced nuclease may bind to a second PAM sequence in the target nucleic acid. In certain embodiments, the second PAM sequence may be outward-facing.

[0049] This disclosure also relates to a method for modifying cells, comprising the step of contacting the cells with one of the genome editing systems disclosed herein. In certain embodiments, the step of contacting the cells may include contacting the cells with a solution containing first and second ribonucleoprotein complexes. In certain embodiments, the step of contacting the cells with the solution may further include electroporating the cells, thereby introducing the first and second ribonucleoprotein complexes into the cells.

[0050] In another embodiment, this disclosure relates to cells modified using the methods disclosed herein. Cells containing productive indels resulting in HbF expression are also disclosed herein. In certain embodiments, indels can be generated by contacting cells with dgRNA, a first gRNA containing a second targeting domain sequence longer than 17 nucleotides, a first RNA-inducible nuclease, and a second RNA-inducible nuclease. In certain embodiments, the first RNA-inducible nuclease and dgRNA may be configured to bind to a first target region in the target nucleic acid. In certain embodiments, the second RNA-inducible nuclease and the first gRNA may bind to the second target region, causing a double-strand break (DSB) in the target nucleic acid, thereby generating an indel between the first and second target regions. In certain embodiments, the cells disclosed herein may differentiate into erythroblasts, erythrocytes, or precursors of erythrocytes or erythroblasts. In certain embodiments, the cells may be CD34+ cells.

[0051] A genome editing system or method comprising any of the above features may include a target nucleic acid comprising the human HBG1, HBG2 gene, or a combination thereof. In certain embodiments, the target region may be a CCAAT box target region of the human HBG1, HBG2 gene, or a combination thereof. In certain embodiments, a first targeting domain sequence may be complementary to a first sequence on the side of the CCAAT box target region of the human HBG1, HBG2 gene, or a combination thereof, wherein the first sequence optionally overlaps with the CCAAT box target region of the human HBG1, HBG2 gene, or a combination thereof. In certain embodiments, a second targeting domain sequence may be complementary to a second sequence on the side of the CCAAT box target region of the human HBG1, HBG2 gene, or a combination thereof, wherein the second sequence optionally overlaps with the CCAAT box target region of the human HBG1, HBG2 gene, or a combination thereof. In certain embodiments, the first targeting domain may comprise a truncated gRNA targeting domain. In certain embodiments, the gRNA targeting domain may include gRNAs listed in Table 2, Table 10, or Table 15, and the gRNA targeting domain is truncated from its 5' end. In certain embodiments, the first targeting domain is identical to or differs by 3 nucleotides or less from the dgRNA targeting domain listed in Table 10 or Table 15. In certain embodiments, the second targeting domain differs by 3 nucleotides or less from the gRNA targeting domain listed in Table 2, Table 10, or Table 15. In certain embodiments, the indel may be a modified CCAAT box targeting region indel. In certain embodiments, the indel may be a productive indel resulting in an increased level of fetal hemoglobin expression. In certain embodiments, the gRNA, dgRNA, or both may be synthesized ex vivo or chemically.

[0052] In certain embodiments, the cells may include at least one modified allele of the HBG locus produced by any of the methods for modifying cells disclosed herein, wherein the modified allele of the HBG locus includes modifications of the human HBG1 gene, the HBG2 gene, or a combination thereof.

[0053] In certain embodiments, a population of isolated cells may be modified by any of the methods for modifying cells disclosed herein, wherein the population of cells may include a distribution of indels that may differ from the population of isolated cells that have not been modified by the method or from their offspring of the same cell type.

[0054] In certain embodiments, a plurality of cells may be generated by any of the cell modification methods disclosed herein, wherein at least 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% of the cells may involve sequence modifications within the CCAAT box target region of the human HBG1 gene, HBG2 gene, or a combination thereof.

[0055] In certain embodiments, the cells disclosed herein may be used for pharmaceutical purposes. In certain embodiments, the cells may be used in the treatment of β-hemoglobin disorders. In certain embodiments, β-hemoglobin disorders may be selected from the group consisting of sickle cell disease and β-thalassemia.

[0056] In one embodiment, the present disclosure relates to a plurality of cells generated by a method comprising the dgRNA disclosed above, wherein at least 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% of the cells comprises a modification of the sequence of the CCAAT box target region of the human HBG1 or HBG2 gene, or a composition comprising a plurality of cells generated by a method comprising at least 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% of the cells comprises a modification of the sequence of the CCAAT box target region of human HBG1 or HBG2. In certain embodiments, at least a portion of the plurality of cells may be within the erythrocyte lineage. In certain embodiments, the plurality of cells may be characterized by an increased level of fetal hemoglobin expression compared to an unmodified plurality of cells. In certain embodiments, the level of fetal hemoglobin may be increased by at least 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90%. In certain embodiments, the composition may further comprise a pharmaceutically acceptable carrier.

[0057] In one embodiment, the disclosure relates to a population of cells modified by a genome editing system comprising the above-described dgRNA, wherein the population of cells comprises a higher percentage of productive indels compared to a population of cells not modified by the genome editing system. The disclosure also relates to a population of cells modified by a genome editing system comprising the above-described dgRNA, wherein a higher percentage of the cell population compared to a population of cells not modified by the genome editing system can differentiate into a population of erythrocyte lineage cells expressing HbF. In certain embodiments, the higher percentage can be at least about 15%, at least about 20%, at least about 25%, at least about 30%, or at least about 40% higher. In certain embodiments, the cells may be hematopoietic stem cells. In certain embodiments, the cells may differentiate into erythroblasts, erythrocytes, or precursors of erythrocytes or erythroblasts. In certain embodiments, the indels may be generated by repair mechanisms other than microhomology-mediated end-joint (MMEJ) repair.

[0058] This disclosure also relates to the use of any of the cells disclosed herein in the manufacture of a drug for treating β-abnormal hemoglobinopathy in subjects.

[0059] In one embodiment, the present disclosure relates to a method for treating β-abnormal hemoglobinopathy in a subject requiring such treatment, comprising the step of administering cells disclosed herein to a subject. In a particular embodiment, the method for treating β-abnormal hemoglobinopathy in a subject requiring such treatment may comprise the step of administering a population of modified hematopoietic cells, wherein one or more cells are modified according to a method for modifying cells disclosed herein.

[0060] In one embodiment, the disclosure relates to a genome editing system comprising an RNA-induced nuclease and a first guide RNA, wherein the first guide RNA may comprise a first targeting domain complementary to a first sequence on the side of the CCAAT box target region of the human HBG1, HBG2 gene or a combination thereof, the first sequence optionally overlapping the CCAAT box target region of the human HBG1, HBG2 gene or a combination thereof. In certain embodiments, the genome editing system may further comprise a template nucleic acid encoding a modification of the CCAAT box target region of the human HBG1, HBG2 gene or a combination thereof. In certain embodiments, the template nucleic acid may be a single-stranded oligodeoxynucleotide (ssODN) or a double-stranded oligodeoxynucleotide (dsODN). In certain embodiments, the ssODN may comprise a 5' homology arm, a substitution sequence, and a 3' homology arm. In certain embodiments, the homology arms may be symmetrical in length. In certain embodiments, homology arms may be asymmetric in length. In certain embodiments, ssODN may contain one or more phosphorothioate modifications. In certain embodiments, one or more phosphorothioate modifications may be at the 5' end, the 3' end, or a combination thereof. In certain embodiments, ssODN may be a positive or negative chain. In certain embodiments, modifications may be non-spontaneous modifications. In certain embodiments, modifications may include deletions in the CCAAT box target region. In certain embodiments, deletions may include 18nt deletions, 11nt deletions, 4nt deletions, 1nt deletions, or a combination thereof. In certain embodiments, the CCAAT box target region may include 18nt target regions, 11nt target regions, 4nt target regions, 1nt target regions, or a combination thereof. In certain embodiments, the 5' homology arm may be a nucleotide of about 25 to about 200 or more in length, such as at least about 25, 50, 75, 100, 125, 150, 175, or 200 nucleotides; the substitution sequence may include a length of 0 nucleotides; the 3' homology arm may be a nucleotide of about 25 to about 200 or more in length, such as at least about 25, 50, 75, 100, 125, 150, 175, or 200 nucleotides.In certain embodiments, the 5' homology arm may include homologies of approximately 50-100 bp on the 5' side of the 18nt target region, such as 55-95, 60-90, 70-90, or 80-90 bp, and the 3' homology arm may include homologies of approximately 50-100 bp on the 3' side of the 18nt target region, such as 55-95, 60-90, 70-90, or 80-90 bp. In certain embodiments, the ssODN may include, essentially be, or be derived from, sequence number 974 or sequence number 975. In certain embodiments, the 5' homology arm may include homologies of approximately 50-100 bp on the 5' side of the 11nt target region, such as 55-95, 60-90, 70-90, or 80-90 bp, and the 3' homology arm may include homologies of approximately 50-100 bp on the 3' side of the 11nt target region, such as 55-95, 60-90, 70-90, or 80-90 bp. In certain embodiments, the ssODN may include, essentially be, or may be, sequence number 976 or sequence number 978. In certain embodiments, the 5' homology arm may include homologies of approximately 50-100 bp on the 5' side of the 4nt target region, such as 55-95, 60-90, 70-90, or 80-90 bp, and the 3' homology arm may include homologies of approximately 50-100 bp on the 3' side of the 4nt target region, such as 55-95, 60-90, 70-90, or 80-90 bp. In certain embodiments, the ssODN may include, essentially be, or may be, sequences selected from the group consisting of sequence numbers 984, 985, 986, 987, 988, 989, 990, 991, 992, 993, 994, and 995. In certain embodiments, the 5' homology arm may include homology of approximately 50-100 bp on the 5' side of the 1nt target region, such as 55-95, 60-90, 70-90, or 80-90 bp, and the 3' homology arm may include homology of approximately 50-100 bp on the 3' side of the 1nt target region, such as 55-95, 60-90, 70-90, or 80-90 bp. In certain embodiments, the homology arms may be symmetrical in length.In certain embodiments, ssODN may include, be essentially, or be derived from, sequence number 982 or sequence number 983. In certain embodiments, modifications may be spontaneous modifications. In certain embodiments, modifications may include deletions or mutations of the CCAAT box target region. In certain embodiments, the CCAAT box target region may include the 13nt target region, the -117G>A target region, or a combination thereof. In certain embodiments, modifications may include a 13nt deletion in the 13nt target region or a G to A substitution in the -117G>A target region, or a combination thereof. In certain embodiments, the 5' homology arm may include homologies of approximately 50-100 bp on the 5' side of the 13nt target region, such as 55-95, 60-90, 70-90, or 80-90 bp, and the 3' homology arm may include homologies of approximately 50-100 bp on the 3' side of the 13nt target region, such as 55-95, 60-90, 70-90, or 80-90 bp. In certain embodiments, the ssODN may include, essentially be, or be derived from, sequence number 977 or sequence number 979. In certain embodiments, the 5' homology arm may include approximately 50-100 bp of homology on the 5' side of the 13nt target region, such as 55-95, 60-90, 70-90, or 80-90 bp, and the 3' homology arm may include approximately 50-100 bp of homology on the 3' side of the 13nt target region, such as 55-95, 60-90, 70-90, or 80-90 bp. In certain embodiments, the ssODN may include, essentially be, or be derived from, SEQ ID NO: 980 or SEQ ID NO: 981. In certain embodiments, the RNA-inducing nuclease may be S. pyogenes Cas9. In certain embodiments, the RNA-inducing nuclease may be a Cpf1 variant as disclosed herein. In certain embodiments, the first targeting domain may differ by 3 nucleotides or less from the targeting domains listed in Tables 7, 18, and 19, or from the gRNAs in Tables 12 and 19.In certain embodiments, the genome editing system may further include a second guide RNA, which may include a second targeting domain that is complementary to a second sequence on the side of the CCAAT box target region of the human HBG1, HBG2 gene or a combination thereof, and the second sequence optionally overlaps with the CCAAT box target region of the human HBG1, HBG2 gene or a combination thereof. In certain embodiments, the RNA-induced nuclease may be a nickasase that optionally lacks RuvC activity. In certain embodiments, the genome editing system may include first and second RNA-induced nucleases. In certain embodiments, the first and second RNA-induced nucleases may complex with first and second guide RNAs, respectively, to form first and second ribonucleoprotein complexes. In certain embodiments, the genome editing system may further include a third guide RNA and optionally a fourth guide RNA, wherein the third and fourth guide RNAs may include third and fourth targeting domains complementary to the third and fourth sequences, located opposite the position of the GATA1 binding motif in the BCL11A erythrocyte enhancer (BCL11Ae) of the human BCL11A gene, and either or both of the third and fourth sequences optionally overlap with the GATA1 binding motif in the BCL11Ae of the human BCL11A gene. In certain embodiments, the genome editing system may further include a nucleic acid template encoding the deletion of the GATA1 binding motif in the BCL11Ae. In certain embodiments, the RNA-inducing nuclease may be S. pyogenes Cas9. In certain embodiments, the RNA-inducing nuclease may be a nickase, optionally lacking RuvC activity. In certain embodiments, the third targeting domain may be complementary to a sequence within 1000 nucleotides upstream of the GATA1 binding motif in BCL11Ae. In certain embodiments, the third targeting domain may be complementary to a sequence within 100 nucleotides upstream of the GATA1 binding motif in BCL11Ae. In certain embodiments, one of the third and fourth targeting domains may be complementary to a sequence within 100 nucleotides downstream of the GATA1 binding motif in BCL11Ae.In certain embodiments, the fourth targeting domain may be complementary to a sequence of up to 50 nucleotides downstream of the GATA1 binding motif in BCL11Ae. In certain embodiments, at least one of the third and fourth targeting domains may differ from the targeting domains listed in Table 9 by up to 3 nucleotides. In certain embodiments, the genome editing system may include first and second RNA-inducing nucleases. In certain embodiments, the first and second RNA-inducing nucleases may complex with third and fourth guide RNAs, respectively, to form third and fourth ribonucleoprotein complexes.

[0061] In one embodiment, the present disclosure relates to a method for modifying cells, comprising the step of contacting the cells with a genome editing system. In certain embodiments, the step of contacting the cells with a genome editing system may include the step of contacting the cells with a solution containing first and second ribonucleoprotein complexes. In certain embodiments, the step of contacting the cells with a solution may further include the step of electroporating the cells, thereby introducing the first and second ribonucleoprotein complexes into the cells. In certain embodiments, the method for modifying cells may further include the step of contacting the cells with a genome editing system, and the step of contacting the cells with a genome editing system may include the step of contacting the cells with a solution containing first, second, third and optionally fourth ribonucleoprotein complexes. In certain embodiments, the step of contacting the cells with a solution may further include the step of electroporating the cells, thereby introducing the first, second, third and optionally fourth ribonucleoprotein complexes into the cells. In certain embodiments, the cells may differentiate into erythroblasts, erythrocytes, or precursors of erythrocytes or erythroblasts. In certain embodiments, the cells may be CD34+ cells.

[0062] In one embodiment, the present disclosure relates to a CRISPR-mediated method for modifying cells, comprising the steps of: introducing a first DNA single-strand break (SSB) or double-strand break (DSB) into the genome of a cell at position c.-106 to -120 of the human HBG1 or HBG2 gene; and optionally, introducing a second SSB or DSB into the genome of a cell at position c.-106 to -120 of the human HBG1 or HBG2 gene, wherein the first and second SSBs or DSBs can be repaired by the cell in a manner that modifies the CCAAT box target region of the human HBG1 or HBG2 gene. In certain embodiments, the first and second SSBs or DSBs can be repaired by the cell in a manner that results in a modification of the CCAAT box target region of the human HBG1 or HBG2 gene. In certain embodiments, the CRISPR-mediated method may further include a template nucleic acid encoding a modification of the CCAAT box target region of the human HBG1, HBG2 gene, or a combination thereof. In certain embodiments, the template nucleic acid may be a single-stranded oligodeoxynucleotide (ssODN). In certain embodiments, the ssODN may include a 5' homology arm, a substitution sequence, and a 3' homology arm. In certain embodiments, the ssODN may be a positive or negative strand. In certain embodiments, the modification may be a non-spontaneous modification. In certain embodiments, the first and second SSBs or DSBs may be repaired by cells in a manner that results in the formation of at least one indel, deletion, or insertion in the CCAAT box target region of the human HBG1 or HBG2 gene. In certain embodiments, the CCAAT box target region may include an 18nt target region, an 11nt target region, a 4nt target region, a 1nt target region, or a combination thereof. In certain embodiments, the 5' homology arm may be about 25 to about 200 nucleotides or more in length, for example, at least about 25, 50, 75, 100, 125, 150, 175, or 200 nucleotides; the substitution sequence may include a length of 0 nucleotides; the 3' homology arm may be about 25 to about 200 nucleotides or more in length, for example, at least about 25, 50, 75, 100, 125, 150, 175, or 200 nucleotides.In certain embodiments, the 5' homology arm may include homology of approximately 50-100 bp on the 5' side of the 18nt target region, 11nt target region, 4nt target region, or 1nt target region, such as 55-95, 60-90, 70-90, or 80-90 bp, and the 3' homology arm may include homology of approximately 50-100 bp on the 3' side of the 18nt target region, 11nt target region, 4nt target region, or 1nt target region, such as 55-95, 60-90, 70-90, or 80-90 bp. In certain embodiments, the ssODN may include, essentially be, or may be selected from sequences selected from the group consisting of SEQ ID NOs: 974, 975, 976, 978, 984, 985, 986, 987, 988, 989, 990, 991, 992, 993, 994, 995, 982, and 983. In certain embodiments, the modification may be a non-spontaneous modification. In certain embodiments, the first and second SSBs or DSBs may be repaired by cells in a manner that results in the formation of at least one indel, deletion, or insertion in the CCAAT box target region of the human HBG1 or HBG2 gene. In certain embodiments, the CCAAT box target region may include the 13nt target region, the -117G>A target region, or a combination thereof. In certain embodiments, modifications may include a 13nt deletion in the 13nt target region or a G-to-A substitution in the -117G>A target region, or a combination thereof. In certain embodiments, the 5' homology arm may include homology of approximately 50-100 bp on the 5' side of the 13nt target region or -117G>A target region, such as 55-95, 60-90, 70-90, or 80-90 bp, and the 3' homology arm may include homology of approximately 50-100 bp on the 3' side of the 13nt target region or -117G>A target region, such as 55-95, 60-90, 70-90, or 80-90 bp. In certain embodiments, the ssODN may include, essentially be, or be derived from, a sequence selected from the group consisting of SEQ ID NO: 977 or SEQ ID NO: 979, SEQ ID NO: 980, or SEQ ID NO: 981.

[0063] In one embodiment, the present disclosure relates to a composition which may comprise a plurality of cells produced by a cell modification method disclosed herein, wherein at least 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% of the cells may comprise a modification of the sequence of the CCAAT box target region of the human HBG1 gene, HBG2 gene, or a combination thereof. In certain embodiments, the modification may comprise a G to A substitution at 18nt deletion, 11nt deletion, 4nt deletion, 1nt deletion, 13nt deletion, or -117 in the human HBG1 gene, HBG2 gene, or a combination thereof. In certain embodiments, at least a portion of the plurality of cells may be within the erythrocyte lineage. In certain embodiments, the plurality of cells may be characterized by an increased level of fetal hemoglobin expression compared to an unmodified plurality of cells. In certain embodiments, the level of fetal hemoglobin may be increased by at least 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90%. In certain embodiments, the composition may further comprise a pharmaceutically acceptable carrier.

[0064] In one embodiment, the disclosure relates to cells comprising synthetic genotypes produced by methods for modifying cells disclosed herein, wherein the cells may include 18nt deletions, 11nt deletions, 4nt deletions, 1nt deletions, 13nt deletions, and G-to-A substitutions at -117 in the human HBG1 gene, HBG2 gene, or combinations thereof.

[0065] In one embodiment, the disclosure relates to a cell comprising at least one allele of an HBG locus produced by a cell modification method disclosed herein, wherein the cell may encode G to A substitutions at 18nt deletion, 11nt deletion, 4nt deletion, 1nt deletion, 13nt deletion, and -117 in the human HBG1 gene, the HBG2 gene, or a combination thereof.

[0066] In one embodiment, the disclosure relates to an AAV vector which may include a template nucleic acid encoding a non-spontaneous modification of the CCAAT box target region of the human HBG1, HBG2 genes or a combination thereof. In certain embodiments, the template nucleic acid may be a single-stranded oligodeoxynucleotide (ssODN). In certain embodiments, the CCAAT box target region may include an 18nt target region, an 11nt target region, a 4nt target region, a 1nt target region, or a combination thereof. In certain embodiments, the ssODN may include a 5' homology arm, a substitution sequence, and a 3' homology arm. In certain embodiments, the 5' homology arm may be a nucleotide of about 25 to about 200 or more in length, such as at least about 25, 50, 75, 100, 125, 150, 175, or 200 nucleotides; the substitution sequence may include a length of 0 nucleotides; the 3' homology arm may be a nucleotide of about 25 to about 200 or more in length, such as at least about 25, 50, 75, 100, 125, 150, 175, or 200 nucleotides. In certain embodiments, the 5' homology arm may include homology of approximately 50-100 bp on the 5' side of the 18nt target region, 11nt target region, 4nt target region, or 1nt target region, such as 55-95, 60-90, 70-90, or 80-90 bp, and the 3' homology arm may include homology of approximately 50-100 bp on the 3' side of the 18nt target region, 11nt target region, 4nt target region, or 1nt target region, such as 55-95, 60-90, 70-90, or 80-90 bp. In certain embodiments, the ssODN may include, essentially be, or be derived from, sequences selected from the group consisting of SEQ ID NOs: 974-976, SEQ ID NOs: 978, and SEQ ID NOs: 982-995.

[0067] In one embodiment, the disclosure relates to a nucleotide sequence comprising a template nucleic acid encoding a non-spontaneous modification of a CCAAT box target region of the human HBG1, HBG2 genes, or a combination thereof. In certain embodiments, the template nucleic acid may be a single-stranded oligodeoxynucleotide (ssODN) or a double-stranded oligodeoxynucleotide (dsODN) containing the modification. In certain embodiments, the CCAAT box target region may comprise an 18nt target region, an 11nt target region, a 4nt target region, a 1nt target region, or a combination thereof. In certain embodiments, the ssODN may comprise a 5' homology arm, a substitution sequence, and a 3' homology arm. In certain embodiments, the 5' homology arm may be a nucleotide of about 25 to about 200 or more in length, such as at least about 25, 50, 75, 100, 125, 150, 175, or 200 nucleotides; the substitution sequence may include a length of 0 nucleotides; the 3' homology arm may be a nucleotide of about 25 to about 200 or more in length, such as at least about 25, 50, 75, 100, 125, 150, 175, or 200 nucleotides. In certain embodiments, the 5' homology arm may include homology of approximately 50-100 bp on the 5' side of the 18nt target region, 11nt target region, 4nt target region, or 1nt target region, such as 55-95, 60-90, 70-90, or 80-90 bp, and the 3' homology arm may include homology of approximately 50-100 bp on the 3' side of the 18nt target region, 11nt target region, 4nt target region, or 1nt target region, such as 55-95, 60-90, 70-90, or 80-90 bp. In certain embodiments, the ssODN may include, essentially be, or be derived from, sequences selected from the group consisting of SEQ ID NOs: 974-976, SEQ ID NOs: 978, and SEQ ID NOs: 982-995.

[0068] In one embodiment, the present disclosure relates to cells comprising synthetic genotypes, wherein the cells may include 18nt deletions, 11nt deletions, 4nt deletions, 1nt deletions, 13nt deletions, and G-to-A substitutions at -117 in the human HBG1 gene, the HBG2 gene, or a combination thereof.

[0069] In one embodiment, the disclosure relates to a composition comprising a population of cells produced by a method of modifying cells disclosed herein, wherein the cells include a higher frequency of modification of the sequence of the CCAAT box target region of the human HBG1 gene, HBG2 gene, or a combination thereof, compared to a population of unmodified cells. In certain embodiments, the higher frequency is at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% higher. In certain embodiments, the modification includes 18nt deletions, 11nt deletions, 4nt deletions, 1nt deletions, 13nt deletions, and G to A substitutions at -117 in the human HBG1 gene, HBG2 gene, or a combination thereof. In certain embodiments, at least a portion of the population of cells is within the erythrocyte lineage.

[0070] This list is intended to be illustrative and descriptive, not comprehensive or restrictive. Additional aspects and embodiments may be described in or revealed in the remainder of this disclosure and the claims.

[0071] The accompanying drawings are intended to provide illustrative and schematic examples, rather than comprehensive, of any particular aspects and embodiments of the present disclosure. The drawings are not intended to limit or be bound by any particular theory or model, and are not necessarily to scale. Without limiting them, nucleic acids and polypeptides may be depicted as linear sequences or as schematic two- or three-dimensional structures; these depictions are intended to be illustrative rather than limiting or bound by any particular model or theory relating to their structures. [Brief explanation of the drawing]

[0072] [Figure 1]Figure 1 schematically shows the HBG1 and HBG2 genes in relation to the β-globin gene cluster on human chromosome 11. Each gene within the β-globin gene cluster is transcriptionally regulated by a proximal promoter. While we do not wish to be constrained by any particular theory, it is generally believed that Aγ and / or Gγ expression is activated by engagement between the proximal promoter and the distal strongly erythrocyte-specific locus regulatory region (LCR), which is a strong erythrocyte-specific enhancer. Long-range transactivation by the LCR is thought to be mediated by alterations in chromatin arrangement / conformation. The LCR is marked by four erythrocyte-specific DNase I hypersensitivity sites (HS1-HS4) and two distal enhancer elements (5'HS and 3'HS1). Globin gene expression is regulated in a developmental stage-specific manner, and changes in globin gene expression coincide with changes in major sites of hematopoietic production. [Figure 2A] The study identifies small deletions and point mutations within and upstream of the HBG1 and HBG2 proximal promoters that have been identified in the HBG1 and HBG2 genes, their coding sequences (CDS), and patients, and that are associated with elevated fetal hemoglobin (HbF). It also identifies core elements within the proximal promoter (CAAT box, 13nt sequence) that are deleted in some patients with persistent heritable fetal hemoglobin (HPFH). "Target sequence" regions of each locus screened for gRNA binding target sites have also been identified. [Figure 2B] The study identifies small deletions and point mutations within and upstream of the HBG1 and HBG2 proximal promoters that have been identified in the HBG1 and HBG2 genes, their coding sequences (CDS), and patients, and that are associated with elevated fetal hemoglobin (HbF). It also identifies core elements within the proximal promoter (CAAT box, 13nt sequence) that are deleted in some patients with persistent heritable fetal hemoglobin (HPFH). "Target sequence" regions of each locus screened for gRNA binding target sites have also been identified. [Figure 3A]Figure 3 shows data from gRNA screening for 13nt deletion integration in human K562 erythroleukemia cells. Figure 3A shows gene editing determined by T7E1 endonuclease assay analysis (hereinafter referred to synonymously as "T7E1 analysis") of HBG1 and HBG2 locus-specific PCR products amplified from genomic DNA extracted from K562 cells after electroporation with DNA encoding S. pyogenes-specific gRNA and plasmid DNA encoding S. pyogenes-Cas9. [Figure 3B] Figure 3 shows data from gRNA screening for 13nt deletion integration in human K562 erythroleukemia cells. Figure 3B shows gene editing determined by DNA sequencing analysis of PCR products amplified from the HBG1 locus in genomic DNA extracted from K562 cells after electroporation with the DNA encoding the shown gRNA and Cas9 plasmid. In Figure 3B, the type of editing event (insertion, deletion) and deletion subtype (partial deletion of the 13nt target [12nt HPFH] or complete deletion [13-26nt HPFH], deletion of other sequences [other deletion]) are indicated by bars with different shaded / patterned colors. [Figure 3C] Figure 3 shows data from gRNA screening for 13nt deletion integration in human K562 erythroleukemia cells. Figure 3C shows gene editing determined by DNA sequencing analysis of PCR products amplified from the HBG2 locus in genomic DNA extracted from K562 cells after electroporation with the DNA encoding the shown gRNA and Cas9 plasmid. In Figure 3C, the type of editing event (insertion, deletion) and deletion subtype (partial deletion of the 13nt target [12nt HPFH] or complete deletion [13-26nt HPFH], deletion of other sequences [other deletion]) are indicated by bars with different shaded / patterned colors. [Figure 4A]Figure 4 shows the results of gene editing in human umbilical cord blood (CB) and human adult CD34+ cells after electroporation with RNP complexed with in vitro transcribed S. pyogenes gRNA, targeting deletions of specific 13nt sequences (HBG sgRNA Sp35 and Sp37). Figure 4A shows the percentage of indels detected by T7E1 analysis of HBG1 and HBG2-specific PCR products amplified from gDNA extracted from CB CD34+ cells treated with the indicated RNP or from untreated donor-matched control cells (n=3 types of CB CD34+ cells, 3 separate experiments). The data shown represent the mean values ​​from 3 separate donors / 3 experiments, and the error bars correspond to the standard deviation. [Figure 4B] Figure 4 shows the results of gene editing in human umbilical cord blood (CB) and human adult CD34+ cells after electroporation with RNP complexed with in vitro transcribed S. pyogenes gRNA, targeting deletions of specific 13nt sequences (HBG sgRNA Sp35 and Sp37). Figure 4B shows the percentage of indels detected by T7E1 analysis of HBG2-specific PCR products amplified from gDNA extracted from CB CD34+ cells, adult CD34+ cells, or untreated donor-matched control cells treated with the indicated RNP (n=3 CB CD34+ cells, n=3 types of recruited peripheral blood (mPB) CD34+ cells, 3 separate experiments). The data shown represent the mean values ​​from 3 separate donors / 3 experiments, and the error bars represent the standard deviation. [Figure 4C]Figure 4 shows the results of gene editing in human umbilical cord blood (CB) and human adult CD34+ cells after electroporation with RNP complexed with in vitro transcribed S. pyogenes gRNA, targeting specific 13nt sequences (HBG sgRNA Sp35 and Sp37) for deletion. Figure 4C (upper panel) shows indels detected by T7E1 analysis of HBG2 PCR products amplified from gDNA extracted from human CB CD34+ cells electroporated with HBG Sp35RNP or HBG Sp37RNP+ / -ssODN (unmodified or with PhTx-modified 5' and 3' ends). The lower left panel shows the level of gene editing determined by Sanger DNA sequencing analysis of gDNA from cells edited with HBG Sp37RNP and ssODN. The lower right panel shows specific types of deletions detected within the total deletions. [Figure 5A] Figure 5 shows gene editing of HBG in adult human recruited peripheral blood (mPB) CD34+ cells and induction of fetal hemoglobin in erythrocyte offspring of RNP-treated cells after electroporation of mPB CD34+ cells with HBG Sp37RNP+ / -ssODN encoding a 13nt deletion. Figure 5A shows the percentage of indels detected by T7E1 analysis of HBG2 PCR products amplified from gDNA extracted from RNP-treated mPB CD34+ cells or donor-matched untreated control cells. [Figure 5B] Figure 5 shows gene editing of HBG in adult human recruited peripheral blood (mPB) CD34+ cells and induction of fetal hemoglobin in erythrocyte offspring of RNP-treated cells after electroporation of mPB CD34+ cells with HBG Sp37RNP+ / -ssODN encoding a 13nt deletion. Figure 5B shows the ratio of change in HBG mRNA expression in erythroblasts differentiated from RNP-treated and untreated donor-matched control mPB CD34+ cells at day 7. mRNA levels are normalized to GAPDH and calibrated to the levels detected in the untreated control at the corresponding differentiation day. [Figure 6A]Figure 6 shows the in vitro differentiation potential of RNP-treated and untreated mPB CD34+ cells from the same donor. Figure 6A shows the potential of hematopoietic myelocytes / erythrocyte colony-forming cells (CFCs), with the number and subtype of colonies indicated (GEMM: granulocyte-erythrocyte-monocyte-macrophage colony, E: erythrocyte colony, GM: granulocyte-macrophage colony, M: macrophage colony, G: granulocyte colony). [Figure 6B] Figure 6 shows the in vitro differentiation potential of RNP-treated and untreated mPB CD34+ cells from the same donor. Figure 6B shows the percentage of glycophorin A expressed over the course of erythrocyte differentiation, as determined by flow cytometry analysis for the indicated time points and samples. [Figure 7A] Figure 7A shows indels detected by T7E1 analysis of HBG PCR products amplified from gDNA extracted from HBG RNP (D10A paired nickase)-treated human mPB CD34+ cells. In a subset of the samples, cells were also administered ssODN encoding silent SNPs in addition to 13nt deletions, and HDR(ssODN) was monitored. [Figure 7B] Figure 7B shows DNA sequencing analysis for a selected subset of the samples shown in Figure 7A. Indels were subdivided according to indel type (insertion, 13nt deletion, or other deletion). [Figure 8A] Figure 8A shows indels at HBG target sites in mPB CD34+ cells after electroporation by the indicated gRNA pair complexed in D10A nickase and WT RNP pairs. [Figure 8B] Figure 8B shows large deletion events (e.g., HBG2 deletion) in mPB CD34+ cells after electroporation, mediated by the indicated gRNA pair complexed in D10A nickase and WT RNP pairs. [Figure 8C] Figure 8C shows the DNA sequencing analysis and subtypes of events (insertions, deletions) detected in gDNA from mPB CD34+ cells treated with paired D10A nickase pairs. [Figure 8D]Figure 8D shows the DNA sequencing analysis and subtypes of events (insertions, deletions) detected in gDNA from mPB CD34+ cells treated with paired WT RNP pairs. [Figure 9] Figures 7 and 8 outline HbF protein and mRNA expression in progeny of mPB CD34+ cells treated with HBG-targeting paired RNPs for the experiments shown. HbF protein (by HPLC analysis) and HbF mRNA expression (by ddPCR analysis) were evaluated in erythrocyte progeny of RNP-treated human mPB CD34+ cells (the background level of HbF detected in untreated donor-matched controls was subtracted from the levels detected in the progeny of RNP-treated CD34+ cells). [Figure 10A] Figure 10 shows the indel frequency and in vitro and in vitro short-term hematopoietic capacity of CD34+ cells after treatment with paired D10A nickase RNP (SpA+Sp85) at different concentrations (0, 2.5, and 3.7 μM). Indels were evaluated by T7E1 analysis (Figure 10A) and Illumina sequence analysis (insertions and deletions, Figure 10B). [Figure 10B] Figure 10 shows the indel frequency and in vitro and in vitro short-term hematopoietic capacity of CD34+ cells after treatment with paired D10A nickase RNP (SpA+Sp85) at different concentrations (0, 2.5, and 3.7 μM). Indels were evaluated by T7E1 analysis (Figure 10A) and Illumina sequence analysis (insertions and deletions, Figure 10B). [Figure 10C] Figure 10 shows the indel frequency and in vitro and in vitro short-term hematopoietic capacity of CD34+ cells after treatment with paired D10A nickase RNP (SpA + Sp85) at different concentrations (0, 2.5, and 3.7 μM). Figure 10C shows the percentage of HbF protein detected by HPLC analysis (%HbF = 100% × HbF / (HbF + HbA)). [Figure 10D]Figure 10 shows the indel frequency and in vitro and in vitro short-term hematopoietic capacity of CD34+ cells after treatment with paired D10A nickase RNP (SpA+Sp85) at different concentrations (0, 2.5, and 3.7 μM). Figure 10D shows the hematopoietic activity of RNP-treated and donor-matched untreated CD34+ cells in a colony-forming cell (CFC) assay. CFCs are shown per 1000 seeded CD34+ cells. [Figure 10E] Figure 10 shows the indel frequency and in vitro and in vitro short-term hematopoietic capacity of CD34+ cells after treatment with paired D10A nickase RNP (SpA+Sp85) at different concentrations (0, 2.5, and 3.7 μM). Figure 10E shows the CD45+ cell rearrangement of peripheral human blood from immunodeficient mice (NSGs) one month after transplantation of donor-matched human mPB CD34+ cells, treated with either untreated (0 μM) or two doses (2.5 and 3.75 μM) of one D10A RNP and paired gRNA. [Figure 10F] Figure 10 shows the indel frequency and in vitro and in vitro short-term hematopoietic capacity of CD34+ cells after treatment with paired D10A nickase RNP (SpA+Sp85) at different concentrations (0, 2.5, and 3.7 μM). Figure 10F shows the rearrangement of human blood CD45+ cells in peripheral blood of immunodeficient mice (NSG) two months after transplantation. [Figure 10G] Figure 10 shows the indel frequency and in vitro and in vitro short-term hematopoietic capacity of CD34+ cells after treatment with paired D10A nickase RNP (SpA+Sp85) at different concentrations (0, 2.5, and 3.7 μM). Figure 10G shows the strain distribution following human CD45+ blood cell rearrangement of NSG mice at 1 month. [Figure 10H] Figure 10 shows the indel frequency and in vitro and in vitro short-term hematopoietic capacity of CD34+ cells after treatment with paired D10A nickase RNP (SpA+Sp85) at different concentrations (0, 2.5, and 3.7 μM). Figure 10H shows the strain distribution following human CD45+ blood cell rearrangement of NSG mice at 2 months. [Figure 11A]Figure 11A correlates HbF levels assayed by HPLC with indel frequencies assessed by T7E1 analysis for two D10A nickase RNP pairs (SP37+SPB and SP37+SPA) delivered to mPB CD34+ cells at the indicated concentrations. HbF levels were analyzed in erythrocyte progeny (day 18) of edited CD34+ cells. HbF protein detected in untreated donor-matched controls was subtracted from the edited samples. [Figure 11B] Figure 11B shows the indel rate superimposed on hematopoietic colony-forming cell (CFC) activity associated with CD34+ cells treated with the indicated D10A nickase pair or untreated controls. [Figure 11C] Figure 11C shows the rearrangement of human CD45+ blood cells in immunodeficient NSG mice one month after transplantation of mPB CD34+ cells, either treated with a given concentration of D10 RNP nickase pair or with an untreated donor-matched control. [Figure 11D] Figure 11D shows the distribution of human blood lineages detected in the human CD45+ fraction of mouse peripheral blood one month after transplantation. [Figure 12] This shows the target site for HbF desuppression, which is the GATA1 motif of the erythrocyte-specific enhancer at the +58 DNaseI hypersensitive site (DHS) of BCL11A (BCL11Ae) (genomic coordinates: chr2:60,495,265~60,495,270). [Figure 13A] Figure 13A shows the percentage of indels detected by T7E1 endonuclease analysis of BCL11A PCR products amplified from gDNA extracted from CB CD34+ cells treated with the indicated RNP+ / -ssODN or untreated donor-matched control cells. The data shown represent the average values ​​from three separate donors / three experiments. [Figure 13B]Figure 13B shows the indels detected by T7E1 endonuclease analysis of BCL11A PCR products amplified from gDNA extracted from CB CD34+ cells, and the hematopoietic activity of the cells under each condition. These products were treated with either WT RNP (a single gRNA targeting the BCL11A erythrocyte enhancer, complexed with WT S. pyogenes Cas9 having both RuvC and HNH activity) or paired nickase RNP (a paired gRNA (e.g., D10A) targeting the BCL11A erythrocyte enhancer, complexed with S. pyogenes Cas9 nickase sharing the same HNH single-strand cleavage activity). [Figure 14A] Figure 14A shows the editing frequency of BCL11Ae in adult BM CD34+ cells (using a single gRNA approach targeting the GATA1 motif). [Figure 14B] Figure 14B shows the single-allele and bi-allele edits detected in hematopoietic colonies (GEMM, clonal offspring of BCL11Ae RNP-treated CD34+ cells) based on DNA sequencing analysis. [Figure 14C] Figure 14C shows the dynamics of erythroblast maturation (nuclear removal determined by DRAQ5- cells detected by flow cytometry). [Figure 14D] Figure 14D shows the acquisition of the erythrocyte phenotype (glycofolin A+ cells) in differentiated control cells and RNP-treated BM CD34+ cells. [Figure 14E] Figure 14E shows the multiplier of HbF+ cells determined by flow cytometry analysis, compared to HbF+ cells in untreated donor-matching control samples. [Figure 15A]Figure 15 shows gene editing of BCL11Ae in adult mPB CD34+ cells; and induction of fetal hemoglobin in erythrocyte offspring of RNP and ssODN-treated cells after electroporation of mPB CD34+ cells by BCL11Ae RNP + nonspecific ssODN. Figure 15A shows the percentage of indels detected by T7E1 analysis of HBG2 PCR products amplified from gDNA extracted from mPB CD34+ cells treated with BCL11Ae RNP and nonspecific ssODN or from donor-matched untreated control cells. [Figure 15B] Figure 15 shows gene editing of BCL11Ae in adult mPB CD34+ cells; and induction of fetal hemoglobin in erythrocyte offspring of RNP and ssODN-treated cells after electroporation of mPB CD34+ cells by BCL11Ae RNP + nonspecific ssODN. Figure 15B shows the ratio of change in HBG mRNA expression in erythroblasts at day 10 differentiated from BCL11Ae RNP-treated and untreated donor-matched control mPB CD34+ cells (mRNA levels are normalized to GAPDH and calibrated to levels detected in untreated controls at the corresponding differentiation days). [Figure 15C] Figure 15 shows gene editing of BCL11Ae in adult mPB CD34+ cells; and induction of fetal hemoglobin in erythrocyte offspring of RNP and ssODN-treated mPB CD34+ cells after electroporation of mPB CD34+ cells by BCL11Ae RNP + nonspecific ssODN. Figure 15C shows the percentage of glycophorin A expressed over time course of erythrocyte differentiation of mPB CD34+ cells treated with BCL11Ae RNP and nonspecific ssODN, as determined by flow cytometry analysis for the indicated time points and samples. [Figure 16]The percentages of indels detected by next-generation sequencing (NGS) of HBG PCR products amplified from gDNA extracted from hematopoietic stem / progenitor cells (HSPCs) treated with Cas9 ("OLI7066-RNP") complexed with chemically synthesized guide RNA OLI7066 (column 970 in Table 10) at a concentration of 16 μM are shown. Various indels were identified, including HBGΔ-102:-121, HBGΔ-114:-124, HBGΔ-116, HBG-114+T, HBG-116+G, HBGΔ-112:-115, HBGΔ-113:-115, HBGΔ-114:-115, HBGΔ-115, and HBGΔ-102:-114 (spontaneous 13nt deletion). [Figure 17A] Figure 17 shows the expression levels of Gγ-globin and Aγ-globin chains (or AGγ-globin obtained from a 4.9kb deletion) as determined by [γ chain] / [total γ chain + β chain], compared to relative indels held in HBG1 or HBG2, as measured by UPLC analysis in erythrocyte progeny of single HSPCs electroporated with Cas9 ("OLI7066-RNP") complexed with gRNA OLI7066 (SEQ ID NO: 970) (Table 10). Figure 17A shows the expression of AγT-globin chains, determined by [AγT-globin chain] / [total γ chain + β chain], for clones carrying indels indicated on the corresponding HBG1 alleles (HBG1Δ-115, HBG1Δ-114:-115, HBG1Δ-113:-115, HBG1Δ-112:-115, HBG1Δ-102:-114, HBG1Δ-104:-121, HBG1Δ-116). [Figure 17B]Figure 17 shows the expression levels of Gγ-globin and Aγ-globin chains (or AGγ-globin obtained from a 4.9kb deletion) as determined by [γ chain] / [total γ chain + β chain], compared to relative indels held in HBG1 or HBG2, as measured by UPLC analysis in erythrocyte progeny of single HSPCs electroporated with Cas9 ("OLI7066-RNP") complexed with gRNA OLI7066 (SEQ ID NO: 970) (Table 10). Figure 17B shows Gγ-globin chain expression, determined by [Gγ-γ chain] / [total γ chain + β chain], for clones carrying indels indicated on the HBG2 alleles (HBG2Δ-115, HBG2Δ-114:-115, HBG2Δ-113:-115, HBG2Δ-112:-115, HBG2Δ-102:-114, HBG2Δ-104:-121, HBG2Δ-116). To ensure that the analysis of Gγ-globin induction is the result of a single edited allele, only clones with a single allele edit of HBG2 or clones with a deletion of one of the HBG2 alleles (obtained from a 4.9kb deletion) were analyzed. [Figure 17C] Figure 17 shows the expression levels of Gγ-globin and Aγ-globin chains (or AGγ-globin obtained from a 4.9kb deletion) as determined by [γ chain] / [total γ chain + β chain], compared to relative indels held in HBG1 or HBG2, as measured by UPLC analysis in erythrocyte progeny of single HSPCs electroporated with Cas9 ("OLI7066-RNP") complexed with gRNA OLI7066 (SEQ ID NO: 970) (Table 10). Figure 17C shows the AGγT-globin chain expression determined by [AGγT-γ chain] / [total γ chain + β chain] for clones carrying indels indicated on the corresponding HBG1 / 2 alleles (HBG1 / 2Δ-115, HBG1 / 2Δ-114:-115, HBG1 / 2Δ-113:-115, HBG1 / 2Δ-112:-115, HBG1 / 2Δ-102:-114, HBG1 / 2Δ-104:-121, HBG1 / 2Δ-116). [Figure 17D]Figure 17 shows the expression levels of Gγ-globin and Aγ-globin chains (or AGγ-globin obtained from a 4.9kb deletion) as determined by [γ chain] / [total γ chain + β chain], compared to relative indels held in HBG1 or HBG2, as measured by UPLC analysis in erythrocyte progeny of single HSPCs electroporated with Cas9 ("OLI7066-RNP") complexed with gRNA OLI7066 (SEQ ID NO: 970) (Table 10). Figure 17D shows the AGγI-globin chain expression determined by [AGγI-γ chain] / [total γ chain + β chain] for clones carrying indels indicated on the corresponding HBG1 / 2 alleles (HBG1 / 2Δ-115, HBG1 / 2Δ-114:-115, HBG1 / 2Δ-113:-115, HBG1 / 2Δ-112:-115, HBG1 / 2Δ-102:-114, HBG1 / 2Δ-104:-121, HBG1 / 2Δ-116). [Figure 18] The HBG1 and HBG2 genes are schematically shown in relation to the β-globin gene cluster on human chromosome 11. The schematic diagram shows the CCAAT box target sites for HBG1 and HBG2. Due to homology within this region, a single guide RNA, such as OLI8394 (SEQ ID NO: 971), complexed with an RNA nuclease (e.g., Cas9), is cleaved in both HBG1 and HBG2. The editing results following delivery of Cas9 complexed with OLI8394 ("OLI8394-RNP") vary, resulting in deletions or insertions of different sizes. Single-stranded oligodeoxynucleotides (ssODNs) are designed to provide templates for copying the desired indel signature in the CCAAT box (Table 11). The ssODNs "encode" each deletion with sequence homology arms adjacent to the non-existent sequence, generating complete deletions. [Figure 19A]Figure 19 shows the results from gene editing of the CCAAT box target region of HBG in adult CD34+ cells from mPBs ("mPB CD34+ cells") electroporated with 2 μM OLI8394-RNP or OLI7066-RNP and 2.5 μM of various ssODNs (Table 11). Figure 19A shows the percentage of indels detected by sequencing of HBG PCR products 72 hours after electroporation with OLI8394-RNP and ssODN OLI16413 ("-11nt+ strand") or ssODN OLI16411 ("-11nt- strand"). [Figure 19B] Figure 19 shows the results from gene editing of the CCAAT box target region of HBG in adult CD34+ cells from mPBs ("mPB CD34+ cells") electroporated with 2 μM OLI8394-RNP or OLI7066-RNP and 2.5 μM of various ssODNs (Table 11). Figure 19B shows the percentage of exact "-11nt deletions" relative to total indels, detected by sequencing of HBG PCR products 72 hours after electroporation with OLI8394-RNP and ssODN OLI16413 ("-11nt+ strand") or ssODN OLI16411 ("-11nt- strand"). [Figure 19C] Figure 19 shows the results from gene editing of the CCAAT box target region of HBG in adult CD34+ cells from mPB cells ("mPB CD34+ cells") electroporated with 2 μM OLI8394-RNP or OLI7066-RNP and 2.5 μM of various ssODNs (Table 11). Figure 19C shows the percentage of indels detected by sequencing of HBG PCR products 72 hours after electroporation with OLI8394-RNP and ssODN OLI16430 ("-4nt+ strand") or ssODN OLI16424 ("-4nt- strand"). The percentage of exact -4nt deletions (i.e., Δ-112:-115) is distinguished from other indels. [Figure 19D]Figure 19 shows the results from gene editing of the CCAAT box target region of HBG in adult CD34+ cells from mPB cells ("mPB CD34+ cells") electroporated with 2 μM OLI8394-RNP or OLI7066-RNP and 2.5 μM of various ssODNs (Table 11). Figure 19D shows the percentage of indels detected by sequencing of HBG PCR products 72 hours after electroporation with OLI8394-RNP and ssODN OLI16418 ("-1nt+ strand") or ssODN OLI16417 ("-1nt- strand"). The percentage of exact -1nt deletions (i.e., Δ-116) is distinguished from other indels. [Figure 19E] Figure 19 shows the results from gene editing of the CCAAT box target region of HBG in adult CD34+ cells from mPB cells ("mPB CD34+ cells") electroporated with 2 μM OLI8394-RNP or OLI7066-RNP and 2.5 μM of various ssODNs (Table 11). Figure 19E shows the percentage of indels detected by sequencing of HBG PCR products 72 hours after electroporation with OLI8394-RNP and ssODN OLI16409 ("-18nt+ strand") or ssODN OLI16410 ("-18nt- strand"). The percentage of exact -18nt deletions (i.e., Δ-104:-121) is distinguished from other indels. [Figure 19F] Figure 19 shows the results from gene editing of the CCAAT box target region of HBG in adult CD34+ cells from mPB cells ("mPB CD34+ cells") electroporated with 2 μM OLI8394-RNP or OLI7066-RNP and 2.5 μM of various ssODNs (Table 11). Figure 19F shows the percentage of indels detected by sequencing of HBG PCR products 72 hours after electroporation with OLI7066-RNP and ssODN OLI16414 ("-13nt+ strand") or ssODN OLI16412 ("-13nt- strand"). The percentage of exact -13nt deletions (i.e., Δ-102:-114) is distinguished from other indels. [Figure 19G]Figure 19 shows the results from gene editing of the CCAAT box target region of HBG in adult CD34+ cells from mPBs ("mPB CD34+ cells") electroporated with 2 μM OLI8394-RNP or OLI7066-RNP and 2.5 μM of various ssODNs (Table 11). Figure 19G shows the percentage of indels detected by sequencing of HBG PCR products 72 hours after electroporation with OLI8394-RNP and ssODN OLI16416 ("-117G>A+ strand") or ssODN OLI16415 ("-117G>A- strand"). Regardless of the presence or absence of indels, the percentage of reads with -117G>A substitutions is distinguished from other reads. [Figure 20A] Figure 20 shows the expression levels of the γ-globin chain relative to the total β-like globin chain (γ-chain / [γ-chain+β-chain]), as measured by UPLC analysis on erythrocyte progeny of mPB CD34+ cells electroporated with OLI8394-RNP or OLI7066-RNP and various ssODNs (Table 11). Figure 20A shows the percentage of γ-globin chains relative to total β-like globin chains (γ-chain / [γ-chain+β-chain]) measured by ULC after electroporation with (i) OLI8394-RNP and OLI7066-RNP alone, (ii) OLI8394-RNP and ssODN OLI16430 ("-4nt+ chain"), ssODN OLI16424 ("-4nt- chain"), ssODN OLI16413 ("-11nt+ chain") or ssODN OLI16411 ("-11nt- chain"), and (iii) OLI7066-RNP and ssODN OLI16414 ("-13nt+ chain") or ssODN OLI16412 ("-13nt- chain"). [Figure 20B]Figure 20 shows the expression levels of the γ-globin chain relative to the total β-like globin chain (γ-chain / [γ-chain+β-chain]), as measured by UPLC analysis on erythrocyte progeny of mPB CD34+ cells electroporated with OLI8394-RNP or OLI7066-RNP and various ssODNs (Table 11). Figure 20B shows the percentage of γ-globin chains relative to total β-like globin chains (γ-chain / [γ-chain+β-chain]), measured by UPLC after electroporation with OLI8394-RNP and ssODN OLI16418 ("-1nt+ chain"), ssODN OLI16417 ("-1nt- chain"), ssODN OLI16416 ("-117G>A+ chain"), ssODN OLI16415 ("-117G>A- chain"), ssODN OLI16409 ("-18nt+ chain"), or ssODN OLI16410 ("-18nt- chain"). [Figure 21A] Figure 21 shows the results from gene editing of mPB CD34+ cells electroporated with 2 μM OLI8394-RNP and ssODN OLI16424 ("-4nt- strand") (Table 11) at doses ranging from 0.625 μM to 10 μM. Figure 21A shows the percentage of indels detected by sequencing of HBG PCR products 72 hours after electroporation. [Figure 21B] Figure 21 shows the results from gene editing in mPB CD34+ cells electroporated with 2 μM OLI8394-RNP and ssODN OLI16424 ("-4nt- strand") (Table 11) at doses ranging from 0.625 μM to 10 μM. Figure 21B shows the frequency of 4.9 kb deletions detected by ddPCR between HBG1 and HBG2 after electroporation. [Figure 21C] Figure 21 shows the results of gene editing from mPB CD34+ cells electroporated with 2 μM OLI8394-RNP and ssODN OLI16424 ("-4nt- strand") (Table 11) at doses ranging from 0.625 μM to 10 μM. Figure 21C shows the survival percentage of adult CD34+ cells from mPB 48 hours after electroporation. [Figure 21D]Figure 21 shows the results from gene editing of mPB CD34+ cells electroporated with 2 μM OLI8394-RNP and ssODN OLI16424 ("-4nt-chain") (Table 11) at doses ranging from 0.625 μM to 10 μM. Figure 21D shows the percentage of γ-globin chains relative to total β-like globin chains (γ-chain / [γ-chain+β-chain]), as measured by UPLC analysis of cytolysates of erythrocyte progeny from electroporated cells. [Figure 22A] Figure 22 shows the results from gene editing of mPB CD34+ cells electroporated with the indicated doses of OLI8394-RNP and ssODN OLI16424 ("-4nt-chain") (Table 11). Figure 22A shows the percentage of indels detected by sequencing of HBG PCR products 72 hours after electroporation with the indicated doses (2, 4, or 8 μM) of OLI8394-RNP and ssODN OLI16424 ("-4nt-chain") (0, 1.25, 2.5, or 5 μM). [Figure 22B] Figure 22 shows the results from gene editing of mPB CD34+ cells electroporated with the indicated doses of OLI8394-RNP and ssODN OLI16424 ("-4nt-chain") (Table 11). Figure 22B shows the percentage of -4nt deletions ("-112:-115 deletions") detected by next-generation sequencing (NGS) of HBG PCR products after electroporation with the indicated doses (2, 4, or 8 μM) of OLI8394-RNP and ssODN OLI16424 ("-4nt-chain") (0, 1.25, 2.5, or 5 μM). [Figure 22C] Figure 22 shows the results from gene editing of mPB CD34+ cells electroporated with the indicated doses of OLI8394-RNP and ssODN OLI16424 ("-4nt-chain") (Table 11). Figure 22C shows the frequency of 4.9kb deletions between HBG1 and HBG2 after electroporation with the indicated doses of OLI8394-RNP (2, 4, or 8 μM) and ssODN OLI16424 ("-4nt-chain") (0, 1.25, 2.5, or 5 μM). Deletions were measured via ddPCR. [Figure 22D] Figure 22 shows the results from gene editing of mPB CD34+ cells electroporated with the indicated doses of OLI8394-RNP and ssODN OLI16424 ("-4nt-chain") (Table 11). Figure 22D shows the percentage of γ-globin chains relative to total β-like globin chains (γ-chain / [γ-chain+β-chain]), measured by UPLC analysis of cytolysates from erythrocyte progeny of mPB CD34+ cells after electroporation with the indicated doses of OLI8394-RNP (2, 4, or 8 μM) and ssODN OLI16424 ("-4nt-chain") (0, 1.25, 2.5, or 5 μM). [Figure 23A] Figure 23 shows schematic diagrams of ssODN templates with symmetric and asymmetric homology arms and the results provided thereby. Figure 23A shows the CCAAT box target sites in HBG1 and HBG2, targeted by OLI8394 (SEQ ID NO: 971) and OLI7066 (SEQ ID NO: 970). ssODNs with symmetric or asymmetric arms are designed to provide templates for copying the -4nt deletion (HBG-112:-115) in HBG1 and HBG2 (Table 11). The ssODN "encodes" each deletion by a sequence homology arm adjacent to a non-existent sequence, generating a complete deletion in HBG-112:-115. [Figure 23B] Figure 23 shows schematic diagrams of ssODN templates with symmetric and asymmetric homology arms and the results provided thereby. Figure 23B shows the percentage of indels detected by sequencing of HBG PCR products 72 hours after electroporation of mPB CD34+ cells with 2 μM OLI8394-RNP and 2.5 μM of various ssODNs, OLI16424 ("90 / 90"), OLI16419 ("40 / 80"), or OLI16421 ("50 / 50") (HBG-112:-115) (Table 11) that "encode" a 4nt deletion. [Figure 24]The percentage of indels detected by sequencing of HBG PCR products 72 hours after electroporation of mPB CD34+ cells using D10A Cas9 complexed with Sp37 and SpA gRNA ("sp37-D10A-RNP+spA-D10A-RNP") alone or in combination with ssODN OLI16424 ("-4nt-chain") (Table 11) is shown. The percentage of indels with exact -4nt deletions (i.e., Δ-112:-115) is distinguished from other indels. [Figure 25] This shows gene editing of HBG in mPB CD34+ cells electroporated with mutants of the Acidaminococcus species Cpf1 ("AsCpf1"), namely His-AsCpf1-nNLS (SEQ ID NO: 1000) and His-AsCpf1-sNLS-sNLS (SEQ ID NO: 1001) ("His-AsCpf1-nNLS_HBG1-1 RNP" and "His-AsCpf1-sNLS-sNLS_HBG1-1 RNP") complexed with guide RNA HBG1-1 (OLI13620) (Table 13). RNPs were electroporated at 5 μM or 20 μM. [Figure 26A] Figure 26 shows gene editing of HBG in mPB CD34+ cells electroporated with His-AsCpf1-sNLS-sNLS_HBG1-1 RNP alone or in combination with various ssODNs. Figure 26A shows the percentage of indels detected by sequencing of HBG PCR products 72 hours after electroporation with His-AsCpf1-sNLS-sNLS_HBG1-1 RNP alone or in combination with OLI164324 ("-4nt- strand"), OLI16430 ("-4nt+ strand"), OLI16410 ("-18nt- strand"), or OLI16409 ("-18nt+ strand"). [Figure 26B]Figure 26 shows gene editing of HBG in mPB CD34+ cells electroporated with His-AsCpf1-sNLS-sNLS_HBG1-1 RNP alone or in combination with various ssODNs. Figure 26B shows the percentage of exact 18-nucleotide deletion indels detected by sequencing of HBG PCR products 72 hours after electroporation with His-AsCpf1-sNLS-sNLS_HBG1-1 RNP alone or in combination with OLI16410 ("-18nt- strand") or OLI16409 ("-18nt+ strand"). [Figure 26C] Figure 26 shows gene editing of HBG in mPB CD34+ cells electroporated with His-AsCpf1-sNLS-sNLS_HBG1-1 RNP alone or in combination with various ssODNs. Figure 26C shows the percentage of exact 18nt deletions in all indels detected by sequencing of HBG PCR products 72 hours after electroporation with His-AsCpf1-sNLS-sNLS_HBG1-1 RNP alone or in combination with OLI16410 ("-18nt- strand") or OLI16409 ("-18nt+ strand"). [Figure 27]Figure 27A shows a schematic diagram of the pair of HBG1-1 target region and S. pyogenes Cas9 gRNA used in combination. Figure 27A shows the target region of HBG1-1 gRNA (including the RNA targeting domain described in SEQ ID NO: 1002 in Table 15). The distal CCAAT box of the HBG promoter (i.e., HBG1 / 2 c.-111~-115) is shown in gray. Figure 27B shows the target region of HBG1-1, the distal CCAAT box of HBG, and the target region SpA gRNA (including the targeting domain of SEQ ID NO: 941 in Table 15). Figure 27C shows the target region of HBG1-1, the distal CCAAT box of HBG, and the target region SpG gRNA (including the targeting domain of SEQ ID NO: 359 in Table 15). Figure 27D shows the target region of HBG1-1, the distal CCAAT box of HBG, the target region of tSpA death gRNA ("dgRNA") (including the targeting domain of SEQ ID NO: 326 in Table 15), and the target region of Sp182 dgRNA (including the targeting domain of SEQ ID NO: 1028 in Table 15). Figure 27E shows the target region of HBG1-1, the distal CCAAT box of HBG, and tSpA dgRNA (including the targeting domain of SEQ ID NO: 326 in Table 15). Figure 27F shows the target region of HBG1-1, the distal CCAAT box of HBG, and the target region of Sp182 dgRNA (including the targeting domain of SEQ ID NO: 1028 in Table 15). [Figure 28A]Figure 28 shows HbF expression achieved by in vitro editing of mPB CD34+ cells using HBG1-1-AsCpf1-RNP, which targets the HBG promoter region. Figure 28A shows the results of editing in the HBG promoter region within mPB CD34+ cells following delivery of 5 μM or 20 μM of HBG1-1-AsCpf1-RNP ("HBG-1-1") via Amaxa electroporation. Delivery of 20 μM of HBG1-1-AsCpf1-RNP via Amaxa electroporation results in up to approximately 43% editing and 21% HbF induction (above background levels). HbF levels are represented by black circles indicating γ-globin chain expression levels relative to total β-like globin chains (γ-chain / [γ-chain+β-chain]), as measured by ULC analysis on erythrocyte progeny of mPB CD34+ cells. The gray bars indicate the percentage of indels detected by next-generation sequencing (NGS) of HBG PCR products 72 hours after electroporation. [Figure 28B] Figure 28 shows HbF expression achieved by in vitro editing of mPB CD34+ cells using HBG1-1-AsCpf1-RNP, which targets the HBG promoter region. Figure 28B shows the results of editing in the HBG promoter region within mPB CD34+ cells following delivery of 5 μM or 20 μM of HBG1-1-AsCpf1-RNP ("HBG-1-1") via MaxCyte electroporation. Delivery of 20 μM of HBG1-1-AsCpf1-RNP via MaxCyte electroporation results in up to approximately 16% editing and 7% HbF induction (above background levels). HbF levels are represented by black circles indicating γ-globin chain expression levels relative to total β-like globin chains (γ-chain / [γ-chain+β-chain]), as measured by ULC analysis on erythrocyte progeny of mPB CD34+ cells. The gray bars indicate the percentage of indels detected by NGS of the HBG PCR product 72 hours after electroporation. [Figure 29]This shows enhanced editing of the HBG promoter region on a MaxCyte instrument by HBG1-1-AsCpf1H800A-RNP upon simultaneous delivery of various S. pyogenes Cas9 WT or Cas9 D10A RNPs. "S.Py D10A" represents the Cas9 D10A nickase protein, and "S.Py WT" represents the Cas9 WT protein. RNPs tested include SpA-D10A-RNP, SpG-D10A-RNP, tSpA-Cas9-RNP, Sp182-Cas9-RNP, and tSpA-Cas9-RNP+Sp182-Cas9-RNP (Table 14). The diagram shows total editing (gray bars) and associated HbF protein induction (black circles) in the HBG promoter region following delivery of HBG1-1-AsCpf1H800A-RNP ("HBG1-1") alone or in combination with S. pyogenes Cas9 RNP or an RNP pair. HbF levels are represented by black circles indicating the γ-globin chain expression level relative to the total β-like globin chain (γ-chain / [γ-chain+β-chain]), as measured by ULC analysis on erythrocyte progeny of mPB CD34+ cells. The gray bars show the percentage of indels detected by NGS of the HBG PCR product. [Figure 30] This report shows the viability of mPB CD34+ cells following MaxCyte delivery with HBG1-1-AsCpf1H800A-RNP ("HBG1-1") alone or in combination with various S. pyogenes Cas9 WT or Cas9 D10A RNPs. "S.Py D10A" represents the Cas9 D10A nickase protein, and "S.Py WT" represents the Cas9 WT protein. RNPs tested included SpA-D10A-RNP, SpG-D10A-RNP, tSpA-Cas9-RNP, Sp182-Cas9-RNP, and tSpA-Cas9-RNP+Sp182-Cas9-RNP (Table 14). Viability was measured 24 hours after electroporation by DAPI staining and flow cytometry analysis. [Figure 31A]Figure 31 shows the cleavage sites of HBG1-1-AsCpf1H800A-RNP and D10A-Cas9 RNP in the target region, as well as the editing profile obtained from the simultaneous delivery of HBG1-1-AsCpf1H800A-RNP and D10A RNP. Figure 31A shows the locations of the HBG1-1-AsCpf1H800A-RNP cleavage sites on each strand of the target region (light gray arrows) and the locations of the nicking sites targeted by SpG-D10A-RNP and SpA-D10A-RNP (Table 14) (dark arrows). [Figure 31B] Figure 31 shows the cleavage sites of HBG1-1-AsCpf1H800A-RNP and D10A-Cas9 RNP in the target region, as well as the editing profiles obtained from the simultaneous delivery of HBG1-1-AsCpf1H800A-RNP and D10A RNP. Figure 31B shows the editing profiles obtained from the simultaneous delivery of HBG1-1-AsCpf1H800A-RNP ("HBG1-1 RNP") and SpG-D10A-RNP ("spG RNP") or SpA-D10A-RNP ("spA RNP"), as detected by NGS analysis of HBG PCR products 72 hours after electroporation of mPB CD34+ (Table 14). The X-axis represents the genomic position of the center of the indel relative to the HBG1-1-AsCpf1H800A-RNP plus-strand cleavage site. The Y-axis represents the length of the indel, with deletions represented by negative values ​​and insertions by positive values. The total frequency of each indel is represented by the area of ​​the symbol. Indels occurring at a frequency of 0.1% or more are shown. SpG and SpA target sites are indicated by dashed lines. [Figure 32A]Figure 32 shows that simultaneous delivery of HBG1-1-AsCpf1H800A-RNP and Sp182-Cas9-RNP results in an increase in total indels and distal CCAAT box-interfered indels without substantial alteration of the indel profile, as detected by NGS analysis of HBG PCR products 72 hours after electroporation. Figure 32A shows the indel profile following editing with HBG1-1-AsCpf1H800A-RNP ("HBG1-1 RNP") alone or in combination with Sp182-Cas9-RNP ("sp182 RNP"). The X-axis represents the genomic position of the center of the indel relative to the HBG1-1-AsCpf1H800A-RNP plus-strand break site. The Y-axis represents the length of the indel, with deletions represented as negative values ​​and insertions as positive values. The total frequency of each indel is represented by the area of ​​the symbol. Indels occurring at a frequency of 0.1% or higher are shown. The Sp182 target site is indicated by a dotted line. [Figure 32B] Figure 32 shows that the simultaneous delivery of HBG1-1-AsCpf1H800A-RNP and Sp182-Cas9-RNP results in an increase in total indels and distal CCAAT box-interfering indels without substantial alteration of the indel profile, as detected by NGS analysis of the HBG PCR product 72 hours after electroporation. Figure 32B shows the frequency of indels interfering with any of the entire 0nt, 1nt, 2nt, 3nt, 4nt, or 5nt distal CCAAT box sequence. [Figure 33]Table 14 shows that the optimal dose of HBG1-1-AsCpf1H800A-RNP delivered co-administered with Sp182-Cas9-RNP results in increased total editing and HbF production. Delivery of HBG1-1-AsCpf1H800A-RNP ("HBG1-1") together with Sp182-Cas9-RNP ("Sp182") as an RNP pair achieved over 92% editing (gray bars) in the HBG promoter region, with up to 34% HbF induction (above background) (black circles). No editing was observed when Sp182-Cas9-RNP was delivered alone at 12 μM. HbF levels are represented by black circles indicating γ-globin chain expression levels relative to total β-like globin chains (γ-chain / [γ-chain+β-chain]), measured by ULC analysis on erythrocyte progeny of mPB CD34+ cells. The gray bars indicate the percentage of indels detected by NGS of HBG PCR products. [Figure 34] Table 14 shows the distribution level of γ-chain expression relative to the total β-like chain (γ-chain / [γ-chain+β-chain]) in clonal erythrocyte progeny of single human mPB CD34+ cells edited in the HBG promoter region with HBG1-1-AsCpf1H800A-RNP combined with Sp182-Cas9-RNP. Each black circle represents the level of γ-globin protein detected in a population of clonal erythrocytes derived from single cells isolated by FACS sorting 48 hours after electroporation. [Figure 35A]Figure 35 shows the total editing, HbF production, viability, and colony formation ability after simultaneous delivery of RNP containing modified HBG1-1 gRNA (sequence number 1041 in Table 14) (represented as "His-AsCpf1-sNLS-sNLS H800A_HBG1-1RNP" and "HBG1-1" in Figures 35A-C) complexed with His-AsCpf1-sNLS-sNLS H800A (sequence number 1032 in Table 14) and increasing concentrations of ssODN OLI16431 (sequence number 1040 in Table 11) (represented as "OLI16431" in Figures 35A-C) (Table 11). Figure 35A shows distal CAATT box (gray bars) editing and HbF induction (black circles) after co-delivery of 6 μM His-AsCpf1-sNLS-sNLS H800A_HBG1-1 RNP and increasing concentrations of ssODN OLI16431. HbF levels are represented by black circles indicating γ-globin chain expression levels relative to total β-like globin chains (γ chain / [γ chain + β chain]), measured by ULC analysis on erythrocyte progeny of mPB CD34+ cells. Gray bars indicate the percentage of indels detected by NGS of HBG PCR products. [Figure 35B] Figure 35 shows the total editing, HbF production, viability, and colony formation ability after simultaneous delivery of RNP containing modified HBG1-1 gRNA (sequence number 1041 in Table 14) (represented as "His-AsCpf1-sNLS-sNLS H800A_HBG1-1RNP" and "HBG1-1" in Figures 35A-C) complexed with His-AsCpf1-sNLS-sNLS H800A (sequence number 1032 in Table 14) and increasing concentrations of ssODN OLI16431 (sequence number 1040 in Table 11) (represented as "OLI16431" in Figures 35A-C) (Table 11). Figure 35B shows the viability of mPB CD34+ cells following delivery of His-AsCpf1-sNLS-sNLS H800A_HBG1-1 RNP alone or in combination with increasing doses of ssODN OLI16431. Viability was measured by DAPI exclusion 72 hours after electroporation. [Figure 35C]Figure 35 shows the total editing, HbF production, viability, and colony-forming ability after simultaneous delivery of RNP containing modified HBG1-1 gRNA (SEQ ID NO: 1041 in Table 14) (represented as "His-AsCpf1-sNLS-sNLS H800A_HBG1-1RNP" and "HBG1-1" in Figures 35A-C) complexed with His-AsCpf1-sNLS-sNLS H800A (SEQ ID NO: 1032 in Table 14) and increasing concentrations of ssODN OLI16431 (SEQ ID NO: 1040 in Table 11) (represented as "OLI16431" in Figures 35A-C) (Table 11). Figure 35C shows the hematopoietic activity of "HBG1-1" RNP and ssODN OLI16431-treated cells and donor-matched untreated control CD34+ cells in a colony-forming cell (CFC) assay. CFC indicates the number per 800 seeded CD34+ cells. The number of colonies and subtypes are displayed (GEMM: granulocyte-erythrocyte-monocyte-macrophage colonies (black), GM: granulocyte-macrophage colonies (dark gray), E: erythrocyte colonies (light gray)). [Figure 36A]The results for overall editing, HbF production, and viability are shown using different concentrations of RNP containing unmodified HBG1-1 gRNA (SEQ ID NO: 1022 in Table 14) ("His-AsCpf1-sNLS-sNLS H800A_HBG1-1 RNP", represented as "HBG1-1" in Figure 36A) complexed with His-AsCpf1-sNLS-sNLS H800A (SEQ ID NO: 1032 in Table 14) ("His-AsCpf1-sNLS-sNLS H800A_HBG1-1 RNP"), which was delivered simultaneously with various concentrations of ssODN OLI16431 (SEQ ID NO: 1040 in Table 11) (represented as "OLI16431" in Figure 36A) (Table 11). Figure 36A shows the editing (black bars (48 hours) and light gray bars (14 days into erythrocyte culture)) and HbF induction (black circles) in the distal CAATT box after co-delivery of various concentrations of His-AsCpf1-sNLS-sNLS H800A_HBG1-1 RNP ("HBG1-1") and ssODN OLI16431. HbF levels are represented by black circles indicating the γ-globin chain expression level relative to the total β-like globin chain (γ chain / [γ chain + β chain]) as measured by ULC analysis on erythrocyte progeny of mPB CD34+ cells. The bars show the percentage of indels detected by NGS of HBG PCR products at 48 hours (black) and 14 days (light gray) into erythrocyte culture. [Figure 36B]The results for overall editing, HbF production, and viability are shown using different concentrations of RNP containing unmodified HBG1-1 gRNA (SEQ ID NO: 1022 in Table 14) ("His-AsCpf1-sNLS-sNLS H800A_HBG1-1 RNP", represented as "HBG1-1" in Figure 36B) complexed with His-AsCpf1-sNLS-sNLS H800A (SEQ ID NO: 1032 in Table 14) ("His-AsCpf1-sNLS-sNLS H800A_HBG1-1 RNP"), which was co-delivered with various concentrations of ssODN OLI16431 (SEQ ID NO: 1040 in Table 11) (represented as "OLI16431" in Figure 36B) (Table 11). Figure 36B shows the viability of mPB CD34+ cells following delivery of His-AsCpf1-sNLS-sNLS H800A_HBG1-1 RNP alone ("HBG1-1"), ssODN OLI16431 alone, or in combination with various doses of HBG1-1 and OLI16431. Viability was measured by DAPI exclusion 48 hours after electroporation and 14 days after erythrocyte culture. Following editing of mPB CD34+ cells, in vitro differentiation into erythrocyte lineages was carried out over 18 days (Giarratana 2011). A subset of cells was isolated on day 14 of culture, and viability (DAPI exclusion) and editing (NGS of HB GPCR products) were measured. [Figure 37] This shows RNP-mediated editing in mPB CD34+ cells. As shown in Table 21, the RNP contained gRNA complexed with the Cpf1 protein. 72 hours after electroporation, the isolated genomic DNA was sequenced using Illumina sequencing. [Figure 38A] Figure 38 shows the bulk editing of CD34+ cell populations (black bars), progenitor cells (light gray bars), and HSCs (dark gray bars) determined by Illumina sequencing 48 hours after electroporation. Figure 38A shows RNP33 (Table 21) delivered alone or co-delivered with ssODN OLI16431 (SEQ ID NO: 1040 in Table 11). [Figure 38B]Figure 38 shows the bulk editing of CD34+ cell populations (black bars), progenitor cells (light gray bars), and HSCs (dark gray bars) as determined by Illumina sequencing 48 hours after electroporation. Figure 38B shows RNP33 (Table 21) delivered alone or co-delivered with Sp182 RNP (dead gRNA including SEQ ID NO: 1027 (Table 14) complexed with S. pyogenes Cas9 (SEQ ID NO: 1033)). [Figure 39] The bulk editing of CD34+ cell populations (black bars), progenitor cells (light gray bars), and HSCs (dark gray bars) determined by Illumina sequencing 48 hours after electroporation is shown. RNP34, RNP33, and RNP43 (Table 21) were delivered alone or in combination with Sp182 RNP (dead gRNA including SEQ ID NO: 1027 (Table 14) complexed with S. pyogenes Cas9 (SEQ ID NO: 1033)) or ssODN OLI16431 (SEQ ID NO: 1040 in Table 11). [Figure 40] Editing of CD34+ cells determined by Illumina sequencing 72 hours after electroporation is shown. RNP64, RNP63, and RNP45 (Table 21) were delivered in either a stoichiometric ratio of 2 or 4 (gRNA:Cpf1 complex formation ratio) with a molar excess of gRNA. [Figure 41] Editing of CD34+ cells determined by Illumina sequencing 72 hours after electroporation is shown. RNP33, RNP64, RNP63, and RNP45 (Table 21) were delivered alone or in combination with Sp45 RNP (dead gRNA containing SEQ ID NO: 1027 (Table 14) complexed with S. pyogenes Cas9 (SEQ ID NO: 1033)) or ssODN OLI16431 (SEQ ID NO: 1040 in Table 11). [Figure 42] Editing in CD34+ cells, as determined by Illumina sequencing, is shown. RNPs containing Cpf1 (sequence number 1094) complexed with gRNAs having various 5' DNA extensions (Table 21) were delivered alone or in combination with 8 μM OLI16431 (sequence number 1040 in Table 11). [Figure 43] Editing in CD34+ cells, as determined by Illumina sequencing, is shown. RNPs containing gRNAs with matching 5' ends (RNP49 vs. RNP58 and RNP59 vs. RNP60, Table 21) were delivered to CD34+ cells, and the effect of 3' modification was evaluated. In all comparisons, gRNAs with 3'PS-OMe were superior to the unmodified 3' version 24 hours after electroporation. [Figure 44A] Editing in CD34+ cells, as determined by Illumina sequencing 24 and 48 hours after electroporation, is shown. RNP58 (Table 21) was delivered to CD34+ cells in stoichiometric ratios (gRNA:Cpf1 complex formation ratio) of 2:1, 1:1, or 0.5:1. At all doses tested, editing was best when RNP complexed at a ratio of 2:1. [Figure 44B] Editing in CD34+ cells, as determined by Illumina sequencing 24 and 48 hours after electroporation, is shown. RNP58 (Table 21) was delivered to CD34+ cells in stoichiometric ratios (gRNA:Cpf1 complex formation ratio) of 2:1, 1:1, or 0.5:1. At all doses tested, editing was best when RNP complexed at a ratio of 2:1. [Figure 45A] Editing in CD34+ cells as determined by Illumina sequencing is shown. Figure 45A shows RNPs (Table 21) containing gRNAs with matching 5' ends but different 3' modifications, delivered to CD34+ cells to evaluate the effects of 3' modification or extension. [Figure 45B] Editing in CD34+ cells, as determined by Illumina sequencing, is shown. Figure 45B shows RNPs (Table 21) containing gRNAs with matching 5' ends but different 3' modifications, delivered to CD34+ cells to evaluate the effects of 3' modification or extension. [Figure 46A]Figure 46 shows editing in CD34+ cells and their erythrocyte progeny, and HbF levels in erythrocyte progeny, following delivery of RNPs targeting various cleavage sites within the HBG gene locus. Figure 46A shows RNPs containing guide RNA with 1xPS-Ome at the unmodified 5' and 3' ends (Table 21). [Figure 46B] Figure 46 shows editing in CD34+ cells and their erythrocyte progeny, and HbF levels in the erythrocyte progeny, following delivery of RNPs targeting various cleavage sites within the HBG gene locus. Figure 46B shows the RNP (Table 21) containing guide RNA with a 5' 2PS+20 DNA extension and a 3' 1xPS-Ome extension. [Figure 46C] Figure 46 shows editing in CD34+ cells and their erythrocyte progeny, and HbF levels in the erythrocyte progeny, following delivery of RNPs targeting various cleavage sites within the HBG gene locus. Figure 46C shows the RNP (Table 21) containing guide RNA with a 25-DNA extension at the 5' end and 1xPS-Ome at the 3' end. [Figure 47] This shows editing and HbF levels in erythrocyte progeny of CD34+ cells following delivery of 1 μM, 2 μM, and 4 μM of RNP58. [Figure 48] This document describes editing in CD34+ cells following MaxCyte electroporation of RNPs. RNP58, RNP26, RNP27, and RNP28 (Table 21), containing gRNA SEQ ID NO: 1051 complexed with different Cpf1 proteins (SEQ ID NOs: 1094, 1096, 1107, and 1108), were delivered to CD34+ cells. Editing was determined by Illumina sequencing 24 and 48 hours after electroporation. [Figure 49] This report describes editing in CD34+ cells following MaxCyte electroporation of RNPs. RNP58, RNP29, RNP30, and RNP31 (Table 21), containing Cpf1 protein SEQ ID NO: 1094, complexed with guide RNAs having various 5' extensions, were delivered to CD34+ cells. Editing was determined by Illumina sequencing 24 and 48 hours after electroporation. RNP30 was not tested at 1 μM due to limited cell numbers (nt). [Figure 50] The bulk editing of CD34+ cell populations (black bars), progenitor cells (dark gray bars), and HSCs (light gray bars) determined by Illumina sequencing 48 hours after electroporation is shown. RNP58, RNP27, and RNP26 (Table 21) were delivered to CD34+ cells at 2 μM or 4 μM. [Figure 51] The bulk editing of CD34+ cell populations (black bars), progenitor cells (dark gray bars), and HSCs (light gray bars) determined by Illumina sequencing 48 hours after electroporation is shown. RNP61, RNP62, and RNP34 (Table 21) (8 μM) were co-delivered to CD34+ cells along with ssODN OLI16431 (SEQ ID NO: 1040 in Table 11) (8 μM). [Figure 52] The bulk editing of CD34+ cell populations (black bars), progenitor cells (dark gray bars), and HSCs (light gray bars) determined by Illumina sequencing 48 hours after electroporation is shown. RNP58 and RNP32 (Table 21) were delivered to CD34+ cells at 2 μM. [Figure 53] The bulk editing of CD34+ cell populations (black bars), progenitor cells (dark gray bars), and HSCs (light gray bars) determined by Illumina sequencing 48 hours after electroporation is shown. RNP58 and RNP1 (Table 21) were delivered to CD34+ cells at 2 μM, 4 μM, or 8 μM. Cells edited with 2 μM RNP1 were not sorted (NS), and therefore, editing data is not available. [Figure 54] The image shows indels of mPB CD34+ cells transplanted from BM of "NBSGW" mice 8 weeks after infusion of electroporated cells. RNP34 and RNP33 (Table 21) (8 μM) were co-delivered to CD34+ cells along with ssODN OLI16431 (SEQ ID NO: 1040 in Table 11) (6 μM). [Figure 55A]Figure 55 shows indels in transplanted mPB CD34+ cells and HbF expression in chimeric BM-derived erythrocytes from "NBSGW" mice 8 weeks after infusion of electroporated cells. RNP33 or RNP34 (Table 21) was co-delivered to CD34+ cells with Sp182 RNP (dead gRNA including SEQ ID NO: 1027 (Table 14) complexed with S. pyogenes Cas9 (SEQ ID NO: 1033)) (16 μM total RNP). Figure 55A shows indel frequencies in unfractionated bone marrow or fluid-sorted individual populations of CD15+, CD19+, GlyA+, and Lin-CD34+ cells in pseudo-transferred (without RNP) or RNP-transferred cells. Lin-CD34+ cells are defined as CD34+ cells that are negative for CD3, CD14, CD15, CD16, CD19, CD20, and CD56 from the bone marrow (BM) of non-irradiated NOD, B6.SCID Il2rγ- / -Kit(W41 / W41) ("NBSGW") mice infused with mPB CD34+ cells that have been mimicked (without RNP) or RNP-transplanted. Indels were determined for each cell population by Illumina sequencing. [Figure 55B] Figure 55 shows HbF expression in indels of transplanted mPB CD34+ cells and erythrocytes derived from chimeric BM of "NBSGW" mice 8 weeks after infusion of electroporated cells. RNP33 or RNP34 (Table 21) was co-delivered to CD34+ cells with Sp182 RNP (dead gRNA including SEQ ID NO: 1027 (Table 14) complexed with S. pyogenes Cas9 (SEQ ID NO: 1033)) (16 μM total RNP). Figure 55B shows HbF expression calculated by UPLC as γ / β-like (%) from erythrocyte lysates following 18 days of erythrocyte differentiation culture from whole chimeric BM. [Figure 56A]Figure 56 shows indels of transplanted mPB CD34+ cells and HbF expression in chimeric BM-derived erythrocytes from "NBSGW" mice 8 weeks after infusion of electroporated cells. RNP61 or RNP62 (Table 21) (8 μM) was co-delivered to CD34+ cells along with ssODN OLI16431 (SEQ ID NO: 1040 in Table 11) (8 μM). Figure 56A shows indels of unfractionated bone marrow or fluid-sorted individual populations of CD15+, CD19+, GlyA+, and Lin-CD34+ cells in simulated transfusion (without RNP) or RNP-transfusion cells. Lin-CD34+ cells are defined as CD34+ cells that are negative for CD3, CD14, CD15, CD16, CD19, CD20, and CD56 from non-irradiated NBSGW mouse bone marrow (BM) infused with mPB CD34+ cells that have been mimicked (without RNP) or transfused with RNP. Indels were determined for each cell population by Illumina sequencing. [Figure 56B] Figure 56 shows HbF expression in indels of transplanted mPB CD34+ cells and erythrocytes derived from chimeric BM of "NBSGW" mice 8 weeks after infusion of electroporated cells. RNP61 or RNP62 (Table 21) (8 μM) was co-delivered to CD34+ cells along with ssODN OLI16431 (SEQ ID NO: 1040, Table 11) (8 μM). Figure 56B shows HbF expression calculated by UPLC as γ / β-like (%) by erythrocytes following 18 days of erythrocyte differentiation culture from whole chimeric BM. [Figure 57] Table 21 shows the human chimeric phenomenon in the bone marrow 8 weeks after infusion of mPB CD34+ cells edited with simulated transfection (without RNP) or mPB CD34+ cells edited with RNP1 (4 or 8 μM) or RNP58 (2, 4, or 8 μM). The human chimeric phenomenon and strain rearrangement (CD45+, CD14+, CD19+, glycophorin A (GlyA, CD235a+), strain, and CD34+ and mouse CD45+ marker expression) in BM were determined by flow cytometry. [Figure 58]Table 21 shows indels in unclassified bulk bone marrow 8 weeks after infusion of mPB CD34+ cells edited with simulated transfusion (without RNP) or mPB CD34+ cells edited with RNP1 (4 or 8 μM) or RNP58 (2, 4, or 8 μM). Indels were determined by Illumina sequencing. [Figure 59] This report shows the indel frequencies in unfractionated bone marrow or fluid-sorted individual populations of CD15+, CD19+, GlyA+, and Lin-CD34+ cells in simulated transfusion (without RNP) or RNP-transfusiond cells. Lin-CD34+ cells are defined as CD34+ cells negative for CD3, CD14, CD15, CD16, CD19, CD20, and CD56 from the bone marrow (BM) of non-irradiated NOD,B6.SCID Il2rγ- / -Kit(W41 / W41) ("NBSGW") mice infused with simulated transfusion (without RNP) or RNP-transfusiond mPB CD34+ cells. Indels were determined for each cell population by Illumina sequencing. [Figure 60] HbF is shown from the GlyA+ fraction isolated from bone marrow 8 weeks after infusion of simulated (RNP-free) mPB CD34+ cells or mPB CD34+ cells edited with RNP58 (2, 4, or 8 μM) (Table 21). [Figure 61] This shows the colony-forming ability of cells from a bone marrow flush, collected 8 weeks after infusion of simulated or edited human-mobilized CD34+ cells. The number and subtype of colonies are shown (GEMM: granulocyte-erythrocyte-monocyte-macrophage colonies (black), GM: granulocyte-macrophage colonies (dark gray), E: erythrocyte colonies (light gray)). [Figure 62-1]Table 20 shows the sequences of Cpf1 protein variants. Nuclear localized sequences are shown in bold, and 6-histidine sequences are shown underlined. For example, the addition of two or more nNLS sequences or combinations of nNLS and sNLS sequences (or other NLS sequences) to either the N-terminal or C-terminal position, and the addition of sequences containing and not containing purified sequences such as 6-histidine sequences, are within the scope of the currently disclosed subject matter. [Figure 62-2] Table 20 shows the sequences of the Cpf1 protein variants. (The explanation below is the same as for Figure 62-1.) [Figure 62-3] Table 20 shows the sequences of the Cpf1 protein variants. (The explanation below is the same as for Figure 62-1.) [Figure 62-4] Table 20 shows the sequences of the Cpf1 protein variants. (The explanation below is the same as for Figure 62-1.) [Figure 62-5] Table 20 shows the sequences of the Cpf1 protein variants. (The explanation below is the same as for Figure 62-1.) [Figure 62-6] Table 20 shows the sequences of the Cpf1 protein variants. (The explanation below is the same as for Figure 62-1.) [Figure 62-7] Table 20 shows the sequences of the Cpf1 protein variants. (The explanation below is the same as for Figure 62-1.) [Figure 62-8] Table 20 shows the sequences of the Cpf1 protein variants. (The explanation below is the same as for Figure 62-1.) [Figure 62-9] Table 20 shows the sequences of the Cpf1 protein variants. (The explanation below is the same as for Figure 62-1.) [Figure 62-10] Table 20 shows the sequences of the Cpf1 protein variants. (The explanation below is the same as for Figure 62-1.) [Figure 62-11] Table 20 shows the sequences of the Cpf1 protein variants. (The explanation below is the same as for Figure 62-1.) [Figure 62-12] Table 20 shows the sequences of the Cpf1 protein variants. (The explanation below is the same as for Figure 62-1.) [Figure 62-13] Table 20 shows the sequences of the Cpf1 protein variants. (The explanation below is the same as for Figure 62-1.) [Modes for carrying out the invention]

[0073] Definitions and Abbreviations Unless otherwise specified, the following terms have the meanings associated with them in this section.

[0074] The indefinite articles "a" and "an" refer to at least one of the related nouns and are used interchangeably with the terms "at least one" and "one or more." For example, "module" means at least one module or one or more modules.

[0075] The conjunctions "or" and "and / or" are used synonymously as non-exclusive disjunctions.

[0076] The term "domain" is used to represent a segment of a protein or nucleic acid. Unless otherwise specified, a domain does not need to possess any specific functional properties.

[0077] The term “exogenous transacting factor” refers to any peptide or nucleotide component of a genome editing system that (a) interacts with an RNA-inducible nuclease or gRNA by modification means such as the insertion or fusion of a peptide or nucleotide into the RNA-inducible nuclease or gRNA, and (b) interacts with target DNA to alter its helical structure. Peptide or nucleotide insertions or fusions may include, without limitation, direct covalent bonds between the RNA-inducible nuclease or gRNA and the exogenous transacting factor, and / or non-covalent bonds mediated by the insertion or fusion of RNA / protein interaction domains such as the MS2 loop and protein / protein interaction domains such as the DZ, Lim, or SH1, 2, or 3 domains. Other specific RNA and amino acid interaction motifs will be well known to those skilled in the art. Transacting factors may generally include transcription activators.

[0078] The term “booster element” refers to an element that, when co-delivered with a ribonucleoprotein (RNP) complex containing gRNA complexed with an RNA-inducing nuclease (“gRNA-nuclease-RNP”), enhances the editing of the target nucleic acid compared to editing without the booster element. In certain embodiments, co-delivery may be sequential or simultaneous. In certain embodiments, the booster element may be an RNP complex containing death guide RNA complexed with a WT Cas9 protein, a Cas9 nickase protein (e.g., Cas9 D10A protein), or an enzymatically inactive Cas9 (eiCas9) protein. In certain embodiments, the booster element may be an RNP complex containing guide RNA complexed with a Cas9 nickase protein (e.g., Cas9 D10A protein) or an enzymatically inactive Cas9 (eiCas9) protein. In certain embodiments, the booster element may be single-stranded or double-stranded donor template DNA. In certain embodiments, one or more booster elements may be delivered co-delivered with a gRNA-nuclease-RNP to enhance editing of the target nucleic acid. In certain embodiments, the booster elements may be delivered co-delivered with an RNP containing a gRNA complexed with a Cpf1 molecule ("gRNA-Cpf1-RNP") to enhance editing of the target nucleic acid.

[0079] A "productive indel" refers to an indel (deletion and / or insertion) that results in HbF expression. In certain embodiments, a productive indel may induce HbF expression. In certain embodiments, a productive indel may result in an increase in the level of HbF expression.

[0080] An "indel" is an insertion and / or deletion in a nucleic acid sequence. Indels may be products of DNA double-strand break repair, such as double-strand breaks formed by the genome editing systems of this disclosure. Indels are most commonly formed when breaks are repaired by "erroneous" repair pathways, such as the NHEJ pathway described below.

[0081] "Genetic transformation" refers to the modification of a DNA sequence by incorporating endogenous homologous sequences (e.g., homologous sequences within a gene array). "Genetic modification" refers to the modification of a DNA sequence by incorporating exogenous homologous sequences, such as exogenous single-stranded or double-stranded donor template DNA. Genetic transformation and genetic modification are products of DNA double-strand break repair via HDR pathways, such as those described below.

[0082] Indel, gene transformation, gene modification, and other genome editing results are typically evaluated by sequencing (most commonly by "next-generation" or "synthetic sequencing" methods, although Sanger sequencing may still be used) and quantified by the relative frequency of numerical changes in all sequencing reads (e.g., ±1, ±2 or more bases). DNA samples for sequencing can be prepared by a variety of methods known in the art, including amplification of the target site by polymerase chain reaction (PCR), capture of DNA ends resulting from double-strand breaks, such as in the GUIDEseq process described in Tsai 2016 (incorporated herein by reference), or by other means known in the art. Genome editing results can also be evaluated by in-situ hybridization methods such as the FiberComb® system commercialized by Genomic Vision (Bagneux, France) and any other suitable methods known in the art.

[0083] "Alt-HDR," "alternative homology-directed repair," or "alternative HDR" are used synonymously to refer to a process that repairs DNA damage using homologous nucleic acids (e.g., endogenous homologous sequences such as sister chromatids or exogenous nucleic acids such as template nucleic acids). Alt-HDR differs from standard HDR in that the process utilizes a different pathway and can be inhibited by standard HDR mediators, RAD51 and BRCA2. Alt-HDR is also distinguished by the involvement of single-stranded or nicked homologous nucleic acid templates, whereas standard HDR generally involves double-stranded homologous templates.

[0084] Standard HDR, standard homology-directed repair, or cHDR refers to the process of repairing NA damage using homologous nucleic acids (e.g., endogenous homologous sequences such as sister chromatids or exogenous nucleic acids such as template nucleic acids). Standard HDR typically functions when there is significant excision at double-strand breaks and at least one single-stranded portion of DNA is formed. In normal cells, cHDR typically involves a series of steps including break recognition, break stabilization, excision, single-stranded DNA stabilization, DNA crossover intermediate formation, crossover intermediate degradation, and ligation. The process requires RAD51 and BRCA2, and the homologous nucleic acid is typically double-stranded.

[0085] Unless otherwise specified, the term "HDR" in this specification encompasses both standard HDR and alt-HDR.

[0086] "Non-homologous end joining" or "NHEJ" refers to ligation-mediated repairs and / or non-template-mediated repairs such as standard NHEJ (cNHEJ) and alternative NHEJ (altNHEJ), which then include microhomology-mediated end joining (MMEJ), single-strand annealing (SSA), and synthesis-dependent microhomology-mediated end joining (SD-MMEJ).

[0087] When used in relation to the modification of a molecule (e.g., nucleic acid or protein), "substitution" or "substituted" simply indicates the presence of a substituted entity, without requiring any process restrictions.

[0088] "Subject" means human, mouse, or non-human primate. Human subjects can be of any age (e.g., infant, child, young adult, or adult) and may have a disease or require genetic modification.

[0089] "To treat," "to treat," and "treatment" mean the treatment of a disease in an object (e.g., a human subject), including suppressing the disease, i.e., preventing or stopping its onset or progression; alleviating the disease, i.e., causing regression of the diseased state; alleviating one or more symptoms of the disease; and curing the disease.

[0090] "Preventing," "preventing," and "prevention" refer to the prevention of a disease in a subject, including (a) avoiding or eliminating the disease; (b) influencing predisposition to the disease; or (c) preventing or delaying the onset of at least one symptom of the disease.

[0091] A “kit” means any set of two or more components that together constitute a functional unit that can be used for a particular purpose. For illustrative purposes (and not limited to), a kit according to this disclosure may include a guide RNA that is or can be complexed with an RNA-inducing nuclease and accompanied by a pharmaceutically acceptable carrier (e.g., suspended or suspendable therein). In certain embodiments, a kit may include a booster element. For example, the kit may be used to introduce the complex into cells or subjects for the purpose of causing a desired genomic modification in such cells or subjects. The components of the kit may be packaged together or they may be packaged separately. A kit according to this disclosure may optionally also include, for example, instructions for use (DFU) describing the use of the kit according to the methods of this disclosure. The DFU may be physically packaged with the kit or it may be provided to the kit user by, for example, electronic means.

[0092] The terms “polynucleotide,” “nucleotide sequence,” “nucleic acid,” “nucleic acid molecule,” “nucleic acid sequence,” and “oligonucleotide” refer to a series of nucleotide bases (also called “nucleotides”) in DNA and RNA, and mean any chain of two or more nucleotides. Polynucleotides, nucleotide sequences, nucleic acids, etc., can be single-stranded or double-stranded chimeric mixtures or derivatives or modified versions thereof. They can be modified, for example, with base moieties, sugar moieties, or phosphate backbone to improve molecular stability, their hybridization parameters, etc. Nucleic acid sequences typically contain genetic information, including, but not limited to, information used by cellular mechanisms to produce proteins and enzymes. These terms include double-stranded or single-stranded genomic DNA, RNA, any synthetic and genetically engineered polynucleotides, and both sense and antisense polynucleotides. These terms also include nucleic acids containing modified bases.

[0093] As shown in Table 1 below, the conventional IUPAC notation is used in the nucleotide sequences presented herein (see also Cornish-Bowden A, Nucleic Acids Res. 1985 May 10; 13(9):3021-30, which is incorporated herein by reference). However, it should be noted that if the sequence can be encoded by either DNA or RNA, for example in a gRNA targeting domain, “T” indicates “thymine or uracil”.

[0094] [Table 1]

[0095] The terms “protein,” “peptide,” and “polypeptide” are used synonymously and refer to a continuous chain of amino acids linked together via peptide bonds. This term includes individual proteins, groups or complexes of proteins linked together, and fragments or parts, variants, derivatives, and analogues of such proteins. Peptide sequences are presented herein using conventional notation, starting from the amino or N-terminus on the left and progressing to the carboxyl or C-terminus on the right. Standard one- or three-letter abbreviations may be used.

[0096] The term "CCAAT box target region" refers to the sequence located 5' to the transcription start site (TSS) of the HBG1 and / or HBG2 genes. The CCAAT box is a highly conserved motif within the promoter regions of α-like and β-like globin genes. Regions within or near the CCAAT box play a crucial role in the regulation of globin genes. For example, the distal CCAAT box of γ-globin is associated with the heritability persistence of fetal hemoglobin. Several transcription factors have been reported to bind to overlapping CCAAT box regions of the γ-globin promoter, such as NF-Y, COUP-TFII (NF-E3), CDP, GATA1 / NF-E1, and DRED (Martyn 2017). While we do not wish to impose theoretical constraints, the binding site of the transcription activator NF-Y is thought to overlap with that of transcription repressors of the γ-globin promoter. For example, HPFH mutations located within the distal γ-globin promoter region, such as within or near the CCAAT box, can alter the competitive binding of these factors and thus contribute to increased γ-globin expression and elevated HbF levels. The genomic locations provided herein for HBG1 and HBG2 are based on the coordinates provided in the NCBI reference sequence NC_000011 “Homo sapiens chromosome 11, GRCh38.p12 Primary Assembly,” (Version NC_000011.10). The distal CCAAT boxes for HBG1 and HBG2 are located at c.-111 to -115 (genomic locations are Hg38 Chr11:5,249,968 to Chr11:5,249,972 and Hg38 Chr11:5,254,892 to Chr11:5,254,896, respectively). The HBG1 c.-111~-115 region is exemplified in sequence number 902 (HBG1) at positions 2823~2827, and the HBG2 c.-111~-115 region is exemplified in sequence number 903 (HBG2) at positions 2747~2751.In certain embodiments, the “CCAAT box target region” refers to a region of the distal CCAAT box or a region adjacent thereto, and includes the nucleotides of the distal CCAAT box and 25 nucleotides upstream (5') and downstream (3') of the distal CCAAT box (i.e., HBG1 / 2 c.-86~-140) (the genomic locations are Hg38 Chr11:5249943~Hg38 Chr11:5249997 and Hg38 Chr11:5254867~Hg38 Chr11:5254921, respectively). The HBG1 c.-86~-140 region is exemplified in SEQ ID NO: 902 (HBG1) at locations 2798~2852, and the HBG2 c.-86~-140 region is exemplified in SEQ ID NO: 903 (HBG2) at locations 2723~2776. In other embodiments, the “CCAAT box target region” refers to a region of the distal CCAAT box or a region adjacent thereto, and includes the nucleotides of the distal CCAAT box and the five nucleotides upstream (5') and downstream (3') of the distal CCAAT box (i.e., HBG1 / 2 c.-106~-120 (the genomic positions are Hg38 Chr11:5249963~Hg38 Chr11:5249977 (HGB1) and Hg38 Chr11:5254887~Hg38 Chr11:5254901), respectively). The HBG1 c.-106~-120 region is exemplified in SEQ ID NO: 902 (HBG1) at positions 2818~2832, and HBG2 The c.-106~-120 region is exemplified in SEQ ID NO: 903 (HBG2) at positions 2742~2756. The term "CCAAT box target site modification" refers to modifications (e.g., deletions, insertions, mutations) of one or more nucleotides in the CCAAT box target region. Exemplary examples of CCAAT box target region modifications include, but are not limited to, 1nt deletions, 4nt deletions, 11nt deletions, 13nt deletions, and 18nt deletions, as well as the -117G>A modification. In the use of this specification, the terms "CCAAT box" and "CAAT box" may be used synonymously.

[0097] The notations "c.-114~-102 region", "c.-102~-114 region", "-102:-114", and "13nt target region" refer to the 5' sequence of the transcription start site (TSS) of the HBG1 and / or HBG2 genes at genomic locations Hg38 Chr11:5,249,959~Hg38 Chr11:5,249,971 and Hg38 Chr11:5,254,883~Hg38 Chr11:5,254,895, respectively. The HBG1 c.-102~-114 region is exemplified at positions 2824~2836 in SEQ ID NO: 902 (HBG1), and the HBG2 c.-102~-114 region is exemplified at positions 2748~2760 in SEQ ID NO: 903 (HBG2). Terms such as "13nt deletion" refer to the deletion of the target region of 13nt.

[0098] The notations "c.-121~-104 region", "c.-104~-121 region", "-104:-121", and "18nt target region" refer to the 5' sequence of the transcription start site (TSS) of the HBG1 and / or HBG2 genes at genomic locations Hg38 Chr11:5,249,961~Hg38 Chr11:5,249,978 and Hg38 Chr11:5,254,885~Hg38 Chr11:5,254,902, respectively. The HBG1 c.-104~-121 region is exemplified at positions 2817~2834 in SEQ ID NO: 902 (HBG1), and the HBG2 c.-104~-121 region is exemplified at positions 2741~2758 in SEQ ID NO: 903 (HBG2). Terms such as "18nt deletion" refer to the deletion of the target region of 18nt.

[0099] The notations "c.-105~-115 region", "c.-115~-105 region", "-105:-115", and "11nt target region" refer to the 5' sequence of the transcription start site (TSS) of the HBG1 and / or HBG2 genes at genomic locations Hg38 Chr11:5,249,962~Hg38 Chr11:5,249,972 and Hg38 Chr11:5,254,886~Hg38 Chr11:5,254,896, respectively. The HBG1 c.-105~-115 region is exemplified at positions 2823~2833 in SEQ ID NO: 902 (HBG1), and the HBG2 c.-105~-115 region is exemplified at positions 2747~2757 in SEQ ID NO: 903 (HBG2). Terms such as "11nt deletion" refer to the deletion of the target region of 11nt.

[0100] The terms "c.-115~-112 region," "c.-112~-115 region," "-112:-115," and "4nt target region" refer to the 5' sequences of the transcription start sites (TSS) of the HBG1 and / or HBG2 genes at genomic locations Hg38 Chr11:5,249,969~Hg38 Chr11:5,249,972 and Hg38 Chr11:5,254,893~Hg38 Chr11:5,254,896, respectively. The HBG1 c.-112~-115 region is exemplified at locations 2823~2826 in SEQ ID NO: 902, and the HBG2 c.-112~-115 region is exemplified at locations 2747~2750 in SEQ ID NO: 903 (HBG2). Terms such as "4nt deletion" refer to the deletion of the target region of 4nt.

[0101] The terms "c.-116 region," "HBG-116," and "1nt target region" refer to the 5' sequence of the transcription start site (TSS) of the HBG1 and / or HBG2 genes at genomic locations Hg38 Chr11:5,249,973 and Hg38 Chr11:5,254,897, respectively. The HBG1 c.-116 region is exemplified at location 2822 in SEQ ID NO: 902, and the HBG2 c.-116 region is exemplified at location 2746 in SEQ ID NO: 903 (HBG2). Terms such as "1nt deletion" refer to the deletion of a 1nt target region.

[0102] The terms "c.-117G>A region," "HBG-117G>A," and "-117G>A target region" refer to the 5' sequence of the transcription start site (TSS) of the HBG1 and / or HBG2 genes at genomic locations Hg38 Chr11:5,249,974~Hg38 Chr11:5,249,974 and Hg38 Chr11:5,254,898~Hg38 Chr11:5,254,898, respectively. The HBG1 c.-117G>A region is exemplified by the substitution of guanine (G) to adenine (A) at position 2821 in SEQ ID NO: 902, and the HBG2 c.-117G>A region is exemplified by the substitution of G to A at position 2745 in SEQ ID NO: 903 (HBG2). Terms such as "-117G>A modification" refer to the substitution of G to A in the -117G>A target region.

[0103] The term “proximal HBG1 / 2 promoter target sequence” refers to a region within 50, 100, 200, 300, 400, or 500 bp proximal to the HBG1 / 2 promoter sequence containing the 13nt target region. Modifications by genome editing systems in accordance with this disclosure facilitate the upregulation of HbF production in erythrocyte offspring (e.g., tend to induce, promote, or increase the likelihood of HbF production).

[0104] The term "GATA1 binding motif in BCL11Ae" refers to the sequence that is the GATA1 binding motif in the erythrocyte-specific enhancer of BCL11A (BCL11Ae), located in the +58 DNaseI hypersensitivity site (DHS) region of intron 2 of the BCL11A gene. The genomic coordinates of the GATA1 binding motif in BCL11Ae are chr2:60,495,265~60,495,270. The +58 DHS site contains a 115 base pair (bp) sequence described in sequence number 968. The +58 DHS site sequence, including approximately 500 bp upstream and approximately 200 bp downstream, is described in sequence number 969.

[0105] Where a range is provided herein, it includes endpoints. Furthermore, unless otherwise specified or is evident from the context and / or the understanding of those skilled in the art, values ​​expressed as a range are understood to take any specific value within the range described, up to one-tenth of the unit of the lower limit of the range, unless the context explicitly states otherwise. Also, unless otherwise specified or is evident from the context and / or the understanding of those skilled in the art, values ​​expressed as a range are understood to take any partial range within a given range, and the endpoints of the partial range are understood to be expressed with the same precision as one-tenth of the unit of the lower limit of the range.

[0106] overview Various embodiments of this disclosure generally relate to genome editing systems configured to introduce modifications (e.g., deletions or insertions or other mutations) into chromosomal DNA to enhance the transcription of the HBG1 and / or HBG2 genes encoding the Aγ and Gγ subunits of hemoglobin, respectively. In certain embodiments, increased expression of one or more γ-globin genes (e.g., HBG1, HBG2) using the methods provided herein results in the preferential formation of HbF over HbA and / or an increase in HbF levels as a percentage of total hemoglobin. In certain embodiments, this disclosure generally relates to the use of an RNP complex comprising a gRNA complexed with a Cpf1 molecule. In certain embodiments, the gRNA may be unmodified or modified, and the Cpf1 molecule may be a wild-type Cpf1 protein or a modified Cpf1 protein. In certain embodiments, the gRNA may comprise sequences listed in Table 13, Table 18, or Table 19. In certain embodiments, modified Cpf1 may be encoded by sequences described in SEQ ID NOs: 1000, 1001, 1008-1018, 1032, 1035-39, 1094-1097, 1107-09 (Cpf1 polypeptide sequences) or SEQ ID NOs: 1019-1021, 1110-17 (Cpf1 polynucleotide sequences). In certain embodiments, the RNP complex may comprise the RNP complexes described in Table 21. For example, the RNP complex may comprise a gRNA containing the sequence described in SEQ ID NO: 1051 and a modified Cpf1 protein encoded by the sequence described in SEQ ID NO: 1097 (RNP32 in Table 21).

[0107] Patients with the hereditary persistent condition of fetal hemoglobin (HPFH) have mutations in gamma-globin regulators, which have been previously shown to result in lifelong expression of fetal gamma-globin that is not suppressed before or after birth (Martyn 2017). This leads to elevated expression of fetal hemoglobin (HbF). HPFH mutations can be deletional or non-deletional (e.g., point mutations). Subjects with HPFH exhibit lifelong HbF expression, i.e., they do not undergo globin switching or undergo it only partially without symptoms of anemia.

[0108] HbF expression is, for example, HBG1 c.-114 C>T;c.-117 G>A;c.-158 C>T;c.-167 C>T;c.-170 G>A;c.-175 T>G;c.-175 T>C;c.-195 C>G;c.-196 C>T;c.-197 C>T;c.-198 T>C;c.-201 C>T;c.-202 C>T;c.-211 C>T, c.-251 T>C; or c.-499 T>A; or HBG2 c.-109 G>T;c.-110 A>C;c.-114 C>A;c.-114 C>T;c.-114 C>G;c.-157 C>T;c.-158 These mutations can be induced through point mutations in γ-globin regulatory elements associated with spontaneously occurring HPFH variants, including C>T;c.-167, C>T;c.-167, C>A;c.-175, T>C;c.-197, C>T;c.-200+C;c.-202, C>G;c.-211, C>T;c.-228, T>C;c.-255, C>G;c.-309, A>G;c.-369, C>G; or c.-567 T>G.

[0109] Spontaneous mutations in the distal CCAAT box motif within the promoters of the HBG1 and / or HBG2 genes (i.e., HBG1 / 2 c.-111~-115) have also been shown to result in continued γ-globin expression and the pathological symptoms of HPFH. Modifications (mutations or deletions) of the CCAAT box are thought to interfere with the binding of one or more transcriptional repressors, potentially leading to continued γ-globin gene expression and elevated HbF expression (Martyn 2017). For example, spontaneous 13-base pair del c.-114~-102 ("13nt deletion") has been shown to be associated with elevated HbF levels (Martyn 2017). The distal CCAAT box is likely to overlap with the binding motifs within and around the CCAAT box of negative regulatory transcription factors that are expressed in adulthood and repress HBG (Martyn 2017).

[0110] The gene editing strategies disclosed herein involve increasing HbF expression by interfering with one or more nucleotides within and / or around the distal CCAAT box. In certain embodiments, the “CCAAT box target region” may be a region of the distal CCAAT box or a region in its vicinity, comprising the nucleotides of the distal CCAAT box and 25 nucleotides upstream (5') and downstream (3') of the distal CCAAT box (i.e., HBG1 / 2 c.-86 to -140). In other embodiments, the “CCAAT box target region” may be a region of the distal CCAAT box or a region in its vicinity, comprising the nucleotides of the distal CCAAT box and 5 nucleotides upstream (5') and downstream (3') of the distal CCAAT box (i.e., HBG1 / 2 c.-106 to -120). Without limitation, the following intrinsic, non-spontaneous modifications of the CCAAT box target region that induce HBG expression are disclosed herein, including HBG del c.-104 to -121 ("18nt deletion"), HBG del c.-105 to -115 ("11nt deletion"), HBG del c.-112 to -115 ("4nt deletion"), and HBG del c.-116 ("1nt deletion"). In certain embodiments, the modifications may be introduced into the CCAAT box target region of HBG1 and / or HBG2 using the genome editing systems disclosed herein. In certain embodiments, the genome editing system may include one or more DNA donor templates that encode the modification (such as deletion, insertion, or mutation) of the CCAAT box target region. In certain embodiments, the modification may be a non-spontaneous or spontaneous modification. In certain embodiments, the donor template may encode a 1nt deletion, a 4nt deletion, an 11nt deletion, a 13nt deletion, an 18nt deletion, or a c.-117 G>A modification. In certain embodiments, the genome editing system may include RNA-inducing nucleases, including Cas9, modified Cas9, Cpf1, or modified Cpf1. In certain embodiments, the genome editing system may include an RNP containing gRNA and a Cpf1 molecule.In certain embodiments, the gRNA may be unmodified or modified, and the Cpf1 molecule may be wild-type Cpf1 protein, modified Cpf1 protein, or a combination thereof. In certain embodiments, the gRNA may contain sequences listed in Table 13, Table 18, or Table 19. In certain embodiments, modified Cpf1 may be encoded by sequences listed in SEQ ID NOs: 1000, 1001, 1008-1018, 1032, 1035-39, 1094-1097, 1107-09 (Cpf1 polypeptide sequences) or SEQ ID NOs: 1019-1021, 1110-17 (Cpf1 polynucleotide sequences). In certain embodiments, the RNP complex may contain RNP complexes listed in Table 21. For example, the RNP complex may contain a gRNA containing the sequence listed in SEQ ID NO: 1051 and a modified Cpf1 protein encoded by the sequence listed in SEQ ID NO: 1097 (RNP32 in Table 21).

[0111] HbF expression can also be induced through targeted interference of erythrocyte-specific expression of BCL11A, a transcriptional repressor that encodes repressors that terminate the expression of HBG1 and HBG2 (Canvers 2015). Another gene editing strategy disclosed herein is to increase HbF expression by targeted interference of the erythrocyte-specific enhancer of BCL11A (BCL11Ae) (discussed in its entirety in International Publication 2015 / 148860 by Friedland et al. ("Friedland"), published on 1 October 2015 and transferred to the assignee of the present invention, which is incorporated herein by reference). In certain embodiments, the region of BCL11Ae targeted for interference may be the GATA1 binding motif in BCL11Ae. In certain embodiments, modifications can be introduced to the GATA1 binding motif, the CCAAT box target region, the 13nt target region of HBG1 and / or HBG2, or a combination thereof, using the genome editing system disclosed herein.

[0112] The genome editing systems of this disclosure may include, but are not limited to, RNA-induced nucleases such as Cas9 or Cpf1, one or more gRNAs having a targeting domain complementary to a sequence in or near a target region, and one or more DNA donor templates that selectively encode specific mutations (such as deletions or insertions) in or near a target region, and / or agents that enhance the efficiency of such mutations, including random oligonucleotides, small molecule agonists or antagonists or peptides of gene products involved in DNA repair or DNA damage responses.

[0113] Various approaches to introducing mutations into the CCAAT box target region, the 13nt target region, the proximal HBG1 / 2 promoter target sequence, and / or the GATA1 binding motif in BCL11Ae may be used in embodiments of this disclosure. One approach involves generating a single modification, such as a double-strand break, within the CCAAT box target region, the 13nt target region, the proximal HBG1 / 2 promoter target sequence, and / or the GATA1 binding motif in BCL11Ae, which is then repaired in a manner that disrupts the function of the region, for example, by the formation of an indel or the incorporation of a donor template sequence that encodes the deletion of the region. A second approach involves generating two or more modifications on either side of the region, resulting in the deletion of intervening sequences, including the CCAAT box target region, the 13nt target region, and the GATA1 binding motif in BCL11Ae.

[0114] The treatment of abnormal hemoglobin disorders by gene therapy and / or genome editing is complicated by the fact that the cells phenotypically affected by the disease, namely erythrocytes or RBCs, are enucleated and do not contain the genetic material encoding either the abnormal hemoglobin protein (Hb) subunit or the Aγ or Gγ subunit that is targeted in the exemplary genome editing approaches described above. This complicated situation is addressed in certain embodiments of the present disclosure by modifying cells that are capable of differentiating into erythrocytes or otherwise capable of producing erythrocytes. Cells within erythrocyte lineages that can be modified according to various embodiments of the present disclosure include, but are not limited to, hematopoietic stem and progenitor cells (HSCs), erythroblasts (including basophilic, polychromatic, and / or orthochromatic erythroblasts), proerythroblasts, polychromatic erythrocytes or reticulocytes, embryonic stem (ES) cells and / or induced pluripotent stem (iPSC) cells. These cells can be modified in situ (e.g., within the tissue of interest) or in vitro. The implementation forms of genome editing systems for in-situ and in vitro cell modification are described below under the heading "Implementation Forms of Genome Editing Systems: Delivery, Formulation, and Administration Routes."

[0115] In certain embodiments, modifications resulting in the induction of Aγ and / or Gγ expression are obtained through the use of a genome editing system comprising an RNA-inducing nuclease and at least one gRNA having a targeting domain complementary to a sequence within or adjacent to the CCAAT box target region of HBG1 and / or HBG2 (e.g., within 10, 20, 30, 40 or 50, 100, 200, 300, 400 or 500 bases of the CCAAT box target region). As will be discussed in more detail below, the RNA-inducing nuclease and gRNA form a complex that can bind to and modify the CCAAT box target region or a region adjacent thereto. Examples of suitable gRNAs and gRNA targeting domains directed to the CCAAT box target region or adjacent region of HBG1 and / or HBG2 for use in the embodiments disclosed herein include, but are not limited to, those described in SEQ ID NOs: 251-901, 940-942, 970, 971, 996, 997, 1002 and 1004.

[0116] In certain embodiments, modifications resulting in the induction of Aγ and / or Gγ expression are obtained through the use of a genome editing system comprising an RNA-inducing nuclease and at least one gRNA having a targeting domain complementary to a sequence within or adjacent to the 13nt target region of HBG1 and / or HBG2 (e.g., within 10, 20, 30, 40 or 50, 100, 200, 300, 400 or 500 bases of the 13nt target region). As will be discussed in more detail below, the RNA-inducing nuclease and gRNA form a complex that can bind to and modify the 13nt target region or adjacent region. Examples of suitable gRNAs and gRNA targeting domains directed to the 13nt target region or adjacent region of HBG1 and / or HBG2 for use in the embodiments disclosed herein include, but are not limited to, those described in SEQ ID NOs: 251-901, 940-942, 970, 971, 996, 997, 1002 and 1004.

[0117] In certain embodiments, modifications resulting in the induction of HbF expression are obtained through the use of a genome editing system comprising an RNA-inducing nuclease and at least one gRNA having a targeting domain complementary to a sequence within or adjacent to the GATA1 binding motif in BCL11Ae (e.g., within 10, 20, 30, 40 or 50, 100, 200, 300, 400 or 500 bases of the GATA1 binding motif in BCL11Ae). In certain embodiments, the RNA-inducing nuclease and the gRNA form a complex that can bind to and modify the GATA1 binding motif in BCL11Ae. Examples of suitable targeting domains directed to the GATA1 binding motif in BCL11Ae for use in the embodiments disclosed herein include, but are not limited to, those described in SEQ ID NOs. 952-955.

[0118] Genome editing systems can be implemented in various ways, as will be discussed in detail below. As an example, the genome editing system of this disclosure may be implemented as a ribonucleoprotein complex or a group of complexes in which multiple gRNAs are used. This ribonucleoprotein complex can be introduced into target cells using methods known in the art, including electroporation, as described in International Publication No. 2016 / 182959, co-transferred by Jennifer Gori (“Gori”), published on 17 November 2016, which is incorporated herein by reference in its entirety.

[0119] The ribonucleoprotein complexes in these compositions are introduced into target cells by methods known in the art, including, but not limited to, electroporation (e.g., using Nucleofection®, a technology commercialized by Lonza, Basel, Switzerland, or similar technologies commercialized by MaxCyte Inc., Gaithersburg, Maryland, etc.) and lipofection (e.g., using the Lipofectamine® reagent commercialized by Thermo Fisher Scientific, Waltham Massachusetts). Alternatively, or in addition, the ribonucleoprotein complexes are formed within the target cells themselves following the introduction of RNA-inducible nucleases and / or nucleic acids encoding gRNA. These and other delivery modes are described below in general terms and Gori.

[0120] Cells modified in vitro in accordance with this disclosure may be manipulated before their delivery to a target (e.g., proliferation, passage, freezing, differentiation, dedifferentiation, transduction by transgenes, etc.). The cells are delivered in various ways: to a target from which they are derived ("autologous" transplantation) or to a recipient immunologically different from the cell donor ("allogeneic" transplantation).

[0121] In some cases, autologous transplantation includes the steps of obtaining multiple cells from a subject that are circulating in the peripheral blood or within the bone marrow or other tissue (e.g., spleen, skin, etc.), and enriching these cells with cells within the erythrocyte lineage by manipulating them (e.g., by induction to generate iPSCs, purification of cells expressing specific cell surface markers such as CD34, CD90, CD49f, and / or purification of cells that do not express surface markers characteristic of non-erythrocyte lineages such as CD10, CD14, CD38). The cells are optionally or additionally grown before transduction by a genome editing system targeting the CCAAT box target region, the 13nt target region, the proximal HBG1 / 2 promoter target sequence and / or the GATA1 binding motif in BCL11Ae, transduced with the transgene, exposed to cytokines or other peptides or small molecule drugs, and / or frozen / thawed. Genome editing systems can be implemented or delivered to cells in any suitable form, such as as ribonucleoprotein complexes, as isolated protein and nucleic acid components, and / or as nucleic acids encoding components of the genome editing system.

[0122] In certain embodiments, CD34+ hematopoietic stem and progenitor cells (HSPCs) edited using the genome editing methods disclosed herein may be used for the treatment of abnormal hemoglobinopathy in subjects requiring it. In certain embodiments, abnormal hemoglobinopathy may be severe sickle cell disease (SCD) or thalassemia such as β-thalassemia, δ-thalassemia, or β / δ-thalassemia. In certain embodiments, an exemplary protocol for the treatment of abnormal hemoglobinopathy may include harvesting CD34+ HSPCs from a subject requiring it, ex vivo editing of autologous CD34+ HSPCs using the genome editing methods disclosed herein, and subsequently reinfusion of the edited autologous CD34+ HSPCs into the subject. In certain embodiments, treatment with edited autologous CD34+ HSPCs may result in increased HbF induction.

[0123] In certain embodiments, prior to the collection of CD34+HSPCs, the subject may discontinue hydroxyurea treatment, if applicable, and receive blood transfusions to maintain adequate hemoglobin (Hb) levels. In certain embodiments, the subject may be intravenously administered prelixafor (e.g., 0.24 mg / kg) to mobilize CD34+HSPCs from the bone marrow into the peripheral blood. In certain embodiments, the subject may undergo one or more leukocyte apheresis cycles (e.g., with approximately one month between cycles, and one cycle defined as two prelixafor-mobilized leukocyte apheresis collections performed on consecutive days). In certain embodiments, the number of leukocyte apheresis cycles performed on a subject may be the number required to achieve a dose of edited autologous CD34+HSPC (e.g., ≥1.5 × 10⁶ cells / kg) for reinfusion to the subject, along with a dose of unedited autologous CD34+HSPC / kg for backup storage (e.g., ≥2 × 10⁶ cells / kg, ≥3 × 10⁶ cells / kg, ≥4 × 10⁶ cells / kg, ≥5 × 10⁶ cells / kg, 2 × 10⁶ cells / kg to 3 × 10⁶ cells / kg, 3 × 10⁶ cells / kg to 4 × 10⁶ cells / kg, 4 × 10⁶ cells / kg to 5 × 10⁶ cells / kg), along with a dose of unedited autologous CD34+HSPC / kg for backup storage. In certain embodiments, CD34+HSPCs taken from the subject may be edited using any of the genome editing methods discussed herein. In certain embodiments, any one or more gRNAs and one or more RNA-inducing nucleases disclosed herein may be used in the genome editing method.

[0124] In certain embodiments, the treatment may include autologous stem cell transplantation. In certain embodiments, the subject may undergo myeloablative habituation by busulfan habituation (e.g., dose-adjusted based on initial dose pharmacokinetic analysis at a test dose of 1 mg / kg). In certain embodiments, habituation may be carried out over 4 consecutive days. In certain embodiments, after a 3-day busulfan-free period, edited autologous CD34+HSPCs (e.g., ≥2×10⁶ cells / kg, ≥3×10⁶ cells / kg, ≥4×10⁶ cells / kg, ≥5×10⁶ cells / kg, 2×10⁶ cells / kg to 3×10⁶ cells / kg, 3×10⁶ cells / kg to 4×10⁶ cells / kg, 4×10⁶ cells / kg to 5×10⁶ cells / kg) may be reinfused into the subject (e.g., into peripheral blood). In certain embodiments, edited autologous CD34+HSPCs may be manufactured for a specific subject and cryopreserved. In certain embodiments, subjects may achieve neutrophil transplantation following a series of myeloablative adaptation regimens and infusions of edited autologous CD34+ cells. Neutrophil transplantation may be defined as three consecutive measurements of ANC at ≥0.5 × 10⁹ / L.

[0125] Regardless of how they are implemented, genome editing systems may include, or be delivered together with, one or more factors that improve cell viability during and after editing, including, but not limited to, StemRegenin-1 (SR1), UM171, LGC0006, α-naphthoflavone and CH-223191, as well as aryl hydrocarbon receptor antagonists and / or cyclosporine A, dexamethasone, resveratrol, MyD88 inhibitory peptides, RNAi-targeted Myd88, B18R recombinant protein, glucocorticoids, OxPAPC, TLR antagonists, rapamycin, BX795 and RLR shRNA, and other innate immune response antagonists. These and other factors that improve cell viability during and after editing are described under the heading "I. Optimization of Stem Cells" on pages 36-61 of Gori, which is incorporated herein by reference.

[0126] Following the delivery of the genome editing system, cells are selectively manipulated, for example, to enrich HSCs and / or cells within the erythrocyte lineage, and / or edited cells, and then the cells are prepared in another way for proliferation, freezing / thawing, or returning to the target. The edited cells are then returned to the target, for example, by intravenous delivery or other means of delivery, into the circulatory system or into solid tissue such as bone marrow.

[0127] Functionally, modification of the CCAAT box target region, 13nt target region, proximal HBG1 / 2 promoter target sequence and / or GATA1 binding motif in BCL11Ae using the compositions, methods and genome editing systems of this disclosure results in significant induction of Aγ and / or Gγ subunits (synonymously referred to as HbF expression) among hemoglobin-expressing cells, such as induction of Aγ and / or Gγ subunit expression by at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, or 50% or more compared to an unmodified control. The induction of this protein expression is generally the result of modifications to the CCAAT box target region, 13nt target region, proximal HBG1 / 2 promoter target sequence, and / or GATA1 binding motif in BCL11Ae (expressed as a percentage of the whole genome, including indel mutations in multiple cells) in at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50% of multiple cells, or any portion or all of multiple cells, including, for example, at least one allele containing a modification including an indel, insertion, or deletion within or near the GATA1 binding motif in the CCAAT box target region, 13nt target region, proximal HBG1 / 2 promoter target sequence, and / or GATA1 binding motif in BCL11Ae.

[0128] The functional effects of modifications caused or facilitated by the genome editing systems and methods disclosed herein can be evaluated by any number of suitable methods. For example, the effect of modifications on the expression of fetal hemoglobin can be evaluated at the protein or mRNA level. The expression of HBG1 and HBG2 mRNA can be evaluated by digital droplet PCR (ddPCR) performed on cDNA samples obtained by reverse transcription of mRNA taken from treated or untreated samples. Primers for HBG1, HBG2, HBB, and / or HBA can be used individually or in multiplexing using methods known in the art. For example, ddPCR analysis of samples can be performed using the QX200™ ddPCR system commercialized by BioRad (Hercules, CA) and related protocols published by Bio-Rad. Fetal hemoglobin proteins can be evaluated, for example, by high-pressure liquid chromatography (HPLC) or high-performance protein liquid chromatography (FPLC) using ion-exchange columns and / or reverse-phase columns that separate HbF, HbB, and HbA and / or Aγ and Gγ globin chains, as is known in the Art, according to the methods discussed on pages 143-44 of Chang 2017 (incorporated herein by reference).

[0129] It should be noted that the rate at which the CCAAT box target region (e.g., 18nt, 11nt, 4nt, 1nt, c.-117 G>A target region), 13nt target region, proximal HBG1 / 2 promoter target sequence, and / or GATA1 binding motif in BCL11Ae are modified in target cells can be modified by the use of optional genome editing system components such as oligonucleotide donor templates. Donor template design is described under the heading "Donor Template Design" using the following general terms. Examples of donor templates for use in targeting 13nt target regions include, but are not limited to, donor templates encoding modifications (e.g., deletions) of HBG1 c.-114~-102 (corresponding to nucleotides 2824~2836 of SEQ ID NO: 902), HBG1 c.-225~-222 (corresponding to nucleotides 2716~2719 of SEQ ID NO: 902), and / or HBG2 c.-114~-102 (corresponding to nucleotides 2748~2760 of SEQ ID NO: 903). Exemplary full-length donor templates encoding exemplary 5' and 3' homology arms and deletions such as c.-114~-102 are also presented below (SEQ ID NOs: 904~909). In certain embodiments, donor templates for use in targeting 18nt target regions may include, without limitation, donor templates encoding modifications (e.g., deletions) of HBG1 c.-104~-121, HBG2 c.-104~-121, or combinations thereof. Exemplary full-length donor templates encoding deletions such as c.-104~-121 include sequence numbers 974 and 975. In certain embodiments, donor templates for use in targeting 11nt target regions may include, without limitation, donor templates encoding modifications (e.g., deletions) of HBG1 c.-105~-115, HBG2 c.-105~-115, or combinations thereof. Exemplary full-length donor templates encoding deletions such as c.-105~-115 include sequence numbers 976 and 978.In certain embodiments, donor templates for use in targeting 4nt target regions may include, without limitation, donor templates encoding modifications (e.g., deletions) of HBG1 c.-112~-115, HBG2 c.-112~-115, or combinations thereof. Exemplary full-length donor templates encoding deletions such as c.-112~-115 include sequence numbers 984~995. In certain embodiments, donor templates for use in targeting 1nt target regions may include, without limitation, donor templates encoding modifications (e.g., deletions) of HBG1 c.-116, HBG2 c.-116, or combinations thereof. Exemplary full-length donor templates encoding deletions such as c.-116 include sequence numbers 982 and 983. In certain embodiments, donor templates for use in targeting the c.-117 G>A target region may include, without limitation, donor templates encoding modifications (e.g., deletions) of HBG1 c.-117 G>A, HBG2 c.-117 G>A, or combinations thereof. Exemplary full-length donor templates encoding deletions such as c.-117 G>A include sequence numbers 980 and 981. In certain embodiments, the donor template may be a positive or negative chain.

[0130] The donor templates used herein may be non-specific templates that are non-homologous to a region of DNA within or near the target sequence. In certain embodiments, donor templates for use in targeting a 13nt target region may include, without limitation, non-target-specific templates that are non-homologous to a region of DNA within or near the 13nt target region. For example, a non-specific donor template for use in targeting a 13nt target region may be non-homologous to a region of DNA within or near the 13nt target region and may include a donor template that encodes a deletion in HBG1 c.-225~-222 (corresponding to nucleotides 2716~2719 of SEQ ID NO: 902). In certain embodiments, donor templates for use in targeting a GATA1 binding motif in BCL11Ae may include, without limitation, non-target-specific templates that are non-homologous to a region of DNA within or near the GATA1 binding motif target sequence in BCL11Ae. Other donor templates for use in targeting BCL11Ae include, without limitation, donor templates containing modifications (e.g., deletions) of BCL11Ae, including the GATA1 motif in BCL11Ae.

[0131] The embodiments described herein may be used in all classes of vertebrates, including but not limited to primates, mice, rats, rabbits, pigs, dogs, and cats.

[0132] This summary focuses on a few exemplary embodiments that describe the principles of genome editing systems and CRISPR-mediated methods for modifying cells. However, for clarity, this disclosure includes modifications and variations that are not expressly covered above but would be apparent to those skilled in the art. With this in mind, the following disclosure is intended to illustrate the operating principles of genome editing systems more generally. The following content should not be understood as limiting, but rather as illustrating specific principles of genome editing systems and CRISPR-mediated methods utilizing these systems, which, in combination with this disclosure, will inform those skilled in the art of additional implementations and modifications within its scope.

[0133] RNA-inducing helicase, guide RNA, and dead guide RNA Various embodiments of this disclosure generally relate to genome editing systems, methods, and compositions configured to enhance genome editing of target regions in nucleic acids (e.g., CCAAT box target regions, 13nt target regions, proximal HBG1 / 2 promoter target sequences, and / or GATA1 binding motifs in BCL11Ae) by modifying the helical structure of nucleic acids. Many embodiments relate to the observation that arranging events that modify the helical structure of DNA within or adjacent to a target region in nucleic acids may improve the activity of genome editing systems directed to such target regions. While we do not wish to be constrained by any theory, it is thought that modification of the helical structure within or proximal to a DNA target region (e.g., by unwinding) may induce or increase the accessibility of the genome editing system to the target region, thereby leading to increased editing of the target region by the genome editing system.

[0134] CRISPR nucleases have primarily evolved to defend bacteria against viral pathogens whose genomes are not naturally organized into chromatin. In contrast, eukaryotic genomes are organized into nucleosome units containing genomic DNA segments wrapped around histones. CRISPR nucleases from several bacteriaceae have been found to be inactive in eukaryotic DNA editing, suggesting that the ability to edit nucleosome-bound DNA may vary among enzymes (Ran 2015). Biochemical evidence shows that Cas9 from S. pyogenes can efficiently cleave DNA at the edges of nucleosomes but exhibits reduced activity when the target site is located near the center of the nucleosome dyad (Hinz 2016).

[0135] In many cell types, the target site of interest may either be strongly bound by the nucleosome or only possess adjacent PAMs of enzymes that are not efficiently edited in the presence of the nucleosome. In this case, the problematic nucleosome can be initially replaced by using adjacent target sites that are closer to the nucleosome rim or bound by enzymes that are more effective at binding nucleosomal DNA. However, cleavage at these adjacent sites can be detrimental to the therapeutic strategy. Therefore, a programmable enzyme that binds to these adjacent sites but does not cleave them would enable more efficient functional editing.

[0136] Related strategies utilize the recruitment of exogenous transacting factors to promote nucleosome substitution. However, the systems and methods of this disclosure are advantageous because they do not require gRNA modification beyond targeting domain truncation, do not require the recruitment of exogenous transacting factors, and do not require transcriptional activation to achieve increased editing rates.

[0137] Various approaches for unwinding and modifying nucleic acids are employed in various embodiments of this disclosure. One approach involves unwinding (or opening) a chromatin segment within or proximal to a target region of a nucleic acid in a cell (e.g., a CCAAT box target region, a 13nt target region, a proximal HBG1 / 2 promoter target sequence, and / or a GATA1 binding motif of BCL11Ae) to generate a double-strand break (DSB) within the target region of the nucleic acid, thereby modifying the target region. In certain embodiments, the DSB may be repaired in a manner that modifies the target region. Unwinding a chromatin segment using the methods provided herein may facilitate increased access to chromatin by catalytically active RNPs (e.g., catalytically active RNA-induced nucleases and gRNAs) and enable more efficient editing of DNA. For example, these methods may be used to edit a target region in chromatin that is difficult for ribonucleoproteins (e.g., RNA-induced nucleases complexed with gRNA) to access because the chromatin is occupied by nucleosomes such as closed chromatin. In certain embodiments, chromatin segment unwinding occurs via RNA-induced helicase activity. In certain embodiments, the unwinding step does not require the recruitment of an exogenous trans-acting factor to the chromatin segment. In certain embodiments, the chromatin segment unwinding step does not involve the formation of single-strand or double-strand breaks in the nucleic acids within the chromatin segment.

[0138] In certain embodiments of the above approaches and methods, modification of the DNA helical structure is achieved through the action of an "RNA-induced helicase," a term generally used to refer to a molecule, typically a peptide, that (a) interacts with (e.g., complexes with) gRNA and (b) binds to a target site together with the gRNA and unwinds it. In certain embodiments, the RNA-induced helicase may include an RNA-induced nuclease configured to lack nuclease activity. However, we have observed that even RNA-induced nucleases with cleavage ability can be adapted for use as RNA-induced helicases by complexing with dead gRNA having a truncated targeting domain of nucleotides 15 or less in length. The complex of dead gRNA with a wild-type RNA-induced nuclease shows reduced or absent RNA cleavage activity but appears to maintain helicase activity. RNA-induced helicases and dead gRNAs are described in more detail below.

[0139] With respect to RNA-induced helicases, according to this disclosure, RNA-induced helicases may include, but are not limited to, any of the RNA-induced nucleases disclosed herein and hereafter under the heading “RNA-induced nucleases,” including Cas9 or Cpf1 RNA-induced nucleases. The helicase activity of these RNA-induced nucleases enables DNA unwinding and provides increased access of genome editing system components (e.g., catalytically active RNA-induced nucleases and gRNAs) to desired target regions to be edited (e.g., CCAAT box target regions, 13nt target regions, proximal HBG1 / 2 promoter target sequences and / or GATA1 binding motifs in BCL11Ae). In certain embodiments, the RNA-induced nuclease may be a catalytically active RNA-induced nuclease having nuclease activity. In certain embodiments, the RNA-induced helicase may be configured to lack nuclease activity. For example, in certain embodiments, the RNA-induced helicase may be a non-catalyzed RNA-induced nuclease lacking nuclease activity, such as a catalytically inactive Cas9 molecule, which still provides helicase activity. In certain embodiments, the RNA-induced helicase may complex with dead gRNA to form dead RNPs that cannot cleave nucleic acids. In other embodiments, the RNA-induced helicase may be a catalytically active RNA-induced nuclease that complexes with dead gRNA to form dead RNPs that cannot cleave nucleic acids. In certain embodiments, the RNA-induced nuclease is not configured to recruit an exogenous trans-acting factor to a desired target region to be edited (e.g., a CCAAT box target region, a 13nt target region, a proximal HBG1 / 2 promoter target sequence, and / or a GATA1 binding motif in BCL11Ae).

[0140] When referring to dead gRNAs, these include any of the dead gRNAs discussed herein and below under the heading “Dead gRNA Molecules.” Dead gRNAs (also referred to herein as “dgRNAs”) can be generated by truncating the 5' end of a gRNA targeting domain sequence to yield a targeting domain sequence of 15 nucleotides or less in length. In certain embodiments, dgRNAs can be generated by truncating the 5' end of any one of the gRNA targeting domain sequences disclosed in Table 2 or Table 10 herein. Dead guide RNA molecules according to this disclosure include dead guide RNA molecules having reduced, low, or undetectable cleavage activity. The dead guide RNA targeting domain sequence may be 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides shorter in length than the targeting domain sequence of the active guide RNA. The dead gRNA molecule may contain a targeting domain complementary to a region proximal to or within a target region in the target nucleic acid (e.g., the CCAAT box target region, the 13nt target region, the proximal HBG1 / 2 promoter target sequence, and / or the GATA1 binding motif in BCL11Ae). In certain embodiments, "proximal" may refer to a region within 10, 25, 50, 100, or 200 nucleotides of the target region (e.g., the CCAAT box target region, the 13nt target region, the proximal HBG1 / 2 promoter target sequence, and / or the GATA1 binding motif in BCL11Ae). In certain embodiments, the dead gRNA contains a targeting domain complementary to the transcription or non-transcription strand of DNA. In certain embodiments, the dead guide RNA is not configured to recruit an exogenous trans-acting factor to the target region (e.g., the CCAAT box target region, the 13nt target region, the proximal HBG1 / 2 promoter target sequence, and / or the GATA1 binding motif in BCL11Ae).

[0141] Also provided herein is a method for increasing the rate of indel formation in a target nucleic acid by unwinding DNA within or proximal to a target region (e.g., a CCAAT box target region, a 13nt target region, a proximal HBG1 / 2 promoter target sequence, and / or a GATA1 binding motif in BCL11Ae) using an RNA-induced helicase to generate a double-segment break (DSB) within the target region and forming an indel within the target region through DSB repair. The step of unwinding DNA using an RNA-induced helicase provides increased indel formation compared to indel formation methods that do not use a helicase.

[0142] This disclosure further encompasses a method for deleting a segment of a target nucleic acid within a cell, comprising the step of contacting the cell with an RNA-induced helicase to induce a double-strand break (DSB) within a target region (e.g., a CCAAT box target region, a 13nt target region, a proximal HBG1 / 2 promoter target sequence, and / or a GATA1 binding motif in BCL11Ae). In certain embodiments, the RNA-induced helicase is configured to bind within or proximal to a target region of the target nucleic acid and to unwind double-stranded DNA (dsDNA) within or proximal to the target region. In certain embodiments, the target nucleic acid is a promoter region of a gene, a coding region of a gene, a non-coding region of a gene, an intron of a gene, or an exon of a gene. In certain embodiments, the segment of the target nucleic acid to be deleted may be at least about 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, or 100 base pairs in length. In certain embodiments, the DSB is repaired by deleting a segment of the target nucleic acid.

[0143] Genome editing systems configured to introduce modifications to a helical structure can be implemented in various ways, as will be discussed in detail below. As an example, the genome editing systems of this disclosure may be implemented as a ribonucleoprotein complex or a plurality of complexes, in which a plurality of gRNAs are used. In certain embodiments, the ribonucleoprotein complex of the genome editing system may be an RNA-induced helicase complexed with a death guide RNA. The ribonucleoprotein complex may be introduced into target cells using methods known in the art, including electroporation, as described in Gori. Genome editing systems incorporating RNA-induced helicases may be modified in any appropriate manner, including, but not limited to, the incorporation of one or more DNA donor templates encoding specific mutations (such as deletions or insertions) within or near a target region, and / or, but not limited to, agents that enhance the efficiency of such mutations, including random oligonucleotides, small molecule agonists or antagonists or peptides of gene products involved in DNA repair or DNA damage responses. These modifications will be described in more detail below under the heading "Genome Editing Strategies". For clarity, this disclosure includes compositions comprising one or more gRNAs, dead gRNAs, RNA-induced helicases, RNA-induced nucleases, or combinations thereof.

[0144] While some of the exemplary embodiments described above focus on DNA unwinding, it should be noted that other helical modifications are also within the scope of this disclosure. These include, but are not limited to, overwinding, underwinding, increased or decreased torsional strain (e.g., through topoisomerase activity) on DNA strands within or adjacent to a target region, denaturation or strand separation, and / or other appropriate modifications resulting in chromatin structure modification. Each of these modifications may be catalyzed by RNA-inducing activity or recruitment of endogenous factors to a target region.

[0145] Genome editing systems and methods for modifying one or more indels (e.g., indel signatures) generated by an active guide are also provided herein. As the inventors have discovered herein, pairing of a dead RNP (dRNP) (i.e., a dead guide RNA complexed with an RNA-inducing nuclease) with an active RNP (i.e., an active guide RNA complexed with an RNA-inducing nuclease) can result in a change in the orientation of indels (e.g., indel signatures) generated by active RNPs alone (without dRNP). As shown in the following examples, the use of dead guide RNA can result in an increased frequency of larger deletions extending from the active guide RNA cleavage site toward the dead guide RNA binding site. Thus, deletion editing can be effectively "directed" toward a desired target site using dead guide RNA. In certain embodiments, the use of dead guide RNA in combination with active guide RNA can increase the frequency of deletions that are not related to microhomology.

[0146] While the examples disclosed in the following section on examples focus on the modification of CCAAT box target regions, those skilled in the art will envision that the genome editing systems, methods, cells, and compositions described herein may be used to modify any other target regions, for example, to increase the frequency of deletions in target regions, and to increase the frequency of deletions in non-microhomology-related target regions (e.g., those not repaired via MMEJ).

[0147] This summary focuses on a few exemplary embodiments that describe the principles of genome editing systems and CRISPR-mediated methods for modifying cells. However, for clarity, this disclosure includes modifications and variations that are not expressly covered above but would be apparent to those skilled in the art. With this in mind, the following disclosure is intended to illustrate the operating principles of genome editing systems more generally. The following content should not be understood as limiting, but rather as illustrating specific principles of genome editing systems and CRISPR-mediated methods utilizing these systems, which, in combination with this disclosure, will inform those skilled in the art of additional implementations and modifications within its scope.

[0148] genome editing system The term "genome editing system" refers to any system having RNA-induced DNA editing activity. The genome editing systems of this disclosure include at least two components, a guide RNA (gRNA) and an RNA-induced nuclease, adapted from a naturally occurring CRISPR system. These two components bind to a specific nucleic acid sequence to form a complex capable of editing DNA in or around the nucleic acid sequence by, for example, generating one or more single-strand breaks (SSBs or nicks), double-strand breaks (DSCs), and / or point mutations.

[0149] In certain embodiments, the genome editing system of this disclosure may include a helicase for unwinding DNA. In certain embodiments, the helicase may be an RNA-inducing helicase. In certain embodiments, the RNA-inducing helicase may be an RNA-inducing nuclease as described herein, such as a Cas9 or Cpf1 molecule. In certain embodiments, the RNA-inducing nuclease is not configured to recruit an exogenous trans-acting factor to the target region. In certain embodiments, the RNA-inducing nuclease may be configured to lack nuclease activity. In certain embodiments, the RNA-inducing helicase may complex with a death guide RNA as disclosed herein. For example, the death guide RNA (dgRNA) may include a targeting domain sequence less than 15 nucleotides in length. In certain embodiments, the death guide RNA is not configured to recruit an exogenous trans-acting factor to the target region.

[0150] Naturally occurring CRISPR systems are evolutionarily organized into two classes and five types (Makarova 2011, incorporated herein by reference), and while the genome editing systems of this disclosure may be adapted from components of any type or class of naturally occurring CRISPR system, the embodiments presented herein are generally adapted from Class 2 and Type II or Type V CRISPR systems. Class 2 systems, encompassing Types II and V, are characterized by a relatively large multidomain RNA-induced nuclease protein (e.g., Cas9 or Cpf1) and one or more guide RNAs (e.g., crRNA, optionally tracrRNA) that form a ribonucleoprotein (RNP) complex that associates with (targets) and cleaves a specific locus complementary to the target (or spacer) sequence of the crRNA. The genome editing systems of this disclosure similarly target and edit cellular DNA sequences, but differ significantly from naturally occurring CRISPR systems. For example, the single-molecule guide RNA described herein does not exist in nature, and both the guide RNA and RNA-inducing nuclease described herein may incorporate any number of modifications that do not exist in nature.

[0151] Genome editing systems can be implemented in various ways (e.g., they can be administered or delivered to cells or subjects), and different implementations may be suitable for different applications. For example, in certain embodiments, a genome editing system may be implemented as a protein / RNA complex (ribonucleoprotein or RNP), which may be included in a pharmaceutical composition that optionally includes a pharmaceutically acceptable carrier and / or encapsulant, but is not limited to lipid or polymer microparticles or nanoparticles, micelles or liposomes. In certain embodiments, a genome editing system may be implemented as one or more nucleic acids (optionally with one or more additional components) encoding the RNA-inducing nuclease and guide RNA components described above; in certain embodiments, a genome editing system may be implemented as one or more vectors containing such nucleic acids, such as a viral vector, such as an adeno-associated virus (see the following sections under the heading "Implementations of Genome Editing Systems: Delivery, Formulation and Route of Administration"); in certain embodiments, a genome editing system may be implemented as any combination of the foregoing. Additional or modified implementations operating in accordance with the principles described herein will be apparent to those skilled in the art and are within the scope of this disclosure.

[0152] It should be noted that the genome editing systems of this disclosure can or may target a single specific nucleotide sequence, and that by using two or more guide RNAs, two or more specific nucleotide sequences can be edited in parallel. The use of multiple gRNAs is referred to as “multiplexing” throughout this disclosure and may be used to target multiple unrelated target sequences of interest or to form multiple SSBs or DSBs within a single target domain, and, if applicable, to perform specific edits within such target domains. For example, International Publication No. 2015 / 138510 by Maeder et al. ("Maeder"), incorporated herein by reference, describes a genome editing system for correcting a point mutation in the human CEP290 gene (C.2991+1655A to G) that results in the generation of a potential splice site, thereby reducing or eliminating gene function. Maeder’s genome editing system utilizes two guide RNAs that target sequences on both sides (i.e., flanking) the point mutation to form a DSB located on the mutated side. This then promotes the deletion of the intervening sequence containing the mutation, thereby removing the potential splice site and restoring normal gene function.

[0153] As another example, Cotta-Ramusino et al. ("Cotta-Ramusino"), International Publication No. 2016 / 073990, incorporated herein by reference, describes a genome editing system that utilizes two gRNAs in combination with Cas9 nickase (Cas9 that produces single-stranded nicks such as S. pyogenes D10A), in a configuration referred to as the "double nickase system." The Cotta-Ramusino double nickase system is configured to produce two nicks on the reverse strand of a target sequence offset by one or more nucleotides, and these nicks combine to create a double-strand break with an overhang (5' in this case, but a 3' overhang is also possible). The overhang can then facilitate homology-directed repair events in several situations. As another example, Palestrant et al., International Publication No. 2015 / 070083 (incorporated herein by reference), describes a gRNA (referred to as “control RNA”) that targets a nucleotide sequence encoding Cas9, which may be included in a genome editing system comprising one or more additional gRNAs to enable transient expression of Cas9, which would otherwise be constitutively expressed, in some transgenic cells. These multiplexing applications are intended to be illustrative rather than restrictive, and those skilled in the art will understand that other multiplexing applications are generally compatible with the genome editing systems described herein.

[0154] As disclosed herein, in certain embodiments, the genome editing system may comprise a plurality of gRNAs that can be used to introduce mutations into the GATA1 binding motif in BCL11Ae or the 13nt target region of HBG1 and / or HBG2.

[0155] Genome editing systems may, in some cases, create double-strand breaks that are repaired by cellular DNA double-strand break mechanisms such as NHEJ or HDR. These mechanisms are described in the literature (see, for example, Davis & Maizels 2014 (describes Alt-HDR); Frit 2014 (describes Alt-NHEJ); and Iyama & Wilson 2013 (generally describes the standard HDR and NHEJ pathways)).

[0156] When a genome editing system functions by forming double-strand breaks (DSBs), such a system may optionally include one or more components that promote or facilitate a particular type of double-strand break repair or a particular repair outcome. For example, Cotta-Ramusino also describes a genome editing system in which a single-strand oligonucleotide "donor template" is attached; the donor template can be incorporated into a target region of cellular DNA to be cleaved by the genome editing system, resulting in a change in the target sequence.

[0157] In certain embodiments, genome editing systems modify a target sequence or alter the expression of a gene in or near a target sequence without causing single-strand or double-strand breaks. For example, a genome editing system may include an RNA-inducing nuclease fused to a functional domain that acts on DNA, thereby modifying the target sequence or its expression. As an example, an RNA-inducing nuclease may bind (e.g., fuse) to a cytidine deaminase functional domain and function by generating a targeted C-to-A substitution. An exemplary nuclease / deaminase fusion is described in Komor 2016, which is incorporated herein by reference. Alternatively, a genome editing system may utilize a cleavage-inactivated (i.e., “dead”) nuclease, such as dead Cas9 (dCas9), which forms a stable complex on one or more targeted regions of cellular DNA, thereby functioning by interfering with functions involving the target region, including, but not limited to, mRNA transcription and chromatin remodeling. In certain embodiments, a genome editing system may include an RNA-induced helicase that unwinds DNA within or near a target sequence without causing single-strand or double-strand breaks. For example, a genome editing system may include an RNA-induced helicase configured to bind within or near a target sequence to unwind the DNA and induce access to the target sequence. In certain embodiments, the RNA-induced helicase may complex with a dead guide RNA configured to lack cleavage activity, allowing the DNA to be unwound without causing DNA breaks.

[0158] Guide RNA (gRNA) molecule The terms “guide RNA” and “gRNA” refer to any nucleic acid that facilitates the specific binding (or “targeting”) of an RNA-induced nuclease, such as Cas9 or Cpf1, to a target sequence, such as a cellular genome or episome sequence. gRNAs can be monomolecular (containing a single RNA molecule, also referred to as a chimera) or modular (containing two or more, typically two separate RNA molecules, such as crRNA and tracrRNA, which usually bind to each other by duplication). gRNAs and their components are described throughout the literature (see, for example, Briner 2014; Cotta-Ramusino, cited by reference). Examples of modular and monomolecular gRNAs that may be used according to embodiments of this specification include, but are not limited to, the sequences described in SEQ ID NOs. 29–31 and 38–51. Examples of gRNA proximal and tail domains that may be used according to embodiments of this specification include, but are not limited to, the sequences described in SEQ ID NOs. 32–37.

[0159] In bacteria and archaea, the type II CRISPR system generally includes an RNA-inducing nuclease protein such as Cas9; CRISPR RNA (crRNA) containing a 5' region complementary to the exogenous sequence; and transactivating crRNA (tracrRNA) containing a 5' region complementary to the 3' region of the crRNA and forming a double helix with it. While not intended to be bound by any theory, this double helix is ​​thought to promote the formation of the Cas9 / gRNA complex and is necessary for its activity. While adapting the type II CRISPR system for use in gene editing, in one non-limiting example, it was discovered that crRNA and tracrRNA can be linked to a single monomolecule or chimeric guide RNA by a 4-nucleotide (e.g., GAAA) "tetraloop" or "linker" sequence that bridges the complementary regions of crRNA (its 3' end) and tracrRNA (its 5' end). (Mali 2013; Jiang 2013; Jinek 2012; all incorporated herein by reference).

[0160] Guide RNA, whether monomolecular or modular, contains a “targeting domain” that is fully or partially complementary to the target domain in the target sequence, such as a DNA sequence in the genome of the cell to be edited. The targeting domain is referred to by various names in the literature, including, but is not limited to, “guide sequence” (Hsu 2013, incorporated herein by reference), “complementary region” (Cotta-Ramusino), “spacer” (Briner 2014), and comprehensively “crRNA” (Jiang). Regardless of the name given to it, the targeting domain is typically 10–30 nucleotides long, and in certain embodiments 16–24 nucleotides long (e.g., 16, 17, 18, 19, 20, 21, 22, 23, or 24 nucleotides long), located at or near the 5' end in the case of Cas9 gRNA, and at or near the 3' end in the case of Cpf1 gRNA.

[0161] In addition to the targeting domain, gRNA typically (though not always, as discussed below) contains multiple domains that can influence the formation or activity of the gRNA / Cas9 complex. For example, as described above, the double-stranded structure formed by the first and second complementary domains (also called repeats: anti-repeat double helixes) of the gRNA can interact with the Cas9 recognition (REC) lobe to mediate the formation of the Cas9 / gRNA complex (Nishimasu 2014; Nishimasu 2015; both incorporated herein by reference). It should be noted that the first and / or second complementary domains may contain one or more poly(A) strands that can be recognized as termination signals by RNA polymerase. Therefore, the sequences of the first and second complementary domains can be selectively modified, for example, through the use of AG swaps or AU swaps as described in Briner 2014, to remove these regions and promote complete in vitro transcription of the gRNA. These and other similar modifications to the first and second complementary domains are within the scope of this disclosure.

[0162] Along with the first and second complementary domains, Cas9g RNA typically contains two or more additional double-stranded regions that are involved in nuclease activity in vivo but not necessarily in vitro (Nishimasu 2015). The first stem-loop 1, located near the 3' portion of the second complementary domain, is referred to in various ways, including "proximal domain" (Cotta-Ramusino), "stem-loop 1" (Nishimasu 2014 and 2015), and "nexus" (Briner 2014). One or more additional stem-loop structures are generally located near the 3' end of the gRNA, and their number varies by species. S. pyogenes gRNA typically contains two 3' stem-loop structures (a total of four stem-loop structures, including repeat / anti-repeat double helixes), while S. aureus and other species have only one (a total of three stem-loop structures). A description of conserved stem-loop structures (and more generally, gRNA structures) grouped by species is provided in Briner 2014.

[0163] While the above explanation has focused on gRNAs for use with Cas9, it should be understood that there are other RNA-inducible nucleases that utilize gRNAs that differ in some respects from those described so far. For example, Cpf1 (CRISPR from "Prevotella and Francicella 1") is a recently discovered RNA-inducible nuclease that does not require tracrRNA to function (Zetsche 2015, incorporated herein by reference). gRNAs for use in the Cpf1 genome editing system generally contain a targeting domain and a complementary domain (referred to instead as a "handle"). It should also be noted that in gRNAs for use with Cpf1, the targeting domain is usually located at or near the 3' end, rather than at the 5' end, as described above for Cas9 gRNA (the handle is at or near the 5' end of Cpf1 gRNA). Exemplary targeting domains of Cpf1 gRNAs are listed in Tables 13 and 18.

[0164] However, those skilled in the art will understand that while structural differences may exist between gRNAs from different prokaryotic species or between Cpf1 and Cas9 gRNAs, the principle by which gRNAs operate is generally consistent. Due to this consistency of operation, gRNAs can be defined in a broad sense by their targeting domain sequences, and those skilled in the art will understand that a given targeting domain sequence can be incorporated into any suitable gRNA, including monomolecular or chimeric gRNAs or gRNAs containing one or more chemical modifications and / or sequence modifications (substitutions, additional nucleotides, cleavage, etc.). Therefore, for the economy of presentation in this disclosure, gRNAs may be described only in terms of their targeting domain sequences.

[0165] More generally, those skilled in the art will understand that some aspects of this disclosure relate to systems, methods, and compositions that can be implemented using multiple RNA-inducing nucleases. For this reason, unless otherwise specified, the term gRNA should be understood to encompass not only gRNAs compatible with specific species of Cas9 or Cpf1, but also any suitable gRNAs that can be used with any RNA-inducing nuclease. For example, in certain embodiments, the term gRNA may include gRNAs for use with, or derived from or adapted from, a Class 2 CRISPR system such as Type II or Type V, or any RNA-inducing nuclease present in a CRISPR system.

[0166] gRNA design Methods for selecting and validating target sequences and for off-target analysis have been previously described (see, e.g., Mali 2013; Hsu 2013; Fu 2014; Heigwer 2014; Bae 2014; Xiao 2014). Each of these references is incorporated herein by reference. As a non-limiting example, gRNA design may involve the use of software tools to optimize the selection of potential target sequences corresponding to the user's target sequence in order to minimize total off-target activity across the genome. Off-target activity is not limited to cleavage, but the cleavage efficiency at each off-target sequence can be predicted, for example, using experimentally derived weighting schemes. These and other guided selection methods are described in detail by Maeder and Cotta-Ramusino.

[0167] For the selection of gRNA targeting domain sequences directed towards the HBG1 / 2 target site (e.g., the 13nt target region), an in silico gRNA targeting domain identification tool was used, and hits were stratified into four layers. In S. pyogenes, the first layer targeting domains were selected based on (1) the distance upstream or downstream from any end of the target site (i.e., the HBG1 / 2 13nt target region), specifically within 400 bp of any end of the target site, (2) a high level of orthogonality, and (3) the presence of 5'G. The second layer targeting domains were selected based on (1) the distance upstream or downstream from any end of the target site (i.e., the HBG1 / 2 13nt target region), specifically within 400 bp of any end of the target site, and (2) a high level of orthogonality. The third layer targeting domain was selected based on (1) the distance upstream or downstream from any end of the target site (i.e., the HBG1 / 2 13nt target region), specifically within 400 bp of any end of the target site, and (2) the presence of 5'G. The fourth layer targeting domain was selected based on the distance upstream or downstream from any end of the target site (i.e., the HBG1 / 2 13nt target region), specifically within 400 bp of any end of the target site.

[0168] In S. aureus, the first layer targeting domain was selected based on (1) the distance upstream or downstream from any end of the target site (i.e., the HBG1 / 2 13nt target region), specifically within 400 bp of any end of the target site, (2) a high level of orthogonality, (3) the presence of a 5'G, and (4) the presence of an NNGRRT sequence (sequence number 204) in the PAM. The second layer targeting domain was selected based on (1) the distance upstream or downstream from any end of the target site (i.e., the HBG1 / 2 13nt target), specifically within 400 bp of any end of the target site, (2) a high level of orthogonality, and (3) the presence of an NNGRRT sequence (sequence number 204) in the PAM. The third layer targeting domain was selected based on (1) the upstream or downstream distance from any end of the target site (i.e., the HBG1 / 2 13nt target region), specifically within 400 bp of any end of the target site, and (2) the PAM having the NNGRRT sequence (SEQ ID NO: 204). The fourth layer targeting domain was selected based on (1) the upstream or downstream distance from any end of the target site (i.e., the HBG1 / 2 13nt target), specifically within 400 bp of any end of the target site, and (2) the PAM having the NNGRRV sequence (SEQ ID NO: 205).

[0169] Table 2 below shows the targeting domains of gRNAs of S. pyogenes and S. aureus, classified by (a) layer (1st, 2nd, 3rd, or 4th) and (b) HBG1 or HBG2.

[0170] [Table 2]

[0171] Additional gRNA sequences designed to target modifications to the CCAAT box target region include, but are not limited to, sequences described in SEQ ID NOs: 970 and 971. Examples of targeting domain sequences for gRNAs designed to target interference with the CCAAT box target region include, but are not limited to, sequence number 1002. Examples of targeting domain sequences plus PAM (UUUG) for gRNAs designed to target interference with the CCAAT box target region include, but are not limited to, sequence number 1004. In certain embodiments, gRNAs containing sequences described in SEQ ID NOs: 1002 and / or 1004 may complex with the Cpf1 protein or modified Cpf1 protein to produce modifications to the CCAAT box target region. In certain embodiments, gRNAs containing any of the Cpf1 gRNAs described in Table 15, Table 18, or Table 19 may complex with the Cpf1 protein or modified Cpf1 protein to form an RNP ("gRNA-Cpf1-RNP") to produce modifications to the CCAAT box target region. In certain embodiments, the modified Cpf1 protein may be His-AsCpf1-nNLS (SEQ ID NO: 1000) or His-AsCpf1-sNLS-sNLS (SEQ ID NO: 1001). In certain embodiments, the Cpf1 molecule of gRNA-Cpf1-RNP may be encoded by sequences described in SEQ ID NOs: 1000, 1001, 1008-1018, 1032, 1035-39 (Cpf1 polypeptide sequences) or SEQ ID NOs: 1019-1021 (Cpf1 polynucleotide sequences).

[0172] The gRNA may be designed to target the erythrocyte-specific enhancer of BCL11A (BCL11Ae) to disrupt the expression of the transcriptional repressor BCL11A (as described by Friedland, as incorporated herein by reference). The gRNA is designed to target the GATA1 binding motif in the erythrocyte-specific enhancer of BCL11A (i.e., the GATA1 binding motif in BCL11Ae) located in the +58 DHS region of intron 2, the +58 DHS enhancer region containing the sequence described in SEQ ID NO: 968. Targeting domain sequences of gRNA designed to target disruption of the GATA1 binding motif in BCL11Ae include, but are not limited to, the sequences described in SEQ ID NOs: 952-955. Targeting domain sequences plus PAM (NGG) of gRNA designed to target disruption of the GATA1 binding motif in BCL11Ae include, but are not limited to, the sequences described in SEQ ID NOs: 960-963.

[0173] gRNA modification The activity, stability, or other characteristics of gRNAs can be altered by incorporating specific modifications. For example, transiently expressed or delivered nucleic acids may be susceptible to degradation by cellular nucleases, for instance. Therefore, the gRNAs described herein may contain one or more modified nucleosides or nucleotides that introduce stability against nucleases. While we do not wish to impose theoretical constraints, it is also thought that certain modified gRNAs described herein may, upon introduction into cells, exhibit a reduction in innate immune responses. Those skilled in the art are aware of certain cellular responses commonly observed in cells, such as mammalian cells, in response to exogenous nucleic acids, particularly those of viral or bacterial origin. Such responses, which may include induction of cytokine expression and release and cell death, can be reduced or completely eliminated by the modifications presented herein.

[0174] The specific exemplary modifications discussed in this section may be located at any position in the gRNA sequence, such as the 5' end or its vicinity (e.g., within 1–10, 1–5, or 1–2 nucleotides of the 5' end) and / or the 3' end or its vicinity (e.g., within 1–10, 1–5, or 1–2 nucleotides of the 3' end). In some cases, the modifications are located within functional motifs such as repeat:antirepeat double helixes of Cas9 gRNA, stem-loop structures of Cas9 or Cpf1 gRNA, and / or targeting domains of gRNA.

[0175] As an example, the 5' end of a gRNA may contain a eukaryotic mRNA cap structure or G cap analogue (e.g., G(5')ppp(5')G cap analogue, m7G(5')ppp(5')G cap analogue, or 3'-O-Me-m7G(5')ppp(5')G anti-cap analogue (ARCA)) as shown below. [ka] Caps or cap analogues may be present during the chemical synthesis or in vitro transcription of gRNA.

[0176] A similar approach can result in the absence of a 5' triphosphate group at the 5' end of a gRNA. For example, the 5' triphosphate group can be removed by phosphatase treatment of an in vitro transcribed gRNA (e.g., using calf intestinal alkaline phosphatase).

[0177] Another common modification involves adding a series of adenine (A) residues (e.g., 1–10, 10–20, or 25–200) to the 3' end of the gRNA, known as a poly-A tract. The poly-A sequence can be added to the gRNA by a polyadenylation sequence either during chemosynthesis following in vitro transcription using a polyadenosine polymerase (e.g., E. coli poly(A) polymerase) or in vivo, as described by Maeder.

[0178] It should be noted that the modifications described herein may be combined in any suitable manner, for example, gRNA transcribed from a DNA vector in vivo or in vitro may contain either or both a 5' cap structure or a cap analogue and a 3' polyA sequence.

[0179] Guide RNA can be modified with 3' terminal U-ribose. For example, the two terminal hydroxyl groups of U-ribose can be oxidized to aldehyde groups, and simultaneous ring-opening of the ribose ring results in the modified nucleoside shown below. [ka] In the formula, "U" may be unmodified or modified uridine.

[0180] The 3'-terminal U-ribose may be modified with a 2'3'-cyclic phosphate ester, as shown below. [ka] In the formula, "U" may be unmodified or modified uridine.

[0181] The guide RNA may contain a 3' nucleotide that can be stabilized against degradation by incorporating, for example, one or more of the modified nucleotides described herein. In certain embodiments, uridine may be substituted with modified uridines such as, for example, 5-(2-amino)propyluridine and 5-bromouridine or any of the modified uridines described herein; adenosine and guanosine may be substituted with modified adenosine and guanosine with, for example, an 8-modified position such as 8-bromoguanosine, or any of the modified adenosine or guanosines described herein.

[0182] In certain embodiments, sugar-modified ribonucleotides may be incorporated into gRNA, for example, the 2'OH group may be substituted with a group selected from H, -OR, -R (wherein R may be, for example, alkyl, cycloalkyl, aryl, aralkyl, heteroaryl, or sugar), halo, -SH, -SR (wherein R may be, for example, alkyl, cycloalkyl, aryl, aralkyl, heteroaryl, or sugar), amino (wherein amino may be, for example, NH2; alkylamino, dialkylamino, heterocyclyl, arylamino, diarylamino, heteroarylamino, diheteroarylamino, or amino acid); or cyano(-CN). In certain embodiments, the phosphate backbone may be modified, for example, with a phosphorothioate (PhTx) group, as described herein. In certain embodiments, one or more nucleotides of the gRNA may independently be modified or unmodified nucleotides, including, but not limited to, 2'-sugar modifications such as 2'-O-methyl, 2'-O-methoxyethyl, or 2'-fluoro modifications such as 2'-F or 2'-O-methyl, adenosine (A), 2'-F or 2'-O-methyl, cytidine (C), 2'-F or 2'-O-methyl, uridine (U), 2'-F or 2'-O-methyl, thymidine (T)1, 2'-F or 2'-O-methyl, guanosine (G), 2'-O-methoxyethyl-5-methyluridine (Teo), 2'-O-methoxyethyladenosine (Aeo), 2'-O-methoxyethyl-5-methylcytidine (m5Ce0), and any combination thereof.

[0183] Guide gRNA may also include "locked" nucleic acid (LNA), where the 2'OH group may be linked, for example, by a C1-6 alkylene or C1-6 heteroalkylene crosslink to the 4' carbon of the same ribose sugar. To provide such crosslinks, any suitable part may be used, but is not limited to methylene, propylene, ether or amino crosslinks; O-amino (wherein amino may be, for example, NH2; alkylamino, dialkylamino, heterocyclyl, arylamino, diarylamino, heteroarylamino or diheteroarylamino, ethylenediamine or polyamino) and aminoalkoxy or O(CH2)n-amino (wherein amino may be, for example, NH2; alkylamino, dialkylamino, heterocyclyl, arylamino, diarylamino, heteroarylamino or diheteroarylamino, ethylenediamine or polyamino).

[0184] In certain embodiments, gRNA may include polycyclic modified nucleotides (e.g., tricyclonucleotides) and "unlocked" forms such as glycol nucleic acids (GNAs) (e.g., R-GNAs or S-GNAs, in which ribose is replaced by a glycol unit attached to a phosphate diester bond) or threose nucleic acids (TNAs, in which ribose is replaced by α-L-treophranosyl-(3'→2')).

[0185] Generally, gRNAs contain ribose, a five-membered ring sugar containing oxygen. Typical modified gRNAs include, but are not limited to, substitution of oxygen in ribose (e.g., by sulfur (S), selenium (Se), or alkylenes such as methylene or ethylene); addition of a double bond (e.g., substitution of ribose with cyclopentenyl or cyclohexenyl); ribose ring contraction (e.g., forming a four-membered ring of cyclobutane or oxetane); and ribose ring expansion (e.g., forming a six- or seven-membered ring with additional carbon or heteroatoms, such as anhydrohexitol, althritol, mannitol, cyclohexanyl, cyclohexenyl, and morpholino, which also has a phosphoramidate skeleton). Most sugar analog modifications are localized at the 2' position, but other sites, including the 4' position, are also receptive to modification. In certain embodiments, gRNAs include 4'-S, 4'-Se, or 4'-C-aminomethyl-2'-O-Me modifications.

[0186] In certain embodiments, deazanucleotides, such as 7-deaza-adenosine, may be incorporated into the gRNA. In certain embodiments, O- and N-alkylated nucleotides, such as N6-methyladenosine, may be incorporated into the gRNA. In certain embodiments, one, several, or all nucleotides in the gRNA are deoxyribonucleotides.

[0187] In certain embodiments, the gRNA may be modified or unmodified gRNA as used herein. In certain embodiments, the gRNA may have one or more modifications. In certain embodiments, one or more modifications may include phosphorothioate-binding modifications, phosphorodithioate (PS2)-binding modifications, 2'-O-methyl modifications, or a combination thereof. In certain embodiments, one or more modifications may be located at the 5' end of the gRNA, the 3' end of the gRNA, or a combination thereof.

[0188] In certain embodiments, the gRNA modification may include one or more phosphorodithioate (PS2) binding modifications.

[0189] In some embodiments, the gRNA used herein includes one or more or a sequence of deoxyribonucleic acid (DNA) bases, also referred herein as “DNA extension.” In some embodiments, the gRNA used herein includes a DNA extension at the 5' end of the gRNA, the 3' end of the gRNA, or a combination thereof. In certain embodiments, the DNA extension is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51 The DNA base length may be 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100. For example, in certain embodiments, the DNA extension may be 1, 2, 3, 4, 5, 10, 15, 20, or 25 DNA base lengths. In certain embodiments, the DNA extension may contain one or more DNA bases selected from adenine (A), guanine (G), cytosine (C), or thymine (T). In certain embodiments, the DNA extension may contain the same DNA base. For example, a DNA extension may include a sequence of adenine (A) bases. In certain embodiments, a DNA extension may include a sequence of thymine (T) bases. In certain embodiments, a DNA extension may include a combination of different DNA bases. In certain embodiments, a DNA extension may include sequences listed in Table 24. For example, a DNA extension may include sequences listed in Sequence IDs 1235-1250. In certain embodiments, the gRNA used herein includes a DNA extension and one or more phosphorothioate-binding modifications, one or more phosphorodithioate (PS2)-binding modifications, one or more 2'-O-methyl modifications, or a combination thereof. In certain embodiments, one or more modifications may be located at the 5' end of the gRNA, the 3' end of the gRNA, or a combination thereof.In certain embodiments, the gRNA containing DNA extension may include the sequences listed in Table 19 that contain DNA extension. In certain embodiments, the gRNA containing DNA extension may include the sequence listed in SEQ ID NO: 1051. In certain embodiments, the gRNA containing DNA extension may include sequences selected from the group consisting of SEQ ID NOs: 1046-1060, 1067, 1068, 1074, 1075, 1078, 1081-1084, 1086-1087, 1089-1090, 1092-1093, 1098-1102, and 1106. While we do not wish to impose theoretical constraints, it is assumed that any DNA extension may be used herein, as long as it does not hybridize with the target nucleic acid targeted by the gRNA and shows increased editing at the target nucleic acid site compared to gRNA without such DNA extension.

[0190] In some embodiments, the gRNA used herein includes one or more or a sequence of ribonucleic acid (RNA) bases, also referred herein as “RNA extension.” In some embodiments, the gRNA used herein includes the RNA extension at the 5' end of the gRNA, the 3' end of the gRNA, or a combination thereof. In certain embodiments, the RNA extension is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51 The RNA length may be 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100 RNA base lengths. For example, in certain embodiments, the RNA extension may be 1, 2, 3, 4, 5, 10, 15, 20, or 25 RNA base lengths. In certain embodiments, the RNA extension may comprise one or more RNA bases selected from adenine (rA), guanine (rG), cytosine (rC), or uracil (rU), where "r" represents RNA, 2'-hydroxyl. In certain embodiments, the RNA extension comprises identical RNA bases. For example, the RNA extension may comprise a sequence of adenine (rA) bases. In certain embodiments, the RNA extension comprises a combination of different RNA bases. In certain embodiments, the RNA extension may comprise sequences listed in Table 24. For example, the RNA extension may comprise sequences listed in 1231-1234 and 1251-1253. In certain embodiments, the gRNA used herein comprises the RNA extension and one or more phosphorothioate-binding modifications, one or more phosphorodithioate (PS2)-binding modifications, one or more 2'-O-methyl modifications, or a combination thereof. In certain embodiments, one or more modifications may be located at the 5' end of the gRNA, the 3' end of the gRNA, or a combination thereof.In certain embodiments, gRNAs containing RNA extensions may include sequences listed in Table 19 that contain RNA extensions. GRNAs containing RNA extensions at their 5' end may include sequences selected from the group consisting of SEQ ID NOs. 1042-1045 and 1103-1105. GRNAs containing RNA extensions at their 3' end may include sequences selected from the group consisting of SEQ ID NOs. 1070-1075, 1079, 1081 and 1098-1100.

[0191] The gRNAs used herein may include RNA extensions and DNA extensions. In certain embodiments, both the RNA extension and the DNA extension may be located at the 5' end of the gRNA, the 3' end of the gRNA, or a combination thereof. In certain embodiments, the RNA extension may be at the 5' end of the gRNA and the DNA extension may be at the 3' end of the gRNA. In certain embodiments, the RNA extension may be at the 3' end of the gRNA and the DNA extension may be at the 5' end of the gRNA.

[0192] In some embodiments, gRNAs containing both 3'-terminal phosphorothioate modification and 5'-terminal DNA extension are complexed with an RNA-induced nuclease, such as Cpf1, to form an RNP, which is then used to edit hematopoietic stem cells (HSCs) or CD34+ cells in vitro (i.e., in vitro from the target from which such cells originate) at the HBG locus.

[0193] In the use of this specification, an example of a gRNA includes the sequence described in SEQ ID NO: 1051.

[0194] Dead gRNA molecules Examples of death guide RNA (dgRNA) molecules according to this disclosure include death guide RNA molecules having reduced, low, or undetectable cleavage activity. The death guide RNA targeting domain sequence is 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides shorter in length than the targeting domain sequence of the active guide RNA. In certain embodiments, the death guide RNA molecule may include a targeting domain having a length of 15 nucleotides or less, 14 nucleotides or less, 13 nucleotides or less, 12 nucleotides or less, or 11 nucleotides or less. In some embodiments, the death guide RNAs are configured so that they do not provide an RNA-induced nuclease cleavage event. The death guide RNA can be generated by removing the 5' end of a gRNA targeting domain sequence, resulting in a truncated targeting domain sequence. For example, if a gRNA sequence configured to provide a cleavage event (i.e., 17 nucleotides or longer) has a targeting domain sequence that is 20 nucleotides long, the death guide RNA can be generated by removing 5 nucleotides from the 5' end of the gRNA sequence. For example, the dgRNAs used herein may include a targeting domain as described in Table 2 or Table 10, which is truncated from the 5' end of a gRNA sequence and contains no more than 15 nucleotides in length. In certain embodiments, the dgRNA may be configured to bind (or associate) with a nucleic acid sequence within or proximal to the target region to be edited (e.g., the CCAAT box target region, the 13nt target region, the proximal HBG1 / 2 promoter target sequence and / or the GATA1 binding motif in BCL11Ae). For example, any of the dgRNAs described in Table 10 may be used to bind to a nucleic acid sequence proximal to the 13nt target region or the CCAAT box target region. In certain embodiments, "proximal" may refer to a region of 10, 25, 50, 100, or 200 nucleotides within the target region (e.g., the CCAAT box target region, the 13nt target region, the proximal HBG1 / 2 promoter target sequence and / or the GATA1 binding motif in BCL11Ae). In certain embodiments, the death guide RNA is not configured to recruit an exogenous trans-acting factor to the target region.In certain embodiments, dgRNA is configured such that it does not produce a DNA cleavage event when complexed with an RNA-guided nuclease. A person skilled in the art will understand that a dead guide RNA molecule can be designed to comprise a targeting domain complementary to a region proximal to or within the target region of a target nucleic acid. In certain embodiments, the dead guide RNA comprises a targeting domain sequence complementary to the transcribed or non-transcribed strand of double-stranded DNA. The dgRNA herein can comprise modifications at the 5' and 3' ends of the gRNA, as described for guide RNAs in the "gRNA modifications" section herein. For example, in certain embodiments, the dead guide RNA can comprise an anti-reverse cap analog (ARCA) at the 5' end of the RNA. In certain embodiments, the dgRNA can comprise a poly A tail at the 3' end.

[0195] In certain embodiments, use of dead guide RNA in the genome editing systems and methods disclosed herein can increase the overall editing level of active guide RNA. In certain embodiments, use of dead guide RNA in the genome editing systems and methods disclosed herein can increase the frequency of deletions. In certain embodiments, the deletion can extend from the cleavage site of the active guide RNA toward the dead guide RNA binding site. In this way, the dead guide RNA can alter the directionality of the active guide RNA and direct editing toward a desired target region.

[0196] As used herein, the terms "dead gRNA" and "truncated gRNA" are used interchangeably.

[0197] RNA-guided nuclease RNA-guided nucleases according to the present disclosure include, but are not limited to, naturally occurring class 2 CRISPR nucleases such as Cas9 and Cpf1, and other nucleases derived or obtained therefrom. In addition, certain RNA-guided nucleases such as Cas9 have also been shown to have helicase activity that allows unwinding of nucleic acids. In certain embodiments, an RNA-guided helicase according to the present disclosure can be any of the RNA-nucleases described herein and in the section entitled "RNA-guided nuclease" above. In certain embodiments, the RNA-guided nuclease is not configured to recruit exogenous trans-acting factors to a target region. In certain embodiments, the RNA-guided helicase can be an RNA-guided nuclease configured to lack nuclease activity. For example, in certain embodiments, the RNA-guided helicase can be a catalytically inactive RNA-guided nuclease that lacks nuclease activity but still retains its helicase activity. In certain embodiments, the RNA-guided nuclease can be mutated to abolish its nuclease activity (e.g., dead Cas9), generating a catalytically inactive RNA-guided nuclease that cannot cleave a nucleic acid but can still unwind DNA. In certain embodiments, it can be an RNA-guided helicase complexed with any of the dead guide RNAs as described herein. For example, a catalytically active RNA-guided helicase (e.g., Cas9 or Cpf1) can form an RNP complex with a dead guide RNA, resulting in a catalytically inactive dead RNP (dRNP). In certain embodiments, a catalytically inactive RNA-guided helicase (e.g., dead Cas9) and a dead guide RNA can form a dRNP. These dRNPs are unable to provide a cleavage event, but still retain helicase activity that is important for unwinding of nucleic acids.

[0198] Functionally, RNA-inducible nucleases are defined as nucleases that (a) interact with gRNA (e.g., form a complex); and (b) together with gRNA, bind to (i) a sequence complementary to the targeting domain of gRNA, and optionally (ii) an additional sequence referred to as a “protospacer adjacent motif” or “PAM,” as described in more detail below, and optionally cleave or modify it. RNA-inducible nucleases can be defined in a broad sense by their PAM specificity and cleavage activity, even though variations may exist among individual RNA-inducible nucleases sharing the same PAM specificity or cleavage activity, as illustrated by the following examples. Those skilled in the art will understand that some aspects of this disclosure relate to systems, methods, and compositions that can be implemented using any suitable RNA-inducible nuclease having a particular PAM specificity and / or cleavage activity. For this reason, unless otherwise specified, the term RNA-induced nuclease should be understood as a general term and not limited to any particular type of RNA-induced nuclease (e.g., Cpf1 in comparison to Cas9), species (e.g., S. aureus in comparison to S. pyogenes), or variation (e.g., truncated or split forms in comparison to full length; manipulated PAM specificity in comparison to natural PAM specificity, etc.).

[0199] Various RNA-induced nucleases may require different sequence relationships between the PAM and the protospacer. Generally, Cas9 recognizes the PAM sequence at the 3' position of the protospacer, while Cpf1 generally recognizes the PAM sequence at the 5' position of the protospacer.

[0200] In addition to recognizing the orientation of specific PAM and protospacer sequences, RNA-induced nucleases can also recognize specific PAM sequences. For example, S. aureus Cas9 recognizes the NNGRRT or NNGRRV PAM sequence, where the N residue is adjacent to the 3' region recognized by the gRNA targeting domain. S. pyogenes Cas9 recognizes the NGG PAM sequence. Also, F. novicida Cpf1 recognizes the TTN PAM sequence. PAM sequences have been identified for various RNA-induced nucleases, and strategies for identifying novel PAM sequences have been described by Shmakov 2015. It should also be noted that the engineered RNA-inducible nuclease may have a PAM specificity different from that of the reference molecule (for example, in the case of an engineered RNA-inducible nuclease, the reference molecule may be a naturally occurring mutant from which the RNA-inducible nuclease is derived, or a naturally occurring mutant that has maximum amino acid sequence homology with the engineered RNA-inducible nuclease). Examples of PAMs that may be used according to the embodiments of this specification include, but are not limited to, the sequences described in SEQ ID NOs: 199-205.

[0201] In addition to their PAM specificity, RNA-inducible nucleases can be characterized by their DNA cleavage activity. While native RNA-inducible nucleases typically form double-sided subunits (DSBs) in target nucleic acids, engineered mutants have been created that produce only single-sided subunits (SSBs) or do not cleave at all (as discussed above and in Ran & Hsu 2013, which are incorporated herein by reference).

[0202] Cas9 The crystal structures of S. pyogenes Cas9 (Jinek 2014) and S. aureus Cas9, which form complexes with single-molecule guide RNA and target DNA, have been determined (Nishimasu 2014; Anders 2014 and Nishimasu 2015).

[0203] The spontaneously occurring Cas9 protein contains two lobes: a recognition (REC) lobe and a nuclease (NUC) lobe, each containing a specific structural and / or functional domain. The REC lobe contains an arginine-rich cross-linking helix (BH) domain and at least one REC domain (e.g., the REC1 domain, optionally the REC2 domain). The REC lobe does not share structural similarities with other known proteins, suggesting it is a unique functional domain. While we do not wish to be constrained by any theory, mutational analysis suggests specific functional roles for the BH and REC domains. The BH domain appears to play a role in gRNA:DNA recognition, while the REC domain is thought to interact with the repeat:anti-repeat double helix of gRNA, mediating the formation of the Cas9 / gRNA complex.

[0204] The NUC lobe comprises a RuvC domain, an HNH domain, and a PAM interaction (PI) domain. The RuvC domain shares structural similarities with members of the retroviral integrase superfamily and cleaves the non-complementary (i.e., lower) strand of the target nucleic acid. It can be formed from two or more split RuvC motifs (e.g., RuvCI, RuvCII, and RuvCIII in S. pyogenes and S. aureus). The HNH domain, on the other hand, is structurally similar to the HNN endonuclease motif and cleaves the complementary (i.e., upper) strand of the target nucleic acid. As its name suggests, the PI domain contributes to PAM specificity. Examples of polypeptide sequences encoding Cas9 RuvC-like and Cas9 HNH-like domains that may be used according to embodiments herein are described in SEQ ID NOs: 15-23, 52-123 (RuvC-like domain) and SEQ ID NOs: 24-28, 124-198 (HNH-like domain).

[0205] While certain functions of Cas9 are associated with the specific domains described above (though not necessarily entirely determined by them), these and other functions may be mediated or influenced by other Cas9 domains or multiple domains on any of the lobes. For example, in S. pyogenes Cas9, as described by Nishimasu 2014, the repeat:antirepeat double helix of the gRNA enters the groove between the REC and NUC lobes, and the nucleotides in the double helix interact with amino acids in the BH, PI, and REC domains. Some nucleotides in the first stem-loop structure also interact with amino acids in multiple domains (PI, BH, and REC1), as do some nucleotides in the second and third stem-loops (RuvC and PI domains). Examples of polypeptide sequences encoding Cas9 molecules that may be used according to embodiments herein are described in SEQ ID NOs: 1-2, 4-6, 12, and 14.

[0206] Cpf1 The crystal structure of Cpf1, a species of Acidaminococcus, when complexed with crRNA and double-stranded (ds)DNA targets such as the TTTN PAM sequence, was analyzed by Yamano 2016 (incorporated herein by reference). Similar to Cas9, Cpf1 has two lobes: a REC (recognition) lobe and a NUC (nuclease) lobe. The REC lobe contains REC1 and REC2 domains, which are not similar to any known protein structure. The NUC lobe, on the other hand, contains three RuvC domains (RuvC-I, -II, and -III) and one BH domain. However, in contrast to Cas9, the Cpf1 REC lobe lacks an HNH domain and is a structurally unique domain that does not resemble any known protein structure, containing three wedge (WED) domains (WED-I, -II, and -III) and a nuclease (Nuc) domain.

[0207] While Cas9 and Cpf1 share structural and functional similarities, it should be understood that certain Cpf1 activities are mediated by structural domains that are not similar to either Cas9 domain. For example, cleavage of the complementary strand of target DNA appears to be mediated by a Nuc domain that is sequence- and spatially distinct from the HNH domain of Cas9. Furthermore, the untargeting region (handle) of Cpf1 gRNA adopts a pseudo-knot structure rather than a stem-loop structure formed by the repeat:anti-repeat double helix in Cas9 gRNA.

[0208] In certain embodiments, the Cpf1 protein may be a modified Cpf1 protein. In certain embodiments, the modified Cpf1 protein may comprise one or more modifications. In certain embodiments, the modifications may, without limitation, be one or more mutations in the Cpf1 nucleotide sequence or Cpf1 amino acid sequence, one or more additional sequences such as a His tag or nuclear localization signal (NLS), or a combination thereof. In certain embodiments, modified Cpf1 may also be referred to herein as a Cpf1 variant.

[0209] In certain embodiments, the Cpf1 protein may be derived from a Cpf1 protein selected from the group consisting of Acidaminococcus strain BV3L6 Cpf1 protein (AsCpf1), Lachnospiraceae bacterium ND2006 Cpf1 protein (LbCpf1), and Lachnospiraceae bacterium MA2020 (Lb2Cpf1). In certain embodiments, the Cpf1 protein may include a sequence selected from the group consisting of SEQ ID NOs. 1016-1018, each having a codon-optimized nucleic acid sequence of SEQ ID NOs. 1019-1021.

[0210] In certain embodiments, the modified Cpf1 protein may include a nuclear localization signal (NLS). For example, but not limited to, NLS sequences useful in relation to the methods and compositions disclosed herein would include amino acid sequences that can facilitate protein translocation into the cell nucleus. NLS sequences useful in relation to the methods and compositions disclosed herein are known in the art. Examples of such NLS sequences include the nucleoplasmin NLS having the amino acid sequence: KRPAATKKAGQAKKKK (SEQ ID NO: 1006) and the Simian virus 40 "SV40" NLS having the amino acid sequence PKKKRKV (SEQ ID NO: 1007).

[0211] In certain embodiments, the NLS sequence of the modified Cpf1 protein is located at or near the C-terminus of the Cpf1 protein sequence. For example, but not limited to, the modified Cpf1 protein may be selected from His-AsCpf1-nNLS (SEQ ID NO: 1000); His-AsCpf1-sNLS (SEQ ID NO: 1008); and His-AsCpf1-sNLS-sNLS (SEQ ID NO: 1001), where "His" refers to a purified 6-histidine sequence, "AsCpf1" refers to an Acidaminococcus species Cpf1 protein sequence, "nNLS" refers to a nucleoplasmin NLS, and "sNLS" refers to an SV40 NLS. For example, additional permutations of NLS sequence identity and C-terminal position, such as the addition of two or more nNLS sequences or a combination of nNLS and sNLS sequences (or other NLS sequences), and the addition of sequences that include or do not include purified sequences, such as 6-histidine sequences, are within the scope of the currently disclosed subject matter.

[0212] In certain embodiments, the NLS sequence of the modified Cpf1 protein may be located at or near the N-terminus of the Cpf1 protein sequence. For example, but not limited to, the modified Cpf1 protein may be selected from His-sNLS-AsCpf1 (SEQ ID NO: 1009), His-sNLS-sNLS-AsCpf1 (SEQ ID NO: 1010), and sNLS-sNLS-AsCpf1 (SEQ ID NO: 1011). Additional permutations of NLS sequence identity and N-terminal position, such as the addition of two or more nNLS sequences or a combination of nNLS and sNLS sequences (or other NLS sequences), and the addition of sequences that include or do not include purified sequences, such as 6-histidine sequences, are within the scope of the subject matter currently disclosed.

[0213] In certain embodiments, the modified Cpf1 protein may include NLS sequences located at or near both the N-terminus and C-terminus of the Cpf1 protein sequence. For example, but not limited to, the modified Cpf1 protein may be selected from His-sNLS-AsCpf1-sNLS (SEQ ID NO: 1012) and His-sNLS-sNLS-AsCpf1-sNLS-sNLS (SEQ ID NO: 1013). For example, additional permutations of NLS sequence identity and N-terminus / C-terminus positions, such as the addition of two or more nNLS sequences or combinations of nNLS and sNLS sequences (or other NLS sequences) to either the N-terminus or C-terminus position, and the addition of sequences that include or do not include purified sequences such as 6-histidine sequences, are within the scope of the subject matter currently disclosed.

[0214] In certain embodiments, the modified Cpf1 protein may include modifications (e.g., deletions or substitutions) to one or more cysteine ​​residues in the Cpf1 protein sequence. For example, but not limited to, the modified Cpf1 protein may include modifications at positions selected from the group consisting of C65, C205, C334, C379, C608, C674, C1025, and C1248. In certain embodiments, the modified Cpf1 protein may include substitutions of one or more serine or alanine cysteine ​​residues. In certain embodiments, the modified Cpf1 protein may include modifications selected from the group consisting of C65S, C205S, C334S, C379S, C608S, C674S, C1025S, and C1248S. In certain embodiments, the modified Cpf1 protein may include modifications selected from the group consisting of C65A, C205A, C334A, C379A, C608A, C674A, C1025A, and C1248A. In certain embodiments, the modified Cpf1 protein may include modifications at positions C334 and C674 or C334, C379, and C674. In certain embodiments, the modified Cpf1 protein may include modifications of C334S and C674S or C334S, C379S, and C674S. In certain embodiments, the modified Cpf1 protein may include modifications of C334A and C674A or C334A, C379A, and C674A. In certain embodiments, modified Cpf1 proteins may include both one or more cysteine ​​residue modifications, such as His-AsCpf1-nNLS Cys-less (SEQ ID NO: 1014) or His-AsCpf1-nNLS Cys-low (SEQ ID NO: 1015), and the introduction of one or more NLS sequences. In various embodiments, Cpf1 proteins containing one or more cysteine ​​residue deletions or substitutions exhibit reduced aggregation.

[0215] In certain embodiments, other modified Cpf1 proteins known in the Art may be used in conjunction with the methods and systems described herein. For example, in certain embodiments, the modified Cpf1 may be Cpf1 containing the mutation S542R / K548V / N552R ("Cpf1 RVR"). Cpf1 RVR has been shown to cleave target sites having TATV PAM. In certain embodiments, the modified Cpf1 may be Cpf1 containing the mutation S542R / K607R ("Cpf1 RR"). Cpf1 RR has been shown to cleave target sites having TYCV / CCCC PAM.

[0216] In some embodiments, the Cpf1 variant is used herein, where the Cpf1 variant is 11, 12, 13, 14, 15, 16, 17, 34, 36, 39, 40, 43, 46, 47, 50, 54, 57, 58, 111, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 532, 533, 534, 535, 536, 537, 538, 539, 540, 541, 542, 543, 544, 545, 546, 547, 548, 549, 550, 551, 552, 553, 554, 555, 556, 565, 566, 567, 568, 569, 570, 571, 572, 573, 574, 575, 592, 593, 594, 595, 596, 597, 598, 599, 600, 601, 602, 603, 604, 605, 606, 607, 608, 609, 610, 611, 612, 613, 61 4, 615, 616, 617, 618, 619, 620, 626, 627, 628, 629, 630, 631, 632, 633, 634, 635, 636, 637, 638, 642, 643, 644, 645, 646, 647, 648, 649, 651, 652, 653, 654, 655, 656, 676, 679, 680, 682, 683, 684, 685, 686, 687, 688, 689, 690, 691, 692, 693, 707, 711, 714, 715, 716, 717, 718, 719, 720, 721, The mutation is found in one or more residues of AsCpf1 (Acidaminococcus species BV3L6) selected from the group consisting of 722, 739, 765, 768, 769, 773, 777, 778, 779, 780, 781, 782, 783, 784, 785, 786, 870, 871, 872, 873, 874, 875, 876, 877, 878, 879, 880, 881, 882, 883, 884, or 1048, or the corresponding positions of orthologs, homologs, or mutants of AsCpf1.

[0217] In certain embodiments, the Cpf1 variant, as used herein, may comprise any of the Cpf1 proteins described in International Publication No. WO 2017 / 184768 (the "'768 publication") by Zhang et al., which is incorporated herein by reference.

[0218] In certain embodiments, the modified Cpf1 protein (also referred to as a Cpf1 variant) used herein may be encoded by any of the sequences set forth in SEQ ID NOs: 1000, 1001, 1008 to 1018, 1032, 1035-39, 1094 to 1097, 1107 to 1109 (Cpf1 polypeptide sequences) or SEQ ID NOs: 1019 to 1021, 1110 to 1117 (Cpf1 polynucleotide sequences). Table 20 describes exemplary Cpf1 variant amino acid and nucleotide sequences. These sequences are set forth in Figure 62, which details the positions of the 6-histidine sequence (underlined text) and the NLS sequence (bold text). Additional permutations of NLS sequence identity and N-terminal / C-terminal position, including addition of two or more nNLS sequences or a combination of nNLS and sNLS sequences (or other NLS sequences) at either the N-terminal or C-terminal position, and addition of sequences with or without purification sequences such as a 6-histidine sequence, are within the scope of the presently disclosed subject matter.

[0219] In certain embodiments, either the Cpf1 protein or the modified Cpf1 protein disclosed herein may be complexed with one or more gRNAs containing the targeting domains described in SEQ ID NO: 1002 and / or 1004 to modify the CCAAT box target region. In certain embodiments, either the Cpf1 protein or the modified Cpf1 protein disclosed herein may be complexed with one or more gRNAs containing the sequences described in Table 18 or Table 19. In certain embodiments, the modified Cpf1 protein may be His-AsCpf1-nNLS (SEQ ID NO: 1000) or His-AsCpf1-sNLS-sNLS (SEQ ID NO: 1001). In certain embodiments, the modified Cpf1 protein used herein may be encoded by any of the sequences described in SEQ ID NOs: 1000, 1001, 1008-1018, 1032, 1035-39, 1094-1097, 1107-09 (Cpf1 polypeptide sequences) or SEQ ID NOs: 1019-1021, 1110-17 (Cpf1 polynucleotide sequences). In certain embodiments, the modified Cpf1 protein may include the sequence described in SEQ ID NO: 1097.

[0220] Modification of RNA-induced nucleases While the RNA-induced nucleases described above possess activities and properties that may be useful for a variety of applications, those skilled in the art will recognize that RNA-induced nucleases can, in some cases, be modified to alter their cleavage activity, PAM specificity, or other structural or functional characteristics.

[0221] First, looking at modifications that alter cleavage activity, mutations that reduce or eliminate the activity of domains within the NUC lobe are described above. Exemplary mutations that can be generated in the RuvC domain, Cas9 HNH domain, or Cpf1 Nuc domain are described by Ran & Hsu 2013 and Yamano, as well as Cotta-Ramusino. Generally, mutations that reduce or eliminate the activity of one of two nuclease domains result in an RNA-inducible nuclease with nickase activity, but it should be noted that the type of nickase activity changes depending on which domain is inactivated. As an example, inactivation of the RuvC domain of Cas9 results in a nickase that cleaves the complementary or upper strand, as shown below (where C indicates the cleavage site).

[0222] On the other hand, inactivation of the Cas9 HNH domain results in a nickas that cleaves the lower or non-complementary chain.

[0223] Modifications of PAM specificity compared to naturally occurring Cas9 reference molecules have been described by Kleinstiver et al. for both S. pyogenes (Kleinstiver 2015a) and S. aureus (Kleinstiver 2015b). Kleinstiver et al. also describe modifications that improve the targeting fidelity of Cas9 (Kleinstiver 2016). Each of these references is incorporated herein by reference.

[0224] RNA-induced nucleases are divided into two or more parts, as described by Zetsche 2015 and Fine 2015 (both incorporated herein by reference).

[0225] In certain embodiments, RNA-induced nucleases may be size-optimized or shortened through one or more deletions that reduce the size of the nuclease while still maintaining, for example, gRNA binding, target and PAM recognition, and cleavage activity. In certain embodiments, RNA-induced nucleases optionally bind covalently or noncovalently to another polypeptide, nucleotide, or other structure by a linker. Exemplary binding nucleases and linkers are described by Guilinger 2014, which are incorporated herein by reference for all purposes.

[0226] RNA-induced nucleases optionally include, but are not limited to, labels such as nuclear localization signals, to facilitate the transfer of RNA-induced nuclease proteins into the nucleus. In certain embodiments, RNA-induced nucleases may incorporate nuclear localization signals at their C-terminus and / or N-terminus. Nuclear localization sequences are known in the art and are described in Maeder and elsewhere.

[0227] The aforementioned list of modifications is intended to be illustrative in nature, and those skilled in the art will understand, in view of this disclosure, that other modifications may be possible or desirable in specific applications. Therefore, for the sake of brevity, the exemplary systems, methods, and compositions of this disclosure are presented with reference to specific RNA-inducing nucleases, but it should be understood that the RNA-inducing nucleases used may be modified in a manner that does not alter their operating principles. Such modifications are within the scope of this disclosure.

[0228] nucleic acids encoding RNA-induced nucleases For example, nucleic acids encoding RNA-induced nucleases such as Cas9, Cpf1, or functional fragments thereof are provided herein. Examples of nucleic acid sequences encoding the Cas9 molecule that may be used according to embodiments herein are given in SEQ ID NOs: 3, 7-11, 13. Exemplary nucleic acids encoding RNA-induced nucleases have been previously described (see, for example, Cong 2013; Wang 2013; Mali 2013; Jinek 2012).

[0229] In some cases, the nucleic acid encoding the RNA-inducing nuclease may be a synthetic nucleic acid sequence. For example, the synthetic nucleic acid molecule may be chemically modified. In certain embodiments, the mRNA encoding the RNA-inducing nuclease may have one or more (e.g., all) of the following properties: it may be capped; it may be polyadenylated; and it may be substituted with 5-methylcytidine and / or pseudouridine.

[0230] Synthetic nucleic acid sequences can also be codon-optimized, for example, by replacing at least one uncommon or less common codon with a common codon. For example, synthetic nucleic acids can induce the synthesis of optimized messenger mRNA optimized for expression in mammalian expression systems, as described herein. An example of a codon-optimized Cas9 coding sequence is presented in Cotta-Ramusino.

[0231] Furthermore, or alternatively, nucleic acids encoding RNA-induced nucleases may contain a nuclear localization sequence (NLS). Nuclear localization sequences are known in the art.

[0232] Functional analysis of candidate molecules Candidate RNA-induced nucleases, gRNAs, and their complexes can be evaluated by standard methods known in the art. See, for example, Cotta-Ramusino. The stability of the RNP complex can be evaluated by differential scanning fluorescence quantification as described below.

[0233] Differential scanning fluorescence (DSF) The thermal stability of ribonucleoprotein (RNP) complexes containing gRNA and RNA-induced nucleases can be measured by DSF. DSF technology measures the thermal stability of proteins, which can be increased under favorable conditions, such as the addition of binding RNA molecules like gRNA.

[0234] DSF assays can be performed according to any suitable protocol, but are not limited to, and may be used in any suitable setting, including (a) testing different conditions (e.g., different stoichiometric ratios of gRNA:RNA-induced nuclease protein, different buffers, etc.) to identify optimal conditions for RNP formation; and (b) testing modifications of RNA-induced nucleases and / or gRNAs (e.g., chemical modifications, sequence alterations, etc.) to identify modifications that improve RNP formation or stability. One reading of a DSF assay is the change in melting temperature of the RNP complex; a relatively high change suggests that the RNP complex is more stable (and therefore may have greater activity or more favorable formation, degradation, or other functional properties) compared to a standard RNP complex characterized by a lower change. When a DSF assay is deployed as a screening tool, a threshold melting temperature change may be identified, thereby the result being one or more RNPs with a melting temperature change above the threshold. For example, the threshold could be 5-10°C (e.g., 5°C, 6°C, 7°C, 8°C, 9°C, 10°C) or greater, and the result could be one or more RNPs characterized by a change in melting temperature above the threshold.

[0235] Two non-limiting examples of DSF assay conditions are as follows:

[0236] To determine the best solution for RNP complex formation, Cas9+10×SYPRO Orange® (Life Technologies, catalog no. S-6650) at a fixed concentration (e.g., 2 μM) in water is dispensed into a 384-well plate. Next, equimolar amounts of diluted gRNA are added to solutions with varying pH and salt concentrations. After incubation at room temperature for 10 minutes and brief centrifugation to remove any bubbles, a gradient is run from 20°C to 90°C in 1°C increments every 10 seconds using a Bio-Rad CFX384® real-time system C1000 Touch® thermal cycler with Bio-Rad CFX Manager software.

[0237] The second assay consists of mixing various concentrations of gRNA with a fixed concentration (e.g., 2 μM) of Cas9i in the optimal buffer from assay 1 above, and incubating it in a 384-well plate (e.g., at room temperature for 10 minutes). Equivolutes of optimal buffer + 10 × SYPRO Orange® (Life Technologies, catalog no. S-6650) are added, and the plate is sealed with Microseal® B adhesive (MSB-1001). After briefly centrifugating to remove any bubbles, a gradient is run from 20°C to 90°C in increments of 1°C every 10 seconds using a Bio-Rad CFX384® real-time system C1000 Touch® thermal cycler with Bio-Rad CFX Manager software.

[0238] Genome editing strategies Using the genome editing systems described above, editing (i.e., modification) is performed in cells or within target regions of DNA obtained from cells, in various embodiments of this disclosure. Various strategies for performing specific edits are described herein, and these strategies are generally described by the desired repair outcome, the number and location of individual edits (e.g., SSBs or DSBs), and the target sites of such edits.

[0239] Genome editing strategies involving the formation of SSBs or DSBs are characterized by repair outcomes including (a) deletion of all or part of the target region; (b) insertion or substitution within all or part of the target region; or (c) interruption of all or part of the target region. This grouping is not intended to limit or be bound by any particular theory or model, but is provided solely for economics of presentation. Those skilled in the art will understand that the listed outcomes are not mutually exclusive and that some repairs may result in others. Descriptions of particular editing strategies or methods should not be understood as requiring a particular repair outcome unless otherwise specified.

[0240] Target region substitution generally involves the substitution of all or part of a sequence present within the target region by homologous sequences, through, for example, gene modification or gene conversion, which are two repair outcomes mediated by the HDR pathway. HDR is facilitated by the use of donor templates, which may be single-stranded or double-stranded, as described in more detail below. Single-stranded or double-stranded templates may be exogenous, in which case they facilitate gene modification, or they may be endogenous (e.g., homologous sequences in the cellular genome) and facilitate gene conversion. Exogenous templates may have asymmetric overhangs, as described, for example, by Richardson 2016 (incorporated herein by reference) (i.e., the portion of the template complementary to the DSB site may be offset in the 3' or 5' direction rather than located in the center of the template). If the template is single-stranded, it may correspond to either the complementary (upper) or non-complementary (lower) strand of the target region.

[0241] As described by Ran & Hsu and Cotta-Ramusino, gene transformation and gene modification are sometimes facilitated by forming one or more nicks within or around the target region. Sometimes, a double nicking strategy is used to form two offset SSBs, followed by the formation of a single DSB with an overhang (e.g., a 5' overhang).

[0242] Interruption and / or deletion of all or part of a target sequence can be achieved by various repair outcomes. For example, as described by Maeder for the LCA10 mutation, a sequence can be deleted by simultaneously generating two or more double-strand breaks (DSBs) flanking the target region, which are then excised when the DSBs are repaired. Alternatively, a sequence can be interrupted by the formation of a double-strand break with a single-strand overhang, followed by deletion generated by nucleotide terminal hydrolysis processing of the pre-repair overhang.

[0243] One particular subset of target sequence interruptions is mediated by the formation of indels in the target sequence, where the repair outcome is typically mediated by the NHEJ pathway (including Alt-NHEJ). NHEJ is referred to as the “error-prone” repair pathway due to its association with indel mutations. However, in some cases, DSBs are repaired by NHEJ without alteration of the surrounding sequence (so-called “perfect” or “scarless” repair); this generally requires both ends of the DSB to be completely ligated. Indels, on the other hand, are thought to arise from the enzymatic processing of free DNA ends before they are ligated, which adds and / or removes nucleotides from one or both strands of one or both free ends.

[0244] Because the enzymatic processing of free DSB ends can be inherently probabilistic, indel mutations tend to be variable, occurring along distributions and influenced by various factors, including specific target sites, cell types used, and genome editing strategies employed. Nevertheless, it is possible to generalize to a limited extent about indel formation: deletions formed by single DSB repair are most commonly in the range of 1–50 bp, but can exceed 100–200 bp. Insertions formed by single DSB repair tend to be shorter and often contain short duplications of sequences directly surrounding the cleavage site. However, larger insertions are possible, and in these cases, the inserted sequence is often traced back to other regions of the genome or plasmid DNA present within the cell.

[0245] Indel mutations and genome editing systems configured to generate indels are useful for interrupting target sequences, for example, when the generation of a specific final sequence is not required and / or when frameshift mutations are tolerable. They may also be useful in situations where a specific sequence is preferred, as long as the desired specific sequence tends to preferentially arise from the repair of SSBs or DSBs at a given site. Indel mutations are also a useful tool for evaluating or screening the activity of specific genome editing systems and their components. In these and other settings, indels may be characterized by (a) their relative and absolute frequencies in the genome of the cell in contact with the genome editing system, and (b) a distribution of numerical differences, such as ±1, ±2, ±3, relative to the unedited sequence. As an example, in a read discovery setting, multiple gRNAs may be screened to identify the gRNA that most efficiently promotes cleavage at the target site based on indel readings under controlled conditions. Guides that generate indels above a threshold frequency or guides that generate a specific distribution of indels may be selected for further research and development. The frequency and distribution of indels can also be useful as a read for evaluating different genome editing system implementations or formulation and delivery methods, for example, by keeping the gRNA constant while varying other specific reaction conditions or delivery methods.

[0246] Multiple strategies The genome editing systems described herein may also be used for multiplex gene editing to generate two or more double-segmented breaks (DSBs) at either the same or different loci. Any of the RNA-induced nucleases and gRNAs disclosed herein may be used in genome editing systems for multiplex gene editing. Editing strategies involving the formation of multiple DSBs or single-segmented breaks (SSBs) are described, for example, in Cotta-Ramusino.

[0247] As disclosed herein, multiple gRNAs can be used in a genome editing system to introduce modifications (e.g., deletions, insertions) to the 13nt target regions of HBG1 and / or HBG2. In certain embodiments, one or more gRNAs containing the targeting domains described in SEQ ID NOs. 251-901, 940-942 may be used to introduce modifications to the 13nt target regions of HBG1 and / or HBG2. In other embodiments, multiple gRNAs can be used in a genome editing system to introduce modifications to the CCAAT box target region. In certain embodiments, one or more gRNAs containing the sequences described in SEQ ID NOs. 970, 971, 996, 997 may be used to introduce modifications to the CCAAT box target region. In certain embodiments, one or more gRNAs containing the targeting domains described in SEQ ID NOs. 1002, 1004 may be used to introduce modifications to the CCAAT box target region. In other embodiments, multiple gRNAs can be used in a genome editing system to introduce modifications to the GATA1 binding motif in BCL11Ae. In certain embodiments, modifications can be introduced into the GATA1 binding motif of BCL11Ae using one or more gRNAs containing the targeting domains described in SEQ ID NOs.952-955. Multiple gRNAs can also be used in genome editing systems to introduce modifications into the GATA1 binding motif, the CCAAT box target region, the 13nt target region of HBG1 and / or HBG2, or a combination thereof, in BCL11Ae. In certain embodiments, one or more gRNAs containing the targeting domains described in SEQ ID NOs.952-955 may be used to introduce modifications into the GATA1 binding motif in BCL11Ae, and one or more gRNAs containing the targeting domains described in SEQ ID NOs.251-901, 940-942 may be used to introduce modifications into the 13nt target region of HBG1 and / or HBG2. In certain embodiments, one or more gRNAs containing the targeting domains described in SEQ ID NOs: 952-955 may be used to introduce a modification to the GATA1 binding motif in BCL11Ae, and one or more gRNAs containing the targeting domains described in SEQ ID NOs: 970, 971, 996, 997 may be used to introduce a modification to the CCAAT box target region.In certain embodiments, one or more gRNAs containing the targeting domains described in SEQ ID NOs: 952-955 may be used to introduce a modification to the GATA1 binding motif in BCL11Ae, and one or more gRNAs containing the targeting domains described in SEQ ID NOs: 1002, 1004 may be used to introduce a modification to the CCAAT box target region.

[0248] In certain embodiments, multiple gRNAs and RNA-inducing nucleases may be used in a genome editing system to introduce modifications (e.g., deletions, insertions) to the CCAAT box target regions of HBG1 and / or HBG2. In certain embodiments, the RNA-inducing nuclease may be Cas9, modified Cas9 (e.g., D10A), Cpf1, or modified Cpf1.

[0249] Donor mold design Donor template design is described in detail in literature such as Cotta-Ramusino. DNA oligomer donor templates (oligodeoxynucleotides or ODNs), which can be single-stranded (ssODN) or double-stranded (dsODN), can be used to facilitate HDR-based repair of DSBs or to improve the overall edit rate, and are particularly useful for introducing modifications to a target DNA sequence, inserting a new sequence into the target sequence, or completely replacing the target sequence.

[0250] Whether single-stranded or double-stranded, the donor template generally includes regions homologous to the DNA region in or near the target sequence to be cleaved (e.g., lateral or adjacent). These homologous regions are referred to herein as “homology arms” and are schematically shown below. [5' homology arm]--[substitution sequence]--[3' homology arm]

[0251] Homology arms can have any appropriate length (including none if only one homology arm is used), and the 3' and 5' homology arms may have the same length or different lengths. The selection of appropriate homology arm lengths may be influenced by various factors, such as the desire to avoid homology or microhomology with certain sequences, such as Alu repeats or other very common elements. For example, the 5' homology arm may be shortened to avoid sequence repeats. In other embodiments, the 3' homology arm may be shortened to avoid sequence repeats. In some embodiments, both the 5' and 3' homology arms may be shortened to avoid inclusion of certain sequence repeats. Furthermore, some homology arm designs may improve editing efficiency or increase the frequency of desired repair outcomes. For example, Richardson 2016, incorporated herein by reference, found that the relative asymmetry of the 3' and 5' homology arms of a single-stranded donor template affects the repair rate and / or outcome.

[0252] Substitution sequences in donor templates are described elsewhere, including in Cotta-Ramusino. Substitution sequences can be of any appropriate length (including no nucleotides if the desired repair outcome is a deletion) and typically involve one, two, three, or more sequence modifications to the intrinsic intracellular sequence to be edited. One common sequence modification involves altering the intrinsic sequence to repair mutations associated with the disease or condition to be treated. Another common sequence modification involves altering one or more sequences complementary to the PAM sequence of an RNA-induced nuclease or the targeting domain of a gRNA used to generate SSBs or DSBs, thereby reducing or eliminating repeat breaks at the target site after the substitution sequence has been incorporated.

[0253] When a linear ssODN is used, it may be configured to (i) anneal to the nicked strand of the target nucleic acid, (ii) anneal to the intact target nucleic acid strand, (iii) anneal to the positive strand of the target nucleic acid, and / or (iv) anneal to the negative strand of the target nucleic acid. The ssODN may have any suitable length, such as approximately or at least 80 to 200 nucleotides or less (e.g., 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190 or 200 nucleotides).

[0254] It should be noted that the template nucleic acid may be a nucleic acid vector such as a viral genome or a circular double-stranded DNA such as a plasmid. Nucleic acid vectors containing a donor template may contain other coding or non-coding elements. For example, the template nucleic acid may contain specific genomic backbone elements (e.g., reverse terminal repeats in the case of an AAV genome) and may be delivered as part of a viral genome (e.g., in an AAV or lentiviral genome) that optionally contains additional sequences encoding gRNA and / or RNA-induced nucleases. In certain embodiments, the donor template may be adjacent to or sandwiched between target sites recognized by one or more gRNAs, facilitating the formation of free DSBs at one or both ends of the donor template, which may be involved in the repair of corresponding SSBs or DSBs formed in cellular DNA using the same gRNA. Exemplary nucleic acid vectors suitable for use as donor templates are described by Cotta-Ramusino, which are incorporated by reference.

[0255] Regardless of the form used, the template nucleic acid may be designed to avoid undesirable sequences. In certain embodiments, one or both homology arms may be shortened to avoid overlap with certain sequence repeat elements, such as Alu repeats and LINE elements.

[0256] In certain embodiments, silent, non-pathogenic SNPs may be included in the ssODN donor template to enable the identification of gene editing events.

[0257] In certain embodiments, the donor template may be a non-specific template that is non-homologous to a region of DNA within or near the target sequence to be cleaved. In certain embodiments, a donor template for use in targeting the GATA1 binding motif in BCL11Ae may be, without limitation, a non-target-specific template that is non-homologous to a region of DNA within or near the GATA1 binding motif in BCL11Ae. In certain embodiments, a donor template for use in targeting a 13nt target region may be, without limitation, a non-target-specific template that is non-homologous to a region of DNA within or near the 13nt target region.

[0258] In the uses herein, donor template or template nucleic acid refers to a nucleic acid sequence that, when used in combination with an RNA nuclease molecule and one or more gRNA molecules, can modify (e.g., delete, disrupt, or alter) a target DNA sequence. In certain embodiments, the template nucleic acid results in a modification (e.g., deletion) to the CCAAT box target region of HBG1 and / or HBG2. In certain embodiments, the modification is a non-spontaneous modification. In certain embodiments, a non-spontaneous modification in the CCAAT box target region of HBG1 and / or HBG2 may include an 18nt target region, an 11nt target region, a 4nt target region, or a 1nt target region, or a combination thereof. In certain embodiments, the modification is a spontaneous modification. In certain embodiments, a spontaneous modification in the CCAAT box target region of HBG1 and / or HBG2 may include a 13nt target region, a c.-117G>A target region, or a combination thereof. In certain embodiments, the template nucleic acid is ssODN. In certain embodiments, ssODN is either a positive or negative strand.

[0259] For example, a template nucleic acid for introducing an 18nt deletion into an 18nt target region (HBG1 c.-104~-121, HBG2 c.-104~-121, or a combination thereof) may include a 5' homology arm, a substitution sequence, and a 3' homology arm, the substitution sequence being 0 nucleotides or 0 bp. In certain embodiments, the 5' homology arm may be about 25 to about 200 nucleotides or more in length, for example, at least about 25, 50, 75, 100, 125, 150, 175, or 200 nucleotides. In certain embodiments, the 5' homology arm includes about 50 to 100 bp of homology on the 5' side of the 18nt target region, for example, 55~95, 60~90, 70~90, or 80~90 bp. In certain embodiments, the 3' homology arm may be approximately 25 to approximately 200 nucleotides or more in length, for example, at least approximately 25, 50, 75, 100, 125, 150, 175, or 200 nucleotides. In certain embodiments, the 3' homology arm includes approximately 50 to 100 bp of homology on the 3' side of the 18nt target region, for example, 55 to 95, 60 to 90, 70 to 90, or 80 to 90 bp. In certain embodiments, the 5' and 3' homology arms are symmetric in length. In certain embodiments, the 5' and 3' homology arms are asymmetric in length. In certain embodiments, the template nucleic acid is ssODN. In certain embodiments, the ssODN is a positive strand. In certain embodiments, the ssODN is a negative strand. In certain embodiments, ssODN includes, is essentially derived from, or consists of sequence number 974 (OLI16409) or sequence number 975 (OLI16410).

[0260] In certain embodiments, the template nucleic acid for introducing an 11nt deletion into an 11nt target region (HBG1 c.-105~-115, HBG2 c.-105~-115, or a combination thereof) may include a 5' homology arm, a substitution sequence, and a 3' homology arm, the substitution sequence being 0 nucleotides or 0 bp. In certain embodiments, the 5' homology arm may be about 25 to about 200 nucleotides in length, for example, at least about 25, 50, 75, 100, 125, 150, 175, or 200 nucleotides. In certain embodiments, the 5' homology arm includes about 50 to 100 bp of homology on the 5' side of the 11nt target region, for example, 55~95, 60~90, 70~90, or 80~90 bp. In certain embodiments, the 3' homology arm may be approximately 25 to approximately 200 nucleotides in length, for example, at least approximately 25, 50, 75, 100, 125, 150, 175, or 200 nucleotides. In certain embodiments, the 3' homology arm includes approximately 50 to 100 bp of homology on the 3' side of the 11nt target region, for example, 55 to 95, 60 to 90, 70 to 90, or 80 to 90 bp. In certain embodiments, the 5' and 3' homology arms are symmetric in length. In certain embodiments, the 5' and 3' homology arms are asymmetric in length. In certain embodiments, the template nucleic acid is ssODN. In certain embodiments, the ssODN is a positive strand. In certain embodiments, the ssODN is a negative strand. In certain embodiments, ssODN includes, is essentially derived from, or consists of sequence number 976 (OLI16411) or sequence number 978 (OLI16413).

[0261] In certain embodiments, the template nucleic acid for introducing a 4nt deletion into a 4nt target region (HBG1 c.-112~-115, HBG2 c.-112~-115, or a combination thereof) may include a 5' homology arm, a substitution sequence, and a 3' homology arm, the substitution sequence being 0 nucleotides or 0 bp. In certain embodiments, the 5' homology arm may be about 25 to about 200 nucleotides in length, for example, at least about 25, 50, 75, 100, 125, 150, 175, or 200 nucleotides. In certain embodiments, the 5' homology arm includes about 50 to 100 bp of homology on the 5' side of the 4nt target region, for example, 55~95, 60~90, 70~90, or 80~90 bp. In certain embodiments, the 3' homology arm may be approximately 25 to approximately 200 nucleotides in length, for example, at least approximately 25, 50, 75, 100, 125, 150, 175, or 200 nucleotides. In certain embodiments, the 3' homology arm includes approximately 50 to 100 bp of homology on the 3' side of the 4nt target region, for example, 55 to 95, 60 to 90, 70 to 90, or 80 to 90 bp. In certain embodiments, the 5' and 3' homology arms are symmetric in length. In certain embodiments, the 5' and 3' homology arms are asymmetric in length. In certain embodiments, the template nucleic acid is ssODN. In certain embodiments, the ssODN is a positive strand. In certain embodiments, the ssODN is a negative strand. I...

Claims

1. 5' end and 3' end, The 5' end contains a DNA extension selected from the group consisting of sequence numbers 1236 to 1250. A targeting domain complementary to the target site in the HBG gene promoter, and RNA segments capable of binding to Cpf1 RNA-induced nuclease or Cpf1 mutants, Single-molecule gRNAs including, The Cpf1 RNA-induced nuclease or the nucleic acid encoding the Cpf1 mutant, A genome editing system containing [the specified ingredient].

2. The genome editing system according to claim 1, wherein the gRNA includes a 2'-O-methyl modification and a phosphorothioate modification at the 3' end.

3. The genome editing system according to claim 1 or 2, wherein the gRNA includes a 2'-fluoro modification.

4. The targeting domains are sequence numbers 1002, 1004, 1139, 1141, 1143, 1145, 1147, 1149, 1151, 1153, 1155, 1157, 1159, 1161, 1163, 1165, 1167, 1169, 1171, 1173, 1175, 1177, 1179, 1181, 1183, 1185, 1187, 1189, 1191, 1193, A genome editing system according to any one of claims 1 to 3, comprising a sequence selected from the group consisting of 1195, 1197, 1199, 1201, 1203, 1205, 1207, 1209, 1211, 1213, 1215, 1217, 1219, 1221, 1223, 1225, 1227, 1229, 1254, 1256, 1258, 1260, 1262, and 1264.

5. The genome editing system according to any one of claims 1 to 4, wherein the target site includes a nucleotide located at Chr11(NC_000011.10) 5,249,955 to 5,249,987 or Chr11(NC_000011.10) 5,254,879 to 5,254,909.

6. The genome editing system according to any one of claims 1 to 5, wherein the Cpf1 mutant comprises one or more modifications selected from the group consisting of one or more mutations in the wild-type Cpf1 amino acid sequence, one or more nuclear localization signals, one or more purified tags, and combinations thereof.

7. The genome editing system according to any one of claims 1 to 6, wherein the Cpf1 variant includes a sequence selected from the group consisting of SEQ ID NOs: 1000, 1001, 1008-1015, and 1035-1039.

8. A genome editing system according to any one of claims 1 to 7, used for a method of modifying the promoter of the HBG gene in the cell, comprising contacting the cell with the gRNA and the nucleic acid.

9. The genome editing system according to claim 8, wherein the cells are CD34+ cells or hematopoietic stem cells.

10. The genome editing system according to claim 8 or 9, wherein the gRNA and the nucleic acid are delivered to the cells using electroporation or lipid nanoparticles.

11. The genome editing system according to any one of claims 8 to 10, wherein the nucleic acid includes messenger RNA.

12. A genome editing system according to any one of claims 8 to 11, for use in alleviating one or more symptoms of sickle cell disease in a person who requires it.

13. A composition comprising a genome editing system according to any one of claims 8 to 12, further comprising a pharmaceutically acceptable carrier.

14. A genome editing system according to any one of claims 1 to 7, Pharmacologically acceptable carriers A kit that includes this.

Citation Information

Patent Citations

  • Applications of modified crRNA in CRISPR / Cpf1 gene editing system

    CN106244591A

  • Crispr / CAS-related methods and compositions for treating beta hemoglobinopathies

    WO2017160890A1