Systems and methods for the treatment of hemoglobinopathies
Patent Information
- Application Number
- JP2024506528
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-11-12
- Filing Date
- 2022-08-02
- Publication Date
- 2025-08-08
AI Technical Summary
Current treatments for hemoglobinopathies such as sickle cell disease and beta-thalassemia are limited by the risk of complications from blood transfusions, graft-versus-host disease, and the challenge of identifying matched donors for hematopoietic stem cell transplantation, necessitating improved methods for managing these conditions.
A genome editing system using RNP complexes comprising Cpf1 and gRNA is employed to modify the promoter of the HBG gene in CD34+ hematopoietic stem cells, increasing fetal hemoglobin expression by introducing specific indels, thereby reducing symptoms of beta-thalassemia.
The method enhances fetal hemoglobin production in modified stem cells, leading to reduced anemia and other symptoms of beta-thalassemia, potentially improving patient outcomes and reducing the need for frequent blood transfusions.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] Claiming priority This application claims the benefit of U.S. Provisional Application No. 63 / 228,509, filed August 2, 2021, and U.S. Provisional Application No. 63 / 278,899, filed November 12, 2021, both of which are incorporated by reference herein in their entireties.
[0002] Sequence Listing This application contains a Sequence Listing that has been submitted in ASCII format via EFS-Web, which is hereby incorporated by reference in its entirety. The ASCII copy created on November 12, 2021 is named SequenceListing.txt and is 699KB in size.
[0003] Technical Field The present disclosure relates to genome editing systems and methods for altering a target nucleic acid sequence or modulating the expression of a target nucleic acid sequence and their uses in connection with altering genes encoding hemoglobin subunits and / or treating hemoglobinopathies. [Background technology]
[0004] Hemoglobin (Hb) carries oxygen in erythrocytes or red blood cells (RBCs) from the lungs to tissues. During prenatal development and until shortly after birth, hemoglobin exists in the form of fetal hemoglobin (HbF), a tetrameric protein consisting of two alpha (α)-globin chains and two gamma (γ)-globin chains. HbF is largely replaced by adult hemoglobin (HbA), a tetrameric protein in which the γ-globin chains of HbF are replaced by beta (β)-globin chains, through a process known as globin switching. The average adult produces less than 1% HbF from total hemoglobin (Thein 2009). The α-hemoglobin gene is located on chromosome 16, and the β-hemoglobin gene (HBB), the A gamma (Aγ)-globin chain (HBG1, also known as gamma globin A), and the G gamma (Gγ)-globin chain (HBG2, also known as gamma globin G) are located on chromosome 11 within the globin gene cluster (also called the globin locus).
[0005] Mutations in HBB can cause hemoglobin disorders (i.e., hemoglobinopathies), including sickle cell disease (SCD) and beta-thalassemia (β-Thal). Approximately 93,000 people in the United States have been diagnosed with hemoglobinopathies. Worldwide, 300,000 children are born with hemoglobinopathies each year (Angastiniotis 1998). Because these conditions are associated with HBB mutations, their symptoms typically do not manifest until globin switching from HbF to HbA.
[0006] SCD is the most common inherited blood disorder in the United States, affecting approximately 80,000 people (Brousseau 2010). SCD is most common in people of African descent, with a prevalence of SCD of 1 in 500. In Africa, the prevalence of SCD is 15 million (Aliyu 2008). SCD is also common in people of Indian, Saudi, and Mediterranean descent. In Hispanic Americans, the prevalence of sickle cell disease is 1 in 1,000 (Lewis 2014).
[0007] SCD is caused by a single homozygous mutation in the HBB gene, c.17A>T (HbS mutation). The sickle mutation is a point mutation (GAG>GTG) on HBB that results in a substitution of valine with glutamic acid at amino acid position 6 in exon 1. The valine at position 6 of the β-hemoglobin chain is hydrophobic and causes a change in the conformation of the β-globin protein when it is not bound to oxygen. This conformational change causes the HbS protein to polymerize in the absence of oxygen, resulting in a deformation of the RBC (i.e., sickle cells). SCD is inherited in an autosomal recessive manner, so only patients with two HbS alleles have the disease. Heterozygous subjects have the sickle cell trait and can suffer from anemia and / or painful crises if they are severely dehydrated or oxygen deprived.
[0008] Sickled RBCs cause multiple symptoms including anemia, sickle cell crises, vaso-occlusive crises, aplastic crises, and acute chest syndrome. Sickled RBCs are less elastic than wild-type RBCs and therefore cannot easily pass through capillary beds, causing blockage and ischemia (i.e., vaso-occlusive crises). Vaso-occlusive crises occur when sickled RBCs obstruct blood flow in the capillary beds of organs, causing pain, ischemia, and necrosis. These episodes usually last for 5-7 days. The spleen is responsible for removing dysfunctional RBCs and therefore typically becomes enlarged during early childhood and is exposed to frequent vaso-occlusive crises. By the end of childhood, the spleen of SCD patients often becomes infarcted, leading to splenic damage (autosplenectomy). Hemolysis is a permanent feature of SCD and causes anemia. Sickled RBCs survive in the circulation for 10-20 days, whereas healthy RBCs survive for 90-120 days. SCD subjects are transfused as needed to maintain adequate hemoglobin levels. Frequent transfusions put subjects at risk of infection with HIV, Hepatitis B, and Hepatitis C. Subjects may also suffer from acute chest attacks and infarctions of the limbs, end-organs, and central nervous system.
[0009] Subjects with SCD have a decreased life expectancy. The prognosis for patients with SCD has steadily improved with careful lifelong management of attacks and anemia. As of 2001, the average life expectancy for subjects with sickle cell disease was in the mid to late 50s. Current treatment for SCD involves hydration and pain management during attacks, as well as blood transfusions as needed to correct anemia.
[0010] Thalassemia (e.g., β-Thal, δ-Thal, and β / δ-Thal) causes chronic anemia. β-Thal is estimated to affect approximately 1 in 100,000 people worldwide. Its prevalence is higher in certain populations, including those of European descent, where its prevalence is approximately 1 in 10,000. A more severe form of the disease, severe β-Thal, is life-threatening unless treated with lifelong blood transfusions and chelation therapy. There are approximately 3,000 subjects in the United States with severe β-Thal. Intermediate β-Thal does not require blood transfusions, but can cause growth retardation and significant systemic abnormalities, often requiring lifelong chelation therapy. HbA accounts for the majority of hemoglobin in adult RBCs, while approximately 3% of adult hemoglobin is in the form of HbA2 (HbA2 is an HbA variant in which two γ-globin chains are replaced by two delta (Δ)-globin chains). δ-Thal is associated with mutations in the Δ hemoglobin gene (HBD) that cause loss of expression. Co-inheritance of HBD mutations can mask the diagnosis of β-Thal (i.e., β / δ-Thal) by lowering the levels of HbA2 to the normal range (Bouva 2006). β / δ-Thal is usually caused by deletion of HBB and HBD sequences in both alleles. In homozygous (δo / δo βo / βo) patients, HBG is expressed, leading to the production of only HbF.
[0011] Like SCD, β-Thal is caused by mutations in the HBB gene. The most common HBB mutations causing β-Thal are c.-136C>G, c.92+1G>A, c.92+6T>C, c.93-21G>A, c.118C>T, c.316-106C>G, c.25_26delAA, c.27_28insG, c.92+5G>C, c.118C>T, c.135delC, c.315+1G>A, c.-78A>G, c. These and other mutations associated with β-Thal cause mutations or disruption of the β-globin chain, which causes a disruption of the normal Hb α-hemoglobin to β-hemoglobin ratio. Excess α-globin chains are precipitated in erythroid precursors in the bone marrow.
[0012] In severe β-Thal, both alleles of HBB harbor nonsense, frameshift, or splicing mutations (β 0 / β 0 Severe β-Thal results in a severe reduction in β-globin chains, leading to marked precipitation of α-globin chains in RBCs and more severe anemia.
[0013] Intermediate β-Thal results from mutations in the 5' or 3' untranslated regions of HBB, mutations in the promoter region or polyadenylation signal of HBB, or splicing mutations within the HBB gene. Patients' genotypes are designated βo / β+ or β+ / β+. βo represents the absence of expression of the β-globin chain, and β+ represents the presence of a dysfunctional β-globin chain. Phenotypic expression varies from patient to patient. Because there is some production of β-globin, intermediate β-Thal results in less precipitation of α-globin chains in erythroid precursors, resulting in a milder anemia than severe β-Thal. However, there is a more significant impact of proliferation of erythroid lineages secondary to chronic anemia.
[0014] Subjects with severe β-Thal between 6 months and 2 years of age suffer from failure to thrive, fever, hepatosplenomegaly, and diarrhea. Appropriate treatment includes regular blood transfusions. Therapy for severe β-Thal also includes splenectomy and treatment with hydroxyurea. If patients receive regular blood transfusions, they will develop normally until their early 20s. At that point, they require chelation therapy (in addition to continued transfusions) to prevent complications of iron overload. Iron overload may manifest as growth retardation and delayed sexual maturation. In adults, inadequate chelation therapy can lead to cardiomyopathy, cardiac arrhythmias, hepatic fibrosis and / or cirrhosis, diabetes, thyroid and parathyroid abnormalities, thrombosis, and osteoporosis. Frequent blood transfusions also put subjects at risk for infection with HIV, Hepatitis B, and Hepatitis C.
[0015] Subjects with intermediate β-Thal are generally between the ages of 2 and 6 years. They generally do not require blood transfusions. However, bone abnormalities occur due to chronic hypertrophy of the erythroid lineage to compensate for chronic anemia. Subjects may fracture long bones due to osteoporosis. Extramedullary erythropoiesis is common, resulting in enlargement of the spleen, liver, and lymph nodes. It may also lead to spinal cord compression and neurological problems. Subjects may also suffer from lower extremity ulcers and be at high risk for thrombotic events, including stroke, pulmonary embolism, and deep vein thrombosis. Treatment of intermediate β-Thal includes splenectomy, folic acid supplementation, hydroxyurea therapy, and radiation therapy for extramedullary masses. Chelation therapy is used in subjects who develop iron overload.
[0016] Patients with β-Thal often have a shortened life expectancy. Subjects with severe β-Thal who do not receive transfusion therapy typically die in their 20s or 30s. Subjects with severe β-Thal who receive regular transfusions and appropriate chelation therapy can live into their 50s or beyond. Heart failure secondary to iron toxicity is the main cause of death in subjects with severe β-Thal due to iron toxicity.
[0017] A variety of new treatments for SCD and β-Thal are currently being developed. Delivery of anti-sickle HBB genes via gene therapy is currently being investigated in clinical trials. However, the long-term efficacy and safety of this approach are unknown. Transplantation with hematopoietic stem cells (HSCs) from HLA-matched allogeneic stem cell donors has been demonstrated to cure SCD and β-Thal, but this procedure carries risks, including those associated with the ablation therapy required to prepare the subject for transplantation, and increases the risk of life-threatening opportunistic infections and the risk of graft-versus-host disease after transplantation. Furthermore, in many cases, a matched allogeneic donor cannot be identified. Thus, improved methods of managing these and other hemoglobinopathies are needed. Summary of the Invention
[0018] In a particular aspect, a method is provided for alleviating one or more symptoms of beta-thalassemia (β-Thal) in a subject in need thereof.In a particular embodiment, the method comprises: a) isolating a CD34+ or hematopoietic stem cell population from a subject; b) modifying the isolated cell population ex vivo by delivering an RNP complex to the isolated cell population, thereby altering the promoter of HBG gene in one or more isolated cells in the population, wherein the RNP complex comprises Cpf1 and a gRNA, the gRNA comprising a 5'-end and a 3'-end, an RNA or DNA extension at the 5'-end, a modification (e.g., phosphorothioate linkage and / or 2'-O-methyl modification at the 5'-end and / or 3'-end), and a targeting domain that is complementary to a target site in the promoter of HBG gene; and c) administering the modified isolated cell population to the subject, thereby alleviating one or more symptoms of β-Thal in the subject. In certain embodiments, the modification can be a 2'-O-methyl modification (e.g., 2'-O-methyl adenosine) at the 3' terminus, the 5' terminus, or the 3' and 5' terminus. In certain embodiments, the modification can be a phosphorothioate linkage followed by 2'-O-methyl adenosine at the 3' terminus. In certain embodiments, the DNA extension comprises a sequence selected from the group consisting of SEQ ID NOs: 1235-1250. In certain embodiments, the targeting domain comprises or consists of a sequence set forth in Tables 7, 8, 11, or 12. In certain embodiments, the target site comprises a nucleotide located at Chr11(NC_000011.10) 5,249,904-5,249,927 (Table 6, Region 6), Chr11(NC_000011.10) 5,254,879-5,254,909 (Table 6, Region 16), or a combination thereof. In certain embodiments, the Cpf1 comprises one or more modifications selected from the group consisting of one or more mutations in the amino acid sequence of wild-type Cpf1, one or more mutations in the nucleic acid sequence of wild-type Cpf1, one or more nuclear localization signals (NLS), one or more purification tags, and combinations thereof.In certain embodiments, Cpf1 comprises or consists of a sequence selected from the group consisting of SEQ ID NOs: 1000, 1001, 1008-1018, 1032, 1035-39, 1094-1097, and 1107-09. In certain embodiments, Cpf1 comprises or consists of a sequence selected from the group consisting of SEQ ID NOs: 1019-1021 and 1110-17. In certain embodiments, the RNP complex is delivered to the cell using electroporation.
[0019] In certain aspects, methods are provided for inducing expression of hemoglobin (Hb) in a CD34+ or hematopoietic stem cell population from a subject with beta-thalassemia (β-Thal). In certain embodiments, the methods comprise delivering an RNP complex comprising a guide RNA (gRNA) and Cpf1 to an unmodified CD34+ or hematopoietic stem cell population from a subject with β-Thal to generate a modified CD34+ or hematopoietic stem cell population comprising an indel, wherein the gRNA comprises a gRNA targeting domain, each modified CD34+ or hematopoietic stem cell comprises an indel in the HBG gene promoter, and the modified CD34+ or hematopoietic stem cell population comprises a higher Hb level than the unmodified CD34+ or hematopoietic stem cell population. In certain embodiments, the gRNA comprises a DNA extension comprising a sequence selected from the group consisting of SEQ ID NOs: 1235-1250. In certain embodiments, the gRNA targeting domain comprises or consists of a sequence set forth in Tables 7, 8, 11, or 12. In certain embodiments, the gRNA comprises a targeting domain that is complementary to a target site in the promoter of the HBG gene, wherein the target site comprises nucleotides located at Chr11(NC_000011.10) 5,249,904 to 5,249,927 (Table 6, region 6), Chr11(NC_000011.10) 5,254,879 to 5,254,909 (Table 6, region 16), or a combination thereof. In certain embodiments, the RNP complex comprises Cpf1 that comprises one or more modifications selected from the group consisting of one or more mutations in the amino acid sequence of wild-type Cpf1, one or more mutations in the nucleic acid sequence of wild-type Cpf1, one or more nuclear localization signals (NLS), one or more purification tags, and combinations thereof. In certain embodiments, Cpf1 comprises or consists of a sequence selected from the group consisting of SEQ ID NOs: 1000, 1001, 1008-1018, 1032, 1035-39, 1094-1097, and 1107-09. In certain embodiments, Cpf1 comprises or consists of a sequence selected from the group consisting of SEQ ID NOs: 1019-1021 and 1110-17. In certain embodiments, the RNP complex is delivered to the cell using electroporation.
[0020] Provided herein are genome editing systems, ribonucleoprotein (RNP) complexes, guide RNAs, Cpf1 proteins including modified Cpf1 proteins (Cpf1 variants), and CRISPR-mediated methods for altering the promoter region of one or more gamma-globin genes (e.g., HBG1, HBG2, or HBG1 and HBG2) to increase expression of fetal hemoglobin (HbF). In certain embodiments, the RNP complexes may include guide RNAs (gRNAs) complexed to wild-type Cpf1 or modified Cpf1 RNA-guided nucleases (modified Cpf1 proteins).
[0021] In certain embodiments, the gRNA may comprise a sequence set forth in Tables 7, 8, 11, or 12. In certain embodiments, the RNP complex may comprise an RNP complex set forth in Table 10. For example, the RNP complex may comprise a gRNA comprising the sequence set forth in SEQ ID NO: 1051 and a modified Cpf1 protein encoded by the sequence set forth in SEQ ID NO: 1097 (RNP32, Table 10).
[0022] In certain embodiments, the modified Cpf1 protein may include one or more modifications. In certain embodiments, the one or more modifications may include, but are not limited to, one or more mutations in the amino acid sequence of wild-type Cpf1, one or more mutations in the nucleic acid sequence of wild-type Cpf1, one or more nuclear localization signals (NLS), one or more purification tags (e.g., His tags), or combinations thereof. In certain embodiments, the modified Cpf1 may be encoded by the sequence set forth in SEQ ID NOs: 1000, 1001, 1008-1018, 1032, 1035-39, 1094-1097, 1107-09 (Cpf1 polypeptide sequence), or SEQ ID NOs: 1019-1021, 1110-17 (Cpf1 polynucleotide sequence).
[0023] In certain embodiments, the RNP complex comprising the modified Cpf1 protein can increase the editing of the target nucleic acid. In certain embodiments, the RNP complex comprising the modified Cpf1 protein can increase the editing that leads to an increase in productive indels. In various embodiments, the increase in editing of the target nucleic acid can be evaluated by any means known to those skilled in the art, including, but not limited to, PCR amplification of the target nucleic acid and subsequent sequencing analysis (e.g., Sanger sequencing, next generation sequencing).
[0024] In certain embodiments, the gRNA may include one or more modifications including phosphorothioate linkage modifications, phosphorodithioate (PS2) linkage modifications, 2'-O-methyl modifications, one or more or a series of deoxyribonucleic acid (DNA) bases (also referred to herein as "DNA extensions"), one or more or a series of ribonucleic acid (RNA) bases (also referred to herein as "RNA extensions"), or combinations thereof. In certain embodiments, the DNA extension may include a sequence set forth in Table 13. For example, in certain embodiments, the DNA extension may include a sequence set forth in SEQ ID NOs: 1235-1250. In certain embodiments, the RNA extension may include a sequence set forth in Table 13. For example, in certain embodiments, the RNA extension may include a sequence set forth in SEQ ID NOs: 1231-1234, 1251-1253. In certain embodiments, an RNP complex comprising a modified gRNA may increase editing of a target nucleic acid. In certain embodiments, RNP complexes containing modified gRNAs may increase editing leading to increased productive indels.
[0025] In one aspect, the present disclosure relates to an RNP complex comprising a Prevotella and Franciscella derived CRISPR1 (Cpf1) RNA-guided nuclease or variant thereof and a gRNA, wherein the gRNA can bind to a target site in a promoter of an HBG gene in a cell. In certain embodiments, the gRNA can be modified or unmodified. In certain embodiments, the gRNA can include one or more modifications including phosphorothioate bond modification, phosphorodithioate (PS2) bond modification, 2'-O-methyl modification, DNA extension, RNA extension, or combinations thereof. In certain embodiments, the DNA extension can include a sequence shown in Table 13. In certain embodiments, the RNA extension can include a sequence shown in Table 13. In certain embodiments, the gRNA can include a sequence shown in Table 7, 8, 11, or 12. In certain embodiments, the RNP complex can include an RNP complex shown in Table 10. For example, the RNP complex may comprise a gRNA comprising the sequence set forth in SEQ ID NO: 1051 and a Cpf1 variant protein encoded by the sequence set forth in SEQ ID NO: 1097 (RNP32, Table 10). In certain embodiments, the Cpf1 variant protein may comprise one or more modifications. In certain embodiments, the one or more modifications may include, but are not limited to, one or more mutations in the amino acid sequence of wild-type Cpf1, one or more mutations in the nucleic acid sequence of wild-type Cpf1, one or more nuclear localization signals (NLS), one or more purification tags (e.g., His tags), or a combination thereof. In certain embodiments, the Cpf1 variant protein may be encoded by a sequence set forth in SEQ ID NOs: 1000, 1001, 1008-1018, 1032, 1035-39, 1094-1097, 1107-09 (Cpf1 polypeptide sequences), or SEQ ID NOs: 1019-1021, 1110-17 (Cpf1 polynucleotide sequences).
[0026] In one aspect, the disclosure relates to a method of altering a promoter of an HBG gene in a cell, the method comprising contacting the cell with an RNP complex disclosed herein. In certain embodiments, the alteration may comprise an indel within one or more regions set forth in Table 6. In certain embodiments, the alteration may comprise an indel within a CCAAT box target region of the promoter of the HBG gene. For example, in certain embodiments, the alteration may comprise an indel within Chr11(NC_000011.10):5,249,955-5,249,987 (Table 6, region 6), Chr11(NC_000011.10):5,254,879-5,254,909 (Table 6, region 16), or a combination thereof. In certain embodiments, the RNP complex may comprise a gRNA and a Cpf1 protein. In certain embodiments, the gRNA may comprise an RNA targeting domain set forth in Table 8. In certain embodiments, the gRNA targeting domain may comprise a sequence selected from the group consisting of SEQ ID NOs: 1002, 1254, 1258, 1260, 1262, and 1264. In certain embodiments, the gRNA may comprise a gRNA sequence as set forth in Table 8. In certain embodiments, the gRNA may comprise a sequence selected from the group consisting of SEQ ID NOs: 1022, 1023, 1041-1105. In certain embodiments, the gRNA may comprise a sequence selected from the group consisting of SEQ ID NOs: 1022, 1023, 1041-1105. In certain embodiments, the gRNA may comprise a sequence selected from the group consisting of Chr11:5249973, Chr11:5249977 (HBG1); Chr11:5250042, Chr11:5250046 (HBG1); Chr11:5250055, Chr11:5250059 (HBG1); Chr11:5250179, Chr11:5250183 (HBG1); It can be configured to provide editing events at Chr11:5254897, Chr11:5254901 (HBG2); Chr11:5254897, Chr11:5254901 (HBG2); Chr11:5254966, 5254970 (HBG2); Chr11:5254979, 5254983 (HBG2) (Table 6, Table 7).
[0027] In one aspect, the disclosure relates to an isolated cell comprising an alteration in the promoter of an HBG gene generated by delivery of an RNP complex to a cell. In certain embodiments, the RNP complex may comprise a gRNA and a Cpf1 protein. In certain embodiments, the gRNA may be modified or unmodified. In certain embodiments, the gRNA may comprise one or more modifications including phosphorothioate linkage modifications, phosphorodithioate (PS2) linkage modifications, 2'-O-methyl modifications, DNA extensions, RNA extensions, or combinations thereof. In certain embodiments, the DNA extensions may comprise a sequence shown in Table 13. In certain embodiments, the RNA extensions may comprise a sequence shown in Table 13. In certain embodiments, the gRNAs may comprise a sequence shown in Tables 7, 8, 11, or 12. In certain embodiments, the RNP complexes may comprise an RNP complex shown in Table 10. For example, the RNP complex may comprise a gRNA comprising the sequence set forth in SEQ ID NO: 1051 and a Cpf1 variant protein encoded by the sequence set forth in SEQ ID NO: 1097 (RNP32, Table 10). In certain embodiments, the Cpf1 variant protein may comprise one or more modifications. In certain embodiments, the one or more modifications may include, but are not limited to, one or more mutations in the amino acid sequence of wild-type Cpf1, one or more mutations in the nucleic acid sequence of wild-type Cpf1, one or more nuclear localization signals (NLS), one or more purification tags (e.g., His tags), or a combination thereof. In certain embodiments, the Cpf1 variant protein may be encoded by a sequence set forth in SEQ ID NOs: 1000, 1001, 1008-1018, 1032, 1035-39, 1094-1097, 1107-09 (Cpf1 polypeptide sequences), or SEQ ID NOs: 1019-1021, 1110-17 (Cpf1 polynucleotide sequences).
[0028] In one aspect, the present disclosure relates to an ex vivo method of increasing the level of fetal hemoglobin (HbF) in human cells by genome editing using an RNP complex comprising a gRNA and Cpf1 RNA-guided nuclease or a variant thereof to affect a change in the promoter of the HBG gene, thereby increasing the expression of HbF. In certain embodiments, the gRNA may be modified or unmodified. In certain embodiments, the gRNA may include one or more modifications including phosphorothioate bond modifications, phosphorodithioate (PS2) bond modifications, 2'-O-methyl modifications, DNA extensions, RNA extensions, or combinations thereof. In certain embodiments, the DNA extensions may include sequences shown in Table 13. In certain embodiments, the RNA extensions may include sequences shown in Table 13. In certain embodiments, the gRNA may include sequences shown in Tables 7, 8, 11, or 12. In certain embodiments, the RNP complexes may include RNP complexes shown in Table 10. For example, the RNP complex may comprise a gRNA comprising the sequence set forth in SEQ ID NO: 1051 and a Cpf1 variant protein encoded by the sequence set forth in SEQ ID NO: 1097 (RNP32, Table 10). In certain embodiments, the Cpf1 variant protein may comprise one or more modifications. In certain embodiments, the one or more modifications may include, but are not limited to, one or more mutations in the amino acid sequence of wild-type Cpf1, one or more mutations in the nucleic acid sequence of wild-type Cpf1, one or more nuclear localization signals (NLS), one or more purification tags (e.g., His tags), or a combination thereof. In certain embodiments, the Cpf1 variant protein may be encoded by a sequence set forth in SEQ ID NOs: 1000, 1001, 1008-1018, 1032, 1035-39, 1094-1097, 1107-09 (Cpf1 polypeptide sequences), or SEQ ID NOs: 1019-1021, 1110-17 (Cpf1 polynucleotide sequences).
[0029] In one aspect, the disclosure relates to a CD34+ or hematopoietic stem cell population, where one or more cells in the population comprise an alteration in the promoter of an HBG gene, the alteration being generated by delivering an RNP complex comprising a gRNA and a Cpf1 RNA-guided nuclease or variant thereof to the CD34+ or hematopoietic stem cell population. In certain embodiments, the gRNA may be modified or unmodified. In certain embodiments, the gRNA may comprise one or more modifications comprising a phosphorothioate bond modification, a phosphorodithioate (PS2) bond modification, a 2'-O-methyl modification, a DNA extension, an RNA extension, or a combination thereof. In certain embodiments, the DNA extension may comprise a sequence as set forth in Table 13. In certain embodiments, the RNA extension may comprise a sequence as set forth in Table 13. In certain embodiments, the gRNA may comprise a sequence as set forth in Tables 7, 8, 11, or 12. In certain embodiments, the RNP complex may comprise an RNP complex as set forth in Table 10. For example, the RNP complex may comprise a gRNA comprising the sequence set forth in SEQ ID NO: 1051 and a Cpf1 variant protein encoded by the sequence set forth in SEQ ID NO: 1097 (RNP32, Table 10). In certain embodiments, the Cpf1 variant protein may comprise one or more modifications. In certain embodiments, the one or more modifications may include, but are not limited to, one or more mutations in the amino acid sequence of wild-type Cpf1, one or more mutations in the nucleic acid sequence of wild-type Cpf1, one or more nuclear localization signals (NLS), one or more purification tags (e.g., His tags), or a combination thereof. In certain embodiments, the Cpf1 variant protein may be encoded by a sequence set forth in SEQ ID NOs: 1000, 1001, 1008-1018, 1032, 1035-39, 1094-1097, 1107-09 (Cpf1 polypeptide sequences), or SEQ ID NOs: 1019-1021, 1110-17 (Cpf1 polynucleotide sequences).
[0030] In one aspect, the disclosure relates to a method of alleviating one or more symptoms of beta thalassemia in a subject in need thereof, the method comprising: a) isolating a CD34+ or hematopoietic stem cell population from the subject; b) modifying the isolated cell population ex vivo by delivering an RNP complex comprising a gRNA and a Cpf1 RNA-guided nuclease or a variant thereof to the isolated cell population, thereby affecting a change in the promoter of the HBG gene in one or more cells within the population; and c) administering the modified cell population to the subject, thereby alleviating one or more symptoms of beta thalassemia in the subject. In certain embodiments, the method may further include detecting progeny / daughter cells of the administered modified cells in the subject, e.g., in the form of BM-engrafted CD34+ hematopoietic stem cells or blood cells derived therefrom (e.g., myeloid progenitor cells or differentiated myeloid cells (e.g., erythroid, mast cells, myoblasts), or lymphoid progenitor cells or differentiated lymphoid cells (e.g., T lymphocytes or B lymphocytes, or NK cells), e.g., at least [1, 2, 3, 4, 5, 6, 7, 8, 12, 16, or 20] weeks, or at least [1, 2, 3, 4, 5, or 6] months, or at least [1, 2, 3, 4, or 5] years after administration. In certain embodiments, the method may also include detecting progeny / daughter cells of the administered modified cells, e.g., in the form of BM-engrafted CD34+ hematopoietic stem cells or blood cells derived therefrom (e.g., myeloid progenitor cells or differentiated myeloid cells (e.g., erythroid, mast cells, myoblasts), or lymphoid progenitor cells or differentiated lymphoid cells (e.g., T lymphocytes or B lymphocytes, or NK cells), e.g., at least [1, 2, 3, 4, 5, or 6] weeks, or at least [1, 2, 3, 4, or 5] years after administration. In certain embodiments, the method may also include detecting progeny / daughter cells of the administered modified cells, e.g., in the form of BM-engrafted CD34+ hematopoietic stem cells or blood cells derived therefrom (e.g., myeloid progenitor cells or differentiated myeloid cells (e.g., T lymphocytes or B lymphocytes, or NK cells In certain embodiments, the method may include administering a plurality of edited cells, and the method may result in long-term engraftment of a plurality of [at least 5, 10, 15, 20, 25, ... 100] distinct HSC clones in the BM (e.g., at least [1, 2, 3, 4, 5, 6, 7, 8, 12, 16, or 20] weeks, or at least [1, 2, 3, 4, 5, or 6] months, or at least [1, 2, 3, 4, or 5] years after administration). In certain embodiments, the method may further include detecting a level of total hemoglobin expression in the subject at least [1, 2, 3, 4, 5, 6, 7, 8, 12, 16, or 20] weeks, or at least [1, 2, 3, 4, 5, or 6] months, or at least [1, 2, 3, 4, or 5] years after administration.In certain embodiments, the method may result in prolonged expression (e.g., at least [1, 2, 3, 4, 5, 6, 7, 8, 12, 16, or 20] weeks, or at least [1, 2, 3, 4, 5, or 6] months, or at least [1, 2, 3, 4, or 5] years) of total hemoglobin (e.g., total Hb (e.g., HbA and HbF, if present, combined)) compared to healthy subjects. In certain embodiments, the alteration may include an indel within the CCAAT box target region of the promoter of the HBG gene. In certain embodiments, the RNP complex may be delivered using electroporation. In certain embodiments, at least about 5%, at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, or at least about 90% of the cells in a cell population contain productive indels.
[0031] In one aspect, the disclosure relates to a method of alleviating one or more symptoms of beta-thalassemia (β-Thal) in a subject in need thereof, comprising: a) isolating a CD34+ or hematopoietic stem cell population from the subject; b) modifying the isolated cell population ex vivo by delivering an RNP complex to the isolated cell population, thereby altering the promoter of the HBG gene in one or more isolated cells within the population, the RNP complex comprising Cpf1 and a gRNA comprising a 5'-end and a 3'-end, a DNA extension at the 5'-end, a 2'-O-methyl-3'-phosphorothioate modification at the 3'-end, and a targeting domain that is complementary to a target site in the promoter of the HBG gene; and c) administering the modified isolated cell population to the subject, thereby alleviating one or more symptoms of β-Thal in the subject. In certain embodiments, the DNA extension may comprise a sequence selected from the group consisting of SEQ ID NOs: 1235-1250. In certain embodiments, the targeting domain may comprise a sequence selected from the group consisting of those shown in Tables 7, 8, 11, and 12. In certain embodiments, the target site may comprise a nucleotide located at Chr11(NC_000011.10) 5,249,904 to 5,249,927 (Table 6, Region 6), Chr11(NC_000011.10) 5,254,879 to 5,254,909 (Table 6, Region 16), or a combination thereof. In certain embodiments, Cpf1 may comprise one or more modifications selected from the group consisting of one or more mutations in the amino acid sequence of wild-type Cpf1, one or more mutations in the nucleic acid sequence of wild-type Cpf1, one or more nuclear localization signals (NLS), one or more purification tags, and combinations thereof. In certain embodiments, Cpf1 may be a Cpf1 variant and may comprise or consist of a sequence selected from the group consisting of SEQ ID NOs: 1000, 1001, 1008-1018, 1032, 1035-39, 1094-1097, and 1107-09. In certain embodiments, Cpf1 may be a Cpf1 variant and may comprise or consist of a sequence selected from the group consisting of SEQ ID NOs: 1019-1021 and 1110-17.In certain embodiments, the RNP complexes can be delivered to cells using electroporation.
[0032] In one aspect, the disclosure relates to a method of inducing expression of hemoglobin (Hb) in a first modified cell population from a subject with beta-thalassemia (β-Thal) comprising a plurality of modified CD34+ or hematopoietic stem cells, the method comprising delivering a first RNP complex comprising a first guide RNA (gRNA) and Cpf1 to a first unmodified cell population from a subject with β-Thal comprising a plurality of unmodified CD34+ or hematopoietic stem cells to generate an indel, the first gRNA comprising a first gRNA targeting domain, each modified CD34+ or hematopoietic stem cell comprising an indel in an HBG gene promoter, the first modified cell population comprising a higher Hb level than the first unmodified cell population. In certain embodiments, the first gRNA may comprise a DNA extension comprising a sequence selected from the group consisting of SEQ ID NOs: 1235-1250. In certain embodiments, the first gRNA targeting domain may comprise a sequence selected from the group consisting of those shown in Tables 7, 8, 11, and 12. In certain embodiments, the first gRNA may comprise a targeting domain that is complementary to a target site in the promoter of the HBG gene, the target site comprising nucleotides located at Chr11(NC_000011.10) 5,249,904 to 5,249,927 (Table 6, Region 6), Chr11(NC_000011.10) 5,254,879 to 5,254,909 (Table 6, Region 16), or combinations thereof. In certain embodiments, the first RNP complex may comprise a Cpf1 variant comprising one or more modifications selected from the group consisting of one or more mutations in a wild-type Cpf1 amino acid sequence, one or more mutations in a wild-type Cpf1 nucleic acid sequence, one or more nuclear localization signals (NLS), one or more purification tags, and combinations thereof. In certain embodiments, the Cpf1 variant may comprise or consist of a sequence selected from the group consisting of SEQ ID NOs: 1000, 1001, 1008-1018, 1032, 1035-39, 1094-1097, and 1107-09. In certain embodiments, the Cpf1 variant may comprise or consist of a sequence selected from the group consisting of SEQ ID NOs: 1019-1021 and 1110-17.In certain embodiments, the first RNP complex can be delivered to a cell using electroporation.
[0033] In certain embodiments, the modified CD34+ or hematopoietic stem cells can be erythroblasts differentiated from the modified CD34+ or hematopoietic stem cells. In certain embodiments, the unmodified CD34+ or hematopoietic stem cells can be erythroblasts differentiated from the unmodified CD34+ or hematopoietic stem cells. In certain embodiments, the erythroblasts can include one or more selected from live cells, nucleated cells, cells that fluoresce using anti-human CD235a antibodies via fluorescence activated cell sorting (FACS), or combinations thereof.
[0034] In certain embodiments, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, or 50% or more of the erythroblasts differentiated from modified CD34+ or hematopoietic stem cells may be late stage erythroblasts compared to erythroblasts differentiated from unmodified CD34+ or hematopoietic stem cells. In certain embodiments, late stage erythroblasts may include cells that comprise low or negative CD71 expression.
[0035] In certain embodiments, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, or 50% or more of the erythroblasts differentiated from modified CD34+ or hematopoietic stem cells may be enucleated red blood cells compared to erythroblasts differentiated from unmodified CD34+ or hematopoietic stem cells. In certain embodiments, the enucleated red blood cells may be red blood cells that do not contain a nucleus. In certain embodiments, the enucleated red blood cells may include red blood cells that do not fluoresce (stain) when using a reagent that detects cell nuclei (e.g., NucRed reagent).
[0036] In certain embodiments, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, or 50% or more of the erythroblasts differentiated from unmodified CD34+ or hematopoietic stem cells may be non-viable erythroblasts compared to the erythroblasts differentiated from modified CD34+ or hematopoietic stem cells, in certain embodiments, the non-viable erythroblasts include cells that fluoresce (stain) with 4',6-diamidino-2-phenylindole (DAPI).
[0037] In certain embodiments, erythroblasts differentiated from modified CD34+ or hematopoietic stem cells may have a 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, or 50% higher total hemoglobin content compared to erythroblasts differentiated from unmodified CD34+ or hematopoietic stem cells. In certain embodiments, total hemoglobin content can be measured using reverse phase ultra-performance liquid chromatography (RP-UPLC).
[0038] In one aspect, the disclosure relates to a gRNA comprising a 5' end and a 3' end, comprising a DNA extension at the 5' end and a 2'-O-methyl-3'-phosphorothioate modification at the 3' end, wherein the gRNA comprises an RNA segment capable of hybridizing to a target site and an RNA segment capable of associating with a Cpf1 RNA-guided nuclease. In certain embodiments, the DNA extension may comprise a sequence as set forth in SEQ ID NOs: 1235-1250. In certain embodiments, the gRNA may be modified or unmodified. In certain embodiments, the gRNA may comprise one or more modifications including a phosphorothioate bond modification, a phosphorodithioate (PS2) bond modification, a 2'-O-methyl modification, a DNA extension, an RNA extension, or a combination thereof. In certain embodiments, the DNA extension may comprise a sequence as set forth in Table 13. In certain embodiments, the RNA extension may comprise a sequence as set forth in Table 13. In certain embodiments, the gRNA may comprise a sequence as set forth in Tables 7, 8, 11, or 12.
[0039] In one aspect, the present disclosure relates to an RNP complex comprising a Cpf1 RNA-guided nuclease as disclosed herein and a gRNA as disclosed herein.
[0040] Also provided herein are genome editing systems, guide RNAs, and CRISPR-mediated methods for altering one or more gamma-globin genes (e.g., HBG1, HBG2, or HBG1 and HBG2) to increase expression of fetal hemoglobin (HbF). In certain embodiments, one or more gRNAs comprising a sequence as set forth in Tables 7, 8, 11, or 12 can be used to introduce alterations in the promoter region of the HBG genes. In certain embodiments, the genome editing systems, guide RNAs, and CRISPR-mediated methods can alter a 13 nucleotide (nt) target region ("13 nt target region") 5' to the transcription site of the HBG1, HBG2, or HBG1 and HBG2 genes. In certain embodiments, the genome editing systems, guide RNAs, and CRISPR-mediated methods can alter a CCAAT box target region ("CCAAT box target region") 5' to the transcription site of the HBG1, HBG2, or HBG1 and HBG2 genes. In certain embodiments, the CCAAT box target region may be a region at or near the distal CCAAT box, including the nucleotides of the distal CCAAT box, and 25 nucleotides upstream (5') and 25 nucleotides downstream (3') of the distal CCAAT box (i.e., HBG1 / 2 c.-86 to -140). In certain embodiments, the CCAAT box target region may be a region at or near the distal CCAAT box, including the nucleotides of the distal CCAAT box, and 5 nucleotides upstream (5') and 5 nucleotides downstream (3') of the distal CCAAT box (i.e., HBG1 / 2 c.-106 to -120). In certain embodiments, the CCAAT box target region may include an 18 nt target region, a 13 nt target region, an 11 nt target region, a 4 nt target region, a 1 nt target region, a -117G>A target region, or a combination thereof, as disclosed herein. In certain embodiments, the alteration can be an 18 nt deletion, a 13 nt deletion, an 11 nt deletion, a 4 nt deletion, a 1 nt deletion, a G to A substitution at c.-117, or a combination thereof in the HBG1, HBG2, or HBG1 and HBG2 genes. In certain embodiments, the alteration can be a non-naturally occurring alteration or a naturally occurring alteration.
[0041] In certain embodiments, the genome editing system, guide RNA, and CRISPR-mediated method for altering one or more γ-globin genes (e.g., HBG1, HBG2, or HBG1 and HBG2) may include an RNA-guided nuclease. In certain embodiments, the RNA-guided nuclease may be Cpf1 or a modified Cpf1 as disclosed herein.
[0042] In one aspect, the present disclosure relates to a composition comprising a plurality of cells produced by the method disclosed above, wherein at least 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% of the cells comprise a sequence alteration of the 13 nt target region of the human HBG1 or HBG2 gene or the plurality of cells produced by the method disclosed above, wherein at least 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% of the cells comprise a sequence alteration of the 13 nt target region of the human HBG1 or HBG2 gene. In certain embodiments, at least a portion of the plurality of cells may be in the erythroid lineage. In certain embodiments, the plurality of cells may be characterized by an increased level of fetal hemoglobin expression compared to the plurality of cells that are not modified. In certain embodiments, the level of fetal hemoglobin may be increased by at least 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90%. In certain embodiments, the composition may further comprise a pharma- ceutically acceptable carrier.
[0043] The disclosure herein also relates to a method of altering a cell, the method comprising contacting the cell with any of the genome editing systems disclosed herein. In certain embodiments, the step of contacting the cell may comprise contacting the cell with a solution comprising the first and second ribonucleoprotein complexes. In certain embodiments, the step of contacting the cell with the solution further comprises electroporating the cell, whereby the first and second ribonucleoprotein complexes are introduced into the cell.
[0044] A genome editing system or method including any of all of the above features may include a target nucleic acid including a human HBG1, HBG2 gene, or a combination thereof. In certain embodiments, the target region may be a CCAAT box target region of a human HBG1, HBG2 gene, or a combination thereof. In certain embodiments, the first targeting domain sequence may be complementary to a first sequence on the side of the CCAAT box target region of a human HBG1, HBG2 gene, or a combination thereof, and the first sequence optionally overlaps with the CCAAT box target region of a human HBG1, HBG2 gene, or a combination thereof. In certain embodiments, the second targeting domain sequence may be complementary to a second sequence on the side of the CCAAT box target region of a human HBG1, HBG2 gene, or a combination thereof, and the second sequence optionally overlaps with the CCAAT box target region of a human HBG1, HBG2 gene, or a combination thereof.
[0045] In certain embodiments, the cells may comprise at least one modified allele at the HBG locus generated by any of the methods for altering a cell disclosed herein, wherein the modified allele at the HBG locus comprises an alteration in the human HBG1 gene, the HBG2 gene, or a combination thereof.
[0046] In certain embodiments, the isolated cell population may be modified by any of the methods for altering cells disclosed herein, such that the cell population comprises a distribution of indels that may differ from an isolated cell population of the same cell type that has not been modified by the method, or their progeny.
[0047] In certain embodiments, a plurality of cells may be generated by any of the methods for altering a cell disclosed herein, and at least 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% of the cells may contain a sequence alteration in the CCAAT box target region of the human HBG1 gene, HBG2 gene, or a combination thereof.
[0048] In certain embodiments, the cells disclosed herein may be used in medicine. In certain embodiments, the cells may be for use in the treatment of β-hemoglobinopathies. In certain embodiments, the β-hemoglobinopathies may be selected from the group consisting of sickle cell disease and beta-thalassemia. In certain embodiments, the beta-thalassemia may be transfusion-dependent beta-thalassemia (TDT).
[0049] In one aspect, the present disclosure relates to a composition comprising a plurality of cells produced by the method disclosed above, wherein at least 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% of the cells comprise a sequence alteration in the CCAAT box target region of the human HBG1 or HBG2 gene or the plurality of cells produced by the method disclosed above, wherein at least 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% of the cells comprise a sequence alteration in the CCAAT box target region of human HBG1 or HBG2. In certain embodiments, at least a portion of the plurality of cells may be in the erythroid lineage. In certain embodiments, the plurality of cells may be characterized by an increased level of fetal hemoglobin expression compared to the plurality of cells that are not modified. In certain embodiments, the level of fetal hemoglobin may be increased by at least 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90%. In certain embodiments, the composition may further comprise a pharma- ceutically acceptable carrier.
[0050] In one aspect, the disclosure relates to a cell population modified by the above genome editing system, the cell population comprises a higher percentage of productive indels compared to a cell population not modified by the genome editing system. The disclosure also relates to a cell population modified by the genome editing system, the cell population can differentiate into a cell population of erythroid lineage expressing HbF compared to a cell population not modified by the genome editing system. In certain embodiments, the higher percentage can be at least about 15%, at least about 20%, at least about 25%, at least about 30%, or at least about 40% higher. In certain embodiments, the cells can be hematopoietic stem cells. In certain embodiments, the cells can differentiate into erythroblasts, erythrocytes, or precursors of erythrocytes or erythroblasts. In certain embodiments, the indels can be generated by a repair mechanism other than microhomology-mediated end joining (MMEJ) repair.
[0051] The present disclosure also relates to the use of any of the cells disclosed herein in the manufacture of a medicament for treating a β-hemoglobinopathy in a subject.
[0052] In one aspect, the disclosure relates to a method of treating a β-hemoglobinopathy in a subject in need thereof, comprising administering to the subject a cell as disclosed herein. In certain embodiments, a method of treating a β-hemoglobinopathy in a subject in need thereof can comprise administering to the subject a modified hematopoietic cell population, wherein one or more cells have been altered according to the cell altering methods disclosed herein. In certain embodiments, the method may further include detecting progeny / daughter cells of the administered modified cells in the subject, e.g., in the form of BM-engrafted CD34+ hematopoietic stem cells or blood cells derived therefrom (e.g., myeloid progenitor cells or differentiated myeloid cells (e.g., erythroid, mast cells, myoblasts), or lymphoid progenitor cells or differentiated lymphoid cells (e.g., T lymphocytes or B lymphocytes, or NK cells), e.g., at least [1, 2, 3, 4, 5, 6, 7, 8, 12, 16, or 20] weeks, or at least [1, 2, 3, 4, 5, or 6] months, or at least [1, 2, 3, 4, or 5] years after administration. In certain embodiments, the method may also include detecting progeny / daughter cells of the administered modified cells, e.g., in the form of BM-engrafted CD34+ hematopoietic stem cells or blood cells derived therefrom (e.g., myeloid progenitor cells or differentiated myeloid cells (e.g., erythroid, mast cells, myoblasts), or lymphoid progenitor cells or differentiated lymphoid cells (e.g., T lymphocytes or B lymphocytes, or NK cells), e.g., at least [1, 2, 3, 4, 5, or 6] weeks, or at least [1, 2, 3, 4, or 5] years after administration. In certain embodiments, the method may also include detecting progeny / daughter cells of the administered modified cells, e.g., in the form of BM-engrafted CD34+ hematopoietic stem cells or blood cells derived therefrom (e.g., myeloid progenitor cells or differentiated myeloid cells (e.g., T lymphocytes or B lymphocytes, or NK cells In certain embodiments, the method may include administering a plurality of edited cells, and the method may result in long-term engraftment of a plurality of [at least 5, 10, 15, 20, 25, ... 100] distinct HSC clones in the BM (e.g., at least [1, 2, 3, 4, 5, 6, 7, 8, 12, 16, or 20] weeks, or at least [1, 2, 3, 4, 5, or 6] months, or at least [1, 2, 3, 4, or 5] years after administration). In certain embodiments, the method may further include detecting a level of total hemoglobin expression in the subject at least [1, 2, 3, 4, 5, 6, 7, 8, 12, 16, or 20] weeks, or at least [1, 2, 3, 4, 5, or 6] months, or at least [1, 2, 3, 4, or 5] years after administration.In certain embodiments, the method may result in [at least 50%, at least 60%...at least 99%] prolonged expression (e.g., at least [1, 2, 3, 4, 5, 6, 7, 8, 12, 16, or 20] weeks, or at least [1, 2, 3, 4, 5, or 6] months, or at least [1, 2, 3, 4, or 5] years after administration) of total hemoglobin (e.g., total Hb (e.g., HbA and HbF, if present, combined)) compared to healthy subjects. In certain embodiments, the alteration may include an indel within the CCAAT box target region of the promoter of the HBG gene.
[0053] In one aspect, the disclosure relates to a method of altering a cell, the method comprising contacting the cell with a genome editing system. In certain embodiments, the step of contacting the cell with the genome editing system may comprise contacting the cell with a solution comprising a first and a second ribonucleoprotein complex. In certain embodiments, the step of contacting the cell with the solution may further comprise electroporating the cell, whereby the first and the second ribonucleoprotein complexes are introduced into the cell. In certain embodiments, the method of altering a cell may further comprise contacting the cell with a genome editing system, the step of contacting the cell with the genome editing system may comprise contacting the cell with a solution comprising a first, a second, a third, and optionally a fourth ribonucleoprotein complex. In certain embodiments, the step of contacting the cell with the solution may further comprise electroporating the cell, whereby the first, a second, a third, and optionally a fourth ribonucleoprotein complexes are introduced into the cell. In certain embodiments, the cell may be capable of differentiating into an erythroblast, an erythrocyte, or a precursor of an erythrocyte or an erythroblast. In certain embodiments, the cells are CD34 + It may be a cell.
[0054] In one aspect, the present disclosure relates to a composition that may include a plurality of cells generated by the method of altering a cell disclosed herein, where at least 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% of the cells may include an alteration in the sequence of the CCAAT box target region of the human HBG1 gene, HBG2 gene, or a combination thereof. In certain embodiments, at least a portion of the plurality of cells may be in the erythroid lineage. In certain embodiments, the plurality of cells may be characterized by an increased level of fetal hemoglobin expression compared to the plurality of cells that are not modified. In certain embodiments, the level of fetal hemoglobin may be increased by at least 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90%. In certain embodiments, the composition may further include a pharmaceutically acceptable carrier.
[0055] In one aspect, the disclosure relates to a cell comprising a synthetic genotype generated by the methods of altering a cell disclosed herein, wherein the cell can comprise an 18 nt deletion, an 11 nt deletion, a 4 nt deletion, a 1 nt deletion, a 13 nt deletion, a G to A substitution at -117 in the human HBG1 gene, the HBG2 gene, or a combination thereof.
[0056] In one aspect, the disclosure relates to a cell comprising at least one allele of the HBG locus generated by the methods of altering a cell disclosed herein, wherein the cell may encode an 18 nt deletion, an 11 nt deletion, a 4 nt deletion, a 1 nt deletion, a 13 nt deletion, a G to A substitution at -117 in the human HBG1 gene, the HBG2 gene, or a combination thereof.
[0057] In one aspect, the disclosure relates to a composition comprising a cell population produced by the method of altering a cell disclosed herein, wherein the cells comprise a higher frequency of alterations in the sequence of the CCAAT box target region of the human HBG1 gene, HBG2 gene, or a combination thereof, compared to an unmodified cell population. In certain embodiments, the higher frequency is at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% higher. In certain embodiments, at least a portion of the cell population is within the erythroid lineage.
[0058] This list is intended to be exemplary and illustrative, not exhaustive and limiting. Additional aspects and embodiments may be set forth in or apparent from the remainder of the disclosure and the claims. [Brief description of the drawings]
[0059] The accompanying drawings are intended to provide illustrative and schematic examples of certain aspects and embodiments of the present disclosure, rather than being comprehensive. The drawings are not intended to be limited or constrained to any particular theory or model, and are not necessarily drawn to scale. Without limiting the above, nucleic acids and polypeptides may be depicted as linear sequences or as schematic two- or three-dimensional structures. These depictions are intended to be illustrative, rather than limiting or constraining to any particular model or theory regarding their structure.
[0060] [Figure 1]The HBG1 and HBG2 genes are shown in schematic form in the context of the β-globin gene cluster on human chromosome 11. Figure 1. Each gene of the β-globin gene cluster is transcriptionally regulated by a proximal promoter. Without wishing to be bound by any particular theory, it is generally believed that expression of Aγ and / or Gγ is activated by engagement between the proximal promoter and a distal, potent, erythroid-specific enhancer, the locus control region (LCR). Long-range transactivation by the LCR is thought to be mediated by changes in chromatin organization / confirmation. The LCR is characterized by four erythroid-specific DNase I hypersensitive sites (HS1-4) and two distal enhancer elements (5'HS and 3'HS1). β-like gene globin gene expression is regulated in a developmental stage-specific manner, with globin gene expression coinciding with changes in the primary site of blood production. [Figure 2A-2B] The HBG1 and HBG2 genes, coding sequences (CDS), and small deletions and point mutations in and upstream of the HBG1 and HBG2 proximal promoters that have been identified in patients and are associated with elevated fetal hemoglobin (HbF) are shown. A core element (CAAT box, 13 nt sequence) in the proximal promoter is deleted in some patients with genetic persistence of fetal hemoglobin (HPFH). The region of the "target sequence" of each locus that is being screened for gRNA binding target sites is also identified. [Figure 3A] Shown is the percentage of indels in CD34+ cells from three donors with TDT ("B-thal") and three normal healthy donors ("HD"), 1-3 days after electroporation with RNP32 ("Treatment"). "Mock" represents cells electroporated without RNP (unedited cells). Indel = insertion and / or deletion. [Figure 3B]Shown is the percentage of indels in CD34+ cells from one donor with TDT, 1-3 days after electroporation with RNP32 ("Treatment"). "Mock" represents cells electroporated without RNP (unedited cells). Indel = insertion and / or deletion. [Figure 3C] Percentage of viable cells in CD34+ cells from three donors with TDT 1-3 days after electroporation is shown. "Treated" (dashed line) represents cells electroporated with RNP32. "Mock" (solid line) represents cells electroporated without RNP (unedited cells). [Figure 4A] Shown is the percentage of CD235a+ cells from three donors with TDT on day 18 in erythroid cultures. "RNP32" represents erythroblasts differentiated from RNP32-edited CD34+ cells from donors with TDT. "Mock" represents erythroblasts differentiated from cells electroporated without RNP (unedited cells). N=3 independent donors with triplicate cultures. *p<0.05. [Figure 4B] Shown are the percentages of CD235a+ cells from donor 2 on days 7, 11, 14, and 18 in erythroid cultures. "RNP32" (dashed line) represents erythroblasts differentiated from RNP32-edited CD34+ cells from a donor with TDT. "Mock" (solid line) represents erythroblasts differentiated from cells electroporated without RNP (unedited cells). N=3 independent donors with triplicate cultures. [Figure 4C] Percentage of erythroblasts reaching the late erythroblast stage is shown. "RNP32" represents erythroblasts differentiated from RNP32-edited CD34+ cells from donors with TDT. "Mock" represents erythroblasts differentiated from cells electroporated without RNP (unedited cells). N=3 independent donors with triplicate cultures. *p<0.05, **p<0.01, ***p<0.001, ****p<0.0001. [Figure 4D]Percentage of red blood cells that underwent terminal maturation and enucleation is shown. "RNP32" represents erythroblasts differentiated from RNP32-edited CD34+ cells from donors with TDT. "Mock" represents erythroblasts differentiated from cells electroporated without RNP (unedited cells). N=3 independent donors with triplicate cultures. *p<0.05, **p<0.01, ***p<0.001, ****p<0.0001. [Figure 4E] The frequency of cell death of erythroblasts (i.e., % of non-viable erythroblasts) is shown. "RNP32" represents erythroblasts differentiated from RNP32-edited CD34+ cells from donors with TDT. "Mock" represents erythroblasts differentiated from cells electroporated without RNP (unedited cells). N=3 independent donors with triplicate cultures. *p<0.05, **p<0.01, ***p<0.001, ****p<0.0001. [Figure 4F] Shown are the percentages of donor 2-derived erythroblasts on days 7, 11, 14, and 18 in erythroid cultures that reached the late erythroblast stage. Dashed lines represent erythroblasts differentiated from RNP32-edited CD34+ cells from donor 2 with TDT. Solid lines represent erythroblasts differentiated from cells electroporated without RNP (unedited cells). N=1 independent donor with triplicate cultures. *p<0.05, **p<0.01, ***p<0.001, ****p<0.0001. [Figure 4G] Shown are the percentages of donor 2-derived erythroid cells in erythroid cultures that underwent terminal maturation and enucleation on days 7, 11, 14, and 18. Dashed lines represent erythroblasts differentiated from RNP32-edited CD34+ cells from donor 2 with TDT. Solid lines represent erythroblasts differentiated from cells electroporated without RNP (unedited cells). N=1 independent donor with triplicate cultures. *p<0.05, **p<0.01, ***p<0.001, ****p<0.0001. [Figure 4H]Frequency of cell death (i.e., percentage of non-viable erythroblasts) of erythroblasts from donor 2 on days 7, 11, 14, and 18 in erythroid culture. Dashed lines represent erythroblasts differentiated from RNP32-edited CD34+ cells from donor 2 with TDT. Solid lines represent erythroblasts differentiated from cells electroporated without RNP (unedited cells). N=1 independent donor with triplicate cultures. *p<0.05, **p<0.01, ***p<0.001, ****p<0.0001. [Figure 5A] HBG / GAPDH mRNA content of erythroblasts differentiated from RNP32-edited and unedited CD34+ cells from three donors with TDT ("Donor 1", "Donor 2", "Donor 3"). Data from erythroblasts differentiated from RNP32-edited CD34+ cells are shown on the right for each donor, and data from erythroblasts differentiated from cells electroporated without RNP (unedited cells) are shown on the left for each donor. N=3 independent donors (from cultures of three technical replicates). *p<0.05, **p<0.01, ***p<0.001. GAPDH: glyceraldehyde-3-phosphate dehydrogenase, HBG: γ-globin. [Figure 5B] Shown are gamma-globin protein content (picograms (pg) per cell) of erythroblasts differentiated from RNP32-edited and non-edited CD34+ cells from three donors ("Donor 1", "Donor 2", "Donor 3") with TDT. Data from erythroblasts differentiated from RNP32-edited CD34+ cells are shown on the right for each donor, and data from erythroblasts differentiated from cells electroporated without RNP (unedited cells) are shown on the left for each donor. N=3 independent donors (from cultures of 6 technical replicates). *p<0.05, **p<0.01, ***p<0.001. [Figure 5C]Total globin / GAPDH mRNA content of erythroblasts differentiated from RNP32-edited and unedited CD34+ cells from three donors with TDT ("Donor 1", "Donor 2", "Donor 3"). Data from erythroblasts differentiated from RNP32-edited CD34+ cells are shown on the right for each donor, and data from erythroblasts differentiated from cells electroporated without RNP (unedited cells) are shown on the left for each donor. N=3 independent donors (from cultures of 3 technical replicates). *p<0.05, **p<0.01, ***p<0.001. [Figure 5D] Total hemoglobin protein content per cell of erythroblasts differentiated from RNP32-edited and unedited CD34+ cells from three donors with TDT ("Donor 1", "Donor 2", "Donor 3"). Data from erythroblasts differentiated from RNP32-edited CD34+ cells are shown on the right for each donor, and data from erythroblasts differentiated from cells electroporated without RNP (unedited cells) are shown on the left for each donor. N=3 independent donors (from cultures of 6 technical replicates). *p<0.05, **p<0.01, ***p<0.001. [Figure 5E] Total hemoglobin protein content per cell of erythroblasts differentiated from RNP32-edited and unedited CD34+ cells from three donors with TDT ("Thal donor 1", "Thal donor 2", "Thal donor 3"). "RNP32" represents erythroblasts differentiated from RNP32-edited CD34+ cells from donors with TDT. "Mock" represents erythroblasts differentiated from cells electroporated without RNP (unedited cells). Total hemoglobin production was measured and assessed using reversed-phase ultra-performance liquid chromatography (RP-UPLC). [Figure 6]The sequences of the Cpf1 protein variants shown in Table 9 are shown. The nuclear localization sequence is shown in bold and the six histidine sequence is shown in underlined text. Further substitutions of the identity of the NLS sequence and the N-terminal / C-terminal positions, such as combinations of two or more nNLS sequences, or nNLS sequences and sNLS sequences (or other NLS sequences), as well as the addition of sequences with or without purification sequences (e.g., six histidine sequences) at either the N-terminal / C-terminal positions, are within the scope of the subject matter of this disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0061] Definitions and Abbreviations Unless otherwise stated, each of the following terms has the meaning associated with it in this section.
[0062] The indefinite articles "a" and "an" refer to at least one of the associated noun and are used interchangeably with the terms "at least one" and "one or more." For example, "a module" means at least one module or one or more modules.
[0063] The conjunctions "or" and "and / or" are used interchangeably as non-exclusive disjunctions.
[0064] "Domain" is used to describe a segment of a protein or nucleic acid. Unless otherwise indicated, a domain is not required to have any particular functional properties.
[0065] "Productive indels" refer to indels (deletions and / or insertions) that result in HbF expression. In certain embodiments, productive indels can induce HbF expression. In certain embodiments, productive indels can result in increased levels of HbF expression.
[0066] "Indel" is an insertion and / or deletion in a nucleic acid sequence. Indel can be the product of repair of DNA double-strand breaks, such as the double-strand breaks formed by the genome editing system of the present disclosure. Indels are most commonly formed when breaks are repaired by "error-prone" repair pathways, such as the NHEJ pathway described below.
[0067] "Gene conversion" refers to the change in DNA sequence by the incorporation of an endogenous homologous sequence (e.g., a homologous sequence in a gene array). "Gene correction" refers to the change in DNA sequence by the incorporation of an exogenous homologous sequence, such as an exogenous single-stranded or double-stranded donor template DNA. Gene conversion and gene correction are the products of repair of DNA double-strand breaks by HDR pathways, such as those described below.
[0068] Indels, gene conversions, gene corrections, and other genome editing results are typically assessed by sequencing (most commonly by "next generation" or "sequencing by synthesis" methods, although Sanger sequencing can still be used) and quantified by the relative frequency of numerical changes (e.g., ±1, ±2, or more bases) at the site of interest among all sequencing reads. DNA samples for sequencing can be prepared by a variety of methods known in the art, and may involve amplifying the site of interest by polymerase chain reaction (PCR), capturing DNA ends generated by double-strand breaks, as in the GUIDEseq process described in Tsai 2016 (herein incorporated by reference), or by other means known in the art. Genome editing results can also be assessed by in situ hybridization methods, such as the FiberComb™ system commercialized by Genomic Vision (Bagneux, France), and any other suitable methods known in the art.
[0069] "Alt-HDR", "alternative homology-directed repair", or "alternative HDR" are used interchangeably to refer to the process of repairing DNA damage using homologous nucleic acids (e.g., endogenous homologous sequences (e.g., sister chromatids) or exogenous nucleic acids (e.g., template nucleic acids)). Alt-HDR differs from canonical HDR in that it utilizes a different pathway than canonical HDR and can be inhibited by the canonical HDR mediators RAD51 and BRCA2. Alt-HDR is also distinguished by the involvement of single-stranded or nicked homologous nucleic acid templates, whereas canonical HDR generally involves double-stranded homologous templates.
[0070] "Canonical HDR", "canonical homology directed repair" or "cHDR" refers to a process of repairing DNA damage using a homologous nucleic acid (e.g., an endogenous homologous sequence (e.g., sister chromatid) or an exogenous nucleic acid (e.g., a template nucleic acid)). Canonical HDR typically operates when a significant excision occurs at a double-stranded break, forming at least one single-stranded portion of DNA. In normal cells, cHDR typically involves a series of steps including recognition of the break, stabilization of the break, resection, stabilization of single-stranded DNA, formation of a DNA crossover intermediate, division of the crossover intermediate, and ligation. This process requires RAD51 and BRCA2, and the homologous nucleic acid is typically double-stranded.
[0071] Unless otherwise specified, the term "HDR" as used herein encompasses both canonical HDR and alt-HDR.
[0072] "Non-homologous end joining" or "NHEJ" refers to ligation-mediated repair and / or non-template-mediated repair, including canonical NHEJ (cNHEJ) and alternative NHEJ (altNHEJ), which includes microhomology-mediated end joining (MMEJ), single-strand annealing (SSA), and synthesis-dependent microhomology-mediated end joining (SD-MMEJ).
[0073] "Substitution" or "substituted," when used in reference to modification of a molecule (e.g., a nucleic acid or protein), does not require any restriction of the process, but simply indicates that a replacement entity is present.
[0074] "Subject" means a human, mouse, or non-human primate. A human subject can be of any age (e.g., infant, child, adolescent, or adult) and can be afflicted with a disease or in need of a genetic alteration.
[0075] "Treat", "treating", and "treatment" refer to the treatment of a disease in a subject (e.g., a human subject) and include one or more of inhibiting the disease (i.e., halting or preventing its onset or progression), relieving the disease (i.e., causing regression of the disease state), alleviating one or more symptoms of the disease, and curing the disease.
[0076] "Prevent," "preventing," and "prevention" refer to the prevention of a disease in a subject, and include (a) avoiding or eliminating a disease, (b) affecting a predisposition to a disease, or (c) preventing or delaying the onset of at least one symptom of a disease.
[0077] A "kit" refers to any collection of two or more components that together constitute a functional unit that can be used for a particular purpose. By way of example (and not limitation), a kit according to the present disclosure can include a guide RNA that is complexed or can be complexed with an RNA-guided nuclease and is associated with a pharma- ceutically acceptable carrier (e.g., suspended or suspendable). In certain embodiments, the kit can include a booster element. The kit can be used, for example, to introduce the complex into a cell or subject for the purpose of causing a desired genomic change in such cell or subject. The components of the kit can be packaged together or separately. A kit according to the present disclosure also optionally includes a instructions for use (DFU) that describes the use of the kit, for example, according to the methods of the present disclosure. The DFU can be physically packaged with the kit or can be made available to a user of the kit, for example, by electronic means.
[0078] The terms "polynucleotide", "nucleotide sequence", "nucleic acid", "nucleic acid molecule", "nucleic acid sequence", and "oligonucleotide" refer to a series of nucleotide bases (also called "nucleotides") in DNA and RNA, meaning any chain of two or more nucleotides. Polynucleotides, nucleotide sequences, nucleic acids, etc., can be chimeric mixtures, or derivatives, or modified versions thereof, single-stranded or double-stranded. They can be modified at the base moiety, sugar moiety, or phosphate backbone, for example, to improve the stability of the molecule, its hybridization parameters, etc. Nucleotide sequences typically carry genetic information, including, but not limited to, information used by cellular machinery to make proteins and enzymes. These terms include double-stranded or single-stranded genomic DNA, RNA, any synthetic and genetically engineered polynucleotides, and both sense and antisense polynucleotides. These terms also include nucleic acids containing modified bases.
[0079] Conventional IUPAC notation is used in the nucleotide sequences presented herein, as shown in Table 1 below (see also Cornish-Bowden A, Nucleic Acids Res. 1985 May 10;13(9):3021-30, incorporated herein by reference). Note, however, that "T" indicates "thymine or uracil" where the sequence can be encoded by either DNA or RNA, e.g., in the gRNA targeting domain.
[0080] [Table 1]
[0081] The terms "protein", "peptide" and "polypeptide" are used interchangeably to refer to a continuous chain of amino acids linked together through peptide bonds. These terms include individual proteins, groups or complexes of proteins associated together, as well as fragments or portions of such proteins, variants, derivatives and analogs. Peptide sequences are presented herein using conventional notation, starting from the amino or N-terminus on the left and proceeding toward the carboxyl or C-terminus on the right. Standard one-letter or three-letter abbreviations can be used.
[0082] Designations such as "CCAAT box target region" refer to sequences 5' to the transcription start site (TSS) of the HBG1 and / or HBG2 genes. The CCAAT box is a highly conserved motif within the promoter regions of alpha-like and beta-like globin genes. Regions within or near the CCAAT box play important roles in globin gene regulation. For example, the gamma-globin distal CCAAT box is associated with genetic persistence of fetal hemoglobin. Several transcription factors, such as NF-Y, COUP-TFII (NF-E3), CDP, GATA1 / NF-E1, and DRED (Martyn 2017), have been reported to bind to overlapping CCAAT box regions in the gamma-globin promoter. Without wishing to be bound by theory, it is believed that the binding site of the transcriptional activator NF-Y overlaps with a transcriptional repressor at the gamma-globin promoter. HPFH mutations present within the distal γ-globin promoter region (e.g., within or near the CCAAT box) may alter competitive binding of these factors and thus contribute to increased γ-globin expression and elevated HbF levels. The genomic locations of HBG1 and HBG2 provided herein are based on the coordinates provided in NCBI reference sequence NC_000011, "Homo sapiens chromosome 11, GRCh38.p12 Primary Assembly" (version NC_000011.10). The distal CCAAT boxes of HBG1 and HBG2 are located at c.-111 to -115 of HBG1 and HBG2 (genomic locations are Hg38 Chr11:5,249,968 to Chr11:5,249,972 and Hg38 Chr11:5,254,892 to Chr11:5,254,896, respectively). The c.-111 to -115 region of HBG1 is exemplified by positions 2823 to 2827 of SEQ ID NO: 902 (HBG1), and the c.-111 to -115 region of HBG2 is exemplified by positions 2747 to 2751 of SEQ ID NO: 903 (HBG2).In certain embodiments, the "CCAAT box target region" refers to a region at or near the distal CCAAT box, including the nucleotides of the distal CCAAT box, as well as 25 nucleotides upstream (5') and downstream (3') of the distal CCAAT box (i.e., c.-86 to -140 of HBG1 / 2) (genomic locations Hg38 Chr11:5249943 to Hg38 Chr11:5249997 and Hg38 Chr11:5254867 to Hg38 Chr11:5254921, respectively). The c.-86 to -140 region of HBG1 is exemplified at positions 2798 to 2852 of SEQ ID NO: 902 (HBG1), and the c.-86 to -140 region of HBG2 is exemplified at positions 2723 to 2776 of SEQ ID NO: 903 (HBG2). In another embodiment, a "CCAAT box target region" refers to a region at or near the distal CCAAT box, including the nucleotides of the distal CCAAT box, as well as 5 nucleotides upstream (5') and 5 nucleotides downstream (3') of the distal CCAAT box (i.e., c.-106 to -120 of HBG1 / 2 (genomic locations Hg38 Chr11:5249963 to Hg38 Chr11:5249977 (HGB1 and Hg38 Chr11:5254887 to Hg38 Chr11:5254887 (HGB1 and ... Chr11:5254901). The c.-106 to -120 region of HBG1 is exemplified at positions 2818 to 2832 of SEQ ID NO: 902 (HBG1), and the c.-106 to -120 region of HBG2 is exemplified at positions 2742 to 2756 of SEQ ID NO: 903 (HBG2). Terms such as "alteration of a CCAAT box target site" refer to an alteration (e.g., deletion, insertion, mutation) of one or more nucleotides in a CCAAT box target region. Examples of exemplary alterations in a CCAAT box target region include, but are not limited to, a 1 nt deletion, a 4 nt deletion, an 11 nt deletion, a 13 nt deletion, and an 18 nt deletion, and a -117 G>A alteration. As used herein, the terms "CCAAT box" and "CAAT box" may be used interchangeably.
[0083] The terms "c.-114 to -102 region", "c.-102 to -114 region", "-102:-114", "13 nt target region" and the like refer to sequences located 5' to the transcription start site (TSS) of the HBG1 and / or HBG2 genes at genomic positions Hg38 Chr11:5,249,959 to Hg38 Chr11:5,249,971 and Hg38 Chr11:5,254,883 to Hg38 Chr11:5,254,895, respectively. The c.-102 to -114 region of HBG1 is exemplified by positions 2824 to 2836 of SEQ ID NO: 902 (HBG1), and the c.-102 to -114 region of HBG2 is exemplified by positions 2748 to 2760 of SEQ ID NO: 903 (HBG2). The term "13 nt deletion" or the like refers to a deletion of the 13 nt target region.
[0084] The terms "c.-121 to -104 region", "c.-104 to -121 region", "-104:-121", "18 nt target region" and the like refer to sequences located 5' to the transcription start site (TSS) of the HBG1 and / or HBG2 genes at genomic positions Hg38 Chr11:5,249,961 to Hg38 Chr11:5,249,978 and Hg38 Chr11:5,254,885 to Hg38 Chr11:5,254,902, respectively. The c.-104 to -121 region of HBG1 is exemplified by positions 2817 to 2834 of SEQ ID NO: 902 (HBG1), and the c.-104 to -121 region of HBG2 is exemplified by positions 2741 to 2758 of SEQ ID NO: 903 (HBG2). The term "18 nt deletion" and the like refers to a deletion of the 18 nt target region.
[0085] The terms "c.-105 to -115 region", "c.-115 to -105 region", "-105:-115", "11 nt target region" and the like refer to sequences located 5' to the transcription start site (TSS) of the HBG1 and / or HBG2 genes at genomic positions Hg38 Chr11:5,249,962 to Hg38 Chr11:5,249,972 and Hg38 Chr11:5,254,886 to Hg38 Chr11:5,254,896, respectively. The c.-105 to -115 region of HBG1 is exemplified by positions 2823 to 2833 of SEQ ID NO: 902 (HBG1), and the c.-105 to -115 region of HBG2 is exemplified by positions 2747 to 2757 of SEQ ID NO: 903 (HBG2). The term "11 nt deletion" and the like refers to a deletion of the 11 nt target region.
[0086] The terms "c.-115 to -112 region", "c.-112 to -115 region", "-112:-115", "4 nt target region" and the like refer to sequences located 5' to the transcription start site (TSS) of the HBG1 and / or HBG2 genes at genomic positions Hg38 Chr11:5,249,969 to Hg38 Chr11:5,249,972 and Hg38 Chr11:5,254,893 to Hg38 Chr11:5,254,896, respectively. The c.-112 to -115 region of HBG1 is exemplified by positions 2823 to 2826 of SEQ ID NO:902, and the c.-112 to -115 region of HBG2 is exemplified by positions 2747 to 2750 of SEQ ID NO:903 (HBG2). The term "4 nt deletion" and the like refers to a deletion of the 4 nt target region.
[0087] Designations such as "c.-116 region", "HBG-116", "1 nt target region" and the like refer to sequences 5' to the transcription start site (TSS) of the HBG1 and / or HBG2 genes at genomic locations Hg38 Chr11:5,249,973 and Hg38 Chr11:5,254,897, respectively. The c.-116 region of HBG1 is exemplified at position 2822 of SEQ ID NO:902, and the c.-116 region of HBG2 is exemplified at position 2746 of SEQ ID NO:903 (HBG2). Terms such as "1 nt deletion" refer to a deletion of the 1 nt target region.
[0088] The designations "c.-117 G>A region", "HBG-117 G>A", "-117 G>A target region" and the like refer to sequences 5' to the transcription start site (TSS) of the HBG1 and / or HBG2 genes at genomic locations Hg38 Chr11:5,249,974 to Hg38 Chr11:5,249,974 and Hg38 Chr11:5,254,898 to Hg38 Chr11:5,254,898, respectively. The c.-117 G>A region of HBG1 is exemplified by the guanine (G) to adenine (A) substitution at position 2821 of SEQ ID NO:902, and the c.-117 G>A region of HBG2 is exemplified by the G to A substitution at position 2745 of SEQ ID NO:903 (HBG2). Terms such as "-117 G>A change" refer to a G to A substitution at the -117 G>A target region.
[0089] The term "proximal HBG1 / 2 promoter target sequence" refers to a region within 50, 100, 200, 300, 400, or 500 bp of the proximal HBG1 / 2 promoter sequence that includes the 13 nt target region. Alterations made by a genome editing system according to the present disclosure facilitate (e.g., tend to cause, promote, or increase the likelihood of) upregulation of HbF production in erythroid progeny.
[0090] When ranges are provided herein, the endpoints are included. Furthermore, unless otherwise indicated or clear from the context and / or understanding of one of ordinary skill in the art, it should be understood that values expressed as ranges can assume any specific value within the range described in different embodiments (to the tenth of the unit of the lower limit of the range, unless otherwise clearly indicated by the context and / or understanding of one of ordinary skill in the art). It should also be understood that unless otherwise indicated or clear from the context and / or understanding of one of ordinary skill in the art, values expressed as ranges can assume any subrange within the given range, with the endpoints of the subrange being expressed with the same precision as the tenth of the unit of the lower limit of the range.
[0091] overview Various embodiments of the present disclosure generally relate to genome editing systems configured to introduce changes (e.g., deletions or insertions, or other mutations) into chromosomal DNA that enhance transcription of the HBG1 and / or HBG2 genes, which encode the Aγ and Gγ subunits of hemoglobin, respectively. In certain embodiments, increasing expression of one or more γ-globin genes (e.g., HBG1, HBG2) using the methods provided herein results in preferential formation of HbF over HbA, and / or increased HbF levels as a percentage of total hemoglobin. In certain embodiments, the present disclosure generally relates to the use of an RNP complex that includes a gRNA complexed to a Cpf1 molecule. In certain embodiments, the gRNA may be unmodified or modified, and the Cpf1 molecule may be a wild-type Cpf1 protein or a modified Cpf1 protein. In certain embodiments, the gRNA may include a sequence as set forth in Tables 12, 13, 16, or 17. In certain embodiments, the modified Cpf1 can be encoded by a sequence set forth in SEQ ID NO: 1000, 1001, 1008-1018, 1032, 1035-39, 1094-1097, 1107-09 (Cpf1 polypeptide sequences), or SEQ ID NO: 1019-1021, 1110-17 (Cpf1 polynucleotide sequences). In certain embodiments, the RNP complex can comprise an RNP complex set forth in Table 10. For example, the RNP complex can comprise a gRNA comprising a sequence set forth in SEQ ID NO: 1051 and a modified Cpf1 protein (RNP32, Table 10) encoded by a sequence set forth in SEQ ID NO: 1097.
[0092] Previously, patients with the condition of genetic persistence of fetal hemoglobin (HPFH) have been shown to contain mutations in the γ-globin regulatory element that result in expression of fetal γ-globin throughout life, rather than being suppressed around birth (Martyn 2017). This results in elevated fetal hemoglobin (HbF) expression. HPFH mutations can be deletional or non-deletional (e.g., point mutations). Subjects with HPFH exhibit lifelong expression of HbF, i.e., they are not associated with symptoms of anemia and do not undergo globin switching, or undergo only partial globin switching.
[0093] Expression of HbF can be induced through point mutations in the γ-globin regulatory elements associated with naturally occurring HPFH variants, including, for example, c.-114 C>T, c.-117 G>A, c.-158 C>T, c.-167 C>T, c.-170 G>A, c.-175 T>G, c.-175 T>C, c.-195 C>G, c.-196 C>T, c.-197 C>T, c.-198 T>C, c.-201 C>T, c.-202 C>T, c.-211 C>T, c.-251 T>C, or c.-499 T>A in HBG1, or c.-109 G>T, c.-110 A>C, c.-114 C>A, c.-114 C>T, c.-114 C>G, c.-157 C>T, c.-158 C>T, c.-167 C>T, c.-167 C>A, c.-175 T>C, c.-197 C>T, c.-200+C, c.-202 C>G, c.-211 C>T, c.-228 T>C, c.-255 C>G, c.-309 A>G, c.-369 C>G, or c.-567 T>G.
[0094] Naturally occurring mutations in the distal CCAAT box motif found within the promoters of the HBG1 and / or HBG2 genes (i.e., HBG1 / 2 c.-111 to -115) have also been shown to result in continued γ-globin expression and HPFH conditions. It is believed that the alterations (mutations or deletions) in the CCAAT box may disrupt the binding of one or more transcriptional repressors, resulting in continued expression of the γ-globin genes and elevated HbF expression (Martyn 2017). For example, the naturally occurring 13 base pair del c.-114 to -102 (the "13 nt deletion") has been shown to be associated with elevated levels of HbF (Martyn 2017). The distal CCAAT box likely overlaps with binding motifs within and around the CCAAT box of negative regulatory transcription factors expressed during adulthood that repress HBG (Martyn 2017).
[0095] The gene editing strategy disclosed herein is to increase HbF expression by disrupting one or more nucleotides within and / or surrounding the distal CCAAT box. In certain embodiments, the "CCAAT box target region" may be a region at or near the distal CCAAT box, including the nucleotides of the distal CCAAT box, as well as 25 nucleotides upstream (5') and 25 nucleotides downstream (3') of the distal CCAAT box (i.e., HBG1 / 2 c.-86 to -140). In other embodiments, the "CCAAT box target region" may be a region at or near the distal CCAAT box, including the nucleotides of the distal CCAAT box, as well as 5 nucleotides upstream (5') and 5 nucleotides downstream (3') of the distal CCAAT box (i.e., HBG1 / 2 c.-106 to -120).
[0096] Disclosed herein are unique, non-naturally occurring alterations of the CCAAT box target region that induce expression of HBGs, including, but not limited to, HBG del c.-104 to -121 ("18 nt deletion"), HBG del c.105 to -115 ("11 nt deletion"), HBG del c.112 to -115 ("4 nt deletion"), and HBG del c.-116 ("1 nt deletion"). In certain embodiments, the genome editing system disclosed herein can be used to introduce alterations in the CCAAT box target region of HBG1 and / or HBG2. In certain embodiments, the genome editing system can include an RNA-guided nuclease, including Cas9, modified Cas9, Cpf1, or modified Cpf1. In certain embodiments, the genome editing system can include an RNP that includes a gRNA and a Cpf1 molecule. In certain embodiments, the gRNA may be unmodified or modified, and the Cpf1 molecule may be a wild-type Cpf1 protein, or a modified Cpf1 protein, or a combination thereof. In certain embodiments, the gRNA may comprise a sequence set forth in Tables 7, 8, 11, or 12. In certain embodiments, the modified Cpf1 may be encoded by a sequence set forth in SEQ ID NOs: 1000, 1001, 1008-1018, 1032, 1035-39, 1094-1097, 1107-09 (Cpf1 polypeptide sequences), or SEQ ID NOs: 1019-1021, 1110-17 (Cpf1 polynucleotide sequences). In certain embodiments, the RNP complex may comprise an RNP complex set forth in Table 10. For example, the RNP complex may comprise a gRNA comprising the sequence set forth in SEQ ID NO: 1051 and a modified Cpf1 protein (RNP32, Table 10) encoded by the sequence set forth in SEQ ID NO: 1097.
[0097] Genome editing systems of the present disclosure may include an RNA-guided nuclease, such as Cpf1, and one or more gRNAs having a targeting domain that is complementary to a sequence within or near the target region, and optionally one or more DNA donor templates that encode specific mutations (such as deletions or insertions) within or near the target region, and / or agents that increase the efficiency with which such mutations are generated (including, but not limited to, random oligonucleotides, small molecule agonists or antagonists of gene products involved in DNA repair or DNA damage response, or peptide agents).
[0098] In embodiments of the present disclosure, various approaches to the introduction of mutations into the CCAAT box target region, the 13nt target region, and / or the proximal HBG1 / 2 promoter target sequence can be used. In one approach, a single change, such as a double-strand break, is made within the CCAAT box target region, the 13nt target region, and / or the proximal HBG1 / 2 promoter target sequence, and is repaired to disrupt the function of the region, for example, by the formation of an indel or by the incorporation of a donor template sequence that codes for the deletion of the region. In a second approach, two or more changes are made on either side of the region, resulting in the deletion of the intervening sequence that includes the CCAAT box target region and / or the 13nt target region.
[0099] Treatment of hemoglobinopathies by gene therapy and / or genome editing is complicated by the fact that cells, red blood cells or RBCs, phenotypically affected by the disease are enucleated and do not contain any of the genetic material encoding the abnormal hemoglobin protein (Hb) subunits or the Aγ or Gγ subunits targeted in the exemplary genome editing approaches described above. This complexity is addressed in certain embodiments of the present disclosure by altering cells that are differentiated into or that could otherwise give rise to red blood cells. Cells within the erythroid lineage that are altered according to various embodiments of the present disclosure include, but are not limited to, hematopoietic stem and progenitor cells (HSCs), erythroblasts (including basophilic, polychromatic, and / or normochromatic erythrocytes), proerythroblasts, polychromatic erythrocytes or reticulocytes, embryonic stem cells (ES), and / or induced pluripotent stem cells (iPSCs). These cells can be altered in situ (e.g., in a tissue of a subject) or ex vivo. Implementation of genome editing systems for in situ and ex vivo alteration of cells is described below under the heading "Implementation of genome editing systems: delivery, formulations, and routes of administration."
[0100] In certain embodiments, the alteration resulting in the induction of expression of Aγ and / or Gγ is obtained by the use of a genome editing system comprising an RNA-guided nuclease and at least one gRNA having a targeting domain complementary to a sequence within or adjacent to the CCAAT box target region of HBG1 and / or HBG2 (e.g., within 10, 20, 30, 40, or 50, 100, 200, 300, 400, or 500 bases of the CCAAT box target region). As discussed in more detail below, the RNA-guided nuclease and the gRNA associate with the CCAAT box target region or a region adjacent thereto to form a complex that can be altered. Examples of suitable gRNAs and gRNA targeting domains directed to the CCAAT box target region or a region adjacent thereto of HBG1 and / or HBG2 for use in the embodiments disclosed herein include those set forth herein.
[0101] In certain embodiments, the alteration resulting in the induction of expression of Aγ and / or Gγ is obtained by the use of a genome editing system comprising an RNA-guided nuclease and at least one gRNA having a targeting domain complementary to a sequence within or adjacent to the 13 nt target region of HBG1 and / or HBG2 (e.g., within 10, 20, 30, 40, or 50, 100, 200, 300, 400, or 500 bases of the 13 nt target region). As discussed in more detail below, the RNA-guided nuclease and the gRNA associate with the 13 nt target region or a region adjacent thereto to form a complex that can be altered. Examples of suitable gRNAs and gRNA targeting domains directed to the 13 nt target region or a region adjacent thereto of HBG1 and / or HBG2 for use in the embodiments disclosed herein include those set forth herein.
[0102] The genome editing system can be implemented in various ways, as discussed in detail below. As an example, the genome editing system of the present disclosure can be implemented as a ribonucleoprotein complex or multiple complexes in which multiple gRNAs are used. The ribonucleoprotein complex can be introduced into a target cell using methods known to those skilled in the art, including electroporation, as described in commonly assigned International Patent Publication No. 2016 / 182959, published November 17, 2016 by Jennifer Gori ("Gori"), which is incorporated herein by reference in its entirety.
[0103] The ribonucleoprotein complexes within these compositions are introduced into target cells by methods known to those of skill in the art, including, but not limited to, electroporation (e.g., using the Nucleofection™ technology commercialized by Lonza (Basel, Switzerland) or a similar technology commercialized by, e.g., Maxcyte Inc. (Gaithersburg, Maryland)) and lipofection (e.g., using the Lipofectamine™ reagent commercialized by Thermo Fisher Scientific (Waltham, Massachusetts)). Alternatively, or additionally, the ribonucleoprotein complexes are formed within the target cells themselves after introduction of nucleic acids encoding the RNA-guided nuclease and / or gRNA. These and other delivery modalities are described in general terms below and in Gori.
[0104] Cells altered ex vivo according to the present disclosure can be manipulated (e.g., expanded, passaged, frozen, differentiated, dedifferentiated, transduced with a transgene, etc.) prior to their delivery to a subject. The cells are variously delivered to the subject from whom they were obtained (in "autologous" transplantation) or to a recipient that is immunologically distinct from the donor of the cells (in "allogeneic" transplantation).
[0105] In some cases, autologous transplantation involves obtaining a plurality of cells from a subject, either cells circulating in peripheral blood or cells within bone marrow or other tissues (e.g., spleen, skin, etc.), and manipulating those cells to enrich for cells of erythroid lineage (e.g., by induction to generate iPSCs, purification of cells expressing specific cell surface markers such as CD34, CD90, CD49f, and / or not expressing surface markers characteristic of non-erythroid lineages such as CD10, CD14, CD38, etc.). Prior to transduction with a genome editing system targeting the CCAAT box target region, the 13 nt target region, and / or the proximal HBG1 / 2 promoter target sequence, the cells are optionally or additionally expanded, transduced with a transgene, exposed to cytokines or other peptide or small molecule agents, and / or frozen / thawed. The genome editing system can be implemented or delivered to the cells in any suitable format, including as a ribonucleoprotein complex, as separated protein and nucleic acid components, and / or as nucleic acids encoding the components of the genome editing system.
[0106] In certain embodiments, CD34+ hematopoietic stem and progenitor cells (HSPCs) edited using the genome editing methods disclosed herein can be used for treatment of a hemoglobinopathy in a subject in need of treatment. In certain embodiments, the hemoglobinopathy can be severe sickle cell disease (SCD) or thalassemia (e.g., β-thalassemia, δ-thalassemia, or β / δ-thalassemia). In certain embodiments, an exemplary protocol for treatment of a hemoglobinopathy can include harvesting CD34+ HSPCs from a subject in need of the treatment, editing the autologous CD34+ HSPCs ex vivo using the genome editing methods disclosed herein, and then reinfusing the edited autologous CD34+ HSPCs into the subject. In certain embodiments, treatment with the edited autologous CD34+ HSPCs can result in increased HbF induction.
[0107] In certain embodiments, prior to harvesting the CD34+ HSPCs, the subject may discontinue treatment with hydroxyurea, if applicable, and receive a blood transfusion to maintain sufficient hemoglobin (Hb) levels. In certain embodiments, the subject may be administered intravenous plerixafor (e.g., 0.24 mg / kg) to mobilize CD34+ HSPCs from the bone marrow to the peripheral blood. In certain embodiments, the subject may undergo one or more leukapheresis cycles (e.g., with about one month between cycles, one cycle defined as a collection of two plerixafor mobilized leukapheresis performed on consecutive days). In certain embodiments, the number of leukapheresis cycles performed on a subject is determined by the dose of edited autologous CD34+ HSPCs (e.g., ≧2×10) that will be reinfused into the subject. 6 cells / kg, ≧3×10 6 cells / kg, ≧4×10 6 cells / kg, ≧5×10 6 cells / kg, 2×10 6 cells / kg~3×10 6 cells / kg, 3×10 6 cells / kg~4×10 6 cells / kg, 4×10 6 cells / kg~5×10 6 cells / kg), as well as the dose of unedited autologous CD34+ HSPCs / kg for backup storage (e.g., ≥ 1.5 × 10 6 The number of CD34+ HSPCs may be as many as necessary to achieve a total cell / kg (cells / kg). In certain embodiments, CD34+ HSPCs harvested from a subject can be edited using any of the genome editing methods discussed herein. In certain embodiments, any one or more of the gRNAs and one or more of the RNA-guided nucleases disclosed herein can be used in the genome editing method.
[0108] In certain embodiments, the treatment may include autologous stem cell transplantation. In certain embodiments, the subject may undergo myeloablative conditioning with busulfan conditioning (e.g., at a test dose of 1 mg / kg, dose-adjusted based on pharmacokinetic analysis of the first dose). In certain embodiments, conditioning may be performed for 4 consecutive days. In certain embodiments, after a 3-day busulfan washout period, edited autologous CD34+ HSPCs (e.g., ≧2×10 6 cells / kg, ≧3×10 6 cells / kg, ≧4×10 6 cells / kg, ≧5×10 6 cells / kg, 2×10 6 cells / kg~3×10 6 cells / kg, 3×10 6 cells / kg~4×10 6 cells / kg, 4×10 6 cells / kg~5×10 6 The autologous CD34+ HSPCs may be produced for a particular subject and cryopreserved. In certain embodiments, the subject may achieve neutrophil engraftment following sequential myeloablative conditioning regimens and infusion of autologous edited CD34+ cells. Neutrophil engraftment is achieved when ANC > 0.5 x 10 9 / L on three consecutive measurements.
[0109] However, the genome editing system may include or be co-delivered with one or more factors that improve the viability of cells during and after editing, including, but not limited to, aryl hydrocarbon receptor antagonists such as StemRegenin-1 (SR1), UM171, LGC0006, alpha naphthoflavone, and CH-223191, and / or innate immune response antagonists such as cyclosporine A, dexamethasone, resveratrol, MyD88 inhibitor peptides, RNAi agents targeting Myd88, B18R recombinant proteins, glucocorticoids, OxPAPC, TLR antagonists, rapamycin, BX795, and RLR shRNA. These and other factors that improve the viability of cells during and after editing are described under the heading "I. Optimization of Stem Cells" in Gori, pages 36-61, which is incorporated herein by reference.
[0110] After delivery of the genome editing system, the cells are optionally expanded, frozen / thawed, or otherwise manipulated to prepare the cells for return to the subject (e.g., enriching for HSCs, and / or cells of the erythroid lineage, and / or the edited cells). The edited cells are then returned, for example, to the circulatory system, by intravenous delivery or delivery, or into solid tissue such as bone marrow.
[0111] Functionally, alteration of the CCAAT box target region, the 13 nt target region, and / or the proximal HBG1 / 2 promoter target sequence using the disclosed compositions, methods, and genome editing systems results in a significant induction of Aγ and / or Gγ subunits (interchangeably referred to as HbF expression) among hemoglobin-expressing cells, e.g., at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50% or more induction of expression of Aγ and / or Gγ subunits compared to unmodified controls. This induction of protein expression is generally the result of an alteration of the CCAAT box target region, the 13 nt target region, and / or the proximal HBG1 / 2 promoter target sequence in some or all of the treated cells (e.g., expressed in terms of the percentage of the total genome containing an indel mutation in the plurality of cells), e.g., at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50% of the plurality of cells contain at least one allele that contains a sequence alteration, including but not limited to an indel, insertion, or deletion, within or near the CCAAT box target region, the 13 nt target region, and / or the proximal HBG1 / 2 promoter target sequence.
[0112] The functional effect of the alterations caused or promoted by the genome editing systems and methods of the present disclosure can be evaluated in any number of suitable ways. For example, the effect of the alterations on the expression of fetal hemoglobin can be evaluated at the protein or mRNA level. The expression of HBG1 and HBG2 mRNA can be evaluated by digital droplet PCR (ddPCR), which is performed on cDNA samples obtained by reverse transcription of mRNA taken from treated or untreated samples. Primers for HBG1, HBG2, HBB, and / or HBA can be used individually or multiplexed using methods known in the art. For example, ddPCR analysis of samples can be performed using the QX200™ ddPCR system commercialized by Bio Rad (Hercules, CA) and associated protocols published by BioRad. Fetal hemoglobin proteins can be assessed by high pressure liquid chromatography (HPLC) (e.g., according to the methods discussed in Chang 2017, pages 143-44, incorporated herein by reference) or fast protein liquid chromatography (FPLC) using ion exchange and / or reverse phase columns to separate HbF, HbB and HbA, and / or Aγ and Gγ globin chains, as known in the art.
[0113] The embodiments described herein may be used in all classes of vertebrates, including, but not limited to, primates, mice, rats, rabbits, pigs, dogs, and cats.
[0114] This overview focuses on some exemplary embodiments that illustrate the principles of genome editing systems and CRISPR-mediated methods of cell modification. However, for clarity, this disclosure encompasses modifications and variations that are not explicitly mentioned above, but will be apparent to those skilled in the art. With that in mind, the following disclosure is intended to illustrate the principles of operation of genome editing systems more generally. The following should not be understood as limiting, but rather as illustrating the specific principles of genome editing systems and CRISPR-mediated methods that utilize these systems, which, in combination with this disclosure, will provide those skilled in the art with information on additional implementations and modifications that are within its scope.
[0115] This overview focuses on some exemplary embodiments that illustrate the principles of genome editing systems and CRISPR-mediated methods of cell modification. However, for clarity, this disclosure encompasses modifications and variations that are not explicitly mentioned above, but will be apparent to those skilled in the art. With that in mind, the following disclosure is intended to illustrate the principles of operation of genome editing systems more generally. The following should not be understood as limiting, but rather as illustrating the specific principles of genome editing systems and CRISPR-mediated methods that utilize these systems, which, in combination with this disclosure, will provide those skilled in the art with information on additional implementations and modifications that are within its scope.
[0116] Genome editing system The term "genome editing system" refers to any system with RNA-guided DNA editing activity. The genome editing system of the present disclosure includes at least two components adapted from naturally occurring CRISPR systems: guide RNA (gRNA) and RNA-guided nuclease. These two components form a complex that can associate with a specific nucleic acid sequence and edit the DNA within or around the nucleic acid sequence, for example, by creating one or more of single-strand breaks (SSBs or nicks), double-strand breaks (DSBs), and / or point mutations.
[0117] The genome editing system can be implemented (e.g., administered or delivered to a cell or subject) in a variety of ways, and different implementations may be suitable for different applications. For example, the genome editing system is implemented in certain embodiments as a protein / RNA complex (ribonucleoprotein, or RNP), which may be optionally included in a pharmaceutical composition including a pharma- ceutically acceptable carrier and / or encapsulant (e.g., but not limited to, lipid or polymer microparticles or nanoparticles, micelles, or liposomes). In certain embodiments, the genome editing system is implemented as one or more nucleic acids encoding the above-mentioned RNA-guided nuclease and guide RNA components (optionally with one or more additional components). In certain embodiments, the genome editing system is implemented as one or more vectors (e.g., viral vectors, such as adeno-associated viruses) that include such nucleic acids (see the section below under the heading "Implementation of the Genome Editing System: Delivery, Formulations, and Routes of Administration"). In certain embodiments, the genome editing system is implemented as a combination of any of the foregoing. Additional or modified implementations that operate according to the principles set forth herein will be apparent to one of skill in the art and are within the scope of the present disclosure. Exemplary RNPs are shown in Table 10. See WO 2021 / 119040 (see, e.g., Table 15).
[0118] It should be noted that the genome editing system of the present disclosure can target a single specific nucleotide sequence, or can target and edit two or more specific nucleotide sequences in parallel by using two or more guide RNAs. The use of multiple gRNAs is referred to throughout this disclosure as "multiplexing" and can be used to target multiple unrelated target sequences of interest, or to create multiple SSBs or DSBs within a single target domain, and in some cases, to generate specific edits within such target domains. For example, WO 2015 / 138510 by Maeder et al. ("Maeder", incorporated herein by reference) describes a genome editing system that corrects a point mutation (C.2991+1655 A to G) in the human CEP290 gene, resulting in a cryptic splice site that in turn reduces or eliminates the function of the gene. Maeder's genome editing system utilizes two guide RNAs that target sequences on either side (i.e., adjacent) of the point mutation, creating a DSB adjacent to the mutation. This in turn facilitates the deletion of the intervening sequences containing the mutation, thereby eliminating the cryptic splice site and restoring normal gene function.
[0119] As another example, WO2016 / 073990 by Cotta-Ramusino et al. ("Cotta-Ramusino", incorporated herein by reference) describes a genome editing system that utilizes two gRNAs in combination with a Cas9 nickase (a Cas9 that creates a single-stranded nick, such as S. pyogenes D10A), an arrangement referred to as the "dual nickase system." The Cotta-Ramusino dual nickase system is configured to create two nicks on opposite strands of the sequence of interest that are offset by one or more nucleotides, which the nicks join to create a double-stranded break with an overhang (5' in the case of Cotta-Ramusino, but 3' overhangs are also possible). In some circumstances, the overhangs can facilitate homologous recombination repair events. Also, as another example, WO2015 / 070083 (herein incorporated by reference) describes a gRNA that targets a nucleotide sequence encoding Cas9 (referred to as a "governing RNA"), which can be included in a genome editing system that includes one or more additional gRNAs, allowing for transient expression of Cas9 that may otherwise be constitutively expressed, for example, in some virally transduced cells. These multiplexing applications are intended to be illustrative rather than limiting, and one of skill in the art will understand that other applications of multiplexing are generally compatible with the genome editing systems described herein.
[0120] As disclosed herein, in certain embodiments, the genome editing system may include multiple gRNAs that can be used to introduce mutations into the 13 nt target region of HBG1 and / or HBG2. In certain embodiments, the genome editing system disclosed herein may include multiple gRNAs that can be used to introduce mutations into the 13 nt target region of HBG1 and / or HBG2.
[0121] Genome editing systems can, in some cases, create double-strand breaks that are repaired by cellular DNA double-strand break mechanisms, such as NHEJ or HDR. These mechanisms are described throughout the literature (see, e.g., Davis & Maizels 2014 (describing Alt-HDR), Frit 2014 (describing Alt-NHEJ), Iyama & Wilson 2013 (describing generally the canonical HDR and NHEJ pathways)).
[0122] Where a genome editing system operates by forming DSBs, such systems optionally include one or more components that promote or facilitate a particular mode of double-strand break repair or a particular repair outcome. For example, Cotta-Ramusino also describes a genome editing system in which a single-stranded oligonucleotide "donor template" is added, which can be incorporated into a target region of cellular DNA that is cut by the genome editing system, resulting in a change in the target sequence.
[0123] In certain embodiments, the genome editing system modifies the target sequence or modifies the expression of a gene within or near the target sequence without causing single- or double-strand breaks. For example, the genome editing system can include an RNA-guided nuclease fused to a functional domain that acts on DNA, thereby modifying the target sequence or its expression. As an example, the RNA-guided nuclease can be connected (e.g., fused) to a cytidine deaminase functional domain and can operate by generating a targeted C to A substitution. Exemplary nuclease / deaminase fusions are described in Komor 2016 (incorporated herein by reference). Alternatively, the genome editing system can utilize a cleavage-inactivated (i.e., "dead") nuclease, such as inactivated Cas9 (dCas9), that operates by forming a stable complex on one or more targeted regions of cellular DNA, thereby disrupting functions involved in the targeted regions, including, but not limited to, mRNA transcription, chromatin remodeling, and the like. In certain embodiments, the genome editing system may include an RNA-guided helicase that unwinds DNA within or near a target sequence without causing single-strand or double-strand breaks. For example, the genome editing system may include an RNA-guided helicase configured to associate within or near a target sequence to unwind DNA and induce accessibility to the target sequence. In certain embodiments, the RNA-guided helicase may be complexed to an inactive guide RNA that is configured to lack cleavage activity, allowing DNA to unwind without causing DNA cleavage.
[0124] Guide RNA (gRNA) molecules The terms "guide RNA" and "gRNA" refer to any nucleic acid that facilitates the specific association (or "targeting") of an RNA-guided nuclease, such as a Cpf1 molecule, to a target sequence, such as a genomic or episomal sequence in a cell. gRNAs can be unimolecular (comprising a single RNA molecule, alternatively referred to as chimeric) or modular (comprising multiple (typically two) separate RNA molecules, such as crRNA and tracrRNA, that typically associate with each other, e.g., by duplexing). gRNAs and their component parts are described throughout the literature (see, e.g., Briner 2014, Cotta-Ramusino, incorporated by reference). Examples of modular and unimolecular gRNAs that may be used in accordance with embodiments herein include, but are not limited to, the sequences set forth in SEQ ID NOs: 29-31 and 38-51. Examples of gRNA proximal and tail domains that may be used in accordance with embodiments herein include, but are not limited to, the sequences set forth in SEQ ID NOs: 32-37.
[0125] In bacteria and archaea, type II CRISPR systems generally include an RNA-guided nuclease protein such as Cas9, a CRISPR RNA (crRNA) that includes a 5' region that is complementary to a foreign sequence, and a transactivating crRNA (tracrRNA) that includes a 5' region that is complementary to the 3' region of the crRNA and duplexes with the 3' region of the crRNA. Without intending to be bound by any theory, it is believed that this duplex promotes the formation of the Cas9 / gRNA complex and is necessary for its activity. Adapting type II CRISPR systems for use in gene editing, in one non-limiting example, it has been found that the crRNA and tracrRNA can be linked into a single unimolecular or chimeric guide RNA by a four nucleotide (e.g., GAAA) "tetraloop" or "linker" sequence that bridges the complementary regions of the crRNA (at its 3' end) and the tracrRNA (at its 5' end). (Mali 2013, Jiang 2013, Jinek 2012; all incorporated herein by reference)
[0126] Guide RNAs include a "targeting domain," which, whether unimolecular or modular, is fully or partially complementary to a target domain within a target sequence (e.g., a DNA sequence in the genome of a cell in which editing is desired). Targeting domains have been referred to by various names in the literature, including, but not limited to, "guide sequence" (Hsu 2013, incorporated herein by reference), "complementarity region" (Cotta-Ramusino), "spacer" (Briner 2014), and generally "crRNA" (Jiang). Regardless of the name, targeting domains are typically 10-30 nucleotides in length, and in certain embodiments, 16-24 nucleotides in length (e.g., 16, 17, 18, 19, 20, 21, 22, 23, or 24 nucleotides in length), at or near the 5' end in the case of Cas9 gRNAs, and at or near the 3' end in the case of Cpf1 gRNAs.
[0127] In addition to the targeting domain, gRNAs typically (but not necessarily, as discussed below) contain multiple domains that can affect the formation or activity of the gRNA / Cas9 complex. For example, as described above, the duplex structure (also referred to as the repeat:anti-repeat duplex) formed by the first and second complementarity domains of the gRNA can interact with the recognition (REC) lobe of Cas9 and mediate the formation of the Cas9 / gRNA complex (Nishimasu 2014, Nishimasu 2015; both incorporated herein by reference). It should be noted that the first and / or second complementarity domains can contain one or more poly-A tracts that can be recognized by RNA polymerase as termination signals. Thus, the sequences of the first and second complementarity domains are optionally modified to eliminate these tracts and facilitate complete in vitro transcription of the gRNA, for example, by using AG or AU swaps as described in Briner 2014. These and other similar modifications to the first and second complementarity domains are within the scope of the present disclosure.
[0128] Cas9 gRNAs, along with the first and second complementarity domains, typically contain two or more additional duplex regions that are involved in nuclease activity in vivo, but not necessarily in vitro. (Nishimasu 2015). The first stem-loop near the 3' portion of the second complementarity domain is variously referred to as the "proximal domain" (Cotta-Ramusino), "stem-loop 1" (Nishimasu 2014 and 2015), and "nexus" (Briner 2014). One or more additional stem-loop structures are usually present near the 3' end of the gRNA, the number of which varies by species. S. pyogenes gRNAs typically contain two 3' stem-loops (a total of four stem-loop structures including the repeat:anti-repeat duplex), while S. aureus and other species have only one (a total of three stem-loop structures). A description of the conserved stem-loop structures (and gRNA structures more generally) organized by species is provided in Briner 2014.
[0129] While the above description focuses on gRNAs for use with Cas9, it should be understood that there are other RNA-guided nucleases that utilize gRNAs that differ in some respects from those described thus far. For example, Cpf1 ("CRISPR1 from Prevotella and Franciscella") is a recently discovered RNA-guided nuclease that does not require tracrRNA to function. (Zetsche 2015, incorporated herein by reference). gRNAs for use with the Cpf1 genome editing system generally include a targeting domain and a complementarity domain (alternatively referred to as a "handle"). Also, note that in gRNAs for use with Cpf1, the targeting domain is typically at or near the 3' end rather than the 5' end as described above in relation to Cas9 gRNAs (the handle is at or near the 5' end of Cpf1 gRNA). Exemplary targeting domains for Cpf1 gRNAs are shown in Tables 7, 8, 11, or 12. See WO 2021 / 119040 (see, e.g., Tables 12, 13, 16, 17). gRNA sequences targeting several domains of the HBG promoter (Table 6) are provided in Table 7. See WO 2021 / 119040 (see, e.g., Tables 11 and 12).
[0130] However, those skilled in the art will understand that while structural differences may exist between gRNAs derived from different prokaryotic species, or between Cpf1 gRNA and Cas9 gRNA, the principles by which gRNAs operate are generally consistent. Because of this consistency of operation, gRNAs can be broadly defined by their targeting domain sequences, and those skilled in the art will understand that a given targeting domain sequence can be incorporated into any suitable gRNA, including single molecule or chimeric gRNAs, or gRNAs that include one or more chemical and / or sequential modifications (substitutions, additional nucleotides, truncations, etc.). Thus, for economy of expression in this disclosure, gRNAs may be described in terms of their targeting domain sequences.
[0131] More generally, those skilled in the art will understand that some aspects of the present disclosure relate to systems, methods, and compositions that can be implemented using multiple RNA-guided nucleases. Thus, unless otherwise specified, the term gRNA should be understood to encompass any suitable gRNA that can be used with any RNA-guided nuclease, not just gRNAs that are compatible with a particular species of Cas9 or Cpf1. By way of example, in certain embodiments, the term gRNA can include gRNAs for use with any RNA-guided nuclease present in class 2 CRISPR systems, such as type II or type V or CRISPR systems, or RNA-guided nucleases derived or adapted therefrom.
[0132] gRNA design Methods for target sequence selection and validation, as well as off-target analysis, have been previously described (see, e.g., Mali 2013, Hsu 2013, Fu 2014, Heigwer 2014, Bae 2014, Xiao 2014). Each of these references is incorporated herein by reference. As a non-limiting example, gRNA design may involve the use of software tools to optimize the selection of potential target sequences corresponding to the user's target sequence, e.g., to minimize total off-target activity across the genome. Although off-target activity is not limited to cleavage, cleavage efficiency at each off-target sequence can be predicted, e.g., using an experimentally derived weighting scheme. These and other guide selection methods are described in detail in Maeder and Cotta-Ramusino.
[0133] Targeting domain sequences of gRNAs designed to target the disruption of a CCAAT box target region include, but are not limited to, SEQ ID NO: 1002. In certain embodiments, a gRNA comprising a sequence as set forth in SEQ ID NO: 1002 can be complexed with a Cpf1 protein or a modified Cpf1 protein to generate an alteration at a CCAAT box target region. In certain embodiments, a gRNA comprising any of the Cpf1 gRNAs set forth in Tables 7, 8, 11, and 12 can be complexed with a Cpf1 protein or a modified Cpf1 protein to form an RNP ("gRNA-Cpf1-RNP") to generate an alteration at a CCAAT box target region. In certain embodiments, the modified Cpf1 protein can be His-AsCpf1-nNLS (SEQ ID NO: 1000) or His-AsCpf1-sNLS-sNLS (SEQ ID NO: 1001). In certain embodiments, the Cpf1 molecule of the gRNA-Cpf1-RNP may be encoded by the sequence set forth in SEQ ID NOs: 1000, 1001, 1008-1018, 1032, 1035-39 (Cpf1 polypeptide sequences) or SEQ ID NOs: 1019-1021 (Cpf1 polynucleotide sequences).
[0134] gRNA Modification The activity, stability, or other properties of gRNAs may be altered by incorporating certain modifications. As an example, transiently expressed or delivered nucleic acids may be susceptible to degradation, for example, by cellular nucleases. Thus, the gRNAs described herein may include one or more modified nucleosides or nucleotides that introduce stability against nucleases. Without wishing to be bound by theory, it is believed that certain modified gRNAs described herein may exhibit reduced innate immune responses when introduced into cells. Those skilled in the art will recognize certain cellular responses that are commonly observed in cells (e.g., mammalian cells) in response to exogenous nucleic acids, particularly nucleic acids derived from viruses or bacteria. Such responses, which may include cytokine expression and release, as well as induction of cell death, may be reduced or completely eliminated by the modifications presented herein.
[0135] Certain exemplary modifications discussed in this section may be included anywhere within the gRNA sequence, including, but not limited to, at or near the 5' end (e.g., within 1-10, 1-5, or 1-2 nucleotides of the 5' end) and / or at or near the 3' end (e.g., within 1-10, 1-5, or 1-2 nucleotides of the 3' end). In some cases, modifications are located within functional motifs (e.g., the repeat-antirepeat duplex of a Cas9 gRNA, the stem-loop structure of a Cas9 or Cpf1 gRNA, and / or the targeting domain of a gRNA).
[0136] As an example, the 5' end of the gRNA can include a eukaryotic mRNA cap structure or cap analog (e.g., a G(5')ppp(5')G cap analog, an m7G(5')ppp(5')G cap analog, or a 3'-O-Me-m7G(5')ppp(5')G anti-reverse cap analog (ARCA)) as shown below.
[0137] [ka] The cap or cap analog can be included during the chemical synthesis or during in vitro transcription of the gRNA.
[0138] Similarly, the 5' end of the gRNA can lack a 5' triphosphate group. For example, in vitro transcribed gRNA can be phosphatase treated (e.g., using calf intestinal alkaline phosphatase) to remove the 5' triphosphate group.
[0139] Another common modification involves the addition of multiple adenine (A) residues (e.g., 1-10, 10-20, or 25-200) at the 3' end of the gRNA, referred to as a polyA tract. PolyA tracts can be added to gRNAs during chemical synthesis, following in vitro transcription using a polyadenosine polymerase (e.g., E. coli poly(A) polymerase), or in vivo transcription with a polyadenylation sequence, as described by Maeder.
[0140] It should be noted that the modifications described herein can be combined in any suitable manner. For example, a gRNA can include either or both a 5' cap structure or cap analog and a 3' polyA tract, whether in vivo transcribed from a DNA vector or an in vitro transcribed gRNA.
[0141] Guide RNAs can be modified with a 3'-terminal U-ribose. For example, the two terminal hydroxyl groups of U-ribose can be oxidized to aldehyde groups, simultaneously opening the ribose ring to give modified nucleosides as shown below:
[0142] [ka] In this formula, "U" can be an unmodified or modified uridine.
[0143] The 3' terminal U ribose can be modified with a 2'3' cyclic phosphate as shown below:
[0144] [ka] In this formula, "U" can be an unmodified or modified uridine.
[0145] The guide RNA may contain a 3' nucleotide, which may be stabilized against degradation, for example, by incorporating one or more of the modified nucleotides described herein. In certain embodiments, uridine may be replaced with modified uridine, such as 5-(2-amino)propyluridine, and 5-bromouridine, or any of the modified uridines described herein. Adenosine and guanosine may be replaced with modified adenosine and modified guanosine (e.g., a modification at position 8, such as 8-bromoguanosine), or any of the modified adenosines or modified guanosines described herein.
[0146] In certain embodiments, sugar-modified ribonucleotides may be incorporated into the gRNA. Sugar-modified ribonucleotides are those in which, for example, the 2'OH group is replaced by a group selected from H, -OR, -R (where R can be, for example, alkyl, cycloalkyl, aryl, aralkyl, heteroaryl, or sugar), halo, -SH, -SR (where R can be, for example, alkyl, cycloalkyl, aryl, aralkyl, heteroaryl, or sugar), amino (where amino can be, for example, NH2, alkylamino, dialkylamino, heterocyclyl, arylamino, diarylamino, heteroarylamino, diheteroarylamino, or amino acid), or cyano (-CN). In certain embodiments, the phosphate backbone may be modified, for example, with a phosphorothioate (PhTx) group, as described herein. In certain embodiments, one or more of the nucleotides of the gRNA may each independently be a modified or unmodified nucleotide. Modified nucleotides include, but are not limited to, 2'-sugar modifications (e.g., 2'-O-methyl, 2'-O-methoxyethyl, or 2'-fluoro modifications, such as 2'-F or 2'-O-methyl adenosine (A), 2'-F or 2'-O-methyl cytidine (C), 2'-F or 2'-O-methyl uridine (U), 2'-F or 2'-O-methyl thymidine (T), 2'-F or 2'-O-methyl guanosine (G), 2'-O-methoxyethyl-5-methyluridine (Teo), 2'-O-methoxyethyl adenosine (Aeo), 2'-O-methoxyethyl-5-methylcytidine (m5Ceo), and any combination thereof).
[0147] Guide RNAs may also include "locked" nucleic acids (LNAs), in which the 2'OH group may be connected to the 4' carbon of the same ribose sugar, for example, by a C1-6 alkylene or C1-6 heteroalkylene bridge. Any suitable moiety may be used to provide such a bridge, including methylene, propylene, ether, or amino bridges, O-amino (where amino may be, for example, NH2, alkylamino, dialkylamino, heterocyclyl, arylamino, diarylamino, heteroarylamino or diheteroarylamino, ethylenediamine, or polyamino), and aminoalkoxy or O(CH2). n -amino, where amino can be, for example, NH2, alkylamino, dialkylamino, heterocyclyl, arylamino, diarylamino, heteroarylamino or diheteroarylamino, ethylenediamine, or polyamino.
[0148] In certain embodiments, gRNAs can include modified nucleotides that are polycyclic (e.g., tricyclic and "unlocked" forms, such as glycol nucleic acids (GNAs) (e.g., R-GNAs or S-GNAs, where the ribose is replaced with a glycol unit linked to a phosphodiester bond) or threose nucleic acid (TNA, where the ribose is replaced with α-L-threofuranosyl-(3'→2')).
[0149] Generally, gRNAs contain the sugar group ribose, which is a five-membered ring with oxygen. Exemplary modified gRNAs may include, but are not limited to, substitution of the oxygen of ribose (e.g., sulfur (S), selenium (Se), or alkylene, e.g., methylene or ethylene), addition of double bonds (e.g., replacing ribose with cyclopentenyl or cyclohexenyl), ring contraction of ribose (e.g., forming a four-membered ring of cyclobutane or oxetane), ring expansion of ribose (e.g., forming a six- or seven-membered ring with additional carbon or heteroatoms, such as anhydrohexitol, altritol, mannitol, cyclohexanyl, cyclohexenyl, and morpholino, which also has a phosphoramidate backbone). Most of the changes in sugar analogs are localized at the 2' position, but other sites are amenable to modification, including the 4' position. In certain embodiments, the gRNA comprises a 4'-S, 4'-Se, or 4'-C-aminomethyl-2'-O-Me modification.
[0150] In certain embodiments, deazanucleotides, such as 7-deaza-adenosine, may be incorporated into the gRNA. In certain embodiments, O- and N-alkylated nucleotides, such as N6-methyladenosine, may be incorporated into the gRNA. In certain embodiments, one or more or all of the nucleotides in the gRNA are deoxynucleotides.
[0151] In certain embodiments, the gRNA used herein may be a modified gRNA or an unmodified gRNA. In certain embodiments, the gRNA may include one or more modifications. In certain embodiments, the one or more modifications may include a phosphorothioate bond modification, a phosphorodithioate (PS2) bond modification, a 2'-O-methyl modification, or a combination thereof. In certain embodiments, the one or more modifications may be at the 5' end of the gRNA, the 3' end of the gRNA, or a combination thereof.
[0152] In certain embodiments, gRNA modifications may include one or more phosphorodithioate (PS2) linkage modifications.
[0153] In some embodiments, gRNAs as used herein comprise one or more or a series of deoxyribonucleic acid (DNA) bases, also referred to herein as "DNA extensions." In some embodiments, gRNAs as used herein comprise a DNA extension at the 5' end of the gRNA, a DNA extension at the 3' end of the gRNA, or a combination thereof. In certain embodiments, the DNA extension is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, The DNA extension may be 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100 DNA bases long. For example, in certain embodiments, the DNA extension may be 1, 2, 3, 4, 5, 10, 15, 20, or 25 DNA bases long. In certain embodiments, the DNA extension may include one or more DNA bases selected from adenine (A), guanine (G), cytosine (C), or thymine (T). In certain embodiments, the DNA extension comprises the same DNA bases. For example, the DNA extension may comprise a series of adenine (A) bases. In certain embodiments, the DNA extension may comprise a series of thymine (T) bases. In certain embodiments, the DNA extension comprises a combination of different DNA bases. In certain embodiments, the DNA extension may comprise a sequence as set forth in Table 13. For example, the DNA extension may comprise a sequence as set forth in SEQ ID NOs: 1235-1250. In certain embodiments, the gRNA used herein comprises a DNA extension as well as one or more phosphorothioate linkage modifications, one or more phosphorodithioate (PS2) linkage modifications, one or more 2'-O-methyl modifications, or a combination thereof. In certain embodiments, the one or more modifications may be at the 5' end of the gRNA, the 3' end of the gRNA, or a combination thereof.In certain embodiments, the gRNA comprising the DNA extension may comprise a sequence as set forth in Table 13, which comprises the DNA extension. In certain embodiments, the gRNA comprising the DNA extension may comprise a sequence as set forth in SEQ ID NO: 1051. In certain embodiments, the gRNA comprising the DNA extension may comprise a sequence selected from the group consisting of SEQ ID NOs: 1046-1060, 1067, 1068, 1074, 1075, 1078, 1081-1084, 1086-1087, 1089-1090, 1092-1093, 1098-1102, and 1106. Without wishing to be bound by theory, it is contemplated that any DNA extension may be used herein as long as it does not hybridize to the target nucleic acid targeted by the gRNA and exhibits increased editing at the target nucleic acid site as compared to a gRNA not comprising such a DNA extension. Exemplary DNA and RNA extensions are provided in Table 13. See International Publication No. 2021 / 119040 (see, for example, Table 18).
[0154] In some embodiments, a gRNA as used herein comprises one or more or a series of ribonucleic acid (RNA) bases (also referred to herein as "RNA extension"). In some embodiments, a gRNA as used herein comprises an RNA extension at the 5' end of the gRNA, an RNA extension at the 3' end of the gRNA, or a combination thereof. In certain embodiments, the RNA extension is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, The RNA extension may be 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100 RNA bases long. For example, in certain embodiments, the RNA extension may be 1, 2, 3, 4, 5, 10, 15, 20, or 25 RNA bases long. In certain embodiments, the RNA extension may include one or more RNA bases selected from adenine (rA), guanine (rG), cytosine (rC), or uracil (rU), where "r" represents RNA, 2'-hydroxy. In certain embodiments, the RNA extension includes the same RNA bases. For example, the RNA extension may include a series of adenine (rA) bases. In certain embodiments, the RNA extension includes a combination of different RNA bases. In certain embodiments, the RNA extension may include a sequence shown in Table 13. For example, the RNA extension may include a sequence shown in 1231-1234, 1251-1253. In certain embodiments, the gRNA used herein includes an RNA extension and one or more phosphorothioate linkage modifications, one or more phosphorodithioate (PS2) linkage modifications, one or more 2'-O-methyl modifications, or a combination thereof. In certain embodiments, the one or more modifications may be at the 5' end of the gRNA, the 3' end of the gRNA, or a combination thereof.In certain embodiments, the gRNA comprising the RNA extension may comprise a sequence shown in Table 13, including an RNA extension at the 5' end of the gRNA may comprise a sequence selected from the group consisting of SEQ ID NOs: 1042-1045, 1103-1105. The gRNA comprising the RNA extension at the 3' end of the gRNA may comprise a sequence selected from the group consisting of SEQ ID NOs: 1070-1075, 1079, 1081, 1098-1100.
[0155] It is contemplated that gRNA as used herein may also include an RNA extension and a DNA extension. In certain embodiments, the RNA extension and the DNA extension may both be at the 5' end of the gRNA, the 3' end of the gRNA, or a combination thereof. In certain embodiments, the RNA extension is at the 5' end of the gRNA and the DNA extension is at the 3' end of the gRNA. In certain embodiments, the RNA extension is at the 3' end of the gRNA and the DNA extension is at the 5' end of the gRNA.
[0156] In some embodiments, a gRNA containing both a phosphorothioate modification at the 3' end as well as a DNA extension at the 5' end is complexed with an RNA-guided nuclease (e.g., Cpf1) to form an RNP, which is then used to edit hematopoietic stem cells (HSCs) or CD34+ cells ex vivo (i.e., outside the subject from whom such cells are derived) at the HBG locus.
[0157] An example of a gRNA as used herein includes the sequence shown in SEQ ID NO:1051.
[0158] RNA-guided nucleases RNA-guided nucleases according to the present disclosure include, but are not limited to, naturally occurring class 2 CRISPR nucleases, such as Cpf1 and Cas9, as well as other nucleases derived or obtained therefrom. Certain RNA-guided nucleases, such as Cas9, have also been shown to have helicase activity that allows them to unwind nucleic acids. In certain embodiments, the RNA-guided helicase according to the present disclosure can be any of the RNA nucleases described herein and above in the section entitled "RNA-guided nucleases." In certain embodiments, the RNA-guided nuclease is not configured to recruit exogenous trans-acting factors to the target region. In certain embodiments, the RNA-guided helicase can be an RNA-guided nuclease that is configured to lack nuclease activity. For example, in certain embodiments, the RNA-guided helicase can be a catalytically inactive RNA-guided nuclease that lacks nuclease activity but still retains its helicase activity. In certain embodiments, an RNA-guided nuclease can be mutated to destroy its nuclease activity (e.g., inactive Cas9), creating a catalytically inactive RNA-guided nuclease that cannot cleave nucleic acids but can still unwind DNA. In certain embodiments, an RNA-guided helicase can be complexed with any of the inactive guide RNAs described herein. For example, a catalytically active RNA-guided helicase (e.g., Cas9 or Cpf1) can form an RNP complex with an inactive guide RNA, resulting in a catalytically inactive inactive RNP (dRNP). In certain embodiments, a catalytically inactive RNA-guided helicase (e.g., inactive Cas9) and an inactive guide RNA can form a dRNP. These dRNPs cannot provide cleavage events but still retain their helicase activity that is important for unwinding nucleic acids.
[0159] Functionally, an RNA-guided nuclease is defined as a nuclease that (a) interacts with (e.g., complexes with) a gRNA, and (b) along with the gRNA, associates with and, optionally, cleaves or modifies a target region of DNA that includes (i) a sequence complementary to the targeting domain of the gRNA, and, optionally, (ii) an additional sequence referred to as a "protospacer adjacent motif" or "PAM," which is described in more detail below. Although variations may exist between individual RNA-guided nucleases that share the same PAM specificity or cleavage activity, as described in the Examples below, an RNA-guided nuclease can be broadly defined by its PAM specificity and cleavage activity. Those skilled in the art will appreciate that some aspects of the present disclosure relate to systems, methods, and compositions that can be implemented using any suitable RNA-guided nuclease with a particular PAM specificity and / or cleavage activity. Thus, unless otherwise specified, the term RNA-guided nuclease should be understood as a general term and is not limited to any particular type (e.g., Cas9 vs. Cpf1), species (e.g., S. pyogenes vs. S. aureus), or variation (e.g., full length vs. cleavage or split; naturally occurring PAM specificity vs. engineered PAM specificity, etc.) of RNA-guided nuclease. For example, in certain embodiments, the RNA-guided nuclease can be Cas-Φ (Pausch 2020).
[0160] Different RNA-guided nucleases may require different sequential relationships between the PAM and the protospacer. Generally, Cas9 recognizes the PAM sequence 3' to the protospacer, whereas Cpf1 generally recognizes the PAM sequence 5' to the protospacer.
[0161] In addition to recognizing a specific sequential orientation of the PAM and protospacer, RNA-guided nucleases can also recognize specific PAM sequences. S. aureus Cas9, for example, recognizes the NNGRRT or NNGRRV PAM sequence, with the N residue immediately 3' to the region recognized by the gRNA targeting domain. S. pyogenes Cas9 recognizes the NGG PAM sequence. F. novicida Cpf1 recognizes the TTN PAM sequence. PAM sequences have been identified for various RNA-guided nucleases, and a strategy for identifying novel PAM sequences is described by Shmakov 2015. It should also be noted that engineered RNA-guided nucleases may have PAM specificity that differs from that of the reference molecule (e.g., in the case of engineered RNA-guided nucleases, the reference molecule may be a naturally occurring variant from which the RNA-guided nuclease is derived, or may be a naturally occurring variant with the greatest amino acid sequence homology to the engineered RNA-guided nuclease). Examples of PAMs that may be used in accordance with embodiments of the present specification include, but are not limited to, the sequences set forth in SEQ ID NOs: 199-205.
[0162] In addition to their PAM specificity, RNA-guided nucleases can be characterized by their DNA cleavage activity: naturally occurring RNA-guided nucleases typically form DSBs in target nucleic acids, but engineered variants have been created that generate only SSBs (described above and in Ran & Hsu 2013, incorporated herein by reference) or do not cleave at all.
[0163] Cas9 Crystal structures have been determined for S. pyogenes Cas9 (Jinek 2014) and S. aureus Cas9 complexed with a single guide RNA and target DNA (Nishimasu 2014, Anders 2014, and Nishimasu 2015).
[0164] Naturally occurring Cas9 protein contains two lobes, i.e., a recognition (REC) lobe and a nuclease (NUC) lobe, each of which contains specific structural and / or functional domains. The REC lobe contains an arginine-rich bridge helix (BH) domain and at least one REC domain (e.g., a REC1 domain and optionally a REC2 domain). The REC lobe does not share structural similarity with other known proteins, indicating that it is a unique functional domain. Without wishing to be bound by any theory, mutational analysis suggests specific functional roles for the BH and REC domains. The BH domain appears to play a role in gRNA:DNA recognition, while the REC domain is believed to interact with the repeat:anti-repeat duplex of the gRNA and mediate the formation of the Cas9 / gRNA complex.
[0165] The NUC lobe comprises a RuvC domain, an HNH domain, and a PAM interaction (PI) domain. The RuvC domain shares structural similarity with retroviral integrase superfamily members and cleaves the non-complementary (i.e., bottom) strand of the target nucleic acid. It can be formed from two or more split RuvC motifs (such as RuvC I, RuvCII, and RuvCIII in S. pyogenes and S. aureus). On the other hand, the HNH domain is structurally similar to the HNN endonuclease motif and cleaves the complementary (i.e., top) strand of the target nucleic acid. The PI domain, as the name suggests, contributes to PAM specificity. Examples of polypeptide sequences encoding Cas9 RuvC-like domains and Cas9 HNH-like domains that can be used in accordance with embodiments herein are shown in SEQ ID NOs: 15-23, 52-123 (RuvC-like domains) and SEQ ID NOs: 24-28, 124-198 (HNH-like domains).
[0166] Although certain functions of Cas9 have been linked (not necessarily fully determined) to the specific domains described above, these and other functions may be mediated or influenced by other Cas9 domains or multiple domains in either lobe. For example, in S.pyogenes Cas9, as described in Nishimasu 2014, the repeat:anti-repeat duplex of the gRNA enters a groove between the REC and NUC lobes, and nucleotides of the duplex interact with amino acids in the BH, PI, and REC domains. Some nucleotides of the first stem-loop structure also interact with amino acids in multiple domains (PI, BH, and REC1), as do some nucleotides of the second and third stem-loops (RuvC and PI domains). Examples of polypeptide sequences encoding Cas9 molecules that may be used according to embodiments herein are shown in SEQ ID NOs: 1-2, 4-6, 12, and 14.
[0167] Cpf1 The crystal structure of Acidaminococcus sp. Cpf1 complexed with a double-stranded (ds) DNA target containing crRNA and a TTTN PAM sequence has been solved by Yamano 2016 (herein incorporated by reference). Cpf1 and Cas12a are synonymous and can be used interchangeably herein. Cpf1, like Cas9, has two lobes: the REC (recognition) lobe and the NUC (nuclease) lobe. The REC lobe contains the REC1 and REC2 domains, which lack similarity to any known protein structure. Meanwhile, the NUC lobe contains three RuvC domains (RuvC-I, -II, and -III) and a BH domain. However, in contrast to Cas9, the Cpf1 REC lobe lacks the HNH domain and contains other domains that lack similarity to known protein structures, including a structurally unique PI domain, three wedge (WED) domains (WED-I, -II, and -III), and a nuclease (Nuc) domain.
[0168] Although Cas9 and Cpf1 share similarities in structure and function, it should be understood that certain Cpf1 activities are mediated by structural domains that are not similar to any Cas9 domain. For example, cleavage of the complementary strand of the target DNA appears to be mediated by the Nuc domain, which is sequentially and spatially distinct from the HNH domain of Cas9. Furthermore, the non-targeting portion (handle) of the Cpf1 gRNA adopts a pseudoknot structure rather than the stem-loop structure formed by the repeat:anti-repeat duplex in the Cas9 gRNA.
[0169] In certain embodiments, the Cpf1 protein may be a modified Cpf1 protein. In certain embodiments, the modified Cpf1 protein may include one or more modifications. In certain embodiments, the modifications may be, but are not limited to, one or more mutations in the Cpf1 nucleotide sequence or Cpf1 amino acid sequence, one or more additional sequences, such as a His tag or a nuclear localization signal (NLS), or a combination thereof. In certain embodiments, the modified Cpf1 may also be referred to herein as a Cpf1 variant.
[0170] In certain embodiments, the Cpf1 protein may be derived from a Cpf1 protein selected from the group consisting of Acidaminococcus sp. strain BV3L6 Cpf1 protein (AsCpf1), Lachnospiraceae bacterium ND2006 Cpf1 protein (LbCpf1), and Lachnospiraceae bacterium MA2020 (Lb2Cpf1). In certain embodiments, the Cpf1 protein may comprise a sequence selected from the group consisting of SEQ ID NOs: 1016-1018, which have the codon-optimized nucleic acid sequences of SEQ ID NOs: 1019-1021, respectively.
[0171] In certain embodiments, the modified Cpf1 protein may include a nuclear localization signal (NLS). For example, but not limited to, NLS sequences useful in the context of the methods and compositions disclosed herein include amino acid sequences that can facilitate import of a protein into a cell nucleus. NLS sequences useful in the context of the methods and compositions disclosed herein are known in the art. Examples of such NLS sequences include the nucleoplasmin NLS having the following amino acid sequence: KRPAATKKAGQAKKKK (SEQ ID NO: 1006) and the simian virus 40 "SV40" NLS having the amino acid sequence PKKKRKV (SEQ ID NO: 1007).
[0172] In certain embodiments, the NLS sequence of the modified Cpf1 protein is located at or near the C-terminus of the Cpf1 protein sequence. For example, but not limited to, the modified Cpf1 protein may be selected from the following: His-AsCpf1-nNLS (SEQ ID NO: 1000), His-AsCpf1-sNLS (SEQ ID NO: 1008), and His-AsCpf1-sNLS-sNLS (SEQ ID NO: 1001), where "His" refers to the 6-histidine purification sequence, "AsCpf1" refers to the Acidaminococcus sp. Cpf1 protein sequence, "nNLS" refers to the nucleoplasmin NLS, and "sNLS" refers to the SV40 NLS. Further substitutions of the identity and C-terminal position of the NLS sequence, such as two or more nNLS sequences, or combinations of nNLS and sNLS sequences (or other NLS sequences), as well as adding sequences with or without purification sequences (e.g., a six-histidine sequence) are within the scope of the subject matter of this disclosure.
[0173] In certain embodiments, the NLS sequence of the modified Cpf1 protein may be located at or near the N-terminus of the Cpf1 protein sequence. For example, but not limited to, the modified Cpf1 protein may be selected from the following: His-sNLS-AsCpf1 (SEQ ID NO: 1009), His-sNLS-sNLS-AsCpf1 (SEQ ID NO: 1010), and sNLS-sNLS-AsCpf1 (SEQ ID NO: 1011). Further substitutions of the identity and N-terminus position of the NLS sequence, such as two or more nNLS sequences, or a combination of nNLS and sNLS sequences (or other NLS sequences), and adding sequences with or without purification sequences (e.g., six histidine sequences) are within the scope of the subject matter of the present disclosure.
[0174] In certain embodiments, the modified Cpf1 protein may include an NLS sequence located at or near both the N-terminus and C-terminus of the Cpf1 protein sequence. For example, but not limited to, the modified Cpf1 protein may be selected from the following: His-sNLS-AsCpf1-sNLS (SEQ ID NO: 1012) and His-sNLS-sNLS-AsCpf1-sNLS-sNLS (SEQ ID NO: 1013). Further substitutions of the identity of the NLS sequence and the N-terminus / C-terminus positions, such as two or more nNLS sequences, or a combination of nNLS and sNLS sequences (or other NLS sequences), as well as the addition of sequences with or without purification sequences (e.g., six histidine sequences) at either of the N-terminus / C-terminus positions, are within the scope of the subject matter of the present disclosure.
[0175] In certain embodiments, the modified Cpf1 protein may include an alteration (e.g., deletion or substitution) at one or more cysteine residues of the Cpf1 protein sequence. For example, but not limited to, the modified Cpf1 protein may include an alteration at a position selected from the group consisting of: C65, C205, C334, C379, C608, C674, C1025, and C1248. In certain embodiments, the modified Cpf1 protein may include a substitution of one or more cysteine residues with serine or alanine. In certain embodiments, the modified Cpf1 protein may include an alteration selected from the group consisting of: C65S, C205S, C334S, C379S, C608S, C674S, C1025S, and C1248S. In certain embodiments, the modified Cpf1 protein may include an alteration selected from the group consisting of: C65A, C205A, C334A, C379A, C608A, C674A, C1025A, and C1248A. In certain embodiments, the modified Cpf1 protein may include an alteration at positions C334 and C674, or C334, C379, and C674. In certain embodiments, the modified Cpf1 protein may include an alteration at positions C334S and C674S, or C334S, C379S, and C674S. In certain embodiments, the modified Cpf1 protein may include an alteration at positions C334A and C674A, or C334A, C379A, and C674A. In certain embodiments, modified Cpf1 proteins may include both alterations of one or more cysteine residues as well as the introduction of one or more NLS sequences, such as His-AsCpf1-nNLS Cys-less (SEQ ID NO: 1014) or His-AsCpf1-nNLS Cys-low (SEQ ID NO: 1015). In various embodiments, Cpf1 proteins that include deletions or substitutions at one or more cysteine residues exhibit reduced aggregation.
[0176] In certain embodiments, other modified Cpf1 proteins known in the art may be used with the methods and systems described herein. For example, in certain embodiments, the modified Cpf1 may be a Cpf1 containing the mutations S542R / K548V / N552R ("Cpf1 RVR"). Cpf1 RVR has been shown to cleave a target site with a TATV PAM. In certain embodiments, the modified Cpf1 may be a Cpf1 containing the mutations S542R / K607R ("Cpf1 RR"). Cpf1 RR has been shown to cleave a target site with a TYCV / CCCC PAM.
[0177] In some embodiments, Cpf1 variants are used herein, and the Cpf1 variants include 11, 12, 13, 14, 15, 16, 17, 34, 36, 39, 40, 43, 46, 47, 50, 54, 57, 58, 111, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 1 77, 178, 532, 533, 534, 535, 536, 537, 538, 539, 540, 541, 542, 543, 544, 545, 546, 547, 548, 549, 550, 551, 552, 553, 554, 555, 556, 565, 566, 567, 568, 569, 570, 571, 572, 573, 574, 575, 592, 593, 594, 595, 596, 597, 598, 599, 600, 601, 602, 603, 604, 605, 606, 607, 608, 609, 610 , 611, 612, 613, 614, 615, 616, 617, 618, 619, 620, 626, 627, 628, 629, 630, 631, 632, 633, 634, 635, 636, 637, 638, 642, 643, 644, 645, 646, 647, 648, 649, 651, 652, 653, 654, 655, 656, 676, 679, 680, 682, 683, 684, 685, 686, 687, 688, 689, 690, 691, 692, 693, 707, 711, 714, 715, 7 16, 717, 718, 719, 720, 721, 722, 739, 765, 768, 769, 773, 777, 778, 779, 780, 781, 782, 783, 784, 785, 786, 870, 871, 872, 873, 874, 875, 876, 877, 878, 879, 880, 881, 882, 883, 884, or 1048, or the corresponding positions of an ortholog, homolog, or variant of AsCpf1.
[0178] In certain embodiments, a Cpf1 variant as used herein may include any of the Cpf1 proteins described in WO 2017 / 184768A1 by Zhang et al. (the "768 publication"), which is incorporated herein by reference.
[0179] In certain embodiments, modified Cpf1 proteins (also referred to as Cpf1 variants) used herein may be encoded by any of the sequences set forth in SEQ ID NOs: 1000, 1001, 1008-1018, 1032, 1035-39, 1094-1097, 1107-09 (Cpf1 polypeptide sequences), or SEQ ID NOs: 1019-1021, 1110-17 (Cpf1 polynucleotide sequences). Table 9 shows the amino acid and nucleotide sequences of exemplary Cpf1 variants. See WO 2021 / 119040 (see, e.g., Table 14). Figure 6 shows these sequences. The locations of the 6-histidine sequence (underlined letters) and the NLS sequence (bold) are detailed. Further substitutions of the identity of the NLS sequence and the N-terminal / C-terminal positions, such as combinations of two or more nNLS sequences, or nNLS sequences and sNLS sequences (or other NLS sequences), as well as the addition of sequences with or without purification sequences (e.g., a six-histidine sequence) at either the N-terminal or C-terminal positions, are within the scope of the subject matter of this disclosure.
[0180] In certain embodiments, any of the Cpf1 proteins or modified Cpf1 proteins disclosed herein may be complexed with one or more gRNAs comprising a targeting domain as set forth in SEQ ID NO: 1002 and / or 1004 to alter the CCAAT box target region. In certain embodiments, any of the Cpf1 proteins or modified Cpf1 proteins disclosed herein may be complexed with one or more gRNAs comprising a sequence as set forth in Tables 7, 8, 11, or 12. In certain embodiments, the modified Cpf1 protein may be His-AsCpf1-nNLS (SEQ ID NO: 1000) or His-AsCpf1-sNLS-sNLS (SEQ ID NO: 1001). In certain embodiments, a modified Cpf1 protein as used herein may be encoded by any of the sequences set forth in SEQ ID NOs: 1000, 1001, 1008-1018, 1032, 1035-39, 1094-1097, 1107-09 (Cpf1 polypeptide sequences), or SEQ ID NOs: 1019-1021, 1110-17 (Cpf1 polynucleotide sequences). In certain embodiments, a modified Cpf1 protein may comprise the sequence set forth in SEQ ID NO: 1097.
[0181] In certain embodiments, the modified Cpf1 protein may include a Cpf1 variant described in Kleinstiver 2019. For example, but not limited to, in certain embodiments, the modified Cpf1 protein may be enAsCas12a described in Kleinstiver 2019. In certain embodiments, the modified Cpf1 protein may cleave a target site with a TTTV PAM. In certain embodiments, the modified Cpf1 protein may cleave a target site with a NWYN PAM.
[0182] Modification of RNA-guided nucleases While the RNA-guided nucleases described above have activities and properties that may be useful for a variety of applications, one of skill in the art will understand that RNA-guided nucleases may also be modified in certain cases to alter cleavage activity, PAM specificity, or other structural or functional characteristics.
[0183] Turning first to modifications that change cleavage activity, mutations that reduce or eliminate the activity of domains in the NUC lobe are described above. Exemplary mutations that can be made in the RuvC domain, Cas9 HNH domain, or Cpf1 Nuc domain are described in Ran & Hsu 2013 and Yamano 2016, as well as Cotta-Ramusino. In general, mutations that reduce or eliminate activity in one of the two nuclease domains result in an RNA-guided nuclease with nickase activity, but it should be noted that the type of nickase activity varies depending on which domain is inactivated. As an example, inactivation of the RuvC domain of Cas9 will result in a nickase that cuts the complementary strand or the top strand, as shown below (where C indicates the cleavage site).
[0184] On the other hand, inactivation of the Cas9 HNH domain results in a nickase that cleaves the bottom or non-complementary strand.
[0185] Modifications of PAM specificity to the naturally occurring Cas9 reference molecule have been described by Kleinstiver et al. for both S. pyogenes (Kleinstiver 2015a) and S. aureus (Kleinstiver 2015b). Kleinstiver et al. also describe modifications that improve the targeting fidelity of Cas9 (Kleinstiver 2016). Kleinstiver et al. also describe modifications of Cpf1 that result in increased activity and improved targeting range (Kleinstiver 2019). Each of these references is incorporated herein by reference.
[0186] RNA-guided nucleases have been split into two or more parts, as described by Zetsche 2015 and Fine 2015, both of which are incorporated herein by reference.
[0187] In certain embodiments, the RNA-guided nuclease can be size-optimized or truncated, for example, by one or more deletions that reduce the size of the nuclease while still retaining gRNA association, target and PAM recognition, and cleavage activity. In certain embodiments, the RNA-guided nuclease is covalently or non-covalently linked to another polypeptide, nucleotide, or other structure, optionally by a linker. Exemplary linked nucleases and linkers are described by Guilinger 2014 (herein incorporated by reference for all purposes).
[0188] The RNA-guided nuclease also optionally comprises a tag, such as, but not limited to, a nuclear localization signal to facilitate the translocation of the RNA-guided nuclease protein to the nucleus. In certain embodiments, the RNA-guided nuclease can incorporate a C-terminal and / or N-terminal nuclear localization signal. Nuclear localization sequences are known in the art and described in Maeder et al.
[0189] The foregoing list of modifications is intended to be exemplary in nature, and those skilled in the art will understand in light of the present disclosure that other modifications may be possible or desirable in certain applications. Thus, briefly, the exemplary systems, methods, and compositions of the present disclosure are presented with reference to specific RNA-guided nucleases, but it should be understood that the RNA-guided nucleases used may be modified in ways that do not change their principle of operation. Such modifications are within the scope of the present disclosure.
[0190] Nucleic acid encoding an RNA-guided nuclease Nucleic acids encoding RNA-guided nucleases (e.g., Cas9, Cpf1, or functional fragments thereof) are provided herein. Exemplary nucleic acids encoding RNA-guided nucleases have been previously described (see, e.g., Cong 2013, Wang 2013, Mali 2013, Jinek 2012).
[0191] In some cases, the nucleic acid encoding the RNA-guided nuclease can be a synthetic nucleic acid sequence. For example, the synthetic nucleic acid molecule can be chemically modified. In certain embodiments, the mRNA encoding the RNA-guided nuclease can have one or more (e.g., all) of the following characteristics: capping, polyadenylation, and substitution with 5-methylcytidine and / or pseudouridine.
[0192] The synthetic nucleic acid sequence can also be codon-optimized, e.g., at least one uncommon or less common codon is replaced by a common codon. For example, the synthetic nucleic acid can direct the synthesis of an optimized messenger mRNA, e.g., optimized for expression in a mammalian expression system as described herein. An example of a codon-optimized Cas9 coding sequence is provided in Cotta-Ramusino.
[0193] Additionally or alternatively, the nucleic acid encoding the RNA-guided nuclease may contain a nuclear localization sequence (NLS). Nuclear localization sequences are known in the art.
[0194] Functional analysis of candidate molecules Candidate RNA-guided nucleases, gRNAs, and their complexes can be evaluated by standard methods known in the art. For example, see Cotta-Ramusino. The stability of RNP complexes can be evaluated by differential scanning fluorimetry, as described below.
[0195] Differential Scanning Fluorescence (DSF) The thermal stability of a ribonucleoprotein (RNP) complex containing a gRNA and an RNA-guided nuclease can be measured by DSF. The DSF technique measures the thermal stability of a protein, which can be increased under favorable conditions, such as the addition of a binding RNA molecule (e.g., gRNA).
[0196] The DSF assay can be performed according to any suitable protocol and can be used in any suitable setting, including but not limited to (a) testing different conditions (e.g., different stoichiometric ratios of gRNA:RNA-guided nuclease protein, different buffer solutions, etc.) to identify optimal conditions for RNP formation, and (b) testing modifications (e.g., chemical modifications, sequence changes, etc.) of the RNA-guided nuclease and / or gRNA to identify those modifications that improve RNP formation or stability. One readout of the DSF assay is the shift in melting temperature of the RNP complex, with a relatively high shift suggesting that the RNP complex is more stable (and thus may have higher activity or more favorable formation kinetics, disassembly kinetics, or another functional characteristic) compared to a reference RNP complex characterized by a lower shift. When the DSF assay is deployed as a screening tool, a threshold melting temperature shift can be specified such that the output is one or more RNPs with a melting temperature shift equal to or greater than the threshold. For example, the threshold can be 5-10° C. (e.g., 5°, 6°, 7°, 8°, 9°, 10°) or greater, and the output can be one or more RNPs characterized by a melting temperature shift above the threshold.
[0197] Two non-limiting examples of DSF assay conditions are provided below.
[0198] To determine the best solution for forming RNP complexes, a fixed concentration (e.g., 2 μM) of Cas9+10×SYPRO Orange® (Life Technologies, Catalog No. S-6650) in water is dispensed into a 384-well plate. Equimolar amounts of gRNA diluted in solutions with various pH and salt are then added. After 10' incubation at room temperature and brief centrifugation to remove air bubbles, a gradient from 20°C to 90°C with 1°C increase every 10 seconds is run using a Bio-Rad CFX384™ Real-Time System C1000 Touch™ Thermal Cycler with Bio-Rad CFX Manager software.
[0199] The second assay consists of mixing various concentrations of gRNA in optimal buffer from assay 1 above with a fixed concentration (e.g., 2 μM) of Cas9 and incubating in a 384-well plate (e.g., for 10' at room temperature). An equal volume of optimal buffer + 10x SYPRO Orange® (Life Technologies, Cat. No. S-6650) is added and the plate is sealed with Microseal® B adhesive (MSB-1001). After a brief centrifugation to remove air bubbles, a gradient from 20°C to 90°C with 1°C increase every 10 seconds is run using a Bio-Rad CFX384™ Real-Time System C1000 Touch™ thermal cycler with Bio-Rad CFX Manager software.
[0200] Genome editing strategies The above genome editing system is used in various embodiments of the present disclosure to generate edits (i.e., changes) in targeted regions of DNA in a cell or DNA obtained from a cell. Various strategies for generating specific edits are described herein, and these strategies are generally described in terms of the desired repair outcome, the number and positioning of individual edits (e.g., SSBs or DSBs), and the target sites of such edits.
[0201] Genome editing strategies involving the formation of SSBs or DSBs are characterized by repair outcomes including: (a) deletion of all or part of the targeted region, (b) insertion or substitution into all or part of the targeted region, or (c) interruption of all or part of the targeted region. This grouping is not intended to be limited or bound to a particular theory or model, and is provided solely for economy of presentation. Those skilled in the art will understand that the listed outcomes are not mutually exclusive, and some modifications may result in other outcomes. Unless otherwise specified, description of a particular editing strategy or method should not be understood to require a particular repair outcome.
[0202] Replacement of a targeted region generally involves replacing all or part of an existing sequence in the targeted region with a homologous sequence, for example, via gene correction or gene conversion, two repair outcomes mediated by the HDR pathway. HDR is facilitated by the use of a donor template, which can be single-stranded or double-stranded, as described in more detail below. Single-stranded or double-stranded templates can be exogenous, in which case they facilitate gene correction to facilitate gene conversion, or they can be endogenous (e.g., a homologous sequence in the cellular genome). Exogenous templates can have asymmetric overhangs (i.e., the portion of the template that is complementary to the site of the DSB can be offset in the 3' or 5' direction rather than centered in the donor template), as described, for example, by Richardson 2016 (incorporated herein by reference). If the template is single-stranded, it can correspond to either the complementary strand (top strand) or non-complementary strand (bottom strand) of the targeted region.
[0203] In some cases, gene conversion and correction are facilitated by the formation of one or more nicks at or around the targeted region, as described in Ran & Hsu 2013 and Cotta-Ramusino. In some cases, a dual nickase strategy is used to create two offset SSBs that form a single DSB with an overhang (e.g., a 5' overhang).
[0204] The interruption and / or deletion of all or part of the target sequence can be achieved by various repair outcomes.As an example, the sequence can be deleted by simultaneously generating two or more DSBs adjacent to the target region, and then excising them when the DSBs are repaired, as described by Maeder for LCA10 mutation.As another example, the sequence can be interrupted by the formation of a double-stranded break with a single-stranded overhang, followed by deletion generated by exonuclease degradation treatment of the overhang before repair.
[0205] One particular subset of target sequence interruptions is mediated by the formation of indels within the targeted sequence, with the repair outcome typically being mediated by the NHEJ pathway (including Alt-NHEJ). NHEJ is referred to as an "error-prone" repair pathway due to its association with indel mutations. However, in some cases, DSBs are repaired by NHEJ without altering the sequence surrounding them (so-called "complete" or "scarless" repair). This generally requires that both ends of the DSB are completely ligated. On the other hand, indels are believed to result from enzymatic processing of free DNA ends prior to ligation, which adds and / or removes nucleotides from either or both strands of either or both free ends.
[0206] Because enzymatic processing of free DSB ends can be stochastic in nature, indel mutations tend to be variable and occur along a distribution, which can be influenced by a variety of factors including the specific target site, the cell type used, the genome editing strategy used, etc. Still, it is possible to draw limited generalizations about indel formation, and deletions formed by repair of a single DSB are most commonly in the range of 1-50 bp, but can exceed 100-200 bp. Insertions formed by repair of a single DSB tend to be shorter, often containing a short duplication of the sequence immediately surrounding the break site. However, it is possible to obtain large insertions, and in these cases the inserted sequence has often been traced to other regions of the genome or to plasmid DNA present in the cell.
[0207] Indel mutations (and genome editing systems configured to generate indels) are useful for interrupting target sequences, for example, when the generation of a specific final sequence is not required and / or frameshift mutations are tolerated. They may also be useful in settings where a specific sequence is preferred, so long as the desired specific sequence tends to preferentially result from the repair of SSBs or DSBs at a given site. Indel mutations are also useful tools for evaluating or screening the activity of specific genome editing systems and their components. In these and other settings, indels can be characterized by (a) their relative and absolute frequency in the genome of cells contacted with the genome editing system, and (b) the distribution of numerical differences (e.g., ±1, ±2, ±3, etc.) relative to the unedited sequence. As an example, in a lead discovery setting, multiple gRNAs can be screened to identify those gRNAs that most efficiently drive cleavage at the target site based on indel readouts under controlled conditions. Guides that generate indels at or above a threshold frequency or that generate a specific distribution of indels can be selected for further research and development. The frequency and distribution of indels can also be useful as a readout to evaluate different genome editing system implementations or formulations and delivery methods, for example, by keeping the gRNA constant and varying certain other reaction conditions or delivery methods.
[0208] Multiplexing Strategy The genome editing system according to the present disclosure can also be used for multiplex gene editing to generate two or more DSBs either at the same locus or at different loci. Any of the RNA-guided nucleases and gRNAs disclosed herein can be used in the genome editing system for multiplex gene editing. Strategies for editing involving the formation of multiple DSBs or SSBs are described, for example, in Cotta-Ramusino. In certain embodiments, multiple gRNAs and RNA-guided nucleases can be used in the genome editing system to introduce changes (e.g., deletions, insertions) in the CCAAT box target regions of HBG1 and / or HBG2. In certain embodiments, the RNA-guided nuclease can be Cpf1 or modified Cpf1.
[0209] Donor mold design Donor template design is described in detail in, for example, Cotta-Ramusino.DNA oligomer donor template (oligodeoxynucleotide or ODN) can be single-stranded (ssODN) or double-stranded (dsODN) and can be used to promote HDR-based repair of DSB or increase the overall editing rate, and is particularly useful for introducing changes into target DNA sequence, inserting new sequences into target sequence, or completely replacing target sequence.
[0210] Whether single-stranded or double-stranded, the donor template generally contains regions that are homologous to regions of DNA within or near (e.g., flanking or adjoining) the target sequence to be cleaved. These regions of homology are referred to herein as "homology arms" and are shown diagrammatically below: [5' homology arm]-[substitution sequence]--[3' homology arm].
[0211] The homology arms can have any suitable length (including 0 nucleotides if only one homology arm is used), and the 3' and 5' homology arms can have the same length or can be different lengths. The selection of the appropriate homology arm length can be influenced by various factors, such as the desire to avoid homology or microhomology with certain sequences, such as Alu repeats or other very common elements. For example, the 5' homology arm can be shortened to avoid sequence repeat elements. In other embodiments, the 3' homology arm can be shortened to avoid sequence repeat elements. In some embodiments, both the 5' and 3' homology arms can be shortened to avoid including certain sequence repeat elements. In addition, some homology arm designs can improve the efficiency of editing or increase the frequency of the desired repair outcome. For example, Richardson 2016, incorporated herein by reference, found that the relative asymmetry of the 3' and 5' homology arms of a single-stranded donor template affects the repair rate and / or repair outcome.
[0212] The replacement sequence in the donor template is described in other publications, including Cotta-Ramusino. The replacement sequence can be of any suitable length (including zero nucleotides where the desired repair result is a deletion), and typically includes one, two, three, or more sequence modifications relative to the naturally occurring sequence in the cell where editing is desired. One common sequence modification involves a change in the naturally occurring sequence to repair a mutation associated with a disease or condition that is desired to be treated. Another common sequence modification involves a change in one or more sequences that are complementary to, or are then complementary to, the PAM sequence of the RNA-guided nuclease or the targeting domain of the gRNA used to generate the SSB or DSB, in order to reduce or eliminate repeated cuts at the target site after the replacement sequence is incorporated into the target site.
[0213] When a linear ssODN is used, it can be configured to (i) anneal to a nicked strand of a target nucleic acid, (ii) anneal to an intact strand of a target nucleic acid, (iii) anneal to a plus strand of a target nucleic acid, and / or (iv) anneal to a minus strand of a target nucleic acid. The ssODN can have any suitable length, for example, about 80-200 nucleotides, at least 80-200 nucleotides, or 80-200 nucleotides or less (e.g., 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, or 200 nucleotides).
[0214] It should be noted that the template nucleic acid can also be a nucleic acid vector, such as a viral genome or a circular double-stranded DNA, e.g., a plasmid. The nucleic acid vector containing the donor template can include other coding or non-coding elements. For example, the template nucleic acid can be delivered as part of a viral genome (e.g., an AAV or lentiviral genome) that includes certain genomic backbone elements (e.g., in the case of the AAV genome, inverted terminal repeats), and optionally includes additional sequences encoding gRNAs and / or RNA-guided nucleases. In certain embodiments, the donor template can be adjacent to or flanked by target sites recognized by one or more gRNAs, facilitating the formation of free DSBs at one or both ends of the donor template that can participate in the repair of corresponding SSBs or DSBs formed in cellular DNA using the same gRNA. Exemplary nucleic acid vectors suitable for use as donor templates are described in Cotta-Ramusino (incorporated by reference).
[0215] Whatever the format used, the template nucleic acid can be designed to avoid undesired sequences, in certain embodiments, one or both homology arms can be shortened to avoid overlap with certain sequence repeat elements, e.g., Alu repeats, LINE elements, etc.
[0216] In certain embodiments, silent, non-pathogenic SNPs may be included in the ssODN donor template to allow identification of gene editing events.
[0217] Donor template or template nucleic acid as the term is used herein refers to a nucleic acid sequence that can be used in conjunction with an RNA nuclease molecule and one or more gRNA molecules to alter (e.g., delete, disrupt, or modify) a target DNA sequence. In certain embodiments, the template nucleic acid results in an alteration (e.g., deletion) of the CCAAT box target region of HBG1 and / or HBG2. In certain embodiments, the alteration is a non-naturally occurring alteration.
[0218] In certain embodiments, the ssODN comprises, consists essentially of, or consists of one or more sequences selected from the group consisting of SEQ ID NOs: 974-995, 1040. See WO 2021 / 119040 (see, e.g., Examples 2, 9, 10, 11, 12).
[0219] In certain embodiments, the 5' homology arm comprises a 5' phosphorothioate (PhTx) modification. In certain embodiments, the 3' homology arm comprises a 3' PhTx modification. In certain embodiments, the template nucleic acid comprises 5' and 3' PhTx modifications.
[0220] target cell The genome editing system according to the present disclosure can be used to manipulate or alter a cell, for example, to edit or alter a target nucleic acid. Manipulation can, in various embodiments, be performed in vivo or ex vivo.
[0221] Various cell types can be engineered or altered according to embodiments of the present disclosure. In some cases, such as in vivo applications, multiple cell types are altered or engineered, for example, by delivering a genome editing system according to the present disclosure to the multiple cell types. However, in other cases, it may be desirable to limit the engineering or alteration to a particular cell type(s). For example, in some instances, it may be desirable to edit cells with limited differentiation potential or terminally differentiated cells (e.g., photoreceptor cells in the case of Maeder, where modification of the genotype is expected to result in a change in the cell phenotype). However, in other cases, it may be desirable to edit less differentiated, multipotent, or pluripotent stem or progenitor cells. By way of example, the cells may be embryonic stem cells, induced pluripotent stem cells (iPSCs), hematopoietic stem / progenitor cells (HSPCs), or other stem or progenitor cell types that differentiate into cell types relevant to a given application or indication.
[0222] Naturally, the cells that are changed or engineered are dividing or non-dividing cells, depending on the cell type targeted and / or the desired editing outcome.
[0223] If cells are manipulated or altered ex vivo, the cells can be used immediately (e.g., administered to a subject) or they can be maintained or stored for later use. One of skill in the art will understand that cells can be maintained or stored in culture (e.g., frozen in liquid nitrogen) using any suitable method known in the art.
[0224] Implementation of genome editing systems: delivery, formulation, and routes of administration As discussed above, the genome editing system of the present disclosure can be implemented in any suitable manner. This means that the components of such a system, including but not limited to RNA-guided nuclease, gRNA, and optional donor template nucleic acid, can be delivered, formulated, or administered in any suitable form or combination of forms that results in transduction, expression, or introduction of the genome editing system and / or results in a desired repair outcome in a cell, tissue, or subject. Tables 2 and 3 show some non-limiting examples of implementations of the genome editing system. However, those skilled in the art will understand that these lists are not comprehensive and other implementations are possible. With particular reference to Table 2, the table lists some exemplary embodiments of a genome editing system that includes a single gRNA and an optional donor template. However, the genome editing system according to the present disclosure can incorporate multiple gRNAs, multiple RNA-guided nucleases, and other components (e.g., proteins), and various implementations will be apparent to those skilled in the art based on the principles shown in the table. In the table, [N / A] indicates that the genome editing system does not include the indicated component.
[0225] [Table 2-1]
[0226] [Table 2-2]
[0227] Table 3 summarizes various delivery methods for the components of the genome editing system described herein. Again, this list is intended to be illustrative, not limiting.
[0228] [Table 3]
[0229] Nucleic acid-based delivery of genome editing systems Nucleic acids encoding various elements of the genome editing system according to the present disclosure can be administered to a subject or delivered to a cell by methods known in the art or as described herein. For example, DNA encoding an RNA-guided nuclease and / or DNA encoding a gRNA, and a donor template nucleic acid can be delivered, for example, by a vector (e.g., a viral or non-viral vector), a non-vector-based method (e.g., using naked DNA or DNA complexes), or a combination thereof.
[0230] Nucleic acids encoding the genome editing system or its components can be delivered directly to cells as naked DNA or RNA, for example by transfection or electroporation, or can be conjugated to molecules (e.g., N-acetylgalactosamine) that facilitate uptake by target cells (e.g., red blood cells, HSCs). Nucleic acid vectors, such as those summarized in Table 3, can also be used.
[0231] The nucleic acid vector may include one or more sequences encoding genome editing system components, such as an RNA-guided nuclease, a gRNA, and / or a donor template. The vector may also include a sequence encoding a signal peptide (e.g., for nuclear localization, nucleolar localization, or mitochondrial localization) associated with (e.g., inserted or fused to) the protein-encoding sequence. As an example, the nucleic acid vector may include a sequence encoding Cas9 that includes one or more nuclear localization sequences (e.g., a nuclear localization sequence from SV40).
[0232] The nucleic acid vector can also include any suitable number of regulatory / control elements, such as promoters, enhancers, introns, polyadenylation signals, Kozak consensus sequences, or internal ribosome entry sites (IRES). These elements are well known in the art and described in Cotta-Ramusino.
[0233] Nucleic acid vectors according to the present disclosure include recombinant viral vectors. Table 3 shows exemplary viral vectors. Additional suitable viral vectors and their use and production are described in Cotta-Ramusino. Other viral vectors known in the art can also be used. In addition, viral particles can be used to deliver genome editing system components in nucleic acid and / or peptide form. For example, "empty" viral particles can be assembled to include any suitable cargo. Viral vectors and viral particles can also be engineered to incorporate targeting ligands to change target tissue specificity.
[0234] In addition to viral vectors, non-viral vectors can be used to deliver the nucleic acid encoding the genome editing system according to the present disclosure. One important category of non-viral nucleic acid vectors is nanoparticles, which can be organic or inorganic. Nanoparticles are well known in the art and are summarized in Cotta-Ramusino. Any suitable nanoparticle design can be used to deliver the genome editing system components or the nucleic acid encoding such components. For example, organic (e.g., lipid and / or polymer) nonparticles can be suitable for use as delivery vehicles in certain embodiments of the present disclosure. Table 4 shows exemplary lipids for use in nanoparticle formulations and / or gene transfer. Also, Table 5 lists exemplary polymers for use in gene transfer and / or nanoparticle formulations.
[0235] [Table 4-1]
[0236] [Table 4-2]
[0237] [Table 4-3]
[0238] [Table 5]
[0239] Non-viral vectors optionally include targeting modifications to improve uptake and / or selectively target specific cell types. These targeting modifications may include, for example, cell-specific antigens, monoclonal antibodies, single-chain antibodies, aptamers, polymers, sugars (e.g., N-acetylgalactosamine (GalNAc)), and cell-penetrating peptides. Such vectors also optionally use fusogenic and endosome-destabilizing peptides / polymers, undergo acid-induced conformational changes (e.g., to accelerate endosomal escape of the cargo), and / or incorporate stimuli-cleavable polymers, for example, for release within a cellular compartment. For example, disulfide-based cationic polymers can be used that are cleaved in a reducing cellular environment.
[0240] In certain embodiments, one or more nucleic acid molecules (e.g., DNA molecules) other than the components of the genome editing system (e.g., RNA-guided nuclease components and / or gRNA components described herein) are delivered. In certain embodiments, the nucleic acid molecules are delivered simultaneously with one or more of the components of the genome editing system. In certain embodiments, the nucleic acid molecules are delivered before or after (e.g., less than about 30 minutes, 1 hour, 2 hours, 3 hours, 6 hours, 9 hours, 12 hours, 1 day, 2 days, 3 days, 1 week, 2 weeks, or 4 weeks) one or more of the components of the genome editing system are delivered. In certain embodiments, the nucleic acid molecules are delivered by a different means than one or more of the components of the genome editing system (e.g., RNA-guided nuclease components and / or gRNA components) are delivered. The nucleic acid molecules may be delivered by any of the delivery methods described herein. For example, the nucleic acid molecule may be delivered by a viral vector (e.g., an integration-deficient lentivirus), and the RNA-guided nuclease molecule component and / or the gRNA component may be delivered by electroporation, for example, to reduce toxicity caused by the nucleic acid (e.g., DNA). In certain embodiments, the nucleic acid molecule encodes a therapeutic protein, such as a protein described herein. In certain embodiments, the nucleic acid molecule encodes an RNA molecule, such as an RNA molecule described herein.
[0241] Delivery of RNP and / or RNA-encoding genome editing system components RNP (complex of gRNA and RNA-guided nuclease) and / or RNA encoding RNA-guided nuclease and / or gRNA can be delivered to cells or administered to a subject by methods known in the art, some of which are described in Cotta-Ramusino. In vitro, RNA encoding RNA-guided nuclease and / or RNA encoding gRNA can be delivered by, for example, microinjection, electroporation, transient cell compaction or squeezing (see, for example, Lee 2012). Lipid-mediated transfection, peptide-mediated delivery, GalNAc or other conjugate-mediated delivery, and combinations thereof can also be used for delivery in vitro and in vivo. The Protective Interactive Non-Condensing (PINC) system can be used for delivery.
[0242] In vitro delivery by electroporation involves mixing cells with RNA encoding the RNA-guided nuclease and / or gRNA, with or without a donor template nucleic acid molecule, in a cartridge, chamber, or cuvette, and applying one or more electrical impulses of predetermined duration and magnitude. Systems and protocols for electroporation are known in the art, and any suitable electroporation tools and / or protocols can be used in connection with various embodiments of the present disclosure.
[0243] Route of administration The genome editing system, or cells altered or engineered using such a system, can be administered to a subject by any suitable manner or route, whether local or systemic. Systemic administration modes include oral and parenteral routes. Parenteral routes include, by way of example, intravenous, intramedullary, intraarterial, intramuscular, intradermal, subcutaneous, intranasal, and intraperitoneal routes. Components administered systemically can be modified or formulated to target, for example, HSCs, hematopoietic stem / progenitor cells, or erythroid progenitor or precursor cells.
[0244] Local administration modes include, by way of example, intramedullary injection into cancellous bone or intrafemoral injection into the medullary cavity, and injection into the portal vein. In certain embodiments, a significantly smaller amount of the component (compared to a systemic approach) can be effective when administered locally (e.g., directly into the bone marrow) compared to when administered systemically (e.g., intravenously). Local administration modes can reduce or eliminate the incidence of potentially toxic side effects that can occur when a therapeutically effective amount of the component is administered systemically.
[0245] Dosing can be provided as a periodic bolus (e.g., intravenously) or as a continuous infusion from an internal or external reservoir (e.g., an intravenous bag or implantable pump). The components can be administered locally, for example, by continuous release from a sustained release drug delivery device.
[0246] In addition, the components can be formulated to allow release over an extended period of time. The release system can include a matrix of biodegradable materials or materials that release incorporated components by diffusion. The components can be distributed uniformly or non-uniformly within the release system. A variety of release systems can be useful, but the selection of an appropriate system will depend on the release rate required by the particular application. Both non-degradable and degradable release systems can be used. Suitable release systems include, but are not limited to, polymers and polymeric matrices, non-polymeric matrices, or inorganic and organic excipients and diluents, such as calcium carbonate and sugars (e.g., trehalose). The release system can be natural or synthetic. However, synthetic release systems are generally preferred, as they produce more reliable, more reproducible, and more defined release profiles. The release system material can be selected such that components with different molecular weights are released by diffusion or degradation of the material.
[0247] Representative synthetic biodegradable polymers include polyamides (e.g., poly(amino acids) and poly(peptides)), polyesters (e.g., poly(lactic acid), poly(glycolic acid), poly(lactic-co-glycolic acid), and poly(caprolactone)), poly(anhydrides), polyorthoesters, polycarbonates, as well as their chemical derivatives (substitution, addition of chemical groups (e.g., alkyl, alkylene, hydroxylation, oxidation, and other modifications routinely performed by one of skill in the art)), copolymers, and mixtures thereof. Representative synthetic non-degradable polymers include, for example, polyethers (e.g., poly(ethylene oxide), poly(ethylene glycol), and poly(tetramethylene oxide)), vinyl polymers-polyacrylates and polymethacrylates (e.g., methyl, ethyl, other alkyl, hydroxyethyl methacrylate, acrylic, and methacrylic acid, and others (e.g., poly(vinyl alcohol), poly(vinyl pyrrolidone), and poly(vinyl acetate)), poly(urethanes), cellulose and its derivatives (e.g., alkyl, hydroxyalkyl, ethers, esters, nitrocellulose, and various cellulose acetates), polysiloxanes, and any of their chemical derivatives (substitution, addition of chemical groups (e.g., alkyl, alkylene, hydroxylation, oxidation, and other modifications routinely performed by one of ordinary skill in the art)), copolymers, and mixtures thereof.
[0248] Poly(lactide-co-glycolide) microspheres can also be used. Typically, the microspheres are composed of polymers of lactic acid and glycolic acid structured to form hollow spheres. The spheres can be about 15-30 microns in diameter and can be loaded with the components described herein. In some embodiments, the genome editing system, system components and / or nucleic acids encoding the system components are delivered in block copolymers such as poloxamers or poloxamines.
[0249] Multimodal or differential delivery of ingredients Those skilled in the art will understand in light of the present disclosure that the different components of the genome editing systems disclosed herein can be delivered together or separately, simultaneously or non-simultaneously. Separate and / or asynchronous delivery of genome editing system components may be particularly desirable to provide temporal or spatial control over the function of the genome editing system and limit certain effects caused by their activity.
[0250] As used herein, different or differential modality refers to a delivery modality that confers different pharmacodynamic or pharmacokinetic properties to a component molecule of interest (e.g., an RNA-guided nuclease molecule, a gRNA, a template nucleic acid, or a payload). For example, a delivery modality may result in a different tissue distribution, a different half-life, or a different temporal distribution, e.g., in a selected compartment, tissue, or organ.
[0251] Some modes of delivery, such as delivery by nucleic acid vectors that persist in the cell or cell progeny by autonomous replication or insertion into the cell's nucleic acid, result in more sustained expression and presence of the component. Examples include viral (e.g., AAV or lentiviral) delivery.
[0252] By way of example, components of the genome editing system (e.g., RNA-guided nuclease and gRNA) may be delivered by modalities that differ in terms of the half-life or persistence of the delivered components in the body, or in a particular compartment, tissue, or organ. In certain embodiments, the gRNA may be delivered by such a modality. The RNA-guided nuclease molecule component may be delivered by a modality that results in less persistence or less exposure to the body, or in a particular compartment, tissue, or organ.
[0253] More generally, in certain embodiments, a first delivery mode is used to deliver a first component and a second delivery mode is used to deliver a second component. The first delivery mode confers a first pharmacodynamic or pharmacokinetic property. The first pharmacodynamic property can be, for example, the distribution, persistence, or exposure of the component or the nucleic acid encoding the component in the body, compartment, tissue, or organ. The second delivery mode confers a second pharmacodynamic or pharmacokinetic property. The second pharmacodynamic property can be, for example, the distribution, persistence, or exposure of the component or the nucleic acid encoding the component in the body, compartment, tissue, or organ.
[0254] In certain embodiments, a first pharmacodynamic or pharmacokinetic property, eg, distribution, duration, or exposure, is more specific than a second pharmacodynamic or pharmacokinetic property.
[0255] In certain embodiments, the first delivery modality is selected, for example, to optimize (eg, minimize) pharmacodynamic or pharmacokinetic properties, such as distribution, duration, or exposure.
[0256] In certain embodiments, the second delivery modality is selected, for example, to optimize (eg, maximize) pharmacodynamic or pharmacokinetic properties, such as distribution, duration, or exposure.
[0257] In certain embodiments, the first delivery modality involves the use of a relatively persistent element, e.g., a nucleic acid, e.g., a plasmid or a viral vector (e.g., AAV or lentivirus). Such vectors will be relatively persistent because of the relatively persistent products transcribed therefrom.
[0258] In certain embodiments, the second delivery modality comprises a relatively transient element (eg, RNA or protein).
[0259] In certain embodiments, the first component comprises a gRNA and the delivery mode is relatively sustained, e.g., the gRNA is transcribed from a plasmid or viral vector (e.g., AAV or lentivirus). Transcription of these genes will have little physiological consequence, as the genes do not code for protein products and the gRNA cannot act alone. The second component, the RNA-guided nuclease molecule, is delivered in a transient manner, e.g., as an mRNA or protein, ensuring that the complete RNA-guided nuclease molecule / gRNA complex is present and active only for a short period of time.
[0260] Furthermore, the components can be delivered in different molecular forms or in different delivery vectors that complement each other to enhance safety and tissue specificity.
[0261] The use of different delivery modes can enhance performance, safety, and / or efficacy, for example, reducing the possibility of eventual off-target modifications. Delivery of immunogenic components (e.g., Cas9 molecules) in a less persistent manner can reduce immunogenicity, since peptides from bacterial Cas enzymes are presented on the surface of cells by MHC molecules. A two-part delivery system can alleviate these drawbacks.
[0262] Different delivery modalities can be used to deliver components to different but overlapping target regions. Formation active complexes are minimized outside the overlap of target regions. Thus, in certain embodiments, a first component, e.g., gRNA, is delivered by a first delivery modality that results in a first spatial (e.g., tissue) distribution. A second component, e.g., RNA-guided nuclease molecule, is delivered by a second delivery modality that results in a second spatial (e.g., tissue) distribution. In certain embodiments, the first modality includes a first element selected from liposomes, nanoparticles (e.g., polymeric nanoparticles), and nucleic acids (e.g., viral vectors). The second modality includes a second element selected from that group. In certain embodiments, the first delivery modality includes a first targeting element (e.g., a cell-specific receptor or antibody) and the second delivery modality does not include that element. In certain embodiments, the second delivery modality includes a second targeting element (e.g., a second cell-specific receptor or a second antibody).
[0263] When RNA-guided nuclease molecules are delivered in viral delivery vectors, liposomes, or polymer nanoparticles, there is the possibility of delivery and therapeutic activity to multiple tissues when it is desirable to target only a single tissue. A two-part delivery system can solve this problem and increase tissue specificity. When gRNA and RNA-guided nuclease molecules are packaged in separate delivery vehicles with different but overlapping tissue tropisms, a fully functional complex is formed only in tissues targeted by both vectors. EXAMPLES
[0264] The above principles and embodiments are further illustrated by the following non-limiting examples:
[0265] Example 1: Use of ribonucleoproteins for the treatment of β-hemoglobinopathies Described herein is an autologous cell therapy for beta thalassemia, which comprises administering genetically modified CD34+ cells to a subject suffering from beta thalassemia to promote the expression of gamma globin. In certain embodiments, the beta thalassemia can be transfusion-dependent beta thalassemia (TDT). Beta thalassemia is one of the most common recessive blood disorders in the world, with over 200 mutations identified to date. These mutations reduce or completely abolish the expression of beta globin. When beta globin pairs with alpha globin to form adult hemoglobin (HbA, α2β2), the reduction or absence of beta globin leads to excess alpha globin chains, which form toxic aggregates. These aggregates cause maturation blockage and premature death of red blood cell precursors, as well as hemolysis of red blood cells (RBCs), resulting in various degrees of anemia. Patients with the most severe form of beta thalassemia (ie, beta thalassemia major) are transfusion dependent, i.e., they require life-long RBC transfusions with the burden of iron chelation therapy.
[0266] The autologous cell therapy described herein is a therapeutic approach for treating beta thalassemia to promote the expression of fetal hemoglobin by directly targeting the promoters of the HBG1 and HBG2 genes that encode the fetal gamma globin chains. Gamma globin pairs with excess alpha globin chains to form fetal hemoglobin (HbF, α2γ2), thereby reducing the imbalance between alpha globin and beta globin chains in beta thalassemia. Gamma globin induction, and consequently HbF induction, can be achieved via Cpf1 (Cas12a) ribonucleoprotein (RNP)-mediated editing of the distal CCAAT box region of the HBG1 and HBG2 promoters, where naturally occurring genetic persistence of fetal hemoglobin (HFPH) mutations exists.
[0267] RNP32 (Table 10), comprising a gRNA (comprising the sequence set forth in SEQ ID NO: 1051) and a modified Cpf1 protein (comprising the sequence set forth in SEQ ID NO: 1097), highly efficiently and specifically edits the distal CCAAT box of the HBG1 and HBG2 promoters.
[0268] To test whether RNP32 could be an effective therapy for beta thalassemia (e.g., TDT), mPB CD34+ cells from individuals with TDT were electroporated with RNP32 targeting the HBG1 and HBG2 promoters. The efficiency of RNP32 editing for such cell therapy was determined and compared using mPB CD34+ cells obtained from individuals with TDT and normal donors. Briefly, CD34+ cells from normal or TDT donors were pre-stimulated for 2 days in a humidified incubator at 37°C with 5% carbon dioxide (CO2) in a medium consisting of X-Vivo10 supplemented with 1×Glutamax, 100 ng / mL stem cell factor (SCF), 100 ng / mL thrombopoietin (TPO), and 100 ng / mL FMS-like tyrosine kinase 3 ligand (Flt3L). After 2 days of culture, cells were harvested and resuspended in MaxCyte electroporation buffer. RNP32 (6 μM, gRNA / protein molar ratio of 2) was delivered to CD34+ cells via a MaxCyte GT electroporation device. 1 × 10 per OC-100 cartridge. 6 ~6.25×10 6 10 cells can be used for electroporation. Pre-warmed complete medium is then added to the cells to obtain approximately 1 x 10 6A final cell density of 1000 cells / mL was obtained. The electroporated cells were then placed in a humidified incubator at 37° C. and 5% CO2 along with untreated control cells (cells that were not electroporated). One, two, and three days after electroporation, aliquots of cells were taken for further analysis. Extraction of crude genomic deoxyribonucleic acid (gDNA) was performed by subjecting the lysate to the following conditions in a thermocycler: 65° C. for 15 minutes, followed by 95° C. for 10 minutes. The crude gDNA was then analyzed for indels by next generation sequencing using the following primers: forward=CATGGCGTCTGGACTAGGAG (SEQ ID NO: 1266), and reverse=AAACACATTTCACAATCCCTGAAC (SEQ ID NO: 1267).
[0269] As shown in Figure 3A, RNP32 edited mPB CD34+ cells from individuals with TDT as efficiently as CD34+ cells from normal donors. The percentage of indels increased from day 1 post-electroporation to day 3 post-electroporation for cells from both TDT donors (Figures 3A, 3B) and normal donors (Figure 3A). Furthermore, in addition to efficient editing (Figure 3A), RNP32-edited mPB CD34+ cells from individuals with TDT maintained high viability from day 1 post-electroporation to day 3 post-electroporation (Figure 3C).
[0270] We next examined erythroid differentiation of RNP32-edited beta-thalassemia CD34+ cells from three individuals with TDT (donors 1–3) to assess the maturation and health of RNP-edited erythroid cells, as maturation block and premature death of erythroid precursors are hallmarks of TDT.
[0271] Briefly, 1 day after electroporation with RNP32, cells were cultured in erythroid induction medium to generate erythroid cells. CD34+ cells were cultured for 7 days in step 1 medium consisting of Iscove's modified Dulbecco's medium (IMDM) supplemented with 1× GlutaMAX (Gibco), 100 U / mL penicillin, 100 mg / mL streptomycin, 5% human AB+ plasma, 330 μg / mL human holotransferrin, 20 mg / mL human insulin, 2 U / mL heparin, 3 U / mL recombinant human erythropoietin (EPO), 100 ng / mL SCF, and 5 ng / mL interleukin (IL)-3. On day 7, cells were transferred to step 2 medium (identical to step 1 medium except for the absence of IL-3) and cultured for 4 days. Cells were then cultured for 7 days in Step 3 medium (similar to Step-2 medium, but without the addition of SCF and with 5% human AB+ plasma replaced by 5% KnockOut Serum Replacement (Gibco)). At the end of 18 days of culture, the frequencies of erythrocyte maturation, enucleation, and cell death were determined using fluorescence-activated cell sorting (FACS).
[0272] Erythroid differentiation of edited beta thalassemia CD34+ cells showed significant improvements in red blood cell maturation and health. Day 18 erythroid cells were stained with antibodies against CD71 and CD235a, NucRed to stain cells containing nuclei, and DAPI to stain dead cells. Erythroblasts were classified as live, nucleated, and high CD235a populations. Late erythroblasts were classified as erythroblasts with low or negative CD71 expression.
[0273] Beta-thalassemia CD34+ donor cells edited with RNP32 underwent successful erythroid differentiation at a similar rate to unedited control cells (Figure 4A, 4B). Approximately 70% of edited erythroblasts reached the late erythroblast stage compared to approximately 53% of unedited erythroblasts (Figure 4C). Enucleated erythroid cells were classified as NucRed negative cells within the viable and high CD235a population. Approximately 56% of edited erythroid cells were terminally mature and enucleated compared to approximately 28% of unedited erythroid cells (Figure 4D). Those staining positive with DAPI within the nucleated and high CD235a population were classified as nonviable erythroblasts. Nonviable erythroblasts were reduced from approximately 33% to approximately 22% after editing (Figure 4E). Figures 4F-4H show the percentage of cells (edited and unedited) that reached the late erythroblast stage, the percentage of enucleated erythroid cells, and the percentage of non-viable erythroblasts in erythroid cultures from a single donor on days 7, 11, 14, and 18, respectively.
[0274] Changes in γ-globin and total globin production at both the mRNA and protein levels were assessed in erythroid cells differentiated from beta-thalassemia CD34+ donor cells edited with RNP32 or in non-edited cells using reverse transcription droplet digital polymerase chain reaction and reverse phase ultra-performance liquid chromatography (RP-UPLC). The total area under the curve for alpha, beta, and gamma globin was calculated against a standard curve to determine the hemoglobin content per cell. Results showed that improved erythropoiesis was accompanied by significantly increased γ-globin and total hemoglobin levels compared to non-edited controls at both the mRNA and protein levels (Figures 5A-5E). These data strongly support that editing of the CCAAT box in the HBG1 and HBG2 promoters using RNP32 can reverse erythropoiesis defects associated with beta-thalassemia and increase hemoglobin production.
[0275] In summary, the data herein support the use of RNP32 in autologous cell therapy for beta-thalassemia. Erythroid cells differentiated from RNP32-edited beta-thalassemia CD34+ donor cells showed significantly improved erythrocyte maturation and reduced erythrocyte death, thus reversing the maturation block associated with TDT mutations. Erythroid cells differentiated from RNP32-edited beta-thalassemia CD34+ donor cells showed significantly increased γ-globin production and total hemoglobin content per cell. Treatment with RNP32 can help address the underlying disease mechanisms of TDT, showing improved erythropoiesis and increased hemoglobin content in their erythroid progeny. Because edited mPB CD34+ cells retain the ability to engraft and result in robust HbF induction over time, these data support that RNP32 can be used as a one-time effective autologous cell therapy for individuals with TDT to reverse erythropoiesis defects and alleviate anemia.
[0276] Example 2: Treatment of β-hemoglobinopathies using edited hematopoietic stem cells The methods and genome editing systems disclosed herein can be used for the treatment of β-hemoglobinopathies, such as sickle cell disease or beta-thalassemia, in patients in need thereof. For example, genome editing can be performed on cells derived from the patient in an autologous treatment. Modification of the patient's cells ex vivo and reintroduction of the cells into the patient can result in increased HbF expression and treatment of β-hemoglobinopathies.
[0277] For example, HSCs can be extracted from bone marrow of patients with β-hemoglobinopathies using techniques well known to those of skill in the art. HSCs can be modified using the methods disclosed herein for genome editing. For example, HSCs can be edited using an RNP consisting of a guide RNA (gRNA) targeting one or more regions of the HBG gene complexed with an RNA-guided nuclease. In certain embodiments, the RNA-guided nuclease can be a Cpf1 protein. In certain embodiments, the Cpf1 protein can be a modified Cpf1 protein. In certain embodiments, the modified Cpf1 protein can be encoded by the sequence shown in SEQ ID NOs: 1000, 1001, 1008-1018, 1032, 1035-39, 1094-1097, 1107-09 (Cpf1 polypeptide sequence), or SEQ ID NOs: 1019-1021, 1110-17 (Cpf1 polynucleotide sequence). For example, the modified Cpf1 protein may be encoded by the sequence set forth in SEQ ID NO: 1097. In certain embodiments, the gRNA may be a modified gRNA or an unmodified gRNA. In certain embodiments, the gRNA may comprise a sequence set forth in Tables 7, 8, 11, or 12. For example, in certain embodiments, the gRNA may comprise a sequence set forth in SEQ ID NO: 1051. In certain embodiments, the RNP complex may comprise an RNP complex set forth in Table 10. For example, the RNP complex may comprise a gRNA comprising a sequence set forth in SEQ ID NO: 1051 and a modified Cpf1 protein (RNP32, Table 10) encoded by a sequence set forth in SEQ ID NO: 1097. In certain embodiments, the modified HSCs have an increased frequency or level of indels in the human HBG1 gene, the HBG2 gene, or both, compared to unmodified HSCs. In certain embodiments, the modified HSCs are capable of differentiating into erythroid cells expressing increased levels of HbF. A population of modified HSCs can be selected for reintroduction into the patient via transfusion or other methods known to one of skill in the art. A population of modified HSCs for reintroduction can be selected based on, for example, increased HbF expression in erythroid progeny of the modified HSCs, or increased indel frequency in the modified HSCs.In some embodiments, any form of ablation prior to reintroduction of the cells may be used to enhance engraftment of the modified HSCs. In other embodiments, peripheral blood stem cells (PBSCs) may be extracted from a patient with β-hemoglobinopathy using techniques well known to those skilled in the art (e.g., apheresis or leukapheresis), and stem cells may be removed from the PBSCs. The above genome editing methods may be performed on the stem cells, and the modified stem cells may be reintroduced into the patient as described above.
[0278] [Table 6-1]
[0279] [Table 6-2]
[0280] [Table 6-3]
[0281] [Table 6-4]
[0282] [Table 7-1]
[0283] [Table 7-2]
[0284] [Table 7-3]
[0285] [Table 7-4]
[0286]
Table 7-5
[0287]
Table 8-1
[0288]
Table 8-2
[0289]
Table 8-3
[0290]
Table 8-4
[0291]
Table 8-5
[0292]
Table 8-6
[0293]
Table 8-7
[0294]
Table 8-8
[0295]
Table 8-9
[0296]
Table 8-10
[0297]
Table 8-11
[0298]
Table 8-12
[0299]
Table 9
[0300]
Table 10-1
[0301]
Table 10-2
[0302]
Table 11-1
[0303]
Table 11-2
[0304]
Table 12-1
[0305]
Table 12-2
[0306] [Table 13-1]
[0307] [Table 13-2]
[0308] array Genome editing system components according to the present disclosure (including but not limited to RNA-guided nucleases, guide RNAs, donor template nucleic acids, nucleic acids encoding nucleases or guide RNAs, and portions or fragments of any of the foregoing) are exemplified by the nucleotide and amino acid sequences presented in the Sequence Listing. The sequences presented in the Sequence Listing are not intended to be limiting, but rather to illustrate certain principles of genome editing systems and their component parts, and in combination with the present disclosure will inform one of skill in the art of additional implementations and modifications that are within the scope of the present disclosure.
[0309] Incorporation by Reference All publications, patents, and patent applications mentioned in this specification are incorporated herein by reference in their entirety as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. In case of conflict, the present application, including any definitions herein, will control.
[0310] Equivalent Those skilled in the art will recognize, or be able to ascertain using no more than routine experimentation, many equivalents to the specific embodiments described herein which equivalents are intended to be encompassed by the following claims.
[0311] References Ahern et al., Br J Haematol 25(4):437-444 (1973) Akinbami Hemoglobin 40:64-65 (2016) Aliyu et al. Am J Hematol 83:63–70 (2008) Anders et al. Nature 513(7519):569-573 (2014) Angastiniotis & Modell Ann NY Acad Sci 850:251-269 (1998) Bae et al. Bioinformatics 30(10):1473-1475 (2014) Barbosa et al. Braz J Med Bio Res 43(8):705-711 (2010) Bauer et al. Nat. Med. 25(5):776-783 (2019) Bothmer et al. CRISPR J 3(3):177-187 (2020) Bouva Hematologica 91(1):129-132 (2006) Briner et al. Mol Cell 56(2):333-339 (2014) Brousseau Am J Hematol 85(1):77-78 (2010) Caldecott Nat Rev Genet 9(8):619-631 (2008) Canvers et al. Nature 527(12):192-197 (2015) Chang et al. Mol Ther Methods Clin Dev 4:137-148 (2017) Chassanidis Ann Hematol 88(6):549-555 (2009) Chylinski et al. RNA Biol 10(5):726-737 (2013) Cong et al. Science 399(6121):819-823 (2013) Costa et al., Cad Saude Publica 18(5):1469-1471 (2002) Davis & Maizels 2 Proc Natl Acad Sci USA 111(10):E924-932 (2014) Fine et al. SciRep. 5:10777 (2015) Frit et al. DNA Repair (Amst) 17:81-97 (2014) Fu et al. Nat Biotechnol 32:279-284 (2014) Gao et al. Nat Biotechnol.: 35(8):789-792 (2017) Giarratana et al. Nat Biotechnol. 23(1):69-74 (2005) Giarratana et al. Blood:118, 5071‐5079 (2011) Giannoukos et al. BMC Genomics 19(1):212 (2018) Guilinger et al. Nat Biotechnol 32:577-582 (2014) Heigwer et al. Nat Methods 11(2):122-123 (2014) Hsu et al. Nat Biotechnol 31(9):827–832 (2013) Iyama & Wilson DNA Repair (Amst) 12(8):620-636 (2013) Jiang et al. Nat Biotechnol 31(3):233-239 (2013) Jinek et al. Science 337(6096):816-821 (2012) Jinek et al. Science 343(6176):1247997 (2014) Kleinstiver et al. Nature 523(7561):481-485 (2015a) Kleinstiver et al. Nat Biotechnol 33(12):1293-1298 (2015b) Kleinstiver et al. Nature 529(7587):490-495 (2016) Kleinstiver et al. Nat Biotechnol 37(3):276-282 (2019) Komor et al. Nature 533(7603):420-424 (2016) Kosicki et al. Nat Biotechnol 36(8): 765-771 (2018) Lee et al. Nano Lett 12(12):6322-6327 (2012) Lewis "Medical-Surgical Nursing: Assessment and Management of Clinical Problems" (2014) Li Cell Res 18(1):85-98 (2008) Makarova et al. Nat Rev Microbiol 9(6):467-477 (2011) Mali et al. Science 339(6121):823-826 (2013) Mantovani et al. Nucleic Acids Res 16(16):7783-7797 (1988) Masala Methods Enzymol. 231:21-44 (1994) Marteijn et al. Nat Rev Mol Cell Biol 15(7):465-481 (2014) Martyn et al., Biochim Biophys Acta 1860(5):525-536 (2017) Metais et al. Blood Adv. 3(21):3379-92 (2019) Nishimasu et al. Cell 156(5):935-949 (2014) Nishimasu et al. Cell 162:1113-1126 (2015) Notta et al. Science 333(6039):218-21 (2011) Pausch et al. Science 369(6501):333-337 (2020) Ran & Hsu Cell 154(6):1380-1389 (2013) Richardson et al. Nat Biotechnol 34:339-344 (2016) Swarts et al. May 22:e1481. doi: 10.1002 / wrna.1481. Epub ahead of print. PMID: 29790280 (2018) Shmakov et al. Molecular Cell 60(3):385-397 (2015) Sternberg et al. Nature 507(7490):62-67 (2014) Strohkendl et al. Mol. Cell 71:816-824 (2018) Superti-Furga et al. EMBO J 7(10):3099-3107 (1988) Thein Hum Mol Genet 18(R2):R216-223 (2009) Thorpe et al. Br J Haematol. 87(1):125-132 (1994) Tsai et al. Nat Biotechnol 34(5): 483 (2016) Waber et al. Blood 67(2):551-554 (1986) Wang et al. Cell 153(4):910-918 (2013) Weber et al. Sci Adv. 6(7):eaay9392 (2020) Wu et al. Nat. Med. 25(5): 776-83 (2019) Xiao et al. Bioinformatics 30(8):1180–1182 (2014) Xu et al. Genes Dev 24(8):783–798 (2010) Yamano et al. Cell 165(4): 949-962 (2016) Zetsche et al. Nat Biotechnol 33(2):139-42 (2015)
Claims
1. 1. A pharmaceutical composition comprising a modified population of isolated cells for use in a method of alleviating one or more symptoms of beta-thalassemia (β-Thal) in a subject in need thereof, said method comprising: a) providing a population of CD34+ or hematopoietic stem cells isolated from said subject; b) modifying the population of isolated cells ex vivo by delivering to the population of isolated cells an RNP complex, thereby altering the promoter of the HBG gene in one or more isolated cells within the population, wherein the RNP complex: Cpf1 and A gRNA, 5' end and 3' end, an RNA or DNA extension at the 5' end; modifications (e.g., phosphorothioate linkages and / or 2'-O-methyl modifications at the 5' or 3' termini); and a gRNA comprising a targeting domain that is complementary to a target site in the promoter of the HBG gene; c) administering the population of modified isolated cells to the subject, thereby alleviating one or more symptoms of β-Thal in the subject.
2. 2. The pharmaceutical composition of claim 1, wherein the DNA stretch comprises a sequence selected from the group consisting of SEQ ID NOs: 1235-1250.
3. 2. The pharmaceutical composition of claim 1, wherein the targeting domain comprises or consists of a sequence shown in Tables 7, 8, 11, and 12.
4. 2. The pharmaceutical composition of claim 1, wherein the target site comprises a nucleotide located at Chr11 (NC_000011.10) 5,249,904-5,249,927 (Table 6, Region 6), Chr11 (NC_000011.10) 5,254,879-5,254,909 (Table 6, Region 16), or a combination thereof.
5. 2. The pharmaceutical composition of claim 1, wherein the Cpf1 comprises one or more modifications selected from the group consisting of one or more mutations in the wild-type Cpf1 amino acid sequence, one or more mutations in the wild-type Cpf1 nucleic acid sequence, one or more nuclear localization signals (NLS), one or more purification tags, and combinations thereof.
6. 2. The pharmaceutical composition of claim 1, wherein the Cpf1 comprises or consists of a sequence selected from the group consisting of SEQ ID NOs: 1000, 1001, 1008-1018, 1032, 1035-39, 1094-1097, and 1107-09.
7. 2. The pharmaceutical composition of claim 1, wherein the Cpf1 comprises or consists of a sequence selected from the group consisting of SEQ ID NOs: 1019-1021 and 1110-17.
8. The pharmaceutical composition of claim 1 , wherein the RNP complex is delivered to the cell using electroporation.
9. 1. A method for inducing hemoglobin (Hb) expression ex vivo in a population of CD34+ or hematopoietic stem cells from a subject with beta-thalassemia (β-Thal), comprising: delivering an RNP complex comprising a guide RNA (gRNA) and Cpf1 to a population of unmodified CD34+ or hematopoietic stem cells from a subject having β-Thal to generate a population of modified CD34+ or hematopoietic stem cells comprising an indel, wherein the gRNA comprises a gRNA targeting domain; each modified CD34+ or hematopoietic stem cell comprises an indel in the HBG gene promoter; The method, wherein the modified CD34+ or hematopoietic stem cell population comprises a higher Hb level than the unmodified CD34+ or hematopoietic stem cell population.
10. 10. The method of claim 9, wherein the gRNA comprises a DNA stretch comprising a sequence selected from the group consisting of SEQ ID NOs: 1235-1250.
11. 10. The method of claim 9, wherein the gRNA targeting domain comprises or consists of a sequence set forth in Tables 7, 8, 11, and 12.
12. 10. The method of claim 9, wherein the gRNA comprises a targeting domain that is complementary to a target site in the promoter of the HBG gene, and the target site comprises nucleotides located at Chr11 (NC_000011.10) 5,249,904 to 5,249,927 (Table 6, Region 6), Chr11 (NC_000011.10) 5,254,879 to 5,254,909 (Table 6, Region 16), or a combination thereof.
13. 10. The method of claim 9, wherein the RNP complex comprises Cpf1 comprising one or more modifications selected from the group consisting of one or more mutations in a wild-type Cpf1 amino acid sequence, one or more mutations in a wild-type Cpf1 nucleic acid sequence, one or more nuclear localization signals (NLS), one or more purification tags, and combinations thereof.
14. 10. The method of claim 9, wherein the Cpf1 comprises or consists of a sequence selected from the group consisting of SEQ ID NOs: 1000, 1001, 1008-1018, 1032, 1035-39, 1094-1097, and 1107-09.
15. 10. The method of claim 9, wherein the Cpf1 comprises or consists of a sequence selected from the group consisting of SEQ ID NOs: 1019-1021 and 1110-17.
16. 10. The method of claim 9, wherein the RNP complex is delivered to the cell using electroporation.