Compositions targeting HBG1 and HBG2 and methods of use thereof

By using a fusion protein containing inactivated Cas9 and Clo051 peptides combined with gRNA targeting the HBG1 and HBG2 genes, the unintended effects and drug complications of gene editing in existing technologies were resolved. This resulted in the effective increase of γ-globulin and fetal hemoglobin expression in hematopoietic stem cells, thereby improving the treatment efficacy of hemoglobinopathies.

CN121335979APending Publication Date: 2026-01-13POSEIDA THERAPEUTICS INC
View PDF 14 Cites 0 Cited by

Patent Information

Application Number
CN202480032662.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-05-17
Filing Date
2024-05-16
Publication Date
2026-01-13

AI Technical Summary

Technical Problem

Existing gene editing technologies exhibit unintended insertion and deletion effects when treating hemoglobinopathies such as sickle cell disease and β-thalassemia, and the use of existing drugs such as hydroxyurea can lead to serious complications. Therefore, there is a need for improved therapeutic compositions and methods to enhance the precision and safety of genome editing.

Method used

A composition comprising first and second guide RNA (gRNA) and a fusion protein containing inactivated Cas9 peptide and Clo051 peptide is used. Gene modification is performed by targeting the HBG1 and HBG2 loci to introduce insertions/deletions to increase the expression of gamma globulin and fetal hemoglobin. The fusion protein is delivered using lipid nanoparticles.

Benefits of technology

This study significantly increased the expression of gamma globulin and fetal hemoglobin in hematopoietic stem cells, improved the efficacy of hemoglobinopathies treatment, reduced unexpected genomic insertions and deletions, and lowered the risk of drug complications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121335979A_ABST
    Figure CN121335979A_ABST
Patent Text Reader

Abstract

Methods and compositions for functional genetic modification at selected genomic sites, such as the BCL11A gene, the HMG1 promoter, and / or the HMG2 promoter, are disclosed. Also provided are cell populations comprising a functional genetic modification at one or more selected loci.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross Reference to Related Applications

[0002] This application claims the benefit of U.S. Provisional Patent Application No. 63 / 502,903, filed May 17, 2023, which is incorporated by reference herein in its entirety. TECHNICAL FIELD

[0003] The present disclosure relates to the field of gene editing and genome engineering. More specifically, the present disclosure relates to compositions and methods for targeted genetic modification and modulating expression of target nucleic acid sequences and applications thereof.

[0004] INCORPORATION BY REFERENCE OF THE SEQUENCE LISTING

[0005] The instant application contains a Sequence Listing which has been submitted via Patent Center in XML format and is hereby incorporated by reference in its entirety. The XML copy, created on May 8, 2024, is named “POTH-085_001WO_SeqList_ST25.xml” and is 61,664 bytes in size. BACKGROUND

[0006] Genome editing refers to strategies and techniques for targeted, specific modification of the genetic information (genome) of a living organism. Genome engineering is an active area of research because of its broad possible applications, particularly in the field of human health, such as correcting genes that carry deleterious mutations or exploring the function of genes. Early techniques developed to insert transgenes into living cells were generally limited by the randomness of the location in the genome where the new sequence was inserted. In contrast to early techniques, common genome editing strategies allow for modification of specific regions of DNA, thereby increasing the precision of correction or insertion. While these platforms provide greater reproducibility and reduce the level of unintended effects of random insertion and deletion in the genome, limitations remain.

[0007] Hemoglobin (Hb) transports oxygen from the lungs to tissues in red blood cells or erythrocytes (RBCs). During prenatal development and shortly after birth, hemoglobin exists in the form of fetal hemoglobin (HbF), a tetrameric protein composed of two alpha (a)-globin chains and two gamma (y)-globin chains. HbF is largely replaced by adult hemoglobin (HbA), a tetrameric protein, in which the y-globin chains of HbF are replaced by beta (b)-globin chains through a process known as globin switching. The average adult produces less than 1% of total hemoglobin as HbF. The a-hemoglobin gene is located on chromosome 16, while the b-hemoglobin gene (HBB), the A gamma (yA)-globin chain (HBG1, also known as gamma globin A), and the G gamma (yG)-globin chain (HBG2, also known as gamma globin G) are located on chromosome 11 within the globin gene cluster (also known as the globin locus).

[0008] HBB mutations can result in hemoglobin disorders (i.e., hemoglobinopathies), including sickle cell disease (SCD) and b-thalassemia (b-Thal). Approximately 93,000 people in the United States are diagnosed with a hemoglobinopathy. Globally, 300,000 children are born each year with a hemoglobinopathy. Because these disorders are associated with HBB mutations, their symptoms typically do not appear until globin switches from HbF to HbA.

[0009] SCD is the most common inherited blood disorder in the United States, affecting approximately 80,000 people (Brousseau, 2010). SCD is most common in people of African descent, with a prevalence of SCD of 1 in 500 for these populations. In Africa, SCD affects 15 million people. SCD is also more common in people of Indian, Saudi Arabian, and Mediterranean descent. In Hispanic American descent, the prevalence of sickle cell disease is 1 in 1,000.

[0010] SCD is caused by the simple heterozygous mutation c.17A>T in the HBB gene (HbS mutation). The sickle mutation is a point mutation on HBB (GAG>GTG) that results in a valine substitution for glutamic acid at amino acid position 6 in exon 1. The valine at position 6 of the beta-hemoglobin chain is hydrophobic and causes a change in the protein conformation of beta-globin when the protein is not bound to oxygen. This change in conformation results in the HbS protein polymerizing in the absence of oxygen, resulting in RBC deformation (i.e., sickling). SCD is inherited in an autosomal recessive manner, so only patients with two HbS alleles have the disease. Heterozygous subjects have sickle cell trait and can suffer from anemia and / or painful crises if these subjects are severely dehydrated or hypoxic.

[0011] Sickle red RBCs cause a variety of symptoms, including anemia, sickle cell crises, vaso-occlusive crises, aplastic crises, and acute chest syndrome. Sickle RBCs are less elastic than wild-type RBCs and thus cannot easily pass through capillary beds and cause occlusion and ischemia (i.e., vaso-occlusion). Vaso-occlusive crises occur when sickle cells block blood flow in the capillary beds of organs, resulting in pain, ischemia, and necrosis. These episodes typically last 5-7 days. The spleen plays a role in clearing dysfunctional RBCs and thus is often enlarged in early childhood and often presents with vaso-occlusive crises. By the end of childhood, the spleen of SCD patients often infarcts, which leads to autopsplenectomy. Hemolysis is a common feature of SCD and results in anemia. Sickle cells survive in the circulation for 10-20 days, whereas healthy RBCs survive for 90-120 days. SCD subjects are transfused with blood as necessary to maintain adequate hemoglobin levels. Frequent transfusions put subjects at risk for infection with HIV, hepatitis B, and hepatitis C. Subjects can also suffer from acute chest crises and infarction of the extremities, end organs, and central nervous system.

[0012] Subjects with SCD have a decreased life expectancy. Prognosis for patients with SCD is steadily improving through careful, lifelong management of crises and anemia. As of 2001, the average life expectancy for subjects with sickle cell disease was in the 50s. Current treatment for SCD involves hydration and pain management during crises, and transfusion of blood as necessary to correct anemia.

[0013] Thalassemia (e.g., β-Thal, δ-Thal, and β / δ-Thal) causes chronic anemia. It is estimated that 1 in 100,000 people worldwide have β-thalassemia. Its prevalence is higher in certain populations, including people of European descent, with a prevalence of approximately 1 in 10,000. Unless they receive lifelong transfusions and chelation therapy, subjects with severe β-thalassemia, a more severe form of the disease, are at risk of death. In the United States, there are approximately 3000 subjects with severe β-thalassemia. Intermediate β-thalassemia does not require transfusions, but can cause growth retardation and severe systemic abnormalities, and often requires lifelong chelation therapy.

[0014] β-thalassemia is caused by mutations in the HBB gene. The most common HBB mutations that cause β-thalassemia are: c.-136C>G, c.92+1G>A, c.92+6T>C, c.93-21G>A, c.118C>T, c.316-106C>G, c.25_26delAA, c.27_28insG, c.92+5G>C, c.118C>T, c.135delC, c.315+1G>A, c.-78A>G, c.52A>T, c.59A>G, c.92+5G>C, c.124_127delTTCT, c.316-197C>T, c.-78A>G, c.52A>T, c.124_127delTTCT, c.316-197C>T, c.-138C>T, c.-79A>G, c.92+5G>C, c.75T>A, c.316-2A>G, and c.316-2A>C. These mutations, and others associated with β-thalassemia, cause mutations or deletions in the β-globin chain, which disrupts the normal ratio of Hb α- to β-hemoglobin. Excess α-globin chains precipitate in erythroid precursors in the bone marrow.

[0015] In severe β-thalassemia, both alleles of HBB contain nonsense, frameshift, or splicing mutations that result in the complete absence of β-globin production (denoted as β 0 / β 0 ). Severe β-thalassemia results in a severe reduction in β-globin chains, leading to significant precipitation of α-globin chains in RBCs and more severe anemia.

[0016] Intermediate β-thalassemia is caused by mutations in the 5' or 3' untranslated region of HBB, mutations in the promoter region or polyadenylation signal of HBB, or splicing mutations within the HBB gene. The patient genotype is denoted as β0 / β+ or β+ / β+. β0 denotes the absence of expression of β-globin chains; 3+ denotes a dysfunctional but present β-globin chain. Phenotypic expression varies from patient to patient. Because there is some amount of β-globin production, intermediate β-thalassemia causes less precipitation of α-globin chains in erythroid precursor cells, and the anemia symptoms are also less severe than in major β-thalassemia. However, erythroid lineage expansion secondary to chronic anemia has more severe consequences.

[0017] Subjects with major β-thalassemia are between the ages of 6 months and 2 years and present with developmental delay, fever, hepatosplenomegaly, and diarrhea. Adequate treatment includes regular transfusions. Therapy for β-thalassemia also includes splenectomy and treatment with hydroxyurea. If patients are regularly transfused, they will develop normally until the second decade, at which point they require chelation therapy (in addition to continued transfusions) to prevent complications of iron overload. Iron overload can manifest as delayed growth or delayed sexual maturation. In adulthood, inadequate chelation therapy can lead to cardiomyopathy, cardiac arrhythmias, liver fibrosis and / or cirrhosis, diabetes mellitus, thyroid and parathyroid abnormalities, thrombosis, and osteoporosis. Frequent transfusions also put subjects at risk for contracting HIV, hepatitis B, and hepatitis C.

[0018] Subjects with intermediate β-thalassemia generally present between the ages of 2-6 years. They generally do not require transfusions. However, skeletal abnormalities occur due to chronic hypertrophy of the erythroid lineage to compensate for chronic anemia. Subjects can have long bone fractures due to osteoporosis. Extramedullary erythropoiesis is common and leads to splenomegaly, hepatomegaly, and lymphadenopathy. It can also lead to spinal cord compression and neurological problems. Subjects also suffer from lower extremity ulcers and have an increased risk of thrombotic events, including stroke, pulmonary embolism, and deep vein thrombosis. Treatment for intermediate β-thalassemia includes splenectomy, folate supplementation, hydroxyurea therapy, and radiation of extramedullary masses. Chelation therapy is used for subjects who develop iron overload.

[0019] The life expectancy of β-thalassemia patients is generally reduced. Subjects with major β-thalassemia who do not receive transfusion therapy generally die in their twenties or thirties. Subjects with major β-thalassemia who receive regular transfusions and appropriate chelation therapy can live into their fifties or beyond. Heart failure secondary to iron overload is the leading cause of death in β-thalassemia subjects due to iron overload.

[0020] A number of new treatments for SCD and β-Thal are currently being developed. Current clinical trials are investigating delivery of an anti-sickling HBB gene via gene therapy. However, the long-term efficacy and safety of this approach is unknown. Transplantation of hematopoietic stem cells (HSCs) from HLA-matched allogeneic stem cell donors has been shown to cure SCD and β-Thal, but the procedure carries risks, including risks associated with ablative therapy, which is necessary to prepare a subject for transplant, increases the risk of life-threatening opportunistic infections, and the risk of graft versus host disease after transplant. Additionally, matched allogeneic donors are often not available. Reactivation of HBG (HBG1, HBG2) gene expression and induction of fetal hemoglobin (HbF) is an important therapeutic strategy to improve the clinical symptoms and severity of SCD. Hydroxyurea is the only FDA-approved drug that has been shown to induce HbF in patients with SCD, but its use can lead to serious complications. Over the past three decades, a number of other pharmacological agents have been investigated to reactivate HBG transcription in vitro, but few have gained FDA approval beyond arginine butyrate and decitabine; however, neither of these drugs meets the requirements for routine clinical use due to difficulties with oral delivery and failure to achieve therapeutic levels. Thus, improved therapeutic compositions and methods are needed to treat these hemoglobinopathies. Development of gene editing platforms with superior genome editing efficacy is needed to treat these hemoglobinopathies. SUMMARY

[0021] The present disclosure provides a composition comprising: a) a first guide RNA (gRNA) and a first fusion protein configured to form a complex with the first gRNA or a first polynucleotide encoding the first fusion protein, the first fusion protein comprising: a mutant Cas9 (dCas9) polypeptide or a dead nuclease domain thereof and a Clo051 polypeptide or a nuclease domain thereof, and b) a second gRNA and a second fusion protein configured to form a complex with the second gRNA or a second polynucleotide encoding the second fusion protein, the second fusion protein comprising: a dCas9 polypeptide or a dead nuclease domain thereof and a Clo051 polypeptide or a nuclease domain thereof, wherein i) the first gRNA comprises a first targeting sequence comprising a nucleotide sequence selected from SEQ ID NO: 1, 5, 7, 13, 15, or 17; and the second gRNA comprises a second targeting sequence comprising a nucleotide sequence of SEQ ID NO: 3, or ii) the first gRNA comprises a first targeting sequence comprising a nucleotide sequence selected from SEQ ID NO: 7, 9, or 13; and the second gRNA comprises a second targeting sequence comprising a nucleotide sequence of SEQ ID NO: 11.

[0022] In some embodiments, the first gRNA comprises a first scaffold sequence and the second gRNA comprises a second scaffold sequence, wherein the first scaffold sequence and the second scaffold sequence comprise a nucleotide sequence selected from SEQ ID NO: 20 or 21.

[0023] In some embodiments, i) the first gRNA comprises a nucleotide sequence selected from SEQ ID NO: 2, 6, 8, 16, or 18; and the second gRNA comprises a nucleotide sequence of SEQ ID NO: 4, or ii) the first gRNA comprises a nucleotide sequence selected from SEQ ID NO: 10, 14, or 19; and the second gRNA comprises a nucleotide sequence of SEQ ID NO: 12.

[0024] In some embodiments, the first gRNA, the second gRNA, or both the first gRNA and the second gRNA comprise one or more chemical modifications of ribonucleotides, ribonucleotide bases, or phosphodiester bonds. In some embodiments, the one or more chemical modifications comprise at least one chemically modified phosphodiester bond. In some embodiments, the at least one chemically modified phosphodiester bond is a phosphorothioate bond. In some embodiments, the composition comprises at least two phosphorothioate bonds at the 5' terminus of the first gRNA, the second gRNA, or both the first gRNA and the second gRNA. In some embodiments, the composition comprises a 2' O-Me chemical modification at the 3' terminus of the first gRNA, the second gRNA, or both the first gRNA and the second gRNA.

[0025] In some embodiments, the dCas9 of the fusion protein and / or the second fusion protein is derived from a S. pyogenes Cas9 polypeptide or a S. aureus Cas9 polypeptide. In some embodiments, the C-terminus of the dCas9 or its inactivated nuclease domain is joined to the N-terminus of the Clo051 polypeptide or its nuclease domain via a peptide linker sequence selected from GGGGS or SEQ ID NO: 23. In some embodiments, the first fusion protein comprises the amino acid sequence of SEQ ID NO: 39, 41, 42, or 43. In some embodiments, the second fusion protein comprises the amino acid sequence of SEQ ID NO: 39, 41, 42, or 43. In some embodiments, the first fusion protein and the second fusion protein are the same. In some embodiments, the first fusion protein and the second fusion protein are different.

[0026] In some embodiments, the first polynucleotide encoding the first fusion protein, the second polynucleotide encoding the second fusion protein, or both the first polynucleotide and the second polynucleotide are mRNA. In some embodiments, the mRNA comprises a 5'-cap.

[0027] The present disclosure provides a composition comprising: i) a first gRNA comprising a nucleotide sequence of SEQ ID NO: 2, a second gRNA comprising a nucleotide sequence of SEQ ID NO: 4, a first polynucleotide sequence encoding a first fusion protein of SEQ ID NO: 39, and a second polynucleotide sequence encoding a second fusion protein of SEQ ID NO: 39; ii) a first gRNA comprising a nucleotide sequence of SEQ ID NO: 6, a second gRNA comprising a nucleotide sequence of SEQ ID NO: 4, a first polynucleotide sequence encoding a first fusion protein of SEQ ID NO: 39, and a second polynucleotide sequence encoding a second fusion protein of SEQ ID NO: 39; iii) a first gRNA comprising a nucleotide sequence of SEQ ID NO: 8, a second gRNA comprising a nucleotide sequence of SEQ ID NO: 4, a first polynucleotide sequence encoding a first fusion protein of SEQ ID NO: 41, and a second polynucleotide sequence encoding a second fusion protein of SEQ ID NO: 41; iv) a first gRNA comprising a nucleotide sequence of SEQ ID NO: 10, a second gRNA comprising a nucleotide sequence of SEQ ID NO: 12, a first polynucleotide sequence encoding a first fusion protein of SEQ ID NO: 39, and a second polynucleotide sequence encoding a second fusion protein of SEQ ID NO: 42; v) a first gRNA comprising a nucleotide sequence of SEQ ID NO: 14, a second gRNA comprising a nucleotide sequence of SEQ ID NO: 12, a first polynucleotide sequence encoding a first fusion protein of SEQ ID NO: 42, and a second polynucleotide sequence encoding a second fusion protein of SEQ ID NO: 42; vi) a first gRNA comprising a nucleotide sequence of SEQ ID NO: 19, a second gRNA comprising a nucleotide sequence of SEQ ID NO: 12, a first polynucleotide sequence encoding a first fusion protein of SEQ ID NO: 42, and a second polynucleotide sequence encoding a second fusion protein of SEQ ID NO: 42; vii) a first gRNA comprising a nucleotide sequence of SEQ ID NO: 16, a second gRNA comprising a nucleotide sequence of SEQ ID NO: 4, a first polynucleotide sequence encoding a first fusion protein of SEQ ID NO: 41, and a second polynucleotide sequence encoding a second fusion protein of SEQ ID NO: 41;or viii) a first gRNA comprising the nucleotide sequence of SEQ ID NO: 18, a second gRNA comprising the nucleotide sequence of SEQ ID NO: 4, a first polynucleotide sequence encoding the first fusion protein of SEQ ID NO: 41, and a second polynucleotide sequence encoding the second fusion protein of SEQ ID NO: 41.

[0028] In some embodiments, the composition is encapsulated in at least one lipid nanoparticle (LNP) comprising: about 40.75% of a compound of Formula (I), by mole,

[0029] Formula (I)

[0030]

[0031] about 51.75% cholesterol by mole, about 5% DOPC by mole, and about 2.5% DMG-PEG2000 by mole; wherein the first and second polynucleotides are RNA molecules, and wherein the ratio of lipids to RNA molecules in the at least one nanoparticle is about 120: 1 (w / w).

[0032] In some embodiments, the composition is encapsulated in at least one LNP comprising: about 54% SS-OP by mole, about 35% cholesterol by mole, about 5% DOPC by mole, about 5% DSPC by mole, and about 1% DMG-PEG2000 by mole, wherein the first and second polynucleotides are RNA molecules, and wherein the ratio of lipids to RNA molecules in the at least one nanoparticle is about 100: 1 (w / w), and the total lipid is 25 nM.

[0033] In some embodiments, the composition is used to modify a HBG1 gene, a HBG2 gene, a BCL11 A gene, or a combination thereof, in a cell.

[0034] The present disclosure provides a method of modifying a population of cells, the method comprising contacting the population of cells with any of the compositions of the present disclosure, wherein the first gRNA forms a complex with the first targeting sequence and the first fusion protein, and the second gRNA forms a complex with the second targeting sequence and the second fusion protein, thereby generating an insertion / deletion (indel) between the first targeting sequence and the second targeting sequence, and generating a modified population of cells. In some embodiments, the insertion / deletion results in inactivation of a BCL11 A gene.

[0035] In some embodiments, the population of modified cells has increased expression of gamma globin relative to the population of unmodified cells by a fold of about 4-fold to about 9-fold. In some embodiments, the population of modified cells has increased fetal hemoglobin (HbF) expression levels relative to the population of unmodified cells.

[0036] In some embodiments, the cells are hematopoietic stem and progenitor cells (HSPCs). In some embodiments, the HSPCs are capable of differentiating into erythroid progenitor cells.

[0037] The present disclosure provides a population of cells modified according to any of the methods of the present disclosure.

[0038] The present disclosure provides a method of treating a beta-hemoglobinopathy in a subject in need thereof, the method comprising administering to the subject any of the compositions or cells of the present disclosure. In some embodiments, the beta-hemoglobinopathy is beta-thalassemia or sickle cell disease. BRIEF DESCRIPTION OF DRAWINGS

[0039] FIG. 1 is a graph showing indels in HSPCs after editing with compositions comprising Cas-CLOVER and different concentrations of gRNA pairs targeting HBG1 or HBG2. On the y-axis, indels are shown as a percentage of modified reads out of total sequencing reads. The x-axis shows the gRNA pairs tested (Pair #1, Pair #2, Pair #3, and Pair #4). Each gRNA pair was tested at three increasing concentrations (40 pg / ml, 200 pg / ml, and 400 pg / ml).

[0040] FIG. 2 is a graph showing indels in HSPCs after editing with compositions comprising Cas-CLOVER and high concentration (400 pg / ml) or low concentration (200 pg / ml) of gRNA pairs targeting HBG1 or HBG2.

[0041] FIG. 3 is a graph showing indels in HSPCs after editing with compositions comprising Cas-CLOVER and different concentrations of gRNA pairs targeting HBG1 or HBG2. On the y-axis, indels are shown as a percentage of modified reads out of total sequencing reads. The x-axis shows the gRNA pairs tested (Pair #2, Pair #3, and Pair #4).

[0042] FIGS. 4A-4B are a series of graphs showing absolute numbers of colony forming units (CFU) types of HSPCs after editing with compositions comprising Cas-CLOVER and different gRNA pairs targeting HBG1 or HBG2. Three controls were tested: EP only (nucleofection / electroporation only), CC only (Cas-CLOVER mRNA only), and sgRNA only. The following cell types were assayed: CFU-GEMM (CFU- granulocyte / erythrocyte / macrophage / megakaryocyte), CFU-GM (CFU-granulocyte / macrophage), BFU-E (burst-forming unit erythroid), and CFU-E (CFU-erythroid).

[0043] FIG. 5 is a graph showing HBG mRNA expression in HSPCs after editing with compositions comprising Cas-CLOVER and different gRNA pairs targeting HBG1 or HBG2. RT-qPCR results were normalized to HBB at days 14, 18, and 21 during erythroid differentiation. Adult and cord blood samples were used as negative and positive controls, respectively.

[0044] FIG. 6 is a graph showing HbF protein expression levels in HSPCs after editing with compositions comprising Cas-CLOVER and different gRNA pairs targeting HBG1 or HBG2. The left y-axis shows the percentage of F cells determined by flow cytometry. The right y-axis shows the mean fluorescence intensity of HbF signal per F cell relative to the EP only (nucleofection / electroporation only) control.

[0045] FIG. 7 is a schematic of a composition of the present disclosure. A first fusion protein (e.g., Cas-Clover comprising dCas9-linker-Clo051) is complexed with a first gRNA at the 5’ end of a genomic region. A second fusion protein (e.g., Cas-Clover comprising dCas9-linker-Clo051) is complexed with a second gRNA at the 3’ end of the genomic region. Targeting is provided with high efficiency and accuracy using gRNAs. Cleavage of the genomic DNA template only occurs when the Clo051 nuclease of the first fusion protein and the second fusion protein are in close proximity. In some cases, the HBG1 gene, the HBG2 gene, the BCL11A gene, or a combination thereof is cleaved.

[0046] FIG. 8 is a graph showing indels in HSPCs after editing with compositions comprising Cas-CLOVER variants and gRNA pairs targeting HBG1 or HBG2.

[0047] All documents cited herein, including any cross-referenced or related patents or applications, are hereby incorporated by reference in their entirety as to all purposes to the same extent as if each were individually and specifically indicated to be incorporated by reference. The citation of any document is not to be construed as an admission that it is prior art with respect to any application disclosed or claimed herein, or that it has any relevance to the patentability of any application disclosed or claimed herein. Furthermore, to the extent that any meaning or definition of a term in this document conflicts with any meaning or definition of the same term in a document incorporated by reference, the meaning or definition assigned to that term in this document shall control. DETAILED DESCRIPTION

[0048] The present disclosure provides compositions and methods for genetically modifying a genome to include a polynucleotide insertion, deletion, and / or substitution into chromosomal DNA that enhances transcription of HBG1 and / or HBG2 genes, which encode the γΑ and γθ subunits of hemoglobin, respectively. In particular, the present disclosure overcomes problems associated with current technologies by providing a method for efficiently genetically modifying a cell genome to include a polynucleotide insertion, deletion, and / or substitution that facilitates the therapeutic expression of fetal hemoglobin (HbF) for the treatment of hemoglobinopathies.

[0049] Fetal hemoglobin (HbF) expression can be induced using various genomic strategies. For example, HbF expression can be induced by targeted disruption of a region near the HBG1 and HBG2 promoter target sequence and / or erythroid-specific expression of the transcriptional repressor BCL11A (also discussed in commonly-assigned International Patent Publication No. WO 2015 / 148860 to Friedland et al. (“Friedland”), published October 1, 2015, which is incorporated by reference herein in its entirety) (which encodes a repressor that silences HBG1 and HBG2 (Canver 2015)). Genetic mapping and genome-wide association studies have identified a locus in B-cell lymphoma / leukemia 11A (BCL11A) (Xnm1 variant upstream of hemoglobin subunit gamma 1 (HBG1)). Supporting studies have reported that hemoglobin switching is controlled by the activity of epigenetic regulators such as BCL11A. The HBG promoter region of the HBG1 and HBG2 genes contains a DNA binding region for BCL11A, which is an effective silencer of HbF expression. Binding of BCL11A to the HBGB promoter inhibits expression of the gamma subunits 1 and 2, which can result in inhibition of the gamma- to beta-globin switching process.

[0050] The genome editing systems of the present disclosure can include two or more fusion proteins (e.g., Cas-Clover) and two or more gRNAs having targeting domains complementary to sequences in or near the targeted region. In certain embodiments, the DNA binding region of BCL11A is targeted for disruption. In certain embodiments, the promoter region of HBG1 and / or HBG2 is targeted for disruption. In certain embodiments, the genome editing systems disclosed herein can be used to introduce polynucleotide insertions, deletions, and / or substitutions in the targeted region.

[0051] Treating hemoglobinopathies by gene therapy and / or genome editing becomes complicated because the cells affected by the disease phenotype (i.e., red blood cells or RBCs) are enucleated and do not contain genetic material encoding the abnormal hemoglobin protein (Hb) subunits or the γA or γG subunits targeted in the above example genome editing methods. In certain embodiments of the present disclosure, this complication is addressed by altering cells that are capable of differentiating into or otherwise producing red blood cells. Cells within the erythroid lineage that are altered in accordance with various embodiments of the present disclosure include, but are not limited to, hematopoietic stem and progenitor cells (HSPCs), erythroblasts (including basophilic, polychromatic, and / or orthochromatic erythroblasts), proerythroblasts, polychromatic or reticulocytes, embryonic stem (ES) cells, and / or induced pluripotent stem (iPSC) cells. These cells can be altered in situ (e.g., within a subject’s tissue) or ex vivo.

[0052] The present disclosure overcomes problems associated with current technologies by providing compositions comprising genetically engineered fusion molecules (e.g., Cas-Clover) for use in targeting reduction or elimination of gene products in cells for in vivo gene therapy. Compositions comprising genetically engineered fusion molecules of the present disclosure can be used to treat genetic diseases. Non-limiting examples of genetic diseases include hemoglobinopathies, such as sickle cell disease or beta-thalassemia. Accordingly, methods of making genetically engineered fusion molecules and pharmaceutical formulations thereof (e.g., lipid nanoparticle formulations) for in vivo delivery are also provided. As a non-limiting example, the improved magnitude provided by the compositions of the present disclosure can span a critical therapeutic threshold to enable full activation of fetal hemoglobin (HbF) expression, which would provide therapeutic efficacy for functional correction of sickle cell disease or beta-thalassemia.

[0053] Methods for targeted genome editing at selected loci

[0054] Gene editing compositions and methods

[0055] The present disclosure provides a gene editing composition and a cell comprising the same. The gene editing composition can comprise a sequence encoding a DNA localization domain and a sequence encoding a nuclease protein or a nuclease domain thereof. The sequence encoding the nuclease protein or the sequence encoding the nuclease domain thereof can comprise a DNA sequence, an RNA sequence, or a combination thereof. The DNA localization domain can comprise one or more of a CRISPR / Cas protein, a transcription activator-like effector nuclease (TALEN), a zinc finger nuclease (ZFN), and an endonuclease.

[0056] Exemplary dCas9-Clo051 (Cas-CLOVER) fusion proteins

[0057] The nuclease protein or the nuclease domain thereof can comprise a nuclease-inactivated Cas (dCas) protein and an endonuclease. The endonuclease can comprise a Clo051 nuclease or a nuclease domain thereof. The gene editing composition can comprise a fusion protein. The fusion protein can comprise a nuclease-inactivated Cas9 (dCas9) protein and a Clo051 nuclease or a Clo051 nuclease domain. The gene editing composition can further comprise a guide sequence. In some embodiments, the guide sequence comprises an RNA sequence.

[0058] The present disclosure provides a composition comprising a Cas9 operably linked to an effector. The present disclosure provides a fusion protein comprising, consisting essentially of, or consisting of a DNA localization component and an effector molecule, wherein the effector comprises a Cas9. The Cas9 construct of the present disclosure can comprise an effector comprising a type IIS endonuclease. Staphylococcus aureus Cas9 with an active catalytic site comprises the amino acid sequence of SEQ ID NO: 30.

[0059] The present disclosure provides a composition comprising an inactivated Cas9 (dSaCas9) operably linked to an effector. The present disclosure provides a fusion protein comprising, consisting essentially of, or consisting of a DNA localization component and an effector molecule, wherein the effector comprises an inactivated Cas9 (dSaCas9). The inactivated Cas9 (dSaCas9) construct of the present disclosure can comprise an effector comprising a type IIS endonuclease. The dSaCas9 comprises the amino acid sequence of SEQ ID NO: 31, which includes D10A and N580A mutations that inactivate the catalytic site.

[0060] The present disclosure provides compositions comprising an inactivated Cas9 (dCas9) operably linked to an effector. The present disclosure provides a fusion protein comprising, consisting essentially of, or consisting of a DNA localization component and an effector molecule, wherein the effector comprises an inactivated Cas9 (dCas9). The inactivated Cas9 (dCas9) constructs of the present disclosure can comprise an effector comprising a Type IIS endonuclease.

[0061] The dCas9 can be isolated from or derived from Streptoccocus pyogenes. The dCas9 can comprise a dCas9 having substitutions at amino acid positions 10 and 840, which inactivate the catalytic site. In some aspects, the substitutions are D10A and H840A. The dCas9 can comprise the amino acid sequence of SEQ ID NO: 32 or SEQ ID NO: 33.

[0062] An exemplary Clo051 nuclease domain comprises, consists essentially of, or consists of the amino acid sequence of SEQ ID NO: 34. In some aspects, the Clo051 nuclease domain comprises at least one amino acid substitution. In some aspects, the amino acid substitution is in the a-helix loop domain of Clo051 nuclease. In some aspects, the amino acid substitution is at position 35, 37, 60, 92, 98, 100, or 146 of SEQ ID NO: 34. In some aspects, the amino acid substitution is at position 37 of SEQ ID NO: 34. In some aspects, the amino acid substitution is at positions 37 and 92 of SEQ ID NO: 34.

[0063] An exemplary dCas9-Clo051 (Cas-CLOVER) fusion protein can comprise, consist essentially of, or consist of the amino acid sequence of SEQ ID NO: 35. An exemplary dCas9-Clo051 fusion protein can be encoded by a polynucleotide comprising, consisting essentially of, or consisting of the nucleic acid sequence of SEQ ID NO: 36. The nucleic acid encoding the dCas9-Clo051 fusion protein can be DNA or RNA.

[0064] An exemplary dCas9-Clo051 (Cas-CLOVER) fusion protein can comprise, consist essentially of, or consist of the amino acid sequence of SEQ ID NO: 37. An exemplary dCas9-Clo051 fusion protein can be encoded by a polynucleotide comprising, consisting essentially of, or consisting of the nucleic acid sequence of SEQ ID NO: 38. The nucleic acid encoding the dCas9-Clo051 fusion protein can be DNA or RNA.

[0065] An exemplary dCas9-Clo051 fusion protein (Cas-CLOVER) of the present disclosure can further comprise at least one nuclear localization sequence (NLS). In some embodiments, a dCas9-Clo051 fusion protein of the present disclosure comprises at least two nuclear localization sequences. In some embodiments, the NLS is on the N' terminus of the dCas9-Clo051 fusion protein (NLS-dCas9-Clo051). In some embodiments, the NLS is on the C terminus of the dCas9-Clo051 fusion protein (dCas9-Clo051-NLS). In some embodiments, the NLS is on the N' terminus and C' terminus of the dCas9-Clo051 fusion protein ("NLS-dCas9-Clo051-NLS" or "wild type Cas-CLOVER" or "dspCas9 Cas-CLOVER").

[0066] An NLS-dCas9-Clo051-NLS ("wild type Cas-CLOVER" or "dspCas9 Cas-CLOVER") fusion protein can comprise, consist essentially of, or consist of the amino acid sequence of SEQ ID NO: 39.

[0067] The dspCas9 Cas-CLOVER amino acid sequence (NLS amino acid sequence is bold and underlined; linker is bold and italicized)

[0068] MA PKKKRKVPKKKRK V SS (SEQ ID NO: 39).

[0069] A nucleic acid encoding a NLS-dCas9-Clo051-NLS ("wild type Cas-CLOVER" or "dspCas9 Cas-CLOVER") fusion protein can be DNA or RNA. In some embodiments, a dCas9-Clo051 fusion protein comprising two NLS regions is encoded by an mRNA sequence comprising, consisting essentially of, or consisting of SEQ ID NO: 40.

[0070] A NLS-dCas9-Clo051-NLS ("dspCas9-XL Cas-CLOVER" or "dspCas9 5xGGGGS Cas-CLOVER") fusion protein can comprise, consist essentially of, or consist of the amino acid sequence of SEQ ID NO: 41.

[0071] The "dspCas9-XL Cas-CLOVER" amino acid sequence (NLS amino acid sequences are bold and underlined; linkers are bold and italicized)

[0072] MA PKKKRKVPKKKRKV SS (SEQ ID NO: 41)

[0073] NLS-dCas9-Clo051-NLS (“dsaCas9 Cas-CLOVER”) fusion proteins can comprise, consist essentially of, or consist of the amino acid sequence of SEQ ID NO: 42.

[0074] “dsaCas9 Cas-CLOVER” amino acid sequence (NLS amino acid sequence is bold and underlined; linker is bold and italicized)

[0075] MA PKKKRKVPKKKRKV SS (SEQ IDNO: 42)

[0076] NLS-dCas9-Clo051-NLS (“dsaCas9-XL Cas-CLOVER” or “dsaCas9 5xGGGGS Cas-CLOVER”) fusion proteins can comprise, consist essentially of, or consist of the amino acid sequence of SEQ ID NO: 43.

[0077] “dsaCas9-XL Cas-CLOVER” amino acid sequence (NLS amino acid sequence is bold and underlined; linker is bold and italicized)

[0078] MA PKKKRKVPKKKRKV SS (SEQ ID NO: 43)

[0079] A cell comprising a gene editing composition can stably or transiently express the gene editing composition.

[0080] gRNA

[0081] As used herein, the term "guide sequence" or "spacer" in the context of a Cas-Clover system or a CRISPR-Cas9 system comprises any polynucleotide molecule that has sufficient complementarity to a target nucleic acid sequence to hybridize with the target nucleic acid sequence and direct sequence-specific binding of a nucleic acid-targeting complex to the target nucleic acid sequence. A guide sequence can comprise RNA and DNA polynucleotides. A guide sequence can form a duplex with a target sequence. The duplex can be a DNA duplex, an RNA duplex, or an RNA / DNA duplex. The terms "guide molecule," "guide RNA," "gRNA," "single guide RNA," and "sgRNA" are used interchangeably herein to refer to an RNA-based molecule that is capable of forming a complex with a Cas-Clover or CRISPR-Cas protein and comprises a guide sequence that has sufficient complementarity to a target nucleic acid sequence to hybridize with the target nucleic acid sequence and direct sequence-specific binding of the complex to the target nucleic acid sequence. A guide molecule or guide RNA can encompass an RNA-based molecule that has one or more chemical modifications (e.g., by chemically linking two ribonucleotides or by replacing one or more ribonucleotides with one or more deoxyribonucleotides), as described herein. A guide sequence can also comprise, in part, RNA- and DNA-based nucleotides, where the molecule is chimeric for RNA and DNA nucleobases (e.g., contains ribose or deoxyribose sugars).

[0082] The term "target region," "target sequence," or "protospacer" as used interchangeably herein refers to a region of a target gene or genomic target site that is targeted by a Cas-Clover system or a CRISPR / Cas9-based system. A Cas-Clover or CRISPR / Cas9-based system can include at least one gRNA, where the gRNA targets a different DNA sequence. The target DNA sequences can overlap. A Cas-Clover system can include at least two gRNAs, where the gRNAs target different DNA sequences. A target sequence or protospacer can be followed by a PAM sequence at the 3' end of the protospacer. Different Type II CRISPR systems have different PAM requirements. For example, the S. pyogenes Type II system uses a "NGG" sequence, where "N" can be any nucleotide.

[0083] A guide RNA or guide RNA of a Cas-Clover protein or CRISPR-Cas protein can comprise a tracr mate sequence (encompassing a "direct repeat sequence" in the context of an endogenous CRISPR system) and a guide sequence (also referred to as a "spacer" in the context of an endogenous CRISPR system). In some embodiments, a Cas-Clover or CRISPR-Cas system or complex as described herein does not comprise and / or is not dependent on the presence of a tracr sequence. In certain embodiments, a guide molecule can comprise, consist essentially of, or consist of a direct repeat sequence fused or linked to a guide sequence or spacer sequence.

[0084] In some embodiments, a guide RNA comprises a guide sequence and a scaffold sequence. In some embodiments, the scaffold sequence is isolated from S. pyogenes. In some embodiments, the S. pyogenes scaffold sequence comprises the nucleic acid sequence: GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUU (SEQ ID NO: 19). In some embodiments, the scaffold sequence is isolated from S. aureus. In some embodiments, the S. aureus scaffold sequence comprises the nucleic acid sequence:

[0085] GUUUUAGUACUCUGGAAACAGAAUCUACUAAAACAAGGCAAAAUGCCGUGUUUAUCUCGUCAACUUGUUGGCGAGAUUUU (SEQ ID NO: 20).

[0086] In certain embodiments, a guide sequence or spacer of a guide molecule is 15 to 50 nucleotides in length. In certain embodiments, a spacer of a guide RNA is at least 15 nucleotides in length. In certain embodiments, a spacer is 15 to 17 nucleotides in length, 17 to 20 nucleotides in length, 20 to 24 nucleotides in length, 23 to 25 nucleotides in length, 24 to 27 nucleotides in length, 27 to 30 nucleotides in length, 30 to 35 nucleotides in length, or greater than 35 nucleotides in length.

[0087] In some embodiments, the guide sequence is 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, or 150 nucleotides in length.

[0088] In some embodiments, the sequence of the guide molecule (direct repeat sequence and / or spacer) is selected to reduce the extent of secondary structure within the guide molecule. In some embodiments, about 75%, 50%, 40%, 30%, 25%, 20%, 15%, 10%, 5%, 1%, or less of the nucleotides of the nucleic acid-targeting guide RNA are involved in self-complementary base pairing upon optimal folding. Optimal folding can be determined by any suitable polynucleotide folding algorithm. Some programs are based on calculating the minimum Gibbs free energy. An example of such an algorithm is mFold, as described by Zuker and Stiegler (Nucleic Acids Res. 9 (1981), 133-148). Another example of a folding algorithm is the online web server RNAfold developed at the Institute for Theoretical Chemistry of the University of Vienna, which uses a centroid structure prediction algorithm (see, e.g., A.R. Gruber et al., 2008, Cell 106(1): 23-24; and PA Carr and GM Church, 2009, Nature Biotechnology 27(12): 1151-62).

[0089] As described above, the Cas-Clover system and the CRISPR / Cas9 system utilize one or more targeting gRNAs that provide targeting for the Cas-Clover system and the CRISPR / Cas9-based system. The gRNA can be a fusion of two non-coding RNAs: a crRNA and a tracrRNA. The sgRNA can target any desired DNA sequence by swapping the sequence encoding the 20 bp protospacer sequence, which confers targeting specificity by base pairing with the complementary bases of the desired DNA target. The gRNA mimics the naturally occurring crRNA:tracrRNA duplex involved in Type II effector systems. The duplex can include, for example, a 42 nucleotide crRNA and a 75 nucleotide tracrRNA, which functions as a guide for Cas9 to cleave the target nucleic acid.

[0090] In some embodiments, the gRNA targets a BCL11 A binding site in the HBG promoter region upstream of the target gene (e.g., the HBG1 or HBG2 locus), e.g., between 0 and 1000 bp upstream of the target gene. In some embodiments, the gRNA targets a BCL11 A binding site in the HBG promoter region upstream of the HBG1 locus or the HBG2 locus. In some embodiments, the gRNA targets a region between 0 and 50 bp, 0 and 100 bp, 0 and 150 bp, 0 and 200 bp, or 0 and 250 bp upstream or downstream of the BCL11 A binding site. In some embodiments, the gRNA targets a region between 0 and 50 bp, 0 and 100 bp, 0 and 150 bp, 0 and 200 bp, 0 and 250 bp, 0 and 300 bp, 0 and 350 bp, 0 and 400 bp, 0 and 450 bp, 0 and 500 bp, 0 and 550 bp, 0 and 600 bp, 0 and 650 bp, 0 and 700 bp, 0 and 750 bp, 0 and 800 bp, 0 and 850 bp, 0 and 900 bp, 0 and 950 bp, or 0 and 1000 bp upstream of the transcription start site of the target gene. In some embodiments, the gRNA targets a region within about 100 bp, about 200 bp, about 300 bp, about 400 bp, about 500 bp, about 600 bp, about 700 bp, about 800 bp, about 900 bp, about 1000 bp, about 1100 bp, about 1200 bp, about 1300 bp, about 1400 bp, or about 1500 bp upstream of the target gene.

[0091] A gRNA can be divided into a target-binding region and a Cas9-binding region. The target-binding region hybridizes to a target region in a target gene or intergenic region. Methods for designing such target-binding regions are known in the art, see, e.g., Doench et al., Nat Biotechnol. (2014) 32: 1262-7; and Doench et al., Nat Biotechnol. (2016) 34: 184-91, which are incorporated by reference herein in their entireties. Design tools are available, e.g., at target Finder by the Feng Zhang Lab, Target Finder (E-CRISP) by the Michael Boutros Lab, RGEN Tools (Cas-OF Finder), CasFinder, and CRISPR Optimal Target Finder. In certain embodiments, the target-binding region can be between about 15 and about 50 nucleotides in length (about 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or about 50 nucleotides in length). In certain embodiments, the target-binding region can be between about 19 and about 21 nucleotides in length. In one embodiment, the target-binding region is 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 nucleotides in length.

[0092] In one embodiment, the target-binding region is complementary, e.g., fully complementary, to a target region in a target gene. In one embodiment, the target-binding region is substantially complementary to a target region in a target gene. In one embodiment, the target-binding region comprises no more than 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides that are not complementary to a target region in a target gene.

[0093] Exemplary gRNAs of the present disclosure include, but are not limited to, sequences for targeting the HBG1 locus, the HBG2 locus, or the BCL11A binding site in the HBG promoter region of the HBG1 or HBG2 locus.

[0094] In some embodiments, a first gRNA (also referred to as a “left gRNA”) binds to a template sequence at the 5’ end of a target locus, and a second gRNA (also referred to as a “right gRNA”) binds to a template sequence at the 3’ end of the target locus. A schematic is shown in FIG. 7.

[0095] Exemplary gRNAs of the disclosure comprise, consist essentially of, or consist of the target sequences and full-length sequences shown in Table 1.

[0096] Table 1. Exemplary gRNAs of the disclosure

[0097]

[0098] gRNA modifications

[0099] The activity, stability, or other characteristics of a gRNA can be altered by the incorporation of certain modifications. As one example, transiently expressed or delivered nucleic acids can be susceptible to degradation by, for example, cellular nucleases. Thus, the gRNAs described herein can contain one or more modified nucleosides or nucleotides introduced to confer stability against nucleases. While not wishing to be bound by theory, it is also believed that certain modified gRNAs described herein can exhibit reduced innate immune responses when introduced into a cell. Those skilled in the art will be aware of certain cellular responses that are commonly observed in cells (e.g., mammalian cells) in response to exogenous nucleic acids, particularly nucleic acids of viral or bacterial origin. Such reactions, which can include induction of cytokine expression and release as well as cell death, can be reduced or completely eliminated by the modifications presented herein.

[0100] Certain exemplary modifications discussed in this section can be included at any position within the gRNA sequence, including but not limited to at or near the 5' end (e.g., within 1-10, 1-5, or 1-2 nucleotides of the 5' end) and / or at or near the 3' end (e.g., within 1-10, 1-5, or 1-2 nucleotides of the 3' end). In some cases, the modification is positioned within a functional motif, such as the repeat-anti-repeat duplex of a Cas9 gRNA, the stem-loop structure of a Cas9 or Cpf1 gRNA, and / or the targeting domain of a gRNA.

[0101] As one example, the 5' end of a gRNA can include a eukaryotic mRNA cap structure or cap analog (e.g., a G(5)ppp(5)G cap analog, a m7G(5)ppp(5)G cap analog, or a 3'-O-Me-m7G(5)ppp(5)G anti-reverse cap analog (ARCA)), as shown below:

[0102]

[0103] The cap or cap analog can be included during chemical synthesis or in vitro transcription of the gRNA.

[0104] Similarly, the 5' end of a gRNA can lack a 5' triphosphate group. For example, in vitro transcribed gRNAs can be phosphatase treated (e.g., using calf intestinal alkaline phosphatase) to remove the 5' triphosphate group.

[0105] Another common modification involves the addition of a plurality (e.g., 1-10, 10-20, or 25-200) of adenine (A) residues, called a polyA bundle, to the 3' end of a gRNA. The polyA bundle can be added to a gRNA in vitro during chemical synthesis, after in vitro transcription using a polyadenyl polymerase (e.g., E. coli Poly(A) polymerase), or in vivo with the aid of a polyadenylation sequence, as described in Maede.

[0106] It should be noted that the modifications described herein can be combined in any suitable manner, e.g., a gRNA, whether transcribed in vivo from a DNA vector or in vitro transcribed, can include either or both of a 5' cap structure or cap analog and a 3' polyA bundle.

[0107] A guide RNA can be modified at the 3' terminal U ribose. For example, the two terminal hydroxyl groups of the U ribose can be oxidized to aldehyde groups with opening of the ribose ring to give a modified nucleoside as shown below:

[0108]

[0109] wherein "U" can be unmodified or modified uridine.

[0110] A 3' terminal U ribose can be modified with a 2'3' cyclic phosphate as shown below:

[0111]

[0112] wherein "U" can be unmodified or modified uridine.

[0113] A guide RNA can contain 3' nucleotides that can be stabilized, e.g., by incorporation of one or more modified nucleotides described herein, to prevent degradation. In certain embodiments, uridines can be replaced with modified uridines (e.g., 5-(2-amino)propyl uridine and 5-bromo uridine) or any modified uridine described herein; adenosines and guanosines can be replaced with modified adenosines and guanosines, e.g., modified at position 8, e.g., 8-bromoguanosine, or any modified adenosine or guanosine described herein.

[0114] In certain embodiments, sugar modified ribonucleotides can be incorporated into a gRNA, e.g., where the 2' OH-group is replaced with a group selected from H,—OR,—R (where R can be, e.g., alkyl, cycloalkyl, aryl, aralkyl, heteroaryl, or sugar), halo,—SH,—SR (where R can be, e.g., alkyl, cycloalkyl, aryl, aralkyl, heteroaryl, or sugar), amino (where amino can be, e.g., NH2; alkylamino, dialkylamino, heterocyclyl, arylamino, diarylamino, heteroarylamino, diheteroarylamino, or an amino acid), or cyano (—CN). In certain embodiments, the phosphate backbone can be modified as described herein, e.g., with a phosphorothioate (PhTx) group. In certain embodiments, one or more nucleotides of a gRNA can each independently be a modified or unmodified nucleotide, including but not limited to 2'-sugar modified nucleotides such as 2'-O-methyl, 2'-O-methoxyethyl, or 2'-fluoro modified nucleotides including, e.g., 2'-F or 2'-O-methyl, adenosine (A), 2'-F or 2'-O-methyl, cytidine (C), 2'-F or 2'-O-methyl, uridine (U), 2'-F or 2'-O-methyl, thymidine (T), 2'-F or 2'-O-methyl, guanosine (G), 2'-O-methoxyethyl-5-methyluridine (Teo), 2'-O-methoxyethyladenosine (Aeo), 2'-O-methoxyethyl-5-methylcytidine (m5Ceo), and any combination thereof.

[0115] A guide RNA can also include a“locked” nucleic acid (LNA), where the 2' OH group can be linked to the 4' carbon of the same ribose sugar, e.g., by a C1-6 alkylene or C1-6 heteroalkylene bridge. Any suitable moiety can be used to provide such a bridge, including but not limited to a methylene, propylene, ether, or amino bridge; O-amino (where amino can be, e.g., NH2; alkylamino, dialkylamino, heterocyclyl, arylamino, diarylamino, heteroarylamino, or diheteroarylamino, ethylenediamine, or polyamino) and aminoalkoxy or O(CH2) n amino (where amino can be, e.g., NH2; alkylamino, dialkylamino, heterocyclyl, arylamino, diarylamino, heteroarylamino, or diheteroarylamino, ethylenediamine, or polyamino).

[0116] In certain embodiments, a gRNA can comprise a polycycle (e.g., a tricycle; and “unlocked” forms, such as glycol nucleic acids (GNAs) (e.g., R-GNAs or S-GNAs, in which ribose is replaced by glycol units attached to phosphodiester bonds) or threose nucleic acids (TNAs, in which ribose is replaced by a-L-threofuranosyl-(3’→2’ modified nucleotides.

[0117] Generally, gRNAs include a sugar group, ribose, which is a 5-membered ring with an oxygen. Exemplary modified gRNAs can include, but are not limited to, replacement of the oxygen in ribose (e.g., with sulfur (S), selenium (Se), or an alkylene such as, for example, methylene or ethylene); addition of a double bond (e.g., replacement of ribose with a cyclopentenyl or cyclohexenyl); ring contraction of ribose (e.g., forming a 4-membered ring of cyclobutane or oxetane); ring expansion of ribose (e.g., forming a 6- or 7-membered ring with an additional carbon or heteroatom such as, for example, an anhydrohexitol, altritol, mannitol, cyclohexyl, cyclohexenyl, and morpholino also with phosphoramidate backbones). While most sugar analogs alter the position located at the 2’ position, other positions are also suitable for modification, including the 4’ position. In certain embodiments, a gRNA comprises a 4’-S, 4’-Se, or 4’-C-aminomethyl-2’-O-Me modification.

[0118] In certain embodiments, a deazanucleotide, e.g., 7-deaza-adenosine, can be incorporated into a gRNA. In certain embodiments, an O- and N-alkylated nucleotide, e.g., N6-methyladenosine, can be incorporated into a gRNA. In certain embodiments, one or more or all of the nucleotides in a gRNA are deoxy nucleotides.

[0119] In some embodiments, a gRNA comprises one or more chemical modifications of a ribonucleotide, a ribonucleotide base, or a phosphodiester linkage. In some embodiments, the one or more chemical modifications comprises at least one chemically modified phosphodiester linkage. In some embodiments, the at least one chemically modified phosphodiester linkage is a phosphorothioate linkage.

[0120] In some embodiments, a gRNA comprises three phosphorothioate linkages at the 5’ end of the gRNA. In some embodiments, a gRNA comprises two phosphorothioate linkages at the 3’ end of the gRNA. In some embodiments, a gRNA comprises a 2’ O-Me chemical modification at the 3’ end of the gRNA.

[0121] Exemplary gRNA targeting sequences

[0122] In some embodiments, the gRNA comprises a targeting sequence comprising a nucleotide sequence that is at least 95%, 96%, 97%, 98%, or 99% (or any percentage in between) identical to SEQ ID NO: 1. In some embodiments, the gRNA comprises a targeting sequence comprising the nucleotide sequence of SEQ ID NO: 1.

[0123] In some embodiments, the gRNA comprises a targeting sequence comprising a nucleotide sequence that is at least 95%, 96%, 97%, 98%, or 99% (or any percentage in between) identical to SEQ ID NO: 3. In some embodiments, the gRNA comprises a targeting sequence comprising the nucleotide sequence of SEQ ID NO: 3.

[0124] In some embodiments, the gRNA comprises a targeting sequence comprising a nucleotide sequence that is at least 95%, 96%, 97%, 98%, or 99% (or any percentage in between) identical to SEQ ID NO: 5. In some embodiments, the gRNA comprises a targeting sequence comprising the nucleotide sequence of SEQ ID NO: 5.

[0125] In some embodiments, the gRNA comprises a targeting sequence comprising a nucleotide sequence that is at least 95%, 96%, 97%, 98%, or 99% (or any percentage in between) identical to SEQ ID NO: 7. In some embodiments, the gRNA comprises a targeting sequence comprising the nucleotide sequence of SEQ ID NO: 7.

[0126] In some embodiments, the gRNA comprises a targeting sequence comprising a nucleotide sequence that is at least 95%, 96%, 97%, 98%, or 99% (or any percentage in between) identical to SEQ ID NO: 9. In some embodiments, the gRNA comprises a targeting sequence comprising the nucleotide sequence of SEQ ID NO: 9.

[0127] In some embodiments, the gRNA comprises a targeting sequence comprising a nucleotide sequence that is at least 95%, 96%, 97%, 98%, or 99% (or any percentage in between) identical to SEQ ID NO: 11. In some embodiments, the gRNA comprises a targeting sequence comprising the nucleotide sequence of SEQ ID NO: 11.

[0128] In some embodiments, the gRNA comprises a targeting sequence comprising a nucleotide sequence that is at least 95%, 96%, 97%, 98%, or 99% (or any percentage in between) identical to SEQ ID NO: 13. In some embodiments, the gRNA comprises a targeting sequence comprising the nucleotide sequence of SEQ ID NO: 13.

[0129] In some embodiments, the gRNA comprises a targeting sequence comprising a nucleotide sequence that is at least 95%, 96%, 97%, 98%, or 99% (or any percentage in between) identical to SEQ ID NO: 15. In some embodiments, the gRNA comprises a targeting sequence comprising the nucleotide sequence of SEQ ID NO: 15.

[0130] In some embodiments, the gRNA comprises a targeting sequence comprising a nucleotide sequence that is at least 95%, 96%, 97%, 98%, or 99% (or any percentage in between) identical to SEQ ID NO: 17. In some embodiments, the gRNA comprises a targeting sequence comprising the nucleotide sequence of SEQ ID NO: 17.

[0131] Exemplary gRNA sequences

[0132] In some embodiments, the gRNA comprises a nucleotide sequence that is at least 95%, 96%, 97%, 98%, or 99% (or any percentage in between) identical to SEQ ID NO: 2. In some embodiments, the gRNA comprises the nucleotide sequence of SEQ ID NO:2.

[0133] In some embodiments, the gRNA comprises a nucleotide sequence that is at least 95%, 96%, 97%, 98%, or 99% (or any percentage in between) identical to SEQ ID NO: 4. In some embodiments, the gRNA comprises the nucleotide sequence of SEQ ID NO:4.

[0134] In some embodiments, the gRNA comprises a nucleotide sequence that is at least 95%, 96%, 97%, 98%, or 99% (or any percentage in between) identical to SEQ ID NO: 6. In some embodiments, the gRNA comprises the nucleotide sequence of SEQ ID NO:6.

[0135] In some embodiments, the gRNA comprises a nucleotide sequence that is at least 95%, 96%, 97%, 98%, or 99% (or any percentage in between) identical to SEQ ID NO: 8. In some embodiments, the gRNA comprises the nucleotide sequence of SEQ ID NO: 8.

[0136] In some embodiments, the gRNA comprises a nucleotide sequence that is at least 95%, 96%, 97%, 98%, or 99% (or any percentage in between) identical to SEQ ID NO: 10. In some embodiments, the gRNA comprises the nucleotide sequence of SEQ ID NO: 10.

[0137] In some embodiments, the gRNA comprises a nucleotide sequence that is at least 95%, 96%, 97%, 98%, or 99% (or any percentage in between) identical to SEQ ID NO: 12. In some embodiments, the gRNA comprises the nucleotide sequence of SEQ ID NO: 12.

[0138] In some embodiments, the gRNA comprises a nucleotide sequence that is at least 95%, 96%, 97%, 98%, or 99% (or any percentage in between) identical to SEQ ID NO: 14. In some embodiments, the gRNA comprises the nucleotide sequence of SEQ ID NO: 14.

[0139] In some embodiments, the gRNA comprises a nucleotide sequence that is at least 95%, 96%, 97%, 98%, or 99% (or any percentage in between) identical to SEQ ID NO: 16. In some embodiments, the gRNA comprises the nucleotide sequence of SEQ ID NO: 16.

[0140] In some embodiments, the gRNA comprises a nucleotide sequence that is at least 95%, 96%, 97%, 98%, or 99% (or any percentage in between) identical to SEQ ID NO: 19. In some embodiments, the gRNA comprises the nucleotide sequence of SEQ ID NO: 19.

[0141] Exemplary gRNA compositions

[0142] In certain compositions of the disclosure, the composition comprises a first gRNA (also referred to as a “left gRNA”) and a second gRNA (also referred to as a “right gRNA”). The first gRNA can comprise a first targeting sequence. The second gRNA can comprise a second targeting sequence.

[0143] In some embodiments, the composition comprises: a first gRNA comprising a first targeting sequence comprising the nucleotide sequence of SEQ ID NO: 1 and a second gRNA comprising a second targeting sequence comprising the nucleotide sequence of SEQ ID NO: 3. In some embodiments, the composition comprises: a first gRNA comprising the nucleotide sequence of SEQ ID NO: 2 and a second gRNA comprising the nucleotide sequence of SEQ ID NO: 4.

[0144] In some embodiments, the composition comprises: a first gRNA comprising a first targeting sequence comprising the nucleotide sequence of SEQ ID NO: 5 and a second gRNA comprising a second targeting sequence comprising the nucleotide sequence of SEQ ID NO: 3. In some embodiments, the composition comprises: a first gRNA comprising the nucleotide sequence of SEQ ID NO: 6 and a second gRNA comprising the nucleotide sequence of SEQ ID NO: 4.

[0145] In some embodiments, the composition comprises: a first gRNA comprising a first targeting sequence comprising the nucleotide sequence of SEQ ID NO: 7 and a second gRNA comprising a second targeting sequence comprising the nucleotide sequence of SEQ ID NO: 3. In some embodiments, the composition comprises: a first gRNA comprising the nucleotide sequence of SEQ ID NO: 8 and a second gRNA comprising the nucleotide sequence of SEQ ID NO: 4.

[0146] In some embodiments, the composition comprises: a first gRNA comprising a first targeting sequence comprising the nucleotide sequence of SEQ ID NO: 9 and a second gRNA comprising a second targeting sequence comprising the nucleotide sequence of SEQ ID NO: 11. In some embodiments, the composition comprises: a first gRNA comprising the nucleotide sequence of SEQ ID NO: 10 and a second gRNA comprising the nucleotide sequence of SEQ ID NO: 12.

[0147] In some embodiments, the composition comprises: a first gRNA comprising a first targeting sequence comprising the nucleotide sequence of SEQ ID NO: 13 and a second gRNA comprising a second targeting sequence comprising the nucleotide sequence of SEQ ID NO: 11. In some embodiments, the composition comprises: a first gRNA comprising the nucleotide sequence of SEQ ID NO: 14 and a second gRNA comprising the nucleotide sequence of SEQ ID NO: 12.

[0148] In some embodiments, the composition comprises: a first gRNA comprising a first targeting sequence comprising the nucleotide sequence of SEQ ID NO: 7 and a second gRNA comprising a second targeting sequence comprising the nucleotide sequence of SEQ ID NO: 11. In some embodiments, the composition comprises: a first gRNA comprising the nucleotide sequence of SEQ ID NO: 19 and a second gRNA comprising the nucleotide sequence of SEQ ID NO: 12.

[0149] In some embodiments, the composition comprises: a first gRNA comprising a first targeting sequence comprising the nucleotide sequence of SEQ ID NO: 15 and a second gRNA comprising a second targeting sequence comprising the nucleotide sequence of SEQ ID NO: 3. In some embodiments, the composition comprises: a first gRNA comprising the nucleotide sequence of SEQ ID NO: 16 and a second gRNA comprising the nucleotide sequence of SEQ ID NO: 4.

[0150] In some embodiments, the composition comprises: a first gRNA comprising a first targeting sequence comprising the nucleotide sequence of SEQ ID NO: 17 and a second gRNA comprising a second targeting sequence comprising the nucleotide sequence of SEQ ID NO: 3. In some embodiments, the composition comprises: a first gRNA comprising the nucleotide sequence of SEQ ID NO: 18 and a second gRNA comprising the nucleotide sequence of SEQ ID NO: 4.

[0151] Exemplary Cas-CLOVER and gRNA compositions

[0152] Gene editing compositions comprising Cas-CLOVER and methods of gene editing using these compositions are described in detail in PCT Application Nos. PCT / US2016 / 037922, PCT / US2018 / 066941, PCT / US2017 / 054799, U.S. Patent Publication Nos. 2017 / 0107541, 2017 / 0114149, 2018 / 0187185, and U.S. Patent No. 10,415,024, each of which is incorporated by reference in its entirety herein, e.g., gene editing compositions useful in the methods disclosed herein. Exemplary gene editing compositions comprising Cas-CLOVER and methods of gene editing using these compositions are described herein.

[0153] In certain compositions of the disclosure, the composition comprises a first gRNA (also referred to as a“left gRNA”), a first fusion protein or a first polynucleotide encoding a first fusion protein (e.g., Cas-Clover), a second gRNA (also referred to as a“right gRNA”), and a second fusion protein or a second polynucleotide encoding a second fusion protein (e.g., Cas-Clover). The first gRNA comprises a first targeting sequence. The second gRNA comprises a second targeting sequence.

[0154] In some embodiments, the first gRNA and the first fusion protein complex at the 5’ end of the target DNA to be modified. In some embodiments, the second gRNA and the second fusion protein complex at the 3’ end of the target DNA to be modified. A schematic of the composition complexed with a target DNA is shown in FIG. 7.

[0155] In some embodiments, the first fusion protein, the second fusion protein, or both the first fusion protein and the second fusion protein comprise a dCas9 derived from a Streptococcus pyogenes Cas9 polypeptide. In some embodiments, the first fusion protein, the second fusion protein, or both the first fusion protein and the second fusion protein comprise a dCas9 derived from a Staphylococcus aureus Cas9 polypeptide. Exemplary compositions of the disclosure are shown in Table 2.

[0156] Table 2. Exemplary compositions of the disclosure

[0157]

[0158] In some embodiments, the composition comprises a first gRNA comprising the nucleotide sequence of SEQ ID NO: 2, a first fusion protein comprising the polypeptide sequence of SEQ ID NO: 39, a second gRNA comprising the nucleotide sequence of SEQ ID NO: 4, and a second fusion protein comprising the polypeptide sequence of SEQ ID NO: 39.

[0159] In some embodiments, the composition comprises a first gRNA comprising the nucleotide sequence of SEQ ID NO: 6, a first fusion protein comprising the polypeptide sequence of SEQ ID NO: 39, a second gRNA comprising the nucleotide sequence of SEQ ID NO: 4, and a second fusion protein comprising the polypeptide sequence of SEQ ID NO: 39.

[0160] In some embodiments, the composition comprises: a first gRNA comprising the nucleotide sequence of SEQ ID NO: 8, a first fusion protein comprising the polypeptide sequence of SEQ ID NO: 41, a second gRNA comprising the nucleotide sequence of SEQ ID NO: 4, and a second fusion protein comprising the polypeptide sequence of SEQ ID NO: 41.

[0161] In some embodiments, the composition comprises: a first gRNA comprising the nucleotide sequence of SEQ ID NO: 10, a first fusion protein comprising the polypeptide sequence of SEQ ID NO: 39, a second gRNA comprising the nucleotide sequence of SEQ ID NO: 12, and a second fusion protein comprising the polypeptide sequence of SEQ ID NO: 42.

[0162] In some embodiments, the composition comprises: a first gRNA comprising the nucleotide sequence of SEQ ID NO: 14, a first fusion protein comprising the polypeptide sequence of SEQ ID NO: 42, a second gRNA comprising the nucleotide sequence of SEQ ID NO: 12, and a second fusion protein comprising the polypeptide sequence of SEQ ID NO: 42.

[0163] In some embodiments, the composition comprises: a first gRNA comprising the nucleotide sequence of SEQ ID NO: 19, a first fusion protein comprising the polypeptide sequence of SEQ ID NO: 42, a second gRNA comprising the nucleotide sequence of SEQ ID NO: 12, and a second fusion protein comprising the polypeptide sequence of SEQ ID NO: 42.

[0164] In some embodiments, the composition comprises: a first gRNA comprising the nucleotide sequence of SEQ ID NO: 16, a first fusion protein comprising the polypeptide sequence of SEQ ID NO: 41, a second gRNA comprising the nucleotide sequence of SEQ ID NO: 4, and a second fusion protein comprising the polypeptide sequence of SEQ ID NO: 41.

[0165] In some embodiments, the composition comprises: a first gRNA comprising the nucleotide sequence of SEQ ID NO: 18, a first fusion protein comprising the polypeptide sequence of SEQ ID NO: 41, a second gRNA comprising the nucleotide sequence of SEQ ID NO: 4, and a second fusion protein comprising the polypeptide sequence of SEQ ID NO: 41.

[0166] Delivery of gRNA and gene editing compositions

[0167] Gene editing tools can also be delivered to cells using one or more poly(histidine)-based micelles. Poly(histidine) (e.g., poly(L-histidine)) is a pH-sensitive polymer due to the imidazole ring providing a lone pair of electrons on the unsaturated nitrogen. That is, poly(histidine) has amphoteric properties through protonation-deprotonation. In particular, at certain pHs, a triblock copolymer containing poly(histidine) can assemble into micelles with positively charged poly(histidine) units on the surface, thereby enabling complexation with negatively charged gene editing molecules. The use of these nanoparticles to bind and release proteins and / or nucleic acids in a pH-dependent manner can provide an efficient and selective mechanism to make desired genetic modifications. In particular, this micelle-based delivery system provides substantial flexibility with respect to charged materials, as well as large payload capacity and targeted release of nanoparticle payloads. In one example, site-specific cleavage of double-stranded DNA is achieved by delivering a nuclease using poly(histidine)-based micelles. Without wishing to be bound by a particular theory, it is believed that in micelles formed from various triblock copolymers, the hydrophobic block aggregates to form a core, leaving the hydrophilic block and the poly(histidine) block at the ends to form one or more surrounding layers.

[0168] In one aspect, the present disclosure provides a triblock copolymer made from a hydrophilic block, a hydrophobic block, and a charged block. In some aspects, the hydrophilic block can be poly(ethylene oxide) (PEO), and the charged block can be poly(L-histidine). An exemplary triblock copolymer that can be used is PEO-b-PLA-b-PHIS, where the variable number of repeat units in each block varies by design.

[0169] A diblock copolymer that can be used as an intermediate to make the triblock copolymer can have a hydrophilic biocompatible poly(ethylene oxide) (PEO), which is chemically synonymous with PEG, coupled to various hydrophobic aliphatic poly(anhydrides), poly(nucleic acids), poly(esters), poly(orthoesters), poly(peptides), poly(phosphazenes), and poly(saccharides), including but not limited to poly(lactide) (PLA), poly(glycolide) (PLGA), poly(lactic-co-glycolic acid) (PLGA), poly(ε-caprolactone) (PCL), and poly(trimethylene carbonate) (PTMC). Polymer micelles composed of 100% polyethyleneglycolated surfaces have improved in vitro chemical stability, enhanced in vivo bioavailability, and prolonged blood circulation half-life.

[0170] Polymeric vesicles, polymeric bodies, and poly(histidine)-based micelles, including those comprising triblock copolymers, and methods of making them, are described in further detail in U.S. Patent Nos. 7,217,427; 7,868,512; 6,835,394; 8,808,748; 10,456,452; U.S. Publication Nos. 2014 / 0363496; 2017 / 0000743; and 2019 / 0255191; and PCT Publication No. WO 2019 / 126589, each of which is incorporated herein by reference in its entirety, e.g., polymers useful for delivering the compositions disclosed herein.

[0171] Gene editing compositions (e.g., mutant Cas-Clover) can also be delivered to cells using one or more lipid nanoparticle compositions and methods of making them, as described in PCT Publication Nos. WO 2022 / 182792 and WO 2023 / 141576, each of which is incorporated herein by reference in its entirety, e.g., lipid nanoparticles useful for delivering the gene editing compositions disclosed herein.

[0172] In some aspects, the composition is encapsulated in at least one lipid nanoparticle comprising: about 40.75% by molar of a terpene-based lipid compound, about 51.75% by molar of cholesterol, about 5% by molar of DOPC, and about 2.5% by molar of DMG-PEG2000, wherein the polynucleotide encoding the mutant Cas-Clover is an RNA molecule, and wherein the ratio of lipids to RNA molecules in the at least one nanoparticle is about 120:1 (w / w).

[0173] In some aspects, the terpene-based lipid compound is HMA-404:

[0174]

[0175] Thus, in some aspects, the gene editing composition is encapsulated in at least one lipid nanoparticle comprising: about 40.75% by molar of HMA-404, about 51.75% by molar of cholesterol, about 5% by molar of DOPC, and about 2.5% by molar of DMG-PEG2000, wherein the polynucleotide encoding the mutant Cas-Clover is an RNA molecule, and wherein the ratio of lipids to RNA molecules in the at least one nanoparticle is about 120:1 (w / w).

[0176] In some aspects, the composition is encapsulated in at least one lipid nanoparticle comprising: about 54% by mole SS-OP, about 35% by mole cholesterol, about 5% by mole DOPC, about 5% by mole DSPC, and about 1% by mole DMG-PEG2000. The ratio of lipid to nucleic acid in the nanoparticle is about 100: 1 (weight / weight), and the total lipid is 25 mM.

[0177] Cells and modified cells of the disclosure

[0178] Cells and modified cells of the disclosure can be mammalian cells. The cells and modified cells are human cells. In some embodiments, the cells can include hematopoietic progenitor cells (HPCs). In some embodiments, the cells can include hematopoietic stem cells (HSCs). In some embodiments, the cells are hematopoietic stem and progenitor cells (HSPCs). In some embodiments, the HSPCs are capable of differentiating into erythroid progenitor cells. In certain embodiments, at least a portion of the plurality of cells can belong to the erythroid lineage. In some embodiments, the HSPCs are capable of differentiating into erythroid progenitor cells.

[0179] Cells that have been altered ex vivo according to the disclosure can be manipulated (e.g., expanded, passaged, frozen, differentiated, dedifferentiated, transduced with a transgene, etc.) prior to their delivery to a subject. The cells are delivered to the subject from which they were obtained in various ways (“autologous” transplantation), or to a recipient that is immunologically different from the cell donor (“allogeneic” transplantation).

[0180] In some embodiments, autologous transplantation includes the steps of obtaining a plurality of cells from a subject that are circulating in the peripheral blood or in bone marrow or other tissues (e.g., spleen, skin, etc.), and manipulating these cells to enrich for cells of the erythroid lineage (e.g., by inducing production of iPSCs, purifying cells that express certain cell surface markers such as CD34, CD90, CD49f, and / or do not express surface markers characteristic of non-erythroid lineages such as CD10, CD14, CD38, etc.). Optionally or additionally, the cells are expanded, the cells are transduced with a transgene, the cells are exposed to a cytokine or other peptide or small molecule agent, and / or the cells are frozen / thawed prior to transduction with a genome editing system targeting a BCL11A gene or a HBG1 and / or HBG2 promoter target sequence. The genome editing system can be implemented or delivered to the cells in any suitable format, including as a ribonucleoprotein complex, as separate protein and nucleic acid components, and / or as nucleic acids encoding components of the genome editing system.

[0181] Following delivery of the genome editing system, the cells are optionally manipulated, e.g., enriched for HSCs and / or cells in the erythroid lineage and / or the cells are edited, to expand these cells, frozen / thawed, or otherwise prepared for return to the subject. The edited cells are then returned to the subject, e.g., in the circulatory system by way of intravenous delivery or delivery or into solid tissue, such as bone marrow.

[0182] Modified cells of the disclosure

[0183] The present disclosure provides a method of modifying a population of cells, the method comprising contacting the population of cells with a composition of the present disclosure (e.g., a first Cas-Clover fusion protein and a first gRNA, and a second Cas-Clover fusion protein and a second gRNA composition), wherein the first gRNA forms a complex with a first targeting sequence and the first fusion protein, and the second gRNA forms a complex with a second targeting sequence and the second fusion protein, thereby generating an insertion or a deletion (indel) between the first targeting sequence and the second targeting sequence and generating a modified population of cells. In some embodiments, the indel is generated at the BCL11A gene, the HMG1 promoter region, the HMG2 promoter region, or a combination thereof. In some embodiments, the indel results in inactivation of the BCL11A gene.

[0184] In some embodiments, the present disclosure relates to a composition comprising a plurality of cells produced by the methods disclosed above, wherein at least 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% of the cells comprise an indel between the first targeting sequence and the second targeting sequence. In some embodiments, the present disclosure relates to a composition comprising a plurality of cells produced by the methods disclosed above, wherein at least 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% of the cells comprise an indel at the BCL11A gene. In some embodiments, the present disclosure relates to a composition comprising a plurality of cells produced by the methods disclosed above, wherein at least 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% of the cells comprise an indel at the HMG1 promoter region. In some embodiments, the present disclosure relates to a composition comprising a plurality of cells produced by the methods disclosed above, wherein at least 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% of the cells comprise an indel at the HMG2 promoter region.

[0185] In certain embodiments, the plurality of cells can be characterized by an increased level of fetal hemoglobin (HbF) expression relative to the unmodified plurality of cells. In certain embodiments, the level of fetal hemoglobin can be increased by at least 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90%. In certain embodiments, the level of fetal hemoglobin can be increased by about 1-fold to about 10-fold. In certain embodiments, the level of fetal hemoglobin can be increased by about 4-fold to about 9-fold. In certain embodiments, the level of fetal hemoglobin can be increased by about 1-fold to about 2-fold, about 2-fold to about 3-fold, about 3-fold to about 4-fold, about 4-fold to about 5-fold, about 5-fold to about 6-fold, about 6-fold to about 7-fold, about 7-fold to about 8-fold, about 8-fold to about 9-fold, about 9-fold to about 10-fold. In certain embodiments, the level of fetal hemoglobin can be increased by about 1-fold, about 2-fold, about 3-fold, about 4-fold, about 5-fold, about 6-fold, about 7-fold, about 8-fold, about 9-fold, or about 10-fold.

[0186] In certain embodiments, the plurality of cells can be characterized by an increased level of gamma globulin expression relative to the unmodified plurality of cells. In certain embodiments, the level of gamma globulin can be increased by at least 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90%. In certain embodiments, the level of gamma globulin can be increased by about 1-fold to about 10-fold. In certain embodiments, the level of gamma globulin can be increased by about 4-fold to about 9-fold. In certain embodiments, the level of gamma globulin can be increased by about 1-fold to about 2-fold, about 2-fold to about 3-fold, about 3-fold to about 4-fold, about 4-fold to about 5-fold, about 5-fold to about 6-fold, about 6-fold to about 7-fold, about 7-fold to about 8-fold, about 8-fold to about 9-fold, about 9-fold to about 10-fold. In certain embodiments, the level of gamma globulin can be increased by about 1-fold, about 2-fold, about 3-fold, about 4-fold, about 5-fold, about 6-fold, about 7-fold, about 8-fold, about 9-fold, or about 10-fold.

[0187] The cells and modified immune cells of the present disclosure can be autologous cells or allogeneic cells. Allogeneic cells are engineered to prevent adverse reactions to the transplant upon administration to a subject. The allogeneic cells can be any type of cell. The allogeneic cells can be stem cells, or can be derived from stem cells. The allogeneic cells can be differentiated somatic cells.

[0188] In certain aspects, the cells of the present disclosure are modified to recombinantly express dihydrofolate reductase (DHFR), which advantageously renders the cells resistant to methotrexate (MTX). The MTX-resistant cells can be used in methods of treating a subject in need thereof, in combination with a subsequent administration of MTX, to deplete activated T cells and NK cells that target the modified cells or therapeutic cells, thereby increasing the in vivo persistence and efficacy of the modified cells.

[0189] Formulations, dosages, and modes of administration

[0190] Genome editing systems, or cells altered or manipulated using such systems, can be administered to a subject by any suitable mode or route, whether local or systemic. Systemic modes of administration include oral and parenteral routes. For example, parenteral routes include intravenous, intraosseous, intraarterial, intramuscular, intradermal, subcutaneous, intranasal, and intraperitoneal routes. Components of systemic administration can be modified or formulated to be targeted, for example, to HSCs, hematopoietic stem / progenitor cells, or erythroid progenitor or precursor cells.

[0191] For example, local modes of administration include intraosseous injection into the trabecular bone or intramedullary injection into the medullary space of the femur, as well as infusion into the portal vein. In certain embodiments, significantly less of a component can be effective when administered locally (e.g., directly into the bone marrow) than when administered systemically (e.g., intravenously) as compared to systemic methods. Local modes of administration can reduce or eliminate the incidence of potential toxic side effects that can occur when a therapeutically effective amount of a component is administered systemically.

[0192] Administration can be provided as periodic boluses (e.g., intravenously) or as continuous infusion from an internal reservoir or from an external reservoir (e.g., from an intravenous bag or an implanted pump). Components can be administered locally, for example, by continuous release from a slow-release drug delivery device.

[0193] Additionally, components can be formulated to allow release over an extended period of time. Release systems can include biodegradable materials or matrices of materials that release incorporated components by diffusion. Components can be distributed uniformly or heterogeneously in the release system. A variety of release systems can be useful, however, the selection of an appropriate system will depend on the release rate required for a particular application. Both non-degradable and degradable release systems can be used. Suitable release systems include polymeric and polymeric matrices, non-polymeric matrices, or inorganic and organic excipients and diluents, such as, but not limited to, calcium carbonate and sugars (e.g., trehalose). Release systems can be natural or synthetic. However, synthetic release systems are preferred because they are generally more reliable, more reproducible, and produce more defined release profiles. Release system materials can be selected such that components having different molecular weights are released by diffusion or degradation of the material.

[0194] Representative synthetic biodegradable polymers include, for example: polyamides, such as poly(amino acids) and poly(peptides); polyesters, such as poly(lactic acid), poly(glycolic acid), poly(lactic-co-glycolic acid), and poly(caprolactone); poly(anhydrides); poly(orthoesters); polycarbonates; and chemical derivatives thereof (substitution, addition of chemical groups, such as alkyl, alkylene, hydroxylation, oxidation, and other modifications routinely made by those skilled in the art), copolymers, and mixtures thereof. Representative synthetic non-degradable polymers include, for example: polyethers, such as poly(ethylene oxide), poly(ethylene glycol), and poly(tetramethylene oxide); vinyl polymers - polyacrylates and polymethacrylates, such as methyl, ethyl, other alkyl, hydroxyethyl methacrylate, acrylic acid, and methacrylic acid, and other polymers, such as poly( vinyl alcohol), poly(vinyl pyrrolidone), and poly(vinyl acetate); poly(urethanes); celluloses and their derivatives, such as alkyl, hydroxyalkyl, ether, ester, nitrocellulose, and various cellulose acetates; polysiloxanes; and any chemical derivatives thereof (substitution, addition of chemical groups, such as alkyl, alkylene, hydroxylation, oxidation, and other modifications routinely made by those skilled in the art), copolymers, and mixtures thereof.

[0195] Poly(lactide-co-glycolide) microspheres can also be used. Typically, the microspheres are composed of polymers of lactic acid and glycolic acid structured to form hollow spheres. The spheres can be about 15-30 microns in diameter and can be loaded with the components described herein. In some embodiments, the genome editing system, system components, and / or nucleic acids encoding system components are delivered with a block copolymer such as a poloxamer or a poloxamine.

[0196] Methods of using the compositions of the disclosure

[0197] The disclosure provides use of the disclosed compositions or pharmaceutical compositions for treating a disease or condition in a cell, tissue, organ, animal, or subject, as known in the art or as described herein, using the disclosed compositions and pharmaceutical compositions, e.g., administering or contacting a therapeutically effective amount of the composition or pharmaceutical composition to the cell, tissue, organ, animal, or subject. In one aspect, the subject is a mammal. Preferably, the subject is a human. The terms “subject” and “patient” are used interchangeably herein.

[0198] The disclosure provides a method for modulating or treating at least one disease or condition in a cell, tissue, organ, animal, or subject. Preferably, the malignant disease is a β- hemoglobinopathy. Non-limiting examples of β-hemoglobinopathies include sickle cell disease and β-thalassemia.

[0199] The compositions of the present disclosure can be used to treat a disease or condition by using a therapeutic transgene encoding an exogenous nucleic acid sequence or an exogenous amino acid sequence. For certain diseases or conditions, the therapeutic transgene can include [disease] (therapeutic transgene): [beta-thalassemia] (HBB T87Q, BCL11A shRNA, IGF2BP1), [sickle cell disease] (HBB T87Q, BCL11A shRNA, IGF2BP1.

[0200] Genome editing systems, or cells altered or manipulated using such systems, can be administered to a subject by any suitable mode or route, whether local or systemic. Systemic administration modes include oral and parenteral routes. For example, parenteral routes include intravenous, intraosseous, intra-arterial, intramuscular, intradermal, subcutaneous, intranasal, and intraperitoneal routes. Components of systemic administration can be modified or formulated to be targeted, for example, to HSCs, hematopoietic stem / progenitor cells, or erythroid progenitor or precursor cells.

[0201] For example, local administration modes include intraosseous injection into the trabecular bone or intramedullary injection into the medullary space of the femur, and infusion into the portal vein. In certain embodiments, significantly less amount of components can function when administered locally (e.g., directly into the bone marrow) as compared to when administered systemically (e.g., intravenously). Local administration modes can reduce or eliminate the incidence of potential toxic side effects that can occur when a therapeutically effective amount of components is administered systemically.

[0202] Administration can be provided as periodic boluses (e.g., intravenously) or as continuous infusion from an internal reservoir or from an external reservoir (e.g., from an intravenous bag or an implanted pump). Components can be administered locally, for example, by continuous release from a slow-release drug delivery device.

[0203] Additionally, components can be formulated to allow release over an extended period of time. Release systems can include matrices of biodegradable materials or materials that release incorporated components by diffusion. Components can be distributed uniformly or heterogeneously in the release system. A variety of release systems can be useful, however, the selection of an appropriate system will depend on the release rate required for a particular application. Both non-degradable and degradable release systems can be used. Suitable release systems include polymeric and polymeric matrices, non-polymeric matrices, or inorganic and organic excipients and diluents, such as, but not limited to, calcium carbonate and sugars (e.g., trehalose). Release systems can be natural or synthetic. However, synthetic release systems are preferred because they are generally more reliable, more reproducible, and produce more defined release profiles. Release system materials can be selected such that components having different molecular weights are released by diffusion or degradation of the material.

[0204] Representative synthetic biodegradable polymers include, for example: polyamides, such as poly(amino acids) and poly(peptides); polyesters, such as poly(lactic acid), poly(glycolic acid), poly(lactic-co-glycolic acid), and poly(caprolactone); poly(anhydrides); poly(ortho esters); polycarbonates; and chemical derivatives thereof (substitution of chemical groups, addition of chemical groups, such as alkyl, alkylene, hydroxylation, oxidation, and other modifications routinely made by those skilled in the art), copolymers, and mixtures thereof. Representative synthetic non-degradable polymers include, for example: polyethers, such as poly(ethylene oxide), poly(ethylene glycol), and poly(tetramethylene oxide); vinyl polymers - polyacrylates and polymethacrylates, such as methyl, ethyl, other alkyl, hydroxyethyl methacrylate, acrylic acid, and methacrylic acid, and other polymers, such as poly( vinyl alcohol), poly(vinyl pyrrolidone), and poly(vinyl acetate); poly(urethanes); celluloses and their derivatives, such as alkyl, hydroxyalkyl, ether, ester, nitrocellulose, and various cellulose acetates; polysiloxanes; and chemical derivatives thereof (substitution of chemical groups, addition of chemical groups, such as alkyl, alkylene, hydroxylation, oxidation, and other modifications routinely made by those skilled in the art), copolymers, and mixtures thereof.

[0205] Poly(lactide-co-glycolide) microspheres can also be used. Typically, the microspheres are composed of polymers of lactic acid and glycolic acid structured to form hollow spheres. The spheres can be about 15-30 microns in diameter and can be loaded with the components described herein. In some embodiments, the genome editing system, system components, and / or nucleic acids encoding system components are delivered with a block copolymer such as a poloxamer or poloxamine.

[0206] Definitions

[0207] As used throughout this disclosure, the singular forms “a,” “and,” and “the” include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to “a method” includes a plurality of such methods, and reference to “a dose” includes reference to one or more doses and equivalents thereof known to those skilled in the art, and so forth.

[0208] The term "about" or "approximately" means within an acceptable error range for the particular value as determined by one of ordinary skill in the art, which will depend in part on how the value is measured or determined, e.g., the limitations of the measurement system. For example, "about" can mean within 1 or more multiples of the standard deviation determined for the particular measurement method. Alternatively, "about" can mean ranges approximately 20% or approximately 10% or approximately 5% or approximately 1% of a given value. Alternatively, particularly with respect to biological systems or processes, the term can mean within an order of magnitude, preferably within 5-fold and more preferably within 2-fold of a given value. Where particular values described in the application and claims depend on measurements, unless otherwise stated, the term "about" should be considered to mean within an acceptable error range for the particular measurement.

[0209] The present disclosure provides isolated or substantially purified polynucleotide or protein compositions. An "isolated" or "purified" polynucleotide or protein, or biologically active portion thereof, is substantially or essentially free of components of the environment as can be found in nature. Thus, when produced by recombinant techniques, an isolated or purified polynucleotide or protein is substantially free of cellular material or culture medium, or when produced via chemical synthesis, is substantially free of chemical precursors or other chemicals. Optimally, an "isolated" polynucleotide is free of sequences (optimally protein encoding sequences) that naturally flank the polynucleotide in the genomic DNA of the organism from which the polynucleotide is derived (i.e., sequences located at the 5' and 3' ends of the polynucleotide). For example, in various aspects, an isolated polynucleotide can contain less than about 5 kb, 4 kb, 3 kb, 2 kb, 1 kb, 0.5 kb, or 0.1 kb of nucleotide sequences that naturally flank the polynucleotide in genomic DNA of the cell from which the polynucleotide is derived. A protein that is substantially free of cellular material includes preparations of protein having less than about 30%, 20%, 10%, 5%, or 1% (by dry weight) of contaminant proteins. When the protein or biologically active portion thereof of the disclosure is produced recombinantly, optimally, the culture medium represents less than about 30%, 20%, 10%, 5%, or 1% (by dry weight) of chemical precursors or non- protein chemicals.

[0210] The present disclosure provides fragments and variants of the disclosed DNA sequences and proteins encoded by these DNA sequences. As used throughout the present disclosure, the term "fragment" refers to a portion of a DNA sequence or a portion of an amino acid sequence, and thus to a portion of a protein encoded thereby. Fragments of DNA sequences comprising coding sequences can encode protein fragments that retain the biological activity of the native protein, and thus retain DNA recognition or binding activity for a target DNA sequence as described herein. Alternatively, fragments of DNA sequences useful as hybridization probes generally do not encode proteins that retain biological activity or do not retain promoter activity. Thus, fragments of DNA sequences can range from at least about 20 nucleotides, about 50 nucleotides, about 100 nucleotides, and up to the full-length polynucleotide of the present disclosure.

[0211] Nucleic acids or proteins of the present disclosure can be constructed by a modular approach that includes pre-assembly of monomeric units and / or repeat units in a target vector, which can then be assembled into the final target vector. Polypeptides of the present disclosure can comprise repeat monomers of the present disclosure, and can be constructed by a modular approach by pre-assembly of repeat units in a target vector, which can then be assembled into the final target vector. The present disclosure provides polypeptides produced by this approach as well as nucleic acid sequences encoding these polypeptides. The present disclosure provides host organisms and cells comprising nucleic acid sequences encoding polypeptides produced by this modular approach.

[0212] The term "antibody" is used in the broadest sense, and specifically covers single monoclonal antibodies, including agonist and antagonist antibodies, as well as antibody compositions with polyepitopic specificity. The use of natural or synthetic analogs, mutants, variants, alleles, homologs and orthologs of antibodies as defined herein (collectively referred to herein as "analogues") are also within the scope of the present application. Thus, according to an aspect of the present application, the term "antibody herein" also encompasses such analogues in the broadest sense of the term. Generally, in such analogues, one or more amino acid residues can be substituted, deleted and / or added compared to an antibody as defined herein.

[0213] The term "binding" refers to a sequence-specific noncovalent interaction between macromolecules (e.g., between a protein and a nucleic acid). Not all components of a binding interaction need be sequence-specific (e.g., contact with a phosphate residue in the DNA backbone), so long as the overall interaction is sequence-specific.

[0214] The term "comprising" is intended to mean that the compositions and methods include the recited elements, but not excluding others. "Consisting essentially of" when used to define compositions and methods, shall mean excluding other elements of any essential significance to the composition and method. Thus, compositions consisting essentially of the elements as defined herein, do not exclude trace contaminants or inert carriers. "Consisting of" shall mean excluding more than trace elements of other ingredients. Each of these transitional terms is herein intended to, independently, convey the inclusion of the elements with which it is associated to the exclusion of other elements not specifically listed.

[0215] As used herein, "expression" refers to the process by which polynucleotides are transcribed into mRNA and / or the process by which the transcribed mRNA is subsequently translated into peptides, polypeptides, or proteins. If the polynucleotide is derived from genomic DNA, expression can include splicing of the mRNA in a eukaryotic cell.

[0216] "Gene expression" refers to the conversion of the information contained in a gene into a gene product. A gene product can be the direct transcriptional product of a gene (e.g., mRNA, tRNA, rRNA, antisense RNA, ribozyme, shRNA, microRNA, structural RNA, or any other type of RNA) or a protein produced by translation of mRNA. Gene products also include RNAs and proteins modified by processes such as capping, polyadenylation, methylation, and editing, and proteins modified by processes such as methylation, acetylation, phosphorylation, ubiquitination, ADP-ribosylation, myristoylation, and glycosylation.

[0217] "Modulation" or "modulation" of gene expression refers to a change in the gene's activity. Modulation of expression can include, but is not limited to, gene activation and gene repression.

[0218] "Operably linked" or its equivalent (e.g., "operably linked") means that two or more molecules are positioned relative to each other such that they are able to interact to affect the function possessed by one or both molecules or their combination.

[0219] Non-covalently linked components and methods of making and using non-covalently linked components are disclosed. Various components can take various different forms as described herein. For example, non-covalently linked (i.e., operably linked) proteins can be used to achieve transient interactions to avoid one or more problems in the art. The ability of non-covalently linked components, such as proteins, to associate and dissociate allows for functional association only or primarily where such association is needed to achieve a desired activity. This bonding can be sustained for a sufficient length of time to achieve a desired effect.

[0220] A method of directing a protein to a specific locus in the genome of an organism is disclosed. The method can include the steps of providing a DNA localization component and providing an effector molecule, wherein the DNA localization component and the effector molecule are operably linked via a non-covalent bond.

[0221] A "target site" or "target sequence" is a nucleic acid sequence defining the portion of a nucleic acid to which a binding molecule will bind, provided that sufficient binding conditions are present.

[0222] The term "nucleic acid" or "oligonucleotide" or "polynucleotide" refers to at least two nucleotides covalently linked together. Depiction of a single strand also defines the sequence of the complementary strand. Thus, a nucleic acid can also encompass the complementary strand of a depicted single strand. Nucleic acids of the disclosure also encompass substantially identical nucleic acids that retain the same structure or encode the same protein and their complements.

[0223] A probe of the disclosure can comprise a single-stranded nucleic acid capable of hybridizing to a target sequence under stringent hybridization conditions. Thus, a nucleic acid of the disclosure can refer to a probe that hybridizes under stringent hybridization conditions.

[0224] A nucleic acid of the disclosure can be single-stranded or double-stranded. A nucleic acid of the disclosure can contain double-stranded sequences even if the majority of the molecule is single-stranded. A nucleic acid of the disclosure can contain single-stranded sequences even if the majority of the molecule is double-stranded. A nucleic acid of the disclosure can include genomic DNA, cDNA, RNA, or hybrids thereof. A nucleic acid of the disclosure can contain a combination of deoxyribonucleotides and ribonucleotides. A nucleic acid of the disclosure can contain a combination of bases including uracil, adenine, thymine, cytosine, guanine, inosine, xanthine, hypoxanthine, isocytosine, and isoguanine. A nucleic acid of the disclosure can be synthesized to contain non-natural amino acid modifications. A nucleic acid of the disclosure can be obtained by chemical synthesis methods or by recombinant methods.

[0225] A nucleic acid of the disclosure (whether its entire sequence or any portion thereof) can be non-naturally occurring. A nucleic acid of the disclosure can contain one or more non-naturally occurring mutations, substitutions, deletions, or insertions such that the entire nucleic acid sequence is non-naturally occurring. A nucleic acid of the disclosure can contain one or more repeated, inverted, or repeated sequences such that the resulting sequence is not naturally occurring, such that the entire nucleic acid sequence is non-naturally occurring. A nucleic acid of the disclosure can contain non-naturally occurring modified, artificial, or synthetic nucleotides such that the entire nucleic acid sequence is non-naturally occurring.

[0226] Due to the redundancy of the genetic code, multiple nucleotide sequences can encode any particular protein. All such nucleotide sequences are contemplated herein.

[0227] As used throughout this disclosure, the term "promoter" refers to a molecule of synthetic or natural origin capable of conferring, activating or enhancing expression of a nucleic acid in a cell. A promoter can comprise one or more specific transcriptional regulatory sequences to further enhance expression and / or alter its spatial expression and / or temporal expression. A promoter can also comprise a distal enhancer or suppressor element, which can be located thousands of base pairs from the transcriptional start site. Promoters can be derived from sources including viruses, bacteria, fungi, plants, insects, and animals. Promoters can constitutively or differentially regulate expression of genomic components in terms of the cell, tissue, or organ in which expression occurs, or in terms of the developmental stage at which expression occurs, or in response to external stimuli such as physiological stress, pathogens, metal ions, or inducers. Representative examples of promoters include the bacteriophage T7 promoter, the bacteriophage T3 promoter, the SP6 promoter, the lac operator promoter, the tac promoter, the SV40 late promoter, the SV40 early promoter, the RSV-LTR promoter, the CMV IE promoter, the EF-1 alpha promoter, the CAG promoter, the SV40 early promoter, or the SV40 late promoter and the CMV IE promoter.

[0228] As used throughout this disclosure, the term "substantially complementary" refers to a first sequence being at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, or 99% identical to the complement of a second sequence over a region of 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 180, 270, 360, 450, 540, or more nucleotides or amino acids, or the two sequences hybridize under stringent hybridization conditions.

[0229] As used throughout this disclosure, the term "substantially identical" refers to a first sequence and a second sequence being at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, or 99% identical over a region of 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 180, 270, 360, 450, 540, or more nucleotides or amino acids (or in the case of nucleic acids, if the first sequence is substantially complementary to the complement of the second sequence).

[0230] As used throughout this disclosure, the term "variant" when used to describe a nucleic acid refers to (i) a portion or fragment of a reference nucleotide sequence; (ii) a complement of a reference nucleotide sequence or portion thereof; (iii) a nucleic acid that is substantially identical to a reference nucleic acid or complement thereof; or (iv) a nucleic acid that hybridizes under stringent conditions to a reference nucleic acid, complement thereof, or to a sequence substantially identical thereto.

[0231] As used throughout this disclosure, the term "vector" refers to a nucleic acid sequence that contains an origin of replication. The vector can be a viral vector, a bacteriophage, a bacterial artificial chromosome, or a yeast artificial chromosome. The vector can be a DNA or RNA vector. The vector can be a self-replicating extrachromosomal vector, and is preferably a DNA plasmid. The vector can comprise a combination of amino acid and DNA sequences, RNA sequences, or both DNA and RNA sequences.

[0232] As used throughout this disclosure, the term "variant" when used to describe a peptide or polypeptide refers to a peptide or polypeptide that differs in amino acid sequence by an insertion, deletion, or conservative substitution of amino acids, but which retains at least one biological activity. Variant can also mean a protein having an amino acid sequence substantially identical to a reference protein, wherein the amino acid sequence of the reference protein retains at least one biological activity.

[0233] Conservative substitution of an amino acid, i.e., replacing an amino acid with a different amino acid of similar properties (e.g., hydrophilicity, degree and distribution of charged regions), is recognized in the art as a minor change. As understood in the art, these minor changes can be identified, in part, by considering the hydropathic index of amino acids. Kyte et al., J. Mol. Biol. 157: 105-132 (1982). The hydropathic index of an amino acid is based on a consideration of its hydrophobicity and charge. Amino acids of similar hydropathic index can be substituted and still retain protein function. In one aspect, amino acids having hydropathic indices of ± 2 are substituted. The hydrophilicity of amino acids can also be used to reveal substitutions that will result in proteins retaining biological function. Consideration of the hydrophilicity of amino acids in the context of a peptide permits the identification of substitutions that will result in proteins retaining biological function. U.S. Pat. No. 4,554,101, which is incorporated herein by reference in its entirety.

[0234] Substitution of amino acids with similar hydrophilicity values can result in peptides retaining biological activity, e.g., immunogenicity. Substitutions can be made with amino acids having hydrophilicity values within ±2 of each other. Both the hydropathic and hydrophilic indices of amino acids are influenced by the particular side chain of the amino acid. Consistent with this observation, amino acid substitutions that are compatible with biological function are understood to depend on the relative similarity of the amino acids, and particularly the side chains of those amino acids, as revealed by the hydropathic, hydrophilic, electropositive, electronegative, and other properties.

[0235] As used herein, "conservative" amino acid substitutions can be defined as set forth in Table 3, Table 4, and Table 5. In some aspects, the fusion polypeptides and / or nucleic acids encoding such fusion polypeptides include conservative substitutions that have been introduced by modifying a polynucleotide encoding a polypeptide of the disclosure. Amino acids can be categorized according to physical properties and contributions to secondary and tertiary protein structure. A conservative substitution is the replacement of one amino acid with another having similar properties. Exemplary conservative substitutions are set forth in Table 3.

[0236] Table 3 - Conservative Substitutions I

[0237]

[0238] Alternatively, conservative amino acids can be in accordance with Lehninger, Biochemistry, Second Edition; Worth Publishers, Inc. NY, N.Y. (1975), pp. 71-77, as set forth in Table 4.

[0239] Table 4 - Conservative Substitutions II

[0240]

[0241] Alternatively, exemplary conservative substitutions are set forth in Table 5.

[0242] Table 5 - Conservative Substitutions III

[0243]

[0244] It is understood that the polypeptides of the disclosure are intended to include polypeptides that carry an insertion, deletion, or substitution of one or more amino acid residues, or any combination thereof, as well as modifications other than an insertion, deletion, or substitution of amino acid residues. The polypeptides or nucleic acids of the disclosure can contain one or more conservative substitutions.

[0245] As used throughout this disclosure, the term "more than one" of the above amino acid substitutions refers to 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 or more of the amino acid substitutions. The term "more than one" can refer to 2, 3, 4, or 5 of the amino acid substitutions.

[0246] The polypeptides and proteins of the present disclosure, whether the entire sequence thereof or any portion thereof, can be non-naturally occurring. The polypeptides and proteins of the present disclosure can contain one or more non-naturally occurring mutations, substitutions, deletions, or insertions such that the entire amino acid sequence is non-naturally occurring. The polypeptides and proteins of the present disclosure can contain one or more repeated, inverted, or repeated sequences such that the resulting sequence is not naturally occurring, such that the entire amino acid sequence is non-naturally occurring. The polypeptides and proteins of the present disclosure can contain non-naturally occurring modified, artificial, or synthetic amino acids such that the entire amino acid sequence is non-naturally occurring.

[0247] As used throughout this disclosure, "sequence identity" can be determined by using the standalone executable BLAST engine program (bl2seq) for blasting two sequences using default parameters, which is available from the National Center for Biotechnology Information (NCBI) ftp site (Tatusova and Madden, FEMS Microbiol Lett., 1999, 174, 247-250; incorporated by reference herein in its entirety). When the terms "identical" or "identity" are used in the context of two or more nucleic acids or polypeptide sequences, it is recognized that residues in specific positions can be varied and the term "percent identical" is used to describe the sequences. The percent identity is calculated by optimally aligning the two sequences, comparing the identical residues in the designated regions of the two sequences, determining the number of positions in which the residues are identical, dividing the number of identical positions by the total number of positions in the designated region, and multiplying the result by 100 to yield the percent sequence identity. In cases where the two sequences have different lengths or the alignment results in one or more staggered ends and the designated comparison region includes only a single sequence, the residues of the single sequence are included in the denominator of the calculation but not in the numerator. When comparing DNA and RNA, thymine (T) and uracil (U) can be considered equivalent. Identity can be determined manually or using a computer sequence algorithm such as BLAST or BLAST 2.0.

[0248] As used throughout this disclosure, the term "endogenous" refers to a nucleic acid or protein sequence that is naturally associated with a target gene or a host cell into which the target gene is introduced.

[0249] The present disclosure provides methods of introducing a polynucleotide construct comprising a DNA sequence into a host cell. By "introducing" is intended to present the polynucleotide construct to the cell in a manner that the construct can enter the interior of a host cell. The methods of the present disclosure do not depend on the particular method used to introduce the polynucleotide construct into the host cell, so long as the polynucleotide construct can enter the interior of a cell of the host. Methods for introducing polynucleotide constructs into bacteria, plants, fungi, and animals are known in the art, including but not limited to stable transformation methods, transient transformation methods, and virus-mediated methods.

[0250] As used herein, the terms "substantially" or "essentially" mean a quantity, level, value, number, frequency, percent, dimension, size, amount, weight, or length that is about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or greater of a reference quantity, level, value, number, frequency, percent, dimension, size, amount, weight, or length. In one embodiment, the terms "substantially identical" or "essentially identical" mean that a range of quantities, levels, values, numbers, frequencies, percentages, dimensions, sizes, amounts, weights, or lengths is approximately the same as a reference quantity, level, value, number, frequency, percent, dimension, size, amount, weight, or length.

[0251] As used herein, the terms "substantially free" and "essentially free" are used interchangeably and, when used in reference to a composition such as a cell population or a culture medium, mean a composition that is free of a specified material or source thereof, such as 95% free, 96% free, 97% free, 98% free, 99% free of the specified material or source thereof, or not detectable by conventional methods of measurement. The terms "free of" or "essentially free of" a certain ingredient or material in a composition also mean that such ingredient or material is not included in the composition (1) at any concentration, or (2) at a concentration that is functionally inert but low. Similar meanings can apply to the term "absence" in reference to the absence of a particular material or source thereof in a composition.

[0252] Throughout this specification, unless the context requires otherwise, the words "comprise," "comprises" and "comprising" will be understood to imply the inclusion of a stated step or element or group of steps or elements but not the exclusion of any other step or element or group of steps or elements. In particular embodiments, the terms "comprising," "having," "including," and "containing" are synonymous with each other.

[0253] "Consisting of means including, and limited to, whatever follows the phrase "consisting of." Thus, the phrase "consisting of indicates that the listed elements are required or mandatory, and that no other elements can be present. "Consisting essentially of means including, and limited to, whatever follows the phrase "consisting essentially of or any variations thereof, wherein additional elements are not excluded necessitating use of the term "comprising."

[0254] "Consisting essentially of means that any elements included in the phrase are present, and further that no other elements are present in the specified amount or at all, that interfere with or contribute to the activity or action specified in the disclosure for the listed elements. Thus, the phrase "consisting essentially of indicates that the listed elements are essential or mandatory, but that other elements are optional and can or can not be present depending upon whether or not they affect the activity or action of the listed elements.

[0255] References to "one embodiment", "an embodiment”, "certain embodiments”, "some embodiments”, "additional embodiments” or "further embodiments” or combinations thereof throughout this disclosure, or any variations thereof, mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the disclosure. Thus, the appearances of the above- described phrases or combinations thereof in various places throughout the specification are not necessarily all referring to the same embodiment. Furthermore, the particular features, structures, or characteristics can be combined in any suitable manner in one or more embodiments.

[0256] The term "ex vivo” generally refers to activities that occur outside of an organism, such as experiments or measurements performed in or on living tissue in an artificial environment outside of the organism, preferably with minimal changes to natural conditions. In particular embodiments, an "ex vivo” procedure involves living cells or tissue that is taken from an organism and cultured in a laboratory setting, typically under sterile conditions, and typically for a few hours or up to about 24 hours, but including up to 48 or 72 hours or more, depending on the circumstances. In certain embodiments, such tissue or cells can be collected and frozen, and then thawed for ex vivo processing. The use of living cells or tissue for tissue culture experiments or procedures that last longer than a few days is generally considered to be "in vitro”, although in certain embodiments the term can be used interchangeably with ex vivo.

[0257] The term "in vivo” generally refers to activities that occur within an organism.

[0258] As used herein, the term "reprogramming” or "de-differentiation” or "increasing cell potency” or "increasing developmental potency” refers to a method of increasing cell potency or de-differentiating a cell to a less differentiated state. For example, a cell having increased cell potency has more developmental plasticity (i.e., can differentiate into more cell types) than the same cell in a non-reprogrammed state. In other words, a reprogrammed cell is a cell in a less differentiated state than the same cell in a non-reprogrammed state.

[0259] As used herein, the term "differentiation" is the process by which a nonspecialized ("uncommitted") or less specialized cell acquires the characteristics of a specialized cell, such as, for example, a blood cell or a muscle cell. A differentiated or differentiation-induced cell is a cell that occupies a more specialized ("committed") position within a cell lineage. The term "committed" (when used in reference to the process of differentiation) means that the cell has proceeded to a certain point in the differentiation pathway where, under normal circumstances, it will continue to differentiate into a particular cell type or subset of cell types, and under normal circumstances cannot differentiate into a different cell type or revert to a less differentiated cell type. As used herein, the term "pluripotent" refers to the ability of a cell to form all lineages of the body or somatic body (i.e., the blastema). For example, embryonic stem cells are a type of pluripotent stem cell that are capable of forming cells from each of the three germ layers (ectoderm, mesoderm, and endoderm). Pluripotency is a spectrum of developmental potential, from incomplete or partial pluripotent cells (e.g., epiblast stem cells or EpiSCs) that are unable to generate a complete organism to more primitive, more potent cells (e.g., embryonic stem cells) that are capable of generating a complete organism.

[0260] As used herein, the term "induced pluripotent stem cell" or iPSC means a stem cell produced from a differentiated adult, neonatal, or fetal cell that has been induced or altered, i.e., reprogrammed, to be a cell that is capable of differentiating into tissues of all three germ layers or dermal layers: mesoderm, endoderm, and ectoderm. The iPSC produced is not a cell found in nature.

[0261] As used herein, the term "subject" refers to any animal, preferably a human patient, livestock, or other domesticated animal.

[0262] A "pluripotency factor" or "reprogramming factor" refers to an agent that is capable of increasing the developmental potential of a cell, alone or in combination with other agents. Pluripotency factors include, but are not limited to, polynucleotides, polypeptides, and small molecules that are capable of increasing the developmental potential of a cell. Exemplary pluripotency factors include, for example, transcription factors and small molecule reprogramming agents.

[0263] "Culturing" or "cell culturing" refers to the maintenance, growth, and / or differentiation of cells in an in vitro environment. "Cell culture medium," "medium" (singular "medium" in each case), "supplement," and "medium supplement" refer to a nutritional composition that cultures a cell culture.

[0264] "Culturing" or "maintaining" refers to the sustaining, propagating (growing), and / or differentiating of cells outside of a tissue or body, for example, in a sterile plastic (or coated plastic) cell culture dish or flask. "Culturing" or "maintaining" can utilize a medium as a source of nutrients, hormones, and / or other factors that aid in the propagation and / or maintenance of cells.

[0265] The terms “hematopoietic stem and progenitor cells,” “hematopoietic stem cells,” “hematopoietic progenitor cells,” or “hematopoietic precursor cells” refer to cells committed to the hematopoietic lineage but capable of further hematopoietic differentiation and include pluripotent hematopoietic stem cells (hemocytoblasts), myeloid progenitor cells, megakaryocyte progenitor cells, erythroid progenitor cells, and lymphoid progenitor cells. Hematopoietic stem and progenitor cells (HSCs) are multipotent stem cells that give rise to all blood cell types, including myeloid (monocytes and macrophages, neutrophils, basophils, eosinophils, erythrocytes, megakaryocytes / platelets, dendritic cells) and lymphoid lineages (T cells, B cells, NK cells). As used herein, the term “definitive hematopoietic stem cells” refers to CD34+ hematopoietic cells that are capable of generating both types of mature myeloid and lymphoid cells, including T cells, NK cells, and B cells. Hematopoietic cells also include various primitive hematopoietic cell subpopulations that give rise to primitive red blood cells, macrocytes, and macrophages.

[0266] As used herein, the term “isolated” and the like refer to a cell or population of cells that has been removed from its original environment, i.e., the environment of an “unisolated” reference cell is substantially free of at least one component found in the environment in which the reference cell exists. The term includes cells removed from some or all components when the cell is found in its natural environment, e.g., tissue, biopsy. The term also includes cells removed from at least one, some, or all components when the cell is found in an environment that does not naturally occur, e.g., culture, cell suspension. Thus, an isolated cell is partially or completely separated from at least one component, including other substances, cells, or populations of cells, when the cell is found in nature or grown, stored, or otherwise present in an environment that does not naturally occur. Particular examples of isolated cells include partially purified cells, substantially purified cells, and cells cultured in a non-naturally occurring medium. Isolated cells can be obtained by separating the desired cell or population thereof from other substances or cells in the environment, or removing one or more other cell populations or subpopulations from the environment. As used herein, the term “purified” and the like refer to an increase in purity. For example, the purity can be increased to at least 50%, 60%, 70%, 80%, 90%, 95%, 99%, or 100%.

[0267] As used herein, the term "encoding" refers to the inherent property of specific sequences of nucleotides in a polynucleotide, such as a gene, a cDNA or an mRNA, to serve as templates for synthesis of other polymers and macromolecules in biological processes having defined sequences of nucleotides (i.e., rRNA, tRNA, and mRNA) or of amino acids and that are relied on by biological machinery for their role in protein synthesis, or that are otherwise useful in a biological context. Thus, if a mRNA corresponding to a gene is transcribed and translated, it produces a protein, the gene encodes the protein. Either the coding strand or the non-coding strand of a gene or cDNA may be referred to as encoding a protein or other product of that gene or cDNA, both the nucleotide sequence of the coding strand and the nucleotide sequence of the non-coding strand being provided in the sequence listing.

[0268] A "construct" refers to a macromolecule or molecular complex comprising a polynucleotide to be delivered to a host cell in vitro or in vivo. As used herein, a "vector" refers to any nucleic acid construct capable of delivering or transferring foreign genetic material to a target cell, which foreign genetic material can be replicated and / or expressed in the target cell. As used herein, the term "vector" includes a construct to be delivered. A vector can be linear or circular molecule. A vector can be integrative or non-integrative. Major types of vectors include, but are not limited to, plasmids, episomes, viral vectors, cosmids, and artificial chromosomes. Viral vectors include, but are not limited to, adenoviral vectors, adeno-associated viral vectors, retroviral vectors, lentiviral vectors, Sendai virus vectors, and the like.

[0269] "Integration" means the stable insertion of one or more nucleotides of a construct into the genome of a cell, i.e., covalent linkage to nucleic acid sequences within the chromosomal DNA of the cell. "Targeted integration" means the insertion of nucleotides of a construct into a preselected site or "integration site" within the chromosomal or mitochondrial DNA of a cell. As used herein, the term "integration" further refers to the process involving the insertion of one or more exogenous sequences or nucleotides of a construct, with or without the deletion of endogenous sequences or nucleotides at the integration site. In cases where a deletion is present at the site of insertion, "integration" can further include the replacement of the endogenous sequence or deleted nucleotides with the one or more inserted nucleotides.

[0270] As used herein, the term "exogenous" means the introduction of a reference molecule or reference activity into a host cell. For example, a molecule can be introduced into a host genetic material by integration into a host chromosome or as non-chromosomal genetic material such as a plasmid. Thus, when the term is used in reference to the expression of a coding nucleic acid, it refers to the introduction of the coding nucleic acid in expressible form into a cell. The term "endogenous" refers to a reference molecule or activity present in a host cell. Similarly, when the term is used in reference to the expression of a coding nucleic acid, it refers to the expression of a coding nucleic acid contained within a cell and not introduced exogenously.

[0271] As used herein, a "gene of interest" or "polynucleotide of interest" is a DNA sequence that, when placed under the control of appropriate regulatory sequences, is transcribed into RNA and, in some cases, translated into a polypeptide in vivo. The gene or polynucleotide of interest can include, but is not limited to, prokaryotic sequences, cDNA from eukaryotic mRNA, genomic DNA sequences from eukaryotic (e.g., mammalian) DNA, and synthetic DNA sequences. For example, the gene of interest can encode a miRNA, shRNA, a natural polypeptide (i.e., a polypeptide found in nature) or fragment thereof; a variant polypeptide (i.e., a mutant of a natural polypeptide having less than 100% sequence identity to the natural polypeptide) or fragment thereof; an engineered polypeptide or peptide fragment, a therapeutic peptide or polypeptide, an imaging marker, a selectable marker, etc.

[0272] As used herein, the term "polynucleotide" refers to a polymeric form of nucleotides of any length, either deoxyribonucleotides or ribonucleotides, or their analogs. Polynucleotide sequences are composed of four nucleotide bases: adenine (A); cytosine (C); guanine (G); thymine (T); and uracil (U) in place of thymine (T) when the polynucleotide is RNA. Polynucleotides can include a gene or gene fragment (for example, a probe, primer, EST or SAGE tag), exons, introns, messenger RNA (mRNA), transfer RNA, ribosomal RNA, ribozymes, cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, nucleic acid probes, and primers. Polynucleotide also refers to both double- and single-stranded molecules.

[0273] As used herein, the terms "peptide," "polypeptide," and "protein" can be used interchangeably and refer to molecules having amino acid residues covalently linked by peptide bonds. A polypeptide must contain at least two amino acids, and there is no maximum number of amino acids to a polypeptide. As used herein, the term refers to both short chains (e.g., also commonly referred to in the art as peptides, oligopeptides, and oligomers) and long chains (also commonly referred to in the art as polypeptides or proteins). "Polypeptide" includes, for example, biologically active fragments, substantially homologous polypeptides, oligopeptides, homodimers, heterodimers, variants of polypeptides, modified polypeptides, derivatives, analogs, fusion proteins, and the like. Polypeptides include natural polypeptides, recombinant polypeptides, synthetic polypeptides, or combinations thereof.

[0274] As used herein, the term "engager" refers to a molecule, e.g., a fusion polypeptide, capable of forming a linkage between an immune cell (e.g., T cell, NK cell, NKT cell, B cell, macrophage, neutrophil) and a tumor cell; and activating the immune cell. Examples of engagers include, but are not limited to, bispecific T cell engagers (BiTEs), bispecific killer cell engagers (BiKEs), trispecific killer cell engagers, or multispecific killer cell engagers, or universal engagers compatible with multiple immune cell types.

[0275] As used herein, the term "specific" or "specificity" can be used to refer to the ability of a molecule (e.g., a receptor or an engager) to selectively bind to a target molecule, in contrast to non-specific or non-selective binding.

[0276] "Functional" as used in the context of genomic editing or modification of iPSCs, and derivative non-pluripotent cells differentiated therefrom or non-pluripotent cells reprogrammed therefrom, and genomic editing or modification of derivative iPSCs, means (1) successful knock-in, knock-out, knock-down of gene expression, transgene or controlled gene expression at the genetic level, such as inducible or transient expression at a desired cell developmental stage, achieved by direct genomic editing or modification or by "passage" via differentiation or reprogramming of a starting cell from which the initial genomic engineering was performed; or (2) at the cellular level, successful removal, addition or alteration of cell function / characteristics via (i) gene expression modifications obtained in the cell by direct genomic editing, (ii) gene expression modifications maintained in the cell by "passage" via differentiation or reprogramming of a starting cell from which the initial genomic engineering was performed, (iii) downstream gene regulation in the cell resulting from gene expression modifications that only occur in early developmental stages of the cell, or only in the starting cell from which the cell was generated via differentiation or reprogramming, or (iv) enhanced or newly acquired cell functions or attributes originally derived from the genomic editing or modification performed at the iPSC, progenitor cell or dedifferentiated cell of origin, exhibited in the mature cell product.

[0277] Examples

[0278] The examples in this section are provided for illustration and are not intended to limit the present application.

[0279] Example 1: Methods of genetic engineering of hematopoietic stem and progenitor cells

[0280] Exemplary methods of isolating human hematopoietic stem and progenitor cells (HSPCs) from peripheral blood

[0281] Human CD34 was isolated from mobilized peripheral blood from healthy donors using a Miltenyi Biotec automated CliniMACS Prodigy instrument and TS 510 procedure equipped with CD34 GMP microbeads (Miltenyi Biotec, #170-076-711) according to the manufacturer's instructions. + HSPC. The purity of isolated cells was assessed by flow cytometry using CD34-AF647 (BioLegend, clone 561, #343618) and CD45-PE (BioLegend, clone HI30, #304008) and DAPI (BioLegend, #422801) as viability stains. On the same day as isolation, cells were frozen in CryoStor CS10 medium (STEMCELL Technologies, #07930).

[0282] Exemplary methods of introducing nucleic acids into HSPCs using electroporation

[0283] Frozen CD34 separated using the above method + HSPC was thawed and incubated at 37°C in a humidified incubator under a 5% CO2 atmosphere at a rate of 1x10. 6 HSPCs were cultured at 100 ng / mL in StemSpan SFEM II (STEMCELL Technologies, #09655) medium supplemented with 100 ng / mL recombinant human stem cell factor, 100 ng / mL recombinant human Flt3 / Flk-2 ligand, and 100 ng / mL recombinant human thrombopoietin (STEMCELL Technologies catalog numbers 78062, 78009, and 78210.1, respectively). After 24 hours of culture, HSPCs were electroporated using the P3 PrimaryCell 4D-Nucleofector X Kit (Lonza, #V4XP-3032) with program EO-100 according to the manufacturer's instructions. The electroporated cells were then cultured at 7.5 x 10⁻⁶ cells / mL. 5 Cells / mL were resuspended in the culture medium. One day after electroporation, HSPCs underwent a complete culture medium replacement and were cultured for another 3 days for further downstream analysis and assays.

[0284] Exemplary methods for erythroid differentiation of HSPCs

[0285] Four days after electroporation, HSPC was applied at 2 x 10 5Cells were counted every three to four days and expansion media was added to the cultures to maintain 2 x 10 5 Cells were counted every three to four days and expansion media was added to the cultures to maintain 2 x 10 6 Cells were counted every three to four days and expansion media was added to the cultures to maintain 2 x 10 6 Cells were counted every three to four days and expansion media was added to the cultures to maintain 2 x 10

[0286] Example 2: Methods for analyzing gene edited hematopoietic stem and progenitor cells

[0287] Exemplary methods for detecting and quantifying insertions or deletions (indels) in HSPCs and erythroid differentiated cells for detection and quantification of HBG gene editing Exemplary methods of analyzing and quantifying HSPC and erythroid progenitor cell populations using flow cytometry

[0288] Genomic DNA (gDNA) was extracted from control and HBG gene edited HSPCs using Quick-DNA Microprep Kit (Zymo, #D3020) according to manufacturer’s instructions. PCR amplification of HBG1 / 2 genes was performed on extracted genomic DNA using Platinum™ SuperFi II PCR Master Mix (ThermoFisher, #12368050) according to manufacturer’s instructions using forward primer: ACACTCTTTCCCTACACGACGCTCTTCCGATCTGCAGTATCCTCTTGGGGG (SEQ ID NO: 24), and reverse primer: GACTGGAGTTCAGACGTGTGCTCTTCCGATCTACCTCAGACGTTCCAGAAGC (SEQ ID NO: 25) flanking the HBG gene region of interest.

[0289] PCR products were further purified using Select-a-Size DNA Clean & Concentrator MagBead Kit (Zymo, #D4085) according to manufacturer’s instructions, followed by next generation sequencing (NGS) library preparation using NEBNext® Ultra™ II Q5® Master Mix (New England BioLabs, #M0544X) and NEBNext® Multiplex Oligos for Illumina® (96 Index Primers) (New England BioLabs, #E6609S) for sample multiplexing. Libraries were pooled and double-end sequenced on an Illumina MiSeq with a depth of at least 20,000 reads per sample according to manufacturer’s instructions. Generated FASTQ files were used as input for indel quantification using CRISPResso2 software with the following options: -min_average_read_quality 30, -min_single_bp_quality 10, -exclude_bp_from_left 15, -exclude_bp_from_right 15, -ignore_substitutions TRUE, -amplicon_min_alignment_score 60, -quantification_window_size 90.

[0290] To perform separate indel measurements at the single HBG1 or HBG2 loci, a pre- amplification step was added to the above procedure using Platinum™ SuperFi II PCR Master Mix with the following PCR primers.

[0291] HBG1 Forward: TCCACAGTACCTGCCAAAGA (SEQ ID NO: 26);

[0292] HBG1 Reverse: GCCTACCTTCCCAGGGTTTC (SEQ ID NO: 27);

[0293] HBG2 Forward: GGCCTAAAACCACAGAGAGTAT (SEQ ID NO: 28); and

[0294] HBG2 Reverse: CCCCACAGGCTTGTGATAGT (SEQ ID NO: 29). Indel measurements at the single HBG1 or HBG2 loci were determined using the above method.

[0295] Exemplary methods of measuring indels in cells differentiated from unedited and HBG edited HSPCs using RT-qPCR

[0296] Prior to erythroid differentiation, HSPCs were analyzed by flow cytometry to assess their expression of stemness markers including CD34-BV785 (BioLegend, clone 561, #343626), CD90-APC (BioLegend, clone 5E10, #328114), CD38-PE (BioLegend, clone HB-7, #356604), and CD133-PE (BD, clone 293C3, #567917), as well as the viability dye Fixable Viability Stain 450 (BD, #562247). Cells analyzed during erythroid differentiation were harvested and stained for cell surface with the erythroid lineage markers CD235a-APC (BioLegend, clone HI264, #349114) and CD71-PerCP / Cy5.5 (BioLegend, clone CY1G4, #334114), as well as the viability dye Fixable Viability Stain 450. Subsequently, differentiated cells were fixed and permeabilized using Transcription Factor Buffer Set (BD, #562725) followed by intracellular staining for fetal hemoglobin using HbF-PE (BD, clone 2D12, #560041). Stained cells were measured on a BD FACS Celesta or Agilent NovoCyte Quanteon according to the manufacturer’s instructions.

[0297] Exemplary methods of measuring percent gamma globin gene expression. Exemplary methods of analyzing and quantifying cells differentiated from unedited and HBG edited HSPC cells using colony forming unit (CFU) assay

[0298] Harvested HSPC differentiated cells and extracted mRNA using RNeasy Mini Kit (QIAGEN, #74106) and performed on-column DNA digestion according to manufacturer’s instructions (RNase-Free DNase Set, QIAGEN, #79254). mRNA reverse transcription was performed using iScript™ Reverse Transcription Supermix (Bio-Rad, # 1708841). qPCR was performed on the resulting cDNA using SsoAdvanced Universal Probes Supermix (Bio-Rad, # 1725281) on a Bio-Rad CFX Opus 384 Real-Time PCR instrument according to manufacturer’s instructions. TaqMan probe-based assays were ordered from ThermoFisher to detect gene expression of HBG1 / 2 (Hs00361131_g1), HBB (Hs00747223_g1), HBA1 (Hs07292163_s1), and RPP30 (Hs01124518_m1).

[0299] A. Indels via next generation sequencing (NGS) B. Colony forming unit assay

[0300] Four days post electroporation, control and edited HSPCs were frozen in CryoStor CS10 media. Frozen HSPCs were used for colony forming unit (CFU) assays according to the manufacturer’s instructions. After thawing, cells were added to aliquots of MethoCult™ GF+ H4435 (StemCell Technologies, #04445). To maximize the number of well-isolated colonies, cells were plated at 3 densities (250, 500, and 1500 cells / well), each with 3 replicate wells on 6-well STEMvision™ SmartDish™ (StemCell Technologies, #27371). Cells were placed in a humidified incubator at 37°C, 5% CO2 atmosphere for 12-14 days. Subsequently, the total number of erythroid (BFU-E), myeloid (CFU-GM), and multipotent progenitor (CFU-GEMM) colonies were evaluated by trained personnel and counted based on morphology. For each sample, well-isolated BFU-E and CFU-GM colonies were harvested from the culture and genomic DNA was extracted from the harvested colonies using QuickExtract DNA Extraction Solution (Biosearch Technologies, #QE0905T). Remaining colonies were also bulk harvested from each harvested culture and genomic DNA was extracted using the Quick-DNA Microprep Kit. Indel detection and quantification was performed on gDNA from individual colonies and bulk cultures.

[0301] Example 3: Compositions comprising Cas-CLOVER and exemplary HBG gRNA pairs result in dose-dependent editing of the HBG locus in HSPCs

[0302] Exemplary compositions comprising Cas-CLOVER and gRNA pairs were designed to edit the BCL11 A binding site upstream of the HBG1 and HBG2 genes in HSPCs. HSPCs were isolated and frozen according to the methods described above. HSPCs were then thawed and cultured for 24 hours according to the methods described above. After 24 hours of culture, HSPCs were nucleofected with mRNA encoding Cas-CLOVER and gRNA pairs according to Table 6.

[0303] Table 6 Cas-CLOVER and gRNA pairs targeting HBG1 or HBG2

[0304]

[0305] Electroporated cells were resuspended in media at 7.5 x 10 5 The HBG genomic region encompassing the Cas-CLOVER targeting site was PCR amplified and subjected to next generation sequencing (NGS). Cas-CLOVER edit-induced small insertions and deletions (indels) were quantified using the CRISPResso2 bioinformatics pipeline described above, as measured by the percentage of modified reads over total sequencing reads. All four gRNA pairs produced indels at the HBG1 and HBG2 loci in electroporated HSPCs in a dose-dependent manner (Figure 1). gRNA pairs #2, #3, and #4 produced indels significantly higher than pair #1. At a concentration of about 400 pg / ml, pair #3 resulted in greater than 30% modified reads.

[0306] The experimental procedure was repeated using additional gRNA pairs and concentrations. HSPCs were nucleofected with mRNA encoding Cas-CLOVER and gRNA pairs according to Table 7.

[0307] Table 7 Cas-CLOVER and gRNA pairs targeting HBG1 or HBG2

[0308]

[0309] The results of this experiment are shown in Figure 2. gRNA pairs #2, #3, #4, and #5 exhibited significant editing of the HBG locus, resulting in indel percentages between about 20-35%, and an increasing trend in editing percentage at higher gRNA concentrations. In comparison to gRNA pairs #2, #3, #4, and #5, gRNA pairs #6, #7, and #8 exhibited reduced levels of HBG editing.

[0310] The experimental procedure was repeated using certain gRNA pairs and Cas-Clover variants. HSPCs were nucleofected with mRNA encoding Cas-CLOVER (100 pg / mL each variant) and gRNA pairs (200 pg / mL each guide) according to Table 8.

[0311] Table 8 Cas-CLOVER and gRNA pairs targeting HBG1 or HBG2

[0312]

[0313] The results of this experiment are shown in Figure 8. gRNA pairs 11, 11.2, 12, and 12.2 exhibited clear editing of the HBG locus, resulting in indel percentages between about 25-50%, comparable or better than benchmark pair ID 4 in Table 8.

[0314] Example 4: HBG-edited HSPCs maintain multilineage differentiation potential and increase HBG mRNA and HbF protein levels

[0315] HSPCs were nucleofected with 200 pg / ml Cas-CLOVER mRNA and 400 pg / ml of each gRNA pair #2, #3, or #4 according to the above-described method. Four days later, 1 x 105 5 electroporated cells and gDNA was extracted essentially according to the above-described method. The remaining HPSCs were split into two independent groups: the first group was converted to erythroid differentiation as described above; and the second group was subjected to colony forming unit (CFU) assay on MethoCult-based cultures essentially as described above.

[0316] C. Measuring HBG mRNA expression in edited HSPCs by RT-qPCR

[0317] HBG editing-induced indels in HSPCs, erythroid progenitors differentiated from edited HSPCs, and bulk colonies collected at the end of CFU assay were measured via NGS according to the above-described method and quantified from cells at different time points post-nucleofection by CRISPResso2. At day 4, HSPCs showed strong editing at the HBG locus, as evidenced by NGS sequencing, with an editing rate of approximately 40-50% of cells for all three tested gRNA pairs (Figure 3). The percentage of HBG-edited cells remained unchanged at day 11 post-erythroid differentiation, demonstrating that HBG-edited HSPCs can differentiate into erythroid progenitors that retain the edited HBG locus. Additionally, bulk colonies collected at the end of CFU assay also maintained similar editing percentages as HSPCs and erythroid progenitors, demonstrating that HBG editing is maintained upon differentiation of HSPCs along the erythropoietic lineage.

[0318] D. HbF protein expression levels

[0319] At day 20, the absolute number of CFU cell types for HSPC control samples and HSPC edited using gRNA pair #2, #3, or #4 for HBG editing were shown (Figure 4A). There was no change in the absolute number of cell types CFU-GEMM (CFU- granulocyte / erythroid / macrophage / megakaryocyte), CFU-GM (CFU-granulocyte / macrophage), BFU-E (burst-forming unit erythroid), and CFU-E (CFU-erythroid) in the CFU assay for HSPC edited using the three HBG sgRNA pairs compared to the three controls: EP only (nucleofection / electroporation only), CC only (Cas-CLOVER mRNA only), and sgRNA only. Furthermore, the relative distribution of CFU cell types between control and HBG edited cells using pair #2, #3, or #4 also remained unchanged (Figure 4B), demonstrating that editing using the compositions of the present disclosure did not affect HSPC multi-lineage colony-forming ability.

[0320]

[0321] HBG mRNA expression was measured by RT-qPCR according to the method described above and normalized to HBB at days 14, 18, and 21 during erythroid differentiation. Adult and cord blood samples were used as negative and positive controls, respectively. Fold change was calculated by normalizing the observed expression level to the EP only control expression level at each time point.

[0322] An increase in HBG mRNA expression was observed in HSPC cells edited using HBG gRNA pairs #2, #3, and #4 at all time points tested, with HBG mRNA expression in cells edited using each gRNA pair increasing from about 4-fold to about 9-fold from day 14 to day 21 (Figure 5).

[0323]

[0324] HbF protein was detected in erythroid differentiated cells of control cells and HBG edited HSPC by intracellular flow cytometry according to the method described above. The percentage of F cells (as defined by HbF positive staining (HbF+)) determined by intracellular flow cytometry is plotted on the left y-axis (Figure 6). All three gRNA pairs demonstrated an increase in the number of HbF positive cells, with the percentage of HbF positive cells increasing by about 2-fold to about 5-fold compared to the EP control.

[0325] The median fluorescence intensity (MFI) of HbF signal per F cell was quantified relative to EP only control and plotted on the right y-axis (Figure 6). There was an increase in MFI of HBG-edited F cells compared to EP control, with #3 demonstrating the highest MFI per F cell among the three gRNA pairs tested.

Claims

1. A composition comprising a) A first guide RNA (gRNA) and a first fusion protein configured to form a complex with the first gRNA or a first polynucleotide encoding the first fusion protein, the first fusion protein comprising: a mutant Cas9 (dCas9) polypeptide or its inactivating nuclease domain and a Clo051 polypeptide or its nuclease domain, and b) A second gRNA and a second fusion protein configured to form a complex with the second gRNA or a second polynucleotide encoding the second fusion protein, the second fusion protein comprising: a dCas9 polypeptide or an inactivating nuclease domain thereof and a Clo051 polypeptide or a nuclease domain thereof. in i) The first gRNA contains a first targeting sequence, the first targeting sequence comprising a nucleotide sequence selected from SEQ ID NO: 1, 5, 7, 13, 15 or 17; and The second gRNA contains a second targeting sequence, which contains the nucleotide sequence of SEQ ID NO: 3, or ii) The first gRNA contains a first targeting sequence, the first targeting sequence comprising a nucleotide sequence selected from SEQ ID NO: 7, 9, or 13; and The second gRNA contains a second targeting sequence, which contains the nucleotide sequence of SEQ ID NO:

11.

2. The composition of claim 1, wherein the first gRNA comprises a first scaffold sequence and the second gRNA comprises a second scaffold sequence, wherein the first scaffold sequence and the second scaffold sequence comprise nucleotide sequences selected from SEQ ID NO: 20 or 21.

3. The composition according to claim 2, wherein... a) The first gRNA contains a nucleotide sequence selected from SEQ ID NO: 2, 6, 8, 16, or 18; and the second gRNA contains a nucleotide sequence selected from SEQ ID NO: 4, or b) The first gRNA contains a nucleotide sequence selected from SEQ ID NO: 10, 14 or 19; and the second gRNA contains a nucleotide sequence of SEQ ID NO:

12.

4. The composition according to any one of claims 1 to 3, wherein the first gRNA, the second gRNA, or both the first gRNA and the second gRNA comprise one or more chemical modifications of ribonucleotides, ribonucleotide bases, or phosphodiester bonds.

5. The composition according to claim 4, wherein the one or more chemical modifications comprise at least one chemically modified phosphodiester bond.

6. The composition according to claim 5, wherein the at least one chemically modified phosphate diester bond is a thiophosphate bond.

7. The composition of claim 6, comprising at least two consecutive phosphate thioester bonds at the 5' end of the first gRNA, the second gRNA, or both the first gRNA and the second gRNA.

8. The composition according to any one of claims 4 to 6, comprising a 2' O-Me chemical modification at the 3' end of the first gRNA, the second gRNA, or both the first gRNA and the second gRNA.

9. The composition according to any one of claims 1 to 8, wherein dCas9 is derived from Streptococcus pyogenes Cas9 polypeptide or Staphylococcus aureus Cas9 polypeptide.

10. The composition according to any one of claims 1 to 8, wherein the C-terminus of dCas9 or its inactivated nuclease domain is attached to the N-terminus of the Clo051 polypeptide or its nuclease domain via a peptide linker sequence selected from GGGGS and SEQ ID NO:

23.

11. The composition according to any one of claims 9 to 10, wherein the first fusion protein comprises the amino acid sequence of SEQ ID NO: 39, 41, 42 or 43.

12. The composition according to any one of claims 9 to 11, wherein the second fusion protein comprises the amino acid sequence of SEQ ID NO: 39, 41, 42 or 43.

13. The composition according to any one of claims 9 to 12, wherein the first polynucleotide, the second polynucleotide, or both the first polynucleotide and the second polynucleotide are mRNA.

14. The composition of claim 13, wherein the mRNA comprises a 5'-cap.

15. A composition comprising: i) A first gRNA containing the nucleotide sequence of SEQ ID NO: 2, a second gRNA containing the nucleotide sequence of SEQ ID NO: 4, a first polynucleotide sequence encoding a first fusion protein of SEQ ID NO: 39, and a second polynucleotide sequence encoding a second fusion protein of SEQ ID NO: 39; ii) A first gRNA containing the nucleotide sequence of SEQ ID NO: 6, a second gRNA containing the nucleotide sequence of SEQ ID NO: 4, a first polynucleotide sequence encoding a first fusion protein of SEQ ID NO: 39, and a second polynucleotide sequence encoding a second fusion protein of SEQ ID NO: 39; iii) A first gRNA containing the nucleotide sequence of SEQ ID NO: 8, a second gRNA containing the nucleotide sequence of SEQ ID NO: 4, a first polynucleotide sequence encoding a first fusion protein of SEQ ID NO: 41, and a second polynucleotide sequence encoding a second fusion protein of SEQ ID NO: 41; iv) A first gRNA containing the nucleotide sequence of SEQ ID NO: 10, a second gRNA containing the nucleotide sequence of SEQ ID NO: 12, a first polynucleotide sequence encoding the first fusion protein of SEQ ID NO: 39, and a second polynucleotide sequence encoding the second fusion protein of SEQ ID NO: 42; v) A first gRNA containing the nucleotide sequence of SEQ ID NO: 14, a second gRNA containing the nucleotide sequence of SEQ ID NO: 12, a first polynucleotide sequence encoding a first fusion protein of SEQ ID NO: 42, and a second polynucleotide sequence encoding a second fusion protein of SEQ ID NO: 42; vi) A first gRNA containing the nucleotide sequence of SEQ ID NO: 19, a second gRNA containing the nucleotide sequence of SEQ ID NO: 12, a first polynucleotide sequence encoding the first fusion protein of SEQ ID NO: 42, and a second polynucleotide sequence encoding the second fusion protein of SEQ ID NO: 42; vii) A first gRNA containing the nucleotide sequence of SEQ ID NO: 16, a second gRNA containing the nucleotide sequence of SEQ ID NO: 4, a first polynucleotide sequence encoding a first fusion protein of SEQ ID NO: 41, and a second polynucleotide sequence encoding a second fusion protein of SEQ ID NO: 41; or viii) A first gRNA containing the nucleotide sequence of SEQ ID NO: 18, a second gRNA containing the nucleotide sequence of SEQ ID NO: 4, a first polynucleotide sequence encoding the first fusion protein of SEQ ID NO: 41, and a second polynucleotide sequence encoding the second fusion protein of SEQ ID NO:

41.

16. The composition according to any one of claims 1 to 15, wherein the composition is encapsulated in at least one lipid nanoparticle (LNP), said at least one lipid nanoparticle comprising: The compound of formula (I) contains approximately 40.75% by molar amount. Formula (I) It contains approximately 51.75% cholesterol (calculated as a percentage). Approximately 5% DOPC, and Approximately 2.5% DMG-PEG2000 (by molar amount); The first and second polynucleotides are RNA molecules, and the ratio of lipids to RNA molecules in at least one nanoparticle is approximately 120:1 (w / w).

17. The composition according to any one of claims 1 to 15, wherein the composition is encapsulated in at least one LNP, said at least one LNP comprising: It contains approximately 54% SS-OP, approximately 35% cholesterol, approximately 5% DOPC, approximately 5% DSPC, and approximately 1% DMG-PEG2000 (by molar weight). The first and second polynucleotides are RNA molecules, and the ratio of lipids to RNA molecules in at least one nanoparticle is about 100:1 (w / w) and the total lipids are 25 nM.

18. The composition according to any one of claims 1 to 17, used for modifying the HBG1 gene, HBG2 gene, BCL11A gene or a combination thereof in cells.

19. A method for modifying a cell population, the method comprising: Contact the cell population with the composition according to any one of claims 1 to 18. The first gRNA forms a complex with the first target sequence and the first fusion protein, and the second gRNA forms a complex with the second target sequence and the second fusion protein. This results in an insertion / deletion between the first and second target sequences, and generates a modified cell population.

20. The method of claim 19, wherein the insertion / deletion results in the inactivation of the BCL11A gene.

21. The method according to any one of claims 19 to 20, wherein the modified cell population has increased γ-globulin expression relative to the unmodified cell population by a factor of about 4 to about 9.

22. The method according to any one of claims 19 to 20, wherein the modified cell population has an increased level of fetal hemoglobin (HbF) expression relative to the unmodified cell population.

23. The method according to any one of claims 19 to 21, wherein the cells are hematopoietic stem cells and progenitor cells (HSPCs).

24. The method of claim 23, wherein the HSPC is capable of differentiating into erythroid progenitor cells.

25. A cell population modified by the method according to any one of claims 19 to 24.

26. A method of treating β-hemoglobinopathies in a subject with this need, the method comprising administering to the subject the composition according to any one of claims 1 to 18 or the cell population according to claim 25.

27. The method of claim 26, wherein the β-hemoglobinopathy is β-thalassemia or sickle cell disease.

Citation Information

Patent Citations

  • Site-specific enzymes and methods of use

    US10415024B2

  • Compositions and methods for improved encapsulation of functional proteins in polymeric vesicles

    US10456452B2

  • A method for directing proteins to specific loci in the genome and uses thereof

    US20170107541A1

  • Methods and compositions for in vivo non-covalent linking

    US20170114149A1

  • Compositions and methods for directing proteins to specific loci in the genome

    US20180187185A1