Compositions and methods for treating hemoglobin diseases
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANGHAI VITALGEN BIOPHARMA CO LTD
- Filing Date
- 2024-09-27
- Publication Date
- 2026-05-08
AI Technical Summary
The prior art has limitations in the treatment of hemoglobin diseases, especially in the difficulty of providing safe and effective treatments, and requires a fully matched bone marrow or cord blood donor to avoid the risk of post-transplant GVHD.
Increase the level of HbF in human erythrocyte progeny of human CD34+HSPCs by genome editing, and targeted editing in the transcriptional regulatory regions of HBG1 and HBG2 using guide RNA (gRNA) and CRISPR/Cas systems to treat β-thalassemia and sickle cell disease.
This method can significantly increase the expression level of HbF, reduce dependence on fully matched donors, reduce the risk of GVHD after transplantation, and provide a potential cure for beta-hemoglobinosis.
Smart Images

Figure 00000041_0000 
Figure 00000042_0000 
Figure 00000042_0001
Abstract
Description
Compositions and methods for treating hemoglobin disorders Technical Field
[0001] The present application relates to the field of biomedicine, and specifically to a molecule for treating hemoglobin diseases, a composition containing the molecule, and methods and uses thereof. Background Art
[0002] Hemoglobinopathies are a highly heterogeneous group of inherited anemias, classified according to the deficient globin chain and encompassing α- and β-hemoglobinopathies, as well as quantitative hemoglobin abnormalities. β-Thalassemia is a heterogeneous autosomal recessive anemia characterized by reduced or absent β-globin chain synthesis, resulting in a hemoglobin A (HbA)-related disorder. Excessive hemoglobin A chains can lead to impaired erythropoiesis, intramedullary apoptosis of erythroid precursors, and hemolytic anemia. β-Thalassemia is clinically divided into two types, differentiated by the severity of symptoms: β-Thalassemia major (or β0, in which β-globin chain production is abolished by a mutation) and β-Thalassemia intermedia (or β+, in which β-globin chain production is reduced). β-Thalassemia major is a serious medical condition requiring regular blood transfusions. Sickle cell disease (SCD) is an inherited blood disorder caused by a defect in the production of β-globin chains. It includes sickle cell anemia, sickle hemoglobin C disease (HbSC), sickle β+ thalassemia (HbS / β+), and sickle β0 thalassemia (HbS / β0).
[0003] Research has shown that all forms of β-hemoglobinopathy are caused by mutations in the structural β-globin gene (HBB). β-thalassemia is caused by over 200 different β-globin gene mutations. SCD is caused by an A-to-T point mutation in codon 6, which results in an E6V substitution in the defective β-globin (βs). The human β-globin locus, comprised of five genes located within a short region of chromosome 11, is responsible for producing the β-globin chains of hemoglobin. In addition to the β-globin genes, this locus also encompasses the δ, γ-A (HBG1), γ-G (HBG2), and ε globins. A 16-kb locus control region (LCR), located approximately 40-60 kb upstream of the HBB locus, is thought to regulate the differential expression of β-globin-like genes throughout development. In the late neonatal period, fetal γ-A and γ-G expression is repressed, while adult β-globin gene expression is activated; this process is known as the γ-to-β globin expression switch. Hereditary persistence of fetal hemoglobin (HbF) (HPFH) is a benign condition in which high levels of HbF are produced into adulthood. Patients heterozygous for HPFH and β-hemoglobinopathies have mild clinical symptoms, while homozygous patients for these disorders have severe symptoms. HPFH is typically caused by mutations in the β-globin gene cluster or the γ-globin promoter region that either create new binding sites for erythroid activators or disrupt binding sites for repressors. Point mutations in the proximal promoters of HBG1 and HBG2 fall into two distinct groups: approximately 115 bp and 200 bp upstream of the duplicated γ-globin gene transcription start site (TSS), respectively. The major repressor in γ-globin gene silencing is BCL11A, which binds to the -115 bp site. Recently, ZBTB7A, also known as LRF or FBI-1, has been identified as a second major repressor binding to the -200 bp site. These findings suggest that introducing artificial HFPH-like mutations into the promoter region to reactivate γ-globin gene expression is a promising therapeutic strategy for β-hemoglobinopathies.
[0004] In China, although improvements in clinical treatment have reduced the mortality rate of children with β-hemoglobinopathies, supportive care is still the main treatment for most patients. Current treatments aim to relieve symptoms and treat complications, including regular blood transfusions, suppression of erythropoiesis, iron removal, and analgesia. Human leukocyte antigen (HLA)-identical sibling bone marrow transplantation (BMT) or HLA-matched donor umbilical cord blood transplantation (CBT) is another feasible treatment option, but carries the risk of graft-versus-host disease (GVHD). However, allogeneic BMT is only suitable for a small proportion of patients with β-hemoglobinopathies because most patients cannot find a fully matched unrelated donor.
[0005] Novel therapeutic approaches utilizing autologous genetically modified HSPCs may obviate the need for fully matched bone marrow or umbilical cord blood donors, thereby avoiding the risk of GVHD after transplantation. Genome-wide association studies have shown that several single nucleotide polymorphisms (SNPs) located in the erythroid-specific enhancer of the BCL11A locus are associated with downregulation of BCL11a expression, persistent γ-globin expression in adulthood, and mild clinical manifestations of β-hemoglobinopathies. Gene therapy using corrected β-globin is also considered a possible cure for β-hemoglobinopathies. Lentiviral vectors (LVs) without HIV-1 replication elements are capable of transferring exogenous gene fragments into quiescent hematopoietic stem cells. The BB305 lentiviral vector encoding fully functional βA-T87Q-globin was used. Autologous HSPC gene therapy has freed patients with severe β-hemoglobinopathies from regular blood transfusions. However, the safety of such gene therapy is a concern, as the random integration of the vector gene into the genome and the introduction of strong promoter and enhancer elements within the LV vector may lead to activation of oncogenes or other deleterious events caused by such effects.
[0006] Overall, despite recent progress in developing genetic approaches to address beta-hemoglobinopathies, several unmet clinical needs remain in providing safe and effective treatments for these genetic disorders.
[0007] Summary of the Invention
[0008] To overcome the limitations of existing technologies for treating hemoglobinopathies, the present application provides a method for increasing the level of HbF in human erythrocyte progeny of human CD34+ HSPCs by genome editing. The two γ-globin chains of HbF are expressed by HBG1 and HBG2. The molecules, compositions, cells (including autologous CD34+ HSPCs that can be infused into patients with severe hemoglobinopathies), kits, and methods of the present application act on the transcriptional regulatory regions of HBG1 and HBG2 and can be used to treat hemoglobinopathies such as β-thalassemia and sickle cell disease.
[0009] On the one hand, the present application provides a guide RNA (gRNA) molecule, which comprises a targeting domain complementary to a target sequence, wherein the target sequence is located within the HBG1 promoter region or within the HBG2 promoter region, wherein the HBG1 promoter region is located between chr11:5,248,269 and 5,249,857 of the human genome hg19, and the HBG2 promoter region is located between chr11:5,253,188 and 5,254,781 of the human genome hg19.
[0010] In some embodiments, the target sequence is located between about 210 bp upstream to about 100 bp upstream of the transcription start site of the HBG1 and the HBG2.
[0011] In some embodiments, the gRNA comprises a tracr sequence and a crRNA sequence.
[0012] In some embodiments, the gRNA is a two-component guide RNA molecule.
[0013] In some embodiments, the gRNA is a single-component guide RNA molecule (sgRNA).
[0014] In some embodiments, wherein (1) a CRISPR / Cas system comprising the gRNA molecule is contacted with the target sequence, or (2) a CRISPR / Cas system comprising the gRNA molecule is introduced into a cell where a gene comprising the target sequence is located, a base insertion or deletion (indel) can be formed at or near the target sequence.
[0015] In some embodiments, the CRISPR / Cas system comprises a Cas12b protein and / or a functionally active fragment thereof.
[0016] In some embodiments, the CRISPR / Cas system comprises a Cas12i protein and / or a functionally active fragment thereof.
[0017] In some embodiments, the N-terminus and / or C-terminus of the nuclease is linked to one or more nuclear localization signals.
[0018] In some embodiments, the nuclear localization signal is linked to a His tag.
[0019] In some embodiments, the tracr sequence comprises the nucleotide sequence of SEQ ID NO: 60. For example, the tracr sequence comprises the following sequence (m represents a 2-O-methyl modification, * represents a 3' phosphorothioate modification): mG*mG*mT*CGTCTATAGGACGGCGAGTTTTTCAACGGGTGTGCCAATGGCCACTTTCCAGGTGGCAAAGCCCGTTGAACTTCTCAAAAAGAACGCTCGCTCAGTGTTCT*mG*mA*mC
[0020] In some embodiments, the crRNA sequence comprises a nucleotide sequence selected from any one of SEQ ID NOs: 61-66. For example, the crRNA sequence comprises a sequence as shown in the following table:
[0021] In some embodiments, it comprises a nucleotide sequence selected from any one of SEQ ID NOs: 12-59 and 67-78. For example, the gRNA sequence comprises a sequence as shown in the following table:
[0022] On the other hand, the present application provides a nucleic acid molecule encoding the gRNA molecule described in the present application.
[0023] In another aspect, the present application provides a vector comprising the nucleic acid molecule described in the present application.
[0024] In some embodiments, the vector is selected from the group consisting of a lentiviral vector, an adenoviral vector, an adeno-associated virus (AAV) vector, a herpes simplex virus (HSV) vector, a plasmid, a minicircle, a nanoplasmid, and an RNA vector.
[0025] On the other hand, the present application provides a ribonucleoprotein (RNP) complex comprising the gRNA molecule described in the present application and the nuclease of the CRISPR / Cas system.
[0026] In some embodiments, the nuclease is a class II type V Cas enzyme.
[0027] In some embodiments, the nuclease is Cas12b nuclease.
[0028] In some embodiments, the Cas12b nuclease comprises a variant thereof, an ortholog thereof, and / or a functionally active fragment thereof.
[0029] In some embodiments, the Cas12b nuclease is the Cas12b nuclease from Alicyclobacillus acidiphilus (AaCas12b).
[0030] In some embodiments, the AaCas12b nuclease comprises one or more of the following mutations relative to the wild-type AaCas12b nuclease: (1) replacing one or more amino acid residues that interact with the pre-spacer adjacent motif (PAM) in the wild-type AaCas12b nuclease with a positively charged amino acid residue; (2) replacing one or more amino acid residues involved in opening the DNA double strand in the wild-type AaCas12b nuclease with an amino acid residue having an aromatic ring; and / or (3) replacing one or more amino acid residues in the RuvC domain of the wild-type AaCas12b nuclease that interacts with a single-stranded DNA substrate with a positively charged amino acid residue or a hydrophobic amino acid residue, wherein the amino acid sequence of the wild-type AaCas12b nuclease is as set forth in SEQ ID NO: 91.
[0031] In some embodiments, the one or more amino acid residues that interact with a PAM are located at one or more of the following positions: 116, 123, 130, 132, 144, 145, 153, 173, 222, 395, 400, and 475; and wherein the amino acid residues are numbered according to SEQ ID NO:91.
[0032] In some embodiments, the positively charged amino acid residue replacing one or more amino acid residues that interact with the PAM in the wild-type AaCas12b nuclease is one or more of the following substitutions: D116R and E475R; and wherein the amino acid residues are numbered according to SEQ ID NO: 91.
[0033] In some embodiments, the one or more amino acid residues involved in opening the DNA double strand are located at one or more of the following positions: 118, and 119; and wherein the amino acid residues are numbered according to SEQ ID NO:91.
[0034] In some embodiments, the substitution of one or more amino acid residues involved in opening the DNA double strand in the wild-type AaCas12b nuclease with an amino acid residue having an aromatic ring is a Q119Y, Q119F or Q119W substitution; and wherein the amino acid residues are numbered according to SEQ ID NO: 91.
[0035] In some embodiments, the one or more amino acid residues located in the RuvC domain and interacting with the single-stranded DNA substrate are located at one or more of the following positions: 300, 301, 304, 329, 636, 639, 647, 682, 757, 758, 761, 764, 768, 852, 854, 856, 857, 858, 860, 862, 863, 865, 866, 867, 869, 938, 956, 957, 958, 994, 1093, and 1097; and wherein the amino acid residues are numbered according to SEQ ID NO: 91.
[0036] In some embodiments, the substitution of one or more amino acid residues in the wild-type AaCas12b nuclease that are located in the RuvC domain and interact with the single-stranded DNA substrate is one or more of the following substitutions: E636R, Q639R, T647R, Q682R, I757R, E758R, E761R, Q854R, N857K, D858R, I994R, Q1093R and W1097R; and wherein the amino acid residues are numbered according to SEQ ID NO: 91.
[0037] In some embodiments, the gRNA molecule comprises the nucleotide sequence of any one of SEQ ID NOs: 12-59.
[0038] In some embodiments, the nuclease is a Cas12i nuclease.
[0039] In some embodiments, the Cas12i nuclease comprises a variant thereof, an ortholog thereof, and / or a functionally active fragment thereof.
[0040] In some embodiments, the Cas12i nuclease comprises one or more of the following mutations relative to the wild-type Cas12i2 nuclease: (1) one or more amino acids that interact with PAM in the wild-type Cas12i2 nuclease are replaced with positively charged amino acids; (2) one or more amino acids involved in opening the DNA double strand in the wild-type Cas12i2 nuclease are replaced with amino acids with aromatic rings; (3) one or more amino acids in the wild-type Cas12i2 nuclease that are located in the RuvC domain and interact with the single-stranded DNA substrate are replaced with positively charged amino acids; and / or (4) one or more amino acids that interact with the DNA-RNA double helix in the wild-type Cas12i2 nuclease are replaced with positively charged amino acids, wherein the amino acid sequence of the wild-type Cas12i2 nuclease is as described in SEQ ID NO: 92.
[0041] In some embodiments, the one or more amino acids that interact with a PAM are located at one or more of the following positions: 176, 238, 447, and / or 563; and wherein the positions are numbered according to SEQ ID NO:92.
[0042] In some embodiments, wherein (1) the positively charged amino acid is R or K.
[0043] In some embodiments, the Cas12i nuclease comprises any one or combination of the following mutations: (1) E563R; (2) E176R, T447R, E176R, and E563R; (3) K238R and E563R; (4) E176R, K238R, and T447R; (5) E176R, K238R, and E563R; (6) E176R, T447R, and E563R; and / or (7) E176R, K238R, T447R, and E563R; and wherein the amino acid residues are numbered according to SEQ ID NO: 92.
[0044] In some embodiments, the one or more amino acids involved in opening the DNA double strand are located at one or more of the following positions: 163 and / or 164; and wherein the positions are numbered according to SEQ ID NO:92.
[0045] In some embodiments, the amino acid with an aromatic ring in (2) is F, Y or W.
[0046] In some embodiments, the one or more amino acids involved in opening the DNA double strand are replaced with amino acids with aromatic rings, namely: Q163F, Q163Y, Q163W, and / or N164F or N164Y; and the amino acid residues are numbered according to SEQ ID NO: 92.
[0047] In some embodiments, the one or more amino acids located in the RuvC domain and interacting with the single-stranded DNA substrate are located at one or more of the following positions: 323, 362, 425, 925, 926, 391, 424 and / or 929; and wherein the positions are numbered according to SEQ ID NO: 92.
[0048] In some embodiments, wherein (3) the positively charged amino acid is R or K.
[0049] In some embodiments, the Cas12i nuclease comprises any one or combination of the following mutations: (1) E323R; (2) D362R; (3) Q425R; (4) N925R; (5) I926R; (6) E323R and D362R; (7) E323R and Q425R; (8) E323R and I926R; (9) Q425R and I926R; (10) D362R and I926R; (11) N925R and I926R; (12) E323R, D362R and Q425R; (13) E323R, D362R and I926R; (14) E323R, Q425R and I926R; (15) D362R, N925R and I926R; and / or (16) E323R, D362R, Q425R and I926R; and wherein the amino acid residues are numbered according to SEQ ID NO:92.
[0050] In some embodiments, the one or more amino acids that interact with the DNA-RNA double helix are located at one or more of the following positions: 116, 117, 159, 161, 319, 343 and / or 958; and wherein the positions are numbered according to SEQ ID NO:92.
[0051] In some embodiments, wherein (4) the positively charged amino acid is R or K.
[0052] In some embodiments, the Cas12i nuclease comprises any one or combination of the following mutations: G116R, E117R, T159R, S161R, E319R, E343R, and / or D958R; and wherein the amino acid residues are numbered according to SEQ ID NO: 92.
[0053] In some embodiments, the Cas12i nuclease further comprises one or more mutations in the flexible region, which can increase the flexibility of the flexible region of the Cas12i nuclease compared to the wild-type Cas12i2 nuclease.
[0054] In some embodiments, the flexible region is selected from the group corresponding to amino acid residues 228-232, amino acid residues 439-443, amino acid residues 478-482, amino acid residues 500-504, amino acid residues 775-779, and amino acid residues 925-929; and wherein the amino acid residues are numbered according to SEQ ID NO:92.
[0055] In some embodiments, the mutation of the one or more flexible regions comprises insertion of one or more G residues in the flexible region.
[0056] In some embodiments, the one or more G residues are inserted at the N-terminus of the flexible amino acid residues in the flexible region, wherein the flexible amino acid residues are selected from the group consisting of G, S, N, D, H, M, T, E, Q, K, R, A, and P.
[0057] In some embodiments, the mutation of one or more flexible regions comprises replacing a hydrophobic amino acid residue in the flexible region with a G residue, wherein the hydrophobic amino acid residue is selected from the group consisting of L, I, V, C, Y, F, and W.
[0058] In some embodiments, the mutation in one or more flexible regions is: (1) I926G; and / or (2) 439G or 439GG, and wherein the amino acid residues are numbered according to SEQ ID NO:92.
[0059] In some embodiments, the Cas12i nuclease comprises the mutations: E176R, K238R, T447R, E563R, N164Y, E323R, and D362R; and wherein the amino acid residues are numbered according to SEQ ID NO: 92.
[0060] In some embodiments, the gRNA molecule comprises the nucleotide sequence of any one of SEQ ID NOs: 67-78.
[0061] In some embodiments, the N-terminus and / or C-terminus of the nuclease is linked to one or more nuclear localization signals.
[0062] In some embodiments, the nuclear localization signal is linked to a His tag.
[0063] In some embodiments, the nuclease comprises an amino acid sequence selected from any one of SEQ ID NOs: 1-11.
[0064] On the other hand, the present application provides a composition comprising: (1) one or more gRNA molecules according to any one of claims 1 to 11, and a nuclease of the CRISPR / Cas system; (2) a nucleic acid encoding one or more gRNA molecules according to any one of claims 1 to 11, and a nuclease of the CRISPR / Cas system; (3) one or more gRNA molecules according to any one of claims 1 to 11, and a nucleic acid encoding a nuclease of the CRISPR / Cas system; or (4) a nucleic acid encoding one or more gRNA molecules according to any one of claims 1 to 11, and a nucleic acid encoding a nuclease of the CRISPR / Cas system.
[0065] In some embodiments, the one or more gRNA molecules and the nuclease of the CRISPR / Cas system in (1) are present in a ribonucleoprotein (RNP) complex.
[0066] In some embodiments, the nuclease is a class II type V Cas enzyme.
[0067] In some embodiments, the nuclease is Cas12b nuclease.
[0068] In some embodiments, the Cas12b nuclease comprises a variant thereof, an ortholog thereof, and / or a functionally active fragment thereof.
[0069] In some embodiments, the Cas12b nuclease is the Cas12b nuclease from Alicyclobacillus acidiphilus (AaCas12b).
[0070] In some embodiments, the AaCas12b comprises one or more of the following mutations relative to the wild-type AaCas12b nuclease: (1) replacing one or more amino acid residues that interact with the pre-spacer adjacent motif (PAM) in the wild-type AaCas12b nuclease with a positively charged amino acid residue; (2) replacing one or more amino acid residues involved in opening the DNA double strand in the wild-type AaCas12b nuclease with an amino acid residue having an aromatic ring; and / or (3) replacing one or more amino acid residues in the RuvC domain of the wild-type AaCas12b nuclease that interacts with a single-stranded DNA substrate with a positively charged amino acid residue or a hydrophobic amino acid residue, wherein the amino acid sequence of the wild-type AaCas12b nuclease is as set forth in SEQ ID NO: 91.
[0071] In some embodiments, the one or more amino acid residues that interact with a PAM are located at one or more of the following positions: 116, 123, 130, 132, 144, 145, 153, 173, 222, 395, 400, and 475; and wherein the amino acid residues are numbered according to SEQ ID NO:91.
[0072] In some embodiments, the positively charged amino acid residue replacing one or more amino acid residues that interact with the PAM in the wild-type AaCas12b nuclease is one or more of the following substitutions: D116R and E475R; and wherein the amino acid residues are numbered according to SEQ ID NO: 91.
[0073] In some embodiments, the one or more amino acid residues involved in opening the DNA double strand are located at one or more of the following positions: 118, and 119; and wherein the amino acid residues are numbered according to SEQ ID NO:87.
[0074] In some embodiments, the substitution of one or more amino acid residues involved in opening the DNA double strand in the wild-type AaCas12b nuclease with an amino acid residue having an aromatic ring is a Q119Y, Q119F or Q119W substitution; and wherein the amino acid residues are numbered according to SEQ ID NO: 91.
[0075] In some embodiments, the one or more amino acid residues located in the RuvC domain and interacting with the single-stranded DNA substrate are located at one or more of the following positions: 300, 301, 304, 329, 636, 639, 647, 682, 757, 758, 761, 764, 768, 852, 854, 856, 857, 858, 860, 862, 863, 865, 866, 867, 869, 938, 956, 957, 958, 994, 1093, and 1097; and wherein the amino acid residues are numbered according to SEQ ID NO: 91.
[0076] In some embodiments, the substitution of one or more amino acid residues in the wild-type AaCas12b nuclease that are located in the RuvC domain and interact with the single-stranded DNA substrate is one or more of the following substitutions: E636R, Q639R, T647R, Q682R, I757R, E758R, E761R, Q854R, N857K, D858R, I994R, Q1093R and W1097R; and wherein the amino acid residues are numbered according to SEQ ID NO: 91.
[0077] In some embodiments, the gRNA molecule comprises the nucleotide sequence of any one of SEQ ID NOs: 12-59.
[0078] In some embodiments, the nuclease is a Cas12i nuclease.
[0079] In some embodiments, the Cas12i nuclease comprises a variant thereof, an ortholog thereof, and / or a functionally active fragment thereof.
[0080] In some embodiments, the Cas12i nuclease comprises one or more of the following mutations relative to the wild-type Cas12i2 nuclease: (1) one or more amino acids that interact with PAM in the wild-type Cas12i2 nuclease are replaced with positively charged amino acids; (2) one or more amino acids involved in opening the DNA double strand in the wild-type Cas12i2 nuclease are replaced with amino acids with aromatic rings; (3) one or more amino acids in the wild-type Cas12i2 nuclease that are located in the RuvC domain and interact with the single-stranded DNA substrate are replaced with positively charged amino acids; and / or (4) one or more amino acids that interact with the DNA-RNA double helix in the wild-type Cas12i2 nuclease are replaced with positively charged amino acids, wherein the amino acid sequence of the wild-type Cas12i2 nuclease is as described in SEQ ID NO: 92.
[0081] In some embodiments, the one or more amino acids that interact with a PAM are located at one or more of the following positions: 176, 238, 447, and / or 563; and wherein the positions are numbered according to SEQ ID NO:92.
[0082] In some embodiments, wherein (1) the positively charged amino acid is R or K.
[0083] In some embodiments, the Cas12i nuclease comprises any one or combination of the following mutations: (1) E563R; (2) E176R, T447R, E176R, and E563R; (3) K238R and E563R; (4) E176R, K238R, and T447R; (5) E176R, K238R, and E563R; (6) E176R, T447R, and E563R; and / or (7) E176R, K238R, T447R, and E563R; and wherein the amino acid residues are numbered according to SEQ ID NO: 92.
[0084] In some embodiments, the one or more amino acids involved in opening the DNA double strand are located at one or more of the following positions: 163 and / or 164; and wherein the positions are numbered according to SEQ ID NO:92.
[0085] In some embodiments, the amino acid with an aromatic ring in (2) is F, Y or W.
[0086] In some embodiments, the one or more amino acids involved in opening the DNA double strand are replaced with amino acids with aromatic rings, namely: Q163F, Q163Y, Q163W, and / or N164F or N164Y; and the amino acid residues are numbered according to SEQ ID NO: 92.
[0087] In some embodiments, the one or more amino acids located in the RuvC domain and interacting with the single-stranded DNA substrate are located at one or more of the following positions: 323, 362, 425, 925, 926, 391, 424 and / or 929; and wherein the positions are numbered according to SEQ ID NO: 92.
[0088] In some embodiments, wherein (3) the positively charged amino acid is R or K.
[0089] In some embodiments, the Cas12i nuclease comprises any one or combination of the following mutations: (1) E323R; (2) D362R; (3) Q425R; (4) N925R; (5) I926R; (6) E323R and D362R; (7) E323R and Q425R; (8) E323R and I926R; (9) Q425R and I926R; (10) D362R and I926R; (11) N925R and I926R; (12) E323R, D362R and Q425R; (13) E323R, D362R and I926R; (14) E323R, Q425R and I926R; (15) D362R, N925R and I926R; and / or (16) E323R, D362R, Q425R and I926R; and wherein the amino acid residues are numbered according to SEQ ID NO:92.
[0090] In some embodiments, the one or more amino acids that interact with the DNA-RNA double helix are located at one or more of the following positions: 116, 117, 159, 161, 319, 343 and / or 958; and wherein the positions are numbered according to SEQ ID NO:92.
[0091] In some embodiments, wherein (4) the positively charged amino acid is R or K.
[0092] In some embodiments, the Cas12i nuclease comprises any one or combination of the following mutations: G116R, E117R, T159R, S161R, E319R, E343R, and / or D958R; and wherein the amino acid residues are numbered according to SEQ ID NO: 92.
[0093] In some embodiments, the Cas12i nuclease further comprises one or more mutations in the flexible region, which can increase the flexibility of the flexible region of the Cas12i nuclease compared to the wild-type Cas12i2 nuclease.
[0094] In some embodiments, the flexible region is selected from the group corresponding to amino acid residues 228-232, amino acid residues 439-443, amino acid residues 478-482, amino acid residues 500-504, amino acid residues 775-779, and amino acid residues 925-929; and wherein the amino acid residues are numbered according to SEQ ID NO:92.
[0095] In some embodiments, the mutation of the one or more flexible regions comprises insertion of one or more G residues in the flexible region.
[0096] In some embodiments, the one or more G residues are inserted at the N-terminus of the flexible amino acid residues in the flexible region, wherein the flexible amino acid residues are selected from the group consisting of G, S, N, D, H, M, T, E, Q, K, R, A, and P.
[0097] In some embodiments, the mutation of one or more flexible regions comprises replacing a hydrophobic amino acid residue in the flexible region with a G residue, wherein the hydrophobic amino acid residue is selected from the group consisting of L, I, V, C, Y, F, and W.
[0098] In some embodiments, the mutation in one or more flexible regions is: (1) I926G; and / or (2) 439G or 439GG, and wherein the amino acid residues are numbered according to SEQ ID NO:92.
[0099] In some embodiments, the Cas12i nuclease comprises the mutations: E176R, K238R, T447R, E563R, N164Y, E323R, and D362R; and wherein the amino acid residues are numbered according to SEQ ID NO: 92.
[0100] In some embodiments, the gRNA molecule comprises the nucleotide sequence of any one of SEQ ID NOs: 67-78.
[0101] In some embodiments, the N-terminus and / or C-terminus of the nuclease is linked to one or more nuclear localization signals.
[0102] In some embodiments, the nuclear localization signal is linked to a His tag.
[0103] In some embodiments, the nuclease comprises an amino acid sequence selected from any one of SEQ ID NOs: 1-11.
[0104] In some embodiments, the composition optionally further comprises a pharmaceutically acceptable carrier.
[0105] On the other hand, the present application provides a cell comprising the gRNA molecule described in the present application, the nucleic acid molecule described in the present application, the vector described in the present application, the ribonucleoprotein (RNP) complex described in the present application, and / or the composition described in the present application.
[0106] In another aspect, the present application provides a cell population or progeny thereof, comprising a gene product produced by administering the ribonucleoprotein (RNP) complex, the composition, or the cell described herein.
[0107] In some embodiments, the gene product comprises a modified HBG gene.
[0108] In some embodiments, the modification is a base insertion or deletion (indel) located at or near the target sequence; wherein the target sequence is located in the HBG1 promoter region or the HBG2 promoter region, the HBG1 promoter region is located between chr11:5,248,269 and 5,249,857 of the human genome hg19, and the HBG2 promoter region is located between chr11:5,253,188 and 5,254,781 of the human genome hg19.
[0109] In some embodiments, the target sequence is located between about 210 bp upstream to about 100 bp upstream of the transcription start site of the HBG1 and the HBG2.
[0110] In some embodiments, the cell population is capable of differentiating into differentiated cells of the erythroid lineage, and the differentiated cells have increased expression levels of fetal hemoglobin compared to a cell population without modification of the HBG gene.
[0111] On the other hand, the present application provides a kit comprising the gRNA molecule described in the present application, the nucleic acid molecule described in the present application, the vector described in the present application, the ribonucleoprotein (RNP) complex described in the present application, the composition described in the present application, the cell described in the present application, and / or the cell population or its progeny described in the present application.
[0112] On the other hand, the present application provides a method for regulating the expression of fetal hemoglobin (HbF) in a cell, a cell population, or their progeny, the method comprising administering the ribonucleoprotein (RNP) complex described herein, the composition described herein, the cell described herein, the cell population or their progeny described herein, and / or the kit described herein.
[0113] In some embodiments, the cell or cell population is a hematopoietic stem / progenitor cell (HSPC).
[0114] In some embodiments, the cell or cell population is a CD34+ HSPC.
[0115] On the other hand, the present application provides a method for treating, preventing or alleviating hemoglobinopathy and / or its related disorders, comprising administering to a subject in need thereof an effective amount of the ribonucleoprotein (RNP) complex described herein, the composition described herein, the cell described herein, the cell population described herein or its progeny, and / or the kit described herein.
[0116] In some embodiments, the hemoglobin and / or its associated disorder is selected from the group consisting of sickle cell disease, sickle cell anemia, hemoglobin C disease, hemoglobin C trait, hemoglobin S / C disease, hemoglobin D disease, hemoglobin E disease, thalassemia, hypooxygenated hemoglobinopathy, and unstable hemoglobinopathy.
[0117] On the other hand, the present application provides the use of the gRNA molecules described herein, the nucleic acid molecules described herein, the vectors described herein, the ribonucleoprotein (RNP) complexes described herein, the compositions described herein, the cells described herein, and / or the cell populations described herein or their progeny for preparing a drug for treating hemoglobinopathies and / or related disorders thereof.
[0118] On the other hand, the present application provides the gRNA molecules described herein, the nucleic acid molecules described herein, the vectors described herein, the ribonucleoprotein (RNP) complexes described herein, the compositions described herein, the cells described herein, the cell populations or their progeny described herein, and / or the kits described herein, for use in treating hemoglobinopathies and / or their related disorders.
[0119] Those skilled in the art can easily discern other aspects and advantages of the present application from the detailed description below. In the detailed description below, only exemplary embodiments of the present application are shown and described. As will be appreciated by those skilled in the art, the content of this application enables those skilled in the art to modify the disclosed specific embodiments without departing from the spirit and scope of the invention to which this application relates. Accordingly, the descriptions in the drawings and specification of this application are merely exemplary and not restrictive. BRIEF DESCRIPTION OF THE DRAWINGS
[0120] The specific features of the invention involved in this application are shown in the appended claims. The features and advantages of the invention involved in this application can be better understood by referring to the exemplary embodiments described in detail below and the accompanying drawings. A brief description of the drawings is as follows:
[0121] Figure 1A shows the distribution of key motifs in the HBG promoter region and the gRNA targeting locations. Figures 1B-1C show the screening results of gRNAs with a binding region length of 20 nt; Figures 1D-1E show the screening results of crRNAs and sgRNAs with different binding region lengths.
[0122] Figure 2A shows eight Cas12b Max Schematic diagram of the structure of the construct. Figures 2B-2C show the structure of Cas12b Max The editing efficiency and HbF+ erythrocyte induction efficiency of constructs M1-M7 under the guidance of sgRNA1. Figure 2D shows the Cas12b Max Figure 2E shows the editing efficiency of target genes in hematopoietic stem cells using constructs M1 and M2 compared to SpCas9 in the HBG promoter region.
[0123] Figures 3A-3C show the guidance of Cas12b Max Optimization scheme and screening results of sgRNA single-molecule backbone.
[0124] Figures 4A-4D show Cas12b Max -sgRNA G1 complex (RNP) editing of hematopoietic stem cells, screening results of RNP dose and Cas enzyme:sgRNA ratio, hematopoietic stem cell editing efficiency in healthy controls and thalassemia patients, distribution statistics of the main editing sequences in the target genomic region and the length of the editing sequences.
[0125] Figures 5A-5B show Cas12b Max Evaluation of off-target effects when editing hematopoietic stem cells with sgRNA G1 complex (RNP).
[0126] Figures 6A-6D show Cas12b Max -sgRNA G1 complex (RNP) induction results on the mRNA levels of γ-globin and fetal hemoglobin, the levels of HbF-positive red blood cells and the protein levels.
[0127] Figures 7A-7B show Cas12b Max -sgRNA G1 complex (RNP) affects the differentiation potential of hematopoietic stem cells when editing them in vitro and its effects on cell innate immunity.
[0128] Figures 8A-8F show that Cas12b Max In vivo efficacy evaluation of hematopoietic stem cells edited with sgRNA G1 complex (RNP). DETAILED DESCRIPTION
[0129] The following describes the implementation of the present invention through specific embodiments. People familiar with this technology can easily understand other advantages and effects of the present invention from the contents disclosed in this specification.
[0130] Definition of terms
[0131] In this application, the term "guide RNA" can be used interchangeably with "guide RNA (molecule)", "gRNA (molecule)", and generally refers to a group of nucleic acid molecules that facilitate the specific guidance of an RNA-guided nuclease or other effector molecule (usually in complex with a gRNA molecule) to a target sequence. In some embodiments, the guidance is achieved by hybridizing a portion of the gRNA with DNA (e.g., via a gRNA targeting domain) and by binding a portion of the gRNA molecule to an RNA-guided nuclease or other effector molecule (e.g., at least via tracr RNA). In certain embodiments, the gRNA molecule consists of a single continuous polynucleotide molecule, referred to herein as a "single-component guide RNA" or "sgRNA", etc. In other embodiments, the gRNA molecule consists of multiple, typically two, polynucleotide molecules that are themselves capable of associating, typically by hybridization, referred to herein as a "two-component guide RNA" or "dgRNA", etc. The gRNA molecule generally comprises a targeting domain and a tracr. In an embodiment, the targeting domain and tracr are disposed on a single polynucleotide. In other embodiments, the targeting domain and tracr are disposed on separate polynucleotides.
[0132] In the present application, the term "target sequence" generally refers to a nucleic acid sequence that is complementary to a gRNA targeting domain, such as a completely complementary nucleic acid sequence. In an embodiment, the target sequence is arranged on genomic DNA. In one embodiment, the target sequence is adjacent to a protospacer adjacent motif (PAM) sequence (e.g., a PAM sequence recognized by a Cas enzyme) recognized by a protein with nuclease or other effector activity (on the same chain of DNA or on the complementary chain of DNA). In an embodiment, the target sequence is a target sequence in a gene or locus that affects globin gene expression, such as β-globin or fetal hemoglobin (HbF) expression. In an embodiment, the target sequence is a target sequence in a non-deleted HPFH region, such as in an HBG1 and / or HBG2 promoter region. And in the present application, the term "targeting domain" (when the term is used in conjunction with gRNA) refers to a portion of a gRNA molecule that recognizes a target sequence (e.g., a target sequence in a cellular nucleic acid, such as in a gene), such as complementary thereto.
[0133] In this application, the term "crRNA" (when the term is used in conjunction with a gRNA molecule) generally refers to the portion of the gRNA molecule that comprises the targeting domain and the domain that interacts with tracr to form a dgRNA.
[0134] In this application, the term "tracr" (when used in conjunction with a gRNA molecule) generally refers to the portion of the gRNA that binds to a nuclease or other effector molecule. In an embodiment, tracr comprises a nucleic acid sequence that specifically binds to a Cas enzyme. In an embodiment, tracr comprises a nucleic acid sequence that forms part of a dgRNA.
[0135] The term "CRISPR system", "Cas system" or "CRISPR / Cas system" refers to a group of molecules comprising RNA-guided nucleases or other effector molecules and gRNA molecules, which together are necessary and sufficient to guide and achieve the modification of nucleic acids at target sequences by RNA-guided nucleases or other effector molecules. In one embodiment, the CRISPR system comprises gRNA and Cas protein (e.g., Cas9 protein), in which case such systems comprising Cas9 or modified Cas9 molecules are referred to as "Cas9 systems" or "CRISPR / Cas9 systems". In one example, the gRNA molecule and the Cas molecule can be complexed to form a ribonucleoprotein (RNP) complex.
[0136] In this application, the term "indel" generally refers to a nucleic acid comprising one or more nucleotide insertions, one or more nucleotide deletions, or a combination of nucleotide insertions and deletions relative to a reference nucleic acid, which is generated after exposure to a composition comprising a gRNA molecule (e.g., a CRISPR system). Insertions / deletions can be determined by sequencing the nucleic acid after exposure to a composition comprising a gRNA molecule, such as by NGS. With respect to the site of an indel, an indel is said to be located "at or near" a reference site (e.g., a site that is complementary to a targeting domain of a gRNA molecule) if the indel comprises at least one insertion or deletion within about 10, about 20, about 30, about 40, about 50, about 60, about 70, about 80, about 90, or about 100 nucleotides of the reference site, or the indel partially or completely overlaps with the reference site (e.g., comprises at least one insertion or deletion that overlaps by about 10, about 20, about 30, about 40, about 50, about 60, about 70, about 80, about 90, or about 100 nucleotides of a site that is complementary to a targeting domain of a gRNA molecule as described herein, or at least one insertion or deletion that is within about 10, about 20, about 30, about 40, about 50, about 60, about 70, about 80, about 90, or about 100 nucleotides of a site that is complementary to a targeting domain of a gRNA molecule as described herein). For example, the indels described herein may be insertions of 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides at or near the target sequence; or deletions of 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75 53, 54, 55, 56, 57, 58, 59, or 60 nucleotides.
[0137] In this application, the term "functionally active fragment" generally refers to a fragment that has a partial region of a full-length protein or nucleic acid but retains or partially retains the biological activity or function of the full-length protein or nucleic acid. For example, a functionally active fragment can retain or partially retain the ability of the full-length protein to bind to another molecule.
[0138] In this application, the term "nuclear localization sequence" or "nuclear localization signal" or "NLS" generally refers to a peptide that directs a protein to the cell nucleus. For example, an NLS can be located at any position on the peptide chain. NLSs can vary in length and / or sequence, and many specific NLS sequences have been described. For example, it has been found that peptides containing positively charged amino acids, particularly lysine (K), arginine (R) and / or histidine (H), can generally be used as NLSs. Thus, an exemplary NLS can be, for example, a peptide of 4 to 20 amino acids, more specifically 4 to 15, 4 to 12, 4 to 10 or 4 to 8 amino acids, wherein at least 4 amino acids (more specifically 60%, 70%, 75%, 80%, 85% or 90% of the amino acid residues in the NLS peptide) are positively charged amino acids, preferably selected from K, R or H. NLS sequences that can be used in conjunction with the nucleases and / or constructs thereof described herein are known in the art. Non-limiting examples of such NLS sequences include the nucleoplasmin NLS having the following amino acid sequence (KRPAATKKAGQAKKKK), the Simian Virus 40 "SV40" NLS (PKKKRKV), the c-Myc NLS, and / or the BP SV40 NLS.
[0139] In this application, the term "His tag" generally refers to an epitope tag consisting of 6 to 10 histidine residues that is easily recognized by tag-specific antibodies. The His tag is one of the commonly used tags for protein purification and detection. Due to its small molecular weight (only 0.84 kD), fusion to recombinant proteins does not affect the structure and biochemical properties of the tagged protein. His antibodies can be used to detect the expression and intracellular localization of His-tag fusion proteins, as well as to purify, qualitatively, or quantitatively detect His-fusion proteins, and are widely used.
[0140] In this application, the term "nucleic acid" is used interchangeably with "polynucleotide", "nucleotide", "nucleotide sequence" and "oligonucleotide" and generally refers to nucleotides (e.g., deoxyribonucleotides or ribonucleotides) and polymers thereof in single-stranded, double-stranded or multi-stranded form or their complements. For example, a nucleotide can be a ribonucleotide, a deoxyribonucleotide or a modified version thereof. For example, a nucleotide can be a single-stranded and double-stranded DNA, a single-stranded and double-stranded RNA, and a hybrid molecule having a mixture of single-stranded and double-stranded DNA and RNA. For example, a nucleotide can include, but is not limited to, any type of RNA, such as mRNA, siRNA, miRNA, sgRNA and guide RNA, and any type of DNA, genomic DNA, plasmid DNA and minicircle DNA, and any fragments thereof. The term also encompasses nucleic acids containing known nucleotide analogs or modified backbone residues or bonds, which are synthetic, naturally occurring, and non-naturally occurring.
[0141] As used herein, the term "sequence encoding" or "nucleic acid encoding" generally refers to a nucleic acid (RNA or DNA molecule) comprising a nucleotide sequence encoding a protein. The coding sequence may also include start and stop signals operably linked to regulatory elements, including a promoter and polyadenylation signal, capable of directing expression in the cells of an individual or mammal to which the nucleic acid is administered. The coding sequence may be codon-optimized.
[0142] In this application, the term "pharmaceutically acceptable carrier" generally refers to a carrier for administering therapeutic agents, such as antibodies or polypeptides, genes, and other therapeutic agents. The term refers to any pharmaceutical carrier that does not itself induce the production of antibodies harmful to the individual receiving the composition and that can be administered without excessive toxicity. Suitable carriers can be large, slowly metabolized macromolecules, such as proteins, polysaccharides, polylactic acids, polyglycolic acids, polyamino acids, amino acid copolymers, lipid aggregates, and inactivated viral particles. These carriers are well known to those skilled in the art. Pharmaceutically acceptable carriers in therapeutic compositions can include liquids such as water, saline, glycerol, and ethanol. Auxiliary substances, such as wetting agents or emulsifiers, pH buffering substances, etc., may also be present in these carriers.
[0143] In this application, the term "kit" generally refers to any collection of two or more components, which together constitute a functional unit that can be used for a special purpose. By way of illustration (and not limitation), a kit according to the present disclosure may include a guide RNA that is complexed with an RNA-guided nuclease or that is capable of complexing with the nuclease, and is accompanied by (e.g., suspended in, or suspendable in) a pharmaceutically acceptable carrier. The kit can be used to introduce the complex into, for example, a cell or a subject, for the purpose of causing a desired genomic change in such a cell or subject. The components of the kit can be packaged together, or the components can be packaged separately. The kit according to the present disclosure also optionally includes instructions for use, which describe, for example, the use of the kit according to the method of the present application. The instructions for use can be physically packaged with the kit, or the instructions for use can be made available to the user of the kit, for example, by electronic means.
[0144] In this application, the term "cell population" generally refers to eukaryotic mammalian (preferably human) cells isolated from a biological source (eg, a blood product or tissue) and derived from more than one cell type.
[0145] In this application, the terms "hematopoietic stem and progenitor cells" or "HSPCs" are used interchangeably and refer to a cell population comprising both hematopoietic stem cells ("HSCs") and hematopoietic progenitor cells ("HPCs"). Such cells are characterized as, for example, CD34+. In exemplary embodiments, HSPCs are isolated from bone marrow. In other exemplary embodiments, HSPCs are isolated from peripheral blood. In other exemplary embodiments, HSPCs are isolated from umbilical cord blood. In one embodiment, HSPCs are characterized as CD34+ / CD38- / CD90+ / CD45RA-. In an embodiment, HSPCs are characterized as CD34+ / CD90+ / CD49f+ cells. In an embodiment, HSPCs are characterized as CD34+ cells. In an embodiment, HSPCs are characterized as CD34+ / CD90+ cells. In an embodiment, HSPCs are characterized as CD34+ / CD90+ cells. In an embodiment, HSPCs are characterized as CD34+ / CD90+ cells.
[0146] In this application, the term "treating", such as a disease, means that when a complex, composition, cell, etc. as described herein is administered, for example, a subject (e.g., a person) who has the disease, is at risk of having the disease, and / or experiences symptoms of the disease will experience less severe symptoms and / or will recover faster in one embodiment than when the complex, composition, cell, etc. have never been administered. In some embodiments, "treating" may refer to reducing or alleviating the progression, severity, and / or duration of a condition, such as a hemoglobinopathy, or alleviating one or more symptoms (preferably, one or more discernible symptoms) of a condition, such as a hemoglobinopathy, caused by the administration of one or more therapies (e.g., one or more therapeutic agents, such as the gRNA molecules, CRISPR systems, or modified cells of the present invention). In specific embodiments, the term "treating" refers to improving at least one measurable physical parameter of a hemoglobinopathy condition that is not discernible to the patient. In other embodiments, the term "treating" refers to inhibiting the progression of a condition, such as physically, by stabilizing discernible symptoms, physiologically, by stabilizing physical parameters, or both. In other embodiments, the term "treating" refers to reducing or stabilizing the symptoms of a hemoglobinopathy, such as sickle cell disease or beta-thalassemia.
[0147] In this application, the term "subject" generally refers to an animal, typically a mammal, such as a human, non-human primate (apes, gibbons, gorillas, chimpanzees, orangutans, macaques), livestock (dogs and cats), farm animals (poultry such as chickens and ducks, horses, cattle, goats, sheep, pigs), and laboratory animals (mice, rats, rabbits, guinea pigs). Human subjects include fetuses, newborns, infants, adolescents, and adult subjects. Subjects include animal disease models, for example, mice and other animal models of blood coagulation diseases (such as HeA), and other animal models known to those skilled in the art.
[0148] In this application, the term "comprising" generally means including the features specifically stated, but not excluding other elements.
[0149] In this application, the term "selected from" generally refers to the selected objects and all combinations thereof. For example, "selected from (:) A, B and C" means all combinations of A, B and C, for example, A, B, C, A+B, A+C, B+C or A+B+C.
[0150] Without intending to be bound by any theory, the following examples are merely intended to illustrate the fusion protein, preparation method, and use of the present application, and are not intended to limit the scope of the present invention.
[0151] Example
[0152] Materials and methods
[0153] (1) Cell culture
[0154] Human CD34+ HSPCs were isolated from mobilized peripheral blood of deidentified healthy donors, and CD34+ HSPCs from β-thalassemia patients were isolated from mobilized peripheral blood of patients treated with plerixafor. CD34+ HSPCs were enriched using the Miltenyi CD34 Microbead Kit (Miltenyi Biotec). CD34+ HSPCs were isolated using StemSpan TM Electroporated HSPCs were maintained and expanded in SFEMII (Stemcell Technologies). Erythroid differentiation of the electroporated HSPCs was performed in vitro using a three-step serum-free suspension protocol. Erythroid phenotype and γ-globin induction were assessed on day 18.
[0155] Semi-solid cell culture for colony-forming unit assays: seeding CD34+ HSPCs into MethoCult® TM Medium (Stemcell Technologies) and cultured in culture dishes. On day 14 of semi-solid culture, HSPC-derived colonies were counted or picked for further evaluation.
[0156] (2) RNP electroporation
[0157] Chemically synthesized sgRNAs (the first three nucleotides and the last three nucleotides were modified with 2-O-methyl and 3' phosphorothioate) were purchased from IDT Technologies. AaCas12b variants were expressed and purified by Cellnuo Biotec (prokaryotic expression sequences shown in SEQ ID NOs: 79-86). During electroporation with 20 μl Nucleocuvette Strips, Cas12b (60 to 180 pmol) and sgRNA or dgRNA were mixed at a specific ratio and incubated at room temperature for 15 minutes before preparing RNP complexes. 5×10 4 to 5×10 5 CD34+ HSPCs were resuspended in 20 μl of P3 solution and mixed with RNPs. 20 μl of the HSPC / RNP mixture was transferred to a cuvette and electroporated using the DZ-100 program. The electroporated cells were transferred to SFEMII medium overnight before in vitro differentiation or transplantation experiments.
[0158] (3) Measuring Indel frequency using next-generation sequencing (NGS)
[0159] HSPC genomic DNA was extracted using a mouse tail DNA extraction kit (Beyotime, D7283M). The promoter regions upstream of the HBG1 and HBG2 transcription start sites were amplified in a two-step method using Phanta Max Super-Fidelity DNA Polymerase (Vazyme, P525-03) and primers. PCR products were sequenced by next-generation sequencing (NGS). Sequencing data were imported into the CRISPR RGEN Tools Cas-Analyzer online software for indel frequency measurement and indel pattern recognition using a 70-bp resolution window.
[0160] (4) RT-qPCR quantification
[0161] Total RNA from HSPCs was isolated using the FastPure Cell / Tissue Total RNA Isolation Kit V2 (Vazyme, RC112-01), reverse transcribed using the HiScript II 1st Strand cDNA Synthesis Kit (Vazyme, R212-01), and RT-qPCR was performed using the AceQ Universal SYBR qPCR Aster Mix (Vazyme, Q511-02).
[0162] (5) Measure HbF+ cells using flow cytometry
[0163] After staining with a human CD235a+ antibody (clone GA-R2 with PE; BD Pharmingen), in vitro differentiated cells were fixed with 4% paraformaldehyde for 10 minutes and then permeabilized with 0.1% Triton X-100 for 5 minutes at room temperature. Subsequently, cells were stained with a human HbF antibody (clone HBF-1 conjugated to APC; purchased from Invitrogen) for 30 minutes in the dark. Prior to FACS analysis, cells were washed twice to remove unbound antibody.
[0164] (6) Determination of hemoglobin by HPLC
[0165] Hemolysis was performed using a hemolysis reagent to prepare erythrocytes differentiated from CD34+ HSPCs in vitro. Lysates were analyzed using a D-10 hemoglobin analyzer (Bio-Rad) using clinically calibrated human hemoglobin as a standard.
[0166] Example 1
[0167] Screening of Cas12b / Cas12i constructs and their gRNA targeting binding regions
[0168] Cas12bMax or Cas12i2 Max The expression plasmid of the gene and the gRNA expression plasmid were co-transfected into HEK293 cells to screen out the gRNA that can effectively target the cis-elements located upstream of the HBG1 and HBG2 TSS to switch the expression of β / γ globin in mature red blood cells. Figure 1A shows the genomic location of the target site where the CRISPR-Cas12 DNA endonuclease generates HPFH-like mutations in the promoter region of HBG1 and HBG2 (human γ-globin locus), where the expected cutting site is marked with an arrow. Cas12b Max The sgRNA sequence screening results are shown in SEQ ID NOs: 12-59, and Figure 1B shows that Cas12b Max The gRNA was screened in the 293HEK cell line to target the HBG1 / 2 primordium region for efficient editing. Cas12b Max The bi-molecular gRNA consists of tracrRNA and crRNA, and their sequence screening results are shown in SEQ ID NOs: 60-66. Figure 1C shows that Cas12i Max Screening results of gRNAs for efficient editing of the HBG1 / 2 protomer region in 293HEK cell lines. Max The sgRNA sequence screening results are shown in SEQ ID NOs: 67-78.
[0169] In addition, to improve the editing efficiency of the target genome responsible for erythroid γ-globin expression, single-molecule gRNA or double-molecule gRNA with different spacer lengths generated by in vitro transcription were used to optimize RNP-mediated disruption of the LRF binding motif in the HBG1 / 2 promoter region in primary human CD34+ HSPCs. NGS data showed that the target indel frequency in RNP-electroporated HSPCs was related to the spacer length of the sgRNA (Figures 1D and 1E).
[0170] Further optimization introduced a group of engineered Cas protein constructs with different arrangements and combinations of NLS at the end. The amino acid sequences of these constructs are identified as SEQ ID NOs: 1-7. The experimental results showed that when RNPs containing different NLS arrangements were electroporated into HSPCs, their editing efficiency of the target genome varied, ranging from 30% to 85% at a dose of 3 μM (Figure 2B). Similar HbF induction amplitudes were detected by FACS between red blood cells differentiated in vitro from electroporated HSPCs (Figure 2C). This example also selected M1 and M2 constructs with higher editing efficiency for the target genome for comparison with SpCas9. The results showed that the RNP provided in this application was superior to the editing efficiency of Cas9 (Figure 2D). Based on the above protein-level results and RNA-level results, and combined with the comparison of the recovery rate of the Cas protein purification scheme, we further introduced the optimized Cas protein construct M8, whose amino acid sequence was identified as SEQ ID NO: 8. In this example, we compared the M1 construct with the new M8 construct. The results showed that, under different RNP electroporation dosages, the M1 and M8 RNPs had similar target gene editing efficiencies in the hematopoietic stem cells tested (Figure 2E).
[0171] Example 2
[0172] Cas12b Max Optimization and screening results of sgRNA single-molecule backbone
[0173] This example tests Cas12b targeting the HBG promoter region. Max Cas12b with different sgRNA backbones Max The cleavage activity of the complex in vitro (Figure 3A), as well as the terminal extension schemes of different Artsg13 backbone sequences, for Cas12b targeting the HBG promoter region Max The in vitro cleavage activity of the complex was affected differently (Figure 3B). The optimal Artsg13 backbone sequence end extension scheme (double-end extension of 113nt or single-end extension of 116nt) was finally obtained, and the editing efficiency was compared with the unoptimized scheme as shown in Figure 3C. The 110nt-nm sequence and 113nt sequence in Figure 3C are shown in SEQ ID NOs: 13 and 90 (the sequence shown in SEQ ID NO: 90 was chemically modified: mG*mG*mT*CGTCTATAGGACGGCGAGTTTTTCAACGGGTGTGCCAATGGCCACTTTCCAGGTGGCAAAGCCCGTTGAGCTTCAAAGAAGTGGCACAGATAGTGTGGGGAAGGGGC*mC*mC*mC).
[0174] Example 3
[0175] Selected Cas12b Max Gene editing of hematopoietic stem cells with sgRNA-G1
[0176] In the electroporation HSPCs experiment, the Cas:sgRNA molar ratio was set at 1:1 to 1:3, and the dependence of different RNP doses on editing efficiency was observed. The results showed that when the RNP dose was 180 pmol and the Cas:sgRNA molar ratio was 1:3, the indel frequency on the target LRF binding motif exceeded 85% (Figure 4A). NGS results showed that in the HSPCs from healthy donors and β-thalassemia (β 0 β 0 ) patients ( Figure 4B ), and showed that the most common indels were deletions longer than 10 bp, generated by microhomology-mediated end joining (MMEJ) repair ( Figures 4C and 4D ).
[0177] Example 4
[0178] Selected Cas12b Max Evaluation of off-target effects of gene editing with sgRNA-G1 on hematopoietic stem cells
[0179] To evaluate Cas12b Max -Specificity of efficient on-target editing by RNP. This example uses two unbiased whole-genome methods, namely modified Site-seq and Tag-seq, to screen naked genomic DNA and non-target sequences that are easily removed by RNP in the cell environment, respectively; and combined with bioinformatics prediction, a total of 4 potential off-target sites (OTS) were found to overlap with the list screened out by off-target sequences. These 4 potential off-target sites and 23 sites with highly similar on-target sequences were verified by deep sequencing of amplified fragments in highly edited CD34+HSPCs, and their targeted editing indel frequency was ≥90%. No RNP-mediated mutations were detected for each candidate OTS, and its indel frequency was lower than the NGS detection limit of 0.1% (Figure 5B). In summary, these data show that the CD34+HSPCs editing method provided in this application is highly specific to the target sequence and has no detected genetic or genomic-related toxicity.
[0180] Off-target effect analysis method:
[0181] Site-seq experiment: Briefly, purified genomic DNA was cleaved by Cas12b in the reaction buffer. Max-RNP digestion. The digested gDNA is hybridized with biotinylated, sticky-end, Illumina-compatible adapters and ligated. The adapter-ligated gDNA is fragmented, end-repaired, and ligated to a second adapter. After purification of the dual-adapter-ligated DNA, a two-step PCR is performed to construct a sequencing library. Sequencing data are analyzed using the Perl software package.
[0182] Tag-seq experiment: Cas12b was transferred to Max HEK293 cells were co-transfected with a plasmid encoding a guide RNA and a dsTag. Genomic DNA was purified, sheared to an average length of 300 bp, end-repaired, A-tailed, ligated to a half-adapter, and an 8-nt random molecular tag was added. Two rounds of nested anchored PCR were used to construct target-enriched sequencing libraries. Sequencing data were analyzed using an open-source Python package.
[0183] To validate off-target effects of HSPC editing by next-generation sequencing (NGS), a two-step PCR protocol was used to amplify potential off-target sites. Amplicons were sequenced using the MiSeq Sequencing System (Illumina) with 2×150 paired-end reads. Deep sequencing data were analyzed offline using CRISPResso software.
[0184] Example 5
[0185] Selected Cas12b Max Induction of γ-globin fetal hemoglobin with sgRNA-G1
[0186] To test whether disruption of the LRF binding motif in the HBG gene promoter region leads to significant γ-globin reactivation, HSPCs from healthy donors or β-thalassemia patients were electroporated with RNPs and then induced to mature erythrocytes in vitro. The results showed that compared with erythrocytes from unedited HSPCs, γ-globin mRNA levels were increased and β-globin mRNA levels were decreased in erythrocyte progeny from HSPCs with a high rate of HBG promoter LRF motif disruption (Figure 6A). HbF protein expression was analyzed by high-performance liquid chromatography (HPLC) or low-power cell counting. In in vitro generated erythrocytes, disruption of the HBG promoter LRF binding motif significantly increased the number of F cells (HbF / CD235a double positive) (Figure 6B). The results also observed that the HbF protein content in total hemoglobin extracted from erythrocytes with HBG promoter LRF motif disruption was significantly increased compared with controls from healthy donors (Figures 6C and 6D).
[0187] Example 6
[0188] Cas12b MaxOther effects of gene editing with sgRNA-G1 on hematopoietic stem cells
[0189] The multilineage differentiation potential of CD34+HSPCs electroporated with Cas12bMax-RNP was analyzed by colony-forming unit assay. The results in Figure 7A showed that the gene editing process had little effect on the differentiation potential of hematopoietic stem cells. In addition, the dynamics of P21 expression in CD34+HSPCs after electroporation with Cas12bMax-RNP were analyzed by RT-PCR, as shown in Figure 7B.
[0190] Example 7
[0191] Cas12b Max In vivo efficacy of hematopoietic stem cells edited with sgRNA-G1
[0192] from β-thalassemia major (β 0 β 0 ) patients were injected into immunodeficient NCG-X mice to test the effects of the editing method of the present application on their transplantation ability and blood cell differentiation potential. Immunodeficient NCG-X mice support hematopoietic stem cell transplantation without the need for myeloablative drug treatment. Compared with the natural CD34+ HSPCs transplant group, the RNP electroporated hCD34+ HSPCs transplant group had a lower chimerism rate of hCD45+ cells in the peripheral circulation after transplantation. At week 20, the levels of human lymphoid and myeloid (hCD45+) or erythrocyte (hCD235a+) reconstitution in the bone marrow of the two groups of subjects were similar (Figure 8A, Figure 8B and Figure 8D). The number of hCD34+ cells in the bone marrow of recipient mice that received RNP electroporated hCD34+ HSPCs was lower (Figure 8C). At week 20, hCD235+ cells in the bone marrow of mice infused with edited hCD34+ HSPCs induced significant HbF+ cells (Figure 8E). The frequency of targeted mutagenesis detected in human multilineage cells derived from peripheral blood or bone marrow was reduced compared to when the edited HSPCs were immediately infused (Fig. 8F).
[0193] Xenotransplantation of human CD34+ HSPCs:
[0194] NCG-X mice were purchased from GemPharatech Co., Ltd. Non-irradiated NCG-X female mice (4-5 weeks old) were injected with 0.5-0.6×10 6CD34+ HSPCs from healthy donors. For transplantation experiments, CD34+ HSPCs were electroporated with RNPs and cultured overnight. Twenty weeks after transplantation, bone marrow was isolated for human cell chimerism and HbF+ cell analysis. A secondary transplantation was performed via tail vein injection of hCD34+ HSPCs from the primary recipient's bone marrow. For flow cytometric analysis of bone marrow, bone marrow cells were incubated with the following: V450 mouse anti-human CD45 clone HI30 (560367, BD Biosciences), PE-eFluor 610 mCD45 monoclonal antibody (30-F11) (61-0451-82, Thermo Fisher), FITC anti-human CD235a antibody (349104, BioLegend), PE anti-human CD33 antibody (366608, BioLegend), APC anti-human CD19 antibody (302212, BioLegend), and the fixable viability dye eFluor 780 (650865-14, Thermo Fisher) for live / dead staining. The percentage of human engraftment was calculated as hCD45+ cells / (hCD45+ cells + mCD45+ cells) × 100. To assess the engraftment capacity of human CD34+ HSPCs, CD34+ HSPCs were incubated with anti-human CD34 antibodies (343512, Biolegend), anti-mouse CD34 (303508, Biolegend), and the dye eFluor 780 (650865-14, Thermo Fisher) for live / dead staining. For analysis of HbF+ cells, human CD235a+ cells were isolated from the bone marrow using the hCD234a+ microbead kit (Miltenyi). FACS analysis of HbF+ cells was performed using the same method as described above for flow cytometry measurement of HbF+ cells.
Claims
1. A guide RNA (gRNA) molecule, the gRNA molecule comprising a targeting domain complementary to a target sequence, wherein the target sequence is located in the HBG1 promoter region or in the HBG2 promoter region, the HBG1 promoter region is located between chr11:5,248,269 and 5,249,857 of the human genome hg19, and the HBG2 promoter region is located between chr11:5,253,188 and 5,254,781 of the human genome hg19.
2. According to the gRNA molecule according to claim 1, the target sequence is located between about 210bp upstream to about 100bp upstream of the transcription start site of the HBG1 and the HBG2.
3. The gRNA molecule according to any one of claims 1-2, comprising a tracr sequence and a crRNA sequence.
4. The gRNA molecule according to any one of claims 1-3, which is a two-component guide RNA molecule.
5. The gRNA molecule according to any one of claims 1 to 3, which is a single-component guide RNA molecule (sgRNA).
6. The gRNA molecule according to any one of claims 1 to 5, wherein (1) a CRISPR / Cas system comprising the gRNA molecule is contacted with the target sequence, or (2) a CRISPR / Cas system comprising the gRNA molecule is introduced into a cell where a gene comprising the target sequence is located, and a base insertion or deletion (indel) can be formed at or near the target sequence.
7. gRNA molecule according to claim 6, described CRISPR / Cas system comprises Cas12b nuclease and / or its functionally active fragment.
8. gRNA molecule according to claim 6, described CRISPR / Cas system comprises Cas12i nuclease and / or its functionally active fragment.
9. The gRNA molecule of claim 3, wherein the tracr sequence comprises the nucleotide sequence of SEQ ID NO:
60.
10. The gRNA molecule according to claim 3, wherein the crRNA sequence comprises a nucleotide sequence selected from any one of SEQ ID NOs: 61-66.
11. The gRNA molecule according to any one of claims 1-10, comprising a nucleotide sequence selected from any one of SEQ ID NOs: 12-59 and 67-78.
12. A nucleic acid molecule encoding the gRNA molecule according to any one of claims 1-11. A vector comprising the nucleic acid molecule according to claim 12 .
14. The vector according to claim 13, which is selected from the group consisting of a lentiviral vector, an adenoviral vector, an adeno-associated virus (AAV) vector, a herpes simplex virus (HSV) vector, a plasmid, a minicircle, a nanoplasmid and an RNA vector.
15. A ribonucleoprotein (RNP) complex comprising the gRNA molecule of any one of claims 1-11 and the nuclease of the CRISPR / Cas system.
16. The RNP complex according to claim 15, wherein the nuclease is a class II type V Cas enzyme.
17. The RNP complex according to claim 16, wherein the nuclease is a Cas12b nuclease.
18. according to the RNP complex described in claim 17, the Cas12b nuclease includes its variant, its ortholog and / or its functionally active fragment.
19. The RNP complex according to claim 17 or 18, wherein the Cas12b nuclease is a Cas12b nuclease (AaCas12b) from Alicyclobacillus acidiphilus.
20. The RNP complex according to claim 19, wherein the AaCas12b nuclease comprises one or more of the following mutations relative to the wild-type AaCas12b nuclease: (1) replacing one or more amino acid residues that interact with the pre-spacer adjacent motif (PAM) in the wild-type AaCas12b nuclease with positively charged amino acid residues; (2) replacing one or more amino acid residues involved in opening the DNA double helix in the wild-type AaCas12b nuclease with amino acid residues having an aromatic ring; and / or (3) replacing one or more amino acid residues in the RuvC domain of the wild-type AaCas12b nuclease that interact with the single-stranded DNA substrate with positively charged amino acid residues or hydrophobic amino acid residues, in, The amino acid sequence of the wild-type AaCas12b nuclease is as described in SEQ ID NO:
91.
21. The RNP complex of claim 20, wherein the one or more amino acid residues that interact with PAM are located at one or more of the following positions: 116, 123, 130, 132, 144, 145, 153, 173, 222, 395, 400, and 475; and wherein the amino acid residues are numbered according to SEQ ID NO:
91.
22. According to the RNP complex according to claim 20 or 21, the positively charged amino acid residue replacing one or more amino acid residues that interact with PAM in the wild-type AaCas12b nuclease is one or more of the following substitutions: D116R and E475R; and wherein the amino acid residues are numbered according to SEQ ID NO:
91.
23. The RNP complex according to claim 20, wherein the one or more amino acid residues involved in opening the DNA double strand are located at one or more of the following positions: 118, and 119; and wherein the amino acid residues are numbered according to SEQ ID NO:
91.
24. According to the RNP complex according to claim 20 or 23, the substitution of one or more amino acid residues involved in opening the DNA double strand in the wild-type AaCas12b nuclease with an amino acid residue having an aromatic ring is Q119Y, Q119F or Q119W substitution; and wherein the amino acid residues are numbered according to SEQ ID NO:
91.
25. The RNP complex of claim 20, wherein the one or more amino acid residues located in the RuvC domain and interacting with the single-stranded DNA substrate are located at one or more of the following positions: 300, 301, 304, 329, 636, 639, 647, 682, 757, 758, 761, 764, 768, 852, 854, 856, 857, 858, 860, 862, 863, 865, 866, 867, 869, 938, 956, 957, 958, 994, 1093 and 1097; and wherein the amino acid residues are numbered according to SEQ ID NO:
91.
26. The RNP complex of claim 20 or 25, wherein the one or more amino acid residues that are substituted for the wild-type AaCas12b nuclease in the RuvC domain and that interact with the single-stranded DNA substrate are one or more of the following substitutions: E636R, Q639R, T647R, Q682R, I757R, E758R, E761R, Q854R, N857K, D858R, I994R, Q1093R, and W1097R; and wherein the amino acid residues are numbered according to SEQ ID NO:
91.
27. The RNP complex according to any one of claims 17-26, wherein the gRNA molecule comprises a nucleotide sequence according to any one of SEQ ID NOs: 12-59.
28. The RNP complex according to claim 16, wherein the nuclease is a Cas12i nuclease.
29. The RNP complex of claim 28, wherein the Cas12i nuclease comprises a variant thereof, an ortholog thereof and / or a functionally active fragment thereof.
30. The RNP complex of claim 28 or 29, wherein the Cas12i nuclease comprises one or more of the following mutations relative to the wild-type Cas12i2 nuclease: (1) replacing one or more amino acids that interact with PAM in the wild-type Cas12i2 nuclease with positively charged amino acids; (2) replacing one or more amino acids involved in opening the DNA double strand in the wild-type Cas12i2 nuclease with amino acids with aromatic rings; (3) replacing one or more amino acids in the wild-type Cas12i2 nuclease located in the RuvC domain and interacting with the single-stranded DNA substrate with positively charged amino acids; and / or (4) replacing one or more amino acids in the wild-type Cas12i2 nuclease that interact with the DNA-RNA double helix with positively charged amino acids, in, The amino acid sequence of the wild-type Cas12i2 nuclease is as described in SEQ ID NO:
92.
31. The RNP complex of claim 30, wherein the one or more amino acids that interact with a PAM are located at one or more of the following positions: 176, 238, 447, and / or 563; and wherein the positions are numbered according to SEQ ID NO:
92.
32. The RNP complex according to claim 30, wherein (1) the positively charged amino acid is R or K.
33. The RNP complex of any one of claims 28-32, wherein the Cas12i nuclease comprises any one or combination of mutations: (1) E563R; (2) E176R, T447R, E176R and E563R; (3) K238R and E563R; (4) E176R, K238R and T447R; (5) E176R, K238R and E563R; (6) E176R, T447R and E563R; and / or (7) E176R, K238R, T447R and E563R; and wherein the amino acid residues are numbered according to SEQ ID NO:
92.
34. The RNP complex of claim 30, wherein the one or more amino acids involved in opening the DNA double strand are located at one or more of the following positions: 163 and / or 164; and wherein the positions are numbered according to SEQ ID NO:
92.
35. The RNP complex according to claim 30, wherein (2) the amino acid with an aromatic ring is F, Y or W.
36. The RNP complex according to claim 34 or 35, wherein the one or more amino acids involved in opening the DNA double helix are replaced with amino acids with aromatic rings, namely: Q163F, Q163Y, Q163W, and / or N164F or N164Y; and wherein the amino acid residues are numbered according to SEQ ID NO:
92.
37. The RNP complex of claim 30, wherein the one or more amino acids located in the RuvC domain and interacting with the single-stranded DNA substrate are located at one or more of the following positions: 323, 362, 425, 925, 926, 391, 424 and / or 929; and wherein the positions are numbered according to SEQ ID NO:
92.
38. The RNP complex according to claim 30, wherein (3) the positively charged amino acid is R or K.
39. The RNP complex of claim 37 or 38, wherein the Cas12i nuclease comprises any one of the following mutations or combinations of mutations: (1) E323R; (2) D362R; (3) Q425R; (4) N925R; (5) I926R; (6) E323R and D362R; (7) E323R and Q425R; (8) E323R and I926R; (9) Q425R and I926R; (10 )D362R and I926R; (11) N925R and I926R; (12) E323R, D362R and Q425R; (13) E323R, D362R and I926R; (14) E323R, Q425R and I926R; (15) D362R, N925R and I926R; and / or (16) E323R, D362R, Q425R and I926R; and wherein the amino acid residues are numbered according to SEQ ID NO:
92.
40. The RNP complex of claim 30, wherein the one or more amino acids that interact with the DNA-RNA double helix are located at one or more of the following positions: 116, 117, 159, 161, 319, 343 and / or 958; and wherein the positions are numbered according to SEQ ID NO:
92.
41. The RNP complex according to claim 30, wherein (4) the positively charged amino acid is R or K.
42. The RNP complex of claim 40 or 41, wherein the Cas12i nuclease comprises any one or combination of mutations: G116R, E117R, T159R, S161R, E319R, E343R, and / or D958R; and wherein the amino acid residues are numbered according to SEQ ID NO:
92.
43. According to the RNP complex of any one of claims 30-42, the Cas12i nuclease further comprises a mutation in one or more flexible regions, wherein the mutation enables the flexible region of the Cas12i nuclease to have increased flexibility compared to the wild-type Cas12i2 nuclease.
44. The RNP complex of claim 43, wherein the flexible region is selected from the group corresponding to amino acid residues 228-232, amino acid residues 439-443, amino acid residues 478-482, amino acid residues 500-504, amino acid residues 775-779, and amino acid residues 925-929; and wherein the amino acid residues are numbered according to SEQ ID NO:
92.
45. The RNP complex of claim 43 or 44, wherein the mutation of the one or more flexible regions comprises inserting one or more G residues in the flexible region.
46. The RNP complex of claim 45, wherein the one or more G residues are inserted at the N-terminus of the flexible amino acid residues in the flexible region, wherein the flexible amino acid residues are selected from the group consisting of G, S, N, D, H, M, T, E, Q, K, R, A and P.
47. The RNP complex of claim 43 or 44, wherein the mutation of the one or more flexible regions comprises replacing a hydrophobic amino acid residue in the flexible region with a G residue, wherein the hydrophobic amino acid residue is selected from the group consisting of L, I, V, C, Y, F and W.
48. The RNP complex according to any one of claims 43-47, wherein the mutations in the one or more flexible regions are: (1) I926G; and / or (2) 439G or 439GG, and wherein the amino acid residues are numbered according to SEQ ID NO:
92.
49. The RNP complex of any one of claims 43-48, wherein the Cas12i nuclease comprises mutations: E176R, K238R, T447R, E563R, N164Y, E323R, and D362R; and wherein the amino acid residues are numbered according to SEQ ID NO:
92.
50. According to the RNP complex described in any one of claims 28-49, the gRNA molecule comprises the nucleotide sequence described in any one of SEQ ID NOs:67-78.
51. The RNP complex according to any one of claims 15-49, wherein the N-terminus and / or C-terminus of the nuclease is linked to one or more nuclear localization signals.
52. The RNP complex of claim 51, wherein the nuclear localization signal is linked to a His tag.
53. The RNP complex of any one of claims 15-52, wherein the nuclease comprises an amino acid sequence selected from any one of SEQ ID NOs: 1-11.
54. A composition comprising: (1) one or more gRNA molecules according to any one of claims 1 to 11, and a nuclease of the CRISPR / Cas system; (2) a nucleic acid encoding one or more gRNA molecules according to any one of claims 1 to 11, and a nuclease of the CRISPR / Cas system; (3) one or more gRNA molecules as described in any one of claims 1 to 11, and a nucleic acid encoding a nuclease of the CRISPR / Cas system; or (4) A nucleic acid encoding one or more gRNA molecules as described in any one of claims 1 to 11, and a nucleic acid encoding a nuclease of the CRISPR / Cas system.
55. A composition according to claim 54, wherein the one or more gRNA molecules and the nuclease of the CRISPR / Cas system in (1) are present in a ribonucleoprotein (RNP) complex.
56. The composition of claim 55, wherein the nuclease is a class II type V Cas enzyme.
57. The composition of claim 56, wherein the nuclease is a Cas12b nuclease.
58. according to claim 57 compositions, described Cas12b nuclease comprises its variant, its orthologue and / or its functionally active fragment.
59. according to the composition described in claim 57 or 58, described Cas12b nuclease is the Cas12b nuclease (AaCas12b) from Alicyclobacillus acidiphilus.
60. The composition of claim 59, wherein the AaCas12b comprises one or more of the following mutations relative to the wild-type AaCas12b nuclease: (1) replacing one or more amino acid residues that interact with the pre-spacer adjacent motif (PAM) in the wild-type AaCas12b nuclease with positively charged amino acid residues; (2) replacing one or more amino acid residues involved in opening the DNA double helix in the wild-type AaCas12b nuclease with amino acid residues having an aromatic ring; and / or (3) replacing one or more amino acid residues in the RuvC domain of the wild-type AaCas12b nuclease that interact with the single-stranded DNA substrate with positively charged amino acid residues or hydrophobic amino acid residues, in, The amino acid sequence of the wild-type AaCas12b nuclease is as described in SEQ ID NO:
91.
61. The composition of claim 60, wherein the one or more amino acid residues that interact with a PAM are located at one or more of the following positions: 116, 123, 130, 132, 144, 145, 153, 173, 222, 395, 400, and 475; and wherein the amino acid residues are numbered according to SEQ ID NO:
91.
62. According to the composition of claim 60 or 61, the positively charged amino acid residue replacing one or more amino acid residues that interact with PAM in the wild-type AaCas12b nuclease is one or more of the following substitutions: D116R and E475R; and wherein the amino acid residues are numbered according to SEQ ID NO:
91.
63. The composition of claim 60, wherein the one or more amino acid residues involved in opening the DNA double strand are located at one or more of the following positions: 118, and 119; and wherein the amino acid residues are numbered according to SEQ ID NO:
91.
64. according to the composition described in claim 60 or 63, the one or more amino acid residues involved in opening the DNA double-stranded in the wild-type AaCas12b nuclease replaced with an amino acid residue having an aromatic ring are Q119Y, Q119F or Q119W substitutions; and wherein the amino acid residues are numbered according to SEQ ID NO:
91.
65. The composition of claim 60, wherein the one or more amino acid residues located in the RuvC domain and interacting with the single-stranded DNA substrate are located at one or more of the following positions: 300, 301, 304, 329, 636, 639, 647, 682, 757, 758, 761, 764, 768, 852, 854, 856, 857, 858, 860, 862, 863, 865, 866, 867, 869, 938, 956, 957, 958, 994, 1093, and 1097; and wherein the amino acid residues are numbered according to SEQ ID NO:
91.
66. according to the composition described in claim 60 or 65, described substitution one or more amino acid residues in described wild-type AaCas12b nuclease that are located in RuvC domain and interact with single-stranded DNA substrate are one or more of following substitutions: E636R, Q639R, T647R, Q682R, I757R, E758R, E761R, Q854R, N857K, D858R, I994R, Q1093R and W1097R; and wherein said amino acid residue is numbered according to SEQ ID NO:
91.
67. A composition according to any one of claims 57-66, wherein the gRNA molecule comprises a nucleotide sequence according to any one of SEQ ID NOs: 12-59.
68. The composition of claim 56, wherein the nuclease is a Cas12i nuclease.
69. according to the composition described in claim 68, described Cas12i nuclease comprises its variant, its ortholog and / or its functionally active fragment.
70. according to the composition described in claim 68 or 69, the Cas12i nuclease comprises one or more of the following mutations relative to the wild-type Cas12i2 nuclease: (1) replacing one or more amino acids that interact with PAM in the wild-type Cas12i2 nuclease with positively charged amino acids; (2) replacing one or more amino acids involved in opening the DNA double strand in the wild-type Cas12i2 nuclease with amino acids with aromatic rings; (3) replacing one or more amino acids in the wild-type Cas12i2 nuclease located in the RuvC domain and interacting with the single-stranded DNA substrate with positively charged amino acids; and / or (4) replacing one or more amino acids in the wild-type Cas12i2 nuclease that interact with the DNA-RNA double helix with positively charged amino acids, in, The amino acid sequence of the wild-type Cas12i2 nuclease is as described in SEQ ID NO:
92.
71. The composition of claim 70, wherein the one or more amino acids that interact with a PAM are located at one or more of the following positions: 176, 238, 447, and / or 563; and wherein the positions are numbered according to SEQ ID NO:
92.
72. The composition of claim 70, wherein (1) the positively charged amino acid is R or K.
73. A composition according to claim 71 or 72, wherein the Cas12i nuclease comprises any one or combination of mutations: (1) E563R; (2) E176R, T447R, E176R and E563R; (3) K238R and E563R; (4) E176R, K238R and T447R; (5) E176R, K238R and E563R; (6) E176R, T447R and E563R; and / or (7) E176R, K238R, T447R and E563R; and wherein the amino acid residues are numbered according to SEQ ID NO:
92.
74. The composition of claim 70, wherein the one or more amino acids involved in opening the DNA double strand are located at one or more of the following positions: 163 and / or 164; and wherein the positions are numbered according to SEQ ID NO:
92.
75. The composition according to claim 70, wherein (2) the amino acid with an aromatic ring is F, Y or W.
76. The composition of claim 74 or 75, wherein the one or more amino acids involved in opening the DNA double helix are replaced with amino acids with aromatic rings, namely: Q163F, Q163Y, Q163W, and / or N164F or N164Y; and wherein the amino acid residues are numbered according to SEQ ID NO:
92.
77. The composition of claim 70, wherein the one or more amino acids located in the RuvC domain and interacting with a single-stranded DNA substrate are located at one or more of the following positions: 323, 362, 425, 925, 926, 391, 424 and / or 929; and wherein the positions are numbered according to SEQ ID NO:
92.
78. The composition of claim 70, wherein (3) the positively charged amino acid is R or K.
79. The composition of claim 77 or 78, wherein the Cas12i nuclease comprises any one or combination of the following mutations: (1) E323R; (2) D362R; (3) Q425R; (4) N925R; (5) I926R; (6) E323R and D362R; (7) E323R and Q425R; (8) E323R and I926R; (9) Q425R and I926R; (10) D362R and I926R; (11) N925R and I926R; (12) E323R, D362R and Q425R; (13) E323R, D362R and I926R; (14) E323R, Q425R and I926R; (15) D362R, N925R and I926R; and / or (16) E323R, D362R, Q425R and I926R; and wherein the amino acid residues are numbered according to SEQ ID NO:
92.
80. The composition of claim 70, wherein the one or more amino acids that interact with the DNA-RNA double helix are located at one or more of the following positions: 116, 117, 159, 161, 319, 343 and / or 958; and wherein the positions are numbered according to SEQ ID NO:
92.
81. The composition of claim 70, wherein (4) the positively charged amino acid is R or K.
82. according to the composition described in claim 80 or 81, wherein the Cas12i nuclease comprises any one or combination of mutations: G116R, E117R, T159R, S161R, E319R, E343R and / or D958R; and wherein the amino acid residues are numbered according to SEQ ID NO:
92.
83. According to the composition of any one of claims 70-82, the Cas12i nuclease further comprises a mutation in one or more flexible regions, wherein the mutation enables the flexible region of the Cas12i nuclease to have increased flexibility compared to the wild-type Cas12i2 nuclease.
84. The composition of claim 83, wherein the flexible region is selected from the group corresponding to amino acid residues 228-232, amino acid residues 439-443, amino acid residues 478-482, amino acid residues 500-504, amino acid residues 775-779, and amino acid residues 925-929; and wherein the amino acid residues are numbered according to SEQ ID NO:
92.
85. The composition of claim 83 or 84, wherein the mutation of the one or more flexible regions comprises insertion of one or more G residues in the flexible region.
86. The composition of claim 85, wherein the one or more G residues are inserted at the N-terminus of the flexible amino acid residues in the flexible region, wherein the flexible amino acid residues are selected from the group consisting of G, S, N, D, H, M, T, E, Q, K, R, A and P.
87. The composition of claim 85 or 86, wherein the mutation of the one or more flexible regions comprises replacing a hydrophobic amino acid residue in the flexible region with a G residue, wherein the hydrophobic amino acid residue is selected from the group consisting of L, I, V, C, Y, F and W.
88. The composition of any one of claims 83-87, wherein the mutation in the one or more flexible regions is: (1) I926G; and / or (2) 439G or 439GG, and wherein the amino acid residues are numbered according to SEQ ID NO:
92.
89. according to the composition described in any one of claims 68-88, described Cas12i nuclease comprises mutation: E176R, K238R, T447R, E563R, N164Y, E323R and D362R; and wherein amino acid residues are numbered according to SEQ ID NO:
92.
90. A composition according to any one of claims 68-89, wherein the gRNA molecule comprises a nucleotide sequence according to any one of SEQ ID NOs:67-78.
91. The composition of any one of claims 54-90, wherein the N-terminus and / or C-terminus of the nuclease is linked to one or more nuclear localization signals.
92. The composition of claim 91, wherein the nuclear localization signal is linked to a His tag.
93. The composition of any one of claims 54-92, wherein the nuclease comprises an amino acid sequence selected from any one of SEQ ID NOs: 1-11.
94. The composition of any one of claims 54-93, further optionally comprising a pharmaceutically acceptable carrier.
95. A cell comprising a gRNA molecule as described in any one of claims 1-11, a nucleic acid molecule as described in claim 12, a vector as described in any one of claims 13-14, a ribonucleoprotein (RNP) complex as described in any one of claims 15-53, and / or a composition as described in any one of claims 54-94.
96. A cell population or progeny thereof comprising a gene product produced by administering a ribonucleoprotein (RNP) complex as described in any one of claims 15-53, a composition as described in any one of claims 54-94, or a cell as described in claim 95.
97. The cell population or progeny thereof of claim 96, wherein the gene product comprises a modified HBG gene.
98. According to the cell population or its progeny described in claim 97, the modification is a base insertion or deletion (indel) located at or near the target sequence; wherein the target sequence is located in the HBG1 promoter region or in the HBG2 promoter region, the HBG1 promoter region is located between chr11:5,248,269 and 5,249,857 of the human genome hg19, and the HBG2 promoter region is located between chr11:5,253,188-5,254,781 of the human genome hg19.
99. The cell population or progeny thereof of claim 98, wherein the target sequence is located between about 210 bp upstream to about 100 bp upstream of the transcription start site of the HBG1 and the HBG2.
100. The cell population or progeny thereof according to any one of claims 96-99, wherein the cell population is capable of differentiating into differentiated cells of the erythroid cell lineage, and wherein the differentiated cells have increased fetal hemoglobin expression levels compared to a cell population in which the HBG gene is not modified.
101. A kit comprising the gRNA molecule described in any one of claims 1-11, the nucleic acid molecule described in claim 12, the vector described in any one of claims 13-14, the ribonucleoprotein (RNP) complex described in any one of claims 15-53, the composition described in any one of claims 54-94, the cell described in claim 95, and / or the cell population or its progeny described in any one of claims 96-100.
102. A method for regulating the expression of fetal hemoglobin (HbF) in a cell, a cell population, or their progeny, the method comprising administering a ribonucleoprotein (RNP) complex as described in any one of claims 15-53, a composition as described in any one of claims 54-94, a cell as described in claim 95, a cell population or their progeny as described in any one of claims 96-100, and / or a kit as described in claim 101.
103. The method of claim 102, wherein the cell or cell population is a hematopoietic stem / progenitor cell (HSPC).
104. The method of claim 103, wherein the cell or cell population is a CD34+ HSPC.
105. A method for treating, preventing or alleviating hemoglobinopathy and / or its related disorders, the method comprising administering to a subject in need thereof an effective amount of a ribonucleoprotein (RNP) complex as described in any one of claims 15-53, a composition as described in any one of claims 54-94, a cell as described in claim 95, a cell population as described in any one of claims 96-100 or its progeny, and / or a kit as described in claim 101.
106. The method of claim 105, wherein the hemoglobin and / or its related disorders are selected from the group consisting of sickle cell disease, sickle cell anemia, hemoglobin C disease, hemoglobin C trait, hemoglobin S / C disease, hemoglobin D disease, hemoglobin E disease, thalassemia, hypooxygenated hemoglobinopathy, and unstable hemoglobinopathy.