Method and composition for activating zeta-globin gene expression
By artificially forming specific sequence enhancer elements in the non-coding region of the ζ-globin gene, and activating ζ-globin gene expression using technologies such as CRISPR-Cas, the safety and efficacy issues of α-thalassemia treatment in existing technologies have been resolved, and the risk of gene integration has been reduced.
Patent Information
- Application Number
- CN202510613958.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-05-17
- Filing Date
- 2025-05-13
- Publication Date
- 2025-11-18
AI Technical Summary
There is a lack of effective and safe gene therapy for alpha thalassemia in the current technology, especially given the possibility that the gene copy may be randomly integrated into different sites in the genome, leading to insertion mutations and malignant tumors.
Enhancer elements containing specific sequences are artificially created in the non-coding region of the ζ-globin gene using gene editing technology. Homologous recombination repair, insertion/deletion mutations are then performed using CRISPR-Cas, TALEN, ZFN, or Argonaute editing technologies to activate ζ-globin gene expression.
Safe and effective expression of the ζ-globin gene was achieved, improving the safety and efficacy of α-thalassemia treatment and reducing the risk of gene integration.
Smart Images

Figure CN120966920A_ABST
Abstract
Description
[0001] This application claims priority to Chinese patent application CN2024106186116 with the filing date of May 17, 2024. This application incorporates the entire text of the aforementioned Chinese patent application. TECHNICAL FIELD
[0002] The present application belongs to the field of biotechnology, and specifically relates to a method and composition for activating the expression of zeta-globin gene. BACKGROUND
[0003] Thalassemia is a hereditary hemolytic anemia. According to the gene typing, clinically, there are mainly alpha thalassemia and beta thalassemia. Alpha thalassemia is caused by deletion or mutation of alpha gene located on chromosome 16, which reduces the synthesis of alpha globin chain and causes hemoglobin (Hb) synthesis disorder. When there is only one alpha gene deletion or mutation, the patient usually has no obvious symptoms; when there are two alpha gene deletions or mutations, the patient can have mild microcytic hypochromic anemia. When there are three alpha gene deletions or mutations, i.e., the intermediate type of alpha thalassemia, also known as HbH disease, the patient can have mild to moderate anemia, among which the non-deletion type of alpha thalassemia such as Hb-CS type patient has the relatively most serious condition, low hemoglobin, and low survival rate. When four alpha genes are completely deleted, it leads to severe alpha thalassemia, also known as Hb Bart's fetal hydrops syndrome, and the fetus usually dies in the uterus or within a short period after birth.
[0004] Currently, the treatment strategies for moderate and severe alpha thalassemia mainly include regular blood transfusion and iron removal therapy, splenectomy, drug therapy, and hematopoietic stem cell transplantation. Regular blood transfusion and iron removal therapy can effectively alleviate the symptoms caused by anemia by infusing normal red blood cells or plasma to provide normal hemoglobin and improve anemia, but it cannot cure the disease. Splenectomy can reduce the destruction of red blood cells by removing the spleen, prolong the life of red blood cells, and improve anemia, but it also cannot fundamentally cure the disease. Drug therapy such as roxadustat also cannot be cured, and it needs to be taken continuously, and its effectiveness is limited, i.e., even if it is effective, it can only reduce the frequency of blood transfusion. Hematopoietic stem cell transplantation is currently the effective means and only way to cure severe thalassemia, but the lack of HLA-matched healthy donors, immune complications, and viral vector safety issues limit its use.
[0005] With the continuous progress of gene technology, gene therapy has made some important developments, and there are also some gene therapy methods for alpha thalassemia, such as overexpression of alpha globin, zeta globin, etc. However, such methods have significant safety problems, such as random integration of overexpressed globin gene copies into thousands of different sites in the genome, which may lead to insertional mutations and the possibility of malignancy. SUMMARY
[0006] To solve the technical problem of lacking an effective and safe gene therapy method for treating alpha thalassemia in the prior art, the present application provides a method and composition for activating expression of a zeta-globin gene.
[0007] To solve the above technical problem, one of the technical solutions provided by the present application is: a method for activating expression of a zeta-globin gene, the method comprising: using a gene editing technology to artificially form an enhancer element comprising a NTG-N(7-8)-WGATAR sequence, a NAG-N(7-8)-WGATAR sequence, a YTATCW-N(7-8)-CAN sequence and / or a YTATCW-N(7-8)-CTN sequence in the sense strand or the antisense strand of the non-coding region of the zeta-globin gene;
[0008] The N is A, G, C or T, W is T or A, R is A or G, and Y is T or C.
[0009] In a specific embodiment of the present application, the method is not for diagnostic / therapeutic purposes.
[0010] In a specific embodiment of the present application, the way of forming the enhancer element is selected from one or more of the following: deletion of a sequence, insertion of a sequence and replacement of a sequence.
[0011] In a specific embodiment of the present application, the gene editing technology is used to achieve homologous recombination repair and / or insertion / deletion mutation; the gene editing technology is selected from one or more of the following: CRISPR-Cas editing, TALEN editing, ZFN editing and Argonaute editing.
[0012] In a specific embodiment of the present application, the CRISPR-Cas editing is selected from one or more of the following: homologous recombination repair editing, Indel editing, prime editing and single base editing.
[0013] In a specific embodiment of the present application, the non-coding region of the zeta-globin gene is selected from one or more of the following: a sequence within 2 kb upstream of the transcription start site, a 5' UTR region, an intron region, a 3' UTR region and a sequence within 1 kb downstream of the 3' UTR region of the zeta-globin gene.
[0014] In a specific embodiment of the present application, the non-coding region of the zeta-globin gene comprises: a sequence within 2 kb upstream of the transcription start site and a 5' UTR region.
[0015] In a specific embodiment of the present application, the non-coding region of the zeta-globin gene comprises: a promoter region and a 5' UTR region of the zeta-globin gene.
[0016] In particular embodiments of the application, the sequence of the non-coding region of the zeta-globin gene is located at Chr 16: 149855-152909.
[0017] In particular embodiments of the application, the non-coding region of the zeta-globin gene is selected from one or more of the sequences located at Chr 16: 150752-151071, Chr 16: 152474-152793, Chr 16: 152522-152841, Chr 16: 152578-152897, Chr 16: 152487-152806, Chr 16: 152428-152747, Chr 16: 152399-152718, Chr 16: 152253-152572, Chr 16: 152178-152498, and Chr 16: 152444-152758.
[0018] In particular embodiments of the application, the non-coding region of the zeta-globin gene is selected from one or more of the sequences located at Chr 16: 150852-150971, Chr 16: 152574-152693, Chr 16: 152622-152741, Chr 16: 152678-152797, Chr 16: 152587-152706, Chr 16: 152528-152647, Chr 16: 152499-152618, Chr 16: 152353-152472, Chr 16: 152278-152398, and Chr 16: 152544-152658.
[0019] In particular embodiments of the application, the non-coding region of the zeta-globin gene is selected from one or more of the sequences located at Chr 16: 150892-150921, Chr 16: 152615-152644, Chr 16: 152671-152700, Chr 16: 152731-152760, Chr 16: 152629-152658, Chr 16: 152570-152599, Chr 16: 152550-152579, Chr 16: 152401-152430, Chr 16: 152323-152352, and Chr 16: 152587-152616.
[0020] In particular embodiments of the application, the non-coding region of the zeta-globin gene is selected from one or more of the sequences located at: Chr 16: 150901-150916, Chr 16: 152620-152636, Chr 16: 152680-152696, Chr 16: 152737-152753, Chr 16: 152636-152651, Chr 16: 152577-152592, Chr 16: 152557-152573, Chr 16: 152408-152424, Chr 16: 152330-152347, and Chr 16: 152593-152609.
[0021] In particular embodiments of the application, the non-coding region of the a- globin gene is selected from one or more of the sequences located at: Chr 16: 152474-152793, Chr 16: 152494-152793, Chr 16: 152514-152793, Chr 16: 152534-152793, Chr 16: 152554-152793, Chr 16: 152574-152793, Chr 16: 152594-152793, Chr 16: 152614-152793, Chr 16: 152634-152793, Chr 16: 152654-152793, Chr 16: 152674-152793, Chr 16: 152694-152793, Chr 16: 152714-152793, Chr 16: 152734-152793, Chr 16: 152754-152793, Chr 16: 152774-152793, Chr 16: 152593-152609, Chr 16: 152620-152636, Chr 16: 152680-152696, Chr 16: 152737-152753, Chr 16: 152474-152773, Chr 16: 152474-152753, Chr 16: 152474-152733, Chr 16: 152474-152713, Chr 16: 152474-152693, Chr 16: 152474-152673, Chr 16: 152474-152653, Chr 16: 152474-152633, Chr 16: 152474-152613, Chr 16: 152474-152593, Chr 16: 152474-152573, Chr 16: 152474-152553, Chr 16: 152474-152533, Chr 16: 152474-152513, and Chr 16: 152474-152493.
[0022] In particular embodiments of the application, the non-coding region of the zeta-globin gene is selected from one or more of the sequences located at: Chr 16: 152565-152645, Chr 16: 152568-152645, Chr 16: 152571-152645, Chr 16: 152574-152645, Chr 16: 152577-152645, Chr 16: 152580-152645, Chr 16: 152583-152645, Chr 16: 152586-152645, Chr 16: 152589-152645, Chr 16: 152592-152645, Chr 16: 152595-152645, Chr 16: 152598-152645, Chr 16: 152601-152645, Chr 16: 152604-152645, Chr 16: 152607-152645, Chr 16: 152610-152645, Chr 16: 152613-152645, Chr 16: 152616-152645, Chr 16: 152619-152645, Chr 16: 152622-152645, Chr 16: 152625-152645, Chr 16: 152628-152645, Chr 16: 152577-152592, Chr 16: 152593-152609, Chr 16: 152620-152636, Chr 16: 152565-152642, Chr 16: 152565-152639, Chr 16: 152565-152636, Chr 16: 152565-152633, Chr 16: 152565-152630, Chr 16: 152565-152627, Chr 16: 152565-152624, Chr 16: 152565-152621, Chr 16: 152565-152618, Chr 16: 152565-152615, Chr 16: 152565-152612, Chr 16: 152565-152609, Chr 16: 152565-152606, Chr 16: 152565-152603, Chr 16: 152565-152600, Chr 16: 152565-152597, Chr 16: 152565-152594, Chr 16: 152565-152591, Chr 16: 152565-152588, Chr 16: 152565-152585, Chr 16: 152565-152582, and Chr 16: 152577-152636.
[0023] In particular embodiments of the application, the non-coding region of the zeta-globin gene comprises: an intronic region, the sequence of which is located at Chr 16: 153005-153891.
[0024] In particular embodiments of the application, the sequence of the non-coding region of the zeta-globin gene is located at Chr 16: 153061-153380 and / or Chr 16: 153221-153540.
[0025] In particular embodiments of the application, the sequence of the non-coding region of the zeta-globin gene is located at Chr 16: 153161-153280 and / or Chr 16: 153321-153440.
[0026] In particular embodiments of the application, the sequence of the non-coding region of the zeta-globin gene is located at Chr 16: 153213-153242 and / or Chr 16: 153371-153400.
[0027] In particular embodiments of the application, the sequence of the non-coding region of the zeta-globin gene is located at Chr 16: 153219-153234 and / or Chr 16: 153378-153393.
[0028] In particular embodiments of the application, the non-coding region of the zeta-globin gene comprises: a 3' UTR region and sequences within 1 kb downstream of the 3' UTR region.
[0029] In particular embodiments of the application, the sequence of the non-coding region of the zeta-globin gene is located at Chr 16: 154401-155505.
[0030] In particular embodiments of the application, the non-coding region of the zeta-globin gene is selected from one or more of the sequences located at Chr 16: 154667-154986, Chr 16: 154917-155236, and Chr 16: 154976-155295.
[0031] In particular embodiments of the application, the non-coding region of the zeta-globin gene is selected from one or more of the sequences located at Chr 16: 154767-154886, Chr 16: 155017-155136, and Chr 16: 155076-155195.
[0032] In a specific embodiment of the application, the non-coding region of the zeta-globin gene is selected from one or more of the sequences located at: Chr 16: 154815-154844, Chr 16: 155057-155086, and Chr 16: 155127-155156 of the locus.
[0033] In a specific embodiment of the application, the non-coding region of the zeta-globin gene is selected from one or more of the sequences located at: Chr 16: 154823-154838, Chr 16: 155064-155080, and Chr 16: 155133-155149 of the locus.
[0034] In a specific embodiment of the application, the sequence as set forth in SEQ ID NO: 1 in the sense strand of the non-coding region of the zeta-globin gene is edited to the sequence as set forth in any one of SEQ ID NOs: 2-5; preferably the non-coding region of the zeta-globin gene is located at Chr 16: 150852-150971 of the locus.
[0035] In a specific embodiment of the application, the sequence as set forth in SEQ ID NO: 6 in the sense strand of the non-coding region of the zeta-globin gene is edited to the sequence as set forth in any one of SEQ ID NOs: 7-10; preferably the non-coding region of the zeta-globin gene is located at Chr 16: 152574-152693 of the locus.
[0036] In a specific embodiment of the application, the sequence as set forth in SEQ ID NO: 11 in the sense strand of the non-coding region of the zeta-globin gene is edited to the sequence as set forth in SEQ ID NO: 12 or 13; preferably the non-coding region of the zeta-globin gene is located at Chr 16: 152622-152741 of the locus.
[0037] In a specific embodiment of the application, the sequence as set forth in SEQ ID NO: 14 in the sense strand of the non-coding region of the zeta-globin gene is edited to the sequence as set forth in SEQ ID NO: 15; preferably the non-coding region of the zeta-globin gene is located at Chr 16: 152678-152797 of the locus.
[0038] In a specific embodiment of the application, the sequence as set forth in SEQ ID NO: 16 in the sense strand of the non-coding region of the zeta-globin gene is edited to the sequence as set forth in any one of SEQ ID NOs: 17-19; preferably the non-coding region of the zeta-globin gene is located at Chr 16: 153161-153280 of the locus.
[0039] In a specific embodiment of the application, the sequence as set forth in SEQ ID NO: 20 in the sense strand of the non-coding region of the zeta-globin gene is edited to the sequence as set forth in SEQ ID NO: 21 ; preferably the non-coding region of the zeta-globin gene is located at the locus Chr 16: 153321 -153440.
[0040] In a specific embodiment of the application, the sequence as set forth in SEQ ID NO: 22 in the antisense strand of the non-coding region of the zeta-globin gene is edited to the sequence as set forth in SEQ ID NO: 23; preferably the non-coding region of the zeta-globin gene is located at the locus Chr 16: 154767-154886.
[0041] In a specific embodiment of the application, the sequence as set forth in SEQ ID NO: 24 in the sense strand of the non-coding region of the zeta-globin gene is edited to the sequence as set forth in SEQ ID NO: 25; preferably the non-coding region of the zeta-globin gene is located at the locus Chr 16: 155017-155136.
[0042] In a specific embodiment of the application, the sequence as set forth in SEQ ID NO: 26 in the sense strand of the non-coding region of the zeta-globin gene is edited to the sequence as set forth in SEQ ID NO: 27 or 28; preferably the non-coding region of the zeta-globin gene is located at the locus Chr 16: 155076-155195.
[0043] In a specific embodiment of the application, the sequence as set forth in SEQ ID NO: 71 in the sense strand of the non-coding region of the zeta-globin gene is edited to the sequence as set forth in SEQ ID NO: 72; preferably the non-coding region of the zeta-globin gene is located at the locus Chr 16: 152587-152706.
[0044] In a specific embodiment of the application, the sequence as set forth in SEQ ID NO: 73 in the sense strand of the non-coding region of the zeta-globin gene is edited to the sequence as set forth in SEQ ID NO: 74; preferably the non-coding region of the zeta-globin gene is located at the locus Chr 16: 152528-152647.
[0045] In a specific embodiment of the application, the sequence as set forth in SEQ ID NO: 75 in the sense strand of the non-coding region of the zeta-globin gene is edited to the sequence as set forth in SEQ ID NO: 76; preferably the non-coding region of the zeta-globin gene is located at the locus Chr 16: 152499-152618.
[0046] In a specific embodiment of the present application, the sequence as shown in SEQ ID NO: 77 in the sense strand of the non-coding region of the zeta-globin gene is edited into the sequence as shown in SEQ ID NO: 78; preferably the non-coding region of the zeta-globin gene is located at the locus of Chr 16: 152353-152472.
[0047] In a specific embodiment of the present application, the sequence as shown in SEQ ID NO: 79 in the sense strand of the non-coding region of the zeta-globin gene is edited into the sequence as shown in SEQ ID NO: 80; preferably the non-coding region of the zeta-globin gene is located at the locus of Chr 16: 152278-152398.
[0048] To solve the above technical problems, the second technical solution of the present application provides a gRNA, wherein the gRNA comprises a recognition site sequence and a backbone sequence, the recognition site sequence is partially or completely complementary to the DNA sense strand or antisense strand of the non-coding region of the zeta-globin gene, and the gRNA is used to artificially form an enhancer element comprising a NTG-N(7-8)-WGATAR sequence, a NAG-N(7-8)-WGATAR sequence, a YTATCW-N(7-8)-CAN sequence and / or a YTATCW-N(7-8)-CTN sequence in the sense strand or antisense strand of the non-coding region of the zeta-globin gene.
[0049] The N is A, G, C or T, W is T or A, R is A or G, and Y is T or C.
[0050] In a specific embodiment of the present application, the non-coding region of the zeta-globin gene is defined in the method described in the first technical solution of the present application; and / or, the gRNA further comprises a chemical modification, preferably one or more of 3'-phosphorothioate, 2'-O-methyl ester, 2'-O-methyl, 2'-F modification, 2'-ribose 3'-phosphorothioate, deoxy and 5' phosphate modification; more preferably, the three ribonucleotides at the 5' end and the 3' end of the gRNA are 2'-O-methyl modified, and the phosphorothioate bond modification is performed between the four ribonucleotides at the 5' end and the 3' end.
[0051] In a specific embodiment of the present application, the backbone sequence is as shown in SEQ ID NO: 47 or 116; and / or, the recognition site sequence corresponds to a DNA sequence that is partially or completely identical to any one of the sequences as shown in SEQ ID NO: 29-37, SEQ ID NO: 81-85 and SEQ ID NO: 114.
[0052] In a specific embodiment of the present application, the nucleotide sequence of the gRNA is as shown in any one of SEQ ID NO: 48-56, SEQ ID NO: 86-90 and SEQ ID NO: 115.
[0053] To solve the above technical problems, the third technical solution provided by the present application is a gRNA, which comprises a recognition site sequence and a backbone sequence, the DNA sequence corresponding to the recognition site sequence is partially or completely identical to the sequence selected from any one of SEQ ID NO: 29-37, SEQ ID NO: 81-85 and SEQ ID NO: 114; and the backbone sequence is preferably as shown in SEQ ID NO: 47 or 116.
[0054] In a specific embodiment of the present application, the nucleotide sequence of the gRNA is as shown in any one of SEQ ID NO: 48-56, SEQ ID NO: 86-90 and SEQ ID NO: 115; and / or, the gRNA further comprises chemical modification, preferably one or more of 3'-phosphorothioate, 2'-O-methyl ester, 2'-O-methyl, 2'-F modification, 2'-ribose 3'-phosphorothioate, deoxy and 5' phosphate modification; more preferably, the three ribonucleotides at the 5' end and the 3' end of the gRNA are 2'-O-methyl modified, and the phosphorothioate bond modification is performed between the four ribonucleotides at the 5' end and the 3' end.
[0055] To solve the above technical problems, the fourth technical solution provided by the present application is an ssODN for gene editing of the non-coding region of the zeta-globin gene, the ssODN comprises a NTG-N(7-8)-WGATAR sequence, a NAG-N(7-8)-WGATAR sequence, a YTATCW-N(7-8)-CAN sequence and / or a YTATCW-N(7-8)-CTN sequence.
[0056] The N is A, G, C or T, W is T or A, R is A or G, and Y is T or C.
[0057] In a specific embodiment of the present application, the ssODN comprises a 5' homologous arm, a replacement sequence and a 3' homologous arm.
[0058] In a specific embodiment of the present application, the ssODN satisfies one or more of the following conditions:
[0059] (1) the NTG-N(7-8)-WGATAR sequence, NAG-N(7-8)-WGATAR sequence, YTATCW-N(7-8)-CAN sequence or YTATCW-N(7-8)-CTN sequence in the ssODN is located at any position on the ssODN, including the replacement sequence, the 5' homologous arm, the 3' homologous arm, the junction of the 5' homologous arm and the replacement sequence, or the junction of the 3' homologous arm and the replacement sequence;
[0060] (2) the 5' homologous arm and the 3' homologous arm are symmetrical, or are asymmetrical;
[0061] (3) the 5' homologous arm and / or the 3' homologous arm is selected from the sense strand or the antisense strand of the non-coding region of the ζ-globin gene;
[0062] (4) the 3' end of the 5' homologous arm and the 5' end of the 3' homologous arm are 0-20 bases apart in the ssODN;
[0063] (5) the length of the 5' homologous arm and / or the 3' homologous arm is 20-300 nt; and,
[0064] (6) the number of bases in the replacement sequence is 0-6.
[0065] In specific embodiments of the present application, the non-coding region of the ζ-globin gene is defined as in the method described in one of the technical solutions of the present application; and / or, the ssODN is chemically modified, for example, phosphorothioate modification; the chemical modification is preferably modification of the 5' end and the 3' end of the ssODN; more preferably, phosphorothioate modification between the first four nucleotides of the 5' end and the last four nucleotides of the 3' end of the ssODN.
[0066] In specific embodiments of the present application, the ssODN comprises a sequence selected from any one of SEQ ID NO: 2-5, SEQ ID NO: 7-10, SEQ ID NO: 12-13, SEQ ID NO: 15, SEQ ID NO: 17-19, SEQ ID NO: 21, SEQ ID NO: 23, SEQ ID NO: 25, SEQ ID NO: 27-28, SEQ ID NO: 72, SEQ ID NO: 74, SEQ ID NO: 76, SEQ ID NO: 78, SEQ ID NO: 80 and SEQ ID NO: 113, or a complementary sequence thereof.
[0067] In a specific embodiment of the present application, the ssODN comprises a sequence selected from the group consisting of any one of SEQ ID NO: 38-46, SEQ ID NO: 91-95 and SEQ ID NO: 117, or a complement thereof.
[0068] To solve the above technical problems, the sixth technical solution provided by the present application is a composition for gene editing in the non-coding region of the zeta-globin gene, which is used to artificially form an enhancer element comprising a NTG-N(7-8)-WGATAR sequence, a NAG-N(7-8)-WGATAR sequence, a YTATCW-N(7-8)-CAN sequence and / or a YTATCW-N(7-8)-CTN sequence in the sense strand or the antisense strand of the non-coding region of the zeta-globin gene.
[0069] In a specific embodiment of the present application, the ssODN comprises a sequence selected from the group consisting of any one of SEQ ID NO: 38-46, SEQ ID NO: 91-95 and SEQ ID NO: 117, or a complement thereof; and / or, the ssODN is chemically modified, for example, phosphorothioate modification; preferably, the chemical modification is modification of the 5' end and the 3' end of the ssODN; more preferably, phosphorothioate modification between the first four nucleotides of the 5' end and the last four nucleotides of the 3' end of the ssODN.
[0070] To solve the above technical problems, the sixth technical solution provided by the present application is a composition for gene editing in the non-coding region of the zeta-globin gene, which is used to artificially form an enhancer element comprising a NTG-N(7-8)-WGATAR sequence, a NAG-N(7-8)-WGATAR sequence, a YTATCW-N(7-8)-CAN sequence and / or a YTATCW-N(7-8)-CTN sequence in the sense strand or the antisense strand of the non-coding region of the zeta-globin gene.
[0071] The N is A, G, C or T, W is T or A, R is A or G, and Y is T or C.
[0072] In a specific embodiment of the present application, the non-coding region of the zeta-globin gene is defined in the method described in one of the technical solutions of the present application.
[0073] In particular embodiments of the application, the composition comprises one or more selected from the group consisting of: a CRISPR-Cas editing system, a TALEN editing system, a ZFN editing system, and an Argonaute editing system.
[0074] In particular embodiments of the application, the CRISPR-Cas editing system is used to effect one or more of homologous recombination repair editing, Indel editing, prime editing, and single base editing.
[0075] In particular embodiments of the application, the composition comprises:
[0076] (a) an ssODN containing the enhancer element;
[0077] (b) a gRNA targeting a non-coding region of the zeta-globin gene; and,
[0078] (c) a CRISPR / Cas nuclease, an mRNA encoding the CRISPR / Cas nuclease, and / or a plasmid expressing the CRISPR / Cas nuclease.
[0079] In particular embodiments of the application, the composition comprises: an ssODN containing the enhancer element, a gRNA targeting a non-coding region of the zeta-globin gene, and a CRISPR / Cas nuclease; an ssODN containing the enhancer element, a gRNA targeting a non-coding region of the zeta-globin gene, and an mRNA encoding a CRISPR / Cas nuclease; and / or, an ssODN containing the enhancer element, a gRNA targeting a non-coding region of the zeta-globin gene, and a plasmid expressing a CRISPR / Cas nuclease.
[0080] In particular embodiments of the application, the composition satisfies one or more conditions selected from the group consisting of:
[0081] (1) the CRISPR / Cas nuclease comprises at least one or a combination of the following group: Casl, CaslB, Cas2, Cas3, Cas4, Cas5, Cas5d, Cas5t, Cas5h, Cas5a, Cas6, Cas7, Cas8, Cas8a, Cas8b, Cas8c, Cas9, CaslO, CaslOd, Casl2a / Cpfl, Casl2b / C2cl, Casl2c / C2c3, Casl2d / CasY, Casl2e / CasX, Casl2f / CasZ, Casl2g, Casl2h, Casl2i, Csy1, Csy2, Csy3, Csy4, Cse1, Cse2, Cse3, Cse4, Cse5e, Csc1, Csc2, Csa5, Csn1, Csn2, Csm1, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx1S, Csx11, Csf1, Csf2, CsO, Csf4, Csd1, Csd2, Cst1, Cst2, Csh1, Csh2, Csa1, Csa2, Csa3, Csa4, Csa5, Casl3a, Casl3b, Casl3c, Casl3d, Casl3e, Casl3f, fragments thereof, and variants or fragments of variants thereof; the variants are preferably at least one selected from the group consisting of: SpRY variants, NG-nCas9 variants, and NGG-nCas9 variants; (2) the gRNA is as described in the second or third aspect of the present invention; and, (3) the ssODN is as described in the fourth or fifth aspect of the present invention.
[0082] In a specific embodiment of the present invention, the composition comprises:
[0083] (a) a gRNA targeting a non-coding region of the zeta-globin gene; and,
[0084] (b) a CRISPR / Cas nuclease, an mRNA encoding the CRISPR / Cas nuclease, and / or a plasmid expressing the CRISPR / Cas nuclease.
[0085] In specific embodiments of the application, the composition comprises: a gRNA targeting a non-coding region of the zeta-globin gene and a CRISPR / Cas nuclease; a gRNA targeting a non-coding region of the zeta-globin gene and an mRNA encoding a CRISPR / Cas nuclease; and / or, a gRNA targeting a non-coding region of the zeta-globin gene and a plasmid expressing a CRISPR / Cas nuclease.
[0086] In specific embodiments of the application, the CRISPR / Cas nuclease comprises at least one or a combination of the following group: Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5d, Cas5t, Cas5h, Cas5a, Cas6, Cas7, Cas8, Cas8a, Cas8b, Cas8c, Cas9, Cas10, Cas10d, Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12f / CasZ, Cas12g, Cas12h, Cas12i, Csy1, Csy2, Csy3, Csy4, Cse1, Cse2, Cse3, Cse4, Cse5e, Csc1, Csc2, Csa5, Csn1, Csn2, Csm1, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx1S, Csxl l, Csf1, Csf2, CsO, Csf4, Csd1, Csd2, Cst1, Cst2, Csh1, Csh2, Csa1, Csa2, Csa3, Csa4, Csa5, Cas13a, Cas13b, Cas13c, Cas13d, Cas13e, Cas13f, fragments thereof, and variants thereof or fragments of variants; the variants are preferably at least one selected from the group consisting of: SpRY variants, NG-nCas9 variants and NGG-nCas9 variants; and / or, the gRNA is as described in the second or third aspect of the application.
[0087] In specific embodiments of the application, the composition comprises:
[0088] (a) a PEgRNA; and,
[0089] (b) an nCas9 and a reverse transcriptase, an mRNA encoding the nCas9 and reverse transcriptase and / or a plasmid expressing the nCas9 and reverse transcriptase;
[0090] wherein the PEgRNA comprises the enhancer element and targets a non-coding region of the zeta-globin gene.
[0091] In particular embodiments of the application, the composition comprises a PEgRNA comprising the enhancer element and targeting a non-coding region of the zeta-globin gene, and nCas9 and a reverse transcriptase.
[0092] In particular embodiments of the application, the composition comprises a PEgRNA comprising the enhancer element and targeting a non-coding region of the zeta-globin gene, and an mRNA encoding nCas9 or a reverse transcriptase.
[0093] In particular embodiments of the application, the composition comprises a PEgRNA comprising the enhancer element and targeting a non-coding region of the zeta-globin gene, and a plasmid expressing nCas9 or a reverse transcriptase.
[0094] In particular embodiments of the application, the composition comprises:
[0095] (a) a gRNA targeting a non-coding region of the zeta-globin gene; and,
[0096] (b) a single base editor, an mRNA encoding the single base editor, and / or a plasmid expressing the single base editor;
[0097] wherein the single base editor comprises a target nucleic acid binding structural unit and a deaminase structural unit.
[0098] In particular embodiments of the application, the composition comprises: a gRNA targeting a non-coding region of the zeta-globin gene, and a single base editor comprising a target nucleic acid binding structural unit and a deaminase structural unit.
[0099] In particular embodiments of the application, the composition comprises: a gRNA targeting a non-coding region of the zeta-globin gene, and an mRNA encoding a single base editor comprising a target nucleic acid binding structural unit and a deaminase structural unit.
[0100] In particular embodiments of the application, the composition comprises: a gRNA targeting a non-coding region of the zeta-globin gene, and a plasmid expressing a single base editor comprising a target nucleic acid binding structural unit and a deaminase structural unit.
[0101] In particular embodiments of the application, the target nucleic acid binding structural unit is selected from one or more of: a CRISPR / Cas nuclease, a ZFN, a TALEN, and an Argonaute; and / or, the gRNA is as described in Clause II or Clause III of the application.
[0102] In particular embodiments of the application, the CRISPR / Cas enzyme comprises at least one or a combination of the following group: Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5d, Cas5t, Cas5h, Cas5a, Cas6, Cas7, Cas8, Cas8a, Cas8b, Cas8c, Cas9, Cas10, Cas10d, Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12f / CasZ, Cas12g, Cas12h, Cas12i, Csy1, Csy2, Csy3, Csy4, Cse1, Cse2, Cse3, Cse4, Cse5e, Csc1, Csc2, Csa5, Csn1, Csn2, Csm1, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx1S, Csx11, Csf1, Csf2, CsO, Csf4, Csd1, Csd2, Cst1, Cst2, Csh1, Csh2, Csa1, Csa2, Csa3, Csa4, Csa5, Cas13a, Cas13b, Cas13c, Cas13d, Cas13e, Cas13f, fragments thereof, and variants or fragments of variants thereof; and / or, the deaminase structural unit is selected from: an APOBEC1 deaminase, an APOBEC2 deaminase, an APOBEC3A deaminase, an APOBEC3B deaminase, an APOBEC3C deaminase, an APOBEC3D deaminase, an APOBEC3F deaminase, an APOBEC3G deaminase, an APOBEC3H deaminase, an APOBEC4 deaminase, an activation-induced deaminase, and pmCDA1, and variants or combinations thereof.
[0103] In specific embodiments of the application, the variant of the CRISPR / Cas enzyme is selected from at least one of: a SpRY variant, an NG-nCas9 variant, and an NGG-nCas9 variant; and / or, the variant of the deaminase structural unit has at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity to the amino acid sequence of a wild-type deaminase structural unit, and retains deaminase activity.
[0104] To solve the above technical problems, the seventh technical solution of the present application provides a cell, wherein the cell comprises the composition according to the sixth technical solution of the present application; and / or,
[0105] The cell is edited by the composition according to the sixth technical solution of the present application.
[0106] In specific embodiments of the application, the target cell is an erythroid progenitor cell; and / or, the target cell is a mammalian cell, such as a human cell.
[0107] In specific embodiments of the application, the erythroid progenitor cell is an umbilical cord blood stem cell, an induced pluripotent stem cell, a hematopoietic stem / progenitor cell, a myeloid progenitor cell, a burst-forming unit-erythroid cell / erythrocyte, a spleen colony-forming cell, a blast cell colony-forming cell, and / or a megakaryocyte-erythroid progenitor cell.
[0108] To solve the above technical problems, the eighth technical solution of the present application provides use of the gRNA according to the second or third technical solution of the present application, the ssODN according to the fourth or fifth technical solution of the present application, and / or the composition according to the sixth technical solution of the present application in the preparation of a cell for treating alpha-thalassemia.
[0109] In specific embodiments of the application, the alpha-thalassemia is moderate or severe alpha-thalassemia.
[0110] To solve the above technical problems, the ninth technical solution of the present application provides use of the gRNA according to the second or third technical solution of the present application, the ssODN according to the fourth or fifth technical solution of the present application, the composition according to the sixth technical solution of the present application, and / or the cell according to the seventh technical solution of the present application in the preparation of a medicament for treating alpha-thalassemia.
[0111] In specific embodiments of the application, the alpha-thalassemia is moderate or severe alpha-thalassemia; and / or, the medicament is a cell.
[0112] To solve the above technical problems, a tenth technical solution of the present application provides a method for treating alpha-thalassemia, which comprises administering to a subject in need an effective amount of the gRNA according to the second or third technical solution of the present application, the ssODN according to the fourth or fifth technical solution of the present application, the composition according to the sixth technical solution of the present application, and / or the cell according to the seventh technical solution of the present application.
[0113] In specific embodiments of the present application, the alpha-thalassemia is moderate or severe alpha-thalassemia.
[0114] To solve the above technical problems, an eleventh technical solution of the present application provides the gRNA according to the second or third technical solution of the present application, the ssODN according to the fourth or fifth technical solution of the present application, the composition according to the sixth technical solution of the present application, and / or the cell according to the seventh technical solution of the present application for use in treating alpha-thalassemia.
[0115] In specific embodiments of the present application, the alpha-thalassemia is moderate or severe alpha-thalassemia.
[0116] On the basis of common general knowledge in the art, the above-mentioned preferred conditions can be combined in any manner, thereby obtaining preferred embodiments of the present application.
[0117] The reagents and raw materials used in the present application are commercially available.
[0118] The positive progress effect of the present application is that:
[0119] The present application artificially forms the NTG-N(7-8)-WGATAR sequence, the NAG-N(7-8)-WGATAR sequence, the YTATCW-N(7-8)-CAN sequence, and / or the YTATCW-N(7-8)-CTN sequence in the sense strand or the antisense strand of the non-coding region of the zeta-globin gene by gene editing technology, so that the non-coding region of the zeta-globin gene forms an enhancer element containing the NTG-N(7-8)-WGATAR sequence, the NAG-N(7-8)-WGATAR sequence, the YTATCW-N(7-8)-CAN sequence, and / or the YTATCW-N(7-8)-CTN sequence, thereby activating the expression of the zeta-globin gene, forming zeta-globin, increasing the expression of alpha-globin, and further alleviating the alpha-thalassemia phenotype caused by the deletion or mutation of the HBA1 / HBA2 gene, thereby treating alpha-thalassemia. The present application does not need to perform overexpression of an exogenous globin gene, thereby reducing the safety risk of gene therapy. Moreover, activating the expression of the zeta-globin gene can be applied to various types of alpha-thalassemia patients, including patients with alpha-gene deletion or mutation, and is not limited to alpha-thalassemia caused by a certain mutation site. BRIEF DESCRIPTION OF DRAWINGS
[0120] Figure 1 Schematic diagram of distribution position of nine regions in HBZ gene for Example 1.
[0121] Figures 2-1 to 2-9 Position map of gRNA targeting sequence for gene editing of each region for Example 1.
[0122] Figure 3 Editing efficiency result map of gene editing of each region in K562 cells for Example 1.
[0123] Figure 4 Editing efficiency result map of gene editing of each region in CD34 + HSPC for Example 2.
[0124] Figure 5-1 Editing efficiency result map of gene editing of each region in CD34 + HSPC for Example 2. Figure 5-2 Gene expression result map of ζ-globin after gene editing of each region in CD34 + HSPC for Example 1 (n=3).
[0125] Figure 6 Protein expression result map of ζ-globin after gene editing of each region in CD34 + HSPC for Example 2.
[0126] Figure 7 Editing efficiency result map of gene editing of each region in CD34 + HSPC for Example 3 (left) and protein expression result map of ζ-globin after gene editing of each region in CD34 + HSPC for Example 3 (right).
[0127] Figure 8 Flow cytometry result map of ζ-globin after gene editing of each region in CD34 + HSPC for Example 3.
[0128] Figure 9 Editing efficiency result map of gene editing in CD34 + HSPC for Example 4 (n=3).
[0129] Figure 10 Gene expression result map of ζ-globin after gene editing in CD34 + HSPC for Example 4 (n=3). DETAILED DESCRIPTION
[0130] In the present application, the scientific and technical terms used herein have the meanings commonly understood by a person skilled in the art, unless otherwise specified. Also, the molecular genetic, nucleic acid chemistry, chemical, molecular biology, biochemistry, cell culture, microbiology, cell biology, genomics and recombinant DNA procedures used herein are conventional procedures widely used in the corresponding fields. Meanwhile, in order to better understand the present application, the definitions of key terms are provided as follows:
[0131] The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
[0132] The locus, also known as the locus, refers to the specific position of the specific fragment DNA sequence on the chromosome. A person skilled in the art can locate and determine the specific sequence according to the specific locus. In the present application, the complete sequence of human chromosome 16 corresponds to Homosapiens chromosome 16, GRCh38.p14 Primary Assembly (NC_000016.10) in NCBI (https: / / www.ncbi.nlm.nih.gov / ); therefore, the sequence defined by the locus position number range on Chr16 used in the present application refers to the sequence within the corresponding Location range on NC_000016.10.
[0133] In the present application, the antisense complementary sequence corresponding to the NTG-N(7-8)-WGATAR sequence or the NAG-N(7-8)-WGATAR sequence is the YTATCW-N(7-8)-CAN sequence or the YTATCW-N(7-8)-CTN sequence. Wherein, W is T or A, R is A or G, Y is T or C, and N is A, G, C or T.
[0134] In the present application, "Coding region" only refers to the discontinuous DNA sequence with protein coding function (i.e. the exon part without 5'UTR region and 3'UTR region). Therefore, the "Non-coding region" described in the present application includes the non-coding sequence between exons (i.e. the intron part), as well as the region on both sides of the first and last exons that can regulate the expression and strength of the gene, including the promoter region, the terminator region, the enhancer region, the untranslated region (including the 5'UTR region and the 3'UTR region).
[0135] It is known in the art that the promoter region of the gene is contained in the 2kb region upstream of the transcription start site (TSS).
[0136] In this invention, the application scenarios of the term "non-diagnostic / therapeutic purpose" include, but are not limited to: for scientific research purposes, using the method described in one of the technical solutions of this invention to activate the expression of the ζ-globin gene in target cells in vitro, such as the application scenarios exemplified in the embodiments of this invention or similar application scenarios.
[0137] This invention provides a method for efficient gene editing of CD34+ hematopoietic stem cells and other stem cells and progenitor cells capable of erythrocyte differentiation using gene editing technology. A mutation is created in the non-coding region of the HBZ gene (located on human chromosome 16) that encodes human ζ-globin. An NTG-N(7-8)-WGATAR or NAG-N(7-8)-WGATAR sequence structure is artificially created on the sense or antisense strand of the non-coding region. This sequence acts as an enhancer and can recruit transcription activators such as GATA1 to promote ζ-globin expression after the target cells differentiate into erythrocytes.
[0138] This invention applies to site-specific nuclease-mediated gene editing systems such as CRISPR-Cas editing, TALEN editing, ZFN editing, and Argonaute editing. "CRISPR-Cas" is a gene editing technology, including but not limited to various naturally occurring or artificially designed CRISPR-Cas systems, such as the CRISPR-Cas9 system and the CRISPR-Cas12 system. The working principle of CRISPR-Cas9 is that crRNA (CRISPR-derived RNA) binds to tracrRNA (trans-activating RNA) through base pairing to form a tracrRNA / crRNA complex. This complex guides the Cas9 nuclease protein to cleave double-stranded DNA at the target site paired with the crRNA. The roles of tracrRNA and crRNA can also be replaced by an artificially synthesized guide RNA (sgRNA). When using other CRISPR-Cas systems, corresponding sgRNA or crRNA needs to be designed. When using systems such as TALEN or ZFN, corresponding TALEN or ZFN nucleases need to be designed according to the editing sites disclosed in this invention. When using the Argonaute editing system, corresponding 5′ phosphorylated guide DNA needs to be designed according to the editing sites disclosed in this invention.
[0139] In this invention, "homlogous-directed repair (HDR)" refers to the process of repairing DNA damage using homologous nucleic acids (e.g., endogenous homologous sequences (e.g., sister chromatids)) or exogenous nucleic acids (e.g., donor templates)). "Homologous recombination repair editing" refers to the editing method in which a donor template (e.g., ssODN) is used in cells to guide repair and produce specific sequence changes in the genome, including targeted additions to the entire gene. If the donor template is provided with a site-specific nuclease, such as with a CRISPR / Cas9-based system or a CRISPR / Cpf1-based system, the cell will repair breaks via homologous recombination, an improvement several orders of magnitude in the presence of DNA damage such as double-strand breaks.
[0140] ssODN (single-stranded oligonucleotides) structurally contain a 5′ homologous arm, a substitution sequence, and a 3′ homologous arm. Homologous arms are sequences homologous to the DNA regions flanking the target site, used to locate the target site on the chromosome. The corresponding sequences of the 5′ and 3′ homologous arms on the chromosome are usually not contiguous, with a gap of 1-20 nucleotides in between. These gap sequences are the targets of gene editing (the intended edit sequence), i.e., the sequence that is to be replaced by gene editing. The substitution sequence between the 3′ end of the 5′ homologous arm and the 5′ end of the 3′ homologous arm on the ssODN is the desired result of gene editing. Adding ssODN during gene editing can induce homologous recombination repair during gene editing, resulting in the intended edit sequence in the cell genome being replaced by a substitution sequence. The length of the substitution sequence in the ssODN can be shorter than the intended edit sequence; in this case, the result of gene editing is the deletion and / or replacement of the intended edit sequence in the original genome. In ssODN, the replacement sequence can be 0 bases long, in which case the intended edited sequence is deleted from the genome. The replacement sequence in ssODN can also be longer than the intended edited sequence, in which case the gene editing result is the replacement of the intended edited sequence or / and the insertion of a new sequence.
[0141] In this invention, "insertion-deletion (Indel) editing" refers to creating one or more nucleotide insertions, one or more nucleotide deletions, or a combination of nucleotide insertions and deletions in the edited nucleic acid relative to the unedited nucleic acid. In some of these methods, Indel editing is performed using a composition containing gRNA molecules (e.g., a CRISPR system). The results of Indel editing can be determined by sequencing the nucleic acid after exposure to the composition containing gRNA molecules, for example, via NGS.
[0142] In this invention, "Prime Editing (PE)" refers to a method of gene editing using a programmable DNA-binding protein (napDNAbp, such as Cas enzyme, TALEN, zinc finger nuclease, etc.), a polymerase (e.g., reverse transcriptase), and a specialized guide RNA. That is, gene editing based on a Prime Editing (PE) system, wherein the specialized guide RNA contains a DNA synthesis template for encoding desired new genetic information (or deleting genetic information) on the extension arm of a traditional guide RNA or sgRNA. The new genetic information is then introduced into the target DNA.
[0143] The PE system comprises a leader editor and specialized guide RNA. In the PE system, the polymerase (or a functional derivative thereof) and napDNAbp (or a functional derivative thereof) are collectively referred to as the leader editor. "NapDNAbp," or nucleoprogrammable DNA-binding protein, is a protein that targets and binds to specific sequences in a DNA molecule through nucleic acid hybridization. Each napDNAbp associates with at least one guide nucleic acid (e.g., guide RNA), which positions the napDNAbp to a DNA sequence containing a DNA strand complementary to the guide nucleic acid or a portion thereof (e.g., the protospacer region of the guide RNA). In other words, by "programming" the guide nucleic acid, napDNAbp (e.g., Cas9 or a functional derivative thereof) can be positioned and bound to different target DNAs. In some of these schemes, napDNAbp includes one or more nuclease activities that, upon cleaving the DNA, leave various types of DNA damage. For example, napDNAbp may contain nuclease activities that cleave a non-target strand at a first position and / or cleave a target strand at a second position. Depending on the nuclease activity, the target DNA can be cleaved to form a "nick" on a single strand or a "double-strand break." Exemplary napDNAbp with different nuclease activities include “Cas9 nickase” (“nCas9”) and inactivated Cas9 (“dead Cas9” or “dCas9”) without nuclease activity.
[0144] In some of these schemes, the polymerase and napDNAbp exist as a fusion protein. In some of these schemes, napDNAbp is a Cas protein. In some of these schemes, the polymerase is a reverse transcriptase. In some of these schemes, the Cas protein-reverse transcriptase fusion protein or related system utilizes a guide RNA to target a specific DNA sequence, creating a single-stranded nick at the target site (i.e., the target sequence or its complementary strand), and uses the nicked DNA as a primer to perform reverse transcription based on an engineered reverse transcription template integrated with the guide RNA. It should be noted that the polymerase in this application is not limited to reverse transcriptase, but can be selected from almost any DNA polymerase. On the one hand, the leader editor may contain Cas9 (or its equivalent napDNAbp), which is programmed to target the DNA sequence by associating it with a specialized guide RNA (when napDNAbp is Cas9, the specialized guide RNA is PEgRNA). PEgRNA contains a spacer sequence that anneals (or binds to) a protospacer complementary to the target DNA, and carries new genetic information on an extension arm at one end. This new genetic information is encoded by a substitution sequence (DNA synthesis template) containing the desired genetic change, which replaces the corresponding endogenous DNA strand at the target site. To transfer information from the PEgRNA to the target DNA, lead editing involves creating a nick in one strand of DNA at the target site (located in the DNA strand containing the target sequence or its complementary sequence) to expose a 3'-hydroxyl group. In some protocols, the PEgRNA further includes a specific sequence on its extension arm that hybridizes to the strand containing the exposed 3'-hydroxyl terminus to initiate reverse transcription or DNA-dependent DNA synthesis; this is called a "primer binding site." In some protocols, the exposed 3'-hydroxyl terminus triggers DNA polymerization and extension of the sequence at the target site according to the DNA synthesis template contained in the PEgRNA. The resulting DNA strand after DNA polymerization and extension of the sequence at the target site according to the DNA synthesis template contained in the PEgRNA is called the edited strand. In some protocols, the DNA synthesis template can be RNA or DNA. When the DNA synthesis template is RNA, the polymerase of the leader editor can be an RNA-dependent DNA polymerase (e.g., reverse transcriptase). When the DNA synthesis template is DNA, the polymerase of the leader editor can be a DNA-dependent DNA polymerase. In some of these schemes, the "DNA synthesis template" is homologous to the genomic target sequence (i.e., has the same sequence) except that it contains the desired nucleotide changes (e.g., single nucleotide changes, deletions, or insertions, or combinations thereof). The newly polymerized DNA strand, also known as a single-stranded DNA lobe, will compete for hybridization with complementary homologous endogenous DNA strands, thereby replacing the corresponding endogenous sequence. In some of these schemes, the system can be used in combination with a mistake-prone reverse transcriptase (e.g., as a fusion protein with a Cas9 domain).Error-prone reverse transcriptases can introduce changes during the synthesis of single-stranded DNA lobules. Therefore, in some embodiments, error-prone reverse transcriptases can be used to introduce nucleotide changes into the target DNA. The changes induced by this error-prone reverse transcriptase can be random or non-random. In different embodiments, the DNA synthesis template contained in the PEgRNA is located at the 3' or 5' end of its contained conventional guide RNA (or sgRNA) sequence or within the conventional guide RNA sequence.
[0145] PEgRNA is a specialized guide RNA that associates with a lead editor and targets a DNA sequence, guiding the lead editor to perform lead editing on the target DNA sequence. PEgRNA contains a spacer sequence that hybridizes complementary to the target DNA sequence, and extends to one or both sides of the spacer sequence. The extends contain one or more DNA synthesis templates containing novel genetic information used to replace the corresponding endogenous DNA strand at the target site. The DNA synthesis templates include, but are not limited to, single-stranded RNA or DNA. In some embodiments, PEgRNA contains a conventional guide RNA sequence containing a spacer sequence. The DNA synthesis template may be present at the 3' end, 5' end, or within the sequence of the conventional guide RNA. In some embodiments, PEgRNA further includes primer binding sites, adapters, or other additional structural elements on its extends, such as, but not limited to, aptamers, stem-loops, hairpins, and RNA-protein recruitment domains (e.g., MS2 hairpins). The primer binding site (PBS) contains a sequence complementary to the nicked genomic DNA strand. After hybridization with the complementary genomic DNA strand, the PE system can synthesize the edited strand starting from the nick, based on the DNA synthesis template. The primer binding site hybridizes with the strand of the target DNA containing the nick after the leader editor creates a nick in the target sequence.
[0146] In this invention, "single-base editing" refers to a gene editing method that introduces point mutations at target loci using a single-base editing system to convert a specific nucleic acid base into another nucleic acid base. The single-base editing system includes a single-base editor and gRNA. A "single-base editor" refers to a reagent capable of modifying bases (such as A, T, C, G, or U) within a nucleic acid sequence (such as DNA or RNA). In some of these schemes, the single-base editor can deamination of bases within nucleic acids (e.g., DNA molecules), including but not limited to adenine base editors (ABE), cytosine base editors (CBE), cytosine and adenine base editors (CABE), and / or adenine base transversion editors (AYBE, Y = C or T). For example, in the case of an adenine base editor (or adenosine editor), the single-base editor can deamination of adenine (A) in DNA. Such a single-base editor may include a programmable DNA-binding protein (target nucleic acid binding structural unit) fused with an adenosine deaminase. Programmable DNA-binding proteins include CRISPR-mediated Cas effector proteins. In some of these schemes, a single-base editor comprises an inactive Cas9 nuclease (dCas9) fused to a deaminase (a deaminase structural unit) that binds to but does not cleave nucleic acids. For example, the dCas9 domain of a single-base editor may include D10A and H840A mutations. The DNA-cutting domain of Pseudomonas Cas9 comprises two subdomains: the HNH nuclease subdomain and the RuvC I subdomain. The HNH subdomain cleaves the strand complementary to the gRNA (“target strand,” or the strand that is edited or deaminated), while the RuvC I subdomain cleaves the non-complementary strand containing the PAM sequence (“unedited strand”). The RuvC I variant D10A creates a nick in the target strand, while the HNH variant H840A creates a nick in the unedited strand.
[0147] In this invention, a "variant" refers to a protein that contains one or more mutations compared to a wild-type protein (such as a wild-type CRISPR / Cas enzyme or a wild-type deaminase structural unit), such as a single amino acid insertion, a single amino acid deletion, a single amino acid substitution, or a combination thereof. The amino acid sequence of the variant has at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with its corresponding wild-type protein, and retains the activity of the original wild-type protein, or has superior activity to the original wild-type protein.
[0148] In this invention, the term "effective amount" refers to the amount of a drug or agent that elicits a biological or pharmaceutical response in a tissue, system, animal, or human, as sought by, for example, an investigator or clinician. Furthermore, the term "effective amount" refers to the amount that causes improved treatment, cure, prevention, or reduction of disease, symptom, or side effects, or reduces the rate of progression of a disease or condition, compared to a corresponding subject who did not receive that amount. Within its scope, the term also includes amounts that effectively enhance normal physiological function.
[0149] The present invention is further illustrated below by way of embodiments, but the invention is not limited to the scope of the embodiments described herein. Experimental methods in the following embodiments that do not specify specific conditions were performed according to conventional methods and conditions, or as selected according to the product instructions.
[0150] Example 1: Introducing enhancer elements into the non-coding region of the ζ-globin gene in K562 cells
[0151] By analyzing the sequence within approximately 2 kb upstream of the HBZ transcription start site (TSS), introns, and within 1 kb downstream of the 3′UTR tail, the inventors discovered that many sequences could form NTG-N(7-8)–WGATAR or NAG-N(7-8)–WGATAR sequence structures by altering a few bases through substitution, deletion, or insertion. By analyzing whether the target mutant bases in these regions contained NGG (spCas9-recognized PAM) sequences and whether the cleavage site precisely aligned with the target sequence, nine candidate target regions most likely to achieve efficient gene editing were ultimately identified (see the distribution diagram). Figure 1 The location map of the gRNA target sequences used for gene editing in each region is shown below. Figures 2-1 to 2-9 The locations where the reinforcing sub-elements in each region will be formed are as follows:
[0152] Region 1 (TSS1700), upstream of the TSS (promoter) of the HBZ gene, between -1954 and -1939, original sequence ctgtggtaaaggatag (SEQ ID NO: 96), located at locus Chr16:150901-150916 ( Figure 2-1 );
[0153] Region 2 (TSS7), between -235 and -219 of the HBZ gene promoter, original sequence ctatctctcctagactc (SEQ ID NO: 97), located at locus Chr16:152620-152636 ( Figure 2-2 );
[0154] Region 3 (U96), between -175 and -159 of the HBZ gene promoter, original sequence aggaacaggagtgatag (SEQ ID NO: 98), located at locus Chr16:152680-152696 ( Figure 2-3 );
[0155] Region 4 (U39), between -118 and -102 of the HBZ gene promoter, original sequence gtcactggatctgataa (SEQ ID NO: 99), located at locus Chr16:152737-152753 ( Figure 2-4 );
[0156] Region 5 (N200), the first intron of HBZ, located between +365 and +380 downstream of the TSS, with the original sequence cgtgaggacagatag (SEQ ID NO: 100), located at locus Chr16:153219-153234. Figure 2-5 );
[0157] Region 6 (N360), the first intron of HBZ, located between +524 and +539 Å downstream of TSS, with the original sequence cttacagggcagccag (SEQ ID NO: 101), located at locus Chr16: 153378-153393. Figure 2-6 );
[0158] Region 7 (W310), 3′ tail of the HBZ gene, between +1969 and +1984 downstream of the TSS, original sequence ctgatcgttctgaaat (SEQ ID NO: 102), located at locus Chr16:154823-154838 ( Figure 2-7 );
[0159] Region 8 (W560), 3′ tail of the HBZ gene, between +2210 and +2226 downstream of the TSS, original sequence ctgagcctcactcataa (SEQ ID NO:103), located at locus Chr16:155064-155080 ( Figure 2-8 );
[0160] Region 9 (W630), 3′ tail of the HBZ gene, between +2279 and +2295 downstream of the TSS, original sequence tctcacctccctgatag (SEQ ID NO: 104), located at locus Chr16: 155133-155149 ( Figure 2-9 ).
[0161] Based on the sequence characteristics of the nine candidate regions and the cut site locations predicted by the CRISPR-Cas9 gene editing system, the most suitable mutation target types for each region were analyzed, as shown in Table 1. Underlined bases indicate the bases that need to be replaced, representing the desired result after gene editing. In this invention, the replacement base length is 0-6 bases, achieving deletion, substitution, insertion, or a combination of these actions.
[0162] Nine regions of the non-coding region of the ζ-globin gene were altered by substitution, deletion, or insertion of a few bases to form NTG-N(7-8)-WGATAR or NAG-N(7-8)-WGATAR sequence structures.
[0163] Table 1. Wild-type sequences and edited sequences in the HBZ non-coding region.
[0164] Region SEQ ID Number and region Target site and edited sequence (5'-3') TSS1700 SEQ ID NO: 1 Wild type GCACTGTGGCTGTGGTAAAGGATAGACACA TSS1700 SEQ ID NO: 2 Mutant 01 GCACTGTGGCTGTGGTAAA A GATAGACACA]]> TSS1700 SEQ ID NO: 3 Mutant 02 GCACTGTGGCTGTGGTAAA T GATAGACACA]]> TSS1700 SEQ ID NO: 4 Mutant 03 GCACTGTGGCTGTGGTAAAG A GATAGACACA]]> TSS1700 SEQ ID NO: 5 Mutant 04 GCACTGTGGCTGTGGTAAAG T GATAGACACA]]> TSS7 SEQ ID NO: 6 Wild type GGCCCCTATCTCTCCTAGACTCTGTGGTCA TSS7 SEQ ID NO: 7 Mutant 05 GGCCCCTATCTCTCCTAGAC AG TGTGGTCA]]> TSS7 SEQ ID NO: 8 Mutant 06 GGCCCCTATCTCTCCTAGAC AG TCTGTGGTCA TSS7 SEQ ID NO: 9 Mutant 07 GGCCCCTATCTCTCCTAGAC AG CTGTGGTCA]]> TSS7 SEQ ID NO: 10 Mutant 08 GGCCCCTATCTCTCCTAGAC A GTGGTCA U96 SEQ ID NO: 11 Wild type CCACAGGAGAGGAACAGGAGTGATAGCCCC U96 SEQ ID NO: 12 Mutant 09 CCACAGGAG CT GAACAGGAGTGATAGCCCC]]> U96 SEQ ID NO: 13 Mutant 10 CCACAGGAGAG CT GAACAGGAGTGATAGCCCC]]> U39 SEQ ID NO: 14 Wild type CCCTTTGTCACTGGATCTGATAAGAAACAC U39 SEQ ID NO: 15 Mutant 11 CCCTTT CTG ACTGGATCTGATAAGAAACAC]]> N200 SEQ ID NO: 16 Wild type AAGGGACAGTGAGGACAGATAGCGTTCCCT N200 SEQ ID NO: 17 Mutant 12 AAGGGAC T GTGAGGACAGATAGCGTTCCCT]]> N200 SEQ ID NO: 18 Mutant 13 AAGGGAC CT GTGAGGACAGATAGCGTTCCCT]]> N200 SEQ ID NO: 19 Mutant 14 <![CDATA[AAGGGACA CT GTGAGGACAGATAGCGTTCCCCT]]> N360 SEQ ID NO: 20 Wild type GAAGGGCCTTACAGGGCAGCCAGGGCACTA N360 SEQ ID NO: 21 Mutant 15 GAAGGGCCT AT CAGGGCAGCCAGGGCACTA]]> W310 SEQ ID NO: 22 Wild type TTCGTCCTGATCGTTCTGAAATCAGGAAAT W310 SEQ ID NO: 23 Mutant 16 TTCGTCCTGATCGTTCTGATAACAGGAAAT W560 SEQ ID NO: 24 Wild type AAAGAGGCTGAGCCTCACTCATAAGAGAAA W560 SEQ ID NO: 25 Mutant 17 AAAGAGGCTGAGCCTCACT G ATAAGAGAAA]]> W630 SEQ ID NO: 26 Wild type TCCATGTCTCACCTCCCTGATAGGCAAAAA W630 SEQ ID NO: 27 Mutant 18 [TCCTTGTCT G CACCTCCCTGATAGGCAAAAA]]> W630 SEQ ID NO: 28 Mutant 19 [TCCTTGTCTC TG ACCTCCCTGATAGGCAAAAA]]>
[0165] To achieve the goal of cutting the DNA double strand near the editing site in the region selected in Table 1 to form a DSB, this invention selects the spCas9 CRISPR-Cas system (derived from Streptococcus pyogenes), which has a high cutting efficiency, to analyze target sites with potentially high cutting efficiency. Sites 1 to 9 were selected as candidate target sites, and the DNA sequences (SEQ ID NO:29 to SEQ ID NO:37) and PAM sequences (NGG) of the identified target sites are shown in Table 2.
[0166] Table 2. Target sites recognized by sgRNAs used in the HBZ non-coding region
[0167]
[0168] Based on the expected mutation types shown in Table 1 and the DNA strands (sense or antisense strands) recognized by sgRNA in Table 2, corresponding guide gene editing repair ssODNs were designed (Table 3, SEQ ID NO:38 to SEQ ID NO:46). In this invention, the ssODN structure includes a 5′ homologous arm, a substitution sequence, and a 3′ homologous arm. It should be noted that the NTG-N(7-8)-WGATAR sequence, NAG-N(7-8)-WGATAR sequence, or the corresponding antisense complementary sequence in the ssODN structure can be located at any position on the ssODN, including the substitution sequence, the 5′ homologous arm, the 3′ homologous arm, and the junctions between the 5′ homologous arm and the substitution sequence, and between the 3′ homologous arm and the substitution sequence. The homologous arms on both sides of the ssODN can be symmetrical or asymmetrical (different lengths on both sides). For ease of systematic comparison, this invention uniformly uses ssODNs with symmetrical homologous arms for experiments. The length of the homologous arms flanking the ssODN can be 20-300 nt. For ease of system comparison, this embodiment uses an ssODN with a length of approximately 60 nt flanking homologous arms as an example. The ssODN can be the sense or antisense strand of the edited region DNA (the homologous arm sequence is the same as the corresponding sense strand or the antisense strand). In this invention, the DNA strand with the same recognition site as the sgRNA is uniformly selected as the ssODN master sequence. In addition to the flanking homologous arm sequences, the substituted bases in the ssODN sequence can be 0, 1, 2, 3, 4, 5, or 6 bases, achieving the effects of deletion, substitution, insertion, or a combination of deletion, substitution, and insertion. The achieved effect can be a change of more than one base on the genome, such as the deletion or substitution of 1-20 bases. As shown in Table 3, the underlined bases are the flanking homologous arms, approximately 60 nt in length, and the bases between the flanking homologous arms that are not underlined are substituted bases. The first four nucleotides at the 5′ end and the last four nucleotides at the 3′ end of ssODN are modified with phosphate thioesters to enhance the stability of ssODN and improve its activity in gene editing.
[0169] Table 3. ssODN sequences and their applicable regions
[0170]
[0171] Using the sgRNA corresponding to the target sites in Table 4, the first 20 bp is the recognition site sequence of the sgRNA, and the following 80 nt is the universal sgRNA backbone sequence SEQ ID NO:47, (GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUU).
[0172] Table 4. sgRNA sequences and their applicable regions
[0173]
[0174]
[0175] Using commercially available spCas9 protein and chemically synthesized sgRNA (sequences shown in Table 4, with chemically modified ends, 2'-O-methyl modification of the three ribonucleotides at the 5' and 3' ends, and phosphate thioester bond modification between the four ribonucleotides at the 5' and 3' ends), a ribonucleoprotein (RNP) complex was formed in vitro. Corresponding ssODNs for each region were added as homology repair template strands. The RNP complex and ssODN were then delivered into K562 cells via electroporation. After culturing K562 cells for 48 hours, genomic DNA was extracted, and the corresponding gene fragments were amplified by PCR and sequenced. PCR primers for the corresponding sites are shown in Table 5. Sequencing results were analyzed using TIDE software; the editing efficiency results are shown in the figure below. Figure 3 It can be seen that the editing efficiency of the nine selected sites all reached more than 25%.
[0176] Table 5. PCR primer names and sequences for different sites
[0177] SEQ ID Primer name Sequence (5'-3') Applicable region SEQ ID NO: 57 HBZ-TSS1700-F GGAGGCTGAGGTATGAGAATTG TSS1700 SEQ ID NO:58 HBZ-TSS1700-R GTGAAGGGCAGGTCCAGAT TSS1700 SEQ ID NO:59 HBZ-TSS-UTR-F TGCCTCCTCCTGCTTGTCA TSS7, U96, U39 SEQ ID NO:60 HBZ-TSS-UTR-R TGGCCTTGGTAGTGCTCAG TSS7, U96, U39 SEQ ID NO:61 HBZ-N200-F GAGGAGGGAACCGTGGAGAG N200, N360 SEQ ID NO:62 HBZ-N200-R CAGTGCCCTGATCCCAGATG N200, N360 SEQ ID NO:63 HBZ-W310-F CAATGAACGAAGCAGCGTCC W310 SEQ ID NO:64 HBZ-W310-R TTTAGCAAATGAGATGCCCCG W310 SEQ ID NO:65 HBZ-W-F1 CAGAACGATCAGGACGAAGAGG W560, W630 SEQ ID NO:66 HBZ-W-R1 CATGGTGCGGATACCCTTGG W560, W630
[0178] Example 2 on CD34 + In HSPC cells, enhancer elements are introduced into the non-coding region of the ζ-globin gene.
[0179] After reviving peripheral blood-derived human CD34-positive hematopoietic stem cells (HSPCs), they were placed in X-VIVO solution rich in human cytokines (SCF, TPO, Flt3L, 100 ng / ml each). TM After two days of culture in serum-free hematopoietic cell medium (-15%), the corresponding ssODN and spCas9 / sgRNA RNPs (as in Example 1) of each region were delivered to HSPCs using the EO-100 program on the Lonza-4D electroporator. Following electroporation, the cells were cultured in X-VIVO2 enriched with human cytokines (SCF / TPO / Flt3L, 100 ng / ml each). TMAfter one day of recovery culture, the cells were then cultured in IMDM medium containing EPO, SCF, human AB serum, insulin, transferrin, heparin, IL-3, and hydrocortisone for differentiation. Forty-eight hours after electroporation, 2 × 10⁵ cells were harvested to extract genomic DNA, which was then amplified by PCR and sent for Sanger sequencing (sequencing primers for the corresponding sites were the same as those for K562 cells, as shown in Table 5). Sanger sequencing results before and after gene editing are shown in [Table 5]. Figure 4 Editing efficiency results can be found in [link / reference]. Figure 5-1 This demonstrates that high editing efficiency can also be achieved in HSPC.
[0180] After HSPC-induced differentiation into D14 cells, RNA was extracted from 1×10^6 cells and reverse transcribed into cDNA. GAPDH was used as an internal control for qPCR. The qPCR primers for the HBZ gene were: forward primer 5′CCGGTCAACTTCAAGCTCCT3′ (SEQ ID NO:67), reverse primer 5′CTCAGCGGTACTTCTCGGTC3′ (SEQ ID NO:68); and for GAPDH, forward primer 5′CCATGGGGAAGGTGAAGGTC3′ (SEQ ID NO:69), reverse primer 5′GAAGGGGTCATTGATGGCAAC3′ (SEQ ID NO:70). The results showed that the formation of an enhancer element in the TSS7 region significantly increased HBZ expression at the RNA level, achieving an increase of more than 16-fold compared to the control group (see...). Figure 5-2 A slight improvement is visible in the remaining areas.
[0181] On day 21 of HSPC-induced differentiation, 1×10^6 cells were harvested, lysed with 200 μL of RIPA lysis buffer, and incubated on ice for 30 min. The centrifuge was pre-chilled simultaneously. After lysis, the cells were centrifuged at 4°C, 13000g, for 10 min. The lysate was collected, and 5× Loading Buffer was added. The cells were then boiled in a water bath for 10 min to denature the total protein. The denatured protein was used as a backup sample for Western blot experiments, using β-actin as an internal control. The results showed that the formation of an enhancer element in the TSS7 region significantly increased HBZ protein expression (see...). Figure 6 ).
[0182] Example 3: Introducing enhancer elements into the non-coding region of the ζ-globin gene in CD34+HSPC cells
[0183] In this embodiment, NTG-N(7-8)-WGATAR was formed in other regions of the HBZ gene, and the expression level of HBZ was detected.
[0184] The locations where the reinforcing sub-elements in each region will be formed are as follows:
[0185] T0, upstream of the TSS (promoter) of the HBZ gene, between -219 and -204, the original sequence ctgtggtcagactctg (SEQ ID NO:105), is located at the locus Chr16:152636-152651;
[0186] T50, between -278 and -263 upstream of the TSS of the HBZ gene, the original sequence tggactacaaatgcag (SEQ ID NO:106) is located at the locus Chr16:152577-152592;
[0187] T70, between -282 and -298 upstream of the TSS of the HBZ gene, the original sequence gaataaggacggtgcag (SEQ ID NO: 107) is located at locus Chr16:152557-152573;
[0188] T200, between -447 and -431 upstream of the TSS of the HBZ gene, the original sequence caggaatccagagacaa (SEQ ID NO: 108) is located at the locus Chr16:152408-152424;
[0189] T300, between -525 and -508 upstream of the TSS of the HBZ gene, the original sequence ctgcttgtcaggggacag (SEQ ID NO: 109) is located at locus Chr16:152330-152347.
[0190] Table 6 shows the wild-type sequences and edited sequences of other upstream edited regions of the HBZ gene TSS; Table 7 shows the target sites recognized by the sgRNAs used in other upstream edited regions of the HBZ gene TSS; and Table 8 shows the sgRNA sequences used in other upstream edited regions of the HBZ gene TSS. CD34 was used. + HSPC cells were used in experiments, following the same cell differentiation protocol and editing method as in Example 2. The Cas9 protein was combined with the corresponding sgRNA (see Table 8) to form an RNP complex, and an ssODN sequence (see Table 9) was added as a template strand to target CD34. + HSPC was used for electroporation, and DNA was extracted 48 hours later for sequencing analysis. The editing efficiency obtained from the sequencing results is shown in [link to sequencing results]. Figure 7 Left image.
[0191] On day 21 of HSPC-induced differentiation, protein samples were obtained and subjected to Western blot experiments using the same method as in Example 2, with β-actin as the internal control protein. Results are shown below.Figure 7 The right figure shows that the expression of HBZ, a reinforcing element formed in the T50 region, is significantly upregulated.
[0192] On day 21 of HSPC-induced differentiation, cells from each group were collected for flow cytometry analysis. The specific experimental steps were as follows: 1.5 × 10^6 cells were collected from each group, washed with 1 ml PBS, centrifuged, the supernatant was discarded, and the cells were resuspended in 180 μl PBS. 20 μl of 10 × glutaraldehyde fixative was added to each group, mixed well (votex 15 s), and fixed at room temperature in the dark for 10 min. The cells were then centrifuged at 300 g for 10 min, and the supernatant was discarded. The cells were resuspended in 180 μl PBS, and 20 μl of 10 × Triton X-100 / PBS was added. The cells were incubated at room temperature for 5 min, centrifuged at 300 g for 10 min, and the supernatant was discarded. Resuspend the sample in 100 μl PBS, add HBZ antibody (Proteintech, catalog number 17284-1-AP), incubate in the dark for 30 min, then incubate with anti-rabbit AF647 antibody (catalog number #4414S, Cell Signaling) as the conjugate secondary antibody for 30 min. Wash twice with PBS, resuspend, and analyze the results. See attached image. Figure 8 Flow cytometry results showed that the formation of enhancer elements in other regions of the HBZ promoter region could activate the HBZ gene and increase HBZ protein expression, with the T50 region showing the most significant effect.
[0193] Table 6. Wild-type sequences and edited sequences from other upstream edited regions of the HBZ gene TSS.
[0194]
[0195]
[0196] Table 7. Target sites recognized by sgRNAs used in other upstream editing regions of the HBZ gene TSS
[0197] SEQ ID Region Site sgRNA recognition site (5'-3') DNA strand SEQ ID NO:81 T0 site_10 TAGACTCTGTGGTCAGACTC sense strand SEQ ID NO:82 T50 site_11 CAGAACTGGACTACAAATGC sense strand SEQ ID NO:83 T70 site_12 TCGTGATTCTGAAATGAATA sense strand SEQ ID NO:84 T200 site_13 CCACTTTGTCTCTGGATTCC antisense strand SEQ ID NO:85 T300 site_14 TGCCTCCTCCTGCTTGTCAG sense strand
[0198] Table 8. sgRNA sequences used in other upstream editing regions of the HBZ gene TSS.
[0199]
[0200] Table 9. ssODN sequences used in other upstream editing regions of the HBZ gene TSS.
[0201]
[0202]
[0203] Example 4 uses Cas12i to introduce enhancer elements into the non-coding region of the ζ-globin gene in CD34+HSPC cells.
[0204] In this embodiment, Cas12i was used to form a NAG-N(7-8)-WGATAR sequence upstream of the TSS of the HBZ gene, and the expression level of HBZ was detected. The amino acid sequence of the Cas12i protein is SEQ ID NO:110. The position to be formed for the enhancer element is between -262 and -246 upstream of the TSS (promoter) of the HBZ gene, with the original sequence gaggacttcctgggag (SEQ ID NO:111), located at the gene locus Chr16:152593-152609.
[0205] Table 10 shows the wild-type sequence and the sequence obtained after editing at the Cas12i editing site upstream of the TSS of the HBZ gene. Table 11 shows the target sites recognized by the sgRNA used at the Cas12i editing site upstream of the TSS of the HBZ gene. Table 12 shows the sgRNA sequence used at the Cas12i editing site upstream of the TSS of the HBZ gene. Table 13 shows the ssODN sequence used at the Cas12i editing site upstream of the TSS of the HBZ gene.
[0206] Table 10 Wild-type sequence and edited sequence of the Cas12i editing site in the TSS region of the HBZ gene
[0207]
[0208] Table 11 Target sites recognized by sgRNAs used at the Cas12i editing site upstream of the HBZ gene TSS
[0209]
[0210] Table 12. sgRNA sequence used at the Cas12i editing site upstream of the HBZ gene TSS.
[0211]
[0212] The universal sgRNA backbone sequence SEQ ID NO:116: AGAGAAUGUGUGCAUAGUCACAC.
[0213] Table 13 ssODN sequence used for the Cas12i editing site upstream of the TSS in the HBZ gene
[0214]
[0215] After reviving peripheral blood-derived human CD34-positive hematopoietic stem cells (HSPCs), they were placed in X-VIVO solution rich in human cytokines (SCF, TPO, Flt3L, 100 ng / ml each). TM After two days of culture in serum-free hematopoietic cell medium (-15%), the corresponding ssODNs and Cas12i / sgRNA RNPs (prepared using the same method as in Example 1) were delivered to HSPCs using the EO-100 program on a Lonza-4D electroporator. Following electroporation, the cells were then cultured in X-VIVO2+ enriched with human cytokines (SCF / TPO / Flt3L, 100 ng / ml each). TM After one day of recovery culture, the cells were then cultured in IMDM medium containing EPO, SCF, human AB serum, insulin, transferrin, heparin, IL-3, and hydrocortisone for differentiation. Forty-eight hours after electroporation, 2 × 10⁵ cells were harvested to extract genomic DNA, which was then amplified by PCR and sent for Sanger sequencing (sequencing primers for the corresponding sites were the same as those used for K562 cells, as shown in Table 5). The editing efficiency of group C1 was higher than 30%. Figure 9 ).
[0216] HSPC-induced differentiation into D14 cells was performed. RNA was extracted from 1×10^6 cells and reverse transcribed into cDNA. GAPDH was used as an internal control for qPCR. The qPCR primers for HBZ gene and GAPDH were the same as in Example 2. The results are shown in […]. Figure 10 The results showed that the formation of enhancer elements in the C1 region could significantly increase the expression of HBZ at the RNA level, reaching more than 10-fold higher expression compared with the control group.
[0217] The results of the above embodiments demonstrate that the present invention provides a gene editing method for enhancing ζ-globin gene expression, which differs from existing technologies. The present invention forms NTG-N(7-8)-WGATAR or NAG-N(7-8)-WGATAR in different regions of the ζ-globin gene. These NTG-N(7-8)-WGATAR or NAG-N(7-8)-WGATAR elements act as enhancer elements, exerting a positive regulatory effect after gene-edited cells (e.g., hematopoietic stem cells) differentiate into erythrocytes, activating or significantly enhancing the expression of the ζ-globin gene and protein. Therefore, the present invention has potential application value in gene therapy for α-hemoglobinopathies.
[0218] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
Claims
1. A method for activating ζ-globin gene expression for non-diagnostic / therapeutic purposes, characterized in that, The method includes: using gene editing technology to artificially form enhancer elements comprising NTG-N(7-8)-WGATAR, NAG-N(7-8)-WGATAR, YTATCW-N(7-8)-CAN and / or YTATCW-N(7-8)-CTN sequences in the sense or antisense strand of the non-coding region of the ζ-globin gene, wherein N is A, G, C or T, W is T or A, R is A or G, and Y is T or C; The non-coding region of the ζ-globin gene is selected from one or more of the following: a sequence within 2 kb upstream of the transcription start site, the 5′UTR region, the intron region, the 3′UTR region, and a sequence within 1 kb downstream of the 3′UTR region of the ζ-globin gene.
2. The method as described in claim 1, characterized in that, The manner in which the enhanced sub-element is formed is selected from one or more of the following: deletion sequence, insertion sequence, and replacement sequence; Preferably, the gene editing technology is used to achieve homologous recombination repair and / or insertion / deletion mutations; The gene editing technology is selected from one or more of the following: CRISPR-Cas editing, TALEN editing, ZFN editing, and Argonaute editing; More preferably, the CRISPR-Cas editing is selected from one or more of homologous recombination repair editing, indel editing, lead editing, and single-base editing.
3. The method as described in claim 2, characterized in that, The non-coding region of the ζ-globin gene includes: the promoter region and the 5′UTR region of the ζ-globin gene, located at the gene locus Chr16:149855-152909; Preferably, the non-coding region of the ζ-globin gene is selected from the loci located as follows: Chr16:150852-150971, Chr16:152574-152693, Chr16:152622-152741, Chr16:152678-152797, Chr16:152587-152706, Chr16:152528-152647, Chr16:152499-152618, Chr16:152353-152472, Chr16:152278-152398, Chr16:152544-152658, Chr16:152565-15264.
5. Chr16:152568-152645, Chr16:152571-152645, Chr16:152574-152645 , Chr16:152577-152645, Chr16:152580-152645, Chr16:152583-152645, C hr16:152586-152645、Chr16:152589-152645、Chr16:152592-152645、Chr 16:152595-152645, Chr16:152598-152645, Chr16:152601-152645, Chr16 :152604-152645, Chr16:152607-152645, Chr16:152610-152645, Chr16: 152613-152645, Chr16:152616-152645, Chr16:152619-152645, Chr16:15 2622-152645, Chr16:152625-152645, Chr16:152628-152645, Chr16:1525 77-152592, Chr16:152593-152609, Chr16:152620-152636, Chr16:152565 -152642, Chr16:152565-152639, Chr16:152565-152636, Chr16:152565-1 52633, Chr16:152565-152630, Chr16:152565-152627, Chr16:152565-152 624. Chr16:152565-152621, Chr16:152565-152618, Chr16:152565-15261 5. Chr16:152565-152612, Chr16:152565-152609, Chr16:152565-152606,One or more of the sequences Chr16:152565-152603, Chr16:152565-152600, Chr16:152565-152597, Chr16:152565-152594, Chr16:152565-152591, Chr16:152565-152588, Chr16:152565-152585, Chr16:152565-152582, and Chr16:152577-152636; More preferably, the non-coding region of the ζ-globin gene is selected from one or more sequences located at the locus of: Chr16:150892-150921, Chr16:152615-152644, Chr16:152671-152700, Chr16:152731-152760, Chr16:152629-152658, Chr16:152570-152599, Chr16:152550-152579, Chr16:152401-152430, Chr16:152323-152352, and Chr16:152587-152616; More preferably, the non-coding region of the ζ-globin gene is selected from one or more sequences located at the locus of: Chr16:150901-150916, Chr16:152620-152636, Chr16:152680-152696, Chr16:152737-152753, Chr16:152636-152651, Chr16:152577-152592, Chr16:152557-152573, Chr16:152408-152424, Chr16:152330-152347, and Chr16:152593-152609.
4. The method as described in claim 2, characterized in that, The non-coding region of the ζ-globin gene includes one or more of the following: an intron region, a 3′UTR region, and a sequence within 1 kb downstream of the 3′UTR region, located at the gene locus as Chr16:153005-153891 or Chr16:154401-155505. Preferably, the sequence of the non-coding region of the ζ-globin gene is located at the locus as Chr16:153061-153380, Chr16:153221-153540, Chr16:154667-154986, Chr16:154917-155236 and / or Chr16:154976-155295; More preferably, the sequence of the non-coding region of the ζ-globin gene is located at one or more of the following sequences at the locus: Chr16:153161-153280, Chr16:153321-153440, Chr16:154767-154886, Chr16:155017-155136, and Chr16:155076-155195; More preferably, the sequence of the non-coding region of the ζ-globin gene is located at one or more of the following sequences at the locus: Chr16:153213-153242, Chr16:153371-153400, Chr16:154815-154844, Chr16:155057-155086, and Chr16:155127-155156; For example, the sequence of the non-coding region of the ζ-globin gene is located at one or more of the following sequences at the locus: Chr16:153219-153234, Chr16:153378-153393, Chr16:154823-154838, Chr16:155064-155080, and Chr16:155133-155149.
5. The method according to any one of claims 1 to 4, characterized in that, The sequence shown in SEQ ID NO:1 in the positive strand of the non-coding region of the ζ-globin gene is edited to any of the sequences shown in SEQ ID NO:2 to 5; preferably, the non-coding region of the ζ-globin gene is located at Chr16:150852-150971; and / or, The sequence shown in SEQ ID NO:6 in the positive strand of the non-coding region of the ζ-globin gene is edited to any of the sequences shown in SEQ ID NO:7-10; preferably, the non-coding region of the ζ-globin gene is located at Chr16:152574-152693; and / or, The sequence shown in SEQ ID NO:11 in the positive strand of the non-coding region of the ζ-globin gene is edited to the sequence shown in SEQ ID NO:12 or 13; preferably, the non-coding region of the ζ-globin gene is located at Chr16:152622-152741; and / or, The sequence shown in SEQ ID NO:14 in the positive strand of the non-coding region of the ζ-globin gene is edited to the sequence shown in SEQ ID NO:15; preferably, the non-coding region of the ζ-globin gene is located at Chr16:152678-152797; and / or, The sequence shown in SEQ ID NO:16 in the positive strand of the non-coding region of the ζ-globin gene is edited to a sequence shown in any of SEQ ID NO:17-19; preferably, the non-coding region of the ζ-globin gene is located at Chr16:153161-153280; and / or, The sequence shown in SEQ ID NO:20 in the positive strand of the non-coding region of the ζ-globin gene is edited to the sequence shown in SEQ ID NO:21; preferably, the non-coding region of the ζ-globin gene is located at Chr16:153321-153440; and / or, The sequence shown in SEQ ID NO:22 in the antisense strand of the non-coding region of the ζ-globin gene is edited to the sequence shown in SEQ ID NO:23; preferably, the non-coding region of the ζ-globin gene is located at Chr16:154767-154886; and / or, The sequence shown in SEQ ID NO:24 in the positive strand of the non-coding region of the ζ-globin gene is edited to the sequence shown in SEQ ID NO:25; preferably, the non-coding region of the ζ-globin gene is located at Chr16:155017-155136; and / or, The sequence shown in SEQ ID NO:26 in the positive strand of the non-coding region of the ζ-globin gene is edited to the sequence shown in SEQ ID NO:27 or 28; preferably, the non-coding region of the ζ-globin gene is located at Chr16:155076-155195; and / or, The sequence shown in SEQ ID NO:71 in the positive strand of the non-coding region of the ζ-globin gene is edited to the sequence shown in SEQ ID NO:72; preferably, the non-coding region of the ζ-globin gene is located at Chr16:152587-152706; and / or, The sequence shown in SEQ ID NO:73 in the positive strand of the non-coding region of the ζ-globin gene is edited to the sequence shown in SEQ ID NO:74; preferably, the non-coding region of the ζ-globin gene is located at Chr16:152528-152647; and / or, The sequence shown in SEQ ID NO:75 in the positive strand of the non-coding region of the ζ-globin gene is edited to the sequence shown in SEQ ID NO:76; preferably, the non-coding region of the ζ-globin gene is located at Chr16:152499-152618; and / or, The sequence shown in SEQ ID NO:77 in the positive strand of the non-coding region of the ζ-globin gene is edited to the sequence shown in SEQ ID NO:78; preferably, the non-coding region of the ζ-globin gene is located at Chr16:152353-152472; and / or, The sequence shown in SEQ ID NO:79 in the positive strand of the non-coding region of the ζ-globin gene is edited to the sequence shown in SEQ ID NO:80; preferably, the non-coding region of the ζ-globin gene is located at the gene locus Chr16:152278-152398.
6. A gRNA, characterized in that, The gRNA comprises a recognition site sequence and a backbone sequence. The recognition site sequence is partially or completely complementary to the sense or antisense strand of the non-coding region of the ζ-globin gene. The gRNA is used to artificially form NTG-N(7-8)-WGATAR, NAG-N(7-8)-WGATAR, YTATCW-N(7-8)-CAN, and / or YTATCW sequences on the sense or antisense strand of the non-coding region of the ζ-globin gene. - Enhancer elements of the N(7-8)-CTN sequence, wherein N is A, G, C or T, W is T or A, R is A or G, and Y is T or C; The non-coding region of the ζ-globin gene is as defined in the method described in any one of claims 1 to 4.
7. The gRNA as described in claim 6, characterized in that, The backbone sequence is as shown in SEQ ID NO:47 or 116; and / or, the DNA sequence corresponding to the recognition site sequence is partially or completely identical to any of the sequences shown in SEQ ID NO:29-37, SEQ ID NO:81-85 and SEQ ID NO:
114. Preferably, the nucleotide sequence of the gRNA is as shown in any one of SEQ ID NO:48-56, SEQ ID NO:86-90 and SEQ ID NO:115; And / or, the gRNA further comprises chemical modifications, preferably one or more of 3'-thiophosphate, 2'-O-methyl ester, 2'-O-methyl, 2'-F modification, 2'-ribose 3'-thiophosphate, deoxy, and 5' phosphate modification; more preferably, the three ribonucleotides at the 5' and 3' ends of the gRNA are modified with 2'-O-methyl, and the four ribonucleotides at the 5' and 3' ends are modified with thiophosphate bonds.
8. An ssODN for gene editing of the non-coding region of the ζ-globin gene, characterized in that, The ssODN includes NTG-N(7-8)-WGATAR sequence, NAG-N(7-8)-WGATAR sequence, YTATCW-N(7-8)-CAN sequence and / or YTATCW-N(7-8)-CTN sequence; N is A, G, C or T, W is T or A, R is A or G, and Y is T or C; Wherein, the non-coding region of the ζ-globin gene is defined as in the method described in any one of claims 1 to 4; Preferably, the ssODN comprises a sequence selected from any one of SEQ ID NO:2-5, SEQ ID NO:7-10, SEQ ID NO:12-13, SEQ ID NO:15, SEQ ID NO:17-19, SEQ ID NO:21, SEQ ID NO:23, SEQ ID NO:25, SEQ ID NO:27-28, SEQ ID NO:72, SEQ ID NO:74, SEQ ID NO:76, SEQ ID NO:78, SEQ ID NO:80 and SEQ ID NO:113, or a complementary sequence thereof; More preferably, the ssODN comprises a sequence selected from any one of SEQ ID NO:38-46, SEQ ID NO:91-95 and SEQ ID NO:117, or a complementary sequence thereof.
9. A composition for gene editing of the non-coding region of the ζ-globin gene, characterized in that, The composition is used to artificially form enhancer elements comprising NTG-N(7-8)-WGATAR, NAG-N(7-8)-WGATAR, YTATCW-N(7-8)-CAN and / or YTATCW-N(7-8)-CTN sequences in the sense or antisense strand of the non-coding region of the ζ-globin gene, wherein N is A, G, C or T, W is T or A, R is A or G, and Y is T or C; The non-coding region of the ζ-globin gene is as defined in the method described in any one of claims 1 to 4; The composition comprises one or more selected from: CRISPR-Cas editing system, TALEN editing system, ZFN editing system and Argonaute editing system; Preferably, the CRISPR-Cas editing system is used to perform one or more of homologous recombination repair editing, indel editing, lead editing, and single-base editing.
10. The composition according to claim 9, characterized in that, The composition comprises: (a) an ssODN containing the said reinforcing sub-element; (b) targeting the non-coding region of the ζ-globin gene with gRNA; as well as, (c) CRISPR / Cas nuclease, mRNA encoding the CRISPR / Cas nuclease and / or plasmid expressing the CRISPR / Cas nuclease; Preferably, the composition satisfies one or more of the following conditions: (1) The CRISPR / Cas nuclease comprises at least one or a combination of the following groups: Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5d, Cas5t, Cas5h, Cas5a, Cas6, Cas7, Cas8, Cas8a, Cas8b, Cas8c, Cas9, Cas10, Cas10d, Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12f / CasZ, Cas12g, Cas12h, Cas12i, Csy1, Csy2, Csy3, Csy4, Cse1, Cse2, Cse3, Cse4, Cse5e, Csc1, Csc2, Csa5, Csn1, Cs n2, Csm1, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Cs x17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx1S, Csx11, Csf1, Csf2, CsO, Csf4, Csd1, Cs d2, Cst1, Cst2, Csh1, Csh2, Csa1, Csa2, Csa3, Csa4, Csa5, Cas13a, Cas13b, Cas13c, Cas13d, Cas13e, Cas13f, fragments thereof, and variants thereof or fragments of variants thereof; said variants are preferably selected from at least one of: SpRY variant, NG-nCas9 variant and NGG-nCas9 variant; (2) The gRNA as described in claim 6 or 7; and, (3) The ssODN is as described in claim 8.
11. The composition according to claim 9, characterized in that, The composition comprises: (a) a gRNA targeting the non-coding region of the ζ-globin gene; and (b) a CRISPR / Cas nuclease, an mRNA encoding the CRISPR / Cas nuclease, and / or a plasmid expressing the CRISPR / Cas nuclease; and / or, The composition comprises: (a) a PEgRNA containing the enhancer element and targeting the non-coding region of the ζ-globin gene; and (b) nCas9 and reverse transcriptase, mRNA encoding said nCas9 and reverse transcriptase and / or plasmids expressing said nCas9 and reverse transcriptase; and / or, The composition comprises: (a) a gRNA targeting the non-coding region of the ζ-globin gene; and (b) A single-base editor, an mRNA encoding the single-base editor, and / or a plasmid expressing the single-base editor.
12. A cell, characterized in that, The cells comprise the composition as described in any one of claims 9 to 11; and / or, The cells were obtained by editing the composition according to any one of claims 9 to 11; Preferably, the cell is an erythroid progenitor cell; and / or, the cell is a mammalian cell, such as a human cell; More preferably, the erythroid progenitor cells are umbilical cord blood stem cells, induced pluripotent stem cells, hematopoietic stem / progenitor cells, myeloid progenitor cells, bursting unit-erythroid cells / red blood cells, spleen colony-forming cells, embryonic cell colony-forming cells, and / or megakaryocyte-erythroid progenitor cells.
13. The use of the gRNA as described in claim 6 or 7, the ssODN as described in claim 8, the composition as described in any one of claims 9 to 11, and / or the cell as described in claim 12 in the preparation of a medicament for treating α-thalassemia; Preferably, the alpha-thalassemia is moderate or severe; and / or, the drug is cellular.