Engineered gene transcription repressional tool targeting inhibin subunit βe (INHBE) and use thereof
By designing a complex containing DNA methylation and transcriptional repressor domains, which specifically binds to the INHBE gene regulatory element, the unknowns and limitations of existing tools are overcome, achieving efficient silencing regulation of the INHBE gene, reducing the risk of obesity-related diseases and muscle atrophy, and providing a well-tolerated obesity treatment option.
Patent Information
- Application Number
- PCT/CN2025/102298
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-21
- Filing Date
- 2025-06-20
- Publication Date
- 2025-12-26
AI Technical Summary
Existing epigenetic editing tools targeting the INHBE gene have not been able to effectively and sustainably suppress its expression, and there are unknowns and limitations. This leads to side effects such as weight rebound and muscle atrophy in obesity treatment, and GLP-1 receptor agonists have poor tolerability after long-term use.
Design a complex comprising a DNA methylation domain and a transcriptional repressor domain to specifically bind to the transcriptional regulatory elements of the INHBE gene. Utilize the interaction between the nucleic acid binding domain and the recruitment domain to achieve silencing regulation of the INHBE gene, avoiding the risks of DNA cleavage and immunogenicity.
It achieves efficient silencing regulation of the INHBE gene, significantly reduces the risk of obesity-related diseases, provides long-term weight control, and avoids muscle loss and tolerance issues.
Smart Images

Figure PCTCN2025102298-FTAPPB-I100001 
Figure PCTCN2025102298-FTAPPB-I100002 
Figure PCTCN2025102298-FTAPPB-I100003
Abstract
Description
Engineered gene transcriptional repression tools targeting the repressin subunit βE (INHBE) and their applications Technical Field
[0001] This application relates to the field of biomedicine, specifically to a complex for regulating the expression of inhibin subunit βE (INHBE) and its uses. Background Technology
[0002] Obesity is a complex metabolic disease characterized by the excessive accumulation of body fat, which not only affects an individual's appearance and self-perception but also significantly increases the risk of cardiovascular disease, diabetes, and certain types of cancer. According to estimates by the World Health Organization, more than 2 billion adults worldwide are overweight or obese, placing a tremendous burden on global public health systems.
[0003] Currently, treatments for obesity mainly include lifestyle modifications, medication, and surgery. However, these methods often fail to achieve long-term weight control and are frequently accompanied by relapse. In recent years, the application of GLP-1 (glucagon-like peptide-1) receptor agonists in obesity treatment has become an important research area. Clinical trials have confirmed that GLP-1 receptor agonists can effectively reduce weight and improve multiple cardiovascular and metabolic risk factors in obese patients. For example, some GLP-1 receptor agonists on the market, such as liraglutide and semaglutide, have been approved for obesity treatment and have shown good tolerability and safety. Studies have shown that these drugs have some problems during weight loss. For example, semaglutide significantly reduces lean body mass, with approximately 40% of weight loss coming from bone and muscle loss. In addition, GLP-1R agonists suppress the reward system, have poor tolerability, and 68% of people no longer tolerate them after one year. Most importantly, most patients experience rapid weight rebound after discontinuing GLP-1R agonists, leading to long-term drug dependence. Therefore, researchers are still searching for weight-loss drugs that are well-tolerated, have no risk of weight rebound, and do not have side effects such as muscle atrophy.
[0004] Inhibin subunit βE (INHBE) is a member of the transforming growth factor-β (TGF-β) superfamily and is highly specifically expressed in hepatocytes. Early studies found that INHBE may be a liver factor that alters systemic metabolic status under conditions of obesity and insulin resistance. Its liver expression level is positively correlated with insulin resistance and body mass index in humans. In 2022, two papers published by Alnylam Pharmaceuticals and the Regeneron Genetics Center (RGC) formally linked INHBE to weight loss targets. These two articles revealed the close relationship between INHBE and lipid regulation; that is, rare loss-of-function (pLOF) mutations in the INHBE protein protect patients from liver inflammation, dyslipidemia, and type 2 diabetes by promoting healthy fat storage. Patients carrying this mutation have more normal fat distribution, significantly reduced abdominal fat, good metabolic status, and a significantly reduced risk of cardiovascular disease and type 2 diabetes. Therefore, as a liver-specific negative regulator of fat storage, developing nucleic acid drugs targeting INHBE using liver-specific delivery technology to inhibit INHBE expression has become a potentially advantageous strategy for treating metabolic diseases related to improper fat distribution and storage.
[0005] Epigenetic editing of the INHBE gene can be achieved by introducing epigenetic modifications into specific regulatory regions or sites within its genome, altering chromatin structure and thus repressing INHBE transcription, thereby silencing INHBE. More importantly, this process avoids DNA cleavage, preventing the possibility of double-strand breaks and fundamentally eliminating the risk of activating unpredictable DNA repair mechanisms and potentially generating immunogenic truncated or mutant proteins. However, current epigenetic editing tools targeting the INHBE genome are still in the exploratory stage, and the development of epigenetic editing technologies capable of sustaining INHBE suppression faces many unknowns and limitations. Summary of the Invention
[0006] On one hand, this application provides a complex comprising a first fusion and a second fusion, wherein: 1) one of the first fusion and the second fusion comprises a DNA methylation domain and at least one recruitment domain A, and the other fusion comprises a transcriptional repressor domain and at least one recruitment domain A'; and 2) the first fusion or the second fusion comprises a nucleic acid binding domain; and the recruitment domain A and the recruitment domain A' are capable of interacting to recruit a portion of one of the first fusion and the second fusion to the vicinity of the other fusion; the nucleic acid binding domain is capable of specifically binding to an INHBE gene transcriptional regulatory element, the INHBE gene transcriptional regulatory element comprising a transcription start site, a core promoter, a promoter, an enhancer, a silencer, an insulator element, a boundary element, and / or a locus control region, and the nucleic acid binding domain is capable of specifically binding to a target nucleotide sequence within a region 1500 bp upstream to 2000 bp downstream of the INHBE gene transcription start site.
[0007] In some implementations, the nucleic acid binding domain is a DNA binding domain.
[0008] In some implementations, the DNA-binding domain is selected from: TALE domain, zinc finger domain, tetR domain, a wide range of nucleases, Cas protein, Argonaute (Ago) protein, and their homologues, modified forms, or variants.
[0009] In some implementations, the DNA-binding domain is capable of binding the target nucleotide sequence.
[0010] In some implementations, the DNA-binding domain is capable of binding to guide RNA.
[0011] In some implementations, the guide RNA is capable of specifically recognizing and hybridizing with the target nucleotide sequence.
[0012] In some embodiments, the guide RNA is capable of specifically recognizing and hybridizing with target nucleotide sequences within a region approximately 1500 bp, 1000 bp, 750 bp, 500 bp, 250 bp, 200 bp, 100 bp, and 50 bp upstream of the transcription start site of the INHBE gene.
[0013] In some embodiments, the guide RNA is capable of specifically recognizing and hybridizing with target nucleotide sequences within a region approximately 2000 bp, 1500 bp, 1000 bp, 750 bp, 500 bp, 250 bp, 200 bp, 100 bp, and 50 bp downstream of the transcription start site of the INHBE gene.
[0014] In some embodiments, the guide RNA is capable of specifically recognizing and hybridizing with target nucleotide sequences within the regions approximately 1000 bp upstream to approximately 1500 bp downstream, approximately 800 bp upstream to approximately 1500 bp downstream, approximately 800 bp upstream to approximately 1000 bp downstream, approximately 800 bp upstream to approximately 800 bp downstream, and approximately 750 bp upstream to approximately 750 bp downstream of the INHBE gene.
[0015] In some implementations, the guide RNA is capable of specifically recognizing and hybridizing with the target nucleotide sequence within a region approximately 750 bp upstream of the INHBE gene transcription start site.
[0016] In some implementations, the guide RNA is capable of specifically recognizing and hybridizing with the target nucleotide sequence within a region approximately 750 bp downstream of the INHBE gene transcription start site.
[0017] In some implementations, the guide RNA is capable of specifically recognizing and hybridizing with a target nucleotide sequence in a region approximately 400 bp to approximately 500 bp downstream of the transcription start site of the INHBE gene.
[0018] In some implementations, the guide RNA is capable of specifically recognizing and hybridizing with a target nucleotide sequence in a region approximately 450 bp to approximately 500 bp downstream of the transcription start site of the INHBE gene.
[0019] In some implementations, the guide RNA is capable of specifically recognizing and hybridizing with a target nucleotide sequence in a region approximately 750 bp to 1500 bp downstream of the transcription start site of the INHBE gene.
[0020] In some embodiments, the guide RNA comprises SEQ ID NOs: 337-382, 451, 452, 453, 455-458, 465, 466, 468-470, 473, 475, 477, 478, 480-482, 484, 485, 487, 489-495, 497, 498, 502, 503, 505, 507-510, 512, 513, 515, 517-525, 528, 532, 533, 537, 540, 402-406, 413, 416, 423, 424, 426, 42 7. The nucleotide sequence shown in any one of the following: 434-437, 442, 443, 448, 454, 459-464, 467, 472, 474, 476, 479, 488, 496, 499-501, 504, 506, 511, 514, 516, 529-531, 534, 535, 407, 409, 418-422, 425, 428, 430-433, 438-441, 444-447, 449, 541, 542, 545, 547-549.
[0021] In some embodiments, the guide RNA comprises SEQ ID NOs: 337-382, 451, 452, 453, 455-458, 465, 466, 468-470, 473, 475, 477, 478, 480-482, 484, 485, 487, 489-495, 497, 498, 502, 503, 505, 507-510, 512, 513, 515, 517-525, 528, 532, 533, 537, 540, 402-406, 413, 416, 423, 424, 426, 427, 434-437, 44 2. A partial sequence of any one of the nucleotide sequences shown in 443, 448, 454, 459-464, 467, 472, 474, 476, 479, 488, 496, 499-501, 504, 506, 511, 514, 516, 529-531, 534, 535, 407, 409, 418-422, 425, 428, 430-433, 438-441, 444-447, 449, 541, 542, 545, 547-549, wherein the partial sequence is 15-20 base pairs in length.
[0022] In some embodiments, the DNA-binding domain is a Cas protein, and the Cas protein is a type II Cas nuclease.
[0023] In some embodiments, the Cas protein is selected from type II Cas nucleases and type II V Cas nucleases.
[0024] In some embodiments, the Cas protein is a Cas9 or Cas12 protein.
[0025] In some embodiments, the Cas protein is an inactivated Cas9 (dCas9) protein or an inactivated Cas12 (dCas12) protein.
[0026] In some embodiments, the DNA-binding domain comprises an amino acid sequence shown in any one of SEQ ID NO:1-9.
[0027] In some embodiments, the first fusion comprises a DNA methylation domain, a nucleic acid binding domain, and at least one recruitment domain A, and the second fusion comprises a transcriptional repressor domain and at least one recruitment domain A'.
[0028] In some implementations, the first fusion compound comprises, from the N-terminus to the C-terminus, a DNA methylation domain, a nucleic acid binding domain, and a recruitment domain A.
[0029] In some embodiments, the second fusion compound contains a transcriptional repressor domain and a recruitment domain A' from the N-terminus to the C-terminus, or it contains a recruitment domain A' and a transcriptional repressor domain from the N-terminus to the C-terminus.
[0030] In some embodiments, the first fusion comprises a transcriptional repressor domain, a nucleic acid binding domain, and at least one recruitment domain A, and the second fusion comprises a DNA methylation domain and at least one recruitment domain A'.
[0031] In some implementations, the first fusion compound comprises, from the N-terminus to the C-terminus, a recruitment domain A, a nucleic acid binding domain, and a transcriptional repressor domain.
[0032] In some embodiments, the second fusion comprises, from the N-terminus to the C-terminus, a DNA methylation domain and a recruitment domain A', or from the N-terminus to the C-terminus, a recruitment domain A' and a DNA methylation domain.
[0033] In some embodiments, the complex is characterized in that: 1) the first fusion compound comprises, from N-terminus to C-terminus, a DNA methylation domain, a nucleic acid binding domain, and a recruitment domain A, and the second fusion compound comprises, from N-terminus to C-terminus, a transcriptional repressor domain and a recruitment domain A'; or 2) the first fusion compound comprises, from N-terminus to C-terminus, a DNA methylation domain, a nucleic acid binding domain, and a recruitment domain A, and the second fusion compound comprises, from N-terminus to C-terminus, a recruitment domain A' and a transcriptional repressor domain; or 3) the first fusion compound comprises, from N-terminus to C-terminus, a recruitment domain A, a nucleic acid binding domain, and a transcriptional repressor domain, and the second fusion compound comprises, from N-terminus to C-terminus, a DNA methylation domain and a recruitment domain A'; or 4) the first fusion compound comprises, from N-terminus to C-terminus, a recruitment domain A, a nucleic acid binding domain, and a transcriptional repressor domain, and the second fusion compound comprises, from N-terminus to C-terminus, a recruitment domain A' and a DNA methylation domain.
[0034] In some embodiments, the recruitment domain A is selected from one of two groups of domains, and the recruitment domain A' is selected from the other of two groups of domains: 1) universal control non-derepressor protein 4 (GCN4), a GFP11 fragment derived from split green fluorescent protein (GFP), or a GVKESLV polypeptide; and 2) a single-chain antibody (scFv), a GFP1-10 fragment derived from split green fluorescent protein (GFP), or a PDZ protein domain.
[0035] In some implementations, wherein: 1) one of the recruitment domains A and A' is a GCN4 domain and the other is an scFv domain; or 2) one of the recruitment domains A and A' is a GFP11 fragment and the other is a GFP1-10 domain; or 3) one of the recruitment domains A and A' is a GVKESLV domain and the other is a PDZ protein domain.
[0036] In some embodiments, the DNA methylation domain comprises at least one DNA methyltransferase or a functionally active fragment thereof.
[0037] In some embodiments, the DNA methyltransferase is selected from DNMT3A, DNMT3B, DNMT3C, DNMT1, DNMT2, and DNMT3L.
[0038] In some embodiments, the DNA methylation domain comprises at least one DNMT3A and at least one DNMT3L.
[0039] In some embodiments, the DNA methyltransferase comprises the amino acid sequence shown in any one of SEQ ID NO:19-24.
[0040] In some embodiments, the DNA methylation domain comprises a DNMT3A-DNMT3L domain or a DNMT3L-DNMT3A domain; wherein, - indicates that the domains at both ends are directly or indirectly connected in order from the N-terminus to the C-terminus.
[0041] In some embodiments, the transcriptional repressor is selected from one or more of the following domains: KRAB, ZIM3KRAB, ZNF680, ZNF554, ZNF264, ZNF582, ZNF324, ZNF669, ZNF354A, ZNF82, ZNF595, ZNF419, ZNF566, ZIM2, EHMT2, SUV39H1, ZFPM1, TRIM28, EZH2, MXD1, SID, LSD1, HP1a, HDAC3, HDAC1, PRMT1, SETDB1, hSIRT1, ZNF436, ZNF257, ZNF675, ZNF490, ZNF320, ZN F331, ZNF816, ZNF41, ZNF189, ZNF528, ZNF543, ZNF140, ZNF610, ZNF350, ZNF8, ZNF30, ZNF98, ZNF677, ZNF596, ZNF214, ZNF37A, ZNF34, ZNF250, ZNF547, ZNF273, ZFP82, ZNF224, ZNF33A, ZNF45, ZNF175, ZNF184, ZFP28-1, ZFP28-2, ZNF18, ZNF213, ZNF394, ZFP1, ZFP14, ZNF416, ZNF557, ZNF729, ZNF254, ZNF 764, ZNF785, ZNF10, CBX5, RYBP, YAF2, MGA, CBX1, SCMH1, MPP8, SUMO3, HERC2, BIN1, PCGF2, TOX, FOXA1, FOXA2, IRF2BP1, IRF2BP2, IRF2BPLIRF-2BP1_2 N-terminaldomain, HOXA13, HOXB13, HOXC13, HOXA11, HOXC11, HOXC10, HOXA10, HOXB9, HOXA9, ZFP28, ZN334, ZN568, ZN37A, ZN181, ZN510, ZN862, ZN140 , ZN208, ZN248, ZN571, ZN699, ZN726, ZIK1, ZNF2, Z705F, ZNF14, ZN471, ZN6 24. ZNF84, ZNF7, ZN891, ZN337, Z705G, ZN529, ZN729, ZN419, Z705A, ZN302, Z N486, ZN621, ZN688, ZN33A, ZN554, ZN878, ZN772, ZN224, ZN184, ZN544, ZNF 57, ZN283, ZN549, ZN211, ZN615, ZN253, ZN226, ZN730, Z585A, ZN732, ZN681,ZN667,ZN649,ZN470,ZN484,ZN431,ZN382,ZN254,ZN124,ZN607,ZN317,ZN620,ZN141,ZN584,ZN540,ZN75D,ZN555,ZN658,ZN684,RBAK,ZN829,ZN582,ZN112,ZN716,HKR1,ZN350,ZN480,ZN416,ZNF92,ZN100,ZN736,ZNF74,ZN443,ZN195,ZN530,ZN782,ZN791,ZN331,Z354C,ZN157,ZN727,ZN550,ZN793,ZN235,ZN724,ZN573,ZN577,ZN789,ZN718,ZN300,ZN383,ZN429,ZN677,ZN850,ZN454,ZN257,ZN264,ZN485,ZN737,ZNF44,ZN596,ZN565,ZN543,ZFP69,SUMO1,ZNF12,ZN169,ZN433,ZN175,ZN347,ZNF25,ZN519,Z585B,ZN517,ZN846,ZN230,ZNF66,ZN713,ZN816,ZN426,ZN674,ZN627,ZNF20,Z587B,ZN316,ZN233,ZN611,ZN556,ZN234,ZN560,ZNF77,ZN682,ZN614,ZN785,ZN445,ZFP30,ZN225,ZN551,ZN610,ZN528,ZN284,ZN418,ZN490,ZN805,Z780B,ZN763,ZN285,ZNF85,ZN223,ZNF90,ZN557,ZN425,ZN229,ZN606,ZN155,ZN222,ZN442,ZNF91,ZN135,ZN778,ZN534,ZN586,ZN567,ZN440,ZN583,ZN441,ZNF43,ZN589,ZN563,ZN561,ZN136,ZN630,ZN527,ZN333,Z324B,ZN786,ZN709,ZN792,ZN599,ZN613,ZF69B,ZN799,ZN569,ZN564,ZN546,ZFP92,ZN723,ZN439,ZFP57,ZNF19,ZN404,ZN274,CBX3,ZN250,ZN570,ZN675,ZN695,ZN548,ZN132,ZN738,ZN420,ZN626,ZN559,ZN460,ZN268,ZN304,ZN605,ZN844,SUMO5,ZN101,ZN783,ZN417,ZN182,ZN823,ZN177,ZN197,ZN717,ZN669,ZN256,ZN251,CBX4,CDY2,CDYL2,ZN562,ZN461,Z324A,ZN766,ID2,ZN214,CBX7,ID1,CREM,SCX,ASCL1,ZN764,SCML2,TWST1,CREB1,TERF1,ID3,CBX8,GSX1,NKX22,ATF1,TWST2,ZNF17,TOX3,TOX4,ZMYM3,I2BP1,RHXF1,SSX2,I2BPL,ZN680,TRI68,HXA13,PHC3,TCF24,HXB13,HEY1,PHC2,ZNF81,FIGLA,SAM11,KMT2B,HEY2,JDP2,HXC13,ASCL4,HHEX,GSX2,ETV7,ASCL3,PHC1,OTP,I2BP2,VGLL2,HXA11,PDLI4,ASCL2,CDX4,ZN860,LMBL4,PDIP3,NKX25,CEBPB,ISL1,CDX2,PROP1,SIN3B,SMBT1,HXC11,HXC10,PRS6A,VSX1,NKX23,MTG16,HMX3,HMX1,KIF22,CSTF2,CEBPE,DLX2,PPARG,PRIC1,UNC4,BARX2,ALX3,TCF15,TERA,VSX2,HXD12,CDX1,TCF23,ALX1,HXA10,RX,CXXC5,SCML1,NFIL3,DLX6,MTG8,CEBPD,SEC13,FIP1,ALX4,LHX3,PRIC2,MAGI3,NELL1,PRRX1,MTG8R,RAX2,DLX3,DLX1,NKX26,NAB1,SAMD7,PITX3,WDR5,MEOX2,NAB2,DHX8,CBX6,EMX2,CPSF6,HXC12,KDM4B,LMBL3,PHX2A,EMX1,NC2B,DLX4,SRY,ZN777,ZN398,GATA3,BSH,SF3B4,TEAD1,TEAD3,RGAP1,PHF1,GATA2,FOXO3,ZN212,IRX4,ZBED6,LHX4,SIN3A,RBBP7,NKX61,R51A1,MB3L1,DLX5,NOTC1,TERF2,ZN282,RGS12,ZN840,SPI2B,PAX7,NKX62,ASXL2,FOXO1,GATA1,ZMYM5, LRP1, MIXL1, SGT1, LMCD1, CEBPA, SOX14, WTIP, PRP19, NKX11, RBBP4, DMRT2, SMCA2, and their functionally active fragments.
[0042] In some embodiments, the transcriptional repressor domain comprises the amino acid sequence shown in any one of SEQ ID NOs:25-50.
[0043] In some embodiments, wherein: 1) one of the first fusion and the second fusion comprises a DNA methylation domain -dCas9 or dCas12-n×GCN4, and the other fusion comprises a transcriptional repressor domain -scFv; or 2) one of the first fusion and the second fusion comprises a DNA methylation domain -dCas9 or dCas12-scFv, and the other fusion comprises a transcriptional repressor domain -GCN4; or 3) one of the first fusion and the second fusion comprises a DNA methylation domain -dCas9 or dCas12-scFv, and the other fusion comprises a transcriptional repressor domain -GCN4; One of the fusions contains a DNA methylation domain -dCas9 or dCas12-n×GFP11, and the other fusion contains a transcription repressor domain -GFP1-10; or 4) one of the first fusions and the second fusion contains a DNA methylation domain -dCas9 or dCas12-n×GFP11, and the other fusion contains a transcription repressor domain -GFP11; or 5) one of the first fusions and the second fusion contains a DNA methylation domain -dCas9 or dCas12-n×GFP11. Cas12-n×GCN4, and another fusion thereof contains a scFv-transcriptional repressor domain; or 6) a fusion of one of the first and second fusions contains a DNA methylation domain -dCas9 or dCas12-scFv, and another fusion thereof contains a GCN4-transcriptional repressor domain; or 7) a fusion of one of the first and second fusions contains a DNA methylation domain -dCas9 or dCas12-n×GFP11, and another fusion thereof contains GFP1-1 0-transcription repressor domain; or 8) one of the first fusion and the second fusion contains a DNA methylation domain -dCas9 or dCas12-GFP1-10, and the other fusion contains a GFP11-transcription repressor domain; wherein, - indicates that the domains at both ends are directly or indirectly connected in order from the N-terminus to the C-terminus; n×GCN4 or n×GFP11 respectively represent n copies of GCN4 connected by adapter sequences or n copies of GFP11 connected by adapter sequences, where n is selected from any integer from 1 to 20.
[0044] In some embodiments, the first fusion and / or the second fusion comprises the amino acid sequence of any one of SEQ ID NO: 51-76, 78-82, 85-93, 103-105, 110-115, 123 and 124.
[0045] In some embodiments, the complex comprises the amino acid sequence shown in any one of SEQ ID NO:133-142, 153, 154, 158-163, and 168.
[0046] In some embodiments, wherein: 1) one of the first fusion and the second fusion comprises an n×GCN4-dCas9 or dCas12-transcriptional repressor domain, and the other fusion comprises a DNA methylation domain -scFv; or 2) one of the first fusion and the second fusion comprises an scFv-dCas9 or dCas12-transcriptional repressor domain, and the other fusion comprises a DNA methylation domain -GCN4; or 3) one of the first fusion and the second fusion comprises an n×GCN4 ... One of the fusions contains an n×GFP11-dCas9 or dCas12 transcriptional repressor domain, and the other fusion contains a DNA methylation domain -GFP1-10; or 4) one of the first and second fusions contains a GFP1-10-dCas9 or dCas12 transcriptional repressor domain, and the other fusion contains a DNA methylation domain -GFP11; or 5) one of the first and second fusions contains an n×GCN4-dCas9 or dCas12 transcriptional repressor domain. 1) A fusion of the first and second fusions contains a scFv-dCas9 or dCas12-transcriptional repressor domain, and the other fusion contains a GCN4-DNA methylation domain; or 7) A fusion of the first and second fusions contains an n×GFP11-dCas9 or dCas12-transcriptional repressor domain, and the other fusion contains a GFP1-10-transcriptional repressor domain. -DNA methylation domain; or 8) one of the first fusion and the second fusion contains a GFP1-10-dCas9 or dCas12-transcriptional repressor domain, and the other fusion contains a GFP11-DNA methylation domain; wherein, - indicates that the domains at both ends are directly or indirectly connected in order from the N-terminus to the C-terminus; n×GCN4 or n×GFP11 respectively represent n copies of GCN4 connected by adapter sequences or n copies of GFP11 connected by adapter sequences, where n is selected from any integer from 1 to 20.
[0047] In some embodiments, the first fusion and / or the second fusion comprises an amino acid sequence shown in any one of SEQ ID NO: 83, 84, 94-102, 106-109, and 116-122.
[0048] In some embodiments, the complex comprises the amino acid sequence shown in any one of SEQ ID NO:143-152, 155-157 and 164-167.
[0049] In some embodiments, the complex further comprises nuclear localization signal and / or marker domains.
[0050] In some embodiments, the complex is capable of providing modification of at least one nucleotide in the region from 2000 bp upstream to 1000 bp downstream of the transcription start site of the INHBE gene.
[0051] In some embodiments, the complex is capable of providing modification of at least one nucleotide in a region approximately 1500 bp, approximately 1000 bp, approximately 750 bp, approximately 500 bp, approximately 250 bp, approximately 200 bp, approximately 100 bp, and approximately 50 bp upstream of the transcription start site of the INHBE gene.
[0052] In some embodiments, the complex is capable of providing modification of at least one nucleotide in a region approximately 2000 bp, approximately 1500 bp, approximately 1000 bp, approximately 750 bp, approximately 500 bp, approximately 250 bp, approximately 200 bp, approximately 100 bp, and approximately 50 bp downstream of the transcription start site of the INHBE gene.
[0053] In some embodiments, the complex is capable of providing modification of at least one nucleotide in the regions of approximately 1000 bp upstream to approximately 1500 bp downstream, approximately 800 bp upstream to approximately 1500 bp downstream, approximately 800 bp upstream to approximately 1000 bp downstream, approximately 800 bp upstream to approximately 800 bp downstream, and approximately 750 bp upstream to approximately 750 bp downstream of the INHBE gene.
[0054] In some embodiments, the complex is capable of providing modification of at least one nucleotide in a region approximately 750 bp upstream of the transcription start site of the INHBE gene.
[0055] In some embodiments, the complex is capable of providing modification of at least one nucleotide in a region of approximately 750 bp downstream of the transcription start site of the INHBE gene.
[0056] In some embodiments, the complex is capable of providing modification of at least one nucleotide in a region of about 400 bp to about 500 bp downstream of the transcription start site of the INHBE gene.
[0057] In some embodiments, the complex is capable of providing modification of at least one nucleotide in a region of about 450 bp to about 500 bp downstream of the transcription start site of the INHBE gene.
[0058] In some embodiments, the complex is capable of providing modification of at least one nucleotide in a region approximately 750 bp to approximately 1500 bp downstream of the transcription start site of the INHBE gene.
[0059] On the other hand, this application provides a nucleic acid encoding the complex described in this application.
[0060] In some implementations, the nucleic acid is a recombinant vector.
[0061] In some implementations, the recombinant vector also includes a non-coding region.
[0062] In some implementations, the non-coding region is selected from introns, modulating elements, promoters, enhancers, termination sequences, and 5' and 3' untranslated regions.
[0063] In some embodiments, the nucleic acid comprises a first nucleic acid fragment encoding the first fusion compound and a second nucleic acid fragment encoding the second fusion compound.
[0064] In some implementations, the first nucleic acid fragment and the second nucleic acid fragment are linked by a nucleic acid fragment encoding a cleavage peptide.
[0065] In some embodiments, the cleavage peptide is a 2A peptide and / or IRES.
[0066] In some embodiments, the 2A peptide is selected from P2A, T2A, E2A, and F2A.
[0067] In some implementations, the nucleic acid comprises a nucleic acid sequence shown in any one of SEQ ID NO:169-335.
[0068] On the other hand, this application provides a delivery carrier comprising the complex and / or nucleic acid described in this application, and optionally comprising liposomes and / or lipid nanoparticles.
[0069] On the other hand, this application provides a composition comprising the complex described in this application, the nucleic acid described in this application, and / or the delivery vector described in this application.
[0070] On the other hand, this application provides a cell comprising the complex described in this application, the nucleic acid described in this application, the delivery vector described in this application, and / or the composition described in this application.
[0071] On the other hand, this application provides a kit comprising the complex described in this application, the nucleic acid described in this application, the delivery vector described in this application, the composition described in this application, and / or the cells described in this application.
[0072] On the other hand, this application provides a method for regulating the expression of the INHBE gene product, the method comprising administering the complex described in this application, the nucleic acid described in this application, the delivery vector described in this application, the composition described in this application, the cells described in this application, and / or the kit described in this application.
[0073] In some embodiments, the method includes introducing the complex, the nucleic acid, the delivery vector, the composition, the cells, and / or the kit into cells containing the INHBE gene.
[0074] In some embodiments, the method includes contacting the complex, the nucleic acid, the delivery vector, and / or the composition with a target nucleotide sequence near the transcription start site of the INHBE gene.
[0075] In some implementations, the regulatory element includes a core promoter, a proximal promoter, a distal enhancer, a silencer, an insulator element, a boundary element, and / or a locus control region.
[0076] On the other hand, this application provides a method for treating or alleviating a disease or condition associated with abnormal INHBE gene expression and / or abnormal INHBE gene activity, the method comprising administering to a subject in need an effective amount of the complex described in this application, the nucleic acid described in this application, the delivery vector described in this application, the composition described in this application, the cells described in this application, and / or the kit described in this application.
[0077] On the other hand, this application provides the use of the complex described in this application, the nucleic acid described in this application, the delivery vector described in this application, the composition described in this application, the cell described in this application, and / or the kit described in this application for the preparation of a drug for treating or alleviating diseases or symptoms related to abnormal INHBE gene expression and / or abnormal INHBE gene activity.
[0078] In some implementations, the diseases or conditions associated with abnormal INHBE gene expression and / or abnormal INHBE gene activity include obesity and / or metabolic syndrome.
[0079] The complex and its encoded nucleic acid, vector, composition, cell, and other products provided in this application have at least one of the following advantages: significant INHBE gene transcriptional regulation efficiency (up to ~99%), rich INHBE gene regulation range, relatively flexible and diverse connection modes between various regulatory elements, and the significantly improved recruitment effect of the complex peptide based on the SunTag recruitment strategy. Attached Figure Description
[0080] The specific features of the invention involved in this application are shown in the appended claims. The features and advantages of the invention can be better understood by referring to the exemplary embodiments and accompanying drawings described in detail below. A brief description of the drawings is as follows:
[0081] Figure 1 shows the inhibitory effect of the complex described in this application on INHBE gene expression in the mouse hepatocyte line AML12.
[0082] Figure 2 shows the inhibitory effect of the complex described in this application on INHBE gene expression in primary monkey hepatocytes.
[0083] Figure 3 shows the inhibitory effect of the complex described in this application on INHBE gene expression in human Huh7 cells.
[0084] Figure 4 shows the inhibitory effect of candidate sgRNAs with different distances from the transcription start site described in this application on INHBE gene expression in human Huh7 cells. Detailed Implementation
[0085] The following specific embodiments illustrate the implementation of the invention. Those skilled in the art can easily understand other advantages and effects of the invention from the content disclosed in this specification.
[0086] Terminology Definition
[0087] In this application, the term "recruitment" generally refers to recruitment between protein molecules, specifically the recruitment of other molecules by a protein to perform a particular biological function. This recruitment relies primarily on the affinity of intermolecular interactions, which is often considered complex and related to the spatial structure of the protein molecule. Interaction mechanisms may include, but are not limited to, non-covalent interactions such as hydrogen bonds, ionic interactions, hydrophobic interactions, and van der Waals forces. For example, some proteins can recruit enzymes to catalyze chemical reactions or recruit other proteins to form complexes. These recruitment processes are crucial for many cellular processes, such as signal transduction, DNA replication, and gene expression.
[0088] In this application, the term "nucleic acid binding domain" generally refers to a portion of a polypeptide or composition capable of binding to a specific nucleic acid, which may include a region that contacts the nucleic acid, nucleic acid, and / or protein material. Examples of nucleic acid binding domains include, but are not limited to: helical-turn-helical domains, zinc finger domains, leucine zipper (bZIP) domains, winged helical domains, winged helical-turn-helical domains, helical-loop-helical domains, HMG-box domains, Wor3 domains, immunoglobulin domains, B3 domains, TALE domains, and / or domains of CRISPR / CasX proteins, etc.
[0089] In this application, the term "DNA-binding domain" generally refers to a folded protein domain containing at least one motif that recognizes double-stranded or single-stranded DNA. For example, the DNA-binding domain may recognize a specific DNA sequence (recognition or regulatory sequence) or have general affinity for DNA. In some cases, other domains of the DNA-binding domain typically regulate the activity of the DNA-binding domain; the DNA-binding function may be structural or include transcriptional regulation, and sometimes these two functions overlap. In some embodiments of the methods and gene expression regulatory molecules provided in this application, the DNA-binding domain may comprise a (DNA) nuclease, such as a nuclease capable of targeting DNA in a sequence-specific manner or capable of being directed or instructed to target DNA in a sequence-specific manner, such as the CRISPR-Cas system, zinc finger nucleases (ZFNs), transcription activator-like effector nucleases (TALENs), or a broad range of nucleases. In some embodiments, the DNA-binding domain is a DNA nuclease derived from the CRISPR-Cas system. For example, the CRISPR-Cas-derived DNA nuclease is a Cas protein.
[0090] In this application, the term "TALE DNA-binding domain" or "TALE" refers to a polypeptide containing one or more TALE repeating domains / units. Naturally occurring TALEs, or "wild-type TALEs," are nucleic acid-binding proteins secreted by numerous species of Proteobacteria. TALE polypeptides contain a nucleic acid-binding domain consisting of tandem repeats of highly conserved monomeric polypeptides, said monomeric polypeptides being primarily 33, 34, or 35 amino acids in length and differing primarily from each other at amino acid positions 12 and 13. In a preferred embodiment, the nucleic acid is DNA. As used herein, a polypeptide monomer of TALE is used to refer to a highly conserved repeating polypeptide sequence within the TALE nucleic acid-binding domain, and the term "repeated variable diresidue" or "RVD" is used to refer to a highly variable amino acid at positions 12 and 13 of the polypeptide monomer. A general representation of a TALE monomer contained within a DNA-binding domain is X. 1-11 -(X 12 X 13 )-X 14-33或34或35 The subscript indicates the position of the amino acid, and X represents any amino acid. 12 X 13 Indicating RVD. In some TALE polypeptide monomers, the variable amino acid at position 13 is missing or absent, and in such monomers, RVD consists of a single amino acid. In such cases, RVD can alternatively be represented as X*, where X represents X. 12 And (*) indicates X 13 No. The DNA-binding domain contains several repeats of the TALE monomer, and this can be represented as (X1-11 -(X 12 X 13 )-X 14-33或34或35 ) z In a preferred embodiment, z is at least 5-40. In a further preferred embodiment, z is at least 10-26.
[0091] TALE monomers possess nucleotide binding affinity determined by the amino acid types within their RVDs. For example, polypeptide monomers with RVDs containing NI preferentially bind to adenine (A), polypeptide monomers with RVDs containing NG preferentially bind to thymine (T), polypeptide monomers with RVDs containing HD preferentially bind to cytosine (C), and monomers with RVDs containing NN preferentially bind to both adenine (A) and guanine (G). In other embodiments, monomers with RVDs containing IG preferentially bind to T. Therefore, the number and order of repeats of polypeptide monomers within the nucleotide-binding domain of a TALE determine its nucleic acid target specificity. In a further embodiment of this application, monomers with RVDs containing NS recognize all four base pairs and can bind to A, T, G, or C. The structure and function of TALE are further described, for example, in Moscou et al., Science 326:1501 (2009); Boch et al., Science 326:1509-1512 (2009); and Zhang et al., Nature Biotechnology 29:149-153 (2011), each of which is incorporated herein by reference in its entirety. The repeating domains of TALE are involved in the binding of TALE to its homologous target DNA sequences. These repeating units (or “repetitive sequences”) exhibit at least some sequence homology with other TALE repeating sequences within naturally occurring TALE proteins. See, for example, U.S. Patent Publication No. 20110301073. The TALE binding domains involved in this application can be “engineered” to bind to predetermined nucleotide sequences, for example via engineering (changing one or more amino acids) of the recognition helical region of naturally occurring TALE proteins. Thus, engineered DNA-binding proteins (TALEs) are non-naturally occurring proteins. Non-limiting examples of methods for engineering DNA-binding proteins include design and selection. The designed DNA-binding protein is a non-naturally occurring protein, and its design and / or composition are primarily derived from rational design criteria. Rational design criteria include the application of substitution rules and computational algorithms used to process information in an information database storing existing TALE designs and binding data. See, for example, U.S. Patents 6,140,081; 6,453,242; and 6,534,261; also see WO 98 / 53058; WO 98 / 53059; WO 98 / 53060; WO02 / 016536 and WO 03 / 016496 and U.S. Publication No. 20110301073.
[0092] In this application, "Cas enzyme" may be used interchangeably with "Cas protein", "CRISPR protein", "CRISPR enzyme", "CRISPR-Cas protein", "CRISPR-Cas enzyme", "Cas", "CRISPR effector" or "Cas effector protein", which generally refers to a class of enzymes that are complementary to the CRISPR sequence and can use the CRISPR sequence as a guide to recognize and cut specific DNA strands. Non-limiting examples of Cas proteins include: Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csnl and Csxl2), Cas1O, Csyl, Csy2, Csy3, Csel, Cse2, Cscl, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmrl, Cmr3, Cmr4, Cmr5, Cmr6, Csbl, Csb2, Csb3, Csxl7, Csxl4, CsxlO, Csxl6, CsaX, Csx3, Csxl, Csxl5, Csf1, Csf2, Csf3, Csf4, and / or their homologues or modified forms thereof. These proteins are known; for example, the amino acid sequence of the Streptococcus pyogenes Cas9 protein can be found in the SwissProt database accession number Q99ZW2.
[0093] In this application, the term “class II Cas nuclease” generally refers to a class of Cas proteins that perform recognition and / or cleavage functions as a single protein, as defined by the updated classification scheme of CRISPR / Cas loci (Makarova et al., (2015) Nat Rev Microbiol, 13(11):722-36; Shmakov et al., (2015) Mol Cell, 60:385-397).
[0094] In this application, the terms "Type II Cas nucleases and Type II V Cas nucleases" generally refer to single-protein, RNA-directed endonucleases among Type II Cas nucleases. Specifically, Type II and Type V Cas nucleases (Type VB) require both tracrRNA (trans-activating CRISPR RNA) and crRNA (CRISPR RNA) to function properly, and crRNA and tracrRNA can be artificially combined to form a single guide RNA (sgRNA); Type V Cas nucleases (Type VA) require crRNA alone to perform their guide function. Non-restrictive examples of type II Cas nucleases include Cas9 and its family of related nucleases, and non-restrictive examples of type V Cas nucleases include Cas12a (also known as Cpf1), Cas12b (also known as C2c1), Cas12c (also known as C2c3), Cas12d (CasY), Cas12e (CasX), Cas12g, Cas12h, Cas12i, C2c1, C2c4, C2c5, C2c8, C2c9, C2c10, Cas14a, Cas14b, Cas14c nucleases and / or TnpB.
[0095] In this application, the term "dCas" may refer to the dCas protein or a fragment thereof. For example, as used herein, "dCas9" may refer to the dCas9 protein or a fragment thereof. As used herein, the terms "iCas" and "dCas" are used interchangeably to refer to a non-catalytically active CRISPR-related protein. In one embodiment, the dCas protein contains one or more mutations in its DNA cleavage domain. In one embodiment, the dCas protein contains one or more mutations in its RuvC or HNH domain. In one embodiment, the dCas molecule contains one or more mutations in both its RuvC and HNH domains. In one embodiment, the dCas protein is a fragment of the wild-type Cas protein. In one embodiment, the dCas protein contains a functional domain derived from the wild-type Cas protein, wherein the functional domain is selected from the Reel domain, the bridged helical domain, or the PAM interaction domain. In one embodiment, the nuclease activity of dCas is reduced by at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% compared to the nuclease activity of the corresponding wild-type Cas protein.
[0096] In this application, the term "capable of binding" is used interchangeably with "bind to," "specifically recognize," "target," etc., and generally refers to the ability of a binding molecule (e.g., the gene expression regulatory molecule of this application) to interact with nucleotides on a target gene or target site, or the binding molecule (e.g., the gene expression regulatory molecule of this application) to have sufficient affinity for the target gene or target site. Such interaction can be achieved through conjugation, coupling, attachment, providing complementarity, providing covalent or non-covalent forces, or improving binding stability.
[0097] In this application, the terms “guide RNA,” “guide DNA,” and “gRNA” are used interchangeably and generally refer to a DNA molecule capable of directing a nuclease (e.g., Argonaute, or Ago) to bind to and / or cleave a target gene. In some preferred embodiments, the guide DNA may include: a single-stranded DNA molecule (ssDNA), a single-stranded DNA molecule phosphorylated at the 5' end, a single-stranded DNA molecule hydroxylated at the 5' end, a base fragment having a base complement to the target gene, and / or a length of 8-35 nt. In some embodiments of this application, the term “guide RNA” refers to RNA comprising: (1) an “activated” nucleotide sequence of an RNA-directed endonuclease (e.g., a type II Cas nuclease, such as type II, type V, or type VI Cas endonuclease) that binds to the guide RNA-directed endonuclease; and (2) a “target” nucleotide sequence comprising a nucleotide sequence that hybridizes to the target nucleic acid. The “activating” nucleotide sequence and the “target” nucleotide sequence can be on separate RNA molecules (e.g., “two-guide RNA”); or they can be on the same RNA molecule (“single-guide RNA”, also known as sgRNA).
[0098] In this application, the term "DNA methyltransferase" generally refers to an enzyme that catalyzes the transfer of methyl groups to DNA. Non-limiting examples of DNA methyltransferases include DNMT1, DNMT 3A, DNMT 3B, and DNMT 3L. For example, through DNA methylation, DNA methyltransferases can modify the activity of DNA fragments (e.g., regulate gene expression) without altering the DNA sequence. As described herein, gene expression regulatory molecules may include one or more (e.g., two) DNA methyltransferases. When a DNA methyltransferase is included as part of a gene expression regulatory molecule, the DNA methyltransferase may be referred to as a "DNA methyltransferase domain". In all respects, the DNA methyltransferase domain comprises a variant or homolog of DNMT 3A having an amino acid sequence that is at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identical. In all respects, the DNA methyltransferase domain comprises a variant or homolog of DNMT 3L having an amino acid sequence that is at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identical.
[0099] In this application, the term "functionally active fragment" generally refers to a fragment that has a partial region of a full-length protein or nucleic acid but retains or partially retains the biological activity or function of the full-length protein or nucleic acid. For example, a functionally active fragment may retain or partially retain the ability of a full-length protein to bind to another molecule. For example, a functionally active fragment of a DNA methyltransferase may retain or partially retain the biological activity of a full-length DNA methyltransferase in catalyzing the transfer of methyl groups to DNA.
[0100] In this application, the term "transcriptional repressor" generally refers to a substance and / or agent, such as a protein (e.g., a transcription factor or fragment thereof), that binds to a target nucleic acid sequence and causes a decrease in the expression level of a gene product associated with the target nucleic acid sequence. For example, the gene product may be RNA (e.g., mRNA) transcribed from a gene or a polypeptide translated from mRNA transcribed from a gene. Typically, an increase or decrease in mRNA levels leads to an increase or decrease in the level of the polypeptide translated from it. Expression levels can be determined using standard techniques for measuring mRNA or protein. Examples of non-restrictive transcriptional repressors include: mSin3-interacting domain (SID) proteins, methyl-CpG-binding domain 2 (MBD2), MBD3, DNA methyltransferase (DNMT) 1 (DNMT1), DNMT2A, DNMT3A, DNMT3B, DNMT3L, retinoblastoma protein (Rb), methyl-CpG-binding protein 2 (Mecp2), GATA-1 and its cofactor Fog1, MAT2 regulator (ROM2), Arabidopsis HD2A protein (AtHD2A), lysine-specific demethylase 1 (LSD1), and / or Krüppel-related box (KRAB).
[0101] In this application, the term "KRAB" is also referred to as "Krüppel-associated box domain" or "Krüppel-associated box domain," which generally refers to about 45 to about 75 amino acid residues of the transcriptional repressor domain present in transcription factors of human zinc finger proteins. In various aspects, the KRAB domain may include variants or homologs having an amino acid sequence that is at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identical to the ZIM3 KRAB domain or the KOX1 KRAB domain.
[0102] In this application, the term "split green fluorescent protein" generally refers to a polypeptide that is capable of splitting and immediately forming an active green fluorescent protein upon recombination.
[0103] In this application, the term "GCN4" refers to a transcription factor in Saccharomyces cerevisiae, which is a "master regulator" in the yeast genome that regulates nearly one-tenth of the yeast genome. It is a highly conserved protein, and its mammalian homologue is Activating Transcription Factor-4 (ATF4).
[0104] In this application, the term "PDZ protein" generally refers to a naturally occurring protein containing a PDZ domain. Exemplary PDZ proteins include CASK, MPP1, DLG1, DLG2, PSD95, NeDLG, TIP-33, SYN1a, TIP-43, LDP, LIM, LIMK1, LIMK2, MPP2, NOS1, AF6, PTN_4, prIL16, 41.8kD, KIAA0559, RGS12, KIAA0316, DVL1, TIP-40, TIAM1, MINT1, MAGI-I, MAGI-2, MAGI-3, KIAA0303, CBP, MINT3, TIP-2, KIAA0561, and / or TIP-I.
[0105] In this application, the term "single-chain antibody" or "scFv (Single Chain Antibody)" generally refers to a single-chain polypeptide containing one or more antigen-binding sites. Additionally, although the H and L chains of the Fv fragment are encoded by different genes, they can be linked together directly or via peptides. For example, through recombinant methods, the H and L chains can be linked into a single protein chain (called a single-chain antibody, sAb; Bird et al. 1988 Science 242: 423-426; and Huston et al. 1988 PNAS 85: 5879-5883) using synthetic linkers. This single-chain antibody is also included in the term "antibody," which can be used as a binding determinant in the design and manufacture of multispecific binding molecules, and can be prepared by recombinant techniques or by enzymatic or chemical cleavage of intact antibodies.
[0106] In this application, the term "direct or indirect connection" generally refers to the relative terms "direct link" or "indirect link." "Direct link" generally refers to a direct connection. For example, a direct link can be a situation where the linked substances (e.g., amino acid sequence segments) are directly connected without a spacer component (e.g., an amino acid residue or its derivative); for example, amino acid sequence segment X and another amino acid sequence segment Y are directly linked through an amide bond formed by the C-terminal amino acid of amino acid sequence segment X and the N-terminal amino acid of amino acid sequence segment Y. "Indirect link" generally refers to a situation where the linked substances (e.g., amino acid sequence segments) are indirectly connected with a spacer component (e.g., an amino acid residue or its derivative). For example, the spacer component used in this application can be a segment of amino acid residues whose sequence is selected from any one of the amino acid sequences shown in SEQ ID NO:125-132 (SEQ ID NO:126 is GSG).
[0107] In this application, "nuclear localization sequence" or "NLS" generally refers to a peptide that directs a protein to the cell nucleus. In some embodiments, the NLS comprises five basic, positively charged amino acids. The NLS can be located at any position on the peptide chain. In some embodiments, the NLS is an NLS derived from SV40. In some embodiments, the NLS comprises the sequence shown in any one of SEQ ID NO:396-398. In some embodiments, the NLS has an amino acid sequence that is at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identical to any one of SEQ ID NO:396-398.
[0108] In this application, the term "marker" refers to a peptide that can be introduced into an expression vector, which can be used to allow the deletion and / or purification of the expression product of one or more vector insert fragments. Such markers are well-known in the art and comprise radiolabeled amino acids or polypeptides linked to a biotinylated moiety detectable by a labeled avidin (e.g., streptomycin containing a fluorescent label or enzymatic activity detectable by optical or colorimetric methods). Affinity markers such as FLAG, glutathione S-transferase, maltose-binding protein, cellulose-binding domain, thioredoxin, NusA, mistin, chitin-binding domain, keratinase, AGT, GFP, and other widely used markers, such as those used in protein expression and purification systems, are also included. Further non-limiting examples of peptides include, but are not limited to, the following: histidine labels, radioisotopes or radionuclides (e.g., 3H, 14C, 35S, 90Y, 99Tc, 111In, 125I, 177Lu, 166Ho, or 153Sm); fluorescent labels (e.g., FITC, rhodamine, lanthanides); enzyme labels (e.g., horseradish peroxidase, luciferase, alkaline phosphatase); chemiluminescent labels; biotin groups; pendant peptide antigenic determinants recognized by a second reporter (e.g., leucine zipper pairs, binding sites for secondary antibodies, metal-binding domains, antigenic determinant markers); and magnetic agents, such as gadolinium chelates.
[0109] In this application, the term "nucleic acid" is used interchangeably with "polynucleotide," "nucleotide," "nucleotide sequence," and "oligonucleotide," and generally refers to a nucleotide (e.g., deoxyribonucleotide or ribonucleotide) and polymers thereof in single-stranded, double-stranded, or multi-stranded form, or complementary forms thereof. For example, a nucleotide can be a ribonucleotide, a deoxyribonucleotide, or a modified version thereof. For example, a nucleotide can be a single-stranded and double-stranded DNA, a single-stranded and double-stranded RNA, or a hybrid molecule having a mixture of single-stranded and double-stranded DNA and RNA. For example, a nucleotide can be, but is not limited to, any type of RNA, such as mRNA, siRNA, miRNA, sgRNA, and guide RNA, and any type of DNA, genomic DNA, plasmid DNA, and microcircular DNA, and any fragment thereof. The term also covers nucleic acids containing known nucleotide analogs or modified backbone residues or bonds, said nucleic acids being synthetic, naturally occurring, or non-natural.
[0110] In this application, the terms "sequence encoding..." or "nucleic acid encoding..." generally refer to a nucleic acid (RNA or DNA molecule) containing a nucleotide sequence encoding a protein. The coding sequence may also include start and stop signals operatively linked to regulatory elements comprising promoters and polyadenylation signals capable of directing expression in the cells of an individual or mammal to which the nucleic acid has been administered. Codon optimization of the coding sequence is possible. In this application, the term "intron" generally refers to a transcribed DNA fragment that has been removed from the RNA transcript by splicing either end of the sequence (exon) together. Introns are considered interfering sequences within the protein-coding region of a gene and generally do not contain the information represented by the protein produced by that gene.
[0111] In this application, the term "recombinant vector" generally refers to a nucleic acid molecule capable of transporting itself and another nucleic acid linked to it. One type of vector is a "plasmid," which refers to a circular double-stranded DNA loop into which an additional DNA segment can be linked. Alternatively, the vector can be linear. Another type of vector is a viral vector, in which an additional DNA segment can be linked to the viral genome. Certain vectors are capable of autonomous replication within the host cell into which they are introduced (e.g., bacterial vectors with bacterial origins of replication and augmented mammalian vectors). Other vectors (e.g., non-augmented mammalian vectors) can integrate into the host cell's genome after introduction and thus replicate along with the host genome.
[0112] In this application, the term "transcription start site" refers to the first base transcribed at the 5' end of a gene. It is the base on the DNA strand corresponding to the first nucleotide of the mRNA strand during transcription, and is typically a purine (e.g., A or G). The sequence preceding the transcription start site (i.e., the 5' end) is generally referred to as upstream, while the sequence following it (i.e., the 3' end) is referred to as downstream. In this application, the transcription start site is represented by "0". In this application, upstream of the transcription start site is represented by "-". For example, 750 bp upstream of the transcription start site is represented as "-750 bp". In this application, downstream of the transcription start site is represented by "+". For example, 750 bp downstream of the transcription start site can be represented as "+750 bp" or simply "750 bp".
[0113] In this application, the term "regulatory element" refers to a genetic element capable of controlling the expression of a nucleic acid sequence. Examples include splicing signals, promoter sequences, polyadenylation signals, transcription termination sequences, upstream regulatory domains, origins of replication, internal ribosome entry sites ("IRES"), enhancers, etc., which collectively enable the replication, transcription, and translation of the coding sequence in the recipient cell. Not all of these control sequences are required.
[0114] In this application, the term "promoter" generally refers to a nucleotide sequence that controls or regulates transcription of a nucleotide sequence (e.g., a coding sequence) operatively associated with a promoter. The coding sequence controlled or regulated by a promoter may encode a polypeptide and / or functional RNA. Typically, a "promoter" refers to a nucleotide sequence containing an RNA polymerase II binding site and directing transcription initiation. Typically, a promoter is located 5' or upstream of the starting point of the coding region relative to the corresponding coding sequence. A promoter may contain other elements that act as regulators of gene expression; for example, a promoter region. In some embodiments, the promoter region may include at least one intron. Promoters may include, for example, constitutive, inducible, time-regulated, developmentally regulated, chemically regulated, tissue-preferred, and / or tissue-specific promoters for the preparation of recombinant nucleic acid molecules, such as "synthetic nucleic acid constructs" or "protein-RNA complexes." These different types of promoters are known in the art.
[0115] In this application, the term "enhancer" generally refers to a regulatory DNA sequence, for example, 50-1500 bp, that can be bound by proteins (activating proteins) to stimulate or enhance the transcription of one or more genes. These activating proteins (also known as transcription factors) interact with a mediator complex and recruit polymerase II and general transcription factors, which then initiate gene transcription. Enhancers are typically cis-acting but can be located upstream or downstream of the transcription start site of the gene or the genes they regulate. Furthermore, enhancers can be forward or backward and do not need to be located near the transcription start site to influence transcription, as some enhancers have been found located hundreds of thousands of base pairs upstream or downstream of the start site. Enhancers can also be found in introns.
[0116] In this application, the term "cleavage peptide" refers to a class of polypeptides capable of cleaving proteins. For example, the cleavage peptide can achieve protein cleavage via ribosome jumping rather than protease hydrolysis. For example, the cleavage peptide may be a cleavage 2A peptide, which may include T2A, F2A, P2A, and / or E2A.
[0117] In this application, the term "delivery carrier" generally refers to a transfer medium capable of delivering a reagent (e.g., a nucleic acid molecule) to target cells. A delivery carrier can deliver a reagent to a specific cell subclass. For example, the delivery carrier can target certain cell types by means of inherent characteristics of the delivery carrier or by a portion coupled to the carrier, a portion contained therein (or a portion bound to the carrier such that the portion and the delivery carrier remain together, thereby making the portion sufficient to target the delivery carrier). Delivery carriers can also improve the in vivo half-life and / or bioavailability of the reagent to be delivered. Delivery carriers may include viral vectors, virus-like particles, polycationic carriers, peptide carriers, liposomes, and / or hybridization carriers. For example, if the target cells are hepatocytes, the properties of the delivery carrier (e.g., size, charge, and / or pH) can effectively deliver the delivery carrier and / or the molecules encapsulated therein to the target cells, reduce immune clearance, and / or promote residence in the target cells.
[0118] In this application, the term "liposome" generally refers to a vesicle with an internal space that is isolated from an external medium by one or more bilayer membranes. In some embodiments, the bilayer membrane can be formed from amphiphilic molecules, such as synthetic or naturally derived lipids comprising spatially isolated hydrophilic and hydrophobic domains; in other embodiments, the bilayer membrane can be formed from amphiphilic polymers and surfactants. In some embodiments, the liposome is a spherical vesicle structure consisting of a single or multiple lipid bilayer surrounding an internal aqueous compartment and a relatively impermeable outer lipophilic phospholipid bilayer. In some embodiments, the liposome is biocompatible, non-toxic, capable of delivering hydrophilic and lipophilic drug molecules, protecting their carriers from degradation by plasma enzymes, and transporting their load across biological membranes and the blood-brain barrier (BBB). Liposomes can be made from several different types of lipids, such as phospholipids. Liposomes may contain natural phospholipids and lipids such as 1,2-distearate-sn-glycerol-3-phosphatidylcholine (DSPC), sphingomyelin, lecithin, monosialotetrahexosylganglioside, or any combination thereof. Several other additives may be added to liposomes to modify their structure and properties. For example, liposomes may also contain cholesterol, sphingomyelin, and / or 1,2-dioleoyl-sn-glycerol-3-phosphoethanolamine (DOPE), for example, to increase stability and / or prevent leakage of the internal carriers of the liposomes.
[0119] The term "lipid nanoparticle (LNP)" generally refers to a particle containing multiple (i.e., more than one) lipid molecules physically bound together by intermolecular forces (e.g., covalent or non-covalent). LNPs can be, for example, microspheres (including monolayer and multilayer vesicles, such as liposomes), dispersed phases in emulsions, micelles, or internal phases in suspensions. LNPs can encapsulate nucleic acids within cationic lipid particles (e.g., liposomes) and can be delivered to cells relatively easily. In some instances, lipid nanoparticles are free of any viral components, which helps minimize safety and immunogenicity issues. The lipid particles can be used for in vitro, ex vivo, and in vivo delivery. The lipid particles can also be used for cell populations of various sizes. The LNPs of this application can be readily prepared by various methods known in the art, such as by mixing an organic phase with an aqueous phase. Mixing of the two phases can be achieved using microfluidic devices and impinging flow reactors. The more thoroughly the organic and aqueous phases are mixed, the better the encapsulation efficiency and particle size distribution of the obtained LNPs. Preferably, the particle size of the LNPs can be adjusted by varying the mixing rate of the organic and aqueous phases. The faster the mixing rate, the smaller the particle size of the prepared LNPs. Encapsulation efficiency can be optimized by adjusting the N / P (ionizable lipid / nucleic acid) ratio of the LNP system. In some instances, LNPs can be used to deliver DNA molecules and / or RNA molecules (e.g., Cas, sgRNA mRNA). In some cases, LNPs can be used to deliver Cas / gRNA RNP complexes. In some embodiments, LNPs are used to deliver both mRNA and gRNA.
[0120] In this application, the term "subject" generally refers to an animal, typically a mammal such as a human, non-human primates (apes, gibbons, gorillas, chimpanzees, orangutans, macaques), livestock (dogs and cats), farm animals (poultry such as chickens and ducks, horses, cattle, goats, sheep, pigs), and laboratory animals (mice, rats, rabbits, guinea pigs). Human subjects include fetuses, newborns, infants, adolescents, and adult subjects. Subjects include animal disease models, such as mice and other animal models of blood clotting disorders (such as HemA), and other animal models known to those skilled in the art.
[0121] In this application, the term "comprising" generally means including the explicitly specified features, but does not exclude other elements.
[0122] In this application, the term “selected from” generally refers to the selection of objects and all combinations thereof. For example, “selected from (:) A, B and C” means all combinations of A, B and C, such as A, B, C, A+B, A+C, B+C or A+B+C.
[0123] In this application, the term "about" generally refers to a variation within a range of 0.5% to 10% above or below a specified value, such as a variation within a range of 0.5%, 1%, 1.5%, 2%, 2.5%, 3%, 3.5%, 4%, 4.5%, 5%, 5.5%, 6%, 6.5%, 7%, 7.5%, 8%, 8.5%, 9%, 9.5%, or 10% above or below a specified value.
[0124] Invention Details
[0125] On one hand, this application provides a complex comprising a first fusion and a second fusion, wherein: 1) one of the first fusion and the second fusion comprises a DNA methylation domain and at least one recruitment domain A, and the other fusion comprises a transcriptional repressor domain and at least one recruitment domain A'; and 2) the first fusion or the second fusion comprises a nucleic acid binding domain; and the recruitment domain A and the recruitment domain A' are capable of interacting to enable the fusion of one of the first fusion and the second fusion, or a portion thereof, to be recruited to the vicinity of the other fusion; the nucleic acid binding domain is capable of specifically binding to a target nucleotide sequence within a region 2000 bp upstream to 1000 bp downstream of the transcription start site of the INHBE gene. For example, the target nucleotide sequence is located approximately 2000 bp, 1900 bp, 1800 bp, 1700 bp, 1600 bp, 1500 bp, 1400 bp, 1300 bp, 1200 bp, 1100 bp, 1000 bp, 900 bp, 800 bp, 700 bp, 600 bp, 500 bp, 400 bp, 300 bp, 200 bp, and 100 bp upstream of the transcription start site of the INHBE gene and approximately 100 bp downstream of the transcription start site of the INHBE gene.
[0126] For example, the target nucleotide sequence is located within a region approximately 2000 bp, 1500 bp, 1000 bp, 750 bp, 500 bp, 250 bp, 200 bp, 100 bp, and 50 bp upstream of the INHBE transcription start site. For example, the target nucleotide sequence is located within a region approximately 750 bp upstream of the INHBE transcription start site.
[0127] For example, the target nucleotide sequence is located within a region approximately 2000 bp, 1500 bp, 1000 bp, 750 bp, 500 bp, 250 bp, 200 bp, 100 bp, and 50 bp downstream of the INHBE transcription start site. For example, the target nucleotide sequence is located within a region approximately 750 bp downstream of the INHBE transcription start site. For example, the target nucleotide sequence is located within a region approximately 250 bp to approximately 500 bp downstream of the INHBE transcription start site. For example, the target nucleotide sequence is located within a region approximately 400 bp to approximately 500 bp downstream of the INHBE transcription start site. For example, the target nucleotide sequence is located within a region approximately 450 bp to approximately 500 bp downstream of the INHBE transcription start site.
[0128] For example, the target nucleotide sequence is located approximately 1000 bp upstream to approximately 1500 bp downstream of the INHBE gene, approximately 800 bp upstream to approximately 1500 bp downstream, approximately 500 bp upstream to approximately 1500 bp downstream, approximately 200 bp upstream to approximately 1500 bp downstream, approximately 800 bp upstream to approximately 1000 bp downstream, approximately 500 bp upstream to approximately 1000 bp downstream, approximately 200 bp upstream to approximately 1000 bp downstream, approximately 800 bp upstream to approximately 800 bp downstream, and approximately 500 bp upstream of the INHBE gene. Within the regions from 0 bp to approximately 800 bp downstream, from approximately 200 bp upstream to approximately 800 bp downstream, from approximately 750 bp upstream to approximately 750 bp downstream, from approximately 500 bp upstream to approximately 750 bp downstream, from approximately 200 bp upstream to approximately 750 bp downstream, from approximately 500 bp upstream to approximately 500 bp downstream, from approximately 200 bp upstream to approximately 500 bp downstream, from approximately 250 bp upstream to approximately 250 bp downstream, from approximately 250 bp upstream to approximately 100 bp downstream, and from approximately 250 bp upstream to approximately 50 bp downstream. On the other hand, this application provides a nucleic acid encoding the complex described in this application. For example, the nucleic acid comprises DNA and / or mRNA. For example, the nucleic acid can be used to treat or alleviate diseases or symptoms associated with abnormal target gene expression and / or abnormal target gene activity. In some embodiments, the nucleic acid is mRNA; one or more modification techniques may be used to produce more stable mRNA. Known mRNA modification techniques can be broadly categorized into three types: using artificially synthesized non-natural ribonucleic acid (RNA) to replace natural ribonucleic acid (RNA) for mRNA synthesis; adding 5' caps, 3' poly(A) tails, and UTR (untranslated region) sequences; and employing specialized novel formulation techniques to effectively protect mRNA. Among these, the preferred mRNA modification techniques involve using artificially synthesized non-natural ribonucleic acid to replace natural ribonucleic acid for mRNA synthesis. Chemical modifications on eukaryotic mRNA can be broadly classified into three types: methylation, pseudouridine (Ψ), and hypoxanthine. For example, the chemical modification may be selected from: pseudouridine, N1-methylpseuuridine, N1-ethylpseuuridine, 2-thiouridine, 4'-thiouridine, 5-methylcytosine, 2-thio-1-methyl-1-deazo-pseuuridine, 2-thio-1-methylpseuuridine, 2-thio-5-aza-uridine, 2-thio-dihydropseuuridine, 2-thio-dihydrouridine, 2-thio-pseuuridine, 4-methoxy-2-thio-pseuuridine, 4-methoxy-pseuuridine, 4-thio-1-methylpseuuridine, 4-thio-pseuuridine, 5-aza-uridine, dihydropseuuridine, 5-methyluridine, 5-methoxyuridine, and 2'-O-methyluridine. For example, the nucleic acid is a recombinant vector containing a nucleic acid encoding the complex described in this application. For example, a recombinant vector may refer to a nucleic acid molecule capable of transporting another nucleic acid linked thereto.Recombinant vectors may include single-stranded, double-stranded, or partially double-stranded nucleic acid molecules; nucleic acid molecules containing one or more free ends, or without free ends (e.g., circular); nucleic acid molecules containing DNA, RNA, or both; and other types of polynucleotides known in the art. For example, viral vectors may be used. Viral vectors may contain virus-derived DNA or RNA sequences for packaging into viruses (e.g., retroviruses, replication-defective retroviruses, adenoviruses, replication-defective adenoviruses, and adeno-associated virus (AAV)). Viruses and viral vectors may be used for in vitro, ex vivo, and / or in vivo delivery.
[0129] On the other hand, this application provides a delivery carrier comprising the complex and / or nucleic acid described in this application, and optionally liposomes and / or lipid nanoparticles. For example, the delivery carrier can be introduced into cells by physical delivery methods. Examples of physical methods include microinjection, electroporation, and hydrodynamic delivery. For example, LNPs can encapsulate nucleic acids in cationic lipid particles (e.g., liposomes) and can be delivered to cells relatively easily. In some examples, the lipid nanoparticles do not contain any viral components, which helps to minimize safety and immunogenicity issues. Lipid particles can be used for in vitro, ex vivo, and in vivo delivery. The components of LNPs may include cationic lipids, ionizable lipids, polyethylene glycol-modified lipids and / or supporting lipids, and optionally cholesterol components.
[0130] On the other hand, this application provides a composition comprising the complex described in this application, the nucleic acid described in this application, and / or the delivery vector described in this application. For example, the complex, the nucleic acid (or recombinant vector) encoding the complex, and the delivery vector in the composition may be contained simultaneously in one composition or separately in different compositions. For example, when using the complex, the nucleic acid (or recombinant vector) encoding the complex, and / or the delivery vector in the composition, they may be used simultaneously or separately.
[0131] On the other hand, this application provides a cell comprising the complex described in this application, the nucleic acid described in this application, the delivery vector described in this application, and / or the composition described in this application.
[0132] On the other hand, this application provides a kit comprising the complex, nucleic acid, delivery vector, composition, and / or cells described in this application. For example, the kit further comprises at least one container for holding the aforementioned components. For example, the kit comprises more than one of the aforementioned components and further comprises a second, third, and / or other container besides the container, in which the more than one of the aforementioned components can be separately placed. For example, the kit can hold various combinations of the aforementioned components in the containers. For example, the kit further comprises buffer reagents, mixing devices, measuring devices, sorting devices, and / or labeling devices. For example, the kit further comprises packaging for accommodating various containers. For example, the kit further comprises instructions for using the kit components. For example, the instructions may be in physical paper form and / or machine-readable electronic form.
[0133] On the other hand, this application provides a method for regulating the expression of the INHBE gene product, the method comprising administering the complex described in this application, the nucleic acid described in this application, the delivery vector described in this application, the composition described in this application, the cells described in this application, and / or the kit described in this application. For example, the method for inhibiting target gene expression is to introduce the complex, the nucleic acid, the delivery vector, the composition, the cells, and / or the kit into cells containing the INHBE gene. For example, the introduced cells may be introduced using non-viral or virus-based transfection methods. For example, the non-viral transfection method includes any suitable method for introducing cells without using viral DNA or viral particles as a delivery system, and non-limiting examples of non-viral transfection methods include nanoparticle encapsulation of the nucleic acid encoding the complex (e.g., lipid nanoparticles, gold nanoparticles, etc.), calcium phosphate transfection, liposome transfection, nuclear transfection, acoustic perforation, transfection by heat shock, magnetic transfection, and electroporation. For example, virus-based transfection methods include any viral vector suitable for the method described in this application, and non-limiting examples include, but are not limited to, retroviruses, adenoviruses, lentiviruses, and / or adeno-associated virus vectors. For example, the method for inhibiting INHBE gene expression further includes introducing the complex, the nucleic acid, the delivery vector, the composition, the cell, and / or the kit from the external environment into the cell. As another example, the method for inhibiting INHBE gene expression includes contacting the complex, the nucleic acid, the delivery vector, and / or the composition with a target nucleotide sequence near the transcription start site of the INHBE gene. For example, contact refers to contacting the first fusion, the second fusion, and the guide RNA described in this application with the target nucleotide sequence, and the guide RNA forms a complex with the fusion containing a DNA-binding domain, which specifically recognizes and hybridizes to a specific region in the INHBE gene, while the first and second fusions are recruited to the vicinity of the DNA-binding domain through direct or indirect interactions of their recruitment domains A and A', thereby regulating the expression of the target nucleic acid. For example, the method includes presenting the first fusion, the second fusion, and the guide RNA as a complex (e.g., an assembled ribonucleoprotein complex), and contacting this complex with the target nucleotide sequence near the transcription start site of the INHBE gene.
[0134] On the other hand, this application provides a method for treating or alleviating a disease or symptom associated with abnormal INHBE gene expression and / or abnormal INHBE gene activity, the method comprising administering an effective amount of the complex, nucleic acid, delivery vector, composition, cells, and / or kit described in this application to a subject in need. For example, the treatment method comprises mixing the complex, nucleic acid, delivery vector, composition, cells, and / or kit with a therapeutic agent and delivering it systemically to a subject in need, exposing them extensively to a large portion of the body, which can be performed by any means known in the art, including but not limited to intravenous, intra-arterial, subcutaneous, intracavitary, and intraperitoneal delivery. For example, the treatment method comprises mixing the complex, nucleic acid, delivery vector, composition, cells, and / or kit with a therapeutic agent and delivering it locally to a subject in need, allowing them to directly reach a target site within the organism, which can be achieved, for example, by direct injection into a disease site (e.g., a tumor or site of inflammation) or a target organ (e.g., the liver, heart, pancreas, kidney, etc.). For example, the local delivery includes local application or injection techniques, including but not limited to intramuscular, subcutaneous, or intradermal injection. For example, the local delivery does not exclude systemic pharmacological effects.
[0135] On the other hand, this application provides the use of the complex described in this application, the nucleic acid described in this application, the delivery vector described in this application, the composition described in this application, the cell described in this application, and / or the kit described in this application for the preparation of a drug for treating or alleviating diseases or symptoms related to abnormal INHBE gene expression and / or abnormal INHBE gene activity.
[0136] On the other hand, this application provides the complex, nucleic acid, delivery vector, composition, cell, or kit described in this application for the treatment or relief of diseases or symptoms associated with abnormal INHBE gene expression and / or abnormal INHBE gene activity.
[0137] First fusion or second fusion
[0138] In some embodiments, the first and second fusions of the complex of this application can generally be divided into two cases: (1) one of the two fusions contains a nucleic acid binding domain, a DNA methylation domain and a recruitment domain A, and the other fusion contains a transcriptional repressor domain and a recruitment domain A', or (2) one of the two fusions contains a nucleic acid binding domain, a transcriptional repressor and a recruitment domain A, and the other fusion contains a DNA methylation domain and a recruitment domain A'.
[0139] Specifically, in some embodiments of scenario (1) above, one of the two fusions may contain, from N-terminus to C-terminus, a DNA methylation domain, a nucleic acid binding domain, and a recruitment domain A. For example, in some embodiments of scenario (2) above, one of the two fusions may contain, from N-terminus to C-terminus, a recruitment domain A, a nucleic acid binding domain, and a transcriptional repressor domain. For example, in some embodiments of scenario (1) above, the other fusion may contain, from N-terminus to C-terminus, a transcriptional repressor domain and a recruitment domain A', or a recruitment domain A' and a transcriptional repressor domain, i.e., the transcriptional repressor domain and the recruitment domain A' can be connected in an interchangeable order. For example, in some embodiments of scenario (2) above, the other fusion may contain, from N-terminus to C-terminus, a DNA methylation domain and a recruitment domain A', or a recruitment domain A' and a DNA methylation domain, i.e., the DNA methylation domain and the recruitment domain A' can be connected in an interchangeable order.
[0140] In some more specific embodiments, the nucleic acid-binding domain is a DNA-binding domain. For example, the DNA-binding domain may be selected from: TALE domains, zinc finger domains, tetR domains, large-scale nucleases, Cas proteins, Argonaute (Ago) proteins, and their homologues, modified forms, or variants. For example, the DNA-binding domain may be a Cas protein, and the Cas protein may be a type II Cas nuclease. Further, the Cas protein may be selected from type II type II Cas nucleases and type II type V Cas nucleases; for example, the Cas protein may be a Cas9 or Cas12 protein. In some embodiments, the Cas protein may be an inactivated Cas9 (dCas9) protein or an inactivated Cas12 (dCas12) protein. For example, the DNA-binding domain of this application may contain, but is not limited to, the amino acid sequence shown in any one of SEQ ID NO: 1-9. In some more specific embodiments, the DNA-binding domain is capable of binding guide RNA. For example, the guide RNA described in this application may contain, but is not limited to, the nucleotide sequence shown in any one of SEQ ID NO:337-382, 399-1378.
[0141] In some more specific embodiments, the transcriptional repressor is selected from one or more of the following domains: KRAB, ZIM3 KRAB, ZNF680, ZNF554, ZNF264, ZNF582, ZNF324, ZNF669, ZNF354A, ZNF82, ZNF595, ZNF41 9. ZNF566, ZIM2, EHMT2, SUV39H1, ZFPM1, TRIM28, EZH2, MXD1, SID, LSD1, HP1a, HDAC3, HDA C1, PRMT1, SETDB1, hSIRT1, ZNF436, ZNF257, ZNF675, ZNF490, ZNF320, ZNF331, ZNF816, ZN F41, ZNF189, ZNF528, ZNF543, ZNF140, ZNF610, ZNF350, ZNF8, ZNF30, ZNF98, ZNF677, ZNF5 96, ZNF214, ZNF37A, ZNF34, ZNF250, ZNF547, ZNF273, ZFP82, ZNF224, ZNF33A, ZNF45, ZNF1 75, ZNF184, ZFP28-1, ZFP28-2, ZNF18, ZNF213, ZNF394, ZFP1, ZFP14, ZNF416, ZNF557, ZNF 729, ZNF254, ZNF764, ZNF785, ZNF10, CBX5, RYBP, YAF2, MGA, CBX1, SCMH1, MPP8, SUMO3, HERC2, BIN1, PCGF2, TOX, FOXA1, FOXA2, IRF2BP1, IRF2BP2, IRF2BPLIRF-2BP1_2N-terminal domain, HOXA13, HOXB13, HOXC13, HOXA11, HOXC11, HOXC10, HOXA10, HOXB9, HOXA9, ZFP28, ZN334, ZN568, ZN3 7A, ZN181, ZN510, ZN862, ZN140, ZN208, ZN248, ZN571, ZN699, ZN726, ZIK1, ZNF2, Z705F, ZNF14, ZN471, ZN624 , ZNF84, ZNF7, ZN891, ZN337, Z705G, ZN529, ZN729, ZN419, Z705A, ZN302, ZN486, ZN621, ZN688, ZN33A, ZN554, ZN878, ZN772, ZN224, ZN184, ZN544, ZNF57, ZN283, ZN549, ZN211, ZN615, ZN253, ZN226, ZN730, Z585A, ZN732,ZN681,ZN667,ZN649,ZN470,ZN484,ZN431,ZN382,ZN254,ZN124,ZN607,ZN317,ZN620,ZN141,ZN584,ZN540,ZN75D,ZN555,ZN658,ZN684,RBAK,ZN829,ZN582,ZN112,ZN716,HKR1,ZN350,ZN480,ZN416,ZNF92,ZN100,ZN736,ZNF74,ZN443,ZN195,ZN530,ZN782,ZN791,ZN331,Z354C,ZN157,ZN727,ZN550,ZN793,ZN235,ZN724,ZN573,ZN577,ZN789,ZN718,ZN300,ZN383,ZN429,ZN677,ZN850,ZN454,ZN257,ZN264,ZN485,ZN737,ZNF44,ZN596,ZN565,ZN543,ZFP69,SUMO1,ZNF12,ZN169,ZN433,ZN175,ZN347,ZNF25,ZN519,Z585B,ZN517,ZN846,ZN230,ZNF66,ZN713,ZN816,ZN426,ZN674,ZN627,ZNF20,Z587B,ZN316,ZN233,ZN611,ZN556,ZN234,ZN560,ZNF77,ZN682,ZN614,ZN785,ZN445,ZFP30,ZN225,ZN551,ZN610,ZN528,ZN284,ZN418,ZN490,ZN805,Z780B,ZN763,ZN285,ZNF85,ZN223,ZNF90,ZN557,ZN425,ZN229,ZN606,ZN155,ZN222,ZN442,ZNF91,ZN135,ZN778,ZN534,ZN586,ZN567,ZN440,ZN583,ZN441,ZNF43,ZN589,ZN563,ZN561,ZN136,ZN630,ZN527,ZN333,Z324B,ZN786,ZN709,ZN792,ZN599,ZN613,ZF69B,ZN799,ZN569,ZN564,ZN546,ZFP92,ZN723,ZN439,ZFP57,ZNF19,ZN404,ZN274,CBX3,ZN250,ZN570,ZN675,ZN695,ZN548,ZN132,ZN738,ZN420,ZN626,ZN559,ZN460,ZN268,ZN304,ZN605,ZN844,SUMO5,ZN101,ZN783,ZN417,ZN182,ZN823,ZN177,ZN197,ZN717,ZN669,ZN256,ZN251,CBX4,CDY2,CDYL2,ZN562,ZN461,Z324A,ZN766,ID2,ZN214,CBX7,ID1,CREM,SCX,ASCL1,ZN764,SCML2,TWST1,CREB1,TERF1,ID3,CBX8,GSX1,NKX22,ATF1,TWST2,ZNF17,TOX3,TOX4,ZMYM3,I2BP1,RHXF1,SSX2,I2BPL,ZN680,TRI68,HXA13,PHC3,TCF24,HXB13,HEY1,PHC2,ZNF81,FIGLA,SAM11,KMT2B,HEY2,JDP2,HXC13,ASCL4,HHEX,GSX2,ETV7,ASCL3,PHC1,OTP,I2BP2,VGLL2,HXA11,PDLI4,ASCL2,CDX4,ZN860,LMBL4,PDIP3,NKX25,CEBPB,ISL1,CDX2,PROP1,SIN3B,SMBT1,HXC11,HXC10,PRS6A,VSX1,NKX23,MTG16,HMX3,HMX1,KIF22,CSTF2,CEBPE,DLX2,PPARG,PRIC1,UNC4,BARX2,ALX3,TCF15,TERA,VSX2,HXD12,CDX1,TCF23,ALX1,HXA10,RX,CXXC5,SCML1,NFIL3,DLX6,MTG8,CEBPD,SEC13,FIP1,ALX4,LHX3,PRIC2,MAGI3,NELL1,PRRX1,MTG8R,RAX2,DLX3,DLX1,NKX26,NAB1,SAMD7,PITX3,WDR5,MEOX2,NAB2,DHX8,CBX6,EMX2,CPSF6,HXC12,KDM4B,LMBL3,PHX2A,EMX1,NC2B,DLX4,SRY,ZN777,ZN398,GATA3,BSH,SF3B4,TEAD1,TEAD3,RGAP1,PHF1,GATA2,FOXO3,ZN212,IRX4,ZBED6,LHX4,SIN3A,RBBP7,NKX61,R51A1,MB3L1,DLX5,NOTC1,TERF2,ZN282,RGS12,ZN840,SPI2B,PAX7,NKX62,ASXL2,FOXO1,GATA1, ZMYM5, LRP1, MIXL1, SGT1, LMCD1, CEBPA, SOX14, WTIP, PRP19, NKX11, RBBP4, DMRT2, SMCA2, and their functionally active fragments.
[0142] In some more specific embodiments, the DNA methylation domain comprises at least one DNA methyltransferase or a functionally active fragment thereof. For example, the DNA methyltransferase is selected from DNMT3A, DNMT3B, DNMT3c, DNMT1, DNMT2, and DNMT3L. For example, the DNA methylation domain comprises at least one DNMT3A and at least one DNMT3L. For example, the at least one DNMT3A and the at least one DNMT3L may be linked in an interchangeable order. For example, the DNA methylation domain comprises one DNMT3A and one DNMT3L, and they may be linked in an interchangeable order. For example, the DNA methyltransferase comprises the amino acid sequence shown in any one of SEQ ID NO: 19-24.
[0143] The first and second fusions of the complex in this application form aggregated complexes through interactions between their respective recruitment domains. Therefore, this application provides non-limiting examples of combinations of recruitment domain A and recruitment domain A': (1) one of recruitment domain A and recruitment domain A' has a domain of GCN4, and the other domain is scFv; or (2) one of recruitment domain A and recruitment domain A' has a domain of a GFP11 fragment, and the other domain is GFP1-10; or (3) one of recruitment domain A and recruitment domain A' has a domain of GVKESLV, and the other domain is a PDZ protein domain. Similarly, the situation where GFP11 and GFP1-10 are derived from splitting GFP (SEQ ID NO:15) to form recruitment domain A and recruitment domain A', respectively, can also be applied to other classes of fluorescent proteins, such as mCherry (SEQ ID NO:16), eYFP (SEQ ID NO:18), eCFP (SEQ ID NO:17), etc. Different sets of recruitment domain A and recruitment domain A' can be obtained by splitting mCherry, eYFP, or eCFP for use in the complex provided in this application. In some embodiments, one of the first fusion and the second fusion of the complex of this application may contain two or more recruitment domains, which are linked by a linker sequence. The amino acid sequence of an exemplary recruitment domain may include any one of SEQ ID NO:10-14.
[0144] Based on the above, this application may provide the following amino acid sequences of the first or second fusion compound:
[0145] Table 1. Exemplary Fusions
[0146] The embodiments described below are not intended to be limited by any theory, but are merely for illustrating the complex, preparation method and use of this application, and are not intended to limit the scope of the invention.
[0147] Example
[0148] Example 1
[0149] Design and construction of plasmids containing the complex of this application
[0150] The amino acid sequences of the recruitment system with HA epitopes and P2A and the recruited elements (including DNMT3A, DNMT3L, dSpCas9, and KRAB) were optimized by Genscript to be suitable for mammalian expression and synthesized. They were then cloned into the pLV-CAG vector with a CAG promoter and WPRE, and the recruited elements and self-splicing recruitment system fusion protein were expressed by the CAG promoter.
[0151] When optimizing different functional elements, these elements were synthesized by Genscript using nucleic acid sequences optimized for mammalian expression. First, the vector excluding the element to be replaced was amplified by PCR. Then, the element to be replaced was amplified from the synthesized sequence, with homologous arm sequences introduced. Finally, the different elements were recombined into the vector using NEBuilder reagent to construct the final expression plasmid.
[0152] Example 2
[0153] The inhibitory efficiency of the complex in this application on INHBE gene expression in the mouse hepatocyte cell line AML12
[0154] In this embodiment, the mouse N2a cell line was used as the research model. Guide RNAs were designed targeting the region from -1000 to +1000 bp before and after the transcription initiation region of the mouse INHBE genome (SEQ ID NO:393). The mRNA of the epigenetic editing tool (SEQ ID NO:288) and different sgRNAs (SEQ ID NO:337-356) were embedded in LNPs at a mass ratio of 1:1 to obtain LNP test samples targeting different locations (LNPs are cited from the literature: Musunuru, K., Chadwick, AC, Mizoguchi, T. et al. In vivo CRISPR base editing of PCSK9 durably lowers cholesterol in primates. Nature 593,429–434(2021)). The mouse blastoma cell line AML12 was used for testing. Cells were seeded at 50,000 cells / well in 24-well plates. After 12 hours, LNP samples were added to the plates at a dosage of 2.5 μg / ml. The medium was replaced with fresh complete culture medium (ZQ-606) after 4-6 hours, and the cells were cultured for another 3 days. Cells were then collected for mRNA extraction, reverse transcribed into cDNA, and the mRNA expression level of the target gene INHBE was detected using qPCR. The knockdown efficiency of the INHBE gene for each sgRNA was obtained by comparing it with the Non-target group (mRNA and sgRNA without a target site, SEQ ID NO: 336). The qPCR primer sequences are shown in SEQ ID NO: 383, 384, 387, and 388.
[0155] Results 3 days after administration (Figure 1) showed that almost all sgRNAs could achieve some degree of INHBE gene inhibition, with the highest efficiency reaching about 80% (such as sgRNA13 shown in SEQ ID NO:349 and sgRNA18 shown in SEQ ID NO:354).
[0156] Example 3
[0157] The inhibitory efficiency of the complex in this application on INHBE gene expression in primary monkey hepatocytes
[0158] In this embodiment, primary hepatocytes of cynomolgus monkeys were used as the research model. Guide RNAs were designed targeting the region from -2000 to +1000 bp before and after the transcription initiation region of the monkey INHBE gene (SEQ ID NO:394). LNPs were prepared by embedding the mRNA of the epigenetic editing tool (SEQ ID NO:288) and different sgRNAs (SEQ ID NO:357-366) at a mass ratio of 1:1 to obtain LNP test samples targeting different locations (LNPs cited in the literature: Musunuru, K., Chadwick, AC, Mizoguchi, T. et al. In vivo CRISPR base editing of PCSK9 durably lowers cholesterol in primates. Nature 593,429–434(2021)). Primary monkey hepatocytes were used for testing. Cells were seeded at 250,000 cells / well in 24-well plates. After 12 hours, LNP samples were added to the plates at a dosage of 2.5 μg / ml. The culture medium was replaced with fresh maintenance medium after 4-6 hours, and the cells were cultured for another 5 days. Cells were then collected for mRNA extraction, reverse transcribed into cDNA, and the mRNA expression level of the target gene INHBE was detected using qPCR. The knockdown efficiency of each sgRNA was obtained by comparing it with the Non-target group (mRNA and sgRNA without a target site, SEQ ID NO: 336). The qPCR primer sequences are shown in SEQ ID NO: 385, 386, 389, and 390.
[0159] The results after 5 days of drug administration (Figure 2) showed that almost all sgRNAs could achieve some degree of INHBE gene inhibition, with the highest efficiency reaching 88% (such as sgRNA3 shown in SEQ ID NO:359 and sgRNA8 shown in SEQ ID NO:364). The inhibition efficiency of other sgRNAs was also basically maintained at around 80%.
[0160] Example 4
[0161] The inhibitory efficiency of the complex in this application on INHBE gene expression in human Huh7 cells.
[0162] In this embodiment, the human hepatocellular carcinoma cell line Huh7 was used as the research model. Guide RNAs were designed targeting the region from -1500 to +2000 bp before and after the transcription initiation region of the human INHBE gene (SEQ ID NO:395). The mRNA of the epigenetic editing tool (SEQ ID NO:288) and different sgRNAs (SEQ ID NO:358, 362, 364, 365 and 367-382) were embedded in LNPs at a mass ratio of 1:1 to obtain LNP test samples targeting different locations (LNPs cited in the literature: Musunuru, K., Chadwick, AC, Mizoguchi, T. et al. In vivo CRISPR base editing of PCSK9 durably lowers cholesterol in primates. Nature 593,429–434(2021)). Human Huh7 cells were used for testing. 50,000 cells / well were seeded into 24-well plates. After 12 hours, LNP samples were added to the plates at a dosage of 2.5 μg / ml. The culture medium was replaced with fresh medium after 4-6 hours, and the cells were cultured for another 3 days. mRNA was extracted from the cells and reverse transcribed into cDNA. The mRNA expression level of the target gene was detected using qPCR. The knockdown efficiency of each sgRNA was obtained by comparing it with the Non-target group (mRNA and sgRNA without a target site, SEQ ID NO: 336). The qPCR primer sequences are shown in SEQ ID NO: 385, 386, 391, and 392.
[0163] The results 3 days after administration (Figure 3) showed that almost all sgRNAs could achieve some degree of INHBE gene inhibition, with the highest efficiency reaching 99% (such as sgRNA3 shown in SEQ ID NO:368 and sgRNA11 shown in SEQ ID NO:373). The inhibition efficiency of other sgRNAs was also basically maintained at around 90%.
[0164] Example 5
[0165] The inhibitory efficiency of the complex in this application on INHBE gene expression in human Huh7 cells.
[0166] In this embodiment, the human hepatocellular carcinoma cell line Huh7 was used as the research model. Guide RNA was designed targeting the region from -1500 to +2000 bp before and after the transcription initiation region of the human INHBE gene (SEQ ID NO:395). The mRNA of the epigenetic editing tool (SEQ ID NO:288) and different sgRNAs were embedded in LNP at a mass ratio of 1:1 to prepare LNP test samples targeting different locations (LNP reference: Musunuru, K., Chadwick, AC, Mizoguchi, T. et al. In vivo CRISPR base editing of PCSK9 durably lowers cholesterol in primates. Nature 593,429–434(2021)). Human Huh7 cells were used for testing. 50,000 cells / well were seeded into 24-well plates. After 12 hours, LNP samples were added to the wells at a drug dosage of 0.5 μg / ml. The plates were incubated at 37°C for 4 hours, then the medium was replaced with fresh, drug-free medium. Cells were cultured for another 7 days. After incubation, cells were collected for mRNA extraction, reverse transcription into cDNA, and qPCR was used to detect the mRNA expression level of the target gene. The inhibition rate of different candidate sgRNAs was calculated using the following formula: The knockdown efficiency of each sgRNA was obtained by comparing it with the Non-target group (mRNA and sgRNA without a target site, sgRNA sequence: CUGAAGGUGUCUGGCAGAGC, i.e., SEQ ID NO:336). The qPCR primer sequences are shown in SEQ ID NO:385, 386, 391 and 392.
[0167] The results showed that the genomic location of chromosome 12, from 57453807 to 57457306 (reference genome: GRCh38.p14 Primary Assembly), located near the TSS of the INHBE gene, within a range of -1500bp to +2000bp, was divided into three regions based on the overall trend of suppression efficiency: Region 1 was -750bp to 0bp (genomic location: 57454557-57455306), Region 2 was 0 to +750bp (genomic location: 57455307 to 57456056), and Region 3 was +750bp to 1500bp (genomic location: 57456057 to 57456806). The inhibition efficiency of all sgRNAs in the three regions was graded, including the average inhibition efficiency of all sgRNAs in different regions, the proportion of sgRNAs with inhibition efficiencies exceeding 40% and 80%, and the results in Tables 2-4 and Figure 4. The results show that regions 1, 2 and 3 all have good inhibition efficiencies, and region 1 has the best overall inhibition efficiency (as shown in Table 2), which proves that these regions have high gene inhibition efficiency and potential drug development.
[0168] Table 2. Inhibitory effects of different candidate sgRNAs on INHBE expression in region 1 7 days after drug administration.
[0169] Table 3. Inhibitory effects of different candidate sgRNAs on INHBE expression in Region 2 7 days after drug administration.
[0170] Table 4. Inhibitory effects of different candidate sgRNAs on INHBE expression in Region 3 7 days after drug administration.
Claims
1. A complex comprising a first fusion and a second fusion, wherein: 1) A fusion of one of the first fusion and the second fusion comprises a DNA methylation domain and at least one recruitment domain A, and the other fusion comprises a transcriptional repressor domain and at least one recruitment domain A'; and 2) The first fusion compound or the second fusion compound contains a nucleic acid binding domain; Furthermore, the recruitment domain A and recruitment domain A' can interact to recruit a fusion of one of the first fusion and the second fusion, or a portion thereof, to the vicinity of the other fusion; the nucleic acid binding domain can specifically bind to the INHBE gene transcriptional regulatory element, which includes a transcription start site, a core promoter, a promoter, an enhancer, a silencer, an insulator element, a boundary element, and / or a locus control region; the nucleic acid binding domain can specifically bind to the target nucleotide sequence within a region 1500 bp upstream to 2000 bp downstream of the INHBE gene transcription start site.
2. The complex according to claim 1, wherein the nucleic acid binding domain is a DNA binding domain.
3. The complex according to claim 1 or 2, wherein the DNA-binding domain is selected from: TALE domain, zinc finger domain, tetR domain, large-scale nuclease, Cas protein, Argonaute (Ago) protein, and their homologues, modified forms or variants.
4. The complex according to claim 2 or 3, wherein the DNA-binding domain is capable of binding the target nucleotide sequence.
5. The complex according to any one of claims 2-4, wherein the DNA-binding domain is capable of binding to guide RNA.
6. The complex according to claim 5, wherein the guide RNA is capable of specifically recognizing and hybridizing with the target nucleotide sequence.
7. The complex according to any one of claims 5-6, wherein the guide RNA is capable of specifically recognizing and hybridizing with target nucleotide sequences within a region approximately 1500 bp, approximately 1000 bp, approximately 750 bp, approximately 500 bp, approximately 250 bp, approximately 200 bp, approximately 100 bp, and approximately 50 bp upstream of the transcription start site of the INHBE gene.
8. The complex according to any one of claims 5-7, wherein the guide RNA is capable of specifically recognizing and hybridizing with target nucleotide sequences within a region approximately 2000 bp, approximately 1500 bp, approximately 1000 bp, approximately 750 bp, approximately 500 bp, approximately 250 bp, approximately 200 bp, approximately 100 bp, and approximately 50 bp downstream of the transcription start site of the INHBE gene.
9. The complex according to claim 8, wherein the guide RNA is capable of specifically recognizing and hybridizing with target nucleotide sequences within the regions approximately 1000 bp upstream to approximately 1500 bp downstream, approximately 800 bp upstream to approximately 1500 bp downstream, approximately 800 bp upstream to approximately 1000 bp downstream, approximately 800 bp upstream to approximately 800 bp downstream, and approximately 750 bp upstream to approximately 750 bp downstream of the INHBE gene.
10. The complex according to claim 9, wherein the guide RNA is capable of specifically recognizing and hybridizing with a target nucleotide sequence within a region approximately 750 bp upstream of the transcription start site of the INHBE gene.
11. The complex according to claim 9, wherein the guide RNA is capable of specifically recognizing and hybridizing with a target nucleotide sequence within a region approximately 750 bp downstream of the INHBE gene transcription start site.
12. The complex according to claim 11, wherein the guide RNA is capable of specifically recognizing and hybridizing with a target nucleotide sequence in a region approximately 400 bp to approximately 500 bp downstream of the transcription start site of the INHBE gene.
13. The complex according to any one of claims 11-12, wherein the guide RNA is capable of specifically recognizing and hybridizing with a target nucleotide sequence in a region approximately 450 bp to approximately 500 bp downstream of the transcription start site of the INHBE gene.
14. The complex according to claim 9, wherein the guide RNA is capable of specifically recognizing and hybridizing with a target nucleotide sequence in a region approximately 750 bp to approximately 1500 bp downstream of the transcription start site of the INHBE gene.
15. The complex according to any one of claims 5-14, wherein the guide RNA comprises SEQ ID NOs: 337-382, 451, 452, 453, 455-458, 465, 466, 468-470, 473, 475, 477, 478, 480-482, 484, 485, 487, 489-495, 497, 498, 502, 503, 505, 507-510, 512, 513, 515, 517-525, 528, 532, 533, 537, 540, 402-406, 413, 416, 423, 424, 426, 42 7. The nucleotide sequence shown in any one of the following: 434-437, 442, 443, 448, 454, 459-464, 467, 472, 474, 476, 479, 488, 496, 499-501, 504, 506, 511, 514, 516, 529-531, 534, 535, 407, 409, 418-422, 425, 428, 430-433, 438-441, 444-447, 449, 541, 542, 545, 547-549.
16. The complex according to any one of claims 5-15, wherein the guide RNA comprises SEQ ID NOs: 337-382, 451, 452, 453, 455-458, 465, 466, 468-470, 473, 475, 477, 478, 480-482, 484, 485, 487, 489-495, 497, 498, 502, 503, 505, 507-510, 512, 513, 515, 517-525, 528, 532, 533, 537, 540, 402-406, 413, 416, 423, 424, 426, 427, 434-437, 44 2. A partial sequence of any one of the nucleotide sequences shown in 443, 448, 454, 459-464, 467, 472, 474, 476, 479, 488, 496, 499-501, 504, 506, 511, 514, 516, 529-531, 534, 535, 407, 409, 418-422, 425, 428, 430-433, 438-441, 444-447, 449, 541, 542, 545, 547-549, wherein the partial sequence is 15-20 base pairs in length.
17. The complex according to any one of claims 2-16, wherein the DNA-binding domain is a Cas protein, and the Cas protein is a type II Cas nuclease.
18. The complex according to claim 17, wherein the Cas protein is selected from type II Cas nucleases and type II V Cas nucleases.
19. The complex according to claim 17 or 18, wherein the Cas protein is Cas9 or Cas12 protein.
20. The complex according to any one of claims 17-19, wherein the Cas protein is an inactivated Cas9 (dCas9) protein or an inactivated Cas12 (dCas12) protein.
21. The complex according to any one of claims 2-20, wherein the DNA binding domain comprises the amino acid sequence shown in any one of SEQ ID NO:1-9.
22. The complex according to any one of claims 1-21, wherein the first fusion comprises a DNA methylation domain, a nucleic acid binding domain and at least one recruitment domain A, and the second fusion comprises a transcriptional repressor domain and at least one recruitment domain A'.
23. The complex according to any one of claims 1-22, wherein the first fusion comprises, from the N-terminus to the C-terminus, a DNA methylation domain, a nucleic acid binding domain, and a recruitment domain A.
24. The complex according to any one of claims 1-23, wherein the second fusion comprises, from the N-terminus to the C-terminus, a transcriptional repressor domain and a recruitment domain A', or from the N-terminus to the C-terminus, a recruitment domain A' and a transcriptional repressor domain.
25. The complex according to any one of claims 1-24, wherein the first fusion comprises a transcriptional repressor domain, a nucleic acid binding domain and at least one recruitment domain A, and the second fusion comprises a DNA methylation domain and at least one recruitment domain A'.
26. The complex according to any one of claims 1-22, wherein the first fusion comprises, from the N-terminus to the C-terminus, a recruitment domain A, a nucleic acid binding domain, and a transcriptional repressor domain.
27. The complex according to any one of claims 1-22 and 26, wherein the second fusion comprises, from the N-terminus to the C-terminus, a DNA methylation domain and a recruitment domain A', or from the N-terminus to the C-terminus, a recruitment domain A' and a DNA methylation domain.
28. The complex according to any one of claims 1-27, characterized in that: 1) The first fusion compound contains, from N-terminus to C-terminus, a DNA methylation domain, a nucleic acid binding domain, and a recruitment domain A, respectively; the second fusion compound contains, from N-terminus to C-terminus, a transcriptional repressor domain and a recruitment domain A', respectively; or 2) The first fusion compound contains, from N-terminus to C-terminus, a DNA methylation domain, a nucleic acid binding domain, and a recruitment domain A; the second fusion compound contains, from N-terminus to C-terminus, a recruitment domain A' and a transcriptional repressor domain; or 3) The first fusion compound contains, from N-terminus to C-terminus, a recruitment domain A, a nucleic acid binding domain, and a transcriptional repressor domain, respectively; the second fusion compound contains, from N-terminus to C-terminus, a DNA methylation domain and a recruitment domain A', respectively; or 4) The first fusion compound contains, from N-terminus to C-terminus, a recruitment domain A, a nucleic acid binding domain and a transcriptional repressor domain, and the second fusion compound contains, from N-terminus to C-terminus, a recruitment domain A' and a DNA methylation domain.
29. The complex according to any one of claims 1-28, wherein the recruitment domain A is selected from one of two groups of domains, and the recruitment domain A' is selected from the other of two groups of domains: 1) Universally controlled non-derepressor protein 4 (GCN4), a GFP11 fragment derived from splitting green fluorescent protein (GFP), or a GVKESLV polypeptide; and 2) Single-chain antibody (scFv), GFP1-10 fragments derived from split green fluorescent protein (GFP), or PDZ protein domain.
30. The complex according to any one of claims 1-29, wherein: 1) One of the recruitment domains A and A' is a domain GCN4, and the other is a domain scFv; or 2) One of the recruitment domains A and A' is a GFP11 fragment, and the other domain is GFP1-10; or 3) One of the recruitment domains A and A' is GVKESLV, and the other of the domains is the PDZ protein domain.
31. The complex according to any one of claims 1-30, wherein the DNA methylation domain comprises at least one DNA methyltransferase or a functionally active fragment thereof.
32. The complex according to claim 31, wherein the DNA methyltransferase is selected from DNMT3A, DNMT3B, DNMT3C, DNMT1, DNMT2 and DNMT3L.
33. The complex according to any one of claims 1-32, wherein the DNA methylation domain comprises at least one DNMT3A and at least one DNMT3L.
34. The complex according to claim 31 or 32, wherein the DNA methyltransferase comprises the amino acid sequence shown in any one of SEQ ID NO:19-24.
35. The complex according to any one of claims 1-34, wherein the DNA methylation domain comprises a DNMT3A-DNMT3L domain or a DNMT3L-DNMT3A domain; wherein, - indicates that the structural fields at both ends are directly or indirectly connected in order from the N end to the C end.
36. The complex according to any one of claims 1-35, wherein the transcriptional repressor is selected from one or more of the following domains: KRAB, ZIM3 KRAB, ZNF680, ZNF554, ZNF264, ZNF582, ZNF324, ZNF669, ZNF354A, ZNF82, ZNF595, ZNF 419, ZNF566, ZIM2, EHMT2, SUV39H1, ZFPM1, TRIM28, EZH2, MXD1, SID, LSD1, HP1a, HDAC 3. HDAC1, PRMT1, SETDB1, hSIRT1, ZNF436, ZNF257, ZNF675, ZNF490, ZNF320, ZNF331, Z NF816, ZNF41, ZNF189, ZNF528, ZNF543, ZNF140, ZNF610, ZNF350, ZNF8, ZNF30, ZNF98, Z NF677, ZNF596, ZNF214, ZNF37A, ZNF34, ZNF250, ZNF547, ZNF273, ZFP82, ZNF224, ZNF3 3A, ZNF45, ZNF175, ZNF184, ZFP28-1, ZFP28-2, ZNF18, ZNF213, ZNF394, ZFP1, ZFP14, ZN F416, ZNF557, ZNF729, ZNF254, ZNF764, ZNF785, ZNF10, CBX5, RYBP, YAF2, MGA, CBX1, SCMH1, MPP8, SUMO3, HERC2, BIN1, PCGF2, TOX, FOXA1, FOXA2, IRF2BP1, IRF2BP2, IRF2BPL IRF-2BP1_2N-terminal domain, HOXA13, HOXB13, HOXC13, HOXA11, HOXC11, HOXC10, HOXA10, HOXB9, HOXA9, ZFP28, ZN334, ZN568, ZN37A, ZN181, ZN510, ZN862, ZN140, ZN208, ZN248, ZN571, ZN699, ZN726, ZIK1, ZNF2, Z705F, ZNF14, ZN471 , ZN624, ZNF84, ZNF7, ZN891, ZN337, Z705G, ZN529, ZN729, ZN419, Z705A, ZN302, ZN486, ZN621, ZN688, ZN3 3A, ZN554, ZN878, ZN772, ZN224, ZN184, ZN544, ZNF57, ZN283, ZN549, ZN211, ZN615, ZN253, ZN226, ZN730,Z585A,ZN732,ZN681,ZN667,ZN649,ZN470,ZN484,ZN431,ZN382,ZN254,ZN124,ZN607,ZN317,ZN620,ZN141,ZN584,ZN540,ZN75D,ZN555,ZN658,ZN684,RBAK,ZN829,ZN582,ZN112,ZN716,HKR1,ZN350,ZN480,ZN416,ZNF92,ZN100,ZN736,ZNF74,ZN443,ZN195,ZN530,ZN782,ZN791,ZN331,Z354C,ZN157,ZN727,ZN550,ZN793,ZN235,ZN724,ZN573,ZN577,ZN789,ZN718,ZN300,ZN383,ZN429,ZN677,ZN850,ZN454,ZN257,ZN264,ZN485,ZN737,ZNF44,ZN596,ZN565,ZN543,ZFP69,SUMO1,ZNF12,ZN169,ZN433,ZN175,ZN347,ZNF25,ZN519,Z585B,ZN517,ZN846,ZN230,ZNF66,ZN713,ZN816,ZN426,ZN674,ZN627,ZNF20,Z587B,ZN316,ZN233,ZN611,ZN556,ZN234,ZN560,ZNF77,ZN682,ZN614,ZN785,ZN445,ZFP30,ZN225,ZN551,ZN610,ZN528,ZN284,ZN418,ZN490,ZN805,Z780B,ZN763,ZN285,ZNF85,ZN223,ZNF90,ZN557,ZN425,ZN229,ZN606,ZN155,ZN222,ZN442,ZNF91,ZN135,ZN778,ZN534,ZN586,ZN567,ZN440,ZN583,ZN441,ZNF43,ZN589,ZN563,ZN561,ZN136,ZN630,ZN527,ZN333,Z324B,ZN786,ZN709,ZN792,ZN599,ZN613,ZF69B,ZN799,ZN569,ZN564,ZN546,ZFP92,ZN723,ZN439,ZFP57,ZNF19,ZN404,ZN274,CBX3,ZN250,ZN570,ZN675,ZN695,ZN548,ZN132,ZN738,ZN420,ZN626,ZN559,ZN460,ZN268,ZN304,ZN605,ZN844,SUMO5,ZN101,ZN783,ZN417,ZN182,ZN823,ZN177,ZN197,ZN717,ZN669,ZN256,ZN251,CBX4,CDY2,CDYL2,ZN562,ZN461,Z324A,ZN766,ID2,ZN214,CBX7,ID1,CREM,SCX,ASCL1,ZN764,SCML2,TWST1,CREB1,TERF1,ID3,CBX8,GSX1,NKX22,ATF1,TWST2,ZNF17,TOX3,TOX4,ZMYM3,I2BP1,RHXF1,SSX2,I2BPL,ZN680,TRI68,HXA13,PHC3,TCF24,HXB13,HEY1,PHC2,ZNF81,FIGLA,SAM11,KMT2B,HEY2,JDP2,HXC13,ASCL4,HHEX,GSX2,ETV7,ASCL3,PHC1,OTP,I2BP2,VGLL2,HXA11,PDLI4,ASCL2,CDX4,ZN860,LMBL4,PDIP3,NKX25,CEBPB,ISL1,CDX2,PROP1,SIN3B,SMBT1,HXC11,HXC10,PRS6A,VSX1,NKX23,MTG16,HMX3,HMX1,KIF22,CSTF2,CEBPE,DLX2,PPARG,PRIC1,UNC4,BARX2,ALX3,TCF15,TERA,VSX2,HXD12,CDX1,TCF23,ALX1,HXA10,RX,CXXC5,SCML1,NFIL3,DLX6,MTG8,CEBPD,SEC13,FIP1,ALX4,LHX3,PRIC2,MAGI3,NELL1,PRRX1,MTG8R,RAX2,DLX3,DLX1,NKX26,NAB1,SAMD7,PITX3,WDR5,MEOX2,NAB2,DHX8,CBX6,EMX2,CPSF6,HXC12,KDM4B,LMBL3,PHX2A,EMX1,NC2B,DLX4,SRY,ZN777,ZN398,GATA3,BSH,SF3B4,TEAD1,TEAD3,RGAP1,PHF1,GATA2,FOXO3,ZN212,IRX4,ZBED6,LHX4,SIN3A,RBBP7,NKX61,R51A1,MB3L1,DLX5,NOTC1,TERF2,ZN282,RGS12,ZN840,SPI2B,PAX7,NKX62,ASXL2, FOXO1, GATA1, ZMYM5, LRP1, MIXL1, SGT1, LMCD1, CEBPA, SOX14, WTIP, PRP19, NKX11, RBBP4, DMRT2, SMCA2, and their functionally active fragments.
37. The complex according to any one of claims 1-36, wherein the transcriptional repressor domain comprises the amino acid sequence shown in any one of SEQ ID NOs:25-50.
38. The complex according to any one of claims 1-37, wherein: 1) One of the first fusion and the second fusion comprises a DNA methylation domain -dCas9 or dCas12-n×GCN4, and the other fusion comprises a transcriptional repressor domain -scFv; or 2) One of the first fusion and the second fusion contains a DNA methylation domain -dCas9 or dCas12-scFv, and the other fusion contains a transcriptional repressor domain -GCN4; or 3) One of the first fusion and the second fusion contains a DNA methylation domain -dCas9 or dCas12-n×GFP11, and the other fusion contains a transcriptional repressor domain -GFP1-10; or 4) One of the first fusion and the second fusion contains a DNA methylation domain -dCas9 or dCas12-GFP1-10, and the other fusion contains a transcriptional repressor domain -GFP11; or 5) One of the first fusion and the second fusion contains a DNA methylation domain -dCas9 or dCas12-n×GCN4, and the other fusion contains an scFv-transcriptional repressor domain; or 6) One of the first fusion and the second fusion comprises a DNA methylation domain -dCas9 or dCas12-scFv, and the other fusion comprises a GCN4-transcriptional repressor domain; or 7) One of the first fusion and the second fusion comprises a DNA methylation domain -dCas9 or dCas12-n×GFP11, and the other fusion comprises a GFP1-10-transcriptional repressor domain; or 8) One of the first fusion and the second fusion contains a DNA methylation domain -dCas9 or dCas12-GFP1-10, and the other fusion contains a GFP11-transcriptional repressor domain; Wherein, - indicates that the structural domains at both ends are directly or indirectly connected in the order from the N end to the C end; n×GCN4 or n×GFP11 represent n copies of GCN4 connected by the adapter sequence or n copies of GFP11 connected by the adapter sequence, respectively, where n is selected from any integer from 1 to 20.
39. The complex according to any one of claims 1-37, wherein the first fusion and / or the second fusion comprises the amino acid sequence of any one of SEQ ID NO: 51-76, 78-82, 85-93, 103-105, 110-115, 123 and 124.
40. The complex according to any one of claims 1-39, comprising the amino acid sequence shown in any one of SEQ ID NO: 133-142, 153, 154, 158-163 and 168.
41. The complex according to any one of claims 1-37, wherein: 1) One of the first fusion and the second fusion contains an n×GCN4-dCas9 or dCas12-transcriptional repressor domain, and the other fusion contains a DNA methylation domain -scFv; or 2) One of the first fusion and the second fusion contains an scFv-dCas9 or dCas12 transcriptional repressor domain, and the other fusion contains a DNA methylation domain -GCN4; or 3) One of the first fusion and the second fusion contains an n×GFP11-dCas9 or dCas12-transcriptional repressor domain, and the other fusion contains a DNA methylation domain -GFP1-10; or 4) One of the first fusion and the second fusion contains a GFP1-10-dCas9 or dCas12-transcriptional repressor domain, and the other fusion contains a DNA methylation domain-GFP11; or 5) One of the first fusion and the second fusion contains an n×GCN4-dCas9 or dCas12-transcriptional repressor domain, and the other fusion contains an scFv-DNA methylation domain; or 6) One of the first fusion and the second fusion contains an scFv-dCas9 or dCas12-transcriptional repressor domain, and the other fusion contains a GCN4-DNA methylation domain; or 7) One of the first fusion and the second fusion contains an n×GFP11-dCas9 or dCas12-transcriptional repressor domain, and the other fusion contains a GFP1-10-DNA methylation domain; or 8) One of the first fusion and the second fusion contains a GFP1-10-dCas9 or dCas12-transcriptional repressor domain, and the other fusion contains a GFP11-DNA methylation domain; Wherein, - indicates that the structural domains at both ends are directly or indirectly connected in the order from the N end to the C end; n×GCN4 or n×GFP11 represent n copies of GCN4 connected by the adapter sequence or n copies of GFP11 connected by the adapter sequence, respectively, where n is selected from any integer from 1 to 20.
42. The complex according to any one of claims 1-37 and 41, wherein the first fusion and / or the second fusion comprises the amino acid sequence shown in any one of SEQ ID NO: 83, 84, 94-102, 106-109 and 116-122.
43. The complex according to any one of claims 1-37, 41 and 42, comprising the amino acid sequence shown in any one of SEQ ID NO: 143-152, 155-157 and 164-167.
44. The complex according to any one of claims 1-43, wherein the complex further comprises a nuclear localization signal and / or a marker domain.
45. The complex according to any one of claims 1-44, which is capable of providing modification of at least one nucleotide in the region from 2000 bp upstream to 1000 bp downstream of the transcription start site of the INHBE gene.
46. The complex according to claim 45, which is capable of providing modification of at least one nucleotide in the region of about 1500 bp, about 1000 bp, about 750 bp, about 500 bp, about 250 bp, about 200 bp, about 100 bp, and about 50 bp upstream of the transcription start site of the INHBE gene.
47. The complex according to any one of claims 45-46, which is capable of providing modification of at least one nucleotide in a region of about 2000 bp, about 1500 bp, about 1000 bp, about 750 bp, about 500 bp, about 250 bp, about 200 bp, about 100 bp, and about 50 bp downstream of the transcription start site of the INHBE gene.
48. The complex according to claim 47, wherein it is capable of providing modification of at least one nucleotide in the regions of approximately 1000 bp upstream to approximately 1500 bp downstream, approximately 800 bp upstream to approximately 1500 bp downstream, approximately 800 bp upstream to approximately 1000 bp downstream, approximately 800 bp upstream to approximately 800 bp downstream, and approximately 750 bp upstream to approximately 750 bp downstream of the INHBE gene.
49. The complex according to claim 48, which is capable of providing modification of at least one nucleotide in a region of about 750 bp upstream of the transcription start site of the INHBE gene.
50. The complex according to claim 48, which is capable of providing modification of at least one nucleotide in a region of about 750 bp downstream of the transcription start site of the INHBE gene.
51. The complex according to claim 50, which is capable of providing modification of at least one nucleotide in a region of about 400 bp to about 500 bp downstream of the transcription start site of the INHBE gene.
52. The complex according to any one of claims 50-51, which is capable of providing modification of at least one nucleotide in a region of about 450 bp to about 500 bp downstream of the transcription start site of the INHBE gene.
53. The complex according to claim 48, which is capable of providing modification of at least one nucleotide in a region from about 750 bp downstream to about 1500 bp downstream of the transcription start site of the INHBE gene.
54. The nucleic acid encoding the complex of any one of claims 1-53.
55. The nucleic acid according to claim 54, wherein it is a recombinant vector.
56. The nucleic acid according to claim 55, wherein the recombinant vector further comprises a non-coding region.
57. The nucleic acid according to claim 56, wherein the non-coding region is selected from introns, regulatory elements, promoters, enhancers, termination sequences, and 5' and 3' untranslated regions.
58. The nucleic acid according to any one of claims 54-57, comprising a first nucleic acid fragment encoding the first fusion compound and a second nucleic acid fragment encoding the second fusion compound.
59. The nucleic acid according to claim 58, wherein the first nucleic acid fragment and the second nucleic acid fragment are linked by a nucleic acid fragment encoding a cleavage peptide.
60. The nucleic acid according to claim 59, wherein the cleavage peptide is a 2A peptide and / or IRES.
61. The nucleic acid according to claim 60, wherein the 2A peptide is selected from P2A, T2A, E2A and F2A.
62. The nucleic acid according to any one of claims 54-61, comprising the nucleic acid sequence shown in any one of SEQ ID NO:169-335.
63. A delivery vector comprising a complex according to any one of claims 1-53 and / or a nucleic acid according to any one of claims 54-62, and optionally comprising liposomes and / or lipid nanoparticles.
64. A composition comprising the complex of any one of claims 1-53, the nucleic acid of any one of claims 54-62, and / or the delivery vector of claim 63.
65. A cell comprising the complex of any one of claims 1-53, the nucleic acid of any one of claims 54-62, the delivery vector of claim 63, and / or the composition of claim 64.
66. A kit comprising the complex of any one of claims 1-53, the nucleic acid of any one of claims 54-62, the delivery vector of claim 63, the composition of claim 64, and / or the cells of claim 65.
67. A method for regulating the expression of the INHBE gene product, the method comprising administering a complex according to any one of claims 1-53, a nucleic acid according to any one of claims 54-62, a delivery vector according to claim 63, a composition according to claim 64, a cell according to claim 65, and / or a kit according to claim 66.
68. The method of claim 67, the method comprising introducing the complex, the nucleic acid, the delivery vector, the composition, the cell, and / or the kit into cells containing the INHBE gene.
69. The method of claim 67, wherein the method comprises contacting the complex, the nucleic acid, the delivery vector, and / or the composition with a target nucleotide sequence near the transcription start site of the INHBE gene.
70. The method of claim 69, wherein the regulatory element comprises a core promoter, a proximal promoter, a distal enhancer, a silencer, an insulator element, a boundary element, and / or a locus control region.
71. A method for treating or alleviating a disease or symptom thereof associated with abnormal INHBE gene expression and / or abnormal INHBE gene activity, the method comprising administering to a subject in need an effective amount of the complex of any one of claims 1-53, the nucleic acid of any one of claims 54-62, the delivery vector of claim 63, the composition of claim 64, the cell of claim 65, and / or the kit of claim 66.
72. Use of the complex of any one of claims 1-53, the nucleic acid of any one of claims 54-62, the delivery vector of claim 63, the composition of claim 64, the cell of claim 65, and / or the kit of claim 66 for the preparation of a medicament for the treatment or relief of a disease or condition associated with abnormal INHBE gene expression and / or abnormal INHBE gene activity.
73. The method of claim 71, wherein the disease or condition associated with abnormal INHBE gene expression and / or abnormal INHBE gene activity includes obesity and / or metabolic syndrome.
Citation Information
Patent Citations
Methods of treating metabolic disorders and cardiovascular diseases with inhibitors of statin subunit beta E (INHBE)
CN116583291A
Compositions and methods of genome editing
WO2023165597A1
Complex and use thereof
WO2024131917A1