Compositions and methods for epigenetic editing
By developing fusion proteins with DNA binding, methyl transfer, repression and nuclear localization functions, epigenetic modification of target genes is achieved without changing the genome sequence, solving the risk of genome editing and improving the accuracy and safety of editing.
Patent Information
- Application Number
- CN202380060162.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-06-23
- Filing Date
- 2023-06-23
- Publication Date
- 2025-05-06
AI Technical Summary
There are undesirable double-strand breaks, heterologous repair and toxicity possibilities for genome editing, and it is difficult to achieve epigenetic modification without introducing genomic sequence changes.
An epigenetic editing system was developed, including a fusion protein with a DNA binding domain, a DNA methyltransferase (DNMT) domain, a repressor domain, and a nuclear localization sequence (NLS) to achieve epigenetic modifications in the target genome.
This system can achieve epigenetic modification of target genes without changing the genome sequence, reducing the risk of genome editing and improving the accuracy and safety of editing.
Smart Images

Figure BDA0005273779540000261 
Figure BDA0005273779540001051 
Figure BDA0005273779540001061
Abstract
Description
[0001] Cross-references
[0002] This application claims the benefit of U.S. Provisional Application No. 63 / 354,931, filed on June 23, 2022, which is incorporated herein by reference in its entirety. Background Art
[0003] For more than a decade, genome editing has been considered a promising therapeutic approach for treating genetic diseases. However, manipulation at the DNA level remains risky given the potential for undesired double-strand breaks, heterologous repair (including large and small insertions and deletions at the intended site), and toxicity. Summary of the invention
[0004] Provided herein are compositions for epigenetic modification associated with epigenetic editing systems, and methods of using the compositions to produce epigenetic modifications in target genomes, including target genomes in host cells and organisms, without introducing genomic sequence changes.
[0005] The present invention describes an epigenetic editing system, which includes: (a) a fusion protein comprising a DNA binding domain, a DNA methyltransferase (DNMT) domain, a repressor domain and two nuclear localization sequences (NLS), wherein each of the two NLSs is located at the amino (N) terminus or the carboxyl (C) terminus of the fusion protein; or (b) a nucleic acid molecule encoding the fusion protein of (a).
[0006] In some embodiments, the fusion protein comprises one or more NLS at the C-terminus of the fusion protein and one or more NLS at the N-terminus of the fusion protein, optionally wherein the fusion protein comprises two NLS at the N-terminus of the fusion protein and two NLS at the C-terminus of the fusion protein. In some embodiments, the DNMT domain is from a bacterial species. In some embodiments, the DNMT domain is a mammalian DNMT domain. In some embodiments, the DNMT domain is a mouse DNMT domain. In some embodiments, the DNMT domain is a human DNMT domain. In some embodiments, the DNMT domain is a DNMT3A domain. In some embodiments, the DNMT domain is a DNMT3L domain. In some embodiments, the fusion protein comprises a DNMT3A domain and a DNMT3L domain.
[0007] In some embodiments, the DNMT3L domain is from a species selected from the group consisting of giant panda (Ailuropoda melanoleuca), Philippine tarsier (Carlito syrichta), long-clawed gerbil (Meriones unguiculatus), North American pika (Ochotona princeps), Carolina new squirrel (Neosciurus carolinensis), American bison (Bisonbison), Przewalski's horse (Equus przewalskii), Carolinian mouse (Mus caroli) and chimpanzee (Pantroglodytes); the DNMT3L domain optionally comprises one of SEQ ID NOs:72-80 and SEQ ID NOs:101-109, or an amino acid sequence that is at least 90%, optionally at least 95% homologous thereto.
[0008] In some embodiments, the repressor domain includes a KRAB domain of a ZFP28, ZN627, KAP1, MeCP2, HP1b, CBX8, CDYL2, TOX, Tox3, Tox4, EED, RBBP4, RCOR1, or SCML2 protein, or a fusion of the N-terminal and C-terminal regions of the ZIM3 and KOX1 KRAB domains.
[0009] Described herein is an epigenetic editing system comprising: (a) a fusion protein comprising a DNA binding domain and a DNA methyltransferase (DNMT) domain from a bacterial species, wherein the DNMT domain is fused to the N-terminus of the DNA binding domain; or (b) a nucleic acid molecule encoding the fusion protein of (a).
[0010] In some embodiments, the fusion protein plays a role in methylating the mammalian target DNA in the cell. In some embodiments, the DNMT domain of the fusion protein does not include any one of SEQ ID NO:81-93. In some embodiments, the bacterial species is not M.penetrans, S.monbiae, H.parainfluenzae, A.luteus, H.aegyptius, H.haemolyticus, Moraxella, E.coli, T.aquaticus, C.crescentus or C.difficile.
[0011] In some embodiments, the DNMT domain from a bacterial species is derived from (a) M.Sss1, optionally comprising SEQ ID NO:40 or an amino acid sequence at least 90%, optionally at least 95% homologous thereto; NQZ29229, optionally comprising SEQ ID NO:41 or an amino acid sequence at least 90%, optionally at least 95% homologous thereto; WP_131599610, optionally comprising SEQ ID NO:42 or an amino acid sequence at least 90%, optionally at least 95% homologous thereto; or WP_208057179, optionally comprising SEQ ID NO:43 or an amino acid sequence at least 90%, optionally at least 95% homologous thereto.
[0012] The present invention describes an epigenetic editing system, which comprises: (a) a fusion protein comprising a DNA binding domain, and a DNMT3L domain, wherein the DNMT3L domain is from a species selected from the group consisting of giant panda, Philippine tarsier, long-clawed gerbil, North American pika, Carolina new squirrel, American bison, Przewalski's horse, Mouse of Carolinian origin and chimpanzee; the DNMT3L domain optionally comprises one of SEQ ID NOs: 72-80 and SEQ ID NOs: 101-109, or an amino acid sequence at least 90%, optionally at least 95% homologous thereto; or (b) a nucleic acid molecule encoding the fusion protein of (a).
[0013] In some embodiments, (i) the fusion protein further comprises a repressor domain, or (ii) the system further comprises other fusion proteins comprising a DNA binding domain and a repressor domain, or nucleic acid molecules encoding other fusion proteins. In some embodiments, the repressor domain comprises a KRAB domain, which is optionally derived from KOX1, ZIM3, ZFP28 or ZN627.
[0014] In some embodiments, the KRAB domain is derived from human KOX1, which optionally comprises SEQ ID NO: 94, or an amino acid sequence at least 90%, optionally at least 95% homologous thereto, or comprises SEQ ID NO: 100, or an amino acid sequence at least 90%, optionally at least 95% homologous thereto. In some embodiments, the repressor domain is derived from KAP1, MECP2, HP1a / CBX5, HP1b, CBX8, CDYL2, TOX, TOX3, TOX4, EED, EZH2, RBBP4, RCOR1, or SCML2.
[0015] Described herein is an epigenetic editing system comprising (a) a fusion protein comprising: a DNA binding domain, and a repressor domain, wherein the repressor domain comprises a KRAB domain of a ZFP28, ZN627, KAP1, MeCP2, HP1b, CBX8, CDYL2, TOX, Tox3, Tox4, EED, RBBP4, RCOR1 or SCML2 protein, or a fusion of the N-terminal and C-terminal regions of ZIM3 and KOX1 KRAB; or (b) a nucleic acid molecule encoding the fusion protein of (a).
[0016] In some embodiments, the repressor domain comprises one of SEQ ID NO: 44-57, or an amino acid sequence at least 90%, optionally at least 95% homologous thereto. In some embodiments, the fusion protein further comprises a DNA methyltransferase (DNMT) domain, or the system further comprises other fusion proteins comprising a DNA binding domain and a DNMT domain, or nucleic acid molecules encoding other fusion proteins.
[0017] In some embodiments, the fusion protein or system comprises a human DNMT3A domain and a human DNMT3L domain, or a human DNMT3A domain and a mouse DNMT3L domain, optionally (a) wherein the human DNMT3A domain comprises SEQ ID NO: 12, or an amino acid sequence at least 90%, optionally at least 95% homologous thereto; (b) wherein the human DNMT3L domain comprises SEQ ID NO: 13, or an amino acid sequence at least 90%, optionally at least 95% homologous thereto, and / or (c) wherein the mouse DNMT3L domain comprises SEQ ID NO: 15, or an amino acid sequence at least 90%, optionally at least 95% homologous thereto.
[0018] In some embodiments, the epigenetic editing system further comprises a fusion protein comprising one or more NLS, a DNMT3A domain, a DNMT3L domain, a DNA binding domain, a repressor domain and one or more NLS from the N-terminus to the C-terminus, optionally wherein the fusion protein comprises a peptide linker between adjacent domains. In some embodiments, the fusion protein comprises two NLS, a DNMT3A domain, an ADD, a DNMT3L domain, a first peptide linker, a DNA binding domain, a second peptide linker, a repressor domain and two NLS from the N-terminus to the C-terminus.
[0019] In some embodiments, the DNMT3A domain is a human DNMT3A domain and / or the DNMT3L domain is a human DNMT3L domain.In some embodiments, the repressor domain is a KRAB domain of ZFP28, ZF627 or KOX1 from a mammal, optionally from a human.
[0020] In some embodiments, the first peptide linker is XTEN80 (SEQ ID NO: 3) and / or the second peptide linker is XTEN16 (SEQ ID NO: 2). In some embodiments, the system includes an expression construct encoding the fusion protein, wherein the expression construct includes a WPRE sequence located in the 3' non-coding region and upstream of the polyadenylation site.
[0021] In some embodiments, the DNA binding domain is a dCas9 domain. In some embodiments, the dCas9 domain comprises SEQ ID NO: 9, or an amino acid sequence at least 90%, optionally at least 95% homologous thereto. In some embodiments, the epigenetic editing system further comprises one or more guide RNAs (gRNAs) or nucleic acid molecules encoding gRNAs. In some embodiments, the DNA binding domain is a zinc finger protein (ZFP) domain.
[0022] A method for modifying the epigenetic state of a target gene in a mammalian cell is described herein, comprising contacting the cell with an epigenetic editing system disclosed herein. A method for regulating the expression of a target gene in a mammalian cell is described herein, comprising contacting the cell with an epigenetic editing system disclosed herein. A method for treating a disease in a subject in need is described herein, comprising administering an epigenetic editing system disclosed herein to the subject.
[0023] From the following detailed description, additional aspects and advantages of the present disclosure will become apparent to those skilled in the art, wherein only illustrative embodiments of the present disclosure are shown and described. As will be appreciated, the present disclosure is capable of other and different embodiments, and its several details can be modified in various obvious aspects, all without departing from the present disclosure. Therefore, the drawings and description are to be regarded as illustrative in nature, and not restrictive.
[0024] Incorporation by reference
[0025] All publications, patents, and patent applications mentioned in this specification are incorporated herein by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. To the extent that publications and patents or patent applications incorporated by reference conflict with the disclosure contained in this specification, this specification is intended to supersede and / or take precedence over any such conflicting material. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] The novel features of the present invention are particularly set forth in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description and accompanying drawings (hereinafter referred to as "FIG." or "FIGs") which set forth illustrative embodiments in which the principles of the present invention are utilized, in which:
[0027] Figure 1A Schematic representation of fusion protein constructs with different NLS configurations is shown. Figure 1B Schematic representation of other fusion protein constructs with different KRAB domains is shown.
[0028] Figure 2A-2B The results show that the use of 6.25 ng RNA ( Figure 2A ) or 2.5ng RNA ( Figure 2B ) Percentage of PCSK9 protein levels measured after treatment with fusion protein constructs with different NLS positions in HeLa cells. Human and mouse DNMT3L sequences are denoted as h3L and m3L, respectively.
[0029] Figure 3 It was shown that in Hepa1-6 cells, the construct with 2X NLS was 3 times more potent than CRISPRoff in silencing mPcsk9.
[0030] Figure 4A It was shown that in Huh7 cells, the construct with 2X NLS was more efficient than CRISPRoff in silencing mPcsk9. Figure 4B-4C Shown in Huh7 cells, at different doses, on day 5 ( Figure 4B ) and Day 15 ( Figure 4C ), the construct with 2X NLS was also more efficient than CRISPR-Off in silencing mPcsk9.
[0031] Figure 5 It was shown that in Huh7 cells, in a CRISPR-off-like format where dCas9 was replaced by zinc fingers, 2XNLS provided improvements across multiple ZFs.
[0032] Figure 6A-6D It was shown that methylation of the CTLA4 promoter with bacterial DNMT proteins can induce epigenetic silencing of the locus.
[0033] Figure 7 Shown are the methylation profiles at the VIM3 locus for cells treated with different constructs carrying a bacterial DNA methyltransferase fused to dCas9 at day 30. Samples treated with M.SssI were methylated by 20%.
[0034] Figure 8Shown are methylation profiles at the CLTA locus in cells captured by hybridization at day 29, comparing murine DNMT3A / 3L in a fusion of M.SssI and dCas9.
[0035] Figure 9A-9D When 0.5 ng of effector DNA (using CLTA-GFP as a marker) ( Fig.9A ), using 3ng effector DNA (using GFP as a marker) ( Fig. 9B ) and 0.5 ng of effector DNA (using GFP as a marker) ( Fig. 9C ) Alternative KRAB domains tested for apparent silencing activity compared to CRISPRoff. Fig.9D Results after 30 days using different nanogram amounts of effector DNA are shown. DETAILED DESCRIPTION
[0036] Although various embodiments of the present disclosure have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Without departing from the present disclosure, those skilled in the art may conceive of many variations, changes and substitutions. It should be understood that various alternatives to the embodiments of the disclosure described herein may be employed.
[0037] The practice of the present invention will employ, unless otherwise indicated, conventional techniques of chemistry, biochemistry, molecular biology, microbiology and immunology, which are within the capabilities of a person of ordinary skill in the art and are explained in the literature. See, e.g., Sambrook, J., Fritsch, EF, and Maniatis, T. (1989) Molecular Cloning: A Laboratory Manual, 2nd ed., Cold Spring Harbor Laboratory Press; Ausubel, FM et al., (1995 and periodic supplements) Current Protocols in Molecular Biology, Ch. 9, 13, and 16, John Wiley & Sons; Roe, B., Crabtree, J., and Kahn, A. (1996) DNA Isolation and Sequencing: Essential Techniques, John Wiley & Sons; Polak, JM and McGee, J. O'D. (1990) In Situ Hybridization: Principles and Practice, Oxford University Press; Gait, MJ (1984) Oligonucleotide Synthesis: A Practical Approach, IRL Press; and Lilley, DM and Dahlberg, JE (1992) Methods in Enzymology: DNA Structures Part A: Synthesis and Physical Analysis of DNA, Academic Press. Each of these general texts is incorporated herein by reference in its entirety.
[0038] definition
[0039] Whenever the term "at least", "greater than", or "greater than or equal to" precedes the first value in a series of two or more values, the term "at least", "greater than", or "greater than or equal to" applies to each value in the series. For example, greater than or equal to 1, 2, or 3 is equivalent to greater than or equal to 1, greater than or equal to 2, or greater than or equal to 3.
[0040] Whenever the term "not greater than," "less than," or "less than or equal to" precedes the first value in a series of two or more values, the term "not greater than," "less than," or "less than or equal to" applies to every value in the series. For example, less than or equal to 3, 2, or 1 is equivalent to less than or equal to 3, less than or equal to 2, or less than or equal to 1.
[0041] The use of absolute or sequential terms, such as, "shall," "shall not," "should," "should not," "must," "must not," "first," "initially," "next," "subsequently," "before," "after," "last," and "finally," are not meant to limit the scope of the embodiments of the invention disclosed herein, but rather are exemplary.
[0042] As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. In addition, to the extent that the terms "including," "includes," "having," "has," "with," or variations thereof are used in the detailed description and / or claims, such terms are intended to be inclusive in a manner similar to the term "comprising."
[0043] As used herein, the terms "clinic," "clinical setting," "laboratory," or "laboratory setting" refer to a hospital, clinic, pharmacy, research facility, pathology laboratory, or other commercial setting in which trained personnel are employed to process and / or analyze biological and / or environmental samples. These terms are in contrast to point-of-care, remote areas, homes, schools, and other non-commercial, non-institutional settings.
[0044] The terms "determine," "measure," "evaluate," "assess," "measure," and "analyze" are generally used interchangeably herein to refer to forms of measurement. These terms include determining whether an element is present (e.g., detecting). These terms may include quantitative, qualitative, or both quantitative and qualitative determinations. Assessments are relative or absolute. "Detecting the presence of" may include determining the amount of something present in addition to determining whether something is present or not, depending on the context.
[0045] The terms "subject", "patient" or "individual" are often used interchangeably herein. "Subject" can be a biological entity containing expressed genetic material. The biological entity can be a plant, an animal or a microorganism, including, for example, bacteria, viruses, fungi and protozoa. The object can be tissues, cells and their offspring of biological entities obtained in vivo or cultured in vitro. The object can be a mammal. The mammal can be a human. The object may be diagnosed or suspected to be at high risk for the disease. In some cases, the object is not necessarily diagnosed or suspected to be at high risk for the disease. The object may or may not be exposed to the pathogen of interest as described herein, and may be due to symptoms or symptoms of a disease or condition associated with infection or exposure to a pathogen as described herein. In some embodiments, the object is suspected of having been exposed to a pathogen, such as a virus. In some embodiments, the object has been exposed to an antigen or representative protein, or cross-reacts with an antigen of a specific pathogen (such as a virus). In some embodiments, the object has one or more symptoms indicating a disease or condition associated with infection or exposure to a pathogen as described herein. In some embodiments, the object is currently infected by a pathogen (such as a virus as described herein). In some embodiments, the object was previously infected by a pathogen as described herein. In some embodiments, the subject is a carrier of a virus described herein. In some embodiments, the subject is a carrier of a fragment or remnant of a virus described herein. In some cases, the subject is a carrier of adaptive immunity resulting from a previous or current infection with a virus described herein. In some embodiments, the subject is a carrier of adaptive immunity resulting from a previous or current exposure to a different virus or pathogen other than the virus or pathogen of interest.
[0046] The term "subject" encompasses mammals. Examples of mammals include, but are not limited to, any member of the class of mammals: humans, non-human primates such as chimpanzees and other apes and monkey species; livestock such as cattle, horses, sheep, goats, pigs; domestic animals such as rabbits, dogs, and cats; laboratory animals, including rodents such as rats, mice, and guinea pigs, and the like.
[0047] The term "about" or "approximately" means within an acceptable error range for a particular value as determined by one of ordinary skill in the art, which depends in part on how the value is measured or determined, such as limitations of the measurement system. For example, depending on the practice of a given value, "about" can mean within 1 or more than 1 standard deviation. Where specific values are described in the application and claims, unless otherwise stated, the term "about" should be considered to mean an acceptable error range for that specific value.
[0048] As used herein, the phrases "at least one," "one or more," and "and / or" are open-ended expressions that are both conjunctive and disjunctive in operation. For example, each of the expressions "at least one of A, B, and C," "at least one of A, B, or C," "one or more of A, B, and C," "one or more of A, B, or C," and "A, B, and / or C" means A alone, B alone, C alone, A and B together, A and C together, B and C together, or A, B, and C together.
[0049] As used herein, the term "nucleic acid" refers to a polymer containing at least two nucleotides (i.e., deoxyribonucleotides or ribonucleotides) in single-stranded or double-stranded form, and includes DNA and RNA." Nucleotide" comprises sugar deoxyribose (DNA) or ribose (RNA), base and phosphate group. Nucleotides are linked together by phosphate groups. "Base" includes purine and pyrimidine, which further includes natural compounds adenine, thymine, guanine, cytosine, uracil, inosine and natural analogs, and synthetic derivatives of purine and pyrimidine, which include but are not limited to the modification of placing new reactive groups, such as but not limited to amines, alcohols, thiols, carboxylates and alkyl halides. Nucleic acids include nucleic acids containing known nucleotide analogs or modified main chain residues or bonds, which are synthetic, naturally occurring and non-naturally occurring, and have similar binding properties to reference nucleic acids. Examples of such analogs and / or modified residues include but are not limited to phosphorothioates, phosphoramidates, methylphosphonates, chiral methylphosphonates, 2'-O-methyl ribonucleotides and peptide nucleic acids (PNA).
[0050] The term "nucleic acid" includes any oligonucleotide or polynucleotide, wherein the fragment containing up to 60 nucleotides is generally referred to as an oligonucleotide, and longer fragments are referred to as polynucleotides. Deoxyribonucleotides are composed of 5 carbon sugars referred to as deoxyribose, which are covalently linked to phosphate at the 5' and 3' carbons of the sugar to form an alternating unbranched polymer. DNA can be, for example, an antisense molecule, plasmid DNA, pre-concentrated DNA, PCR product, vector, expression cassette, chimeric sequence, chromosomal DNA or the derivative and combination of these groups. Ribo-oligonucleotides are composed of similar repeating structures, wherein 5 carbon sugars are ribose. Therefore, the terms "polynucleotide" and "oligonucleotide" can refer to a polymer or oligomer of nucleotides or nucleoside monomers composed of (main chain) bonds between naturally occurring bases, sugars and sugars. The terms "polynucleotide" and "oligonucleotide" can also include polymers or oligomers containing non-naturally occurring monomers or their parts, which play a role similarly. Such modified or substituted oligonucleotides are often preferred over native forms because of properties such as enhanced cellular uptake, reduced immunogenicity, and increased stability in the presence of nucleases.
[0051] "Nucleic acid" as described herein may include one or more nucleotide variants, including non-standard nucleotides, non-natural nucleotides, nucleotide analogs and / or modified nucleotides. Examples of modified nucleotides include, but are not limited to, diaminopurine, 5-fluorouracil, 5-bromouracil, 5-chlorouracil, 5-iodouracil, hypoxanthine, xanthine, 4-acetylcytosine, 5-(carboxyhydroxymethyl)uracil, 5-carboxymethylaminomethyl-2-thiouridine, 5-carboxymethylaminomethyluracil, dihydrouracil, β-D-galactosylguanosine, inosine, N6-isopentenyladenine, 1-methylguanine, 1-methylinosine, 2,2-dimethylguanine, 2-methyladenine, 2-methylguanine, 3-methylcytosine, 5-methylcytosine, N6-adenine, 7-methylguanine, 5-methylaminomethyl uracil, 5-methoxyaminomethyl-2-thiouracil, β-D-mannosyl braided glycoside, 5'-methoxycarboxymethyl uracil, 5-methoxyuracil, 2-methylthio-N6-isopentenyl adenine, uracil-5-oxyacetic acid (v), wybutoxosine, pseudouracil, braided glycoside, 2-thiocytosine, 5-methyl-2-thiouracil, 2-thiouracil, 4-thiouracil, 5-methyluracil, uracil-5-oxyacetic acid methyl ester, 5-methyl-2-thiouracil, 3-(3-amino-3-N-2-carboxypropyl) uracil, (acp3) w, 2,6-diaminopurine, etc. In some cases, the nucleotide may include modifications in its phosphate portion, including modifications to the triphosphate portion. Non-limiting examples of such modifications include longer length phosphate chains (e.g., phosphate chains having 4, 5, 6, 7, 8, 9, 10 or more phosphate moieties) and modifications with thiol moieties (e.g., α-thiotriphosphates and β-thiotriphosphates).
[0052] Nucleic acid as described herein can be modified at base part (for example, at one or more atoms that can be used to form hydrogen bonds with complementary nucleotides and / or at one or more atoms that can not form hydrogen bonds with complementary nucleotides usually), sugar moiety or phosphate backbone.Main chain modification can include but is not limited to phosphorothioate, phosphorodithioate, selenophosphate, diselenose, thiophosphoramidate, phosphoraniladate, phosphoramidate and diaminophosphorothioate bond.Phosphorothioate bond replaces the non-bridging oxygen atom in the phosphate backbone with sulfur atom, and delays the nuclease degradation of oligonucleotide.Diaminophosphorothioate bond (N3'→P5') allows to prevent the identification and degradation of nuclease.Main chain modification can also be included in the main chain structure with peptide bond instead of phosphorus (for example, N-(2-aminoethyl)-glycine unit connected by peptide bond in peptide nucleic acid), or includes the linking group of carbamate, amide and linear and cyclic hydrocarbon group. Oligonucleotides with modified backbones are reviewed in Micklefield, Backbone modification of nucleic acids: synthesis, structure and therapeutic applications, Curr. Med. Chem., 8(10): 1157-79, 2001 and Lyer et al., Modified oligonucleotides-synthesis, properties and applications, Curr. Opin. Mol. Ther., 1(3): 344-358, 1999. The nucleic acid molecules described herein may comprise sugar moieties comprising ribose or deoxyribose as found in naturally occurring nucleotides, or modified sugar moieties or sugar analogs. The example of modified sugar moiety includes but is not limited to, 2'-O-methyl, 2'-O-methoxyethyl, 2'-O-aminoethyl, 2'-fluorine, N3'→P5' phosphoramidate, 2' dimethylaminooxyethoxy, 2'2' dimethylaminoethoxyethoxy, 2'-guanidinium, 2'-O-guanidinium ethyl, carbamate modified sugar and the sugar of bicyclic modification.2'-O-methyl or 2'-O-methoxyethyl modification promotes A-form or RNA sample conformation in oligonucleotide, increases the binding affinity to RNA, and has enhanced nuclease resistance.Modified sugar moiety can also include other bridge bonds (for example, connecting 2'-O and 4'-C atoms of ribose in locked nucleic acid) or sugar analogs such as morpholine ring (for example, as in diaminophosphorothioate morpholine).
[0053] Unless otherwise indicated, a specific nucleic acid sequence also implicitly encompasses conservatively modified variants thereof (e.g., degenerate codon substitutions), alleles, orthologs, SNPs, and complementary sequences, as well as sequences explicitly indicated. Specifically, degenerate codon substitutions can be achieved by generating sequences in which the third position of one or more selected (or all) codons is replaced by mixed bases and / or deoxyinosine residues (Batzer et al., Nucleic Acid Res., 19: 5081 (1991); Ohtsuka et al., J. Biol. Chem., 260: 2605-2608 (1985); Rossolini et al., Mol. Cell. Probes, 8: 91-98 (1994)).
[0054] The present disclosure encompasses isolated or substantially purified nucleic acid molecules and compositions comprising these molecules. As used herein, "isolated" or "purified" DNA molecules or RNA molecules are DNA molecules or RNA molecules that exist away from their natural environment. Isolated DNA molecules or RNA molecules can exist in purified form, or can exist in non-natural environments, such as transgenic host cells. For example, "isolated" or "purified" nucleic acid molecules or their biologically active parts, when produced by recombinant technology, are substantially free of other cell materials or culture media, or when synthesized by chemosynthesis, are substantially free of chemical precursors or other chemicals. In one embodiment, "isolated" nucleic acids do not contain sequences naturally located at the flanks of the nucleic acid in the genomic DNA of the organism from which the nucleic acid is derived (i.e., sequences located at the 5' and 3' ends of the nucleic acid). For example, in some embodiments, an isolated nucleic acid molecule can include a nucleotide sequence less than about 5kb, 4kb, 3kb, 2kb, 1kb, 0.5kb or 0.1kb, which is naturally located at the flanks of the nucleic acid molecule in the genomic DNA of the cell from which the nucleic acid is derived.
[0055] As used herein, the terms "protein", "polypeptide" and "peptide" are used interchangeably and refer to a polymer of amino acid residues connected via peptide bonds, and it can be composed of two or more polypeptide chains. The terms "polypeptide", "protein" and "peptide" refer to a polymer of at least two amino acid monomers connected together by an amide bond. Amino acids can be L-optical isomers or D-optical isomers. More specifically, the terms "polypeptide", "protein" and "peptide" refer to molecules composed of two or more amino acids in a specific order (e.g., an order determined by the base sequence of nucleotides in a gene or RNA encoding the protein). Proteins are essential for the structure, function and regulation of cells, tissues and organs of the body, and each protein has a unique function. Examples are hormones, enzymes, antibodies and any fragments thereof. In some cases, a protein can be a part of a protein, for example, a domain, a subdomain or a motif of a protein. In some cases, a protein may be a variant (or mutation) of a protein in which one or more amino acid residues are inserted into, deleted from, and / or substituted into a naturally occurring (or at least known) amino acid sequence of a protein. A polypeptide may be a single linear polymer chain of amino acids bound together by peptide bonds between carboxyl and amino groups of adjacent amino acid residues. A polypeptide may be modified, for example, by addition of carbohydrates, phosphorylation, etc. A protein may comprise one or more polypeptides.
[0056] The protein or its variant can be naturally occurring or recombinant. Methods for detecting and / or measuring polypeptides in biological materials are well known in the art and include, but are not limited to, Western blotting, flow cytometry, ELISA, RIA, and various proteomics techniques. An exemplary method for measuring or detecting polypeptides is an immunoassay such as ELISA. This type of protein quantification can be based on antibodies capable of capturing specific antigens and a second antibody capable of detecting the captured antigen. Exemplary assays for detecting and / or measuring polypeptides are described in Harlow, E. and Lane, D. Antibodies: A Laboratory Manual, (1988), Cold Spring Harbor Laboratory Press.
[0057] As used herein, the term "fragment" or equivalent terms may refer to a portion of a protein that is less than the full length of the protein and optionally retains the function of the protein. In addition, when a portion of a protein is blasted against a protein, the portion of the protein sequence may, for example, be aligned with a portion of the protein sequence at least 80% identity.
[0058] Any systems, methods, and platforms described herein are modular and are not limited to sequential steps. Therefore, terms such as "first" and "second" do not necessarily imply priority, order of importance, or order of actions.
[0059] The term "regulation" refers to a change in the quantity, degree or scope of a function. For example, the composition disclosed herein for epigenetic modification can regulate the activity of a promoter sequence by binding to a motif within a promoter, thereby inducing, enhancing or inhibiting the transcription of a gene operably connected to a promoter sequence. Alternatively, regulation may include inhibiting the transcription of a gene, wherein the epigenetic editing system binds to a structural gene and blocks the RNA polymerase reading gene that depends on DNA, thereby inhibiting the transcription of a gene. For example, a structural gene may be a normal cell gene or an oncogene. Alternatively, regulation may include inhibiting the translation of a transcript. Therefore, "regulation" of gene expression includes both gene activation and gene suppression.
[0060] The term "administering" and its grammatical equivalents as used herein can refer to providing one or more replication-competent recombinant adenoviruses or pharmaceutical compositions described herein to a subject or patient. As an example and not limitation, "administering" can be by intravenous (iv) injection, subcutaneous (sc) injection, intradermal (id) injection, intraperitoneal (ip) injection, intramuscular (im) injection, intravascular injection, infusion (inf.), oral route (po), topical (top.) administration or rectal (pr) administration. One or more such routes can be used. Parenteral administration can be, for example, by bolus injection or by gradual infusion over time.
[0061] As used herein, the terms "treat," "treating," or "treatment," and grammatical equivalents, may include alleviating, slowing down, or ameliorating at least one symptom of a disease or condition, preventing other symptoms, inhibiting a disease or condition, e.g., preventing and / or therapeutically arresting the development of a disease or condition, relieving a disease or condition, causing regression of a disease or condition, relieving a condition caused by a disease or condition, or stopping symptoms of a disease or condition. "Treatment" may refer to administering a vector, nucleic acid (e.g., mRNA), or LNP composition to a subject after the onset or suspected onset of a disease or condition. "Treatment" includes the concept of "relief," which refers to reducing the frequency or severity of the occurrence or recurrence of any symptom or other adverse reaction associated with a disease or condition and / or a side effect associated with a disease or condition. The term "treatment" also encompasses the concept of "management," which refers to reducing the severity of a particular disease or condition in a patient or delaying its recurrence, e.g., prolonging the period of remission in a patient suffering from the disease. The term "treatment" further encompasses the concepts of "prevent", "preventing" and "prevention". It should be understood that, although not excluded, the treatment of a disorder or condition does not require the complete elimination of the disorder, condition or symptoms associated therewith. The term "treatment" as used herein covers any treatment of a disease in a mammal, particularly a human, and includes: (a) preventing the disease from occurring in a subject who may be susceptible to the disease but has not yet been diagnosed with the disease; (b) inhibiting the disease, i.e., preventing its development; or (c) alleviating the disease, i.e., alleviating or ameliorating the disease and / or its symptoms or conditions. The term "prevention" is used herein to refer to one or more measures taken to prevent or partially prevent a disease or condition.
[0062] "Treating or preventing a condition" means ameliorating any condition or sign or symptom associated with a condition, either before or after the condition occurs. For example, alleviating the symptoms of a condition may involve a reduction or prevention of at least 3%, 5%, 10%, 20%, 40%, 50%, 60%, 80%, 90%, 95%, 98%, 99%, 99.5%, 99.9% or 100% as compared to an equivalent untreated control, as measured by any standard technique. In some embodiments, alleviating symptoms of a disorder can involve a reduction or prevention of at least 2-fold, at least 3-fold, at least 4-fold, at least 5-fold, at least 10-fold, at least 20-fold, at least 25-fold, at least 30-fold, at least 40-fold, at least 50-fold, at least 60-fold, at least 70-fold, at least 80-fold, at least 90-fold, at least 100-fold, at least 200-fold, at least 300-fold, at least 400-fold, at least 500-fold, at least 600-fold, at least 700-fold, at least 800-fold, at least 900-fold, at least 1000-fold, at least 2000-fold, at least 3000-fold, at least 4000-fold, at least 5000-fold, at least 6000-fold, at least 7000-fold, at least 8000-fold, at least 9000-fold, or at least 10000-fold, as compared to an equivalent untreated control.
[0063] As used herein, the term "pharmaceutical composition" and its grammatical equivalents may refer to a mixture or solution containing a therapeutically effective amount of an active pharmaceutical ingredient and one or more pharmaceutically acceptable excipients, carriers and / or therapeutic agents, which is to be administered to a subject (e.g., a human in need thereof).
[0064] As used herein, the term "pharmaceutically acceptable" and its grammatical equivalents may refer to the properties of materials that can be used to prepare pharmaceutical compositions that are generally safe, non-toxic, and neither biologically nor otherwise undesirable, and are acceptable for veterinary as well as human pharmaceutical uses. "Pharmaceutically acceptable" may refer to materials such as carriers or diluents that do not abrogate the biological activity or properties of the compound, and are relatively non-toxic, i.e., the material can be administered to a subject without causing undesirable biological effects or interacting in a deleterious manner with any component of the pharmaceutical composition in which it is contained.
[0065] A "pharmaceutically acceptable excipient, carrier or diluent" refers to an excipient, carrier or diluent that can be administered to a subject with a pharmaceutical agent and does not destroy its pharmacological activity and is non-toxic when administered in doses sufficient to deliver a therapeutic amount of the pharmaceutical agent.
[0066] "Pharmaceutically acceptable salts" may be acid salts or basic salts that are generally considered in the art to be suitable for use in contact with human or animal tissues without excessive toxicity, irritation, allergic response or other problems or complications. Such salts include inorganic and organic acid salts of basic residues such as amines, and alkali metal or organic salts of acidic residues such as carboxylic acids. Specific pharmaceutical salts include, but are not limited to, salts of acids such as hydrochloric acid, phosphoric acid, hydrobromic acid, malic acid, glycolic acid, fumaric acid, sulfuric acid, sulfamic acid, sulfanilic acid, formic acid, toluenesulfonic acid, methanesulfonic acid, benzenesulfonic acid, ethanedisulfonic acid, 2-hydroxyethylsulfonic acid, nitric acid, benzoic acid, 2-acetoxybenzoic acid, citric acid, tartaric acid, lactic acid, stearic acid, salicylic acid, glutamic acid, ascorbic acid, pamoic acid, succinic acid, fumaric acid, maleic acid, propionic acid, hydroxymaleic acid, hydroiodic acid, phenylacetic acid, alkanoic acids such as acetic acid, HOOC-(CH2)n-COOH, where n is 0-4, etc. Similarly, pharmaceutically acceptable cations include, but are not limited to, sodium, potassium, calcium, aluminum, lithium, and ammonium. Those of ordinary skill in the art will recognize from the present disclosure and the knowledge in the art that additional pharmaceutically acceptable salts include those listed by Remington's Pharmaceutical Sciences, 17th edition, Mack Publishing Company, Easton, PA, page 1418 (1985). Typically, pharmaceutically acceptable acid salts or basic salts can be synthesized from parent compounds containing a basic or acidic moiety by any conventional chemical method. In short, such salts can be prepared by reacting the free acid or base form of these compounds with a stoichiometric amount of a suitable base or acid in a suitable solvent.
[0067] As used herein, the term "therapeutically effective amount" means an amount of an agent to be delivered (e.g., a nucleic acid, a drug, a payload, a composition, a therapeutic agent, a diagnostic agent, a prophylactic agent, etc.) that is sufficient to treat, ameliorate symptoms of, diagnose, prevent and / or delay the onset of an infection, disease, disorder and / or condition when administered to a subject suffering from or susceptible to an infection, disease, disorder and / or condition.
[0068] The term "repressor domain" or "transcriptional repressor domain" refers to a transcriptional repressor protein or a portion thereof, such as a transcription factor, which can be complexed with one or more DNA binding domains to act as a negative regulatory domain. The repressor domain blocks the recruitment of RNA polymerase to inhibit the transcription of certain genes. The repressor domain inhibits transcriptional activation by interacting with other cellular components (such as, but not limited to, basal transcription factors, effector molecules, activators or co-activator proteins, repressors, and co-repressors), thereby allowing precise control of gene expression.
[0069] The term "KRAB" refers to a Krüppel-associated box, a transcriptional repressor protein domain. KRAB refers to homologs, orthologs and mutants of the KRAB domain, which have a conserved or enhanced basic function of inhibiting gene transcription. The KRAB domain is one of a group of transcriptional repressor domains present in approximately 400 human zinc finger protein-based transcription factors. The KRAB domain typically includes about 45 to about 75 amino acid residues. A description of the KRAB domain (including its function and use) can be found in, for example, Ecco, G., Imbeault, M., Trono, D., KRAB zinc finger proteins, Development 144, (15): 2719-2017.
[0070] The term "DNMT" refers to DNA methyltransferase. As used herein, the term encompasses enzymes that catalyze methyl transfer to DNA, such as typical cytosine-5 DNMTs (e.g., DNMT1, DNMT3A, DNMT3B, and DNMT3C) that catalyze methyl addition to genomic DNA. The term also encompasses atypical family members that do not catalyze methylation themselves, but recruit (including activation) catalytically active DNMTs, and non-limiting examples of such DNA methyltransferases include DNMT3L. See, e.g., Lyko, Nat Review. (2018) 19: 81-92. Unless otherwise indicated, a DNMT domain may refer to a polypeptide domain derived from a catalytically active DNMT (e.g., DNMT1, DNMT3A, and DNMT3B) or from a catalytically inactive DNMT (e.g., DNMT3L).
[0071] The term "DNA binding domain" refers to a DNA binding domain from a protein selected from CRISPR proteins, TAL proteins, zinc fingers and other transcriptional regulator families, their homologs, orthologs and mutants, which maintains or enhances the basic function of the DNA binding protein.
[0072] Scope provided herein should be understood as abbreviation of all values in the scope. For example, the scope of 1 to 50 is understood to include any number, combination of numbers or subranges and all intermediate decimal values between the aforementioned integers from 1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23,24,25,26,27,28,29,30,31,32,33,34,35,36,37,38,39,40,41,42,43,44,45,46,47,48,49 or 50, for example 1.1,1.2,1.3,1.4,1.5,1.6,1.7,1.8 and 1.9. About subrange, " nested subrange " extending from any end point of scope is specially considered. For example, nested sub-ranges of the exemplary range of 1 to 50 may include 1 to 10, 1 to 20, 1 to 30, and 1 to 40 in one direction, or 50 to 40, 50 to 30, 50 to 20, and 50 to 10 in the other direction.
[0073] The term "therapeutic agent" may refer to any agent that has a therapeutic, diagnostic and / or prophylactic effect and / or induces a desired biological and / or pharmacological effect when administered to a subject. Therapeutic agents may also be referred to as "actives" or "active agents". Such agents include, but are not limited to, cytotoxins, radioactive ions, chemotherapeutic agents, small molecule drugs, proteins, and nucleic acids.
[0074] The term "ameliorate" as used herein may mean to reduce, inhibit, attenuate, diminish, arrest or stabilize the development or progression of a disease.
[0075] As used herein, "delaying" the development of a disease means postponing, hindering, slowing down, slowing down, stabilizing and / or delaying the progress of a disease. Depending on the history of the disease and / or the individual being treated, this delay can have different time lengths. A method of "delaying" or alleviating the development of a disease or delaying the onset of a disease is a method that, when compared to not using the method, uses the method to reduce the likelihood of developing one or more symptoms of the disease within a given time frame and / or reduce the degree of symptoms within a given time frame. Such comparisons are typically based on clinical studies using a sufficient number of subjects to give statistically significant results.
[0076] "Development" or "progression" of a disease means the initial manifestation and / or subsequent progression of the disease. The development of a disease can be detected and assessed using standard clinical techniques as are well known in the art. However, development also refers to progression that may not be detectable. For purposes of this disclosure, development or progression refers to the biological course of symptoms. "Development" includes onset, recurrence, and onset.
[0077] As used herein, "onset" or "occurrence" of a disease includes initial onset and / or recurrence. Depending on the type of disease to be treated or the site of the disease, the isolated polypeptide or pharmaceutical composition can be administered to the subject using conventional methods known to those of ordinary skill in the medical arts. The composition can also be administered via other conventional routes, such as orally, parenterally, by inhalation spray, topically, rectally, nasally, buccally, vaginally, or via an implanted reservoir.
[0078] The term "parenteral" as used herein includes subcutaneous, intracutaneous, intravenous, intramuscular, intraarticular, intraarterial, intrasynovial, intraspinal, intrathecal, intralesional and intracranial injection or infusion techniques. In addition, it can be administered to a subject via an injectable, long-acting route of administration, such as using 1 month, 3 months or 6 months of long-acting injectable or biodegradable materials and methods.
[0079] It should be understood that, in addition to the specific proteins and nucleotides mentioned herein, the present invention also contemplates the use of variants, derivatives, homologues and fragments thereof. As used herein, the variant of any given sequence is a sequence in which a specific sequence of residues (whether amino acid residues or nucleic acid residues) has been modified in such a way that the polypeptide or polynucleotide in question substantially retains at least one of its endogenous functions. Variant sequences can be obtained by adding, missing, replacing, modifying, replacing and / or changing at least one residue present in a naturally occurring protein. As used herein, derivatives of any given sequence as expected include any replacement, variation, modification, replacement, missing and / or addition from or to one (or more) amino acid residues of the sequence, provided that the resulting protein or polypeptide substantially retains at least one of its endogenous functions. Amino acid substitutions can be performed, such as 1, 2 or 3 to 10 or 20 substitutions, provided that the modified sequence substantially retains the desired activity or ability. Amino acid substitutions can include the use of non-naturally occurring analogs. The protein used in the present disclosure can also have a deletion, insertion or substitution of amino acid residues that do not affect the function of the protein and result in a functionally equivalent protein. Deliberate amino acid substitutions may be made based on similarity in the polarity, charge, solubility, hydrophobicity, hydrophilicity, and / or amphipathic nature of the residues, as long as the intrinsic function is retained. For example, negatively charged amino acids include aspartic acid and glutamic acid; positively charged amino acids include lysine and arginine; and amino acids with uncharged polar head groups with similar hydrophilicity values include asparagine, glutamine, serine, threonine, and tyrosine.
[0080] As used herein, homologs of any protein or nucleic acid sequence contemplated herein include sequences having certain homology to wild-type amino acids and nucleic acid sequences. Homologous sequences may include sequences, such as amino acid sequences, that may be at least 50%, 55%, 65%, 75%, 85% or 90% identical to the subject sequence. In specific embodiments, homologous sequences may include amino acid sequences that are at least 95% or 97% or 99% identical to the subject sequence.
[0081] Sequence identity can be measured using sequence analysis software (e.g., Sequence Analysis Software Package of the Genetics Computer Group at the University of Wisconsin Biotechnology Center, 1710 University Avenue, Madison, Wis. 53705, BLAST, BESTFIT, GAP or PILEUP / PRETTYBOX programs). Such software matches identical or similar sequences by specifying degrees of homology for various substitutions, deletions and / or other modifications. Conservative substitutions typically include substitutions in the following groups: glycine, alanine; valine, isoleucine, leucine; aspartic acid, glutamic acid, asparagine, glutamine; serine, threonine; lysine, arginine; and phenylalanine, tyrosine. In an exemplary method for determining the degree of identity, the BLAST program can be used, with probability scores between e-3 and e-100 representing closely related sequences.
[0082] It will be appreciated that the numbering of specific positions or residues in the corresponding sequence depends on the specific protein and numbering scheme used. The numbering may be different, for example, in the precursor of the mature protein and the mature protein itself, and sequence differences between species may affect the numbering. Those skilled in the art will be able to identify the corresponding residues in any homologous protein and the corresponding encoding nucleic acid by methods well known in the art, such as by sequence alignment and determination of homologous residues.
[0083] Epigenetic editing system
[0084] Epigenetic editing systems for epigenetic modification and target gene expression regulation are described herein. Epigenetic editing systems for gene suppression may include fusion proteins or nucleic acids encoding fusion proteins. In some embodiments, the fusion protein comprises a DNA binding domain. In some embodiments, the fusion protein comprises a DNA methyltransferase (DNMT) domain. In some embodiments, the fusion protein comprises a nuclear localization sequence (NLS). In some embodiments, the fusion protein comprises a repressor domain.
[0085] As used herein, the epigenetic editing system can be any medium that binds to the target polynucleotide and has epigenetic regulatory activity. In some embodiments, the epigenetic editing system uses a DNA binding domain to bind polynucleotides at a specific sequence. In some embodiments, the epigenetic editing system uses a nucleic acid-guided DNA binding protein to bind polynucleotides at a specific sequence. In some embodiments, the epigenetic editing system includes an effector domain that can regulate the epigenetic state of the nucleic acid sequence at the target polynucleotide or adjacent to the target polynucleotide. In some embodiments, the epigenetic editing system can deposit epigenetic editing markers on chromatin regions, nucleic acid sequences, or histone amino acid residues at the target polynucleotide or near the target polynucleotide. For example, the epigenetic editing system can be methylated, demethylated, acetylated, deacetylated, ubiquitinated, or deubiquitinated at the target polynucleotide or near the target polynucleotide to chromatin regions, nucleic acid sequences, or histone amino acid residues. In some embodiments, the epigenetic editing system is capable of recruiting one or more proteins or complexes involved in transcriptional regulation (e.g., transcription factors, transcription activators, transcription repressors, or insulators) to chromatin regions, nucleic acid sequences, or histone amino acid residues at or near a target polynucleotide.
[0086] The epigenetic editing system provided herein may include one or more effector domains as described. In some embodiments, the epigenetic editing system includes multiple effector domains. In some embodiments, the epigenetic editing system includes one effector domain. In some embodiments, the epigenetic editing system includes at least 2, 3, 4, 5, 6, 7, 8, 9, 10 or more effector domains.
[0087] In some embodiments, the epigenetic editing system includes a DNA methylation domain, a repression domain, and a nucleic acid binding domain. In some embodiments, the epigenetic editing system includes a DNA methylation domain and a nucleic acid binding domain. In some embodiments, the nucleic acid binding domain is located at the C-terminus of the fusion protein. In some embodiments, the nucleic acid binding domain is located at the N-terminus of the fusion protein. In some embodiments, the nucleic acid binding domain is located in the middle of the fusion protein. In some embodiments, the DNA methylation domain is located at the C-terminus of the fusion protein. In some embodiments, the DNA methylation domain is located at the N-terminus of the fusion protein. In some embodiments, the DNA methylation domain is located in the middle of the fusion protein. In some embodiments, the repression domain is located at the C-terminus of the fusion protein. In some embodiments, the repression domain is located at the N-terminus of the fusion protein. In some embodiments, the repression domain is located in the middle of the fusion protein. In some embodiments, the epigenetic editing system includes a DNA demethylation domain and a histone acetylation domain. In some embodiments, the epigenetic editing system includes a DNA demethylation domain and an activation domain that recruits other DNA demethylation or histone acetylation proteins. In some embodiments, the epigenetic editing system includes a DNA demethylation domain, a histone acetylation domain, and a scaffold protein that recruits other DNA demethylation or histone acetylation proteins. In some embodiments, the epigenetic editing system includes two or more DNA demethylation domains, two or more histone acetylation domains, and / or two or more scaffold proteins that recruit other DNA demethylation or histone deacetylation proteins.
[0088] A variety of fusion proteins or constructs can be used to achieve activation or inhibition of one or more target genes. For example, an epigenetic editing system includes a fusion protein comprising a DNA binding domain (e.g., a dCas9 domain) and a methylation domain, and another fusion protein comprising a DNA binding domain and a repressor domain can be co-delivered with two or more guide RNAs, each of which targets a different target DNA sequence. The two or more target DNA sequences can be located in the same target gene or can be located in different target genes.
[0089] Effector domain
[0090] The epigenetic editing system provided herein can include one or more effector protein domains for regulating the expression of target genes.The effector domain can be used to contact the target polynucleotide sequence in the target gene to achieve epigenetic modification, for example, the change of the methylation state of DNA nucleotides in the target gene.Therefore, the epigenetic editing system with one or more effector domains can provide the effect of regulating the expression of the target gene without changing the DNA sequence of the target gene.For example, in some embodiments, the effector domain causes the suppression or silencing of the expression of the target gene.In some embodiments, the effector domain causes the activation or increased expression of the target gene.
[0091] Epigenetic effectors can deposit chemical modifications on the chromatin at the position of the target gene. Non-limiting examples of chemical modifications include methylation, demethylation, acetylation, deacetylation, phosphorylation, ubiquitination and / or ubiquitination of DNA or histone residues of chromatin. In some embodiments, epigenetic effectors can modify histone tails. In some embodiments, epigenetic effectors can add or remove active markers on histone tails. In some embodiments, active markers can include H3K4 methylation, H3K9 acetylation, H3K27 acetylation, H3K36 methylation, H3K79 methylation, H4K5 acetylation, H4K8 acetylation, H4K12 acetylation, H4K16 acetylation and / or H4K20 methylation. In some embodiments, epigenetic effectors can add or remove inhibitory markers on histone tails. In some embodiments, these inhibitory markers can include H3K9 methylation and / or H3K27 methylation.
[0092] In some embodiments, the effector domain in the epigenetic editing system changes the chemical modification state of the target gene containing the target sequence.For example, the effector domain can change the chemical modification state of the nucleotides in the target gene.In some embodiments, the effector domain of the epigenetic editing system deposits chemical modifications on the nucleotides in the target gene.In some embodiments, the effector domain of the epigenetic editing system deposits chemical modifications of histones associated with the target gene.In some embodiments, the effector domain of the epigenetic editing system removes the chemical modifications at the nucleotides in the target gene.In some embodiments, the effector domain of the epigenetic editing system removes the chemical modifications of histones associated with the target gene.In some embodiments, chemical modifications increase the expression of the target gene.For example, the epigenetic editing system can include an effector domain with histone acetyltransferase activity.In some embodiments, chemical modifications reduce the expression of the target gene.For example, the epigenetic editing system can include an effector domain with DNA methyltransferase activity.
[0093] The epigenetic modification mediated by the epigenetic editing system can be located near the target gene, or can be distant from the target gene, or extend from an initial epigenetic modification initiated by the epigenetic editing system at one or more nucleotides in the target sequence of the target gene.
[0094] In some embodiments, the change of chemical modification state is DNA methylation state.For example, methylation can be introduced by an effector domain with DNA methyltransferase activity, or can be removed by an effector domain with DNA demethylase activity.In some embodiments, the change of chemical modification (such as methylation) is at a low methylated nucleic acid sequence.For example, compared with a standard control, the sequence of chemical modification in a target gene or chromosome region may lack a methyl group on a 5-methylcytosine nucleotide (for example, in CpG).With respect to younger cells or non-cancerous cells, low methylation may occur in, for example, senescent cells or cancer (for example, early stage of tumor formation) respectively.In some embodiments, the target polynucleotide sequence is in a CpG island.In some embodiments, it is known that the target gene is associated with a disease or condition.In some embodiments, the target gene comprises a specific copy of a disease-related sequence.In some embodiments, the target gene comprises a target sequence related to a disease.In some embodiments, the change of chemical modification (such as methylation) is at a high methylated nucleic acid sequence.In some embodiments, the chemical modification is in a CpG island.
[0095] In some embodiments, the protein fusion construct can have 1 effector domain, 2 effector domains, 3 effector domains, 4 effector domains, 5 effector domains, 6 effector domains, 7 effector domains, 8 effector domains, 9 effector domains, or 10 effector domains.
[0096] Methyltransferase domain
[0097] In some embodiments, the effector domain comprises a histone methyltransferase domain. For example, repression (or silencing) can be caused by repressive chromatin markers on the chromatin containing the target nucleic acid sequence, methylation of DNA, methylation of histone residues (e.g., H3K9, H3K27) or deacetylation of histone residues. It is not intended to be bound by any theory, and the method can be used to, for example, change epigenetic state by closing chromatin via methylation or introducing repressive chromatin markers on the chromatin containing the target nucleic acid sequence (e.g., gene).
[0098] Specific epigenetic imprints guide gene transcription or gene silencing. For example, DNA methylation, histone modification, repressor proteins bound to silenced regions, and other transcriptional activities change gene expression without changing the underlying DNA sequence. Therefore, transcriptional regulation allows specific genes to be expressed in a specific manner while suppressing other genes. In some cases, cell fate or function can be controlled for initial differentiation (e.g., during organism development) or reprogramming cells or cell types (e.g., during diseases such as cancer, chronic inflammation, autoimmune diseases, diseases associated with various microbial groups of organisms, etc.). Histone modification plays a structural and biochemical role in gene transcription, and one approach is to form or destroy nucleosome structures that bind to histones and prevent gene transcription. Histones are alkaline proteins commonly found in the nuclei of eukaryotic cells, ranging from multicellular organisms including humans to unicellular organisms represented by fungi (molds and yeasts), and are ionically bound to genomic DNA. Histones are usually composed of five components (H1, H2A, H2B, H3, and H4) and are highly similar in biological species. For example, in the case of histone H4, budding yeast histone H4 (full-length 102 amino acid sequence) and human histone H4 (full-length 102 amino acid sequence) are identical in 92% of the amino acid sequence and differ in only 8 residues. Among the natural proteins believed to exist in tens of thousands of organisms, histones are known to be the most highly conserved proteins in eukaryotic species. Genomic DNA is folded together with histones by orderly binding, and the complex of the two forms a basic structural unit called a nucleosome. In addition, the aggregation of nucleosomes forms the chromosome chromatin structure. Histones undergo modifications at their N-termini called histone tails, such as acetylation, methylation, phosphorylation, ubiquitination, ubiquitination-like, etc., and maintain or specifically transform the chromatin structure, thereby controlling responses occurring on chromosomal DNA, such as gene expression. Expression, DNA replication, DNA repair, etc. The post-translational modification of histone is a kind of epigenetic regulatory mechanism, and is considered to be necessary for eukaryotic cell genetic regulation. Recent studies have shown that chromatin remodeling factors such as SWI / SNF, RSC, NURF, NRD, etc. (which promote DNA to approach transcription factors by modifying nucleosome structure), histone acetyltransferase (HAT) and histone deacetylase (HDAC) that regulate the acetylation state of histone act as important regulatory factors. DNA methylation mainly occurs in CpG sites (abbreviated as "C-phosphoric acid-G-" or "cytosine-phosphoric acid-guanine"). The transcriptional activity of the highly methylated regions of DNA is often lower than that of the sites with lower methylation. Many mammalian genes have promoter regions (regions with high frequency of CpG sites) close to or including CpG islands.
[0099] Specifically, the unstructured N-terminus of histone can be modified by at least one of acetylation, methylation, ubiquitination, phosphorylation, ubiquitination-like (sumoylation), ribosylation, citrullination, O-GlcNAc glycosylation or crotonylation (crotonylation). For example, the acetylation of K14 and K9 lysine of histone H3 by histone acetyltransferase may be related to the transcriptional ability of mankind. Lysine acetylation can directly or indirectly produce the binding site of the chromatin modifying enzyme regulating transcriptional activation. For example, histone acetyltransferase (HAT) utilizes acetyl-CoA as cofactor, and catalyzes the transfer of acetyl group to the epsilon amino group of lysine side chain. This neutralizes the positive charge of lysine, and weakens the interaction between histone and DNA, thereby opening chromosome for transcription factor binding and initiation of transcription. Similarly, the histone methylation of lysine 9 of histone H3 may be related to the chromatin of heterochromatin or transcriptional silence. A specific DNA methylation pattern can be established and modified by at least one or more, two or more, three or more, four or more, or five or more independent DNA methyltransferases, including DNMT1, DNMT3A, and DNMT3B.
[0100] In some embodiments, the effector domain comprises a histone methyltransferase domain. In some embodiments, the effector domain comprises a DOT1L domain, a SET domain, a SUV39H1 domain, a G9a / EHMT2 protein domain, an EZH1 domain, an EZH2 domain, a SETDB1 domain, or any combination thereof. In some embodiments, the effector domain comprises a histone-lysine-N-methyltransferase SETDB1 domain.
[0101] In some embodiments, the effector domain comprises a DNA methyltransferase domain or a histone methyltransferase domain. The DNA methyltransferase domain can mediate methylation at a DNA nucleotide (e.g., any one of an A, T, G, or C nucleotide). In some embodiments, the methylated nucleotide is N6-methyladenosine (m6A). In some embodiments, the methylated nucleotide is 5-methylcytosine (5MC). In some embodiments, methylation is at a CG (or CpG) dinucleotide sequence. In some embodiments, methylation is at a CHG or CHH sequence, where H is any one of A, T, or C.
[0102] In some embodiments, the effector domain comprises a DNA methyltransferase DNMT domain, which catalyzes the transfer of methyl groups to cytosine, thereby inhibiting the expression of target genes by recruiting repressive regulatory proteins. In some embodiments, the effector domain comprises a DNA methyltransferase (DNMT) family protein domain. In some embodiments, the effector domain comprises a DNMT1 domain. In some embodiments, the effector domain comprises a TRDMT1 domain. In some embodiments, the effector domain comprises a DNMT3 domain. In some embodiments, the effector domain comprises a DNMT3A domain. In some embodiments, the effector domain includes a DNMT3B domain. In some embodiments, the effector domain comprises a DNMT3C domain. In some embodiments, the effector domain comprises a DNMT3L domain. In some embodiments, the effector domain comprises a fusion of a DNMT3A-DNMT3L domain.
[0103] Table 1 below provides exemplary DNA methyltransferases (DNMTs) that can be part of an epigenetic effector domain. In some embodiments, the epigenetic editing system herein contains one or more epigenetic effector domains selected from Table 1, or functional homologs, orthologs, or variants thereof.
[0104] Table 1. Exemplary DNA methyltransferase sequences that can be used in epigenetic effector domains
[0105]
[0106] In some embodiments, the methyltransferase may be a mammalian methyltransferase. In some embodiments, the methyltransferase may be a plant methyltransferase. In some embodiments, the methyltransferase may be a fungal methyltransferase. In some embodiments, the methyltransferase may be a bacterial methyltransferase.
[0107] Bacterial DNA methyltransferases can be obtained from bacterial species. The bacterial species can be cocci. The bacterial species can be bacillus. The bacterial species can be spirochete. The bacteria can be intracellular bacteria, gram-positive bacteria or gram-negative bacteria. Examples of bacterial genera from which suitable DNA methyltransferases can be obtained include, but are not limited to: Acetobacter, Acinetobacter, Actinomyces, Agrobacterium, Anaplasma, Azorhizobium, Azotobacter, Bacillus, Viridans, Bacteroides, Bartonella tonella, Bordetella, Borrelia, Brucella, Brukholderia, Calymmatobacterium, Campylobacter, Chlamydia, Chlamydophila, Clostridium, Corynebacterium, Coxiella ), Ehrlichia, Eikenella, Enterobacter, Enterococcus, Escherichia, Fusobacterium, Gardnerella, Haemophilus, Helicobacter, Klebsiella, Lactobacillus, Legionella onella), Leptospira, Listeria, Methanobacterium, Microbacterium, Micrococcus, Moraxella, Mycobacterium, Mycoplasma, Mycoplasmatales, Neisseria, Pasteurella,Peptostreptococcus, Porphyromonas, Prevotella, Pseudomonas, Rhizobium, Rickettsia, Rochalimaea, Rothia, Salmonella, Shigella, Spirillum, Spiroplasma, Staphylococcus, Stenotrophomonas, Streptococcus In some embodiments, the DNMT or DNMT domain is used as an epigenetic effector sequence in the fusion protein provided herein, and the fusion protein is derived from a bacterial species, wherein the bacterial species is a mycoplasma order bacterium, a marine mycoplasma, or a Chinese spiroplasma. In some embodiments, the bacterial species from which the DNMT is derived is not a penetrant mycoplasma, a monosomal spirochete, a parainfluenzae, a luteobacterium, a haemophilus aegypti, a haemophilus haemolyticus, a moraxella, an escherichia coli, a water-dwelling thermophilus, a crescentus stalked bacillus, or a Clostridium difficile.
[0108] In some embodiments, the DNMT may be an animal DNMT. In some embodiments, the DNMT may be a mammalian DNMT. In some embodiments, the DNMT may be a primate DNMT. In some embodiments, the DNMT may be a human DNMT. In some embodiments, the DNMT may be a chimpanzee DNMT. In some embodiments, the DNMT may be a Philippine tarsier DNMT. In some embodiments, the DNMT may be a rodent DNMT. In some embodiments, the DNMT may be a mouse DNMT. In some embodiments, the DNMT may be a Castorella mouse DNMT. In some embodiments, the DNMT may be a Mus musculus DNMT. In some embodiments, the DNMT may be a Carolina squirrel DNMT. In some embodiments, the DNMT may be a long-clawed gerbil DNMT. In some embodiments, the DNMT may be a horse DNMT. In some embodiments, the DNMT may be a Przewalski's horse DNMT. In some embodiments, the DNMT may be a bovine DNMT. In some embodiments, the DNMT may be a bison DNMT. In some embodiments, the DNMT may be a North American pika DNMT. In some embodiments, the DNMT may be a feline DNMT. In some embodiments, the DNMT may be a canine DNMT. In some embodiments, the DNMT may be a bear DNMT. In some embodiments, the DNMT may be a giant panda DNMT.
[0109] In some embodiments, the DNMT may include SEQ ID NO: 12. In some embodiments, the DNMT may be SEQ ID NO: 12. In some embodiments, the DNMT may be at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% similar to SEQ ID NO: 12.
[0110] In some embodiments, the DNMT may include SEQ ID NO: 13. In some embodiments, the DNMT may be SEQ ID NO: 13. In some embodiments, the DNMT may be at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% similar to SEQ ID NO: 13.
[0111] In some embodiments, the DNMT may include SEQ ID NO: 14. In some embodiments, the DNMT may be SEQ ID NO: 14. In some embodiments, the DNMT may be at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% similar to SEQ ID NO: 14.
[0112] In some embodiments, the DNMT may include SEQ ID NO: 15. In some embodiments, the DNMT may be SEQ ID NO: 15. In some embodiments, the DNMT may be at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% similar to SEQ ID NO: 15.
[0113] In some embodiments, the DNMT may include SEQ ID NO: 16. In some embodiments, the DNMT may be SEQ ID NO: 16. In some embodiments, the DNMT may be at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% similar to SEQ ID NO: 16.
[0114] In some embodiments, the DNMT may include SEQ ID NO: 72. In some embodiments, the DNMT may be SEQ ID NO: 72. In some embodiments, the DNMT may be at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% similar to SEQ ID NO: 72.
[0115] In some embodiments, the DNMT may include SEQ ID NO: 73. In some embodiments, the DNMT may be SEQ ID NO: 73. In some embodiments, the DNMT may be at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% similar to SEQ ID NO: 73.
[0116] In some embodiments, the DNMT may include SEQ ID NO: 74. In some embodiments, the DNMT may be SEQ ID NO: 74. In some embodiments, the DNMT may be at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% similar to SEQ ID NO: 74.
[0117] In some embodiments, the DNMT may include SEQ ID NO: 75. In some embodiments, the DNMT may be SEQ ID NO: 75. In some embodiments, the DNMT may be at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% similar to SEQ ID NO: 75.
[0118] In some embodiments, the DNMT may include SEQ ID NO: 76. In some embodiments, the DNMT may be SEQ ID NO: 76. In some embodiments, the DNMT may be at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% similar to SEQ ID NO: 76.
[0119] In some embodiments, the DNMT may include SEQ ID NO: 77. In some embodiments, the DNMT may be SEQ ID NO: 77. In some embodiments, the DNMT may be at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% similar to SEQ ID NO: 77.
[0120] In some embodiments, the DNMT may include SEQ ID NO: 78. In some embodiments, the DNMT may be SEQ ID NO: 78. In some embodiments, the DNMT may be at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% similar to SEQ ID NO: 78.
[0121] In some embodiments, the DNMT may include SEQ ID NO: 79. In some embodiments, the DNMT may be SEQ ID NO: 79. In some embodiments, the DNMT may be at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% similar to SEQ ID NO: 79.
[0122] In some embodiments, the DNMT may include SEQ ID NO: 80. In some embodiments, the DNMT may be SEQ ID NO: 80. In some embodiments, the DNMT may be at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% similar to SEQ ID NO: 80.
[0123] In some embodiments, the DNMT may include SEQ ID NO: 101. In some embodiments, the DNMT may be SEQ ID NO: 101. In some embodiments, the DNMT may have at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% similarity to SEQ ID NO: 101.
[0124] In some embodiments, the DNMT may include SEQ ID NO: 102. In some embodiments, the DNMT may be SEQ ID NO: 102. In some embodiments, the DNMT may be at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% similar to SEQ ID NO: 102.
[0125] In some embodiments, the DNMT may include SEQ ID NO: 103. In some embodiments, the DNMT may be SEQ ID NO: 103. In some embodiments, the DNMT may be at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% similar to SEQ ID NO: 103.
[0126] In some embodiments, the DNMT may include SEQ ID NO: 104. In some embodiments, the DNMT may be SEQ ID NO: 104. In some embodiments, the DNMT may be at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% similar to SEQ ID NO: 104.
[0127] In some embodiments, the DNMT may include SEQ ID NO: 105. In some embodiments, the DNMT may be SEQ ID NO: 105. In some embodiments, the DNMT may be at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% similar to SEQ ID NO: 105.
[0128] In some embodiments, the DNMT may include SEQ ID NO: 106. In some embodiments, the DNMT may be SEQ ID NO: 106. In some embodiments, the DNMT may be at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% similar to SEQ ID NO: 106.
[0129] In some embodiments, the DNMT may include SEQ ID NO: 107. In some embodiments, the DNMT may be SEQ ID NO: 107. In some embodiments, the DNMT may be at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% similar to SEQ ID NO: 107.
[0130] In some embodiments, the DNMT may include SEQ ID NO: 108. In some embodiments, the DNMT may be SEQ ID NO: 108. In some embodiments, the DNMT may be at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% similar to SEQ ID NO: 108.
[0131] In some embodiments, the DNMT may include SEQ ID NO: 109. In some embodiments, the DNMT may be SEQ ID NO: 109. In some embodiments, the DNMT may be at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% similar to SEQ ID NO: 109.
[0132] In some embodiments, a DNMT (e.g., DNMT3) can comprise an ADD domain. The ADD domain can have a sequence of SEQ ID NO: 98. The ADD domain can inhibit DNMT activity in a specific genomic region (e.g., a region with a specific histone feature).
[0133] Repressor domain
[0134] In some embodiments, the effector domain recruits one or more protein domains that inhibit the expression of a target gene. In some embodiments, the effector domain interacts with a scaffold protein domain that recruits one or more protein domains that inhibit the expression of a target gene. For example, the effector domain can recruit or interact with a scaffold protein domain that recruits a PRMT protein, an HDAC protein, a SETDB1 protein, or a NuRD protein domain. In some embodiments, the effector domain comprises a Krüppel-associated box (KRAB) repressor domain; a repressor element silencing transcription factor (REST) repressor domain, a KRAB-associated protein 1 (KAP1) domain, a MAD domain, a FKHR (forkhead in rhabdomyosarcoma gene) repressor domain, an aEGR-1 (early growth response gene product-1) repressor domain, an ets2 repressor factor repressor domain (ERD), a MAD smSIN3 interaction domain (SID), a WRPW motif of a hair-related basic helix-loop-helix (bHLH) repressor protein; an HP1 alpha chromo-shadow repressor domain (HP1 alpha chromo-shadow repressor domain). In some embodiments, the effector domain comprises a KRAB domain. In some embodiments, the effector domain comprises a tripartite motif containing protein 28 (TRIM28, TIF1-β, or KAP1).
[0135] In some embodiments, the effector domain includes a protein domain that inhibits the expression of a target gene, also referred to herein as a "functional repression domain," "repression domain," or "repressor domain." For example, the effector domain may include a functional repression domain derived from a zinc finger repressor protein.
[0136] In some embodiments, the effector domain comprises a KOX1 / ZNF10 repression domain, a KOX8 / ZNF708 repression domain, a ZNF43 repression domain, a ZNF184 repression domain, a ZNF91 repression domain, a HPF4 repression domain, a HTF10 repression domain, HTF34, or any combination thereof. In some embodiments, the repression domain comprises a ZIM3 repression domain, a ZNF436 repression domain, a ZNF257 repression domain, a ZNF675 repression domain, a ZNF490 repression domain, a ZNF320 repression domain, a ZNF331 repression domain, a ZNF816 repression domain, a ZNF680 repression domain, a ZNF41 repression domain, a ZNF189 repression domain, a ZNF528 repression domain, a ZNF543 repression domain. domain, ZNF554 repression domain, ZNF140 repression domain, ZNF610 repression domain, ZNF264 repression domain, ZNF350 repression domain, ZNF8 repression domain, ZNF582 repression domain, ZNF30 repression domain, ZNF324 repression domain, ZNF98 repression domain, ZNF669 repression domain, ZNF677 repression domain, ZNF596 repression domain, ZNF214 repression domain, ZNF3 7A repression domain, ZNF34 repression domain, ZNF250 repression domain, ZNF547 repression domain, ZNF273 repression domain, ZNF354A repression domain, ZFP82 repression domain, ZNF224 repression domain, ZNF33A repression domain, ZNF45 repression domain, ZNF175 repression domain, ZNF595 repression domain, ZNF184 repression domain, ZNF419 repression domain, ZFP28-1 repression domain In some embodiments, the effector domain comprises a ZIM3 repression domain, a ZNF554 repression domain, a ZNF264 repression domain, a ZNF324 repression domain, a ZNF354A repression domain, a ZNF189 repression domain, a ZNF543 repression domain, a ZFP82, a ZNF669, a ZNF582 repression domain, or any combination thereof.In some embodiments, the effector domain comprises a ZIM3 repression domain, a ZNF554 repression domain, a ZNF264 repression domain, a ZNF324 repression domain, or a ZNF354A repression domain, or any combination thereof. In some embodiments, the effector domain is a ZIM3 repression domain.
[0137] In some embodiments, the repression domain is a KRAB domain. In some embodiments, the effector domain comprises a functional repression domain that comprises or is derived from a KOX1 / ZNF10 KRAB domain, a KOX8 / ZNF708 KRAB domain, a ZNF43 KRAB domain, a ZNF184 KRAB domain, a ZNF91 KRAB domain, a HPF4 KRAB domain, a HTF10 KRAB domain, or a HTF34 KRAB domain, or any combination thereof.In some embodiments, the effector domain comprises a functional repressor domain derived from the following domains: a ZIM3 KRAB domain, a ZNF436 KRAB domain, a ZNF257 KRAB domain, a ZNF675 KRAB domain, a ZNF490 KRAB domain, a ZNF320 KRAB domain, a ZNF331 KRAB domain, a ZNF816 KRAB domain, a ZNF680 KRAB domain, a ZNF41 KRAB domain, a ZNF189 KRAB domain, a ZNF528 KRAB domain, a ZNF543 KRAB domain, a ZNF554 KRAB domain, a ZNF140 KRAB domain, a ZNF610 KRAB domain, a ZNF264 KRAB domain, a ZNF350 KRAB domain, a ZNF8 KRAB domain, a ZNF582 KRAB domain, a ZNF680 KRAB domain, a ZNF41 KRAB domain, a ZNF189 KRAB domain, a ZNF528 KRAB domain, a ZNF543 KRAB domain, a ZNF554 KRAB domain, a ZNF140 KRAB domain, a ZNF610 KRAB domain, a ZNF264 KRAB domain, a ZNF350 KRAB domain, a ZNF8 KRAB domain, a ZNF582 KRAB domain, ZNF30 KRAB domain, ZNF324 KRAB domain, ZNF98 KRAB domain, ZNF669 KRAB domain, ZNF677 KRAB domain, ZNF596 KRAB domain, ZNF214 KRAB domain, ZNF37A KRAB domain, ZNF34 KRAB domain, ZNF250 KRAB domain, ZNF547 KRAB domain, ZNF273 KRAB domain, ZNF354 AKRAB domain, ZFP82 KRAB domain, ZNF224 KRAB domain, ZNF33 AKRAB domain, ZNF45 KRAB domain, ZNF175 KRAB domain, ZNF595 KRAB domain, ZNF184 KRAB domain, ZNF419 KRAB domain, ZFP28-1 KRAB domain, ZFP28-2 KRAB domain, ZNF18 KRAB domain, ZNF213 KRAB domain, ZNF394 KRAB domain, ZFP1 KRAB domain, ZFP14 KRAB domain, ZNF416 KRAB domain, ZNF557 KRAB domain, ZNF566 KRAB domain, ZNF729 KRAB domain, ZIM2 KRAB domain, ZNF254 KRAB domain, ZNF764 KRAB domain, ZNF785 KRAB domain, or any combination thereof.In some embodiments, the domain is a ZIM3 KRAB domain, a ZNF554 KRAB domain, a ZNF264 KRAB domain, a ZNF324 KRAB domain, a ZNF354 AKRAB domain, a ZNF189 KRAB domain, a ZNF543 KRAB domain, a ZFP82 KRAB domain, a ZNF669 KRAB domain, or a ZNF582 KRAB domain, or any combination thereof. In some embodiments, the domain is a ZIM3 KRAB domain, a ZNF554 KRAB domain, a ZNF264 KRAB domain, a ZNF324 KRAB domain, or a ZNF354 AKRAB domain, or any combination thereof. In some embodiments, the domain is a ZIM3 KRAB domain.
[0138] Table 2 below provides the sequence of an exemplary functional repressor domain that reduces or silences target gene expression. In some embodiments, the epigenetic editing system herein comprises one or more repressor domains selected from Table 2, or functional homologs, orthologs, or variants thereof. Further examples of suitable repressors and repressor domains can be found in PCT / US2021 / 030643; and in Tycko et al. High-Throughput Discovery and Characterization of Human Transcriptional Effectors. Cell. 2020 Dec 23; 183 (7): 2020-2035., each of which is incorporated herein by reference in its entirety.
[0139] Table 2. Exemplary protein sequences comprising repression domains suitable for use as epigenetic effector domains that can reduce or silence gene expression
[0140] protein Protein sequence KRAB SEQ ID NO:9 N-terminal KRAB SEQ ID NO:10 C-terminal KRAB SEQ ID NO:11 ZIM3 SEQ ID NO:100 ZFP28 SEQ ID NO:44 ZN627 SEQ ID NO:45 KAP1 SEQ ID NO:46 MeCP2 SEQ ID NO:47 HP1b SEQ ID NO:48 CBX8 SEQ ID NO:49 CDYL2 SEQ ID NO:50 TOX SEQ ID NO:51 TOX3 SEQ ID NO:52 TOX4 SEQ ID NO:53 EED SEQ ID NO:54 RBBP4 SEQ ID NO:55 RCOR1 SEQ ID NO:56 SCML2 SEQ ID NO:57 KOX1(ZNF10) SEQ ID NO:94
[0141] In some embodiments, the effector domain may comprise SEQ ID NO: 9. In some embodiments, the DNMT may be SEQ ID NO: 9. In some embodiments, the DNMT may be at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% similar to SEQ ID NO:9.
[0142] In some embodiments, the effector domain may comprise SEQ ID NO: 10. In some embodiments, the DNMT may be SEQ ID NO: 10. In some embodiments, the DNMT may be at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% similar to SEQ ID NO: 10.
[0143] In some embodiments, the effector domain may comprise SEQ ID NO: 11. In some embodiments, the DNMT may be SEQ ID NO: 11. In some embodiments, the DNMT may be at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% similar to SEQ ID NO: 11.
[0144] In some embodiments, the effector domain may comprise SEQ ID NO: 100. In some embodiments, the DNMT may be SEQ ID NO: 100. In some embodiments, the DNMT may be at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% similar to SEQ ID NO: 100.
[0145] In some embodiments, the effector domain may comprise SEQ ID NO: 44. In some embodiments, the DNMT may be SEQ ID NO: 44. In some embodiments, the DNMT may be at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% similar to SEQ ID NO: 44.
[0146] In some embodiments, the effector domain may comprise SEQ ID NO: 45. In some embodiments, the DNMT may be SEQ ID NO: 45. In some embodiments, the DNMT may be at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% similar to SEQ ID NO: 45.
[0147] In some embodiments, the effector domain may comprise SEQ ID NO: 46. In some embodiments, the DNMT may be SEQ ID NO: 46. In some embodiments, the DNMT may be at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% similar to SEQ ID NO: 46.
[0148] In some embodiments, the effector domain may comprise SEQ ID NO: 47. In some embodiments, the DNMT may be SEQ ID NO: 47. In some embodiments, the DNMT may be at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% similar to SEQ ID NO: 47.
[0149] In some embodiments, the effector domain may comprise SEQ ID NO: 48. In some embodiments, the DNMT may be SEQ ID NO: 48. In some embodiments, the DNMT may be at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% similar to SEQ ID NO: 48.
[0150] In some embodiments, the effector domain may comprise SEQ ID NO: 49. In some embodiments, the DNMT may be SEQ ID NO: 49. In some embodiments, the DNMT may be at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% similar to SEQ ID NO: 49.
[0151] In some embodiments, the effector domain may comprise SEQ ID NO: 50. In some embodiments, the DNMT may be SEQ ID NO: 50. In some embodiments, the DNMT may be at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% similar to SEQ ID NO: 50.
[0152] In some embodiments, the effector domain may comprise SEQ ID NO: 51. In some embodiments, the DNMT may be SEQ ID NO: 51. In some embodiments, the DNMT may be at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% similar to SEQ ID NO: 51.
[0153] In some embodiments, the effector domain may comprise SEQ ID NO: 52. In some embodiments, the DNMT may be SEQ ID NO: 52. In some embodiments, the DNMT may be at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% similar to SEQ ID NO: 52.
[0154] In some embodiments, the effector domain may comprise SEQ ID NO: 53. In some embodiments, the DNMT may be SEQ ID NO: 53. In some embodiments, the DNMT may be at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% similar to SEQ ID NO: 53.
[0155] In some embodiments, the effector domain may comprise SEQ ID NO: 54. In some embodiments, the DNMT may be SEQ ID NO: 54. In some embodiments, the DNMT may be at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% similar to SEQ ID NO: 54.
[0156] In some embodiments, the effector domain may comprise SEQ ID NO: 55. In some embodiments, the DNMT may be SEQ ID NO: 55. In some embodiments, the DNMT may be at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% similar to SEQ ID NO: 55.
[0157] In some embodiments, the effector domain may comprise SEQ ID NO: 56. In some embodiments, the DNMT may be SEQ ID NO: 56. In some embodiments, the DNMT may be at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% similar to SEQ ID NO: 56.
[0158] In some embodiments, the effector domain may comprise SEQ ID NO: 57. In some embodiments, the DNMT may be SEQ ID NO: 57. In some embodiments, the DNMT may be at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% similar to SEQ ID NO: 57.
[0159] In some embodiments, the effector domain may comprise SEQ ID NO: 94. In some embodiments, the DNMT may be SEQ ID NO: 94. In some embodiments, the DNMT may be at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% similar to SEQ ID NO: 94.
[0160] In some embodiments, the effector domain comprises a fusion of two effector domains (e.g., KOX1KRAB and ZIM3). In some embodiments, the effector domain comprises a fusion of 2, 3, 4, 5, 6, 7, 8, 9, or 10 effector domains. In some embodiments, the effector domain comprises a fusion of a truncated form of the effector domain and a second effector domain. In some embodiments, the effector domain comprises a fusion of truncated forms of two effector domains. In some embodiments, the fused effector domain comprises at least one truncated form of the effector domain.
[0161] In some embodiments, the effector domain comprises a functional domain that inhibits or silences gene expression, and the functional domain is a part of a larger protein (e.g., zinc finger repressor). The functional domain that can regulate gene expression (e.g., inhibit or increase gene expression) can be identified from a larger protein using known methods and methods provided herein. For example, a functional effector domain that can reduce or silence target gene expression can be identified based on the sequence of a repressor or activator protein. The amino acid sequence of a protein with the function of regulating gene expression can be obtained from an available genome browser (e.g., UCSD genome browser or Ensembl genome browser).
[0162] Connectors
[0163] The epigenetic editing system provided herein can include one or more joints connecting one or more components of the epigenetic editing system. The joint can be a covalent bond or a polymer joint having many atoms in length. The joint can be a peptide joint or a non-peptide joint.
[0164] In certain embodiments, the joint can be used to connect any one of the peptides or peptide domains of the epigenetic editing system. The joint can be as simple as a covalent bond, or it can be a polymer joint with a length of many atoms. In certain embodiments, the joint is a polypeptide or based on amino acids. In other embodiments, the joint is not peptide-like. In certain embodiments, the joint is a covalent bond (e.g., carbon-carbon bond, disulfide bond, carbon-heteroatom bond, etc.). In certain embodiments, the joint is a carbon-nitrogen bond of an amide bond. In certain embodiments, the joint is a cyclic or acyclic, substituted or unsubstituted, branched or unbranched aliphatic or heteroaliphatic joint. In certain embodiments, the joint is a polymer (e.g., polyethylene, polyethylene glycol, polyamide, polyester, etc.). In certain embodiments, the joint comprises a monomer, dimer or polymer of an aminoalkanoic acid. In certain embodiments, the joint comprises an aminoalkanoic acid (e.g., glycine, acetic acid, alanine, β-alanine, 3-aminopropionic acid, 4-aminobutyric acid, 5-pentanoic acid, etc.). In certain embodiments, the joint comprises a monomer, dimer or polymer of aminocaproic acid (Ahx). In certain embodiments, the joint is based on a carbocyclic moiety (e.g., cyclopentane, cyclohexane). In other embodiments, the joint comprises a polyethylene glycol moiety (PEG). In other embodiments, the joint comprises an amino acid. In certain embodiments, the joint comprises a peptide. In certain embodiments, the joint comprises an aryl or heteroaryl moiety. In certain embodiments, the joint is based on a benzene ring. The joint may include a functionalized portion to promote the connection of a nucleophile (e.g., sulfhydryl, amino) from a peptide to the joint. Any electrophile may be used as a part of a joint. Exemplary electrophiles include, but are not limited to, activated esters, activated amides, Michael acceptors, alkyl halides, aryl halides, acyl halides and isothiocyanates.
[0165] In some embodiments, the linker is a non-peptide linker. For example, the linker can be a carbon bond, a disulfide bond, or a carbon-heteroatom bond. In certain embodiments, the linker is a carbon-nitrogen bond of an amide bond. In certain embodiments, the linker is a cyclic or acyclic, substituted or unsubstituted, branched or unbranched aliphatic or heteroaliphatic linker.
[0166] In certain embodiments, the joint is a polymer (e.g., polyethylene, polyethylene glycol, polyamide, polyester, etc.). In certain embodiments, the joint comprises a monomer, dimer or polymer of an aminoalkanoic acid. In certain embodiments, the joint comprises an aminoalkanoic acid (e.g., glycine, acetic acid, alanine, β-alanine, 3-aminopropionic acid, 4-aminobutyric acid, 5-pentanoic acid, etc.). In certain embodiments, the joint comprises a monomer, dimer or polymer of aminocaproic acid (Ahx). In certain embodiments, the joint is based on a carbocyclic moiety (e.g., cyclopentane, cyclohexane). In other embodiments, the joint comprises a polyethylene glycol moiety (PEG). In other embodiments, the joint comprises an amino acid. In certain embodiments, the joint comprises a peptide. In certain embodiments, the joint comprises an aryl or heteroaryl moiety. In certain embodiments, the joint is based on a benzene ring. The joint may include a functionalized portion to promote the connection of a nucleophile (e.g., sulfhydryl, amino) from a peptide to the joint. Any electrophile may be used as a part of a joint. Exemplary electrophiles include, but are not limited to, activated esters, activated amides, alkyl halides, aryl halides, acyl halides, and isothiocyanates.
[0167] In some embodiments, one or more linkers of the epigenetic editing system provided herein are peptide linkers. For example, a zinc finger array and a repressor domain can be connected by a peptide linker to form a zinc finger-repressor fusion protein. The peptide linker can be any length suitable for the epigenetic editing system fusion protein described herein. In some embodiments, the linker can comprise a peptide of 1 to 200 amino acids. In some embodiments, a DNA binding domain, such as a zinc finger array and an effector domain, is fused via a linker, and the linker comprises a length of 1 to 5, 1 to 10, 1 to 20, 1 to 30, 1 to 40, 1 to 50, 1 to 60, 1 to 80, 1 to 100, 1 to 150, 1 to 200, 5 to 10, 5 to 20, 5 to 30, 5 to 40, 5 to 60, 5 to 80, 5 to 100, 5 to 150, 5 to 200, 10 to 20, 10 to 30, 10 to 40, 10 to 50 , 10 to 60, 10 to 80, 10 to 100, 10 to 150, 10 to 200, 20 to 30, 20 to 40, 20 to 50, 20 to 60, 20 to 80, 20 to 100, 20 to 150, 20 to 200, 30 to 40, 30 to 50, 30 to 60, 30 to 80, 30 to 100, 30 to 150, 30 to 200, 40 to 50, 40 to 60, 40 to 80, 40 to 100, 40 to 150, 40 to 200, 50 to 60 In some embodiments, the peptide linker is a peptide linker having a length of 50 to 80, 50 to 100, 50 to 150, 50 to 200, 60 to 80, 60 to 100, 60 to 150, 60 to 200, 80 to 100, 80 to 150, 80 to 200, 100 to 150, 100 to 200 or 150 to 200 amino acids. Longer or shorter linkers are also contemplated. In some embodiments, the length of the peptide linker is 4, 16, 32 or 104 amino acids. In some embodiments, the peptide linker is a flexible linker. In some embodiments, the peptide linker is a rigid linker.
[0168] The length of the joint can be 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 25, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95 or 100 amino acids. In some embodiments, the length of the joint is 5 amino acids. In some embodiments, the length of the joint is 16 amino acids. In some embodiments, the length of the joint is 20 amino acids. In some embodiments, the length of the joint is 26 amino acids. In some embodiments, the length of the joint is 80 amino acids.
[0169] In some embodiments, the peptide linker comprises the amino acid sequence of SEQ ID NOs: 1-5.
[0170] In some embodiments, the peptide linker is an XTEN linker. In some embodiments, the linker comprises an XTEN16 linker. In some embodiments, the linker comprises an amino acid sequence of SEQ ID NO: 2. In some embodiments, the linker comprises an amino acid sequence having at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91% or 90% sequence similarity to SEQ ID NO: 2. In some embodiments, the peptide linker comprises an XTEN80 linker. In some embodiments, the peptide linker comprises an amino acid sequence of SEQ ID NO: 3. In some embodiments, the linker comprises an amino acid sequence having at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91% or 90% sequence similarity to SEQ ID NO: 3.
[0171] A variety of linker lengths and flexibilities (e.g., ranging from very flexible linkers such as glycine / serine-rich linkers to more rigid linkers) between an effector domain (e.g., a repressor domain) and a DNA binding protein (e.g., a Cas9 domain), between an effector domain and a second effector domain, or between any two components of an epigenetic editing system can be used to achieve an optimal length for effector domain activity for a particular application. In some embodiments, the flexible linker is a glycine / serine-rich linker (GS-rich linker) in which more than 45% (e.g., more than 48%, 50%, 55%, 60%, 70%, 80%, or 90%) of the residues are glycine or serine residues. A non-limiting example of a GS-rich linker is (GGGGS)n ((SEQ ID NO:4) and (G)n. In some embodiments, more rigid linkers include (EAAAK)n, (SGGS)n and (XP)n. In the formulas of the above flexible and rigid linkers, n is any integer between 1 and 30. In some embodiments, n is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14 or 15. In some embodiments, the linker comprises a (GGS)n motif, wherein n is 1, 3 or 7. In some embodiments, the linker comprises a (GGGGS)n motif, wherein n is 4 (SEQ ID NO:4).
[0172] In some embodiments, the linker in the epigenetic editing system comprises a nuclear localization signal, such as the peptide sequence of SEQ ID NO: 7. In some embodiments, the linker in the epigenetic editing system comprises an expression tag, such as a detectable tag, such as green fluorescent protein.
[0173] In some embodiments, the joint comprises a nucleic acid. For example, one or more joints of the epigenetic editing system may include a nucleic acid that can bind to, interact with, associate with, or form a complex with a polypeptide. In some embodiments, the nucleic acid joint can be an RNA joint that can bind to an RNA binding protein domain (e.g., a derived RNA binding domain) and / or interact with it. In some embodiments, the nucleic acid joint can be fused to a guide polynucleotide that can bind to the Cas protein of the epigenetic editing system. In some embodiments, the nucleic acid joint comprises a K homology (KH) domain binding sequence, an MS2 capsid protein binding sequence, a PP7 capsid protein binding sequence, a SfMu COM capsid protein binding sequence, a telomerase Ku binding motif binding sequence, a sm7 protein binding sequence, or other RNA recognition motif binding sequences thereof.
[0174] In some embodiments, the joint comprises an affinity domain that specifically binds to a component of an epigenetic effector. For example, an epigenetic effector may comprise a programmable DNA binding domain-a joint comprising an affinity domain having specific binding affinity to an epigenetic effector domain. The affinity domain may include antibodies, single-chain antibodies, nano antibodies and antigen binding sequences, antibodies, nano antibodies, functional antibody fragments, single-chain variable fragments (scFv), Fab, single domain antibodies (sdAb), VH domains, VL domains, VNAR domains, VHH domains, bispecific antibodies, double antibodies or functional fragments or combinations thereof. In some embodiments, the epigenetic effector domain comprises a programmable DNA binding domain and a KAP1 antibody that is bound to a KAP1 protein. In some embodiments, the epigenetic effector domain comprises a programmable DNA binding domain and a KRAB antibody that is bound to a KRAB protein. In some embodiments, the epigenetic effector domain comprises a programmable DNA binding domain and a DNMT1 antibody that is bound to a DNMT1 protein. In some embodiments, the epigenetic effector domain comprises a programmable DNA binding domain and a DNMT3A antibody that binds to a DNMT3A protein. In some embodiments, the epigenetic effector domain comprises a programmable DNA binding domain and a DNMT3L antibody that binds to a DNMT3L protein. In some embodiments, the epigenetic effector domain comprises a programmable DNA binding domain and a ZIM3 antibody that binds to a ZIM3 protein. In some embodiments, the epigenetic effector domain comprises a programmable DNA binding domain and a TET1 antibody that binds to a TET1 protein. In some embodiments, the epigenetic effector domain comprises a programmable DNA binding domain and a VP16 or VP64 antibody that binds to a VP16 or VP64 protein.
[0175] In some embodiments, the linker comprises a repeating peptide array. In some embodiments, the linker comprises an epitope tag, such as a SunTag. In some embodiments, the epigenetic editing system comprises one or more peptide arrays, the peptide array comprising multiple copies of an epitope tag, the epitope tag can be connected to multiple effector domains, the effector domains are attached to or fused to a peptide that recognizes an epitope tag. For example, an epitope tag array can connect a DNA binding domain and multiple effector domains or multiple copies of an effector domain, the effector domains are fused to or attached to an antibody sequence that recognizes an epitope tag. In some embodiments, the epigenetic editing system comprises at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more epitope tag repeats, which are connected to at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more effector domains or copies of effector domains. In some embodiments, the epigenetic editing system comprises a plurality of epitope tag repeat sequences that are linked to a plurality of effector domains and a detectable expression tag domain, such as GFP. In some embodiments, the repeat peptide array comprises a gene control non-repressible 4 (GCN4) peptide sequence. In some embodiments, the repeat peptide array is further linked by linking a peptide sequence of 15 to 50 amino acids. The repeat peptide arrays as described in U.S. Patent Application No. US20170219596 and U.S. Patent No. 10,612,044 are incorporated herein by reference in their entirety.
[0176] nuclear localization sequence
[0177] In some embodiments, the epigenetic editing system provided herein comprises one or more nuclear targeting sequences. In some embodiments, the epigenetic editing system provided herein comprises at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more nuclear targeting sequences. For example, the zinc finger-repressor fusion protein described herein may also include one or more nuclear targeting sequences, such as a nuclear localization sequence (NLS). In some embodiments, the fusion protein comprises multiple NLS. In some embodiments, the fusion protein comprises NLS at the N-terminus or C-terminus of the fusion protein. In some embodiments, the fusion protein comprises NLS at both the N-terminus and the C-terminus. In some embodiments, the NLS is embedded in the middle of the fusion protein. In some embodiments, the NLS comprises an amino acid sequence that promotes the protein containing the NLS to be imported into the nucleus. In some embodiments, the NLS is fused to the N-terminus of the fusion protein. In some embodiments, the NLS is fused to the C-terminus of the fusion protein. In some embodiments, the NLS is fused to the N-terminus of the nucleic acid binding protein (e.g., Cas9 or zinc finger array). In some embodiments, the NLS is fused to the C-terminus of the nucleic acid binding protein. In some embodiments, the NLS is fused to the N-terminus of an effector domain (e.g., a repressor domain). In some embodiments, the NLS is fused to the C-terminus of an effector domain (e.g., a repressor domain). In some embodiments, the NLS is fused to the fusion protein via one or more joints. In some embodiments, the NLS is fused to the fusion protein without a joint. In some embodiments, the NLS comprises an amino acid sequence of any of the NLS sequences provided or mentioned herein.
[0178] In some embodiments, the epigenetic editing system provided herein comprises two NLS sequences. In some embodiments, the fusion protein of the epigenetic editing system comprises an NLS at the N-terminus and an NLS at the C-terminus. In some embodiments, the fusion protein comprises two NLSs at the N-terminus. In some embodiments, the fusion protein comprises two NLSs at the C-terminus. In some embodiments, one NLS is located at the N-terminus, and one NLS is embedded in the middle of the fusion protein. In some embodiments, one NLS is located at the C-terminus, and one NLS is embedded in the middle of the fusion protein. In some embodiments, both NLSs are embedded in the middle of the fusion protein.
[0179] In some embodiments, the fusion protein of the epigenetic editing system comprises two NLS sequences on either side of the DNMT domain. In some embodiments, the fusion protein comprises two NLS sequences on either side of the fusion DNMT domain. In some embodiments, the fusion protein comprises two NLS sequences on either side of the DNA binding domain. In some embodiments, the fusion protein domain comprises two NLS sequences on either side of the effector domain.
[0180] In some embodiments, the epigenetic editing system provided herein comprises four NLS sequences. In some embodiments, the fusion protein of the epigenetic editing system comprises at least two NLS at the N-terminus. In some embodiments, the fusion protein of the epigenetic editing system comprises at least two NLS at the C-terminus. In some embodiments, the fusion protein of the epigenetic editing system comprises two NLS at the N-terminus and two NLS at the C-terminus. In some embodiments, at least one NLS is embedded in the middle of the fusion protein.
[0181] In some embodiments, the NLS comprises the amino acid sequence SEQ ID NO: 7. In some embodiments, the NLS sequence is an endogenous NLS sequence. Other nuclear localization sequences are known in the art and are apparent to those skilled in the art. Examples of NLS sequences suitable for inclusion in fusion proteins as provided herein include, but are not limited to, those described in Lu et al., Types of nuclear localization signals and mechanisms of protein import into the nucleus, Cell Commun Signal. 2021; 19: 60, 2021; the entire contents of which are incorporated herein by reference.
[0182] In some embodiments, a fusion protein comprising two NLSs at the N-terminus and two NLSs at the C-terminus can increase the efficiency of the epigenetic editing system by at least 5%, at least 10%, at least 15%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 100%, at least 200%, at least 300%, at least 400%, at least 500%, at least 600%, at least 700%, at least 800%, at least 900%, at least 1,000%, at least 5,000%, at least 10,000%, at least 50,000%, at least 100,000% or more compared to an epigenetic editing system that does not have two NLSs at the N-terminus and does not have two NLSs at the C-terminus. In some embodiments, a fusion protein comprising two NLSs at the N-terminus and two NLSs at the C-terminus can increase the efficiency of the epigenetic editing system by up to 100,000%, up to 50,000%, up to 10,000%, up to 5,000%, up to 1,000%, up to 900%, up to 800%, up to 700%, up to 600%, up to 500%, up to 400%, up to 300%, up to 200%, up to 100%, up to 90%, up to 80%, up to 70%, up to 60%, up to 50%, up to 40%, up to 30%, up to 20%, up to 15%, up to 10%, up to 5% or less compared to an epigenetic editing system that does not have two NLSs at the N-terminus and does not have two NLSs at the C-terminus.
[0183] Label
[0184] The epigenetic editing system provided herein may include one or more other sequence domains, tags for tracking, detection and positioning of the editor. In some embodiments, the epigenetic editing system includes one or more detectable tags. In some embodiments, the epigenetic editing system includes 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more detectable tags. Each detectable tag may be the same or different.
[0185] For example, epigenetic editing system fusion protein can include cytoplasmic localization sequence, export sequence, such as nuclear export sequence or other localization sequence, and can be used for the dissolution of fusion protein, purification or detection sequence tag. Suitable protein tags provided herein include but are not limited to biotin carboxylase carrier protein (BCCP) tag, myc tag, calmodulin tag, FLAG tag, hemagglutinin (HA) tag, polyhistidine tag (also referred to as histidine tag or His tag), maltose binding protein (MBP) tag, nus tag, glutathione-S-transferase (GST) tag, green fluorescent protein (GFP) tag, thioredoxin tag, S- tag, Sof tag (for example, Sof tag 1, Sof tag 3), streptococcus tag, biotin ligase tag, FlAsH tag, V5 tag and SBP tag. Other suitable sequences will be apparent to those skilled in the art.
[0186] In some embodiments, the epigenetic editing system comprises 1 to 2 detectable tags. In various aspects, the fusion protein comprises 1 detectable tag. In various aspects, the fusion protein comprises 2 detectable tags. In various aspects, the fusion protein comprises 3 detectable tags. In various aspects, the fusion protein comprises 4 detectable tags. In various aspects, the fusion protein comprises 5 detectable tags.
[0187] Epigenetic editing system structure
[0188] Some aspects of the present disclosure provide epigenetic editing systems. Exemplary, non-limiting structures of such editing systems are described in more detail herein. Based on the present disclosure, other suitable structures will be clear to those skilled in the art, and the present disclosure is not limited in this regard.
[0189] The multiple components of the epigenetic editing system described herein can be arranged in any order. In some embodiments, the epigenetic editing system comprises the following structure: N']-[D1]-[D2]-[C', wherein any one of D1 and D2 is a DNA binding domain, an effector domain, or a nucleic acid binding domain. In these structural examples, N' represents the N-terminus, C' represents the C-terminus; and]-[ represents a connecting element or a linker.
[0190] In some embodiments, the epigenetic editing system comprises the structure: N']-[D1]-[D2]-[D3]-[C', wherein any one of D1, D2, and D3 is a DNA binding domain, an effector domain, or a nucleic acid binding domain. In some embodiments, D1 is a DNA binding domain. In some embodiments, D2 is a DNA binding domain. In some embodiments, D3 is a DNA binding domain. In some embodiments, D1 is the only DNA binding domain. In some embodiments, D2 is the only DNA binding domain. In some embodiments, D3 is the only DNA binding domain.
[0191] In some embodiments, the epigenetic editing system comprises the structure: N']-[D1]-[D2]-[D3]-[D4]-[C', wherein any one of D1, D2, D3, and D4 is a DNA binding domain, an effector domain, or a nucleic acid binding domain. In some embodiments, D1 is a DNA binding domain. In some embodiments, D2 is a DNA binding domain. In some embodiments, D3 is a DNA binding domain. In some embodiments, D4 is a DNA binding domain. In some embodiments, D1 is the only DNA binding domain. In some embodiments, D2 is the only DNA binding domain. In some embodiments, D3 is the only DNA binding domain. In some embodiments, D4 is the only DNA binding domain.
[0192] In some embodiments, the epigenetic editing system comprises the structure: N']-[D1]-[D2]-[D3]-[D4]-[D5]-[C', wherein any one of D1, D2, D3, D4, and D5 is a DNA binding domain, an effector domain, or a nucleic acid binding domain. In some embodiments, D1 is a DNA binding domain. In some embodiments, D2 is a DNA binding domain. In some embodiments, D3 is a DNA binding domain. In some embodiments, D4 is a DNA binding domain. In some embodiments, D5 is a DNA binding domain. In some embodiments, D1 is the only DNA binding domain. In some embodiments, D2 is the only DNA binding domain. In some embodiments, D3 is the only DNA binding domain. In some embodiments, D4 is the only DNA binding domain. In some embodiments, D5 is the only DNA binding domain.
[0193] In some embodiments, the epigenetic editing system comprises at least one effector domain, which is a DNMT domain. In some embodiments, the epigenetic editing system comprises at least one effector domain, which is a KRAB domain. In some embodiments, the epigenetic editing system comprises at least one effector domain, which is a ZIM KRAB domain. In some embodiments, the epigenetic effector comprises at least one effector domain, which is a DNMT3A domain or a truncated version thereof. In some embodiments, the epigenetic effector comprises at least one effector domain, which is a DNMT3L domain or a truncated version thereof.
[0194] The components of the epigenetic editing system can be composed of different configurations. For example, the DNA binding domain can be at the C-terminus, the N-terminus, or between two or more epigenetic effector domains or other domains. In some embodiments, the DNA binding domain is at the C-terminus of the epigenetic editing system. In some embodiments, the DNA binding domain is at the N-terminus of the epigenetic editing system. In some embodiments, the DNA binding domain is connected to one or more nuclear localization signals. In some embodiments, the DNA binding domain is connected to two or more nuclear localization signals. In some embodiments, the flank of the DNA binding domain is an epigenetic effector domain or other domains at both ends. In some embodiments, the epigenetic editing system includes N']-[epigenetic effector domain 1]-[DNA binding domain]-[epigenetic effector domain 2]-[C' configuration. In some embodiments, the epigenetic editing system includes N']-[epigenetic effector domain 1]-[DNA binding domain]-[epigenetic effector domain 2]-[epigenetic effector domain 3]-[C' configuration. In some embodiments, the epigenetic editing system comprises a configuration of N']-[epigenetic effector domain 1]-[epigenetic effector domain 2]-[DNA binding domain]-[epigenetic effector domain 3]-[C'. In some embodiments, the epigenetic editing system comprises a configuration of N']-[epigenetic effector domain 1]-[epigenetic effector domain 2]-[DNA binding domain]-[epigenetic effector domain 3]-[epigenetic effector domain 4]-[C'. In some embodiments, the epigenetic editing system comprises a configuration of N']-[KRAB]-[DNA binding domain]-[Dnmt3A]-[C'. In some embodiments, the epigenetic editing system comprises a configuration of N']-[KRAB]-[DNA binding domain]-[Dnmt3A]-[Dnmt3L]-[C'. In some embodiments, the epigenetic editing system comprises a configuration of N']-[SETDB1]-[DNA binding domain]-[Dnmt3A]-[Dnmt3L]-[C'. In some embodiments, the epigenetic editing system comprises a configuration of N']-[SETDB1]-[DNA binding domain]-[Dnmt3A]-[C'. In some embodiments, the epigenetic editing system comprises a configuration of N']-[KRAB]-[DNA binding domain]-[Dnmt3A-Dnmt3L]-[C', wherein Dnmt3A and Dnmt3L are directly fused via a peptide bond.
[0195] In some embodiments, the connection structure "]-[" in any of the epigenetic editing system structures is a linker, such as a peptide linker. In some embodiments, the connection structure "]-[" in any of the epigenetic editing system structures is a detectable label. In some embodiments, the connection structure "]-[" in any of the epigenetic editing system structures is a peptide bond. In some embodiments, the connection structure "]-[" in any of the epigenetic editing system structures is a nuclear localization signal. In some embodiments, the connection structure "]-[" in any of the epigenetic editing system structures is a promoter or a regulatory sequence. In the epigenetic editing system structure, multiple connection structures "]-[" may be the same, or may each be a different linker, label, NLS, or peptide bond.
[0196] The DNA binding domain (DBD) of the epigenetic editing system may include any of the DNA binding domains described herein or known to those skilled in the art. In some embodiments, the DBD includes one or more zinc finger arrays. In some embodiments, the DBD includes a TALE DNA binding domain. In some embodiments, the DBD is an RNA-guided programmable DNA binding domain, such as a CRISPR-Cas protein domain. Suitable Cas proteins have been provided herein, including nuclease-inactive Cas proteins for the purpose of epigenetic editing without causing target DNA strand breaks. The Cas protein in the epigenetic editing system can be a nuclease-inactive Cas9 ((dCas9), SaCas9d, SpCas9d, dCas9 with modified PAM specificity, high-fidelity dCas9, nuclease-inactive Cpf1 (dCpf1), dCpf1 with modified PAM specificity, high-fidelity dCpf1, dCas12e, dCasY, or any other Cas protein as described herein.
[0197] In some embodiments, the epigenetic editing system comprises a DNA binding domain (DBD) and an effector domain that inhibits or silences the expression of a target gene. In some embodiments, the epigenetic editing system comprises a N']-[repression domain]-[DBD]-[-C' configuration, wherein the connection structure]-[is any one of a linker, a detectable tag, an affinity domain, a peptide bond, a nuclear localization signal, a promoter, and / or a regulatory sequence as described herein. In some embodiments, the epigenetic editing system comprises a N']-[DBD]-[repression domain]-[-C' configuration, wherein the connection structure]-[is any one of a linker, a detectable tag, an affinity domain, a peptide bond, a nuclear localization signal, a promoter, and / or a regulatory sequence as described herein.
[0198] In some embodiments, the epigenetic editing system comprises a DNA binding domain (DBD) and a DNA methyltransferase domain, which deposits one or more methylation markers at the target gene, thereby inhibiting or silencing the expression of the target gene. In some embodiments, the epigenetic editing system comprises a configuration of N']-[DNA methyltransferase domain]-[DBD]-[-C', wherein the connecting structure]-[is any one of a linker, a detectable tag, an affinity domain, a peptide bond, a nuclear localization signal, a promoter, and / or a regulatory sequence as described herein. In some embodiments, the epigenetic editing system comprises a configuration of N']-[DBD]-[DNA methyltransferase domain]-[-C', wherein the connecting structure]-[is any one of a linker, a detectable tag, an affinity domain, a peptide bond, a nuclear localization signal, a promoter, and / or a regulatory sequence as described herein.
[0199] In some embodiments, the epigenetic editing system comprises a DNA binding domain (DBD), a DNA methyltransferase domain, and an effector domain that inhibits or silences the expression of a target gene. In some embodiments, the epigenetic editing system comprises a configuration of N']-[DNA methyltransferase domain]-[DBD]-[repression domain]-[C', wherein the connecting structure]-[is any one of a linker, a detectable tag, an affinity domain, a peptide bond, a nuclear localization signal, a promoter, and / or a regulatory sequence as described herein. In some embodiments, the epigenetic editing system comprises a configuration of N']-[repression domain]-[DBD]-[DNA methyltransferase domain]-[C', wherein the connecting structure]-[is any one of a linker, a detectable tag, an affinity domain, a peptide bond, a nuclear localization signal, a promoter, and / or a regulatory sequence as described herein.
[0200] In some embodiments, the epigenetic editing system comprises a configuration of N']-[DNA methyltransferase domain]-[repression domain]-[DBD]-[C', wherein the connecting structure]-[is any one of a linker, a detectable tag, an affinity domain, a peptide bond, a nuclear localization signal, a promoter, and / or a regulatory sequence as described herein. In some embodiments, the epigenetic editing system comprises a configuration of N']-[repression domain]-[DNA methyltransferase domain]-[DBD]-[C', wherein the connecting structure]-[is any one of a linker, a detectable tag, an affinity domain, a peptide bond, a nuclear localization signal, a promoter, and / or a regulatory sequence as described herein.
[0201] The repression domain in the epigenetic editing system may include any expression repressor protein known to those skilled in the art and as described herein, or any homologue or combination thereof. In some embodiments, the repression domain includes a histone deacetylase domain. In some embodiments, the repression domain interacts with a scaffold protein domain, and the scaffold protein domain recruits one or more protein domains that repress the expression of a target gene. For example, the repression domain can recruit a scaffold protein domain or interact with a scaffold protein domain, and the scaffold protein domain recruits PRMT protein, HDAC protein, SETDB1 protein or NuRD protein domain. In some embodiments, the repression domain interacts with the DNA nucleotides of the epigenetic mark in the target gene, thereby repressing or silencing the expression of the target gene. In some embodiments, the repression domain includes a MECP2 domain. In some embodiments, the repression domain includes a KAP1 domain. In some embodiments, the repression domain includes any one of Table 2 or Table 3, or any combination or homologue thereof.
[0202] The DNA methyltransferase domain in the epigenetic editing system may include any one of the DNA methyltransferase proteins known to those skilled in the art and as described herein, or any homolog or combination thereof. In some embodiments, the effector domain includes a DNMT3 domain. In some embodiments, the DNA methyltransferase domain includes a DNMT3A domain. In some embodiments, the DNA methyltransferase domain includes a DNMT3B domain. In some embodiments, the DNA methyltransferase domain includes a DNMT3C domain. In some embodiments, the DNA methyltransferase domain includes a DNMT3L domain. In some embodiments, the DNA methyltransferase domain includes a fusion of a DNMT3A-DNMT3L domain. As described herein, the DNMT3A-DNMT3L fusion domain can be in any order, such as N-DNMT3A-DNMT3L-C or N-DNMT3L-DNMT3A-C. In some embodiments, the DNA methyltransferase domain includes any one of the domains in Table 1, or any combination or homolog thereof.
[0203] In some embodiments, the epigenetic editing system comprises a DNA binding domain (DBD) and an effector domain that increases the expression of the target gene. In some embodiments, the epigenetic editing system comprises a configuration of N']-[activation domain]-[DBD]-[C', wherein the connection structure]-[is any one of a linker, a detectable tag, an affinity domain, a peptide bond, a nuclear localization signal, a promoter, and / or a regulatory sequence as described herein. In some embodiments, the epigenetic editing system comprises a configuration of N']-[DBD]-[activation domain]-[C', wherein the connection structure]-[is any one of a linker, a detectable tag, an affinity domain, a peptide bond, a nuclear localization signal, a promoter, and / or a regulatory sequence as described herein.
[0204] In some embodiments, the epigenetic editing system comprises a DNA binding domain (DBD) and a DNA demethylation domain, which removes one or more methylation markers at the target gene, thereby increasing the expression of the target gene. In some embodiments, the epigenetic editing system comprises a configuration of N']-[DNA demethylase domain]-[DBD]-[C', wherein the connecting structure]-[is any one of a linker, a detectable tag, an affinity domain, a peptide bond, a nuclear localization signal, a promoter, and / or a regulatory sequence as described herein. In some embodiments, the epigenetic editing system comprises a configuration of N']-[DBD]-[DNA demethylase domain]-[C', wherein the connecting structure]-[is any one of a linker, a detectable tag, an affinity domain, a peptide bond, a nuclear localization signal, a promoter, and / or a regulatory sequence as described herein.
[0205] In some embodiments, the epigenetic editing system comprises a DNA binding domain (DBD), a DNA demethylase domain, and an activation effector domain that increases the expression of the target gene. In some embodiments, the epigenetic editing system comprises a configuration of N']-[DNA demethylase domain]-[DBD]-[activation domain]-[C', wherein the connection structure]-[is any one of a linker, a detectable tag, an affinity domain, a peptide bond, a nuclear localization signal, a promoter, and / or a regulatory sequence as described herein. In some embodiments, the epigenetic editing system comprises a configuration of N']-[activation domain]-[DBD]-[DNA demethylase domain]-[C', wherein the connection structure]-[is any one of a linker, a detectable tag, an affinity domain, a peptide bond, a nuclear localization signal, a promoter, and / or a regulatory sequence as described herein.
[0206] In some embodiments, the epigenetic editing system comprises a configuration of N']-[DNA demethylase domain]-[activation domain]-[DBD]-[C', wherein the connecting structure]-[is any one of a linker, a detectable tag, an affinity domain, a peptide bond, a nuclear localization signal, a promoter, and / or a regulatory sequence as described herein. In some embodiments, the epigenetic editing system comprises a configuration of N']-[activation domain]-[DNA demethylase domain]-[DBD]-[C', wherein the connecting structure]-[is any one of a linker, a detectable tag, an affinity domain, a peptide bond, a nuclear localization signal, a promoter, and / or a regulatory sequence as described herein.
[0207] In some embodiments, the epigenetic editing system that reduces or silences the expression of the target gene comprises a DBD and an affinity domain that specifically binds to a repressor domain. For example, the epigenetic editing system may comprise a DBD and a repressor domain antibody. In some embodiments, the epigenetic editing system comprises a DBD and a KAP1 affinity domain. In some embodiments, the epigenetic editing system comprises a DBD and a KRAB affinity domain. In some embodiments, the epigenetic editing system comprises a DBD and a SETDB1 affinity domain. In some embodiments, the epigenetic editing system comprises a DBD and a MECP2 affinity domain. In some embodiments, the epigenetic editing system comprises a DNA methyltransferase and a repressor domain binding affinity domain.
[0208] In some embodiments, the epigenetic editing system that reduces or silences the expression of the target gene comprises a DBD and an affinity domain that specifically binds to a DNA methyltransferase domain. For example, the epigenetic editing system may comprise a DBD and a DNA methyltransferase antibody. In some embodiments, the epigenetic editing system comprises a DBD and a Dnmt3A affinity domain. In some embodiments, the epigenetic editing system comprises a DBD and a Dnmt3L affinity domain. In some embodiments, the epigenetic editing system comprises a repression domain and a DNA methyltransferase binding affinity domain. In some embodiments, the epigenetic editing system comprises a repression domain and a Dnmt3A binding affinity domain. In some embodiments, the epigenetic editing system comprises a repression domain and a Dnmt3L affinity domain. In some embodiments, the epigenetic editing system comprises one or more of KAP1, KRAB, and MECP2 domains and a Dnmt3A binding affinity domain. In some embodiments, the epigenetic editing system comprises one or more of the KAP1 domain and a Dnmt3A binding affinity domain. In some embodiments, the epigenetic editing system comprises one or more of KAP1, KRAB and MECP2 domains and a Dnmt3L binding affinity domain. In some embodiments, the epigenetic editing system comprises one or more of the KAP1 domains and a Dnmt3L binding affinity domain. The affinity domain can be an antibody, a single-chain antibody, a nanobody and an antigen-binding sequence, an antibody, a nanobody, a functional antibody fragment, a single-chain variable fragment (scFv), a Fab, a single domain antibody (sdAb), a VH domain, a VL domain, a VNAR domain, a VHH domain, a bispecific antibody, a diabody or a functional fragment or a combination thereof.
[0209] In some embodiments, the epigenetic editing system that reduces or silences the expression of the target gene comprises a DBD and a first affinity domain that specifically binds to a DNA methyltransferase domain and a second affinity domain that specifically binds to a repressor domain. For example, the epigenetic editing system may comprise a DBD and a DNA methyltransferase antibody and a repressor domain antibody. In some embodiments, the epigenetic editing system comprises a DBD, a KAP1 affinity domain, and a Dnmt3A affinity domain. In some embodiments, the epigenetic editing system comprises a DBD, a KAP1 affinity domain, and a Dnmt3L affinity domain. In some embodiments, the epigenetic editing system comprises a DBD, a MECP2 affinity domain, and a Dnmt3A affinity domain. In some embodiments, the epigenetic editing system comprises a DBD, a MECP2 affinity domain, and a Dnmt3L affinity domain. In some embodiments, the epigenetic editing system comprises a DBD, a KRAB affinity domain, and a Dnmt3A affinity domain. In some embodiments, the epigenetic editing system comprises a DBD, a KRAB affinity domain and a Dnmt3L affinity domain. The affinity domain can be an antibody, a single chain antibody, a nanobody and an antigen binding sequence, an antibody, a nanobody, a functional antibody fragment, a single chain variable fragment (scFv), a Fab, a single domain antibody (sdAb), a VH domain, a VL domain, a VNAR domain, a VHH domain, a bispecific antibody, a double antibody or a functional fragment or a combination thereof.
[0210] In some embodiments, DNA methyltransferase may include any DNMT domain provided herein, or any combination or homologue thereof. In specific embodiments, DNA methyltransferase domain includes DNMT3A or a truncated version thereof, DNMT3L or a truncated version thereof, or both. In specific embodiments, DBD is a catalytically inactive polynucleotide-guided DNA binding domain (e.g., dCas9) or a ZFP domain. In certain embodiments, the repressor domain includes any repressor domain provided herein, or any combination or homologue thereof. For example, in some embodiments, the repressor domain may be a KRAB domain. In certain embodiments, the repressor domain is a ZFP28, ZN627, KAP1, MeCP2, HP1b, CBX8, CDYL2, TOX, Tox3, Tox4, EED, RBBP4, RCOR1 or SCML2 KRAB domain, or a fusion of two of the domains (e.g., a fusion of the N-terminal and C-terminal regions of the ZIM3 and KOX1 KRAB domains). In specific embodiments, the repressor domain is a KRAB domain from ZFP28, ZF627, ZIM3, or KOX1.
[0211] Specific constructs considered herein include:
[0212] DNMT3A-DNMT3L-XTEN80-NLS-dCas9-NLS-XTEN16-KOX1KRAB (configuration 1),
[0213] DNMT3A-DNMT3L-XTEN80-NLS-ZFP domain-NLS-XTEN16-KOX1 KRAB (Configuration 2),
[0214] NLS-DNMT3A-DNMT3L-XTEN80-dCas9-XTEN16-KOX1 KRAB-NLS (configuration 3),
[0215] NLS-DNMT3A-DNMT3L-XTEN80-ZFP domain-XTEN16-KOX1KRAB-NLS (Configuration 4),
[0216] NLS-NLS-DNMT3A-DNMT3L-XTEN80-dCas9-XTEN16-KOX1KRAB-NLS-NLS (Configuration 5), and
[0217] NLS-NLS-DNMT3A-DNMT3L-XTEN80-ZFP domain-XTEN16-KOX1 KRAB-NLS-NLS (Configuration 6).
[0218] In certain embodiments, the fusion construct may have the following configuration:
[0219] NLS-NLS-hDNMT3A-hDNMT3L-XTEN80-dCas9-XTEN16-KOX1KRAB-NLS-NLS (configuration 7),
[0220] NLS-NLS-DNMT3A-DNMT3L-XTEN80-ZFP domain-XTEN16-KOX1 KRAB-NLS-NLS (Configuration 8),
[0221] NLS-NLS-hDNMT3A-hDNMT3L-XTEN80-dCas9-XTEN16-ZFP28KRAB-NLS-NLS (configuration 9),
[0222] NLS-NLS-DNMT3A-DNMT3L-XTEN80-ZFP domain-XTEN16-ZFP28 KRAB-NLS-NLS (configuration 10),
[0223] NLS-NLS-hDNMT3A-hDNMT3L-XTEN80-dCas9-XTEN16-ZN627KRAB-NLS-NLS (Configuration 11), or
[0224] NLS-NLS-DNMT3A-DNMT3L-XTEN80-ZFP domain-XTEN16-ZN627 KRAB-NLS-NLS (Configuration 12).
[0225] NLS-NLS-DNMT3A-hDNMT3L-XTEN80-dCas9-NLS-XTEN16-ZIM3 KRAB-NLS-NLS (configuration 13).
[0226] NLS-NLS-DNMT3A-hDNMT3L-XTEN80-dCas9-NLS-XTEN16-ZN627 KRAB-NLS-NLS (configuration 14).
[0227] NLS-NLS-DNMT3A-hDNMT3L-XTEN80-dCas9-NLS-XTEN16-KOX1 KRAB-NLS-NLS (configuration 15).
[0228] NLS-DNMT3A-hDNMT3L-XTEN80-dCas9-NLS-XTEN16-KOX1-KRAB (configuration 16).
[0229] DNA binding domain
[0230] The epigenetic editing systems and epigenetic editing complexes described herein can comprise one or more nucleic acid binding protein domains, such as a DNA binding domain, that can direct the epigenetic editing system to a target gene associated with a particular condition.
[0231] As used herein, the target gene can include all nucleotide sequences of a gene of interest. For example, the sequence or nucleotides of the target gene can include coding sequences and non-coding sequences. The sequence of the target gene can include exons or introns. The sequence of the target gene can include a regulatory region, including a promoter, an enhancer, a terminator, a 5' or 3' untranslated region. In some embodiments, the sequence of the target gene includes a long-range enhancer sequence.
[0232] Epigenetic editing systems as described herein can include any polynucleotide binding domain. In some embodiments, the nucleic acid binding domain includes one or more DNA binding proteins, such as zinc finger proteins (ZFPs) or transcription activator-like effectors (TALEs). In some embodiments, the nucleic acid binding domain includes a polynucleotide-guided DNA binding protein, for example, a CRISPR-Cas protein with no nuclease activity guided by a guide RNA.
[0233] The nucleic acid binding domain of the epigenetic editing system described herein can be able to recognize and bind to any gene of interest, for example, a target gene associated with a disease or disorder. In some embodiments, the target gene associated with the disease or disorder comprises a mutation compared to the wild-type gene. In some embodiments, the target gene associated with the disease or disorder comprises a copy containing a mutation associated with the disease or disorder. In some embodiments, the target gene associated with the disease or disorder has one or two copies of a wild-type DNA sequence.
[0234] The DNA binding domain can be modular and / or programmable. In some embodiments, the DNA binding domain comprises a zinc finger domain, a transcription activator-like effector (TALE) domain, a meganuclease DNA binding domain, or a polynucleotide-guided nucleic acid binding domain. Examples of DNA binding domains can be found in U.S. Patent No. 11,162,114, which is incorporated by reference in its entirety.
[0235] Transcription activator-like effectors (TALEs) can be engineered to bind to virtually any desired DNA sequence. Methods for programming TALEs are familiar to those skilled in the art. For example, such methods are described in Carroll et al., Genetics Society of America, 188 (4): 773-782, 2011; Miller et al., Nature Biotechnology 25 (7): 778-785, 2007; Christian et al., Genetics 186 (2): 757-61, 2008; Li et al., Nucleic Acids Res. 39 (1): 359-372, 2010; and Moscou et al., Science 326 (5959): 1501, 2009, each of which is incorporated herein by reference.
[0236] The DNA binding domain can be guided by a nucleic acid sequence (e.g., an RNA sequence) to identify a target gene. In some embodiments, the DNA binding domain comprises a programmable nuclease. In some embodiments, the DNA binding domain comprises a programmable nuclease with reduced or eliminated nuclease activity. For example, a programmable nuclease may contain one or two mutations in its catalytic domain that inactivates the nuclease but maintains the DNA binding activity of the nuclease. In some embodiments, the DNA binding domain comprises a CRISPR-Cas protein domain. In some embodiments, the CRISPR-Cas protein domain lacks nuclease activity or has reduced nuclease activity.
[0237] In some embodiments, the epigenetic editing system provided herein includes a Cas protein, such as a Cas9 protein domain. The Cas9 domain can be any Cas9 domain or Cas9 protein provided herein (e.g., a nuclease-inactive Cas9 or Cas9 nickase, or a Cas9 variant from any species). In some embodiments, any Cas domain or Cas protein provided herein can be fused with one or more effector protein domains as described herein. In some embodiments, any Cas protein domain provided herein can be fused with two or more effector protein domains as described herein. Cas9 can refer to a polypeptide having at least about 50%, 60%, 70%, 80%, 90%, 100% sequence identity and / or sequence similarity with a wild-type exemplary Cas9 polypeptide (e.g., from Streptococcus pyogenes). Cas9 can refer to a wild-type or modified form of a Cas9 protein, which can include amino acid changes, such as deletions, insertions, substitutions, variants, mutations, fusions, chimeric or any combination thereof.
[0238] Cas9 sequences and structures of variant Cas9 orthologs have been described in various species. Exemplary species from which the Cas9 protein or other components may be derived include, but are not limited to, Streptococcus pyogenes, Streptococcus thermophilus, Streptococcus sp.), Staphylococcus aureus, Listeria innocua, Lactobacillus gasseri, Francisella novicida, Wolinella succinogenes, Sutterella wadsworthensis, Gamma proteobacterium, Neisseria meningitidis, Campylobacter jejuni, Pasteurella multocida, Fibrobacter succinogene, Rhodospirillum rubrum, Nocardiopsis dassonvillei, Streptomyces pristinaespiralis, Streptomyces viride viridochromogenes, Streptomyces viridochromogenes, Streptosporangium roseum, Alicyclobacillus acidocaldarius, Bacillus pseudomycoides, Bacillus selenitireducens, Exiguobacterium sibiricum, Lactobacillus delbrueckii, Lactobacillus salivarius, Lactobacillus buchneri, Treponema denticola, Microscilla marina, Burkholderiales bacterium, Polar omonas naphthalenivorans, Polar omonas sp.), marine nitrogen-fixing cyanobacteria (Crocosphaera watsonii), Cyanothece sp., Microcystis aeruginosa, Synechococcus sp., Acetohalobium arabaticum, Ammonifex degensii, Caldicelulosiruptor becscii, Candidatus Desulforudis, Clostridium botulinum, Clostridium difficile, Finegoldia magna, Natranaerobius thermophilus, Pelotomaculum thermopropionium, Acidithiobacillus caldus), Acidithiobacillus ferrooxidans, Allochromatium vinosum, Marinobacter sp., Nitrosococcus halophilus, Nitrosococcus watsoni, Pseudoalteromonas haloplanktis, Ktedonobacter racemifer, Methanohalobium evestigatum, Anabaena variabilis, Nodularia spumigena, Nostoc sp., Arthrospira maxima, Arthrospira platensis, Arthrospira sp., Lyngbya sp.), Microcoleuschthonoplastes, Oscillator sp.), Petrotoga mobilis, Thermosipho africanus, Streptococcus pasteurianus, Neisseria cinerea, Campylobacter lari, Parvibaculum lavamentivorans, Corynebacterium diphtheria, or Acaryochloris marina. In some embodiments, the Cas9 protein is from Streptococcus pyogenes. In some embodiments, the Cas9 protein may be from Streptococcus thermophilus. In some embodiments, the Cas9 protein is from Staphylococcus aureus. .
[0239] Additional suitable Cas9 proteins, orthologs, variants, including nuclease-inactive variants, and sequences will be apparent to those of skill in the art based on this disclosure, and such Cas9 nucleases and sequences include Cas9 sequences from organisms and loci disclosed in Chylinski et al., (2013) RNA Biology 10:5, 726-737, which is incorporated herein by reference.
[0240] Epigenetic editing systems can include nuclease-inactive Cas9 domains (dead Cas9 or dCas9). Compared with wild-type Cas9, dCas9 protein domains can include one, two or more mutations that eliminate its nuclease activity, but retain DNA binding activity. For example, it is known that the DNA cleavage domain of Cas9 includes two subdomains, the HNH nuclease subdomain and the RuvC1 subdomain. The HNH subdomain cuts the strand complementary to the gRNA, while the RuvC1 subdomain cuts the non-complementary strand. Mutations in these subdomains can silence the nuclease activity of Cas9. For example, mutations D10A and H840A completely inactivate the nuclease activity of Streptococcus pyogenes Cas9. In some embodiments, dCas9 includes at least one mutation in the HNH subdomain and the RuvC subdomain, which reduces or eliminates nuclease activity. In some embodiments, dCas9 includes only the RuvC subdomain. In some embodiments, dCas9 includes only the HNR subdomain. It should be understood that any mutation that inactivates the RuvC or HNH domain can be included in dCas9, for example, an insertion, deletion, or single or multiple amino acid substitutions in the RuvC domain and / or the HNH domain.
[0241] Based on the present disclosure and the knowledge in the art, other suitable mutations that inactivate Cas9 will be apparent to those skilled in the art and are within the scope of the present disclosure. In addition, other exemplary suitable nuclease-inactive Cas9 domains include, but are not limited to, D839A, N863A, and / or K603R. Cas9, dCas9, or Cas9 variants also encompass Cas9, dCas9, or Cas9 variants from any organism. It should also be understood that, according to the present disclosure, dCas9, Cas9 nickase, or other suitable Cas9 variants from any organism can be used.
[0242] In some embodiments, the epigenetic editing system comprises a high-fidelity Cas9 domain. For example, a high-fidelity Cas9 domain comprising one or more mutations that reduce the electrostatic interaction between the Cas9 domain and the sugar-phosphate backbone of DNA can be incorporated into the epigenetic editing system to confer increased target binding specificity compared to the corresponding wild-type Cas9 domain. Without wishing to be bound by any particular theory, a high-fidelity Cas9 domain with reduced electrostatic interaction with the sugar-phosphate backbone of DNA may have less off-target effects. In some embodiments, the Cas9 domain comprises one or more mutations that reduce the association between the Cas9 domain and the sugar-phosphate backbone of DNA. In some embodiments, the Cas9 domain comprises one or more mutations that reduce the association between the Cas9 domain and the sugar-phosphate backbone of the DNA by at least 1%, at least 2%, at least 3%, at least 4%, at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70% or more. In some embodiments, the high-fidelity Cas9 domain comprises one or more of the N497X, R661X, Q695X, and / or Q926X mutations numbered as in the wild-type Cas9 amino acid sequence Uniprot reference sequence Q99ZW2, or the corresponding amino acids in another Cas9, wherein X is any amino acid. In some embodiments, the high-fidelity Cas9 domain comprises one or more of the N497A, R661A, Q695A and / or Q926A mutations of the amino acid sequence provided in the wild-type Cas9 sequence, or the corresponding mutations numbered in the wild-type Cas9 amino acid sequence Uniprot reference sequence Q99ZW2 or the corresponding amino acids in another Cas9. It should be understood that any epigenetic editing system provided herein, for example, any epigenetic activator or repressor provided herein, can be converted into a high-fidelity epigenetic editing system by modifying the Cas9 domain as described. In a preferred embodiment, the high-fidelity Cas9 domain is a nuclease-inactive Cas9 domain.
[0243] In some embodiments, the DNA binding domain in the epigenetic editing system is a CRISPR protein that recognizes a protospacer adjacent motif (PAM) sequence in a target gene. The CRISPR protein can recognize a naturally occurring or canonical PAM sequence, or can have a changed PAM specificity. The Cas9 domains that bind to non-canonical PAM sequences have been described in the art and will be apparent to the technician. For example, the Cas9 domains that bind to non-canonical PAM sequences have been described in Kleinstiver, BP et al., "Engineered CRISPR-Cas9 nucleases with altered PAM specificities" Nature 523, 481-485 (2015); and Kleinstiver, BP et al., "Broadening the targeting range of Staphylococcus aureus CRISPR-Cas9 by modifying PAM recognition" Nature Biotechnology 33, 1293-1298 (2015), the entire contents of each document are hereby incorporated by reference.
[0244] In some embodiments, the Cas9 domain is a Cas9 domain (SpCas9) from Streptococcus pyogenes. In some embodiments, SpCas9 recognizes a canonical NGG PAM sequence, wherein the "N" in "NGG" is adenine (A), thymine (T), guanine (G) or cytosine (C), and G is guanine. In some embodiments, the epigenetic editing system or fusion protein provided herein comprises a SpCas9 domain, which is capable of binding to a nucleotide sequence that does not comprise a canonical (e.g., NGG) PAM sequence. In some embodiments, the SpCas9 domain, the SpCas9 domain inactive for nuclease, or the SpCas9 nickase domain can be bound to a nucleic acid sequence having a NGG, NGA or NGCG PAM sequence. In some embodiments, the Cas9 domain is a modified SpCas9 domain specific for a 5'-NGCG-3' PAM sequence, wherein N is any one of nucleotides A, G, C or T. In some embodiments, the Cas9 domain is a modified SpCas9 domain specific for a 5'-NGAN-3' or 5-NGNG-3' PAM sequence, wherein N is any one of the nucleotides A, G, C, or T. In some embodiments, the Cas9 domain is a modified SpCas9 domain specific for a 5'-NGN-3' PAM sequence, wherein N is any one of the nucleotides A, G, C, or T. In some embodiments, the Cas9 domain is a modified SpCas9 domain specific for a 5'-NRN-3' or 5'-NYN-3' PAM sequence, wherein N is any one of the nucleotides A, G, C, or T, wherein R is a nucleotide A or G, and wherein Y is a nucleotide C or T.
[0245] In some embodiments, the Cas9 domain is a Cas9 domain from Staphylococcus aureus (SaCas9). In some embodiments, the SaCas9 domain is a nuclease-inactive SaCas9 (dSacas9). In some embodiments, the SaCas9 domain, the nuclease-inactive SaCas9 domain, or the SaCas9 nickase domain can bind to a nucleic acid sequence having a non-canonical PAM. In some embodiments, the SaCas9 domain, the SaCas9d domain, or the SaCas9n domain can bind to a nucleic acid sequence having a NNGRRT The nucleic acid sequence of the PAM sequence, wherein N=A, T, C or G, and R=A or G. In some embodiments, the Cas9 domain is a Cas9 domain from Neisseria meningitidis (NmeCas9). In some embodiments, the NmeCas9 domain is a nuclease-inactive NmeCas9 ((dNmeCas9). NmeCas9 can be specific for 5'-NNNGATT-3' PAM, wherein N is any one of the nucleotides A, G, C or T. In some embodiments, the Cas9 domain is a Cas9 domain from Campylobacter jejuni (CjCas9). In some embodiments, the CjCas9 domain is a nuclease-inactive CjCas9 (dCjCas9). Cj Cas9 can be specific for 5'-NNNVRYM-3'PAM, wherein N is any one of the nucleotides A, G, C or T, V is a nucleotide A, C or G, R is a nucleotide A or G, Y is a nucleotide C or T, and M is a nucleotide A or C. In some embodiments, the Cas9 domain is a Cas9 domain from Streptococcus thermophilus (StCas9). In some embodiments, StCas9 is encoded by the St CRISPR1 locus (St1Cas9) of Streptococcus thermophilus. In some embodiments, the St1Cas9 domain is a nuclease-inactive St1Cas9 ((dSt1Cas9). St1Cas9 can be specific for 5'-NNAGAAW-3'PAM, wherein N is any one of the nucleotides A, G, C or T, and W is a nucleotide A or T. In some embodiments, StCas9 is encoded by the St CRISPR1 locus (St1Cas9) of Streptococcus thermophilus. The CRISPR3 locus (St3Cas9) is encoded. In some embodiments, the St3Cas9 domain is a nuclease-inactive St3Cas9 (dSt3Cas9). St3Cas9 can be specific for 5'-NGGNG-3'PAM, wherein N is any one of the nucleotides A, G, C or T.
[0246] In some embodiments, the Cas9 domain of any fusion protein provided herein comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any of the Cas9 sequences provided herein.
[0247] In some embodiments, the epigenetic editing system provided herein comprises a Cpf1 (or Cas12a) protein domain. For example, the epigenetic editing system can comprise a nuclease-inactive Cpf1 protein or a variant thereof. The Cpf1 protein has a RuvC-like nuclease domain similar to the RuvC domain of Cas9, but does not have a HNH nuclease domain, and the N-terminus of Cpf1 does not have the α-helical recognition lobe of Cas9.
[0248] In some embodiments, the Cpf1 protein comprises an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to the FnCpf1 sequence provided herein. It should be understood that Cpf1 from other bacterial species may also be used according to the present invention.
[0249] In some embodiments, Cpf1 is a Cpf1 protein from Lachnospiraceae bacterium (LbCpf1). LbCpf1 can be specific for a 5'-TTTV-3' PAM sequence, wherein V is any one of the nucleotides A, G, or C. In some embodiments, the LbCpf1 protein has reduced nuclease activity. In some embodiments, the nuclease activity of the LbCpf1 protein is eliminated (dLbCpf1). In some embodiments, Cpf1 is a Cpf1 protein from Acidaminococcus sp. (AsCpf1). AsCpf1 can be specific for a 5'-TTTV-3' PAM sequence, wherein V is any one of the nucleotides A, G, or C. In some embodiments, the AsCpf1 protein has reduced nuclease activity. In some embodiments, the nuclease activity of the AsCpf1 protein is eliminated (dAsCpf10). In some embodiments, the dAsCpf1 or AsCpf1 protein further comprises a mutation that improves the fidelity of target recognition by the protein. In some embodiments, the dAsCpf1 or AsCpf1 protein further comprises a mutation that results in an altered PAM specificity of the protein.
[0250] In some embodiments, the epigenetic editing system provided herein comprises a Cas protein domain other than Cas9. In some embodiments, the Cas9 protein comprises an inactivated nuclease domain. In some embodiments, the epigenetic editing system comprises a Cas12a, Cas12b, Cas12c, Cas12d, Cas12e, Cas12h or Cas12i domain. In some embodiments, the Cas9 protein is an RNA nuclease or an inactivated RNA nuclease. In some embodiments, the epigenetic editing system comprises a Cas12g, Cas13a, Cas13b, Cas13c or Cas13d domain. In some embodiments, the epigenetic editing system comprises an Argonaut protein domain.
[0251] The CRISPR / Cas system or Cas protein in the epigenetic editing system provided herein may include Class 1 or Class 2 Cas proteins. The nuclease activity of Class 1 or Class 2 proteins used in the epigenetic editing system may be inactivated. In some embodiments, the epigenetic editing system comprises a Cas protein derived from a type II, type IIA, type IIB, type IIC, type V, or type VI Cas nuclease. In some embodiments, the epigenetic editing system comprises a Cas protein derived from a Class 2 Cas nuclease, and the Class 2 Cas nuclease is derived from Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas10, Cas14a, Cas14b, Cas14c, CasX, CasY, CasPhi, C2c4, C2c8, C2c9, C2c10, Csy1, Csy2, Csy 3, Cse1, Cse2, Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx1S, Csf1, Csf2, CsO, Csf4 or a homologue or modified form thereof. In some embodiments, the Cas protein in the epigenetic editing system is a nuclease-inactivated Cas protein.
[0252] In some embodiments, the epigenetic editing system comprises a CasX (Cas12e) protein. The CasX protein can be specific for a 5'-TTCN-3' PAM sequence, wherein N is any of the nucleotides A, G, T, or C. In some embodiments, the CasX protein has reduced or eliminated nuclease activity (dCasX). In some embodiments, the epigenetic editing system comprises a CasY (Cas12d) protein. The CasY protein can be specific for a 5'-TA-3' PAM sequence. In some embodiments, the CasY protein has reduced or eliminated nuclease activity (dCasY). In some embodiments, the epigenetic editing system comprises (CasPhi) protein. The protein may be specific for a 5'-TTN-3' PAM sequence, wherein N is any one of the nucleotides A, T, G, or C. In some embodiments, Proteins with reduced or eliminated nuclease activity
[0253] In some embodiments, the Cas protein is a circular permutant Cas protein. For example, the epigenetic editing system may include a circular permutant Cas9 as described in Oakes et al., Cell 176, 254-267 (2019) (incorporated herein in its entirety). As used herein, the term "circular permutant" refers to a variant polypeptide (e.g., a variant polypeptide of a subject Cas protein), wherein a segment of the primary amino acid sequence has been moved to a different position within the primary amino acid sequence of the polypeptide, but wherein the local order of the amino acids has not changed, and wherein the three-dimensional structure of the protein is conserved. For example, a circular permutant of a wild-type 1000 amino acid polypeptide may have an N-terminal residue numbered 500 (relative to a wild-type protein), wherein residues 1-499 of the wild-type protein are added to the C-terminus. Relative to the wild-type protein sequence, such a circular permutant will have amino acid numbers 500-1000 from the N-terminus to the C-terminus, followed by 1-499, producing a circular permutant protein with amino acid 499 as a C-terminal residue. Thus, such exemplary circular arrays will have the same total number of amino acids as the wild-type reference protein, and the amino acids will be in the same order locally in specific regions of the circular array, but the overall primary amino acid sequence is altered.
[0254] In some embodiments, the epigenetic editing system comprises a circularly arranged Cas protein, such as a circularly arranged Cas9 protein. In some embodiments, the epigenetic editing system comprises a fusion of a circularly arranged Cas protein and an epigenetic effector domain, wherein the epigenetic effector domain is fused to the circularly arranged Cas protein at an N-terminus or C-terminus different from the N-terminus or C-terminus of the wild-type Cas protein.
[0255] In some embodiments, the Cas protein of the circular arrangement comprises the N-terminus of the N-terminal fragment of the wild-type Cas protein fused to the C-terminus of the C-terminal fragment of the wild-type Cas protein, thereby generating a new N-terminus and C-terminus. Without wishing to be bound by any theory, the N-terminus and C-terminus of the wild-type Cas protein may be locked in a small area, which may cause steric hindrance when the Cas protein is fused to the effector domain, and reduce access to the target DNA sequence. In some embodiments, compared with the epigenetic editing system comprising the wild-type Cas protein counterpart, the epigenetic editing system comprising the circular arrangement body Cas protein has reduced spatial incompatibility. In some embodiments, compared with the epigenetic editing system comprising the wild-type Cas protein counterpart, the epigenetic editing system comprising the circular arrangement body Cas protein has improved effectiveness. In some embodiments, compared with the epigenetic editing system comprising the wild-type Cas protein counterpart, the epigenetic editing system comprising the circular arrangement body Cas protein has improved epigenetic editing accuracy. In some embodiments, an epigenetic editing system comprising a circular permutator Cas protein has reduced off-target editing effects compared to an epigenetic editing system comprising a wild-type Cas protein counterpart.
[0256] Guide polynucleotide
[0257] In some embodiments, the epigenetic editing system comprises a guide polynucleotide (or guide nucleic acid). For example, an epigenetic editing system having a DNA binding domain comprising a CRISPR-Cas protein can also include a guide nucleic acid capable of forming a complex with the CRISPR-Cas protein.
[0258] Methods for site-specific DNA targeting (e.g., to modify a genome) using a guide nucleotide sequence programmable DNA binding protein (e.g., Cas9) are known in the art. A guide RNA (gRNA) can guide a programmable DNA binding protein (e.g., a Class 2 Cas protein, such as Cas9) to a target sequence on a target nucleic acid molecule, wherein the gRNA hybridizes with the programmable DNA binding protein and produces a modification at or near the target sequence. In some embodiments, the gRNA and epigenetic editing system fusion protein can form a ribonucleoprotein (RNP), such as a CRISPR / Cas complex.
[0259] A guide nucleotide sequence, such as a guide RNA sequence, may include two parts: 1) a nucleotide sequence having homology to a target nucleic acid (e.g., and guiding the binding of a guide nucleotide sequence programmable DNA binding protein to a target); and 2) a nucleotide sequence of a programmable DNA binding protein (e.g., a CRISPR-Cas protein) guided by a binding nucleic acid. The nucleotide sequence in 1) may include a spacer sequence that hybridizes with a target sequence. The nucleotide sequence in 2) may be referred to as a scaffold sequence of a guide nucleic acid, a tracrRNA, or an activation region of a guide nucleic acid, and may include a stem-loop structure. The scaffold sequence of a guide nucleic acid as described in Jinek et al., Science 337:816-821 (2012), U.S. Patent Application Publication US20160208288, and U.S. Patent Application Publication US20160200779 are incorporated herein by reference in their entirety.
[0260] The guide polynucleotide can be a single molecule, or can comprise two independent molecules. For example, parts 1) and 2) as described above can be fused to form a single guide sequence (e.g., a single guide RNA or sgRNA), or can be two independent molecules. In some embodiments, the guide polynucleotide is a double polynucleotide connected by a joint. In some embodiments, the guide polynucleotide is a double polynucleotide connected by a non-nucleic acid joint, such as a peptide joint or a chemical joint.
[0261] Methods for selecting, designing and verifying gRNA and targeting sequences (or spacer sequences) are described herein and are known to those skilled in the art. Software tools can be used to optimize the gRNA corresponding to the target nucleic acid sequence, for example, to minimize the total off-target activity of the genome. For example, a DNA sequence search algorithm can be used to identify the target sequence in the crRNA of the gRNA used with Cas9. Exemplary gRNA design tools, including such as Bae et al., Cas-OFFinder:A fast and versatile algorithm that searches for potential off-target sites of Cas9 RNA-guided endonucleases.Bioinformatics 30,1473-1475 (2014)) gRNA design tools described therein are incorporated herein in their entirety.
[0262] The guiding polynucleotide can have different lengths. In some embodiments, the length of the spacer or targeting sequence depends on the CRISPR / Cas components of the epigenetic editing system and the components used. For example, different Cas proteins from different bacterial species have different optimal targeting sequence lengths. Therefore, the spacer sequence can include 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more than 50 nucleotide lengths. In some embodiments, the spacer includes 18-24 nucleotide lengths. In some embodiments, the spacer includes 19-21 nucleotide lengths. In some embodiments, the spacer sequence includes 20 nucleotide lengths. In some embodiments, the guide nucleic acid (e.g., guide RNA) is 15-100 nucleotides long and comprises a sequence of at least 10 consecutive nucleotides complementary to the target sequence. In some embodiments, the guide RNA is 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 nucleotides in length. In some embodiments, the guide RNA comprises a sequence of 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49 or 50 consecutive nucleotides, which are complementary to the target sequence. In some embodiments, the target sequence is a DNA sequence. In some embodiments, the degree of complementarity between the targeting sequence of the gRNA and the target sequence on the target nucleic acid molecule is at least about 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99% or 100%. In some embodiments, the targeting sequence of the gRNA and the target sequence on the target nucleic acid molecule can be 100% complementary. In other embodiments, the targeting sequence of the gRNA and the target sequence on the target nucleic acid molecule can include at least one mismatch. For example, the targeting sequence of the gRNA and the target sequence on the target nucleic acid molecule can include 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 mismatches.
[0263] In some embodiments, the target sequence is a sequence in a mammalian genome. In some embodiments, the target sequence is a sequence in a human genome. In some embodiments, the 3' end of the target sequence is adjacent to a canonical PAM sequence (NGG). In some embodiments, the guide nucleic acid (e.g., guide RNA) is complementary to a sequence associated with a disease or disorder.
[0264] In some embodiments, the guide RNA is truncated. The truncation can include any number of nucleotide deletions. For example, the truncation can include 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 40, 50 or more nucleotides. In some embodiments, the guide polynucleotide comprises RNA. In some embodiments, the guide polynucleotide comprises DNA. In some embodiments, the guide polynucleotide comprises a mixture of DNA and RNA.
[0265] The guide polynucleotide can be modified. The modification can include chemical changes, synthetic modifications, nucleotide additions and / or nucleotide reductions. Modified nucleosides or nucleotides can be present in gRNA. For example, gRNA can include one or more non-natural and / or naturally occurring components or configurations, which replace or use in addition to classical A, G, C and U residues. The modified RNA can include one or more changes or replacements of one or two non-connected phosphate oxygens and / or one or more connected phosphate oxygens in the phosphodiester backbone bond, changes in ribose, such as changes in 2' hydroxyl groups on ribose (exemplary sugar modifications), changes in phosphate moieties, modifications or replacements of naturally occurring nucleobases, replacements or modifications of ribose-phosphate backbones, modifications of 3' or 5' ends of oligonucleotides, or replacements of terminal phosphate groups, or conjugation of moieties, caps or linkers, or any combination thereof.
[0266] In some embodiments, the ribose group (or sugar) can be modified. In some embodiments, the modified ribose group can control the oligonucleotide binding affinity of the complementary chain, duplex formation or interaction with nucleases. Examples of chemical modifications to the ribose group include, but are not limited to, 2'-O-methyl (2'-OMe), 2'-fluoro (2'-F), 2'-deoxy, 2'-O-(2-methoxyethyl) (2'-MOE), 2'-NH2, 2'-O-allyl, 2'-O-ethylamine, 2'-O-cyanoethyl, 2'-O-acetyl ester, or bicyclic nucleotides, such as locked nucleic acids (LNA), 2'-(5-constrained ethyl (S-cEt)), constrained MOE, or 2'-0,4'-C-aminomethylene bridged nucleic acids (2',4'-BNANC). In some embodiments, 2'-O-methyl modifications can increase the binding affinity of oligonucleotides. In some embodiments, 2'-O-methyl modifications can enhance the nuclease stability of an oligonucleotide. In some embodiments, 2'-fluoro modifications can increase oligonucleotide binding affinity and nuclease stability.
[0267] In some embodiments, the phosphate group can be chemically modified. The example of chemical modification to the phosphate group includes but is not limited to phosphorothioate (PS), phosphonoacetate (PACE), thiophosphonoacetate (thioPACE), amide, triazole, phosphonate or phosphotriester modification. In some embodiments, the PS key can refer to a key in which sulfur replaces a non-bridging phosphate oxygen in a phosphodiester bond (e.g., between nucleotides). "s" can be used to describe the PS modification in the gRNA sequence. In some embodiments, gRNA or sgRNA can include phosphorothioate (PS) keys at the 5' end or 3' end. In some embodiments, gRNA or sgRNA can include phosphorothioate (PS) keys at the 5' end. In some embodiments, gRNA or sgRNA can include phosphorothioate (PS) keys at the 3' end. In some embodiments, gRNA or sgRNA can include phosphorothioate (PS) keys at the 5' end and at the 3' end. In some embodiments, the gRNA or sgRNA may include one, two, three or more than three phosphorothioate bonds at the 5' end or at the 3' end. In some embodiments, the gRNA or sgRNA may include three phosphorothioate (PS) bonds at the 5' end or at the 3' end. In some embodiments, the gRNA or sgRNA may include three phosphorothioate bonds at the 3' end. In some embodiments, the gRNA or sgRNA may include two and no more than two (i.e., only two) consecutive phosphorothioate (PS) bonds at the 5' end or at the 3' end. In some embodiments, the gRNA or sgRNA may include three consecutive phosphorothioate (PS) bonds at the 5' end or at the 3' end. In some embodiments, the gRNA or sgRNA may include the sequence 5'-UsUsU-3' at the 3' end or at the 5' end, wherein U represents uridine, and wherein s represents a phosphorothioate (PS) bond.
[0268] In some embodiments, the nucleobase can be chemically modified. Examples of chemical modifications to the nucleobase include, but are not limited to, 2-thiouridine, 4-thiouridine, N6-methyladenosine, pseudouridine, 2,6-diaminopurine, inosine, thymidine, 5-methylcytosine, 5-substituted pyrimidines, isoguanine, isocytosine, or halogenated aromatic groups.
[0269] Chemical modifications can be performed on a portion of the guide polynucleotide or the entire guide polynucleotide. In some embodiments, a total of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 base pairs of the guide RNA are chemically modified. In some embodiments, a total of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 base pairs of the guide RNA are chemically modified. In some embodiments, a total of 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105 , 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149 or 150 base pairs are chemically modified. The chemical modification can be in the protospacer, tracr RNA, crRNA, stem loop or any combination thereof.
[0270] Zinc finger proteins
[0271] In some embodiments, the epigenetic editing system described herein comprises a nucleic acid binding domain comprising a zinc finger domain.
[0272] Zinc finger proteins are DNA binding proteins comprising one or more zinc fingers. In some embodiments, zinc fingers (ZF) comprise a relatively small polypeptide domain comprising approximately 30 amino acids. Zinc fingers may comprise an α-helix (referred to as a ββα fold) adjacent to an antiparallel β fold, which may coordinate with a zinc ion between four Cys and / or His residues, as further described below. In some embodiments, the ZF domain recognizes and binds to a nucleic acid triplet or overlapping quadruplex in a double-stranded DNA target sequence. In certain embodiments, ZF may also bind to RNA and proteins.
[0273] As used herein, the term "zinc finger" (ZF) or "zinc finger motif" (ZF motif) refers to an individual "finger" that comprises a β-β-α (ββα)-protein fold stabilized by a zinc ion, as described elsewhere herein. In some embodiments, each finger comprises approximately 30 amino acids. In some embodiments, a ZF protein or ZF protein domain is a protein motif comprising a plurality of fingers or finger-like protrusions that contact its target molecule in series. For example, a ZF finger can bind to a triplet or (overlapping) quadruplex nucleotide sequence. Thus, a tandem array of ZF fingers can be designed for a naturally non-existent ZF protein to bind to a desired target.
[0274] Zinc finger proteins are widely distributed in eukaryotic cells. An exemplary motif characterizing one class of these proteins (the C2H2 class) is -Cys-(X)2-4-Cys-(X)12-His-(X)3-5His, where X is any amino acid. A single finger domain can be about 30 amino acids in length. In some embodiments, a single finger comprises an alpha helix comprising two invariant histidine residues coordinated by zinc to two cysteines of a single beta turn.
[0275] In some embodiments, the amino acid sequence of a zinc finger protein (e.g., Zif268 protein) can be altered by making amino acid substitutions at helical positions on the zinc finger recognition helix (e.g., positions 1, 2, 3, and 6 of Zif268). For example, modified zinc fingers with non-naturally occurring DNA recognition specificity can be generated by phage display and combinatorial libraries with randomized side chains in the first or middle finger of Zif268, and then isolated with altered Zif268 binding sites, in which the appropriate DNA subsite is replaced by an altered DNA triplet.
[0276] In some embodiments, the zinc finger comprises a C2H2 finger. In some embodiments, the zinc finger protein comprises a ZF array comprising a continuous C2H2-ZF, each of which contacts three or more consecutive bases. In some embodiments, the zinc finger protein structure bound to DNA, such as the zinc finger protein Zif268 and its variants, shows a semi-conservative interaction pattern, in which three amino acids from the α-helix of the zinc finger typically contact three adjacent base pairs in the DNA. Therefore, in embodiments, the zinc finger DNA binding domain functions in a modular manner, in which there is a one-to-one interaction between the zinc finger and the three-base pair trinucleotide sequence in the DNA sequence.
[0277] In some embodiments, the epigenetic editing system comprises a zinc finger motif comprising the sequence: N'--(helix 1)--(helix 2)--(helix 3)--(helix 4)--(helix 5)--(helix 6)--C', wherein (helix) is a six consecutive amino acid residue peptide forming a short alpha helix. In some embodiments, the epigenetic editing system comprises a zinc finger motif comprising the sequence: N'--(helix 1)--(helix 2)--(helix 3)--(helix 4)--(helix 5)--C', wherein (helix) is a six consecutive amino acid residue peptide forming a short alpha helix.
[0278] In some embodiments, two or more zinc fingers are connected together in a tandem array to achieve specific recognition and binding to a continuous DNA sequence. The zinc fingers or zinc finger arrays in the epigenetic editing system can be naturally occurring or can be artificially designed for the desired DNA binding specificity. For example, the DNA binding properties of a single zinc finger can be designed by randomizing the amino acids at the α-helical positions of the zinc fingers involved in DNA binding and using selection methods such as phage display to identify the desired variants that can bind to the DNA target site of interest.
[0279] Compared to naturally occurring zinc finger proteins, engineered zinc finger binding domains can have new binding specificities. Zinc fingers with desired DNA binding specificity can be designed and selected via various methods. For example, a database comprising triplet (or quadruple) nucleotide sequences and separate zinc finger amino acid sequences (wherein each triplet or quadruple nucleotide sequence is associated with one or more amino acid sequences of zinc fingers that bind to a specific triplet or quadruple sequence) can be used to design zinc finger arrays for specific DNA sequences. See, for example, U.S. Patents Nos. 6,453,242, 6,534,261, and 8,772,453, which are incorporated herein by reference in their entirety. In some embodiments, zinc finger arrays can be designed and selected from zinc finger libraries (e.g., randomized zinc finger libraries). In some embodiments, zinc fingers with new DNA binding specificity are generated on combinatorial libraries by selection-based methods. For example, zinc fingers can be selected using phage display, which involves displaying zinc finger proteins on the surface of filamentous phage, followed by successive rounds of affinity selection with biotinylated target DNA to enrich for phage-expressed proteins that can bind to a specific target sequence. The bacterial two-hybrid (B2H) system can also be used to select zinc fingers that bind to a specific target site from a randomized library. For example, the zinc finger binding site can be located upstream of a weak promoter that drives the expression of two selective markers in a host cell (e.g., an E. coli cell). A zinc finger library fused to a fragment of a reporter protein (e.g., yeast Gal11P protein) can be expressed in a cell, and the binding of the zinc finger to the target site recruits an RNA polymerase-Gal4 fusion, thereby activating transcription and allowing the cell to survive on a selective medium. Zinc fingers are rationally designed and selected as described in Maeder et al., 2008, Mol. Cell, 31:294-301; Joung et al., 2010, Nat. Methods, 7:91-92; Isalan et al., 2001, Nat. Biotechnol., 19:656-660, Rebar et al., Science 263, 671-673 (1994), and Joung et al., Proc Natl Acad Sci USA 97, 7382-7387 (2000), each of which is incorporated herein by reference in its entirety.
[0280] In some embodiments, zinc fingers can be evolved and selected using a continuous evolution system (PACE), which comprises host cells, such as E. coli cells, a "helper phagemid" present in all host cells and encoding all phage proteins except one phage protein (e.g., a g3p protein), a "helper plasmid" present in all host cells that expresses the g3p protein in response to an active library member; and a "selection phagemid" expressing the library of evolving proteins or nucleic acids, which is replicated and packaged into secreted phage particles. The helper plasmid and the helper plasmid can be combined into a single plasmid. New host cells can only be infected with phage particles containing g3p. Appropriate selection phagemids encoding library members that induce g3p expression from the helper plasmid can be packaged into phage particles containing g3p. Phage particles containing g3p can infect new cells, resulting in further replication of the appropriate selection phagemid, while g3p-deficient phage particles are non-infectious, so low-fitness selection phagemids cannot reproduce. The selection system, combined with continuous flow of host cells through a lagoon that permits replication of the phagemid but not the host cells, can be used to rapidly select zinc fingers. The PACE system as described in US Pat. No. 9,023,594 is incorporated herein by reference in its entirety.
[0281] The zinc finger DNA binding domain of the epigenetic editing system may include one or more zinc fingers. For example, the zinc finger DNA binding domain may include 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12 or more zinc fingers. In some embodiments, the zinc finger DNA binding domain has at least three zinc fingers. In some embodiments, the zinc finger DNA binding domain has at least 4, 5 or 6 zinc fingers. In some embodiments, the zinc finger DNA binding domain has three zinc fingers. In some embodiments, the zinc finger DNA binding domain has at least two zinc fingers. In some embodiments, the zinc finger DNA binding domain has an array of two finger units.
[0282] The zinc finger DNA binding domain of the epigenetic editing system can be designed for optimized specificity. In some embodiments, a sequential selection strategy is used to design a multi-finger ZF domain. For example, in a multi-finger ZF domain, a first finger can be randomly displayed and selected using phage display, and a small batch of selected fingers can be brought to the next stage, where the second finger is randomized and selected. According to the number of fingers in the ZF domain, the process can be repeated many times. In some embodiments, parallel optimization is used to design a multi-finger ZF domain. For example, the B2H system can be used to interrogate the main randomization library under low selection stringency to identify various individual fingers that can bind to each 3 base pair subsite of the target site. Then three selected populations can be randomly reorganized (shuffled) to generate a library of multi-finger proteins, which can then be interrogated under high stringency selection conditions to identify three-finger proteins targeting specific nine base pair sites. In another embodiment, a large number of low stringency selections can be used to generate a master library of a single finger, from which multi-finger proteins, such as three-finger ZF proteins, can be selected. For example, a master library or archive can include a pre-selected zinc finger pool, each zinc finger pool comprising a mixture of fingers targeting different three base pair subsites of a DNA sequence at a defined position within a three-finger ZF protein. In certain embodiments, the zinc finger archive comprises at least 192 finger pools (64 potential three bp target subsites for each position in the three-finger protein). In some embodiments, the zinc finger archive comprises at least one zinc finger pool comprising at least 10, 20, 30, 40, 50, 60, 70, 80, 90, 95, 100 or more different fingers. In some embodiments, a smaller library is created from the archive for interrogation with a reporter system (e.g., a bacterial two-hybrid selection system).
[0283] In some embodiments, two complementary libraries can be used to design and select multi-finger ZF domains, such as three-finger ZF domains. For example, two pre-made zinc finger phage display libraries can be used to design a three-finger ZF domain, wherein the first library contains randomized DNA binding amino acid positions in finger 1 and finger 2, and the second library contains randomized DNA binding amino acid positions in finger 2 and finger 3. The two libraries are complementary because the first library contains randomization in all base contact positions of finger 1 and certain base contact positions of finger 2, while the second library contains randomization in the remaining base contact positions of finger 2 and all base contact positions of finger 3. The selection of "one and a half" fingers from each master library can be carried out in parallel using DNA sequences, in which five nucleotides have been fixed to the sequence of interest. Subsequently, the zinc finger coding sequence can be amplified from the recovered phage using PCR, and the "one and a half" fingers in groups can be paired to produce a recombinant three-finger DNA binding domain.
[0284] In some embodiments, multi-finger ZF domains can be designed based on the environmental impact of adjacent fingers. In some embodiments, multi-finger ZF domains are designed and not selected. For example, a three-finger ZF domain can be assembled using an N-terminal and C-terminal finger identified in other arrays containing a common middle finger, using a library containing an archive of three-finger ZF arrays, including pre-selected and / or tested three-finger arrays.
[0285] Software for designing and selecting ZF arrays, for example, ZiFit (http: / / bindr.gdcb.iastate.edu / ZiFiT / ; http: / / www.zincfingers.org / software-tools.htm) is available and known to those skilled in the art.
[0286] Therefore, the zinc finger DNA binding domain of the epigenetic editing system may include one or more zinc fingers. For example, the zinc finger DNA binding domain may include 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12 or more zinc fingers. In some embodiments, the zinc finger DNA binding domain has at least three zinc fingers. In some embodiments, the zinc finger DNA binding domain has at least 4, 5 or 6 zinc fingers. In some embodiments, the zinc finger DNA binding domain has three zinc fingers. In some embodiments, the zinc finger DNA binding domain comprising at least three zinc fingers recognizes a target DNA sequence of 9 or 10 nucleotides. In some embodiments, the zinc finger DNA binding domain comprising at least four zinc fingers recognizes a target DNA sequence of 12 to 14 nucleotides. In some embodiments, the zinc finger DNA binding domain comprising at least six zinc fingers recognizes a target DNA sequence of 18 to 21 nucleotides.
[0287] In some embodiments, the epigenetic editing system as disclosed herein comprises non-natural, and suitably comprises 3 or more zinc fingers. In some embodiments, the epigenetic editing system comprises 4, 5, 6, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18 or more (e.g., up to about 30 or 32) zinc finger motifs arranged in series adjacent to each other, forming an array of ZF motifs. In some embodiments, the epigenetic editing system includes at least 3 ZF motifs, at least 4 ZF motifs, at least 5 ZF motifs or at least 6 ZF motifs, at least 7 ZF motifs, at least 8 ZF motifs, at least 9 ZF motifs, at least 10 ZF motifs, at least 11 or at least 12 ZF motifs in the nucleic acid binding domain. In some embodiments, the epigenetic editing system includes up to 6, 7, 8, 10, 11, 12, 16, 17, 18, 22, 23, 24, 28, 29, 30, 34, 35, 36, 40, 41, 42, 46, 47, 48, 54, 55, 56, 58, 59, or 60 ZF motifs in the nucleic acid binding domain.
[0288] In some embodiments, zinc fingers or zinc finger arrays targeting specific DNA sequences are designed using a modular assembly approach. For example, two or more pre-selected zinc fingers can be fused in tandem.
[0289] In some embodiments, the zinc finger array comprises a plurality of zinc fingers fused via peptide bonds. In some embodiments, the zinc finger array comprises a plurality of zinc fingers, wherein one or more zinc fingers are connected by a peptide linker. For example, the zinc fingers in the multi-finger array can be connected by a peptide linker of 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15 or more amino acids in length. In some embodiments, the zinc fingers in the multi-finger array are connected by a peptide linker of 5 amino acids in length. In some embodiments, the zinc fingers in the multi-finger array are connected by a peptide linker of 6 amino acids in length.
[0290] In some embodiments, a ZF-containing protein may comprise a ZF array of 2 or more ZF motifs, which may be directly adjacent to each other (i.e., separated by a short (canonical) linker sequence), or may be separated by a longer, flexible or structured polypeptide sequence. In some embodiments, directly adjacent fingers bind to contiguous nucleic acid sequences, i.e., to adjacent trinucleotides / triplets. In some embodiments, adjacent fingers cross-bind between each other's respective target triplets, which may help to strengthen or enhance recognition of the target sequence and result in binding of overlapping quadruplex sequences. In some embodiments, long-range ZF domains within the same protein may recognize (or bind to) discontinuous nucleic acid sequences, or even different molecules (e.g., proteins instead of nucleic acids).
[0291] In some embodiments, the epigenetic editing system comprises a zinc finger containing more than 3 fingers. In some embodiments, the epigenetic editing system comprises at least 6 zinc fingers in the DNA binding domain. In some embodiments, the epigenetic editing system comprises 6 zinc fingers in the DNA binding domain bound to the 18bp target sequence. In some embodiments, the 18bp target sequence is unique in the human genome. In some embodiments, the epigenetic editing system comprises zinc fingers, which include at least 7, 8, 9, 10, 11, 12, 13, 14, 15 or more zinc fingers. In some embodiments, the strong affinity of the three-finger protein will allow a subset of the longer array to bind DNA, and thus reduce specificity. Without wishing to be bound by any theory, a zinc finger protein comprising a plurality of two-finger units or three-finger units connected by an extended linker can confer higher DNA binding specificity compared to an array with fewer fingers or the same number of fingers simply connected via a peptide bond. In some embodiments, the epigenetic editing system comprises at least three two-finger units connected by a peptide linker, wherein each of the two-finger units binds to a subsite in the target DNA sequence. In some embodiments, the epigenetic editing system comprises at least four two-finger units connected by a peptide linker, wherein each of the two-finger units binds to a subsite in the target DNA sequence. In some embodiments, the epigenetic editing system comprises at least five two-finger units connected by a peptide linker, wherein each of the two-finger units binds to a subsite in the target DNA sequence. In some embodiments, the epigenetic editing system comprises at least six, seven, eight, nine, ten or more two-finger units connected by a peptide linker, wherein each of the two-finger units binds to a subsite in the target DNA sequence. In some embodiments, the epigenetic editing system comprises at least two three-finger units connected by a peptide linker, wherein each of the three finger units binds to a subsite in the target DNA sequence. In some embodiments, the epigenetic editing system comprises at least three three-finger units connected by a peptide linker, wherein each of the three finger units binds to a subsite in the target DNA sequence. In some embodiments, the epigenetic editing system comprises at least four three-finger units connected by a peptide linker, wherein each of the three finger units binds to a subsite in the target DNA sequence. In some embodiments, the epigenetic editing system comprises at least five three-finger units connected by a peptide linker, wherein each of the three finger units binds to a subsite in the target DNA sequence. In some embodiments, the epigenetic editing system comprises at least six, seven, eight, nine, ten or more three-finger units connected by a peptide linker, wherein each of the three finger units binds to a subsite in the target DNA sequence.
[0292] In some embodiments, multiple zinc fingers are assembled to target a specific DNA sequence in a target gene, each zinc finger recognizing three specific DNA nucleotides or trinucleotide "subsites". In some embodiments, such DNA subsites are continuous sequences in the target gene. In some embodiments, one or more of the DNA subsites are separated by gaps in the target gene. For example, a multi-finger ZF can recognize a DNA subsite that spans 1, 2, 3 or more base pairs between adjacent subsites. In some embodiments, the zinc fingers in the multi-finger ZF are connected via a peptide linker. The length of the peptide linker can be 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more amino acids. In some embodiments, the linker comprises 5 or more amino acids. In some embodiments, the linker comprises 7-17 amino acids. In some embodiments, the linker is a flexible linker. In some embodiments, the linker is a rigid linker, such as a linker comprising one or more proline.
[0293] Zinc finger arrays with sequence-specific DNA binding activity can be fused to functional effector domains (e.g., epigenetic effector domains as described herein) to confer epigenetic modifications to the associated histones in a DNA sequence or target gene. In some embodiments, the epigenetic editing system described herein comprises a zinc finger array that is specific to a target DNA sequence. In some embodiments, the two connectors of the zinc finger array are the same. In some embodiments, the two connectors of the zinc finger array are different.
[0294] In some embodiments, the programmable DNA binding protein comprises an argonaute protein. An example of such a nucleic acid programmable DNA binding protein is an Argonaute protein from Halobacterium glaucoma (NgAgo). NgAgo is a ssDNA-guided endonuclease. NgAgo binds to the 5' phosphorylated ssDNA of -24 nucleotides (gDNA) to guide it to its target site and will produce a DNA double-strand break at the gDNA site. In contrast to Cas9, the NgAgo-gDNA system does not require a protospacer adjacent motif (PAM). The use of nuclease-inactive NgAgo (dNgAgo) can greatly expand the bases that can be targeted. The characterization and use of NgAgo have been described in Gao et al., Nat Biotechnol., 2016 Jul; 34(7):768-73. PubMed PMID: 27136078; Swarts et al., Nature. 507(7491)(2014):258-61; and Swarts et al., Nucleic Acids Res. 43(10)(2015):5120-9, each of which is incorporated herein by reference.
[0295] In some embodiments, the nucleic acid binding domain comprises a virus-derived RNA binding domain guided by an RNA sequence to bind to a target gene. In some embodiments, the nucleic acid binding domain comprises a K homology (KH) domain, an MS2 capsid protein domain, a PP7 capsid protein domain, a SfMu Com capsid protein domain, a sterile alpha motif, a telomerase Ku binding motif and Ku protein, a telomerase Sm7 binding motif and Sm7 protein, or any other RNA recognition motif.
[0296] In some embodiments, the nucleic acid binding domain comprises an inactivated nuclease, e.g., an inactivated meganuclease. Additional non-limiting examples of DNA binding domains include tetracycline-controlled repressor (tetR) DNA binding domains, leucine zippers, helix-loop helix (HLH) domains, helix-turn-helix domains, zinc fingers, β-sheet motifs, steroid receptor motifs, bZIP domain homology domains, and AT-hook.
[0297] Target sequence
[0298] As used herein, "target polynucleotide sequence" can be a nucleic acid sequence present in a gene of interest. The target sequence can be in the genome of a cell, or can be expressed in a cell. In one aspect, the epigenetic editing system provided herein is used to bind to a target polynucleotide sequence, and to achieve epigenetic modification and / or transcriptional regulation of a target gene. For example, the target sequence can be recognized by a zinc finger array of an epigenetic editing system, or can be hybridized with a guide RNA sequence of a CRISPR protein complexed with a nuclease-inactive epigenetic editing system. In an embodiment in which the epigenetic editing system includes a gRNA-dCas-effector domain complex, the gRNA is designed to have complementarity with the target sequence (or to have identity with the relative strand (opposing strand) of the target sequence (e.g., a protospacer sequence). In some embodiments, the gRNA includes a spacer sequence that is 100% identical to the protospacer sequence in the target sequence. In some embodiments, the gRNA sequence includes a spacer sequence that is about 95%, 90%, 85%, or 80% identical to the protospacer sequence in the target sequence.
[0299] In some embodiments, the target sequence is an endogenous sequence of an endogenous gene of the host cell. In some embodiments, the target sequence is an exogenous sequence.
[0300] The target sequence can be any region of a polynucleotide (e.g., a DNA sequence) suitable for epigenetic editing. For example, the target polynucleotide sequence can be any part of a target gene. In some embodiments, the target polynucleotide sequence is part of a transcriptional regulatory sequence. In some embodiments, the target polynucleotide sequence is part of a promoter, an enhancer, or a silencer. In some embodiments, the target polynucleotide sequence is part of a promoter. In some embodiments, the target polynucleotide sequence is part of an enhancer. In some embodiments, the target polynucleotide sequence is part of a silencer. In some embodiments, the target polynucleotide sequence is within about 3000, 2900, 2800, 2700, 2600, 2500, 2400, 2300, 2200, 2100, 2000, 1900, 1800, 1700, 1600, 1500, 1400, 1300, 1200, 1100, 1000, 900, 800, 700, 600, 500, 400, 300, 200, or 100 base pairs (bp) flanking the transcription start site. In some embodiments, the target polynucleotide sequence is within about 1000, 900, 800, 700, 600, 500, 400, 300, 200, or 100 base pairs (bp) flanking the transcription start site. In some embodiments, the target polynucleotide sequence is within about 500, 400, 300, 200, or 100 base pairs (bp) flanking the transcription start site.
[0301] In some embodiments, the target polynucleotide sequence is within about 100 base pairs (bp) flanking the transcription start site.
[0302] In some embodiments, the target polynucleotide sequence is a low methylated nucleic acid sequence. In some embodiments, the target polynucleotide sequence is a high methylated nucleic acid sequence. In some embodiments, the target polynucleotide sequence is at a promoter sequence, near a promoter sequence, or within a promoter sequence. In some embodiments, the target polynucleotide sequence is at a promoter sequence, near a promoter sequence, or within a promoter sequence. In various aspects, the target polynucleotide sequence is adjacent to a CpG island. In various aspects, it is known that the target polynucleotide sequence is associated with a disease or condition.
[0303] Regulation of target gene expression
[0304] In some embodiments, the present disclosure provides epigenetic editing systems, compositions and methods for epigenetic modification at the target polynucleotide in the target gene encoding protein.In some embodiments, the epigenetic editing system causes epigenetic modification in the coding region of the target gene, such as DNA methylation, so as to reduce or silence the expression of the target gene.In some embodiments, the epigenetic editing system causes epigenetic modification in the promoter or enhancer of the regulatory sequence such as the target gene, such as DNA methylation, so as to reduce or silence the expression of the target gene.In some embodiments, the epigenetic editing system causes transcriptional repression or recruits transcriptional repressors to the coding region of the target gene, so as to reduce or silence the expression of the target gene.In some embodiments, the epigenetic editing system recruits transcriptional repressors to regulatory sequences, such as promoters or enhancers of target genes, so as to reduce or silence the expression of target genes.In some embodiments, the epigenetic editing system causes epigenetic modification (such as DNA demethylation) in the coding region of the target gene, so as to increase the expression of the target gene. In some embodiments, the epigenetic editing system causes epigenetic modification (e.g., DNA demethylation) in a regulatory sequence such as a promoter or enhancer of a target gene, thereby increasing the expression of the target gene. In some embodiments, the epigenetic editing system causes transcriptional activation or recruits transcriptional activators to the coding region of the target gene, thereby increasing the expression of the target gene. In some embodiments, the epigenetic editing system recruits transcriptional activators to a regulatory sequence (e.g., a promoter or enhancer of a target gene), thereby increasing the expression of the target gene.
[0305] In some embodiments, the target gene and / or encoded protein is associated with a disease, disorder, or pathogenic condition.
[0306] The epigenetic modification achieved by the epigenetic editing system described herein is sequence-specific. In some embodiments, the modification is at a specific site of the target polynucleotide. In some embodiments, the modification is at a specific allele of the target gene. Therefore, the epigenetic modification can cause the expression of a copy of the target gene containing a specific allele to be regulated, such as expression reduction or increase, while another copy of the target gene is not regulated. In some embodiments, a specific allele is associated with a disease, condition or illness.
[0307] Epigenetic modification can be carried out on any target gene of the genome of interest (for example, prokaryotic genome, plant genome, mammal or human genome). The target gene can belong to or be derived from any organism and its genome. For example, the target gene can be a prokaryotic gene, a eukaryotic gene, an animal gene, a plant gene, a mouse gene, a rat gene, a rabbit gene, a fish gene, a bird gene, a monkey gene or a human gene. In some embodiments, the target gene is a reporter gene, and the expression of the reporter gene can be easily tracked and monitored. Reporter gene and reporter system include, for example, sequences encoding green fluorescent protein, red fluorescent protein, enhanced yellow or enhanced cyan protein or luciferase protein. In some embodiments, the target gene encodes a selective marker, such as beta-galactosidase, chloramphenicol acetyltransferase or antibiotic resistance marker. In some embodiments, the target gene is associated with one or more mutations or contains one or more mutations, and the mutation is associated with a disease, a disease or a condition. Non-restrictive exemplary target genes include HBB, HBA, hMSH2, HMLH1, growth factor GM-SCF, VEGF, EPO, Erb-B2 and hGH.
[0308] Target genes also include plant genes, the suppression or activation of which results in improved plant properties, such as improved crop yield, disease or herbicide resistance. For example, suppression of the expression of the FAD2-1 gene results in a favorable increase in oleic acid and a decrease in one or more linoleic acids.
[0309] In some embodiments, the epigenetic editing system provided herein realizes epigenetic modification in the gene containing the target sequence. In some embodiments, the epigenetic editing system regulates the expression of the protein encoded by the gene. In some embodiments, the epigenetic editing system reduces the level of the protein encoded by the gene. In some embodiments, the epigenetic editing system increases the level of the protein encoded by the gene.
[0310] In order to produce epigenetic editing at the target gene, the target gene polynucleotide can be contacted with the epigenetic editing composition disclosed herein, the composition comprises a target DNA binding domain, an epigenetic effector domain, such as an epigenetic repressor domain, wherein the DNA binding domain guides the epigenetic effector domain to the target polynucleotide sequence in the target gene, resulting in epigenetic modification (e.g., methylation state modification). In some embodiments, the epigenetic editing system achieves a change in the methylation state of the target DNA sequence in the target gene. In some embodiments, the epigenetic editing system achieves a change in the methylation state of a specific allele in the target gene. In some embodiments, the epigenetic editing system achieves a change in the methylation state of a histone protein associated with the target gene.
[0311] In some embodiments, epigenetic modification reduces the transcription of a target gene containing a target sequence. In some embodiments, epigenetic modification eliminates the transcription of a target gene containing a target sequence. In some embodiments, epigenetic modification reduces the transcription of a copy of a target gene containing a specific allele identified by an epigenetic editing system. In some embodiments, epigenetic modification eliminates the transcription of a copy of a target gene containing a specific allele identified by an epigenetic editing system. In some embodiments, the epigenetic editing system reduces the level of protein encoded by the target gene. In some embodiments, the epigenetic editing system eliminates the expression of protein encoded by the target gene. In some embodiments, the epigenetic editing system reduces the level of protein encoded by a copy of a target gene containing a specific allele identified by an epigenetic editing system. In some embodiments, the epigenetic editing system eliminates the expression of protein encoded by a copy of a target gene containing a specific allele identified by an epigenetic editing system.
[0312] In some embodiments, the epigenetic modification increases the transcription of a target gene containing a target sequence. In some embodiments, the epigenetic modification increases the transcription of a copy of a target gene containing a specific allele identified by an epigenetic editing system. In some embodiments, the epigenetic editing system increases the level of a protein encoded by a target gene. In some embodiments, the epigenetic editing system increases the level of a protein encoded by a copy of a target gene containing a specific allele identified by an epigenetic editing system.
[0313] The target gene can be epigenetically modified in vitro, in vitro or in vivo. Therefore, the epigenetic modification of the target gene can regulate the expression of the target gene or its allele in an ex vivo cell or in an in vivo object. In some embodiments, the target polynucleotide sequence is a locus in the genomic DNA of the cell. In some embodiments, the cell is a cultured cell. In some embodiments, the cell is in vitro. In some embodiments, the cell is in vitro. In some embodiments, the cell is in vivo. For example, an epigenetic editing system, such as a fusion protein comprising a zinc finger array and an effector domain, or a sgRNA compounded with a Cas protein-effector domain fusion, can be expressed in cells requiring regulated expression of the target gene, thereby allowing the target gene to contact the epigenetic editing system described herein. In some embodiments, the cell is from a mammal. In some embodiments, the mammal is a human. In some embodiments, the mammal is a rodent. In some embodiments, the rodent is a mouse. In some embodiments, the rodent is a rat.
[0314] In some embodiments, as measured by transcription of the target gene in a cell, tissue or object, the epigenetic editing system described herein reduces the expression of the target gene by at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 99% or more compared to a control cell, control tissue or control object. In some embodiments, as measured by transcription of a copy of the target gene in a cell, tissue or object, the epigenetic editing system described herein reduces the expression of a copy of the target gene by at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 99% or more compared to a control cell, control tissue or control object. In some embodiments, the copy of the target gene contains a specific sequence or allele recognized by the epigenetic editing system. In some embodiments, the epigenetically modified copy encodes a functional protein. Therefore, in some embodiments, the epigenetic editing system composition disclosed herein reduces or eliminates the expression and / or function of the protein encoded by the target gene by reducing or eliminating the expression of the functional protein encoded by the target gene. For example, compared with a control cell, a control tissue or a control object, the methods and compositions disclosed herein can reduce the expression and / or function of the protein encoded by the target gene in cells, tissues or objects by at least 3 times, at least 4 times, at least 5 times, at least 6 times, at least 7 times, at least 8 times, at least 9 times, at least 10 times, at least 11 times, at least 12 times, at least 13 times, at least 14 times, at least 15 times, at least 20 times, at least 25 times, at least 30 times, at least 35 times, at least 40 times, at least 45 times, at least 50 times, at least 60 times, at least 70 times, at least 80 times, at least 90 times or at least 100 times.
[0315] In some embodiments, the epigenetic editing systems described herein increase expression of a target gene as measured by transcription of the target gene in a cell, tissue or subject by at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 99%, at least 100%, at least 110%, at least 120%, at least 130%, at least 140%, at least 150%, at least 160%, at least 170%, at least 180%, at least 190%, at least 200%, at least 250%, at least 300%, at least 350%, at least 400%, at least 450%, at least 500%, or more as compared to a control cell, a control tissue, or a control subject. In some embodiments, as measured by transcription of a copy of the target gene in a cell, tissue or object, the expression of a copy of the target gene of the epigenetic editing system described herein is increased by at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 99%, at least 100%, at least 110%, at least 120%, at least 130%, at least 140%, at least 150%, at least 160%, at least 170%, at least 180%, at least 190%, at least 200%, at least 250%, at least 300%, at least 350%, at least 400%, at least 450%, at least 500% or more compared to a control cell, control tissue or control object. In some embodiments, the copy of the target gene contains a specific sequence or allele recognized by the epigenetic editing system. In some embodiments, the epigenetically modified copy encodes a functional protein. Therefore, in some embodiments, the epigenetic editing system composition disclosed herein increases the expression and / or function of the protein encoded by the target gene by increasing the expression of the functional protein encoded by the target gene. For example, compared with a control cell, a control tissue or a control object, the methods and compositions disclosed herein can increase the expression and / or function of the protein encoded by the target gene in a cell, tissue or object by at least 3 times, at least 4 times, at least 5 times, at least 6 times, at least 7 times, at least 8 times, at least 9 times, at least 10 times, at least 11 times, at least 12 times, at least 13 times, at least 14 times, at least 15 times, at least 20 times, at least 25 times, at least 30 times, at least 35 times, at least 40 times, at least 45 times, at least 50 times, at least 60 times, at least 70 times, at least 80 times, at least 90 times or at least 100 times.
[0316] Methods for determining the expression level of a gene (such as a target of an epigenetic editing system) are known in the art. For example, reverse transcription PCR, quantitative RT-PCR, droplet digital PCR (ddPCR), Northern blotting, RNA sequencing, DNA sequencing (for example, sequencing of complementary deoxyribonucleic acid (cDNA) obtained from RNA), next generation (Next-Gen) sequencing, nanopore sequencing, pyrophosphate sequencing or nanochain sequencing can be used to determine the transcription level of a gene. The protein level expressed by a gene can be determined by Western blotting, enzyme-linked immunosorbent assay, mass spectrometry, immunohistochemistry or flow cytometry analysis. The gene expression product level can be normalized relative to an internal standard (such as the expression level of total messenger ribonucleic acid (mRNA) or a specific gene (such as a housekeeping gene)).
[0317] In some embodiments, a reporter system can be used to check the effect of the epigenetic editing system in regulating the expression of target genes.For example, the epigenetic editing system can be designed to target the reporter gene encoding reporter protein (e.g., fluorescent protein).The expression of reporter genes in such model systems can be monitored by, for example, flow cytometry, fluorescence activated cell sorting (FACS) or fluorescence microscopy.In some embodiments, a cell colony can be transfected with a vector containing a reporter gene.A vector can be constructed so that when the vector is transfected with cells, the reporter gene is expressed.Suitable reporter genes include genes encoding fluorescent proteins (e.g., green, yellow, cherry, cyan or orange fluorescent proteins).The cell colony carrying the reporter system can be transfected with DNA, mRNA or a vector encoding the epigenetic editing system of the targeted reporter gene.The expression level of the reporter gene can be quantified using suitable technology (such as FACS).
[0318] Epigenetic editing system as described herein can be transiently expressed in host cells, or can be integrated into the genome of host cells. Both the epigenetic editing system of transient expression and the epigenetic editing system of integration can achieve stable epigenetic modification. For example, after the epigenetic editing system comprising a DNA binding domain and an epigenetic repression domain having specificity for the target gene is introduced into the host cell, the target gene in the host cell can be stably or permanently repressed. In some embodiments, compared with the expression level when there is no epigenetic editing system, the expression of the target gene is reduced for at least 1 week, at least 2 weeks, at least 3 weeks, at least 4 weeks, at least 5 weeks, at least 6 weeks, at least 7 weeks, at least 2 months, at least 3 months, at least 5 months, at least 6 months, at least 1 year, at least 2 years, or the whole life cycle of the object of continuous cell or carrying cell. In some embodiments, the expression of the target gene is silenced for at least 1 week, at least 2 weeks, at least 3 weeks, at least 4 weeks, at least 5 weeks, at least 6 weeks, at least 7 weeks, at least 2 months, at least 3 months, at least 5 months, at least 6 months, at least 1 year, at least 2 years, or for the entire life cycle of the cell or the object carrying the cell, compared to the expression level when the epigenetic editing system is not present. In some embodiments, after the epigenetic editing system comprising a DNA binding domain specific to the target gene and an epigenetic activation domain is introduced into the host cell, the target gene in the host cell is stably or permanently activated. In some embodiments, the expression of the target gene is increased for at least 1 week, at least 2 weeks, at least 3 weeks, at least 4 weeks, at least 5 weeks, at least 6 weeks, at least 7 weeks, at least 2 months, at least 3 months, at least 5 months, at least 6 months, at least 1 year, at least 2 years, or for the entire life cycle of the cell or the object carrying the cell, compared to the expression level when the epigenetic editing system is not present.
[0319] Epigenetic modifications described herein can be inherited by the offspring of host cells contacted or introduced into the epigenetic editing system. For example, in some embodiments, after the epigenetic editing system comprising a DNA binding domain and an epigenetic repression domain having specificity for the target gene is introduced into stem cells (e.g., hematopoietic stem cells), compared with cells differentiated from control stem cells when there is no epigenetic editing system, the expression of the target gene in cells differentiated from stem cells is also suppressed. In some embodiments, the expression of the target gene in cells differentiated from stem cells is silenced. In some embodiments, after the epigenetic editing system comprising a DNA binding domain and an epigenetic activation domain having specificity for the target gene is introduced into stem cells (e.g., hematopoietic stem cells), compared with cells differentiated from control stem cells when there is no epigenetic editing system, the expression of the target gene in cells differentiated from stem cells is also increased.
[0320] Regulation of target gene expression can be determined by determining any parameter that is indirectly or directly affected by the expression of the target gene. Such parameters include, for example, changes in RNA or protein levels; changes in protein activity; changes in product levels; changes in downstream gene expression; changes in transcription or activity of reporter genes (such as luciferase, CAT, β-galactosidase or GFP); changes in signal transduction; changes in phosphorylation and dephosphorylation; changes in receptor-ligand interactions; changes in the concentration of second messengers (such as cGMP, cAMP, IP3 and Ca2+); changes in cell growth, changes in new blood vessel formation and / or changes in any functional effects of gene expression. Measurements can be performed in vitro, in vivo and / or ex vivo. Such functional effects can be measured by conventional methods, for example, measuring RNA or protein levels, measuring RNA stability, and / or identifying downstream or reporter gene expression. Readout can be measured by, for example, chemiluminescence, fluorescence, colorimetric reaction, antibody binding, inducible markers, ligand binding assays; changes in intracellular second messengers (such as cGMP and inositol triphosphate (IP3)); changes in intracellular calcium levels; cytokine release, etc.
[0321] To determine the level of gene expression regulation by the ZFP, cells contacted with the ZFP are compared to control cells (e.g., without zinc finger proteins or with non-specific ZFPs) to examine the extent of inhibition or activation. The control sample is assigned a relative gene expression activity value of 100%. Regulation / inhibition of gene expression is achieved when the gene expression activity value relative to the control is about 80%, preferably 50% (i.e., 0.5 times the activity of the control), more preferably 25%, more preferably 5-0%. Regulation / activation of gene expression is achieved when the gene expression activity value relative to the control is 110%, more preferably 150% (i.e., 1.5 times the activity of the control), more preferably 200-500%, more preferably 1000-2000% or more.
[0322] deliver
[0323] In one aspect, there is provided herein a composition for regulating gene expression, comprising an epigenetic editing system as provided herein, wherein the epigenetic editing system produces epigenetic modifications at a target gene. An epigenetic editing system or a nucleic acid encoding an epigenetic editing system or its components (e.g., nucleic acids encoding epigenetic editing system fusion proteins comprising zinc finger-repressor fusions, Cas9-repressor fusions, and / or nucleic acids encoding one or more guide RNAs) can be introduced into a cell via various methods known in the art. For example, in some embodiments, an epigenetic editing system is delivered to a host cell or integrated into the genome of a host cell, or used for transient expression in a host cell.
[0324] In some embodiments, the nucleic acid encoding the epigenetic editing system or its components is operably connected to a promoter and / or a regulatory sequence. As used herein, the term "operably connected" means that the nucleotide sequence of interest is connected to the regulatory sequence in a manner that allows the expression of the nucleotide sequence. As used herein, the term "regulatory sequence" includes but is not limited to promoters, enhancers, and other expression control elements. Such regulatory sequences are well known in the art and are described, for example, in Goeddel; Gene Expression Technology: Methods in Enzymology 185, Academic Press, San Diego, CA (1990).
[0325] In some embodiments, the composition further comprises a vector, and the vector comprises a nucleic acid sequence encoding an epigenetic editing system protein. In some embodiments, the vector may be an expression vector. In some embodiments, the vector is a plasmid or a viral vector. As used herein, the term "vector" refers to a nucleic acid molecule capable of transporting another nucleic acid connected thereto. In some examples, the vector is an expression vector, and the expression vector is capable of guiding the expression of a nucleic acid operably connected thereto. Examples of expression vectors include, but are not limited to, plasmid vectors, based on vaccinia virus, poliovirus, adenovirus, adeno-associated virus, SV40, herpes simplex virus, human immunodeficiency virus, retrovirus (e.g., murine leukemia virus, spleen necrosis virus and vectors derived from retroviruses such as Rous sarcoma virus, Harvey sarcoma virus, avian leukosis virus, slow virus, human immunodeficiency virus, myeloproliferative sarcoma virus and mammary tumor virus) viral vectors and other recombinant vectors. In some embodiments, the vector is a virus-like particle (VLP).
[0326] Non-viral delivery systems include, but are not limited to, DNA delivery methods and RNA delivery methods, such as transfection. Here, transfection includes the process of using non-viral vectors to deliver genes, DNA fragments, gene transcripts, RNA, RNA fragments, circularized DNA or circularized RNA to target cells. Typical transfection methods may include, but are not limited to, electroporation, DNA bio-bombing (biolistics), lipid-mediated transfection, compressed DNA-mediated transfection, liposomes, immunoliposomes, exosomes, lipofection, cationic agent-mediated transfection or cationic surface amphiphiles (CFA).
[0327] In some embodiments, the epigenetic editing system is delivered to the host cell for transient expression, for example, via a transient expression vector. The transient expression of the epigenetic editing system can result in long-term or permanent epigenetic modification of the target gene. For example, after the epigenetic editing system is introduced into the host cell, the epigenetic modification can be stably continued for at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 11, 12 weeks, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12 months or longer. The epigenetic modification can be maintained after one or more mitotic events of the host cell. The epigenetic modification can be maintained after one or more meiotic events of the host cell. In some embodiments, the epigenetic modification is maintained across generations in the offspring produced or derived from the host cell.
[0328] In some embodiments, the nucleic acid sequence encoding the epigenetic editing system or its components is DNA, RNA or mRNA, or a modified nucleic acid sequence. For example, the mRNA sequence encoding the epigenetic editing system fusion protein can be chemically modified, or can include a 5' cap, or one or more 3' modifications.
[0329] The nucleic acid encoding the epigenetic editing system can be directly delivered to the cell as naked DNA or RNA, for example, by transfection or electroporation, or can be conjugated with a molecule (e.g., N-acetylgalactosamine) that promotes uptake by the target cell. Nucleic acid vectors (e.g., vectors) can also be used. In a specific embodiment, polynucleotides (e.g., mRNA encoding an epigenetic editing system or its functional components) can be electroporated with a combination of multiple guide RNAs as described herein.
[0330] Nucleic acid vectors can include one or more sequences encoding the domains of fusion proteins or epigenetic editing systems as described herein. The vector can also include a sequence encoding a signal peptide (e.g., for nuclear localization, nucleolar localization, or mitochondrial localization), which is associated with the sequence encoding the protein (e.g., inserted or fused to the sequence encoding the protein). As an example, a nucleic acid vector can include a Cas9 encoding sequence, which includes one or more nuclear localization sequences (e.g., from SV40 nuclear localization sequences), and one or more effector domains (e.g., repressor domains).
[0331] In certain embodiments, all or part of the fusion protein, protein domain, or epigenetic editing system component is encoded by a polynucleotide present in a suitable capsid protein of a viral vector (e.g., adeno-associated virus (AAV), AAV3, AAV3b, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAVrh8, AAV10, and variants thereof) or any viral vector. Thus, in some aspects, the present disclosure relates to viral delivery of fusion proteins. Examples of viral vectors include retroviral vectors (e.g., Moloney murine leukemia virus, MML-V), adenoviral vectors (e.g., AD100), lentiviral vectors (vectors based on HIV and FIV), herpes virus vectors (e.g., HSV-2).
[0332] In some embodiments, the epigenetic editing system protein is encoded by a polynucleotide present in an adeno-associated virus (AAV) vector. In some embodiments, the epigenetic editing system protein comprises a zinc finger array in a DNA binding domain. Without wishing to be bound by any theory, in view of the small size of zinc fingers, an epigenetic editing system using zinc finger arrays rather than larger DNA binding domains such as Cas protein domains can be conveniently packaged in a viral vector (e.g., AAV vector). In some embodiments, the length of the polynucleotide encoding the epigenetic editing system is about 1000 bp, 1.1 kilobases (kb), 1.2 kb, 1.3 kb, 1.4 kb, 1.5 kb, 1.6 kb, 1.7 kb, 1.8 kb, 1.9 kb, 2.0 kb, 2.1 kb, 2.2 kb, 2.3 kb, 2.4 kb, 2.5 kb, 2.6 kb, 2.7 kb, 2.8 kb, 2.9 kb, 3.0 kb, 3.1 kb, 3.2 kb, 3.3 kb, 3.4 kb, 3.5 kb, 3.6 kb, 3.7 kb, 3.8 kb, 3.9 kb, 4.0 kb or less. In some embodiments, the length of the polynucleotide encoding the epigenetic editing system is about 2.0 kb, 2.1 kb, 2.2 kb, 2.3 kb, 2.4 kb, 2.5 kb, 2.6 kb, 2.7 kb, 2.8 kb, 2.9 kb, 3.0 kb, 3.1 kb, 3.2 kb, 3.3 kb, 3.4 kb, 3.5 kb, 3.6 kb, 3.7 kb, 3.8 kb, 3.9 kb, 4.0 kb, 4.1 kb, 4.2 kb, 4.3 kb, 4.4 kb, 4.5 kb, 4.6 kb, 4.7 kb, 4,8 kb, 4.9 kb, 5 kb or less.
[0333] Any AAV serotype can be used, such as a human AAV serotype, including but not limited to AAV serotype 1 (AAV1), AAV serotype 2 (AAV2), AAV serotype 3 (AAV3), AAV serotype 4 (AAV4), AAV serotype 5 (AAV5), AAV serotype 6 (AAV6), AAV serotype 7 (AAV7), AAV serotype 8 (AAV8), AAV serotype 9 (AAV9), AAV serotype 10 (AAV10), AAV serotype 11 (AAV11), AAV serotype 11 (AAV11), variants thereof, or shuffled variants thereof (e.g., chimeric variants thereof). In some embodiments, the AAV variant has at least 90%, such as 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more amino acid sequence identity with wild-type AAV. AAV1 variants may have at least 90%, such as 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more amino acid sequence identity with wild-type AAV1. AAV2 variants may have at least 90%, such as 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more amino acid sequence identity with wild-type AAV2. AAV3 variants may have at least 90%, such as 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more amino acid sequence identity with wild-type AAV3. AAV4 variants may have at least 90%, such as 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more amino acid sequence identity with wild-type AAV4. AAV5 variants may have at least 90%, such as 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more amino acid sequence identity with wild-type AAV5. AAV6 variants may have at least 90%, such as 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more amino acid sequence identity with wild-type AAV6. The AAV7 variants may have at least 90%, e.g., 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more amino acid sequence identity with wild-type AAV7. The AAV8 variants may have at least 90%, e.g., 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more amino acid sequence identity with wild-type AAV8.AAV9 variants may have at least 90%, such as 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more amino acid sequence identity with wild-type AAV9. AAV10 variants may have at least 90%, such as 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more amino acid sequence identity with wild-type AAV10. AAV11 variants may have at least 90%, such as 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more amino acid sequence identity with wild-type AAV11. The AAV12 variant can have at least 90%, such as 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more amino acid sequence identity with wild-type AAV12.
[0334] In some cases, one or more regions of at least two different AAV serotype viruses are reorganized and reassembled to produce AAV chimeric viruses. For example, chimeric AAV can include inverted terminal repeats (ITRs) of a heterologous serotype compared to the serotype of the capsid. The resulting chimeric AAV virus can have different antigenic reactivity or recognition capabilities compared to its parental serotype. In some embodiments, chimeric variants of AAV include amino acid sequences from 2, 3, 4, 5 or more different AAV serotypes.
[0335] Descriptions of AAV variants and methods for their generation can be found, for example, in Weitzman and Linden. Chapter 1 - Adeno-Associated Virus Biology in Adeno-Associated Virus: Methods and Protocols Methods in Molecular Biology, Vol. 807. Snyder and Moullier, eds., Springer, 2011; Potter et al., Molecular Therapy—Methods & Clinical Development, 2014, 1, 14034; Bartel et al., Gene Therapy, 2012, 19, 694-700; Ward and Walsh, Virology, 2009, 386(2):237-248; and Li et al., Mol Ther, 2008, 16(7):1252-1260, each of which is incorporated herein by reference in its entirety. AAV virions (e.g., viral vectors or viral particles) described herein can be transduced into cells to introduce the epigenetic editing system or any component thereof into cells. The epigenetic editing system can be packaged into an AAV viral vector according to any method known to those skilled in the art. Examples of useful methods are described in McClure et al., J Vis Exp, 2001, 57:3378.
[0336] The nucleic acid vectors described herein may also include any suitable number of regulatory / control elements, such as promoters, enhancers, introns, polyadenylation signals, Kozak consensus sequences, or internal ribosome entry sites (IRES). These elements are well known in the art.
[0337] Nucleic acid vectors according to the present disclosure include recombinant viral vectors. Exemplary viral vectors are described above. Other viral vectors known in the art may also be used. In addition, viral particles may be used to deliver genome editing system components in the form of nucleic acids and / or peptides. For example, "empty" viral particles may be assembled to contain any suitable cargo. Viral vectors and viral particles may also be engineered to incorporate targeting ligands to change target tissue specificity.
[0338] In addition to viral vectors, non-viral vectors can be used to deliver nucleic acids encoding genome editing systems according to the present disclosure. One category of non-viral nucleic acid vectors is nanoparticles, which can be organic or inorganic. Nanoparticles are well known in the art. Any suitable nanoparticle design can be used to deliver genome editing system components or nucleic acids encoding such components. For example, in certain embodiments of the present disclosure, organic (e.g., lipids and / or polymers) nanoparticles can be suitable for use as delivery vectors.
[0339] On the other hand, a lipid nanoparticle (LNP) is provided herein, comprising a composition as provided herein. As used herein, "lipid nanoparticle (LNP) composition" or "nanoparticle composition" is a composition comprising one or more of the lipids described. The size of the LNP composition is generally in the order of micrometer or less, and may include a lipid bilayer. The nanoparticle composition includes lipid nanoparticles (LNP), liposomes (e.g., lipid vesicles) and lipid complexes (lipoplexes). In some embodiments, LNP refers to any particle having a diameter less than 1000nm, 500nm, 250nm, 200nm, 150nm, 100nm, 75nm, 50nm or 25nm. In some embodiments, the size of the nanoparticle can be in the range of 1-1000nm, 1-500nm, 1-250nm, 25-200nm, 25-100nm, 35-75nm or 25-60nm.
[0340] In some embodiments, LNP can be made of cationic lipids, anionic lipids or neutral lipids. In some embodiments, LNP can include neutral lipids (such as fusion phospholipids 1,2-dioleoyl-sn-glycerol-3-phosphoethanolamine (DOPE) or membrane component cholesterol) as auxiliary lipids to enhance transfection activity and nanoparticle stability. In some embodiments, LNP can include hydrophobic lipids, hydrophilic lipids or both hydrophobic lipids and hydrophilic lipids. Any lipid known in the art or the combination of lipids can be used to produce LNP. Examples of lipids used to produce LNPs include, but are not limited to, DOTMA (N[1-(2,3-dioleoyloxy)propyl]-N,N,N-trimethylammonium chloride), DOSPA (N,N-dimethyl-N-([2-sperminamido]ethyl)-2,3-bis(dioleoyloxy)-1-propyliminium pentahydrochloride), DOTAP (1,2-dioleoyl-3-trimethylammonium propane), DMRIE (N-(2-hydroxyethyl)-N,N-dimethyl-2,3-bis(tetradecyloxy-1-propylammonium bromide), DC-cholesterol (3β-[N-(N',N'-dimethylaminoethane)-carbamoyl]cholesterol), DOTAP-cholesterol, GAP-DMORIE-DPyPE, and GL67A-DOPE-DMPE (2-bis(dimethylphosphino)ethane)-polyethylene glycol In some embodiments, the lipid of the present invention can be PEG-modified.In some embodiments, the lipid of the present invention can be PEG-modified.In some embodiments, the lipid of the present invention can be PEG-modified.In some embodiments, the lipid of the present invention can be PEG-modified.In some embodiments, the lipid of the present invention can be PEG-modified.In some embodiments, the lipid of the present invention can be PEG-modified.In some embodiments, the lipid of the present invention can be PEG-modified.In some embodiments, the lipid of the present invention can be PEG-modified.In some embodiments, the lipid of the present invention can be PEG-modified.In some embodiments, the lipid of the present invention can be PEG-modified.In some embodiments, the lipid of the present invention can be PEG-modified.In some embodiments, the lipid of the present invention can be PEG-modified.In some embodiments, the lipid of the present invention can be PEG-modified.In some embodiments, the lipid of the present invention can be PEG-modified.In some embodiments, the lipid of the present invention can be PEG-modified.In some embodiments, the lipid of the present invention can be PEG-modified.In some embodiments, the lipid of the present invention can be PEG-modified.In some embodiments, the lipid of the present invention can be PEG-modified.
[0341] Treatment
[0342] Also provided herein is a method for treating or preventing a condition in a subject in need thereof, the method comprising administering to the subject an epigenetic editing system composition as described herein, wherein the epigenetic editing system complex or protein achieves epigenetic modification of a target polynucleotide in a target gene associated with a disease, condition or disorder in the subject, and modulates the expression of the target, thereby treating or preventing the disease, condition or disorder.
[0343] The epigenetic modification achieved by the epigenetic editing system described herein is sequence-specific. In some embodiments, the modification is at a specific site of the target polynucleotide. In some embodiments, the modification is at a specific allele of the target gene. Therefore, the epigenetic modification can cause the expression of a copy of the target gene containing a specific allele to be regulated, such as expression reduction or increase, while another copy of the target gene is not regulated. In some embodiments, a specific allele is associated with a disease, condition or illness.
[0344] In some embodiments, the epigenetic editing system reduces expression of a target gene associated with a disease, condition, or disorder.
[0345] The epigenetic editing systems described herein can be administered to a subject in need thereof in a therapeutically effective amount to treat a disease, condition, or disorder.
[0346] In another aspect, provided herein is a method for treating or preventing a condition in a subject in need thereof, the method comprising administering to the subject an epigenetic editing complex, vector, nucleic acid, protein or composition as provided herein, wherein the nucleic acid binding domain of the epigenetic editing system directs the effector domain to produce an epigenetic modification in a target polynucleotide sequence in a cell of the subject, thereby regulating the expression of the target gene and treating or preventing the condition.
[0347] In some embodiments, the modification reduces expression of a functional protein encoded by the target gene in the subject.
[0348] Patients being treated for a condition, disease or illness are patients who have been diagnosed by a physician as suffering from such a condition. Diagnosis can be performed by any suitable means. Diagnosis and monitoring can involve, for example, detecting the presence of diseased, dying or dead cells in a biological sample (e.g., a tissue biopsy, a blood test or a urine test), detecting the presence of plaques, detecting the level of a surrogate marker in a biological sample, or detecting symptoms associated with a condition. Patients whose condition is being prevented from developing may or may not have received such a diagnosis. Those skilled in the art will appreciate that these patients may have been subjected to the same standard tests as described above, or may have been identified as patients at high risk without examination due to the presence of one or more risk factors (e.g., a family history or a genetic predisposition).
[0349] The subject may be suffering from a disease, a symptom of a disease, or a disease tendency, and the purpose is to cure, heal, alleviate, mitigate, change, remedy, improve, modify, or affect a disease, a symptom of a disease, or a disease tendency. In some embodiments, the subject suffers from hypercholesterolemia. In some embodiments, the subject suffers from atherosclerotic vascular disease. In some embodiments, the subject suffers from hypertriglyceridemia. In some embodiments, the subject suffers from diabetes. In some embodiments, the subject is a mammal. In some embodiments, the subject is a non-human primate. In some embodiments, the subject is a human. Alleviating a disease includes delaying the development or progression of a disease, or reducing the severity of a disease. Alleviating a disease does not necessarily require a curative outcome.
[0350] As used herein, "delaying" the development of a disease means postponing, hindering, slowing down, slowing down, stabilizing and / or delaying the progress of a disease. Depending on the history of the disease and / or the individual being treated, this delay can have different time lengths. A method of "delaying" or alleviating the development of a disease or delaying the onset of a disease is a method that, when compared to not using the method, uses the method to reduce the likelihood of developing one or more symptoms of the disease within a given time frame and / or reduce the degree of symptoms within a given time frame. Such comparisons are typically based on clinical studies using a sufficient number of subjects to give statistically significant results.
[0351] "Development" or "progression" of a disease means the initial manifestation and / or subsequent progression of the disease. The development of a disease can be detected and assessed using standard clinical techniques as are well known in the art. However, development also refers to progression that may not be detectable. For purposes of this disclosure, development or progression refers to the biological course of symptoms. "Development" includes onset, recurrence, and onset.
[0352] As used herein, "onset" or "occurrence" of a disease includes initial onset and / or recurrence. Depending on the type of disease to be treated or the site of the disease, the isolated polypeptide or pharmaceutical composition can be administered to the subject using conventional methods known to those of ordinary skill in the medical arts. The composition can also be administered via other conventional routes, such as orally, parenterally, by inhalation spray, topically, rectally, nasally, buccally, vaginally, or via an implanted reservoir.
[0353] The therapeutic methods disclosed herein can be implemented on objects showing pathology caused by a disease or condition, objects suspected of showing pathology caused by a disease or condition, and objects at risk of showing pathology caused by a disease or condition. For example, objects with a genetic predisposition to a disease or condition can be treated prophylactically. Objects showing symptoms associated with a condition, disease, or illness can be treated to alleviate symptoms or slow down or prevent further progression of symptoms. Physical changes associated with the increased severity of a disease or condition are shown as progressive herein. Therefore, in embodiments of the present disclosure, objects showing mild signs of pathology associated with a condition or disease can be treated to improve symptoms and / or prevent further progression of symptoms.
[0354] The dosage and frequency (single or multiple doses) of administration to a mammal may vary according to a variety of factors, such as whether the mammal suffers from another disease, and its route of administration; the recipient's size, age, sex, health, weight, body mass index, and diet; the nature and extent of the symptoms of the disease being treated, the type of concurrent treatment, complications of the disease being treated, or other health-related issues. The adjustment and manipulation of the determined dosage (e.g., frequency and duration) are entirely within the capabilities of those skilled in the art. Treatments, such as those disclosed herein, may be administered to a subject daily, twice a day, once every two weeks, once a month, or on any applicable basis that is therapeutically effective. In an embodiment, treatment is performed only on an as-needed basis (e.g., when signs or symptoms of a condition or disease occur).
[0355] The toxicity and therapeutic efficacy of the compositions disclosed herein can be determined by standard pharmaceutical procedures in cell culture or experimental animals, for example, for determining LD50 (a dose lethal to 50% of the population) and ED50 (a dose effective for the treatment of 50% of the population). The dose ratio between toxicity and therapeutic effect (the ratio of LD50 / ED50) is the therapeutic index. Drugs showing a high therapeutic index are preferred. The dosage of the agent is preferably within the range of a circulating concentration of an ED50 that includes little or no toxicity. Although agents showing toxic side effects can be used, it should be noted that the delivery system designed to target such agents to the site of infected tissues is to minimize potential damage to uninfected cells, thereby reducing side effects.
[0356] The skilled person will appreciate that certain factors may affect the dosage and frequency of administration required to effectively treat a subject, including but not limited to the severity of the disease or condition, previous treatment, general characteristics of the subject, including the subject's health, sex, weight and / or age, and other diseases present. In addition, treating a subject with a therapeutically effective amount of a composition may include a single treatment, or preferably may include a series of treatments. It should also be understood that the effective dose of the composition of the present disclosure for treatment may increase or decrease during a particular treatment. Depending on the results of diagnostic assays as described herein, changes in dosage may occur and become apparent. The therapeutically effective dose will generally depend on the patient's condition at the time of administration. The exact amount can be determined by routine experimentation, but may ultimately depend on the judgment of the clinician, for example, by monitoring the patient's signs of disease and adjusting the treatment accordingly.
[0357] The frequency of administration can be determined and adjusted during the course of therapy, and is usually but not necessarily based on the treatment and / or inhibition and / or improvement and / or delay of the disease. Alternatively, the sustained continuous release formulation of the polypeptide or polynucleotide may be suitable. Various formulations and devices for achieving sustained release are known in the art. In some embodiments, the dosage is every day, every other day, every three days, every four days, every five days or every six days. In some embodiments, the frequency of administration is once a week, once every 2 weeks, once every 4 weeks, once every 5 weeks, once every 6 weeks, once every 7 weeks, once every 8 weeks, once every 9 weeks or once every 10 weeks; or once a month, once every 2 months, or once every 3 months, or longer. The progress of this therapy is easily monitored by conventional techniques and determinations.
[0358] Dosage regimens (including compositions disclosed herein) can vary over time. In some embodiments, for adult subjects of normal weight, a dosage in the range of about 0.01 to 1000 mg / kg can be administered. In some embodiments, the dosage is between 1 and 200 mg. The specific dosing regimen, i.e., dosage, time, and repetition, will depend on the specific subject and the subject's medical history, as well as the properties of the polypeptide or polynucleotide (e.g., the half-life of the polypeptide or polynucleotide, and other considerations well known in the art).
[0359] For the purposes of this disclosure, the appropriate therapeutic dose of the compositions as described herein will depend on the specific agent (or composition thereof) employed, the formulation and route of administration, the type and severity of the disease, whether the polypeptide or polynucleotide is being administered for preventive or therapeutic purposes, previous therapy, the subject's clinical history and response to the antagonist, and the judgment of the attending physician. Typically, the clinician will administer the polypeptide until a dose is reached that achieves the desired result.
[0360] Administration of one or more compositions can be continuous or intermittent, depending on, for example, the physiological condition of the recipient, whether the purpose of administration is therapeutic or preventive, and other factors known to skilled practitioners. Administration of the composition can be substantially continuous over a preselected period of time, or can be in a series of spaced doses, such as before, during or after the development of a disease.
[0361] The disclosed methods and compositions (including embodiments thereof) as described herein can be applied together with one or more other treatment schemes or medicaments or treatments, which can be co-administered to mammals."Co-administered" means applying one or more other treatment schemes or medicaments or treatments and compositions of the present disclosure close enough in time to enhance the effect of one or more other therapeutic agents, and vice versa. In this regard, the disclosed compositions as described herein can be applied simultaneously with one or more other treatment schemes or medicaments or treatments, applied at different times, or applied with completely different treatment schemes (for example, the first treatment can be once a day, while other treatments are once a week). For example, in embodiments, secondary treatment schemes or medicaments or treatments are applied while with the disclosed compositions, before the disclosed compositions, or after the disclosed compositions.
[0362] Pharmaceutical compositions, dosage forms and administration
[0363] In some aspects, a pharmaceutical composition for epigenetic modification is provided herein, the pharmaceutical composition comprising an epigenetic editing system as described herein, or one or more nucleic acid sequences encoding epigenetic editing system components (e.g., nucleic acids encoding epigenetic editing system fusion proteins and / or guide RNAs), and a pharmaceutically acceptable carrier. The composition for epigenetic modification described herein can be formulated into a pharmaceutical composition. Pharmaceutical compositions are prepared in a conventional manner using one or more pharmaceutically acceptable inactive ingredients, which contribute to processing active compounds into pharmaceutically usable preparations. Suitable formulations and delivery methods for the present disclosure are generally well known in the art. Suitable formulations depend on the selected route of administration. A summary of the pharmaceutical compositions described herein can be found, for example, in Remington: The Science and Practice of Pharmacy, 19th ed. (Easton, Pa.: Mack Publishing Company, 1995); Hoover, John E., Remington's Pharmaceutical Sciences, Mack Publishing Co., Easton, Pennsylvania 1975; Liberman, H. A. and Lachman, L., eds., Pharmaceutical Dosage Forms, Marcel Decker, New York, NY, 1980; and Pharmaceutical Dosage Forms and Drug Delivery Systems, 7th ed. (Lippincott Williams & Wilkins 1999), the disclosures of which are incorporated herein by reference.
[0364] The pharmaceutical composition can be a mixture of an epigenetic editing system as described herein or a nucleic acid encoding it with one or more other chemical components (i.e., pharmaceutically acceptable ingredients), such as carriers, excipients, adhesives, fillers, suspending agents, flavoring agents, sweeteners, disintegrants, dispersants, surfactants, lubricants, colorants, diluents, solubilizers, wetting agents, plasticizers, stabilizers, penetration enhancers, wetting agents, defoamers, antioxidants, preservatives, or one or more combinations thereof. The pharmaceutical composition facilitates the administration of epigenetic editors, such as nucleic acids encoding zinc finger-epigenetic effector fusion proteins or Cas9-epigenetic effector fusion proteins, and gRNA or sgRNA as described herein to an organism or object in need.
[0365] The pharmaceutical composition of the present disclosure can be applied to the object using any suitable method known in the art. The pharmaceutical composition described herein can be applied to the object in a variety of ways (including parenteral, intravenous, intradermal, intramuscular, colonic, rectal or intraperitoneal). In some embodiments, the pharmaceutical composition can be administered by intraperitoneal injection, intramuscular injection, subcutaneous injection or intravenous injection of the object. In some embodiments, the pharmaceutical composition can be administered by parenteral, intravenous, intramuscular or oral administration.
[0366] For administration by inhalation, the adenoviruses described herein can be formulated for use as an aerosol, mist or powder. For oral or sublingual administration, the pharmaceutical composition can be formulated in the form of tablets, lozenges or gels formulated in a conventional manner. In some embodiments, the adenoviruses described herein can be prepared into transdermal dosage forms. In some embodiments, the adenoviruses described herein can be formulated into pharmaceutical compositions suitable for intramuscular, subcutaneous or intravenous injection. In some embodiments, the adenoviruses described herein can be topically administered and can be formulated into a variety of topically administrable compositions, such as solutions, suspensions, lotions, gels, pastes, medicated sticks, balms, creams or ointments. In some embodiments, the adenoviruses described herein can be formulated in rectal compositions, such as enemas, rectal gels, rectal foams, rectal aerosols, suppositories, jelly suppositories or retention enemas. In some embodiments, the adenoviruses described herein can be formulated for oral administration, such as a liquid in the form of a tablet, capsule, or aqueous suspension or solution selected from, but not limited to, aqueous oral dispersions, emulsions, solutions, elixirs, gels, and syrups.
[0367] In some embodiments, the pharmaceutical composition for epigenetic modification comprising an epigenetic editor as described herein or a nucleic acid sequence encoding it further comprises a therapeutic agent. Other therapeutic agents can regulate different aspects of the disease, disorder or condition being treated, and provide a greater overall benefit than administering a recombinant adenovirus or therapeutic agent with replication capability alone. Therapeutic agents include, but are not limited to, chemotherapeutic agents, radiotherapeutic agents, hormone therapy agents, and / or immunotherapeutic agents. In some embodiments, the therapeutic agent may be a radiotherapeutic agent. In some embodiments, the therapeutic agent may be a hormone therapy agent. In some embodiments, the therapeutic agent may be an immunotherapeutic agent. In some embodiments, the therapeutic agent is a chemotherapeutic agent. The preparation and dosing regimen of other therapeutic agents can be used according to the manufacturer's instructions or as determined empirically by a skilled practitioner. For example, the preparation and dosing regimen of chemotherapy are also described in Chemotherapy Source Book, 4th edition, 2008, MCPerry, editor, Lippincott, Williams & Wilkins, Philadelphia, PA.
[0368] The subject that can be treated with the epigenetic modification composition can be any subject suffering from a disease or condition. For example, the subject can be a eukaryotic subject, such as an animal. In some embodiments, the subject is a mammal, such as a human. In some embodiments, the subject is a human. In some embodiments, the subject is a non-human animal. In some embodiments, the subject is a fetus, an embryo, or a child. In some embodiments, the subject is a non-human primate, such as a chimpanzee and other ape and monkey species; livestock, such as cattle, horses, sheep, goats, pigs; domestic animals, such as rabbits, dogs, and cats; laboratory animals, including rodents, such as rats, mice, and guinea pigs, and the like.
[0369] In some embodiments, the subject is prenatal (e.g., a fetus), a child (e.g., a newborn, an infant, a toddler, a prepubertal child), a teenager, a pubescent teenager, or an adult (e.g., an early adult, a middle-aged adult, a senior citizen). A human subject can be about 0 months to about 120 years old, or older. A human subject can be about 0 to about 12 months old; for example, about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 months old. A human subject can be about 0 to 12 years old; for example, about 0 to 30 days old; about 1 month to 12 months old; about 1 to 3 years old; about 4 to 5 years old; about 4 to 12 years old; about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 years old. A human subject can be about 13 to 19 years old; for example, about 13, 14, 15, 16, 17, 18, or 19 years old. The human subject can be between about 20 and about 39 years old; for example, about 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, or 39 years old. The human subject can be between about 40 and about 59 years old; for example, about 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, or 59 years old. for example, about 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, or 120 years old. The human subject can include a male subject and / or a female subject.
[0370] In certain embodiments, kits and manufactured articles for use with one or more methods described herein are also disclosed herein. Such kits include carriers, packages or containers that are separated to receive one or more containers (such as vials, tubes, etc.), each container including one of the independent elements to be used in the methods described herein. Suitable containers include, for example, bottles, vials, syringes, and test tubes. In one embodiment, the container is formed by a variety of materials (e.g., glass or plastic).
[0371] The articles of manufacture provided herein include packaging materials. Examples of pharmaceutical packaging materials include, but are not limited to, blister packs, bottles, tubes, bags, containers, and any packaging material suitable for a selected formulation and intended mode of administration and treatment.
[0372] For example, the container includes a composition of the present disclosure, and optionally additionally has a therapeutic regimen or agent disclosed herein.Such kits optionally include an identifying description or label or instructions related to its use in the methods described herein.
[0373] The kit typically includes a label listing the contents and / or instructions for use, and a package insert with instructions for use. A set of instructions will also typically be included.
[0374] In an embodiment, the label is on or associated with the container. In one embodiment, the label is on the container when the letters, numbers or other characters forming the label are attached, molded or etched into the container itself; the label is associated with the container when the label is present in a receptacle or carrier that also holds the container, such as as a package insert. In one embodiment, the label is used to indicate that the contents will be used for a specific therapeutic application. The label also indicates instructions for use of the contents, such as in the methods described herein.
[0375] Example
[0376] The following examples are included for illustrative purposes only and are not intended to limit the scope of the present disclosure.
[0377] Example 1: Fusion proteins with different NLS configurations
[0378] Several improved fusion protein constructs were developed using different nuclear localization sequence (NLS) configurations to have significantly higher apparent silencing activity.
[0379] Several constructs with different NLS domain configurations were constructed (Figure 1) and tested in the Pcsk9 locus in HeLa cells ( Figure 2A-2B ). Also in Hepa1-6 ( Figure 3 ) and HuH7 (Figure 4- Figure 5 ) were tested in the Pcsk9 locus in . Exemplary fusion protein construct DNA and amino acid sequences are found in SEQ ID NOs: 17-24, 28-35, 110-117, and 160-167. Figure 4A The sequences in can be found in SEQ ID NOs: 118-130. Figure 5 The sequences in can be found in SEQ ID NOs: 131-150.
[0380] Cell culture and transfection
[0381] HeLa (ATCC-CRM-CCL-2), Hepa1-6 (PCSK9-IRES-TdTomato) Huh7 (SekisuiXenoTech, LLC) and HEK293T Griptite (CLTA-GFP) cells were cultured in DMEM containing 10% FBS. All experiments in HeLa and Huh7 cells were performed using chemically synthesized guide RNAs and in vitro transcribed effector constructs. HeLa cells were reverse transfected using TransIT-X2 transfection reagent from Mirus (Cat#MIR6003). Huh7 cells were reverse transfected using MessengerMAX reagent from Invitrogen (Cat#LMRNA003). TM The secreted PCSK9 levels were measured at the indicated time points by Human PCSK9 ELISA Kit (Cat#443107). All ELISA data were normalized for cell number using Promega's CellTiter-Glo kit (Cat#G7571).
[0382] HEK293T Griptite cells with GFP knocked into the CLTA locus as an in-frame CLTA fusion were co-transfected with plasmids encoding effector constructs and human CLTA guide RNA using Mirus' TransIT-X2 transfection reagent (Cat#MIR6003). GFP expression was measured by FACS as a surrogate for CLTA expression.
[0383] Hepa1-6 was co-transfected with plasmids encoding effector constructs and mouse PCSK9 guide RNA using the SF cell line 96-well nuclear transfection kit (Cat# V4SC-2096, program code: CM-138) in Lonza's Amaxa 4D nuclear transfection device. At the specified time points, FACS analysis of TdTomato expression of cells was performed as a surrogate for PCSK9 levels.
[0384] In vitro transcription of effector constructs and synthetic gRNA
[0385] CellScript T7 mScript was used according to the manufacturer's instructions. TMThe standard mRNA production system (Cat# C-MSC100625) was used to set up an in vitro transcription reaction using 1 ug of linearized effector template to obtain RNA with a Cap 1 structure on the 5' end and 3' polyadenylated. Terminally modified sgRNAs with three 2'O-methyl modified nucleotides and phosphorothioate bonds at both the 5' and 3' ends were obtained by Integrated DNA technologies.
[0386] Methylation profiling
[0387] Genomic DNA was extracted from each well of a 96-well culture plate using the DNAdvance DNA Extraction from Tissue Kit (Beckman Coulter). After quantification of genomic DNA by the High Sensitivity DNA 1X Kit (Quant-IT), each genomic DNA sample was bisulfite converted using the EZ-96 DNA Methylation-Gold MagPrep Kit (Zymo Research) according to the manufacturer's instructions. For hybridization capture experiments, xGen TM DNA libraries were prepared using the Methyl-Seq DNA Library Prep Kit (IDT) and the xGen TM Hybridization Capture is used for hybridization capture. For amplicon sequencing experiments, xGen TM DNA libraries were prepared using the Methyl-Seq DNA Library Prep Kit (IDT) and the xGen TM Hybridization Capture was used for hybridization capture. Using Platinum Taq kit (Invitrogen), bisulfite-converted DNA from each sample was used to inoculate PCR corresponding to each of the two VIM amplicons. The pooled products were cleaned using AMPure XP kit (Beckman Coulter), and the fragment sizes were assessed by D1000 screening tape (screentape) on Tapestation 4200 (Agilent), followed by sequencing by commercial services (Azenta).
[0388] Example 2: Bacterial DNA methyltransferase
[0389] In this experiment, a panel of bacterial proteins was screened for DNA methyltransferase activity in mammalian cells. Using the experimental procedure of Example 1, the epigenetic silencing activity of these bacterial DNA methyltransferases (Table 3) was tested by fusing their N-termini to the dCas9 domain. These constructs were then transfected into a reporter cell line expressing GFP under the control of the mammalian promoter of CTLA4.
[0390] Table 3: Bacterial DNA methyltransferases
[0391] describe SEQ ID NO DNA-(cytosine C5)-methyltransferase, Arthrobacter luteus (M.AluI) 36 Cytosine-specific methyltransferase, Moraxella species (M.MspI) 37 DNA cytosine methyltransferase, Haemophilus influenzae (M.HaeIII) 38 DNA cytosine methyltransferase, Haemophilus haemolyticus (M.HhaI) 39 DNA cytosine methyltransferase, Monospirillum (M.SssI) 40
[0392] M.SssI DNA methyltransferase can effectively methylate DNA in mammalian cells ( Figure 6A-6D ), and remained stable for up to 30 days. Figure 6A-6D The sequences in can be found in SEQ ID NO: 151-158. The methylation profiles of these cells were also analyzed at day 29, confirming 20% methylation of the target gene ( Figure 7-Figure 8 ).
[0393] Three additional orthologous DNA methyltransferases predicted to be closely related to M.SssI were identified and tested for epigenetic silencing activity using the experimental procedure of Example 1 (Table 4).
[0394] Table 4: Bacterial methyltransferases
[0395] describe SEQ ID NO DNA cytosine methyltransferase Mycoplasmaales bacteria 41 DNA cytosine methyltransferase Mycoplasma marineum 42 DNA(cytosine-5-)-methyltransferase Spiroplasma sinensis 43
[0396] The DNA methyltransferases of Table 4 were predicted to have similar or improved functions to M.SssI. The sequences were tested in the context of CRISPR-off in place of murine DNMT3A / DNMT3L and their functions were compared to those of the M.SssI DNA methyltransferase in silencing the Pcsk9 locus in the HeLaTdTomato system to identify novel features and improved functions.
[0397] Example 3: Alternative KRAB domains
[0398] In this example, fusion proteins were constructed with alternative KRAB domains (Table 5) and, when tested using the experimental procedure of Example 1, exhibited improved activity compared to CRISPRoff (Figure 9).
[0399] Table 5: Alternative KRAB domains
[0400] describe SEQ ID NO ZFP28 44 ZN627 45 KAP1 46 MeCP2 47 HP1b 48 CBX8 49 CDYL2 50 TOX 51 TOX3 52 TOX4 53 EED 54 RBBP4 55 RCOR1 56 SCML2 57
[0401] Example 4: ZIM3 fusion constructs
[0402] Novel fusions of ZIM3 and KOX1KRAB were generated. Both ZIM 3 and KOX1KRAB are KRAB family proteins with extensive homology. Therefore, sequences representing the midpoint of ZIM3 and KOX1KRAB were designed. These KOX1KRAB and ZIM3 constructs encode small regions of KOX1KRAB and ZIM3 that are concentrated around the zinc finger domains of the proteins. While the regions used by KOX1KRAB and ZIM3 are very similar within the first approximately 75 bp of their sequences, ZIM3 also has a small alpha helical region at the C-terminus that is not present in KOX1KRAB. The KOX1KRAB-FL sequence includes the KOX1KRAB sequence equivalent to this additional segment, while the ZIM3 truncate has this additional segment removed from the ZIM3 sequence. The ZIM3 / KOX1KRAB chimera is an N-terminal and C-terminal fusion of the two proteins. ZIM3-like KOX1KRAB variants were all assembled in the following manner: first, BLAST was performed on ZIM3 or KOX1KRAB proteins from non-human species to assemble the 100 closest homologs ("families") for each gene; second, the three members of the KOX1KRAB family that were most similar to ZIM3 and the three members of the ZIM3 family that were most similar to KOX1KRAB were identified; third, the KOX1KRAB-FL sequence was rationally modified to make it similar to each group of three sequences (Table 6).
[0403] Table 6: ZIM-KOX1KRAB chimeric proteins
[0404]
[0405] Although preferred embodiments of the present disclosure have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Many variations, changes, and substitutions will be apparent to those skilled in the art without departing from the present disclosure. It should be understood that various alternatives to the embodiments of the present disclosure described herein may be employed in implementing the present disclosure. The appended claims are intended to define the scope of the present disclosure, and thus encompass methods and structures within the scope of these claims and their equivalents.
[0406] Sequence Listing
[0407]
[0408]
[0409]
[0410]
[0411]
[0412]
[0413]
[0414]
[0415]
[0416]
[0417]
[0418]
[0419]
[0420]
[0421]
[0422]
[0423]
[0424]
[0425]
[0426]
[0427]
[0428]
[0429]
[0430]
[0431]
[0432]
[0433]
[0434]
[0435]
[0436]
[0437]
[0438]
[0439]
[0440]
[0441]
[0442]
[0443]
[0444]
[0445]
[0446]
[0447]
[0448]
[0449]
[0450]
[0451]
[0452]
[0453]
[0454]
[0455]
[0456]
[0457]
[0458]
[0459]
[0460]
[0461]
[0462]
[0463]
[0464]
[0465]
[0466]
[0467]
[0468]
[0469]
[0470]
[0471]
[0472]
[0473]
[0474]
[0475]
[0476]
[0477]
[0478]
[0479]
[0480]
[0481]
[0482]
[0483]
[0484]
[0485]
[0486]
[0487]
[0488]
[0489]
[0490]
[0491]
[0492]
[0493]
[0494]
[0495]
[0496]
[0497]
[0498]
[0499]
Claims
1. An epigenetic editing system, comprising: (a) a fusion protein comprising DNA binding domain; DNA methyltransferase (DNMT) domain; repressor domain; and Two nuclear localization sequences (NLS), wherein each of the two NLS is located at the amino (N) terminus or the carboxyl (C) terminus of the fusion protein; or (b) A nucleic acid molecule encoding the fusion protein of (a).
2. The epigenetic editing system according to claim 1, wherein the fusion protein comprises one or more NLS at the C-terminus of the fusion protein and one or more NLS at the N-terminus of the fusion protein, optionally wherein the fusion protein comprises two NLS at the N-terminus of the fusion protein and two NLS at the C-terminus of the fusion protein.
3. The epigenetic editing system of claim 1 or 2, wherein the DNMT domain is from a bacterial species.
4. The epigenetic editing system of claim 1 or 2, wherein the DNMT domain is a mammalian DNMT domain.
5. The epigenetic editing system of claim 1 or 2, wherein the DNMT domain is a mouse DNMT domain.
6. The epigenetic editing system of claim 1 or 2, wherein the DNMT domain is a human DNMT domain.
7. The epigenetic editing system according to claim 1 or 2, wherein the DNMT domain is a DNMT3A domain.
8. The epigenetic editing system according to claim 1 or 2, wherein the DNMT domain is a DNMT3L domain.
9. The epigenetic editing system according to claim 1 or 2, wherein the fusion protein comprises a DNMT3A domain and a DNMT3L domain.
10. The epigenetic editing system according to claim 8 or 9, wherein the DNMT3L domain is from a species selected from the group consisting of giant panda, Philippine tarsier, long-clawed gerbil, North American pika, Carolina squirrel, American bison, Przewalski's horse, Mouse of Castor and chimpanzee; the DNMT3L domain optionally comprises one of SEQ ID NOs: 72-80 and SEQ ID NOs: 101-109, or an amino acid sequence that is at least 90%, optionally at least 95% homologous thereto.
11. The epigenetic editing system according to any one of the preceding claims, wherein the repressor domain comprises The KRAB domain of the ZFP28, ZN627, KAP1, MeCP2, HP1b, CBX8, CDYL2, TOX, Tox3, Tox4, EED, RBBP4, RCOR1, or SCML2 protein, or Fusion of the N-terminal and C-terminal regions of ZIM3 and the KOX1 KRAB domain.
12. An epigenetic editing system comprising: (a) a fusion protein comprising: DNA binding domain; and a DNA methyltransferase (DNMT) domain from a bacterial species, wherein the DNMT domain is fused to the N-terminus of the DNA binding domain; or (b) A nucleic acid molecule encoding the fusion protein of (a).
13. The epigenetic system of claim 12, wherein the fusion protein functions to methylate a mammalian target DNA in a cell.
14. The epigenetic editing system according to claim 12 or 13, wherein the DNMT domain of the fusion protein does not contain any one of SEQ ID NOs: 81-93.
15. The epigenetic editing system of claim 12 or 13, wherein the bacterial species is not Mycoplasma penetrans, Spirillum monocotyledonum, Haemophilus parainfluenzae, Arthrobacter luteus, Haemophilus aegypti, Haemophilus haemolyticus, Moraxella, Escherichia coli, Thermus aquaticus, Caulobacter crescentus, or Clostridium difficile.
16. The epigenetic editing system of any one of claims 3, 11, and 12-15, wherein the bacterial species is a Mycoplasmalales bacterium, Mycoplasma marineum, or Spiroplasma sinensis.
17. The epigenetic editing system according to any one of claims 3, 11 and 12-15, wherein the DNMT domain from the bacterial species is derived from (a) M.Sss1, optionally comprising SEQ ID NO:40 or an amino acid sequence at least 90%, optionally at least 95% homologous thereto; (b) NQZ29229, optionally comprising SEQ ID NO:41 or an amino acid sequence at least 90%, optionally at least 95% homologous thereto; (c) WP_131599610, optionally comprising SEQ ID NO: 42, or an amino acid sequence at least 90%, optionally at least 95% homologous thereto; or (d) WP_208057179, optionally comprising SEQ ID NO: 43 or an amino acid sequence at least 90%, optionally at least 95% homologous thereto.
18. An epigenetic editing system comprising: (a) a fusion protein comprising: DNA binding domain, and A DNMT3L domain, wherein the DNMT3L domain is from a species selected from the group consisting of a giant panda, a Philippine tarsier, a long-clawed gerbil, a North American pika, a Carolina squirrel, an American bison, a Przewalski's horse, a Mouse of the Carolinian genus, and a chimpanzee; the DNMT3L domain optionally comprises one of SEQ ID NOs: 72-80 and SEQ ID NOs: 101-109, or an amino acid sequence that is at least 90%, optionally at least 95%, homologous thereto; or (b) A nucleic acid molecule encoding the fusion protein of (a).
19. The epigenetic editing system according to any one of claims 12 to 18, wherein (i) the fusion protein further comprises a repressor domain, or (ii) The system further comprises other fusion proteins comprising a DNA binding domain and a repressor domain, or a nucleic acid molecule encoding the other fusion protein.
20. The epigenetic editing system of any one of claims 1-10 and 19, wherein the repressor domain comprises a KRAB domain, optionally derived from KOX1, ZIM3, ZFP28 or ZN627.
21. The epigenetic editing system according to claim 20, wherein the KRAB domain is derived from Human KOX1, optionally comprising SEQ ID NO: 94 or an amino acid sequence at least 90%, optionally at least 95% homologous thereto, or Human ZIM3, optionally comprising SEQ ID NO: 100, or an amino acid sequence at least 90%, optionally at least 95% homologous thereto.
22. The epigenetic editing system according to any one of claims 1-10 and 19, wherein the repressor domain is derived from KAP1, MECP2, HP1a / CBX5, HP1b, CBX8, CDYL2, TOX, TOX3, TOX4, EED, EZH2, RBBP4, RCOR1 or SCML2.
23. An epigenetic editing system comprising: (a) a fusion protein comprising: DNA binding domain, and A repressor domain, wherein the repressor domain comprises: The KRAB domain of the ZFP28, ZN627, KAP1, MeCP2, HP1b, CBX8, CDYL2, TOX, Tox3, Tox4, EED, RBBP4, RCOR1, or SCML2 protein, or A fusion of the N-terminal and C-terminal regions of ZIM3 and KOX1 KRAB; or (b) A nucleic acid molecule encoding the fusion protein of (a).
24. The epigenetic editing system of claim 11 or 23, wherein the repressor domain comprises one of SEQ ID NOs: 44-57, or an amino acid sequence at least 90%, optionally at least 95% homologous thereto.
25. The epigenetic editing system according to claim 23 or 24, wherein The fusion protein further comprises a DNA methyltransferase (DNMT) domain, or The system also includes other fusion proteins comprising a DNA binding domain and a DNMT domain, or nucleic acid molecules encoding the other fusion proteins.
26. The epigenetic editing system of any one of claims 18-25, wherein the fusion protein or the system comprises a DNMT3A domain of a mammal, optionally a human or a mouse, and a DNMT3L domain of a mammal, optionally a human or a mouse.
27. The epigenetic editing system according to claim 26, wherein the fusion protein or the system comprises a human DNMT3A domain and a human DNMT3L domain, or a human DNMT3A domain and a mouse DNMT3L domain, optionally (a) wherein the human DNMT3A domain comprises SEQ ID NO: 12, or an amino acid sequence that is at least 90%, optionally at least 95% homologous thereto, (b) wherein the human DNMT3L domain comprises SEQ ID NO: 13, or an amino acid sequence that is at least 90%, optionally at least 95% homologous thereto, and / or (c) wherein the mouse DNMT3L domain comprises SEQ ID NO: 15, or an amino acid sequence that is at least 90%, optionally at least 95% homologous thereto.
28. The epigenetic editing system according to claim 26 or 27, comprising a fusion protein comprising one or more NLSs, the DNMT3A domain, the DNMT3L domain, the DNA binding domain, the repressor domain and one or more NLSs from the N-terminus to the C-terminus, optionally wherein the fusion protein comprises a peptide linker between adjacent domains.
29. The epigenetic editing system of claim 28, wherein the fusion protein comprises, from N-terminus to C-terminus, two NLSs, the DNMT3A domain, the ADD, the DNMT3L domain, a first peptide linker, the DNA binding domain, a second peptide linker, the repressor domain and two NLSs.
30. The epigenetic editing system of claim 29, wherein the DNMT3A domain is a human DNMT3A domain and / or the DNMT3L domain is a human DNMT3L domain.
31. The epigenetic editing system of claim 30, wherein the repressor domain is a KRAB domain from ZFP28, ZF627 or KOX1 from a mammal, optionally a human.
32. The epigenetic editing system of claim 31, wherein the first peptide linker is XTEN80 (SEQ ID NO: 3) and / or the second peptide linker is XTEN16 (SEQ ID NO: 2).
33. The epigenetic editing system of claim 32, wherein the system comprises an expression construct encoding the fusion protein, wherein the expression construct comprises a WPRE sequence located in the 3' non-coding region and upstream of the polyadenylation site.
34. The epigenetic editing system of any one of claims 1-33, wherein the DNA binding domain is a dCas9 domain.
35. The epigenetic editing system of claim 34, wherein the dCas9 domain comprises SEQ ID NO: 9, or an amino acid sequence at least 90%, optionally at least 95% homologous thereto.
36. The epigenetic editing system of claim 34 or 35, further comprising one or more guide RNAs (gRNAs) or nucleic acid molecules encoding the gRNAs.
37. The epigenetic editing system of any one of claims 1-33, wherein the DNA binding domain is a zinc finger protein (ZFP) domain.
38. A method for modifying the epigenetic state of a target gene in a mammalian cell, comprising contacting the cell with an epigenetic editing system according to any one of claims 1-37.
39. A method of regulating target gene expression in a mammalian cell, comprising contacting the cell with an epigenetic editing system according to any one of claims 1-37.
40. A method of treating a disease in a subject in need thereof, comprising administering to the subject the epigenetic editing system of any one of claims 1-37.
Citation Information
Patent Citations
DNA methylation editing kit and DNA methylation editing method
US10612044B2
RNA-guided nucleases and active fragments and variants thereof and methods of use
US11162114B2
Delivery system for functional nucleases
US20160200779A1
Switchable cas9 nucleases and uses thereof
US20160208288A1
A protein tagging system for in vivo single molecule imaging and control of gene transcription
US20170219596A1