Methods and Compositions for Modulating Gene Expression
Site-specific agents targeting anchor sequence-mediated connections provide a means to modulate gene expression, addressing the limitations of current technologies and offering a potential therapeutic approach for disease treatment.
Patent Information
- Application Number
- JP2019533313
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2017-08-08
- Filing Date
- 2017-09-07
- Publication Date
- 2025-05-26
- Estimated Expiration
- 2037-09-07
AI Technical Summary
Current technologies lack effective methods for modulating gene expression by specifically disrupting or modifying anchor sequence-mediated conjunctions, which are crucial in regulating gene expression and treating diseases.
The development of site-specific agents, including DNA-binding moieties and fusion molecules, that target and disrupt or modify anchor sequence-mediated connections, allowing for precise modulation of gene expression.
These site-specific agents effectively compete with endogenous nucleation polypeptides, allowing for controlled modulation of gene expression, which can be used to treat diseases by altering the transcription of target genes.
Smart Images

Figure 0007682604000017 
Figure 0007682604000018 
Figure 0007682604000019
Abstract
Description
[Technical Field]
[0001] Citation of Related Applications This application claims priority to and the benefit of U.S. Provisional Application Nos. 62 / 384,603 (filed September 7, 2016), 62 / 416,501 (filed November 2, 2016), 62 / 439,327 (filed December 27, 2016), and 62 / 542,703 (filed August 8, 2017), the contents of each of which are incorporated herein by reference. [Background technology]
[0002] background Many diseases are caused by defects in the regulation of the expression of certain genes. Summary of the Invention [Means for solving the problem]
[0003] overview In particular, the present disclosure provides various agents, compositions, and methods for modulating gene expression, delivery to cells (e.g., mammalian cells, such as mammalian somatic cells; e.g., delivery across cell membranes), and related treatment methods. To the inventors' knowledge, the present disclosure provides the first disclosure of site-specific agents that physically disrupt and / or modify anchor sequence-mediated connections. The present disclosure also provides, inter alia, site-specific agents that act to disrupt and / or modify anchor sequence-mediated connections by genetic and / or epigenetic methods.
[0004] In some embodiments, the present disclosure provides a site-specific disrupting agent comprising a DNA-binding moiety that specifically binds to one or more target anchor sequences in a cell with sufficient affinity to compete with the binding of endogenous nucleation polypeptides in the cell, and does not bind to non-target anchor sequences in the cell.
[0005] In some embodiments, the present disclosure provides a method of modulating expression of a gene within an anchor sequence-mediated junction comprising a first anchor sequence and a second anchor sequence, the method comprising contacting the first and / or second anchor sequence with a site-specific disrupting agent disclosed herein.
[0006] In some embodiments, the present disclosure provides a method of modulating expression of a gene within 10 kb of a first anchor sequence in an anchor sequence-mediated junction comprising a first anchor sequence and a second anchor sequence, the method comprising contacting the first and / or second anchor sequence with a site-specific disrupting agent disclosed herein.
[0007] In some embodiments, the present disclosure provides a method of increasing expression of a gene within an anchor sequence-mediated junction comprising a first anchor sequence and a second anchor sequence, wherein the first and / or second anchor sequence is located within 10 kb of an external enhancing sequence, the method comprising contacting the first and / or second anchor sequence with a site-specific disrupting agent disclosed herein.
[0008] In some embodiments, the present disclosure provides a method comprising delivering a site-specific disrupting agent disclosed herein to a mammalian cell.
[0009] In some embodiments, the present disclosure provides a fusion molecule comprising: (i) a site-specific targeting moiety; and (ii) a deaminating agent, wherein the site-specific targeting moiety targets the fusion molecule to a target anchor sequence but not to at least one non-target anchor sequence.
[0010] In some embodiments, the disclosure provides a composition comprising: (i) a fusion polypeptide, or a nucleic acid encoding the fusion polypeptide, comprising an enzymatically inactive Cas polypeptide and a deaminating agent; and (ii) a guide RNA that targets the fusion polypeptide to a target anchor sequence but not to at least one non-target anchor sequence.
[0011] In some embodiments, the present disclosure provides a method of modulating expression of a gene within an anchor sequence-mediated junction comprising a first anchor sequence and a second anchor sequence, the method comprising contacting the first and / or second anchor sequence with a site-specific disrupting agent disclosed herein.
[0012] In some embodiments, the present disclosure provides a method of modulating expression of a gene within 10 kb of a first anchor sequence in an anchor sequence-mediated junction comprising a first anchor sequence and a second anchor sequence, the method comprising contacting the first and / or second anchor sequence with a site-specific disrupting agent disclosed herein.
[0013] In some embodiments, the present disclosure provides a method of reducing expression of a gene within an anchor sequence-mediated junction comprising a first anchor sequence, a second anchor sequence, and an internal enhancing sequence, the method comprising contacting the first and / or second anchor sequence with a site-specific disrupting agent disclosed herein.
[0014] In some embodiments, the present disclosure provides a method comprising the step of: (a) delivering a fusion molecule or composition disclosed herein to a mammalian cell.
[0015] In some embodiments, the disclosure provides a method comprising the step of: (a) substituting, adding, or deleting one or more nucleotides of an anchor sequence in a mammalian somatic cell.
[0016] In some embodiments, the present disclosure provides a method comprising delivering mammalian somatic cells to a subject having a disease or condition, wherein one or more nucleotides of an anchor sequence in the mammalian somatic cells have been substituted, added, or deleted.
[0017] In some embodiments, the present disclosure provides a method comprising the step of (a) administering to a subject mammalian somatic cells, wherein the mammalian somatic cells are obtained from the subject, and wherein a fusion molecule or composition disclosed herein is delivered to the mammalian cells ex vivo.
[0018] In some embodiments, the present disclosure provides fusion molecules comprising (i) a site-specific targeting moiety and (ii) an epigenetic modifier, wherein the site-specific targeting moiety targets the fusion molecule to a target anchor sequence but not to at least one non-target anchor sequence.
[0019] In some embodiments, the present disclosure provides a site-specific guide RNA that comprises a targeting domain that is complementary to a target nucleic acid that comprises an anchor sequence.
[0020] In some embodiments, the present disclosure provides a composition comprising: (i) a fusion polypeptide, or a nucleic acid encoding the fusion polypeptide, comprising an enzymatically inactive Cas polypeptide and an epigenetic modifier; and (ii) a guide RNA that targets the fusion polypeptide to a target anchor sequence but not to at least one non-target anchor sequence.
[0021] In some embodiments, the present disclosure provides a method of modulating expression of a gene within an anchor sequence-mediated junction comprising a first anchor sequence and a second anchor sequence, the method comprising contacting the first and / or second anchor sequence with a fusion molecule or composition disclosed herein.
[0022] In some embodiments, the present disclosure provides a method for modulating expression of a gene within 10 kb of a first anchor sequence in an anchor sequence-mediated junction comprising a first anchor sequence and a second anchor sequence, the method comprising contacting the first and / or second anchor sequence with a fusion molecule or composition disclosed herein.
[0023] In some embodiments, the present disclosure provides a method for reducing expression of a gene within an anchor sequence-mediated junction comprising a first anchor sequence, a second anchor sequence, and an internal enhancing sequence, the method comprising contacting the first and / or second anchor sequence with a fusion molecule or composition disclosed herein.
[0024] In some embodiments, the present disclosure provides a method for increasing expression of a gene within an anchor sequence-mediated junction comprising a first anchor sequence and a second anchor sequence, wherein the first and / or second anchor sequence is located within 10 kb of an external enhancing sequence, the method comprising contacting the first and / or second anchor sequence with a fusion molecule or composition disclosed herein.
[0025] In some embodiments, the present disclosure provides a method comprising the step of: (a) delivering a fusion molecule or composition disclosed herein to a mammalian cell.
[0026] In some embodiments, the present disclosure provides an engineered site-specific nucleation agent comprising: an engineered DNA-binding moiety that specifically binds to one or more target sequences within a cell with sufficient affinity to compete with the binding of an endogenous nucleating polypeptide within the cell, and does not bind to non-target sequences within the cell; and a nucleating polypeptide dimerization domain associated with the engineered DNA-binding moiety, wherein when the engineered DNA-binding moiety binds to the at least one target sequence, the nucleating polypeptide dimerization domain is localized thereto, and each of the at least one target sequence is a target anchor sequence, and the at least one or more target anchor sequences are positioned relative to the anchor sequence to which the nucleating polypeptide binds such that when the nucleating polypeptide dimerization domain is localized to the target anchor sequence, interaction between the nucleating polypeptide dimerization domain and the nucleating polypeptide generates an anchor sequence-mediated connection.
[0027] In one aspect, the disclosure includes a pharmaceutical preparation comprising a composition that binds to an anchor sequence of an anchor sequence-mediated junction and alters the formation of the anchor sequence-mediated junction, wherein the composition modulates transcription of a target gene associated with the anchor sequence-mediated junction in a human cell.
[0028] In one aspect, the present disclosure includes compositions comprising a targeting moiety that binds to the anchor sequence of an anchor sequence-mediated connection and alters the formation of the anchor sequence-mediated connection (e.g., alters the affinity of the anchor sequence for a connection nucleating molecule by, e.g., at least 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or more).
[0029] In one aspect, the disclosure includes a pharmaceutical preparation comprising a composition comprising a targeting moiety that binds to the anchor sequence of an anchor sequence-mediated connection and alters formation of the anchor sequence-mediated connection, e.g., the composition modulates transcription of a target gene in an expression unit associated with the anchor sequence-mediated connection in, e.g., a human cell.
[0030] In various aspects of the present disclosure described herein, one or more of the various embodiments described herein may be combined.
[0031] In some embodiments, the targeting moiety (i) is a chemical, e.g., a chemical that modulates cytosine (C) or adenine (A) (e.g., sodium bisulfite, ammonium bisulfite); (ii) has enzymatic activity (methyltransferase, demethylase, nuclease (e.g., Cas9), deaminase); or (iii) comprises an effector moiety that sterically hinders formation of the anchor sequence-mediated connection (e.g., membrane translocating polypeptide + nanoparticle).
[0032] In some embodiments, the anchor sequence-mediated connection is associated with one or more transcription control sequences. In one embodiment, one or more transcription control sequences are located inside the anchor sequence-mediated connection, for example, a type 1 anchor sequence-mediated connection. In another embodiment, one or more transcription control sequences are located outside the anchor sequence-mediated connection, including, for example, a type 2 anchor sequence-mediated connection. In another embodiment, one or more transcription control sequences are located inside, for example, an enhancing sequence, and at least partially outside, for example, a silencing sequence, an anchor sequence-mediated connection, for example, a type 3 anchor sequence-mediated connection. In another embodiment, one or more transcription control sequences are located inside, for example, an enhancing sequence, and at least partially outside, for example, an enhancing sequence, an anchor sequence-mediated connection, for example, a type 4 anchor sequence-mediated connection.
[0033] In some embodiments, the composition disrupts the formation of anchor sequence-mediated connections (e.g., reduces the affinity of the anchor sequence for the connection nucleation molecule by, e.g., at least 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or more). In some embodiments, the composition promotes the formation of anchor sequence-mediated connections (e.g., increases the affinity of the anchor sequence for the connection nucleation molecule by, e.g., at least 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or more). In some embodiments, the target gene is within the anchor sequence-mediated connection. In some embodiments, the target gene is located outside the anchor sequence-mediated connection. In some embodiments, the target gene is located both inside and outside the anchor sequence-mediated connection. In some embodiments, the composition physically disrupts the formation of the anchor sequence-mediated connection, e.g., the composition is both a targeting and an effector, e.g., a transmembrane polypeptide. In some embodiments, the composition comprises a targeting moiety (e.g., gRNA, transmembrane polypeptide) that binds to the anchor sequence, operably linked to an effector moiety that modulates the formation of the anchor sequence-mediated connection. In some embodiments, the effector moiety is a chemical, e.g., a chemical that modulates cytosine (C) or adenine (A) (e.g., sodium bisulfite, ammonium bisulfite). In some embodiments, the effector moiety has enzymatic activity (e.g., methyltransferase, demethylase, nuclease (e.g., Cas9), deaminase). In some embodiments, the effector moiety sterically hinders the formation of the anchor sequence-mediated connection, e.g., the transmembrane polypeptide and / or nanoparticle.
[0034] In some embodiments, the compositions or methods described herein comprise ABX nand C (wherein A is selected from a hydrophobic amino acid or an amide-containing backbone, e.g., aminoethyl-glycine, having nucleic acid side chains; B and C can be the same or different and are each independently selected from arginine, asparagine, glutamine, lysine, and analogs thereof; each X is independently a hydrophobic amino acid, or each X is independently an amide-containing backbone, e.g., aminoethyl-glycine, having nucleic acid side chains; and n is an integer from 1 to 4), which further includes at least one polypeptide that hybridizes to a nucleic acid sequence within the anchor sequence-mediated connection (e.g., an anchor sequence of the anchor sequence-mediated connection, e.g., a CTCF-binding motif, a BORIS-binding motif, a cohesin-binding motif, a USF1-binding motif, a YY1-binding motif, a TATA box, a ZNF143-binding motif, etc.).
[0035] The compositions and methods described in various embodiments of the above aspects can be utilized in any other aspect described herein.
[0036] In one aspect, the present disclosure includes methods of modulating expression of a target gene in an anchor sequence-mediated connection, including targeting a sequence that is external to or not part of the target gene or its associated transcriptional control sequences that affect transcription of the gene, such as modulating expression of the gene by targeting an anchor sequence.
[0037] In one aspect, the disclosure includes a method of modulating transcription of a target gene comprising targeting a sequence that is discontinuous with the target gene or its associated transcriptional control sequence that affects transcription of the target gene, such as by targeting an anchor sequence to alter the formation of an anchor sequence-mediated junction.
[0038] In some embodiments, the method includes an anchor sequence-mediated connection including one or more associated genes and one or more transcriptional regulatory sequences within the anchor sequence-mediated connection. In some embodiments, the anchor sequence-mediated connection includes one or more associated genes and one or more transcriptional regulatory sequences located outside the anchor sequence-mediated connection. In some embodiments, the anchor sequence-mediated connection includes one or more associated genes and one or more transcriptional regulatory sequences located at least partially inside and outside the anchor sequence-mediated connection. For example, one or more inhibitory signals may be located outside the anchor sequence-mediated connection, and one or more enhancing sequences and target genes may be located inside the anchor sequence-mediated connection. In another example, one or more enhancing sequences may be located inside and outside the anchor sequence-mediated connection.
[0039] In some embodiments, the target gene is non-contiguous with one or more anchor sequences. In some embodiments in which the gene is non-contiguous with the anchor sequence, the gene may be separated from the anchor sequence by about 100 bp to about 500 Mb, about 500 bp to about 200 Mb, about 1 kb to about 100 Mb, about 25 kb to about 50 Mb, about 50 kb to about 1 Mb, about 100 kb to about 750 kb, about 150 kb to about 500 kb, or about 175 kb to about 500 kb. In some embodiments, the gene is about 100 bp, 300 bp, 500 bp, 600 bp, 700 bp, 800 bp, 900 bp, 1 kb, 5 kb, 10 kb, 15 kb, 20 kb, 25 kb, 30 kb, 35 kb, 40 kb, 45 kb, 50 kb, 55 kb, 60 kb, 65 kb, 70 kb, 75 kb, 80 kb, 85 kb, 90 kb, 95 kb, 100 kb, 125 kb, 150 kb, 175 kb, 200 kb, 22 The distance from the anchor sequence may be 5 kb, 250 kb, 275 kb, 300 kb, 350 kb, 400 kb, 500 kb, 600 kb, 700 kb, 800 kb, 900 kb, 1 Mb, 2 Mb, 3 Mb, 4 Mb, 5 Mb, 6 Mb, 7 Mb, 8 Mb, 9 Mb, 10 Mb, 15 Mb, 20 Mb, 25 Mb, 50 Mb, 75 Mb, 100 Mb, 200 Mb, 300 Mb, 400 Mb, 500 Mb, or any size in between.
[0040] In some embodiments, anchor sequence mediated connection comprises target gene and is associated with one or more transcriptional regulatory sequences, for example, silencing / suppression sequence and enhancing sequence.In some embodiments, anchor sequence mediated connection comprises one or more, for example, 2, 3, 4, 5 or more genes.In some embodiments, anchor sequence mediated connection is associated with one or more, for example, 2, 3, 4, 5 or more transcriptional regulatory sequences.
[0041] In some embodiments, the target gene is non-contiguous with one or more transcription control sequences. In some embodiments in which the gene is non-contiguous with the transcription control sequence, the gene may be separated from the transcription control sequence by about 100 bp to about 500 Mb, about 500 bp to about 200 Mb, about 1 kb to about 100 Mb, about 25 kb to about 50 Mb, about 50 kb to about 1 Mb, about 100 kb to about 750 kb, about 150 kb to about 500 kb, or about 175 kb to about 500 kb. In some embodiments, the gene is about 100 bp, 300 bp, 500 bp, 600 bp, 700 bp, 800 bp, 900 bp, 1 kb, 5 kb, 10 kb, 15 kb, 20 kb, 25 kb, 30 kb, 35 kb, 40 kb, 45 kb, 50 kb, 55 kb, 60 kb, 65 kb, 70 kb, 75 kb, 80 kb, 85 kb, 90 kb, 95 kb, 100 kb, 125 kb, 150 kb, 175 kb, 200 kb, 22 The distance from the transcriptional regulatory sequence may be 5 kb, 250 kb, 275 kb, 300 kb, 350 kb, 400 kb, 500 kb, 600 kb, 700 kb, 800 kb, 900 kb, 1 Mb, 2 Mb, 3 Mb, 4 Mb, 5 Mb, 6 Mb, 7 Mb, 8 Mb, 9 Mb, 10 Mb, 15 Mb, 20 Mb, 25 Mb, 50 Mb, 75 Mb, 100 Mb, 200 Mb, 300 Mb, 400 Mb, 500 Mb, or any size in between.
[0042] In one aspect, the disclosure includes a pharmaceutical composition comprising: (a) a targeting moiety; and (b) a DNA sequence, including, for example, an anchor sequence.
[0043] In one aspect, the present disclosure includes compositions comprising a targeting moiety that binds to the anchor sequence of an anchor sequence-mediated connection and alters the formation of the anchor sequence-mediated connection (e.g., alters the affinity of the anchor sequence for a connection nucleating molecule by, e.g., at least 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or more).
[0044] In one aspect, the disclosure includes a protein comprising a domain that acts on DNA, e.g., an enzymatic domain (e.g., a nuclease domain, e.g., a Cas9 domain, e.g., a dCas9 domain; a DNA methyltransferase, a demethylase, a deaminase), in combination with at least one guide RNA (gRNA) or antisense DNA oligonucleotide that targets the protein to an anchor sequence of a target anchor sequence-mediated junction, wherein the composition is effective to alter the target anchor sequence-mediated junction in a human cell.
[0045] In some embodiments, the enzymatic domain is Cas9 or dCas9. In some embodiments, the protein comprises two enzymatic domains, e.g., dCas9 and a methylase or demethylase domain.
[0046] The compositions described in various embodiments of the above aspects can be utilized in any other aspect described herein.
[0047] In one aspect, the disclosure includes a composition for introducing a targeted alteration into an anchor sequence-mediated junction to modulate transcription of a nucleic acid sequence, the composition comprising a targeting moiety that binds to the anchor sequence.
[0048] In some embodiments, the targeting moiety comprises a sequence-targeting polypeptide such as an enzyme, e.g., Cas9. In some embodiments, the targeting moiety comprises a fusion of a sequence-targeting polypeptide and a connection nucleation molecule, e.g., a fusion of dCas9 and a connection nucleation molecule. In some further embodiments, the targeting moiety further comprises a guide RNA or a nucleic acid encoding the guide RNA. In some further embodiments, the targeting moiety targets one or more nucleotides of the anchor sequence within the anchor sequence-mediated connection for substitution, addition, or deletion via CRISPR, TALEN, dCas9, recombination, transposon, etc. In some embodiments, the targeting moiety targets one or more DNA methylation sites within the anchor sequence-mediated connection. In some further embodiments, the targeting moiety introduces at least one of the following: at least one exogenous anchor sequence; an alteration in at least one attached nucleation molecule binding site, such as by changing the binding affinity for the attached nucleation molecule; a change in the orientation of at least one consensus nucleotide sequence, such as a CTCF binding motif, a YY1 binding motif, a ZNF143 binding motif, or other binding motif described herein; and a substitution, addition, or deletion in at least one anchor sequence, such as a CTCF binding motif, a YY1 binding motif, a ZNF143 binding motif, or other binding motif described herein.
[0049] In certain embodiments, the composition alters chromatin structure.
[0050] In some embodiments, the composition comprises a vector comprising a targeting moiety, such as a viral vector, e.g., a lentiviral vector.
[0051] In certain embodiments, the targeted modification alters at least one of the binding sites for the splice nucleating molecule, such as binding affinity for an anchor sequence within an anchor sequence-mediated splice, an alternative splice site, and a binding site for untranslated RNA.
[0052] In some embodiments, the present disclosure includes a pharmaceutical composition comprising a composition described herein.
[0053] The compositions described in various embodiments of the above aspects can be utilized in any other aspect described herein.
[0054] In one aspect, the disclosure includes a composition comprising a synthetic junction nucleating molecule that has a selected binding affinity for an anchor sequence within a target anchor sequence mediated junction.
[0055] In some embodiments, the binding affinity may be at least 10%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or more higher or lower than the affinity of the endogenous conjugated nucleating molecule associated with the target anchor sequence. In some embodiments, the synthetic conjugated nucleating molecule has about 30-90%, about 30-85%, about 30-80%, about 30-70%, about 50-80%, or about 50-90% amino acid sequence identity to the endogenous conjugated nucleating molecule.
[0056] In some embodiments, the connected nucleation molecule disrupts the binding of an endogenous connected nucleation molecule to its binding site, such as through competitive binding. In some further embodiments, the connected nucleation molecule is engineered to bind to a target sequence.
[0057] In some embodiments, the composition further comprises a carrier such as a polymeric carrier or a targeting moiety, for example, a liposome, a peptide, an aptamer, or a combination thereof.
[0058] In certain embodiments, the present disclosure includes methods for preparing conjugated nucleating molecules with selected binding affinities.
[0059] The compositions described in various embodiments of the above aspects can be utilized in any other aspect described herein.
[0060] In one aspect, the disclosure includes compositions comprising targeting moieties that bind to specific anchor sequence-mediated connections and alter the topology of the anchor sequence-mediated connections.
[0061] In some embodiments, the targeting moiety is a nucleic acid sequence, a protein, a protein fusion, or a membrane-translocating polypeptide. In some embodiments, the nucleic acid sequence is selected from the group consisting of a gRNA and a sequence complementary to an anchor sequence or a sequence comprising at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% complementary thereto. In some embodiments, the nucleic acid sequence comprises a sequence complementary to a binding motif for a connection nucleation molecule or a consensus sequence or a sequence comprising at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% complementary thereto. In some embodiments, the protein is a connection nucleation molecule, such as CTCF, cohesin, USF1, YY1, TAF3, ZNF143, or another polypeptide, a dominant-negative connection nucleation molecule, a protein comprising a DNA-binding sequence, such as a transcription factor, or a fusion of a sequence-targeting polypeptide and a connection nucleation molecule. In some embodiments, the membrane-translocating polypeptide is ABX. nC (wherein A is selected from a hydrophobic amino acid or an amide-containing backbone, e.g., aminoethyl-glycine, having nucleic acid side chains; B and C may be the same or different and are independently selected from arginine, asparagine, glutamine, lysine, and analogs thereof; each X is independently a hydrophobic amino acid or each X is independently an amide-containing backbone, e.g., aminoethyl-glycine, having nucleic acid side chains; and n is an integer from 1 to 4). In some embodiments, the protein is an epigenetic enzyme (e.g., DNA methylase (e.g., DNMT3a, DNMT3b, DNMTL), DNA demethylase (e.g., TET family), histone methyltransferase, histone deacetylase (e.g., HDAC1, HDAC2, HDAC3), sirtuin 1, 2, 3, 4, 5, 6, or 7, lysine-specific histone demethylase 1 (LSD1), histone-lysine-N-methyltransferase (SSD2), or histone-lysine-N-methyltransferase (SSD3). b1), euchromatic histone-lysine N-methyltransferase 2 (G9a), histone-lysine N-methyltransferase (SUV39H1), enhancer of zeste homolog 2 (EZH2), viral lysine methyltransferase (vSET), histone methyltransferase (SET2), and protein-lysine N-methyltransferase (SMYD2)), a fusion of a sequence-targeting polypeptide and a connecting nucleation molecule.
[0062] In some embodiments, the targeting moiety comprises a sequence-targeting polypeptide, such as Cas9, a fusion of a sequence-targeting polypeptide, such as a fusion of dCas9 and a junction nucleating molecule, or a junction nucleating molecule. In some embodiments, the targeting moiety comprises a guide RNA or a nucleic acid encoding a guide RNA. In some embodiments, the targeting moiety introduces a targeted alteration into the anchor sequence-mediated junction to modulate transcription of the gene at the anchor sequence-mediated junction in human cells.
[0063] In some embodiments, the targeting moiety binds to the anchor sequence of the anchor sequence-mediated connection, and the targeting moiety introduces a targeted modification into the anchor sequence to modulate the transcription of the gene at the anchor sequence-mediated connection in human cells. In some embodiments, the targeted modification includes, for example, at least one of substitution, addition, or deletion of one or more nucleotides in the anchor sequence. In some embodiments, the targeted modification includes, for example, at least one of substitution, addition, or deletion of one or more nucleotides in the anchor sequence, such as those described herein, in a binding motif for a connection nucleation molecule. In some embodiments, the targeted modification includes at least one common nucleotide sequence in the opposite orientation, such as a binding motif for a connection nucleation molecule. In some embodiments, the targeted modification includes a non-naturally occurring anchor sequence that forms or destroys an anchor sequence-mediated connection.
[0064] The compositions described in various embodiments of the above aspects can be utilized in any other aspect described herein.
[0065] In one aspect, the disclosure includes a composition comprising a protein comprising a first polypeptide comprising a Cas or modified Cas protein domain and a second polypeptide comprising a polypeptide having DNA methyltransferase activity (or with demethylating or deaminase activity), in combination with at least one guide RNA (gRNA) or antisense DNA oligonucleotide that targets an anchor sequence of a target anchor sequence-mediated junction, wherein the system is effective to alter the target anchor sequence-mediated junction in a human cell.
[0066] In some embodiments, the composition is effective to alter target anchor sequence-mediated connectivity in human cells.
[0067] The compositions described in various embodiments of the above aspects can be utilized in any other aspect described herein.
[0068] In one aspect, the disclosure includes a pharmaceutical composition comprising a Cas protein and at least one guide RNA (gRNA) that targets the Cas protein to an anchor sequence of a target anchor sequence-mediated junction, wherein the Cas protein is effective to cause a mutation in the target anchor sequence that reduces formation of the anchor sequence-mediated junction associated with the target anchor sequence.
[0069] In one aspect, the present disclosure includes a synthetic nucleic acid comprising a plurality of anchor sequences, gene sequences, and transcription control sequences.
[0070] In some embodiments, the gene sequence and the transcription control sequence are located between multiple anchor sequences. In some embodiments, the nucleic acid comprises, in order: (a) an anchor sequence, a gene sequence, a transcription control sequence, and an anchor sequence, or (b) an anchor sequence, a transcription control sequence, a gene sequence, and an anchor sequence.
[0071] In some embodiments, the sequences are separated by a linker sequence. In some embodiments, the anchor sequence is 7 to 100 nt, 10 to 100 nt, 10 to 80 nt, 10 to 70 nt, 10 to 60 nt, 10 to 50 nt, or 20 to 80 nt. In some embodiments, the nucleic acid is 3,000-50,000 bp, 3,000-40,000 bp, 3,000-30,000 bp, 3,000-20,000 bp, 3,000-15,000 bp, 3,000-12,000 bp, 3,000-10,000 bp, 3,000-8,000 bp, 5,000-30,000 bp, 5,000-20,000 bp, 5,000-15,000 bp, 5,000-12,000 bp, 5,000-10,000 bp or any range therebetween.
[0072] In some embodiments, the vector comprises a nucleic acid described herein.
[0073] In some embodiments, the cells comprise a nucleic acid described herein.
[0074] In some embodiments, the pharmaceutical composition comprises a nucleic acid described herein.
[0075] In some embodiments, the methods of modulating gene expression by administering a composition comprise a nucleic acid described herein.
[0076] The nucleic acids described in various embodiments of the above aspects can be utilized in any other aspect described herein.
[0077] In one aspect, the disclosure includes a kit comprising: (a) a nucleic acid encoding a protein comprising a first polypeptide domain comprising a Cas or modified Cas protein and a second polypeptide domain comprising a polypeptide having DNA methyltransferase activity (or with demethylating or deaminase activity); and (b) at least one guide RNA (gRNA) for targeting the protein to an anchor sequence of a target anchor sequence-mediated junction in a target cell.
[0078] In some embodiments, (a) and (b) are provided in the same vector, e.g., a plasmid, an AAV vector, an AAV9 vector. In some embodiments, (a) and (b) are provided in separate vectors.
[0079] The kits described in various embodiments of the above aspects can be utilized in any other aspect described herein.
[0080] In one aspect, the disclosure includes a method for preparing a conjugated nucleating molecule with a selected binding affinity.
[0081] In one aspect, the disclosure includes a method of (altering gene expression / altering anchor sequence-mediated junction) in a mammalian subject, comprising administering to the subject (separately or in the same pharmaceutical composition) (i) a protein comprising a first polypeptide domain comprising Cas or a modified Cas protein and a second polypeptide domain comprising a polypeptide having DNA methyltransferase activity (or with demethylating or deaminase activity) or (ii) a protein comprising a first polypeptide domain comprising Cas or a modified Cas protein and a second polypeptide domain comprising a polypeptide that has a role in DNA methyltransferase activity (or with demethylating or deaminase activity), and a nucleic acid encoding at least one guide RNA (gRNA) that targets the anchor sequence of the anchor sequence-mediated junction.
[0082] In some embodiments, the anchor sequence is or includes a CTCF binding motif, such as SEQ ID NO: 1 or SEQ ID NO: 2. In some embodiments, the anchor sequence is or includes a CTCF binding motif associated with the target disease gene.
[0083] In some embodiments, the Cas protein is dCas9; the dCas9 is human codon-optimized. In some embodiments, the methyltransferase is a DNMT family methyltransferase. In some embodiments, the polypeptide is a TET family enzyme. In some embodiments, the protein has a linker between the first and second polypeptides.
[0084] In some embodiments, the gRNAs are selected from gRNAs for different diseases.
[0085] The methods described in the various embodiments of the above aspects can be utilized in any other aspect described herein.
[0086] In one aspect, the present disclosure includes a method for modifying chromatin structure, such as a two-dimensional structure, comprising altering the topology of anchor sequence-mediated connections to modulate transcription of a nucleic acid sequence. Altering the topology of anchor sequence-mediated connections, such as loops, modulates transcription of the nucleic acid sequence.
[0087] In another aspect, the present disclosure includes a method of modifying chromatin structure, such as a two-dimensional structure, comprising altering the topology of multiple anchor sequence-mediated connections to modulate transcription of a nucleic acid sequence. Altering the topology of multiple anchor sequence-mediated connections, such as multiple loops, modulates transcription of the nucleic acid sequence.
[0088] In another aspect, the disclosure includes a method of modulating transcription of a nucleic acid sequence comprising altering an anchor sequence-mediated connection, such as a loop, that affects transcription of the nucleic acid sequence. The anchor sequence-mediated connection modulates transcription of the nucleic acid sequence.
[0089] In certain embodiments, altering the anchor sequence-mediated connection alters the chromatin structure. For example, altering the chromatin structure by substituting, adding, or deleting one or more nucleotides within the anchor sequence of the anchor sequence-mediated connection alters the chromatin structure.
[0090] In various embodiments of the above aspects or any other aspects of the disclosure described herein, the topology is altered by substituting, adding, or deleting one or more nucleotides of the anchor sequences within the anchor sequence-mediated connections. For example, the substituted, added, or deleted one or more nucleotides may be within at least one anchor sequence, such as a binding motif for a connection nucleation molecule.
[0091] In some embodiments, the topology is altered by at least one of the following: modulating DNA methylation at one or more sites within the anchor sequence-mediated connection; altering the orientation of at least one common nucleotide sequence, such as a binding motif for a connection nucleation molecule; altering the spatial separation within the anchor sequence-mediated connection; altering the rotational free energy within the anchor sequence-mediated connection; and altering the positional degrees of freedom within the anchor sequence-mediated connection.
[0092] In some further embodiments, the topology is altered by any one or more of the following: disrupting an anchor sequence-mediated connection, forming a non-naturally occurring anchor sequence-mediated connection, forming multiple non-naturally occurring anchor sequence-mediated connections, and introducing an exogenous anchor sequence.
[0093] In certain embodiments, the topology is altered to result in, e.g., stable transcriptional modulation, such as modulation that persists for at least about 1 hour to about 30 days, or at least about 2 hours, 6 hours, 12 hours, 18 hours, 24 hours, 2 days, 3 days, 4 days, 5 days, 6 days, 7 days, 8 days, 9 days, 10 days, 11 days, 12 days, 13 days, 14 days, 15 days, 16 days, 17 days, 18 days, 19 days, 20 days, 21 days, 22 days, 23 days, 24 days, 25 days, 26 days, 27 days, 28 days, 29 days, 30 days, or longer, or any time in between.
[0094] In certain embodiments, the topology is altered to result in, e.g., transient transcriptional modulation, such as modulation that persists from about 30 minutes to about 7 days or less, or about 1 hour, 2 hours, 3 hours, 4 hours, 5 hours, 6 hours, 7 hours, 8 hours, 9 hours, 10 hours, 11 hours, 12 hours, 13 hours, 14 hours, 15 hours, 16 hours, 17 hours, 18 hours, 19 hours, 20 hours, 21 hours, 22 hours, 24 hours, 36 hours, 48 hours, 60 hours, 72 hours, 4 days, 5 days, 6 days, 7 days or less, or any time in between.
[0095] In some embodiments, the method further comprises modulating a connection nucleating molecule that interacts with the anchor sequence-mediated connection, such as its binding affinity to an anchor sequence within the anchor sequence-mediated connection.
[0096] In certain embodiments, the anchor sequence-mediated connection comprises at least a first anchor sequence and a second anchor sequence. In one embodiment, the anchor sequence-mediated connection is mediated by a first connection nucleation molecule bound to the first anchor sequence, a second connection nucleation molecule bound to the second anchor sequence, and an association between the first and second connection nucleation molecules. In another embodiment, the first or second connection nucleation molecule has a binding affinity for the anchor sequence that is higher or lower than a reference value, such as the binding affinity for the anchor sequence in the absence of the alteration.
[0097] In some embodiments, the second anchor sequence is non-contiguous with the first anchor sequence. In one embodiment, the anchor sequence-mediated connection is mediated by a first connection nucleation molecule bound to the first anchor sequence, a second connection nucleation molecule bound to the non-contiguous second anchor sequence, and an association between the first and second connection nucleation molecules. In another embodiment, the first or second connection nucleation molecule has a binding affinity for the anchor sequence that is higher or lower than a reference value, such as the binding affinity for the anchor sequence in the absence of the alteration.
[0098] In some embodiments in which the anchor sequences are non-contiguous with one another, the first anchor sequence is separated from the second anchor sequence by about 500 bp to about 500 Mb, about 750 bp to about 200 Mb, about 1 kb to about 100 Mb, about 25 kb to about 50 Mb, about 50 kb to about 1 Mb, about 100 kb to about 750 kb, about 150 kb to about 500 kb, or about 175 kb to about 500 kb. In some embodiments, the first anchor sequence is about 500bp, 600bp, 700bp, 800bp, 900bp, 1kb, 5kb, 10kb, 15kb, 20kb, 25kb, 30kb, 35kb, 40kb, 45kb, 50kb, 55kb, 60kb, 65kb, 70kb, 75kb, 80kb, 85kb, 90kb, 95kb, 100kb, 125kb, 150kb, 175kb, 200kb, 225kb, 240kb, 250kb, 260kb, 270kb, 280kb, 290kb, 300kb, 310kb, 320kb, 330kb, 340kb, 350kb, 360kb, 370kb, 380kb, 390kb, 400kb, 410kb, 420kb, 430kb, 440kb, 450kb, 460kb, 470kb, 480kb, 490kb, 500kb, 510kb, 520kb, 530kb, 540kb, 550kb, 560kb, 570kb, 580kb, 590kb, 600kb, 610kb, 620kb, 630kb, 640kb, 650kb, 660kb, 670kb, 68 The second anchor sequence may be separated from the second anchor sequence by 50 kb, 275 kb, 300 kb, 350 kb, 400 kb, 500 kb, 600 kb, 700 kb, 800 kb, 900 kb, 1 Mb, 2 Mb, 3 Mb, 4 Mb, 5 Mb, 6 Mb, 7 Mb, 8 Mb, 9 Mb, 10 Mb, 15 Mb, 20 Mb, 25 Mb, 50 Mb, 75 Mb, 100 Mb, 200 Mb, 300 Mb, 400 Mb, 500 Mb, or any size in between.
[0099] In certain embodiments, the first anchor sequence and the second anchor sequence each comprise a common nucleotide sequence, such as a binding motif for a connected nucleation molecule, such as those described herein. In some embodiments, the first anchor sequence and the second anchor sequence comprise different sequences, e.g., the first anchor sequence comprises a binding motif for a connected nucleation molecule and the second anchor sequence comprises a binding motif for another molecule, e.g., another connected nucleation molecule.
[0100] In some embodiments, the anchor sequence-mediated connection comprises multiple anchor sequences, hi one embodiment, at least one of the anchor sequences comprises a CTCF binding motif.
[0101] In some embodiments, the anchor sequence-mediated connection comprises a loop, such as an intrachromosomal loop. In one embodiment, the loop comprises a first anchor sequence, a nucleic acid sequence, a transcriptional regulatory sequence, such as an enhancing sequence or a silencing sequence, and a second anchor sequence. In another embodiment, the loop comprises, in order, a first anchor sequence, a transcriptional regulatory sequence, and a second anchor sequence; or a first anchor sequence, a nucleic acid sequence, and a second anchor sequence. In yet another embodiment, either or both of the nucleic acid sequence and the transcriptional regulatory sequence are located inside or outside the loop.
[0102] In certain embodiments, the anchor sequence-mediated connection has multiple loops. In one embodiment, the anchor sequence-mediated connection comprises multiple loops, and the anchor sequence-mediated connection comprises at least one of an anchor sequence, a nucleic acid sequence, and a transcriptional regulatory sequence in one or more of the loops.
[0103] In some embodiments, transcription of a nucleic acid sequence, such as transcription of a target nucleic acid sequence, is modulated compared to a reference value, e.g., transcription of the target sequence in the absence of an alteration in anchor sequence-mediated connectivity.
[0104] In some embodiments, transcription is activated by the inclusion of an activation loop. In one embodiment, the anchor sequence-mediated connection includes a transcriptional regulatory sequence, such as an enhancing sequence, that increases transcription of the nucleic acid sequence. In some further embodiments, transcription is activated by the elimination of a repression loop. In one embodiment, the anchor sequence-mediated connection excludes a transcriptional regulatory sequence, such as a silencing sequence, that decreases transcription of the nucleic acid sequence.
[0105] In some embodiments, transcription is suppressed by the inclusion of an inhibitory loop. In one embodiment, the anchor sequence-mediated connection includes a transcriptional regulatory sequence, such as a silencing sequence, that reduces transcription of the nucleic acid sequence. In some further embodiments, transcription is suppressed by the elimination of an activation loop. In one embodiment, the anchor sequence-mediated connection excludes a transcriptional regulatory sequence, such as an enhancing sequence, that increases transcription of the nucleic acid sequence.
[0106] In certain embodiments, the anchor sequence-mediated connection is altered in vivo in a subject, such as a human subject. In some embodiments, the methods described herein further include administering to the subject a targeting moiety selected from at least one of an exogenous connection nucleating molecule, a nucleic acid encoding the connection nucleating molecule, and a fusion of a sequence-targeting polypeptide and the connection nucleating molecule. In one embodiment, the connection nucleating molecule disrupts the binding of the endogenous connection nucleating molecule to its binding site, such as through competitive binding. In another embodiment, the targeting moiety comprises a sequence-targeting polypeptide, such as an enzyme, e.g., Cas9. In yet another embodiment, the targeting moiety further comprises a connection nucleating molecule. In yet another embodiment, the targeting moiety further comprises a guide RNA or a nucleic acid encoding the guide RNA.
[0107] In some embodiments, administering comprises administering a vector such as a viral vector, e.g., a lentiviral vector, that includes a nucleic acid encoding a targeting moiety, e.g., a junction nucleating molecule. In some further embodiments, administering comprises administering a formulation such as formulated in a polymeric carrier, e.g., a liposome.
[0108] In one aspect, the disclosure includes engineered cells containing targeted alterations in anchor sequence-mediated connections.
[0109] In another aspect, the disclosure includes engineered nucleic acid sequences comprising anchor sequence-mediated connections with targeted alterations.
[0110] In various embodiments of the above aspects or any other aspects of the disclosure described herein, the targeted alteration comprises any one or more of the following: a substitution, addition, or deletion of one or more nucleotides of an anchor sequence within the anchor sequence-mediated junction; a substitution, addition, or deletion of one or more nucleotides in at least one anchor sequence, e.g., a CTCF binding motif; an alteration of one or more DNA methylation sites within the anchor sequence-mediated junction; and at least one exogenous anchor sequence.
[0111] In some embodiments, the targeted alteration alters at least one connection nucleation molecule binding site, such as changing its binding affinity for the connection nucleation molecule. In some further embodiments, the targeted alteration alters the orientation of at least one consensus nucleotide sequence, e.g., a CTCF binding motif; disrupts an anchor sequence-mediated connection; and causes the formation of a non-naturally occurring anchor sequence-mediated connection.
[0112] In certain embodiments, the anchor sequence-mediated connection comprises at least a first anchor sequence and a second anchor sequence. In one embodiment, the anchor sequence-mediated connection is mediated by a first connection nucleation molecule bound to the first anchor sequence, a second connection nucleation molecule bound to the second anchor sequence, and an association between the first and second connection nucleation molecules. In another embodiment, the first or second connection nucleation molecule has a binding affinity for the anchor sequence that is higher or lower than a reference value, such as the binding affinity for the anchor sequence in the absence of the alteration.
[0113] In some embodiments, the second anchor sequence is non-contiguous with the first anchor sequence. In one embodiment, the anchor sequence-mediated connection is mediated by a first connection nucleation molecule bound to the first anchor sequence, a second connection nucleation molecule bound to the non-contiguous second anchor sequence, and an association between the first and second connection nucleation molecules. In another embodiment, the first or second connection nucleation molecule has a binding affinity for the anchor sequence that is higher or lower than a reference value, such as the binding affinity for the anchor sequence in the absence of the alteration.
[0114] In some embodiments in which the anchor sequences are non-contiguous with one another, the first anchor sequence is separated from the second anchor sequence by about 500 bp to about 500 Mb, about 750 bp to about 200 Mb, about 1 kb to about 100 Mb, about 25 kb to about 50 Mb, about 50 kb to about 1 Mb, about 100 kb to about 750 kb, about 150 kb to about 500 kb, or about 175 kb to about 500 kb. In some embodiments, the first anchor sequence is about 500bp, 600bp, 700bp, 800bp, 900bp, 1kb, 5kb, 10kb, 15kb, 20kb, 25kb, 30kb, 35kb, 40kb, 45kb, 50kb, 55kb, 60kb, 65kb, 70kb, 75kb, 80kb, 85kb, 90kb, 95kb, 100kb, 125kb, 150kb, 175kb, 200kb, 225kb, 240kb, 250kb, 260kb, 270kb, 280kb, 290kb, 300kb, 310kb, 320kb, 330kb, 340kb, 350kb, 360kb, 370kb, 380kb, 390kb, 400kb, 410kb, 420kb, 430kb, 440kb, 450kb, 460kb, 470kb, 480kb, 490kb, 500kb, 510kb, 520kb, 530kb, 540kb, 550kb, 560kb, 570kb, 580kb, 590kb, 600kb, 610kb, 620kb, 630kb, 640kb, 650kb, 660kb, 670kb, 68 The second anchor sequence may be separated from the second anchor sequence by 50 kb, 275 kb, 300 kb, 350 kb, 400 kb, 500 kb, 600 kb, 700 kb, 800 kb, 900 kb, 1 Mb, 2 Mb, 3 Mb, 4 Mb, 5 Mb, 6 Mb, 7 Mb, 8 Mb, 9 Mb, 10 Mb, 15 Mb, 20 Mb, 25 Mb, 50 Mb, 75 Mb, 100 Mb, 200 Mb, 300 Mb, 400 Mb, 500 Mb, or any size in between.
[0115] In certain embodiments, the first anchor sequence and the second anchor sequence each comprise a common nucleotide sequence, such as a CTCF-binding motif. In some embodiments, the first anchor sequence and the second anchor sequence comprise different sequences, for example, the first anchor sequence comprises a CTCF-binding motif, and the second anchor sequence comprises an anchor sequence other than the CTCF-binding motif.
[0116] In some embodiments, the anchor sequence-mediated connection comprises multiple anchor sequences, hi one embodiment, at least one of the anchor sequences comprises a CTCF binding motif.
[0117] In some further embodiments, the anchor sequence-mediated connection comprises a loop, such as an intrachromosomal loop. In one embodiment, the loop comprises a first anchor sequence, a nucleic acid sequence, a transcriptional regulatory sequence, such as an enhancing sequence or a silencing sequence, and a second anchor sequence. In another embodiment, the loop comprises, in order, a first anchor sequence, a transcriptional regulatory sequence, and a second anchor sequence; or a first anchor sequence, a nucleic acid sequence, and a second anchor sequence. In yet another embodiment, either or both of the nucleic acid sequence and the transcriptional regulatory sequence are located inside or outside the loop.
[0118] In certain embodiments, the anchor sequence-mediated connection has multiple loops. In one embodiment, the anchor sequence-mediated connection comprises multiple loops, and the anchor sequence-mediated connection comprises at least one of an anchor sequence, a nucleic acid sequence, and a transcriptional regulatory sequence in one or more of the loops.
[0119] In some embodiments, transcription of a nucleic acid sequence, such as transcription of a target nucleic acid sequence, is modulated compared to a reference value, e.g., transcription of the target sequence in the absence of an alteration in anchor sequence-mediated connectivity.
[0120] In some embodiments, transcription is activated by the inclusion of an activation loop. In one embodiment, the anchor sequence-mediated connection includes a transcriptional regulatory sequence, such as an enhancing sequence, that increases transcription of the nucleic acid sequence. In some further embodiments, transcription is activated by the elimination of a repression loop. In one embodiment, the anchor sequence-mediated connection excludes a transcriptional regulatory sequence, such as a silencing sequence, that decreases transcription of the nucleic acid sequence.
[0121] In some embodiments, transcription is suppressed by the inclusion of an inhibitory loop. In one embodiment, the anchor sequence-mediated connection includes a transcriptional regulatory sequence, such as a silencing sequence, that reduces transcription of the nucleic acid sequence. In some further embodiments, transcription is suppressed by the elimination of an activation loop. In one embodiment, the anchor sequence-mediated connection excludes a transcriptional regulatory sequence, such as an enhancing sequence, that increases transcription of the nucleic acid sequence.
[0122] In some embodiments, the disclosure includes an engineered cell described herein or a pharmaceutical composition comprising an engineered nucleic acid sequence described herein. In some further embodiments, the disclosure includes a plurality of cells comprising an engineered cell described herein. In some further embodiments, the disclosure includes a vector comprising an engineered nucleic acid sequence described herein.
[0123] In one aspect, the present disclosure includes a method of treating a disease or condition comprising administering to a subject a targeting moiety selected from at least one of an exogenous connected nucleating molecule, a nucleic acid encoding a connected nucleating molecule, and a fusion of a sequence-targeting polypeptide and a connected nucleating molecule.
[0124] In certain embodiments, the attached nucleation molecule disrupts the binding of an endogenous attached nucleation molecule to its binding site, such as through competitive binding.
[0125] In some embodiments, the targeting moiety comprises a sequence-targeting polypeptide such as an enzyme, e.g., Cas9. In some embodiments, the targeting moiety further comprises a connection nucleation molecule. In some further embodiments, the targeting moiety further comprises a guide RNA or a nucleic acid encoding the guide RNA. In some further embodiments, the targeting moiety targets one or more nucleotides of an anchor sequence within the anchor sequence-mediated connection for substitution, addition, or deletion, e.g., via CRISPR, TALEN, dCas9, recombination, transposon, etc. In some embodiments, the targeting moiety targets one or more DNA methylation sites within the anchor sequence-mediated connection. In some further embodiments, the targeting moiety introduces at least one of the following: at least one exogenous anchor sequence; an alteration in at least one connection nucleation molecule binding site, such as by changing the binding affinity for the connection nucleation molecule; a change in the orientation of at least one consensus nucleotide sequence, such as a CTCF binding motif; and a substitution, addition, or deletion in at least one anchor sequence, such as a CTCF binding motif.
[0126] In certain embodiments, administering comprises administering a vector, e.g., a viral vector, comprising a nucleic acid encoding a targeting moiety, e.g., a junction nucleating molecule. In some further embodiments, administering comprises administering a formulation, e.g., a liposome.
[0127] In some embodiments, the disease or condition is selected from the group consisting of cancer, trinucleotide repeats (such as Huntington's disease, Fragile X, all spinocerebellar ataxias, Friedreich's ataxia, myotonic dystrophy), autosomal dominant conditions, imprinted gene diseases (such as Prader-Willi syndrome, Angelman syndrome), haploinsufficient diseases, dominant negative mutations (severe congenital neutropenia), viral diseases (such as HIV, HBV, HCV, HPV), and environmentally induced transcriptional-epigenetic modifications (such as smoking, effects on gene expression from maternal diet).
[0128] In one aspect, the present disclosure provides an ABX n and C (wherein A is selected from a hydrophobic amino acid or an amide-containing backbone, e.g., aminoethyl-glycine, having a nucleic acid side chain; B and C may be the same or different and are each independently selected from arginine, asparagine, glutamine, lysine, and analogs thereof; each X is independently a hydrophobic amino acid or each X is independently an amide-containing backbone, e.g., aminoethyl-glycine, having a nucleic acid side chain; and n is an integer from 1 to 4), and the pharmaceutical composition includes at least one polypeptide, e.g., a membrane-translocating polypeptide, each comprising at least one sequence of the formula: ...
[0129] The compositions described in various embodiments of the above aspects can be used in any other aspect described herein. In some embodiments, the targeting moiety of one or more embodiments described herein comprises a membrane translocation polypeptide, for example, a polypeptide described herein.
[0130] In some embodiments, the hydrophobic amino acid is selected from alanine, valine, isoleucine, leucine, methionine, phenylalanine, tyrosine, trytophan, and analogs thereof. In some embodiments, B is selected from arginine or glutamine. In some embodiments, C is arginine. In some embodiments, n is 2.
[0131] In some embodiments, the polypeptides have a size ranging from about 5 to about 50 amino acid units in length.
[0132] In some embodiments, the composition comprises two or more polypeptides linked together. In some embodiments, the polypeptides are linked together, for example, an amino acid on one polypeptide is linked to one or more amino acids or the carboxy or amino terminus of another polypeptide, a branched polypeptide, or linked via a new peptide bond, a linear polypeptide. In some embodiments, the polypeptides are linked by a linker as described herein.
[0133] In some embodiments, the nucleic acid side chains are independently selected from the group consisting of purine side chains, pyrimidine side chains, and nucleic acid analog side chains. In some embodiments, the nucleic acid side chains hybridize to a heterologous moiety comprising a nucleic acid side chain, e.g., a PNA, or a nucleic acid.
[0134] In some embodiments, the composition comprises a membrane permeabilization polypeptide and at least one heterologous moiety. In one embodiment, the heterologous moiety is a connection nucleating molecule that interacts with anchor sequence-mediated connection. In another embodiment, the heterologous moiety is a sequence-targeting polypeptide, such as Cas9. In another embodiment, the heterologous moiety is a guide RNA or a nucleic acid encoding a guide RNA.
[0135] In some embodiments, the heterologous moiety is selected from the group consisting of a small molecule (e.g., a drug), a peptide (e.g., a ligand), and a nucleic acid (e.g., siRNA, DNA, modified RNA, RNA). In another embodiment, the heterologous moiety has at least one effector activity selected from the group consisting of modulating biological activity, binding to a regulatory protein, modulating enzymatic activity, modulating substrate binding, modulating receptor activation, modulating protein stability / degradation, and modulating transcript stability / degradation. In another embodiment, the heterologous moiety has at least one targeted function selected from the group consisting of modulating a function, modulating a molecule (e.g., an enzyme, protein, or nucleic acid), and localizing to a specific location. In another embodiment, the heterologous moiety is a tag or label, e.g., cleavable. In another embodiment, the heterologous moiety is selected from the group consisting of an epigenetic modifier, an epigenetic enzyme, a bicyclic peptide, a transcription factor, a DNA or protein-modifying enzyme, a DNA intercalator, an efflux pump inhibitor, a nuclear receptor activator or inhibitor, a proteasome inhibitor, a competitive inhibitor for an enzyme, a protein synthesis inhibitor, a nuclease, a protein fragment or domain, a tag or marker, an antigen, an antibody or antibody fragment, a ligand or receptor, a synthetic or analog peptide derived from a naturally occurring biologically active peptide, an antimicrobial peptide, a pore-forming peptide, a targeting or cytotoxic peptide, a degradative or self-destruction peptide, a CRISPR system or component thereof, DNA, RNA, an artificial nucleic acid, a nanoparticle, an oligonucleotide aptamer, a peptide aptamer, and a drug with poor pharmacokinetics or pharmacodynamics (PK / PD).
[0136] In some embodiments, the composition further comprises two or more heterologous moieties linked to the polypeptide, e.g., via a linker or directly, at the amino terminus, the carboxy terminus, all termini, a combination of some carboxy termini and some amino termini of the polypeptide, one or more amino acids of the polypeptide, or any combination thereof. In some embodiments, the heterologous moieties are linked to one of the polypeptides, e.g., via a linker or directly, at the amino terminus, the carboxy terminus, both termini, or one or more amino acids of the polypeptide.
[0137] In some embodiments, the composition further comprises a linker, e.g., between the polypeptides or between the polypeptide and the heterologous moiety. The linker may be a chemical bond, e.g., one or more covalent or non-covalent bonds. In some embodiments, the linker is a peptide linker (e.g., a non-ABX n C polypeptide). Such linkers may be 2 to 30 amino acids, or longer. Linkers include flexible, rigid, or cleavable linkers as described herein.
[0138] In some embodiments, the composition modulates DNA methylation at one or more sites within the anchor sequence-mediated junction.
[0139] In some embodiments, the composition modulates transcription transiently, e.g., provides modulation that lasts from about 30 minutes to about 7 days or less, or for about 1 hour, 2 hours, 3 hours, 4 hours, 5 hours, 6 hours, 7 hours, 8 hours, 9 hours, 10 hours, 11 hours, 12 hours, 13 hours, 14 hours, 15 hours, 16 hours, 17 hours, 18 hours, 19 hours, 20 hours, 21 hours, 22 hours, 24 hours, 36 hours, 48 hours, 60 hours, 72 hours, 4 days, 5 days, 6 days, 7 days or less, or any time in between.
[0140] In some embodiments, the composition stably modulates transcription, e.g., provides modulation that lasts for at least about 1 hour to about 30 days, or at least about 2 hours, 6 hours, 12 hours, 18 hours, 24 hours, 2 days, 3 days, 4 days, 5 days, 6 days, 7 days, 8 days, 9 days, 10 days, 11 days, 12 days, 13 days, 14 days, 15 days, 16 days, 17 days, 18 days, 19 days, 20 days, 21 days, 22 days, 23 days, 24 days, 25 days, 26 days, 27 days, 28 days, 29 days, 30 days, or more, or any time in between.
[0141] In some embodiments, the composition modulates the binding affinity of a connection nucleating molecule, eg, an anchor sequence within an anchor sequence-mediated connection, that interacts with the anchor sequence-mediated connection.
[0142] In some embodiments, the composition disrupts the binding of an endogenous connected nucleating molecule to its binding site, for example, by competitive binding.
[0143] In one aspect, the disclosure includes a method of modifying expression of a target gene, comprising altering an anchor sequence-mediated connection associated with the target gene, wherein the alteration modulates transcription of the target gene.
[0144] In one aspect, the present disclosure includes a method of modifying expression of a target gene comprising administering to a cell, tissue, or subject a composition described herein.
[0145] In one aspect, the disclosure includes a method of modulating transcription of a nucleic acid sequence comprising administering a composition described herein to alter an anchor sequence-mediated connection, e.g., a loop, that modulates transcription of the nucleic acid sequence, wherein the alteration of the anchor sequence-mediated connection modulates transcription of the nucleic acid sequence.
[0146] The methods described in the various embodiments of the above aspects can be utilized in any other aspect described herein.
[0147] In some embodiments, the composition modulates DNA methylation at one or more sites within the anchor sequence-mediated junction.
[0148] In some embodiments, the change in anchor sequence-mediated connectivity results in a transient modulation of transcription, e.g., modulation that lasts from about 30 minutes to about 7 days or less, or for about 1 hour, 2 hours, 3 hours, 4 hours, 5 hours, 6 hours, 7 hours, 8 hours, 9 hours, 10 hours, 11 hours, 12 hours, 13 hours, 14 hours, 15 hours, 16 hours, 17 hours, 18 hours, 19 hours, 20 hours, 21 hours, 22 hours, 24 hours, 36 hours, 48 hours, 60 hours, 72 hours, 4 days, 5 days, 6 days, 7 days or less, or any time in between.
[0149] In some embodiments, the change in anchor sequence-mediated connectivity results in stable modulation of transcription, e.g., modulation that persists for at least about 1 hour to about 30 days, or at least about 2 hours, 6 hours, 12 hours, 18 hours, 24 hours, 2 days, 3 days, 4 days, 5 days, 6 days, 7 days, 8 days, 9 days, 10 days, 11 days, 12 days, 13 days, 14 days, 15 days, 16 days, 17 days, 18 days, 19 days, 20 days, 21 days, 22 days, 23 days, 24 days, 25 days, 26 days, 27 days, 28 days, 29 days, 30 days, or longer, or any time in between.
[0150] In some embodiments, the composition modulates the binding affinity of a connection nucleating molecule, eg, an anchor sequence within an anchor sequence-mediated connection, that interacts with the anchor sequence-mediated connection.
[0151] In some embodiments, the composition disrupts the binding of an endogenous connected nucleating molecule to its binding site, for example, by competitive binding.
[0152] In some embodiments, the heterologous moiety is a sequence-targeting polypeptide, such as Cas9. In some embodiments, the heterologous moiety is a guide RNA or a nucleic acid encoding a guide RNA.
[0153] In one aspect, the disclosure includes a method of modulating gene expression comprising providing a composition described herein, e.g., a heterologous moiety that inhibits CpG binding, is an endogenous effector, is an exogenous effector, or is an agonist or antagonist thereof.
[0154] In one aspect, the disclosure includes a method of delivering a therapeutic agent comprising administering to a subject a composition described herein, wherein the heterologous moiety is a therapeutic agent, and the composition increases intracellular delivery of the therapeutic agent compared to the therapeutic agent alone.
[0155] The methods described in the various embodiments of the above aspects can be utilized in any other aspect described herein.
[0156] In some embodiments, the composition is targeted to specific cells or specific tissues.For example, the composition is targeted to epithelial, connective, muscle, or nervous tissue or cell, or a combination thereof.For example, the composition is targeted to specific organ systems, such as the cell or tissue of the cardiovascular system (heart, blood vessels); digestive system (esophagus, stomach, liver, gallbladder, pancreas, intestine, colon, rectum and anus); endocrine system (hypothalamus, pituitary gland, pineal gland or pineal gland, thyroid gland, parathyroid gland, adrenal gland); excretory system (kidney, ureter, bladder); lymphatic system (lymph, lymph nodes, lymphatic vessels, tonsils, adenoids, thymus, spleen); integumentary system (skin, hair, nails); muscular system (for example, skeletal muscle); nervous system (brain, spinal cord, nerve); reproductive system (ovary, uterus, mammary gland, testis, vas deferens, seminal vesicle, prostate); respiratory system (pharynx, larynx, trachea, bronchi, lung, diaphragm); skeletal system (bone, cartilage), and combinations thereof. In some embodiments, the composition crosses the blood-brain barrier, the placental membrane, or the blood-testis barrier.
[0157] In some embodiments, the composition is administered systemically, hi some embodiments, the administration is parenteral and the therapeutic agent is a parenteral therapeutic agent.
[0158] In some embodiments, the composition exhibits improved PK / PD, e.g., increased pharmacokinetics or pharmacodynamics, such as improved targeting, absorption, or transport compared to the therapeutic agent alone (e.g., at least 5%, 10%, 15%, 20%, 30%, 40%, 50%, 60%, 75%, 80%, 90% or more improvement). In some embodiments, the composition exhibits reduced undesirable effects compared to the therapeutic agent alone, such as reduced diffusion to non-target locations, off-target activity, or toxic metabolism (e.g., at least 5%, 10%, 15%, 20%, 30%, 40%, 50%, 60%, 75%, 80%, 90% or more reduction compared to the therapeutic agent alone). In some embodiments, the composition increases the efficacy and / or reduces the toxicity of the therapeutic agent compared to the therapeutic agent alone (e.g., by at least 5%, 10%, 15%, 20%, 30%, 40%, 50%, 60%, 75%, 80%, 90% or more).
[0159] In one aspect, the disclosure includes a method of intracellular delivery of a therapeutic agent comprising contacting a cell with a composition described herein, wherein the heterologous moiety is a therapeutic agent, and the composition increases intracellular delivery of the therapeutic agent compared to the therapeutic agent alone.
[0160] The methods described in the various embodiments of the above aspects can be utilized in any other aspect described herein.
[0161] In some embodiments, the composition has differential PK / PD compared to the therapeutic agent alone, e.g., the composition exhibits increased or decreased absorption or distribution, metabolism or excretion (e.g., at least 5%, 10%, 15%, 20%, 30%, 40%, 50%, 60%, 75%, 80%, 90% or more increase or decrease) compared to the therapeutic agent alone.
[0162] In some embodiments, the composition is administered at a dose sufficient to increase intracellular delivery of the therapeutic agent without significantly increasing endocytosis, e.g., less than about 50%, 40%, 20%, 15%, 10%, 5%, 4%, 3%, 2%, 1%, or any percentage therebetween. In some embodiments, the composition is administered at a dose sufficient to increase intracellular delivery of the therapeutic agent without significantly increasing calcium influx, e.g., less than about 50%, 40%, 20%, 15%, 10%, 5%, 4%, 3%, 2%, 1%, or any percentage therebetween. In some embodiments, the composition is administered at a dose sufficient to increase intracellular delivery of the therapeutic agent without significantly increasing endosomal activity, e.g., less than about 50%, 40%, 20%, 15%, 10%, 5%, 4%, 3%, 2%, 1%, or any percentage therebetween.
[0163] In one aspect, the disclosure includes a method of modulating transcription of a gene in a cell, comprising contacting the cell with a composition described herein, wherein the composition targets the gene and modulates its transcription.
[0164] The methods described in the various embodiments of the above aspects can be utilized in any other aspect described herein.
[0165] In some embodiments, the composition is administered in an amount and for a time sufficient to effect intracellular delivery of the therapeutic agent with reduced off-target transcriptional activity compared to the heterologous moiety alone, e.g., without significantly altering off-target transcriptional activity.
[0166] In one aspect, the disclosure includes a method of modulating membrane proteins, such as ion channels, cell surface receptors, and synaptic receptors, on a cell comprising contacting the cell with a composition described herein, wherein the composition targets the cell and modulates the membrane protein.
[0167] In one aspect, the disclosure includes a method of inducing cell death comprising contacting a cell with a composition described herein, wherein the composition targets the cell and induces apoptosis.
[0168] The methods described in the various embodiments of the above aspects can be utilized in any other aspect described herein.
[0169] In some embodiments, the composition targets cells that carry mutations in viral DNA sequences or genes. In one embodiment, the cells are infected with a virus. In another embodiment, the cells carry genetic mutations. In some embodiments, the composition targets cells that are in the early stages of necrosis, for example, by binding to necrotic cell markers.
[0170] In one aspect, the disclosure includes a method of increasing the bioavailability of a therapeutic agent comprising administering a composition described herein, wherein the therapeutic agent is a heterologous moiety.
[0171] The methods described in the various embodiments of the above aspects can be utilized in any other aspect described herein.
[0172] In some embodiments, the composition improves at least one PK / PD parameter, such as improved targeting, absorption, or transport, compared to the therapeutic agent alone (e.g., by at least 5%, 10%, 15%, 20%, 30%, 40%, 50%, 60%, 75%, 80%, 90% or more). In some embodiments, the composition reduces at least one undesirable parameter, such as reduced diffusion to non-target locations, off-target activity, or toxic metabolism, compared to the therapeutic agent alone (e.g., by at least 5%, 10%, 15%, 20%, 30%, 40%, 50%, 60%, 75%, 80%, 90% or more). In some embodiments, the composition increases the efficacy and / or reduces the toxicity of the therapeutic agent compared to the therapeutic agent alone.
[0173] In one aspect, the present disclosure includes a method of treating an acute or chronic infection comprising administering a composition described herein.
[0174] The methods described in the various embodiments of the above aspects can be utilized in any other aspect described herein.
[0175] In some embodiments, the composition targets infected cells carrying a pathogen. In some embodiments, the infection is caused by a pathogen selected from the group consisting of a virus, a bacterium, a parasite, and a prion. In some embodiments, the composition induces cell death in infected cells, for example, the heterologous moiety is an antibacterial, antiviral, or antiparasitic therapeutic agent.
[0176] In one aspect, the present disclosure includes a method of treating cancer comprising administering a composition described herein.
[0177] The methods described in the various embodiments of the above aspects can be utilized in any other aspect described herein.
[0178] In some embodiments, the heterologous moiety is a therapeutic agent that modulates gene expression of one or more genes.
[0179] In some embodiments, the composition targets cancer cells that carry a mutation in a gene, hi some embodiments, the composition induces cell death in the cancer cells, e.g., the heterologous moiety is a chemotherapeutic agent.
[0180] In one aspect, the present disclosure includes a method of treating a neurological disease or disorder comprising administering a composition described herein.
[0181] The methods described in the various embodiments of the above aspects can be utilized in any other aspect described herein.
[0182] In some embodiments, the composition modulates neuroreceptor activity or activation of a neurotransmitter, neuropeptide, or neuroreceptor.
[0183] In some embodiments, the neurological disease or disorder is Dravet syndrome.
[0184] In one aspect, the disclosure includes a method of treating a disease / disorder / condition in a subject comprising administering a composition described herein, wherein the composition modulates transcription to treat the disease / disorder / condition.
[0185] The methods described in the various embodiments of the above aspects can be utilized in any other aspect described herein.
[0186] In some embodiments, the disease / disorder / condition is a genetic disease.
[0187] In one aspect, the present disclosure includes a method of inducing immune tolerance comprising providing a composition described herein, eg, wherein the heterologous moiety is an antigen.
[0188] In one aspect, the disclosure includes a method of altering expression of a target gene in a genome comprising administering to the genome a pharmaceutical composition comprising a DNA sequence comprising (a) a targeting moiety and (b) an anchor sequence, wherein the anchor sequence promotes the formation of a connection that brings a gene expression factor (enhancing sequence, silencing / repression sequence) into operable linkage with the target gene.
[0189] In one aspect, the disclosure includes a system for pharmaceutical use comprising a protein comprising a first polypeptide domain comprising a Cas or modified Cas protein and a second polypeptide domain comprising a polypeptide having DNA methyltransferase activity (or with demethylating or deaminase activity), in combination with at least one guide RNA (gRNA) or antisense DNA oligonucleotide that targets the protein to an anchor sequence of a target anchor sequence-mediated junction, wherein the system is effective to alter target anchor sequence-mediated junction in a human cell.
[0190] In one aspect, the present disclosure includes a system for altering expression of a target gene in a human cell, comprising a targeting moiety (e.g., gRNA, LDB) associated with an anchor sequence associated with a target gene operably linked to the targeting moiety, and optionally a heterologous moiety (e.g., an enzyme, e.g., a nuclease or deactivated nuclease (e.g., Cas9, dCas9), a methylase, a demethylase, a deaminase), wherein the system is effective to modulate the connection mediated by the anchor sequence and alter expression of the target gene.
[0191] The systems described in various embodiments of the above aspects may be utilized in any other aspect described herein.
[0192] In some embodiments, the targeting moiety and the effector moiety are linked. In some embodiments, the system comprises a synthetic polypeptide comprising the targeting moiety and the heterologous moiety. In some embodiments, the system comprises a nucleic acid vector or vectors encoding at least one of the targeting moiety and the heterologous moiety.
[0193] The aspects described herein can be utilized in conjunction with any one or more of the embodiments described herein.
[0194] definition As used herein, the term "anchor sequence" refers to a sequence recognized by a nucleating agent (e.g., a nucleating protein) that binds sufficiently to form an anchor sequence-mediated connection, e.g., a loop. In some embodiments, the anchor sequence comprises one or more CTCF binding motifs. In some embodiments, the anchor sequence is not located within a gene coding region. In some embodiments, the anchor sequence is located within an intergenic region. In some embodiments, the anchor sequence is not located within either an enhancer or a promoter. In some embodiments, the anchor sequence is located at least 400 bp, at least 450 bp, at least 500 bp, at least 550 bp, at least 600 bp, at least 650 bp, at least 700 bp, at least 750 bp, at least 800 bp, at least 850 bp, at least 900 bp, at least 950 bp, or at least 1 kb away from any transcription start site. In some embodiments, the anchor sequence is located within a region without genomic imprinting, monoallelic expression, and / or monoallelic epigenetic marks. In some embodiments of the present disclosure, techniques are provided that can specifically target a particular anchor sequence or sequences without targeting other anchor sequences (e.g., sequences that may contain a connection nucleator (e.g., CTCF) binding motif in a different context); such targeted anchor sequences can be referred to as "target anchor sequences." In some embodiments, the sequence and / or activity of the target anchor sequence is modulated, but the sequence and / or activity of one or more other anchor sequences that may be present in the same system (e.g., in the same cell and / or in some embodiments, on the same nucleic acid molecule, e.g., on the same chromosome) as the targeted anchor sequence is not modulated.
[0195] As used herein, the phrase "anchor sequence-mediated connection" refers to a DNA structure, in some cases a loop, that is created and / or maintained by the physical interaction or binding of at least two anchor sequences in DNA by one or more proteins, such as nucleation proteins, or one or more proteins and / or nucleic acid entities (such as RNA or DNA) that bind to the anchor sequences in a manner that allows spatial proximity and functional linkage between the anchor sequences (see Figure 1).
[0196] As used herein, the term "associated with" refers to a gene being associated with an anchor sequence-mediated connection when the formation or disruption of the anchor sequence-mediated connection causes a change in the expression (e.g., transcription) of the target gene. For example, the formation or disruption of the anchor sequence-mediated connection causes an enhancing or silencing / repressing sequence to become associated or unassociated with the gene.
[0197] As used herein, the phrase "non-naturally occurring anchor sequence-mediated connection" refers to the formation of an anchor sequence-mediated connection that does not occur in nature. The generation of a non-naturally occurring anchor sequence-mediated connection may be by, but is not limited to, the alteration, addition, or deletion of one or more anchor sequences, and the alteration of one or more connection nucleating molecules.
[0198] As used herein, the term "consensus nucleotide sequence" refers to the connection nucleation molecule binding site in the anchor sequence. Examples of consensus nucleotide sequences include, but are not limited to, CTCF binding motif, USF1 binding motif, YY1 binding motif, TAF3 binding motif, and ZNF143 binding motif.
[0199] As used herein, the term "connected nucleator" refers to a protein that can associate directly or indirectly with an anchor sequence and interact with one or more connected nucleators (by interacting with the anchor sequence or other nucleic acids) to form a dimer (or higher-order structure) comprising two or more such connected nucleators, which may or may not be identical to one another. When connected nucleators associated with different anchor sequences associate with one another to maintain the different anchor sequences in close physical proximity to one another, the resulting structure is an anchor sequence-mediated connection. That is, the close physical proximity of a connected nucleator molecule interacting with another connected nucleator molecule-anchor sequence generates an anchor sequence-mediated connection (e.g., in some cases, a DNA loop) that begins and ends at the anchor sequence (see Figure 2). As those skilled in the art will readily appreciate upon reading this specification, terms such as "nucleating polypeptide," "nucleating molecule," "connected nucleator protein," and the like may also be used to refer to connected nucleators. Similarly, as will be readily understood by those of skill in the art upon reading this specification, an assembly of two or more linked nucleating agents (which in some embodiments may include multiple copies of the same agent, and / or which in some embodiments may include one or more of each of multiple different agents) can be referred to as a "complex," "dimer," "multimer," etc.
[0200] The term "loop" refers to the type of chromatin structure that can be created by the co-localization of two or more anchor sequences as anchor sequence-mediated connections. Thus, a loop is formed as a result of the interaction of at least two anchor sequences in DNA with one or more proteins, such as nucleation proteins, or one or more proteins and / or nucleic acid entities (such as RNA or DNA) that bind to the anchor sequences, allowing spatial proximity and functional connection between the anchor sequences. Those skilled in the art who read this specification will understand that the 2D representation of such a structure can be presented as a loop, for example, as shown in Figure 2. An "activation loop" is a structure that is open to active gene transcription, for example, a structure that includes a transcriptional regulatory sequence (enhancing sequence) that enhances transcription. An "inhibitory loop" is a structure that is closed with respect to active gene transcription, for example, a structure that includes a transcriptional regulatory sequence (silencing sequence) that suppresses transcription.
[0201] As used herein, the term "sequence-targeting polypeptide" refers to a protein, such as an enzyme, e.g., Cas9, that recognizes or specifically binds to a target sequence. In some embodiments, the sequence-targeting polypeptide is a catalytically inactive protein, such as dCas9, that lacks endonuclease activity.
[0202] As used herein, the term "subject" refers to an organism, e.g., a mammal (e.g., a human, a non-human mammal, a non-human primate, a primate, a laboratory animal, a mouse, a rat, a hamster, a gerbil, a cat, or a dog). In some embodiments, the human subject is an adult, adolescent, or pediatric subject. In some embodiments, the subject has had a disease or condition. In some embodiments, the subject is suffering from a disease, disorder, or condition, e.g., a disease, disorder, or condition that can be treated as provided herein. In some embodiments, the subject is predisposed to a disease, disorder, or condition; in some embodiments, a predisposed subject is predisposed to and / or exhibits a high risk (when compared to the average risk observed in a reference subject or population) of developing the disease, disorder, or condition. In some embodiments, the subject exhibits one or more symptoms of a disease, disorder, or condition. In some embodiments, the subject does not exhibit a particular symptom (e.g., clinical sign of disease) or characteristic of a disease, disorder, or condition. In some embodiments, the subject does not exhibit any symptom or characteristic of a disease, disorder, or condition. In some embodiments, the subject is a patient. In some embodiments, the subject is an individual to whom and / or to whom a diagnosis and / or therapy is administered.
[0203] As used herein, the term "targeting moiety" or "targeting element" refers to a molecule that specifically binds to the sequence within or surrounding the anchor sequence-mediated connection. Examples of targeting moieties include, but are not limited to, enzymes, sequence-targeting polypeptides such as Cas9, fusions of sequence-targeting polypeptides with connected nucleation molecules, such as fusions of dCas9 with connected nucleation molecules, or guide RNAs or nucleic acids, such as RNA, DNA, or modified RNA or DNA.
[0204] As used herein, the term "transcriptional control sequence" refers to a nucleic acid sequence that increases or decreases the transcription of a gene. An "enhancing sequence" increases the likelihood of gene transcription. A "silencing or repression sequence" decreases the likelihood of gene transcription. Enhancers and silencing sequences are approximately 50-3500 bp in length and can affect gene transcription from up to 1 Mb away. In certain embodiments, for example, the following are provided: (Item 1) A site-specific disrupting agent comprising a DNA-binding portion that specifically binds to one or more target anchor sequences in a cell with sufficient affinity to compete with the binding of endogenous nucleation polypeptides in the cell, and does not bind to non-target anchor sequences in the cell. (Item 2) 2. The site-specific disrupting agent of item 1, further comprising a negative effector moiety associated with the DNA-binding moiety, wherein when the DNA-binding moiety binds to the one or more target anchor sequences, the negative effector moiety is localized thereto, and wherein the negative effector moiety is characterized by a decrease in dimerization of the endogenous nucleation polypeptide when present compared to when the negative effector moiety is absent. (Item 3) 3. The site-specific disrupting agent of item 2, wherein the negative effector moiety is or comprises a variant of the dimerization domain of the endogenous nucleation polypeptide, or a dimerization portion thereof. (Item 4) The site-specific disrupting agent of any one of the preceding items, wherein the DNA-binding moiety is or comprises a polymer. (Item 5) 5. The site-specific disrupting agent according to item 4, wherein the polymer is or comprises a polyamide. (Item 6) 5. The site-specific disrupting agent according to item 4, wherein the polymer is an oligonucleotide. (Item 7) 7. The site-specific disrupting agent according to item 6, wherein the oligonucleotide has a sequence comprising the complement of the target anchor sequence. (Item 8) 8. The site-specific disrupting agent according to item 6 or 7, wherein the oligonucleotide comprises a chemical modification. (Item 9) 5. The site-specific disrupting agent according to item 4, wherein the polymer is a peptide nucleic acid. (Item 10) 5. The site-specific disrupting agent according to item 4, wherein the DNA binding portion is or comprises a peptide nucleic acid mixture. (Item 11) 5. The site-specific disrupting agent of any one of items 4, wherein the DNA-binding moiety is or comprises a peptide or polypeptide. (Item 12) 12. The site-specific disrupting agent according to item 11, wherein the polypeptide is a zinc finger polypeptide. (Item 13) 14. The site-specific disrupting agent of item 13, wherein the polypeptide is or comprises a transcription activator-like effector nuclease (TALEN) polypeptide. (Item 14) 4. The site-specific disrupting agent according to any one of items 1 to 3, wherein the DNA-binding moiety is or comprises a small molecule. (Item 15) 1. A method for modulating expression of a gene within an anchor sequence-mediated junction comprising a first anchor sequence and a second anchor sequence, the method comprising: A method comprising a step of contacting with the site-specific disrupting agent according to any one of items 1 to 14. (Item 16) 3. The method of claim 2, wherein the anchor sequence-mediated connection comprises at least one internal transcriptional control sequence. (Item 17) 17. The method of item 16, wherein the transcriptional control sequence is an enhancing sequence. (Item 18) 17. The method of item 16, wherein the transcriptional control sequence is a silencing or repression sequence. (Item 19) 19. The method of any one of items 15 to 18, wherein the gene is separated from the internal transcriptional regulatory sequence by at least 300 base pairs. (Item 20) 20. The method according to any one of items 15 to 19, wherein the first and / or second anchor sequence is located within 500 kb of an external transcriptional control sequence. (Item 21) 21. The method of item 20, wherein the external transcriptional control sequence is an enhancing sequence. (Item 22) 21. The method of item 20, wherein the external transcriptional control sequence is a silencing or repression sequence. (Item 23) 1. A method of modulating expression of a gene within 10 kb of a first anchor sequence in an anchor sequence-mediated junction comprising a first anchor sequence and a second anchor sequence, comprising: contacting the first and / or second anchor sequence with the site-specific disrupting agent according to any one of items 1 to 14; A method comprising: (Item 24) 24. The method of claim 23, wherein the anchor sequence-mediated connection comprises at least one internal transcriptional control sequence. (Item 25) 25. The method of item 24, wherein the transcriptional control sequence is an enhancing sequence. (Item 26) 25. The method of claim 24, wherein the transcriptional control sequence is a silencing or repression sequence. (Item 27) 27. The method of any one of items 24 to 26, wherein the gene is separated from the internal transcriptional regulatory sequence by at least 300 base pairs. (Item 28) 28. The method according to any one of items 23 to 27, wherein the first and / or second anchor sequence is located within 500 kb of an external transcriptional control sequence. (Item 29) 29. The method of item 28, wherein the external transcriptional control sequence is an enhancing sequence. (Item 30) 29. The method of item 28, wherein the external transcriptional control sequence is a silencing or repression sequence. (Item 31) 1. A method for reducing expression of a gene within an anchor sequence-mediated junction comprising a first anchor sequence, a second anchor sequence, and an internal enhancing sequence, comprising: contacting the first and / or second anchor sequence with the site-specific disrupting agent according to any one of items 1 to 14; A method comprising: (Item 32) 32. The method of claim 31, wherein the first and / or second anchor sequence is located within 500 kb of an external silencing or suppression sequence. (Item 33) 33. The method of claim 31 or 32, wherein the gene is separated from the internal enhancing sequence by at least 300 base pairs. (Item 34) 1. A method for increasing expression of a gene within an anchor sequence-mediated junction comprising a first anchor sequence and a second anchor sequence, wherein the first and / or second anchor sequence is located within 10 kb of an external enhancing sequence; contacting the first and / or second anchor sequence with the site-specific disrupting agent according to any one of items 1 to 14; A method comprising: (Item 35) 35. The method of claim 34, wherein the anchor sequence-mediated connection further comprises an internal enhancing sequence. (Item 36) 36. The method of claim 35, wherein the gene is separated from the internal enhancing sequence by at least 300 base pairs. (Item 37) (a) delivering the site-specific disrupting agent according to any one of items 1 to 14 to a mammalian cell; A method comprising: (Item 38) 38. The method of item 37, wherein the mammalian cell is a somatic cell. (Item 39) 39. The method of item 37 or 38, wherein the mammalian cells are primary cells. (Item 40) 40. The method of any one of items 37 to 39, wherein the delivering step is carried out ex vivo. (Item 41) 41. The method of claim 40, further comprising the step of removing the mammalian cells from the subject prior to the delivering step. (Item 42) 41. The method of claim 40, further comprising, after the delivering step, (b) administering the mammalian cells to a subject. (Item 43) 40. The method according to any one of items 37 to 39, wherein the delivering step comprises administering to a subject a composition comprising the site-specific destructive agent. (Item 44) 43. The method of item 41 or 42, wherein the subject has a disease or condition. (Item 45) 44. The method of any one of items 37 to 43, wherein the delivering step comprises delivery across a cell membrane. (Item 46) (i) a site-specific targeting moiety; and (ii) a deaminating agent; wherein the site-specific targeting moiety targets the fusion molecule to a target anchor sequence but not to at least one non-target anchor sequence. (Item 47) 47. The fusion molecule of item 46, wherein the target anchor sequence comprises a CTCF binding motif. (Item 48) 48. The fusion molecule of item 47, wherein the at least one non-target anchor sequence also comprises a CTCF binding motif. (Item 49) 49. The fusion molecule of any one of items 46 to 48, wherein the deaminating agent is a deaminase. (Item 50) 50. The fusion molecule of any one of items 46 to 49, wherein the site-specific targeting moiety comprises a Cas polypeptide and a site-specific guide RNA. (Item 51) 51. The fusion molecule of item 50, wherein the Cas polypeptide is enzymatically inactive. (Item 52) 52. The fusion molecule of item 50 or 51, wherein the Cas polypeptide is a Cas9 polypeptide. (Item 53) 49. The fusion molecule of any one of items 46 to 48, wherein the deaminating agent comprises an oligonucleotide. (Item 54) 54. The fusion molecule of item 53, wherein the oligonucleotide is conjugated to sodium bisulfite. (Item 55) 50. The fusion molecule of any one of items 46 to 49, wherein the site-specific targeting moiety is a polymer. (Item 56) 56. The fusion molecule of any one of items 46 to 55, wherein the DNA-binding moiety is or comprises a polymer. (Item 57) 57. The fusion molecule of item 56, wherein the polymer is or comprises a polyamide. (Item 58) 57. The fusion molecule of item 56, wherein the polymer is an oligonucleotide. (Item 59) 59. The fusion molecule of item 58, wherein the oligonucleotide has a sequence comprising the complement of the target anchor sequence. (Item 60) 60. The fusion molecule of item 58 or 59, wherein the oligonucleotide comprises a chemical modification. (Item 61) 57. The fusion molecule of item 56, wherein the polymer is a peptide nucleic acid. (Item 62) 47. The fusion molecule of item 46, wherein the DNA binding moiety is or comprises a peptide nucleic acid mixture. (Item 63) 57. The fusion molecule of item 56, wherein the DNA-binding moiety is or comprises a peptide or polypeptide. (Item 64) 64. The fusion molecule of item 63, wherein the polypeptide is a zinc finger polypeptide. (Item 65) 64. The fusion molecule of item 63, wherein the polypeptide is or comprises a transcription activator-like effector nuclease (TALEN) polypeptide. (Item 66) 49. The fusion molecule of any one of items 46 to 48, wherein the DNA-binding moiety is or comprises a small molecule. (Item 67) (i) a fusion polypeptide comprising an enzymatically inactive Cas polypeptide and a deaminating agent, or a nucleic acid encoding the fusion polypeptide; and (ii) a guide RNA that targets the fusion polypeptide to the target anchor sequence but not to at least one non-target anchor sequence; A composition comprising: (Item 68) 68. A method for modulating expression of a gene within an anchor sequence-mediated junction comprising a first anchor sequence and a second anchor sequence, the method comprising contacting the first and / or second anchor sequence with the fusion molecule of any one of items 46 to 66 or the composition of item 67. (Item 69) Item 69. The method of item 68, wherein the anchor sequence-mediated connection comprises at least one internal transcriptional control sequence. (Item 70) 70. The method of item 69, wherein the transcriptional control sequence is an enhancing sequence. (Item 71) 70. The method of item 69, wherein the transcriptional control sequence is a silencing or repression sequence. (Item 72) 72. The method of any one of items 69 to 71, wherein the gene is separated from the internal transcriptional control sequence by at least 300 base pairs. (Item 73) 73. The method of any one of items 69 to 72, wherein the first and / or second anchor sequence is located within 500 kb of an external transcription control sequence. (Item 74) 74. The method of item 73, wherein the external transcriptional control sequence is an enhancing sequence. (Item 75) 74. The method of item 73, wherein the external transcriptional control sequence is a silencing or repression sequence. (Item 76) 1. A method of modulating expression of a gene within 10 kb of a first anchor sequence in an anchor sequence-mediated junction comprising a first anchor sequence and a second anchor sequence, comprising: Contacting the first and / or second anchor sequence with the fusion molecule according to any one of items 46 to 66 or the composition according to item 67. A method comprising: (Item 77) 77. The method of claim 76, wherein the anchor sequence-mediated connection comprises at least one internal transcriptional control sequence. (Item 78) 78. The method of item 77, wherein the transcriptional control sequence is an enhancing sequence. (Item 79) 78. The method of item 77, wherein the transcriptional control sequence is a silencing or repression sequence. (Item 80) 80. The method of any one of claims 77 to 79, wherein the gene is separated from the internal transcriptional control sequence by at least 300 base pairs. (Item 81) 80. The method of any one of items 76 to 79, wherein the first and / or second anchor sequence is located within 500 kb of an external transcriptional control sequence. (Item 82) 82. The method of item 81, wherein the external transcriptional control sequence is an enhancing sequence. (Item 83) 82. The method of item 81, wherein the external transcriptional control sequence is a silencing or repression sequence. (Item 84) 1. A method for reducing expression of a gene within an anchor sequence-mediated junction comprising a first anchor sequence, a second anchor sequence, and an internal enhancing sequence, comprising: Contacting the first and / or second anchor sequence with the fusion molecule according to any one of items 46 to 66 or the composition according to item 67. A method comprising: (Item 85) 85. The method of item 84, wherein the first and / or second anchor sequence is located within 500 kb of an external silencing or suppression sequence. (Item 86) 86. The method of item 84 or 85, wherein the gene is separated from the internal enhancing sequence by at least 300 base pairs. (Item 87) 1. A method for increasing expression of a gene within an anchor sequence-mediated junction comprising a first anchor sequence and a second anchor sequence, wherein the first and / or second anchor sequence is located within 10 kb of an external enhancing sequence; Contacting the first and / or second anchor sequence with the fusion molecule according to any one of items 46 to 66 or the composition according to item 67. A method comprising: (Item 88) Item 88. The method of item 87, wherein the anchor sequence-mediated connection further comprises an internal enhancing sequence. (Item 89) 89. The method of item 88, wherein the gene is separated from the internal enhancing sequence by at least 300 base pairs. (Item 90) (a) delivering the fusion molecule according to any one of items 46 to 66 or the composition according to item 67 to a mammalian cell; A method comprising: (Item 91) 91. The method of item 90, wherein the mammalian cell is a somatic cell. (Item 92) 92. The method of item 90 or 91, wherein the mammalian cells are primary cells. (Item 93) 93. The method of any one of items 90 to 92, wherein the delivering step is carried out ex vivo. (Item 94) 94. The method of claim 93, further comprising the step of removing the mammalian cells from the subject prior to the delivering step. (Item 95) 95. The method of claim 93 or 94, further comprising (b) administering the mammalian cells to a subject. (Item 96) 93. The method of any one of items 90 to 92, wherein the delivering step comprises administering to the subject a composition comprising the fusion molecule of any one of items 46 to 66 or the composition of item 67. (Item 97) 94. The method of any one of items 90 to 93, wherein the delivering step comprises delivery across a cell membrane. (Item 98) (a) substituting, adding, or deleting one or more nucleotides of an anchor sequence in a mammalian somatic cell; A method comprising: (Item 99) 99. The method of item 98, wherein the mammalian somatic cells are primary cells. (Item 100) 99. The method of item 98, wherein the replacing, adding, or deleting step is carried out in vivo. (Item 101) 99. The method of claim 98, wherein the replacing, adding, or deleting step is carried out ex vivo. (Item 102) 102. The method according to any one of items 98 to 101, wherein the mammalian somatic cell is a non-embryonic cell. (Item 103) 103. The method of any one of Items 98 to 102, wherein the anchor sequence is a genomic anchor sequence. (Item 104) 1. A method comprising the step of delivering mammalian somatic cells to a subject having a disease or condition, wherein one or more nucleotides of an anchor sequence in said mammalian somatic cells have been substituted, added, or deleted. (Item 105) (a) administering mammalian somatic cells to a subject wherein the mammalian somatic cells are obtained from the subject, and the fusion molecule of any one of Items 46 to 66 or the composition of Item 67 is delivered to the mammalian cells ex vivo. (Item 106) 106. The method of any one of items 94 to 96 or 105, wherein the subject is a mammal. (Item 107) 107. The method of claim 106, wherein the subject has a disease or condition. (Item 108) (i) a site-specific targeting moiety, and (ii) epigenetic modifiers wherein the site-specific targeting moiety targets the fusion molecule to a target anchor sequence but not to at least one non-target anchor sequence. (Item 109) 109. The fusion molecule of claim 108, wherein the target anchor sequence comprises a CTCF binding motif. (Item 110) 110. The fusion molecule of claim 109, wherein the at least one non-target anchor sequence also comprises a CTCF binding motif. (Item 111) 111. The fusion molecule of any one of items 108 to 110, wherein the epigenetic modifying agent is selected from the group consisting of DNA methylases, DNA demethylases, histone methyltransferases, histone deacetylases, and combinations thereof. (Item 112) 112. The fusion molecule of any one of items 108 to 111, wherein the site-specific targeting moiety comprises a Cas polypeptide and a site-specific guide RNA. (Item 113) 113. The fusion molecule of item 112, wherein the Cas polypeptide is enzymatically inactive. (Item 114) 114. The fusion molecule of item 112 or 113, wherein the Cas polypeptide is a Cas9 polypeptide. (Item 115) 112. The fusion molecule of any one of items 108 to 111, wherein the site-specific targeting moiety is a polymer. (Item 116) 116. The fusion molecule of item 115, wherein the polymer is or comprises a polyamide. (Item 117) 116. The fusion molecule of item 115, wherein the polymer is an oligonucleotide. (Item 118) 117. The fusion molecule of claim 116, wherein the oligonucleotide has a sequence comprising the complement of the target anchor sequence. (Item 119) 117. The fusion molecule of item 116, wherein the oligonucleotide comprises a chemical modification. (Item 120) 116. The fusion molecule of item 115, wherein the polymer is a peptide nucleic acid. (Item 121) 109. The fusion molecule of item 108, wherein the site-specific targeting moiety is or comprises a peptide nucleic acid mixture. (Item 122) 116. The fusion molecule of item 115, wherein the site-specific targeting binding moiety is or comprises a peptide or polypeptide. (Item 123) 123. The fusion molecule of item 122, wherein the polypeptide is a zinc finger polypeptide. (Item 124) 123. The fusion molecule of item 122, wherein the polypeptide is or comprises a transcription activator-like effector nuclease (TALEN) polypeptide. (Item 125) 111. The fusion molecule of any one of items 108 to 110, wherein the site-specific binding moiety is or comprises a small molecule. (Item 126) A site-specific guide RNA comprising a targeting domain complementary to a target nucleic acid comprising an anchor sequence. (Item 127) 127. The site-specific guide RNA of claim 126, wherein the targeting domain is not complementary to at least one non-target nucleic acid comprising the anchor sequence. (Item 128) 128. The site-specific guide RNA of item 126 or 127, wherein the anchor sequence comprises a CTCF binding motif. (Item 129) (i) a fusion polypeptide comprising an enzymatically inactive Cas polypeptide and an epigenetic modifier, or a nucleic acid encoding the fusion polypeptide; and (ii) a guide RNA that targets the fusion polypeptide to the target anchor sequence but not to at least one non-target anchor sequence; A composition comprising: (Item 130) 1. A method of modulating expression of a gene within an anchor sequence-mediated junction comprising a first anchor sequence and a second anchor sequence, comprising: Contacting the first and / or second anchor sequence with the fusion molecule of any one of items 108 to 125, the site-specific guide RNA of any one of items 126 to 128, or the composition of item 129. A method comprising: (Item 131) Item 131. The method of item 130, wherein the anchor sequence-mediated connection comprises at least one internal transcriptional control sequence. (Item 132) Item 132. The method of item 131, wherein the transcriptional control sequence is an enhancing sequence. (Item 133) 132. The method of claim 131, wherein the transcriptional control sequence is a silencing or repression sequence. (Item 134) 134. The method of any one of items 130 to 133, wherein the gene is separated from the internal transcriptional control sequence by at least 300 base pairs. (Item 135) 135. The method of any one of items 130 to 134, wherein the first and / or second anchor sequence is located within 500 kb of an external transcriptional control sequence. (Item 136) Item 136. The method of item 135, wherein the external transcriptional control sequence is an enhancing sequence. (Item 137) 136. The method of claim 135, wherein the external transcriptional control sequence is a silencing or repression sequence. (Item 138) 1. A method of modulating expression of a gene within 10 kb of a first anchor sequence in an anchor sequence-mediated junction comprising a first anchor sequence and a second anchor sequence, comprising: Contacting the first and / or second anchor sequence with the fusion molecule of any one of items 108 to 125, the site-specific guide RNA of any one of items 126 to 128, or the composition of item 129. A method comprising: (Item 139) 139. The method of item 138, wherein the anchor sequence-mediated connection comprises at least one internal transcriptional control sequence. (Item 140) Item 139. The method of item 139, wherein the internal transcriptional control sequence is an enhancing sequence. (Item 141) 140. The method of claim 139, wherein the internal transcriptional control sequence is a silencing or repression sequence. (Item 142) 142. The method of any one of items 139 to 141, wherein the gene is separated from the internal transcriptional control sequence by at least 300 base pairs. (Item 143) 143. The method of any one of items 138 to 142, wherein the first and / or second anchor sequence is located within 500 kb of an external transcription control sequence. (Item 144) Item 144. The method of item 143, wherein the external transcriptional control sequence is an enhancing sequence. (Item 145) 144. The method of claim 143, wherein the external transcriptional control sequence is a silencing or repression sequence. (Item 146) 1. A method for reducing expression of a gene within an anchor sequence-mediated junction comprising a first anchor sequence, a second anchor sequence, and an internal enhancing sequence, comprising: Contacting the first and / or second anchor sequence with the fusion molecule of any one of items 108 to 125, the site-specific guide RNA of any one of items 126 to 128, or the composition of item 129. A method comprising: (Item 147) 147. The method of claim 146, wherein the first and / or second anchor sequence is located within 500 kb of an external silencing or suppression sequence. (Item 148) 148. The method of claim 146 or 147, wherein the gene is separated from the internal enhancing sequence by at least 300 base pairs. (Item 149) 1. A method for increasing expression of a gene within an anchor sequence-mediated junction comprising a first anchor sequence and a second anchor sequence, wherein the first and / or second anchor sequence is located within 10 kb of an external enhancing sequence, the method comprising: Contacting the first and / or second anchor sequence with the fusion molecule of any one of items 108 to 125, the site-specific guide RNA of any one of items 126 to 128, or the composition of item 129. A method comprising: (Item 150) Item 149. The method of item 149, wherein the anchor sequence-mediated connection further comprises an internal enhancing sequence. (Item 151) 151. The method of claim 150, wherein the gene is separated from the internal enhancing sequence by at least 300 base pairs. (Item 152) (a) delivering the fusion molecule according to any one of Items 108 to 125, the site-specific guide RNA according to any one of Items 126 to 128, or the composition according to Item 129 into a mammalian cell; A method comprising: (Item 153) 153. The method of item 152, wherein the mammalian cell is a somatic cell. (Item 154) 154. The method of item 152 or 153, wherein the mammalian cells are primary cells. (Item 155) 155. The method of any one of items 152 to 154, wherein the delivering step is carried out ex vivo. (Item 156) 156. The method of claim 155, further comprising the step of removing the mammalian cells from the subject prior to the delivering step. (Item 157) 157. The method of claim 155 or 156, further comprising, after the delivering step, (b) administering the mammalian cells to a subject. (Item 158) 156. The method of any one of Items 152 to 155, wherein the delivering step comprises administering to the subject a composition comprising the fusion molecule of any one of Items 108 to 125, the site-specific guide RNA of any one of Items 126 to 128, or the composition of Item 129. (Item 159) 159. The method of any one of items 156, 157, or 158, wherein the subject has a disease or condition. (Item 160) 159. The method of any one of items 152 to 159, wherein the delivering step comprises delivery across a cell membrane. (Item 161) An engineered site-specific nucleating agent, comprising: an engineered DNA-binding moiety that specifically binds to one or more target sequences within a cell with sufficient affinity that it competes with the binding of endogenous nucleation polypeptides within said cell, and does not bind to non-target sequences within said cell; and a nucleation polypeptide dimerization domain associated with said engineered DNA binding moiety; wherein when said engineered DNA-binding moiety binds at said at least one target sequence, a nucleation polypeptide dimerization domain is localized thereto, and each of said at least one target sequence is a target anchor sequence; the at least one or more target anchor sequences are positioned relative to an anchor sequence to which the nucleating polypeptide binds such that, when the nucleating polypeptide dimerization domain is localized to the target anchor sequence, interaction between the nucleating polypeptide dimerization domain and the nucleating polypeptide generates an anchor sequence-mediated connection. Engineered site-specific nucleating agents. (Item 162) 2. The engineered site-specific nucleation agent of item 1, wherein the target anchor sequence does not contain a CTCF binding motif. (Item 163) 1. A method of modulating expression of a gene within an anchor sequence-mediated junction comprising a first anchor sequence and a second anchor sequence, comprising: contacting the first and / or second anchor sequence with an engineered site-specific nucleating agent according to item 161 or 162. A method comprising: (Item 164) Item 164. The method of item 163, wherein the anchor sequence-mediated connection comprises at least one internal transcriptional control sequence. (Item 165) Item 164. The method of item 163, wherein the internal transcriptional control sequence is an enhancing sequence. (Item 166) 164. The method of claim 163, wherein the internal transcriptional control sequence is a silencing or repression sequence. (Item 167) 167. The method of any one of claims 163 to 166, wherein the gene is separated from the internal transcriptional control sequence by at least 300 base pairs. (Item 168) 168. The method of any one of items 163 to 167, wherein the first and / or second anchor sequence is located within 500 kb of an external transcriptional control sequence. (Item 169) Item 169. The method of item 168, wherein the external transcriptional control sequence is an enhancing sequence. (Item 170) 169. The method of claim 168, wherein the external transcriptional control sequence is a silencing or repression sequence. (Item 171) 1. A method of modulating expression of a gene within 10 kb of a first anchor sequence in an anchor sequence-mediated junction comprising a first anchor sequence and a second anchor sequence, comprising: contacting the first and / or second anchor sequence with an engineered site-specific nucleating agent according to item 161 or 162. A method comprising: (Item 172) Item 172. The method of item 171, wherein the anchor sequence-mediated connection comprises at least one internal transcriptional control sequence. (Item 173) Item 173. The method of item 172, wherein the internal transcriptional control sequence is an enhancing sequence. (Item 174) 173. The method of claim 172, wherein the internal transcriptional control sequence is a silencing or repression sequence. (Item 175) 175. The method of any one of items 171 to 174, wherein the gene is separated from the internal transcriptional control sequence by at least 300 base pairs. (Item 176) 176. The method of any one of items 171 to 175, wherein the first and / or second anchor sequence is located within 500 kb of an external transcription control sequence. (Item 177) Item 177. The method of item 176, wherein the external transcriptional control sequence is an enhancing sequence. (Item 178) 177. The method of claim 176, wherein the external transcriptional control sequence is a silencing or repression sequence. (Item 179) 1. A method for reducing expression of a gene within an anchor sequence-mediated junction comprising a first anchor sequence, a second anchor sequence, and an internal enhancing sequence, comprising: contacting the first and / or second anchor sequence with an engineered site-specific nucleating agent according to item 161 or 162. A method comprising: (Item 180) 180. The method of claim 179, wherein the first and / or second anchor sequence is located within 500 kb of an external silencing or suppression sequence. (Item 181) 181. The method of item 179 or 180, wherein the gene is separated from the internal enhancing sequence by at least 300 base pairs. (Item 182) 1. A method for increasing expression of a gene within an anchor sequence-mediated junction comprising a first anchor sequence and a second anchor sequence, wherein the first and / or the second anchor sequence is located within 10 kb of an external enhancing sequence, the method comprising: contacting the first and / or second anchor sequence with an engineered site-specific nucleating agent according to item 161 or 162. A method comprising: (Item 183) Item 183. The method of item 182, wherein the anchor sequence-mediated connection further comprises an internal enhancing sequence. (Item 184) 184. The method of claim 183, wherein the gene is separated from the internal enhancing sequence by at least 300 base pairs. (Item 185) (a) delivering the engineered site-specific nucleating agent according to item 161 or 162 to a mammalian cell; A method comprising: (Item 186) 186. The method of item 185, wherein the mammalian cell is a somatic cell. (Item 187) 187. The method of item 186, wherein the mammalian cells are primary cells. (Item 188) 188. The method of any one of items 185 to 187, wherein the delivering step is carried out ex vivo. (Item 189) 189. The method of claim 188, further comprising the step of removing the mammalian cells from the subject prior to the delivering step. (Item 190) 189. The method of claim 188 or 189, further comprising, after the delivering step, (b) administering the mammalian cells to a subject. (Item 191) 188. The method of any one of items 185 to 187, wherein the delivering step comprises administering to a subject a composition comprising the engineered site-specific nucleating agent of item 161 or 162 to mammalian cells. (Item 192) 192. The method of any one of items 189, 190 or 191, wherein the subject has a disease or condition. (Item 193) 193. The method of any one of items 185 to 192, wherein the delivering step comprises delivery across a cell membrane. (Item 194) 1. A method for modulating expression of a target gene in an expression unit, comprising: A method comprising modulating expression of the target gene by altering formation of an anchor sequence-mediated connection using a targeting moiety that targets the target gene or its associated transcriptional control sequence that affects transcription of the gene, e.g., a sequence outside of, or that is not part of, the anchor sequence. (Item 195) 1. A method of modulating transcription of a nucleic acid sequence, e.g., a target gene in an expression unit, comprising: A method comprising altering formation of the anchor sequence-mediated junction using a targeting moiety that targets the target gene or a sequence that is non-contiguous with its associated transcriptional control sequence that affects transcription of the target gene, so as to alter formation of the anchor sequence-mediated junction. (Item 196) A pharmaceutical preparation comprising a composition comprising a targeting moiety that binds to the anchor sequence of an anchor sequence-mediated connection and alters the formation of the anchor sequence-mediated connection, such that the composition modulates the transcription of a target gene associated with the anchor sequence-mediated connection, for example, in a human cell. (Item 197) A composition comprising a targeting moiety that binds to the anchor sequence of an anchor sequence-mediated connection and alters the formation of said anchor sequence-mediated connection (e.g., alters the affinity of said anchor sequence for a connection nucleating molecule by, e.g., at least 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95% or more). (Item 198) the targeting moiety, (i) a chemical substance, for example, a chemical substance that modulates cytosine (C) or adenine (A) (e.g., sodium bisulfite, ammonium bisulfite); (ii) having enzymatic activity (methyltransferase, demethylase, nuclease (e.g., Cas9), deaminase); (iii) The method or composition of any one of items 194 to 197, comprising an effector moiety that sterically hinders the formation of the anchor sequence-mediated connection, e.g., an ssDNA oligo, a locked nucleic acid (LNA), a peptide oligonucleotide conjugate (e.g., a membrane-translocating polypeptide with nucleic acid side chains), a bridged nucleic acid (BNA), a polyamide, and an antisense oligonucleotide conjugate comprising a DNA-binding molecule. (Item 199) 199. The method or composition of any one of items 194 to 198, wherein one or more transcriptional regulatory sequences are within the anchor sequence-mediated junction, e.g., a type 1 anchor sequence-mediated junction. (Item 200) 200. The method or composition of any one of items 194 to 199, wherein one or more transcription control sequences are external to said anchor sequence mediated junction, including, for example, a type 2 anchor sequence mediated junction. (Item 201) 201. The method or composition of any one of items 194 to 200, wherein the one or more transcription control sequences are internal to, e.g., an enhancing sequence, and at least partially external to, e.g., a silencing sequence, said anchor sequence-mediated connection, e.g., a type 3 anchor sequence-mediated connection. (Item 202) 202. The method or composition of any one of items 194 to 201, wherein the one or more transcription control sequences are, e.g., internal to the enhancing sequence, at least in part, external to, e.g., the enhancing sequence, the anchor sequence-mediated connection, e.g., a type 4 anchor sequence-mediated connection. (Item 203) A pharmaceutical composition comprising (a) a targeting moiety and (b) a DNA sequence, for example, comprising an anchor sequence. (Item 204) A composition comprising a protein comprising a domain that acts on DNA, e.g., an enzymatic domain (e.g., a nuclease domain, e.g., a Cas9 domain, e.g., a dCas9 domain; a DNA methyltransferase, a demethylase, a deaminase), in combination with at least one guide RNA (gRNA) or an antisense DNA oligonucleotide that targets the protein to an anchor sequence of a target anchor sequence-mediated junction, wherein the composition is effective to alter the target anchor sequence-mediated junction in a human cell. (Item 205) A composition comprising a targeting moiety that binds to an anchor sequence in an anchor sequence-mediated junction and alters the topology of said anchor sequence-mediated junction. (Item 206) A composition comprising a protein comprising a first polypeptide comprising a Cas or modified Cas protein domain and a second polypeptide comprising a polypeptide having DNA methyltransferase activity (or with demethylating or deaminase activity), in combination with at least one guide RNA (gRNA) or antisense DNA oligonucleotide that targets the protein to an anchor sequence of a target anchor sequence-mediated junction, wherein the system is effective to alter the target anchor sequence-mediated junction in human cells. (Item 207) A pharmaceutical composition comprising a Cas protein and at least one guide RNA (gRNA) that targets the Cas protein to an anchor sequence of a target anchor sequence-mediated junction, wherein the Cas protein is effective to cause a mutation in the target anchor sequence that reduces formation of the anchor sequence-mediated junction associated with the target anchor sequence. (Item 208) (a) a nucleic acid encoding a protein comprising a first polypeptide domain comprising a Cas or modified Cas protein and a second polypeptide domain comprising a polypeptide having DNA methyltransferase activity [or with demethylating or deaminase activity]; and (b) a kit comprising at least one guide RNA (gRNA) or antisense DNA oligonucleotide for targeting the protein to an anchor sequence of a target anchor sequence-mediated connection in a target cell. (Item 209) 1. A method of altering gene expression / altering anchor sequence-mediated connectivity in a mammalian subject, comprising administering to said subject (separately or in the same pharmaceutical composition): a) a nucleic acid encoding (i) a protein comprising a first polypeptide domain comprising a Cas or modified Cas protein and a second polypeptide domain comprising a polypeptide having DNA methyltransferase activity [or with demethylating or deaminase activity], or (ii) a protein comprising a first polypeptide domain comprising a Cas protein and a second polypeptide domain comprising a polypeptide having DNA methyltransferase activity [or with demethylating or deaminase activity], and b) at least one guide RNA (gRNA) or antisense DNA oligonucleotide targeting the anchor sequence of the anchor sequence-mediated junction The method of claim 1, wherein the (Item 210) altering the topology of the anchor sequence-mediated connection, e.g., looping, to modulate transcription of the nucleic acid sequence; A method of altering chromatin structure, e.g., two-dimensional structure, wherein altering the topology of the anchor sequence-mediated connections modulates transcription of the nucleic acid sequence. (Item 211) including altering the topology of multiple anchor sequence-mediated connections, e.g., multiple loops, to modulate transcription of a nucleic acid sequence; A method of altering chromatin structure, e.g., two-dimensional structure, wherein said change in topology modulates transcription of said nucleic acid sequence. (Item 212) affecting transcription of the nucleic acid sequence, including altering anchor sequence-mediated connections, e.g., loops; A method of modulating transcription of a nucleic acid sequence, wherein alteration of said anchor sequence-mediated connection modulates transcription of said nucleic acid sequence. (Item 213) Engineered cells containing targeted alterations of anchor sequence-mediated connectivity. (Item 214) An engineered nucleic acid sequence comprising an anchor sequence-mediated connection with a targeted alteration. (Item 215) A pharmaceutical composition comprising the engineered cell of any one of the preceding items or the engineered nucleic acid sequence of any one of the preceding items. (Item 216) A plurality of cells comprising the engineered cells of any one of the preceding items. (Item 217) A vector comprising the engineered nucleic acid sequence of any one of the preceding items. (Item 218) A composition for modulating transcription of a nucleic acid sequence by introducing targeted alterations into an anchor sequence-mediated junction, the composition comprising a targeting moiety that binds to said anchor sequence. (Item 219) A composition comprising a synthetic junction nucleating molecule having a selected binding affinity for an anchor sequence in said anchor sequence-mediated junction. (Item 220) A synthetic nucleic acid comprising a plurality of anchor sequences, gene sequences, and transcription modifier sequences. (Item 221) A vector comprising the nucleic acid described in any one of the preceding items. (Item 222) A cell comprising the nucleic acid described in any one of the preceding items. (Item 223) A pharmaceutical composition comprising the nucleic acid described in any one of the preceding items. (Item 224) A method for modulating gene expression by administering a composition comprising the nucleic acid of any one of the preceding items. (Item 225) Methods for preparing conjugated nucleating molecules with selected binding affinities. (Item 226) A method of treating a disease or condition comprising administering to a subject a targeting moiety selected from an exogenous connection nucleating molecule, a nucleic acid encoding said connection nucleating molecule, or a fusion of a sequence-targeting polypeptide with a connection nucleating molecule, wherein the targeting moiety alters anchor sequence-mediated connectivity. (Item 227) ABX n C (where A is selected from a hydrophobic amino acid or an amide-containing backbone, e.g., aminoethyl-glycine, having nucleic acid side chains; B and C may be the same or different and are each independently selected from arginine, asparagine, glutamine, lysine, and analogs thereof; each X is independently a hydrophobic amino acid or each X is independently an amide-containing backbone, e.g., aminoethyl-glycine, having nucleic acid side chains; and n is an integer from 1 to 4, The compositions and methods of any one of the preceding items, further comprising a polypeptide that hybridizes to a nucleic acid sequence within the anchor sequence-mediated junction (e.g., an anchor sequence of the anchor sequence-mediated junction, e.g., a CTCF binding motif, a BORIS binding motif, a cohesin binding motif, a USF1 binding motif, a YY1 binding motif, a TATA box, a ZNF143 binding motif, etc.). (Item 228) 1. A method for modifying expression of a target gene, comprising: Altering the anchor sequence-mediated connection associated with said target gene. wherein said alteration modulates transcription of said target gene. (Item 229) 1. A method for modifying expression of a target gene, comprising: Administering the composition of any one of the preceding items to a cell, tissue, or subject. A method comprising: (Item 230) 1. A method for modulating transcription of a nucleic acid sequence, comprising: A method comprising administering the composition of any one of the preceding items to alter an anchor sequence-mediated connection, e.g., a loop, that modulates transcription of a nucleic acid sequence, wherein the alteration of the anchor sequence-mediated connection modulates transcription of the nucleic acid sequence. (Item 231) A method for altering the expression of a target gene, comprising administering to a genome a pharmaceutical composition comprising (a) a targeting moiety and (b) a DNA sequence comprising an anchor sequence, wherein the anchor sequence promotes the formation of a connection that operably links a gene expression factor (enhancing sequence, silencing / suppression sequence) to the target gene. (Item 232) A method of modulating gene expression comprising providing the composition of any preceding paragraph, wherein, for example, the targeting moiety is an endogenous effector, an exogenous effector, or an agonist or antagonist thereof, including an effector moiety that inhibits CpG binding. (Item 233) 1. A method of delivering a therapeutic agent, comprising: A method comprising administering to a subject the composition of any preceding paragraph, wherein the targeting moiety comprises an effector moiety that is the therapeutic agent, wherein the composition increases intracellular delivery of the therapeutic agent compared to the therapeutic agent alone, and wherein the composition modulates transcription of a gene. (Item 234) 1. A method of modulating a membrane protein on a cell, comprising: contacting the cell with the composition of any one of the preceding items, wherein the composition targets the cell and modulates the membrane protein. (Item 235) A method of inducing cell death comprising contacting a cell with the composition of any one of the preceding items, wherein the composition targets the cell and induces apoptosis. (Item 236) 1. A method for increasing the bioavailability of a therapeutic agent, comprising: A method comprising administering the composition of any preceding item, wherein the therapeutic agent is a heterologous moiety. (Item 237) 1. A method of treating a disease / disorder / condition in a subject, comprising: A method comprising administering a composition according to any preceding item, wherein the composition modulates transcription to treat the disease / disorder / condition. (Item 238) A method of treating acute or chronic infections comprising administering a composition according to any one of the preceding items. (Item 239) A method of treating cancer comprising administering the composition of any one of the preceding items. (Item 240) A method for treating a neurological disease or disorder, comprising administering a composition described in any one of the preceding items. (Item 241) A method of inducing immune tolerance comprising providing a composition according to any preceding item, for example, wherein the heterologous moiety is an antigen. (Item 242) A system for pharmaceutical use comprising a protein comprising a first polypeptide domain comprising a Cas or modified Cas protein and a second polypeptide domain comprising a polypeptide having DNA methyltransferase activity (or with demethylating or deaminase activity), in combination with at least one guide RNA (gRNA) or antisense DNA oligonucleotide that targets the protein to an anchor sequence of a target anchor sequence-mediated junction, wherein the system is effective for altering the target anchor sequence-mediated junction in a human cell. (Item 243) 1. A system for altering expression of a target gene in a human cell, comprising: a targeting moiety (e.g., gRNA, LDB) associated with an anchor sequence associated with the target gene; Optionally, the system includes a heterologous moiety (e.g., an enzyme, e.g., a nuclease or deactivated nuclease (e.g., Cas9, dCas9), a methylase, a demethylase, a deaminase) operably linked to the targeting moiety, which is effective to modulate the connection mediated by the anchor sequence and alter expression of the target gene.
[0205] The following detailed description of embodiments of the present disclosure will be better understood when read in conjunction with the accompanying drawings. For the purpose of illustrating the present disclosure, there are shown in the drawings embodiments which are presently illustrated. It should be understood, however, that the present disclosure is not limited to the precise arrangements and instrumentalities of the embodiments shown in the drawings. [Brief explanation of the drawings]
[0206] [Figure 1]FIG. 1 is a diagram depicting the physical interaction or binding of one connecting nucleation molecule-anchor sequence with another connecting nucleation molecule-anchor sequence to produce an anchor sequence-mediated connection. [Figure 2] FIG. 2 is a diagram depicting a method for targeted disruption and creation of anchor sequence-mediated connections, e.g., loops. [Figure 3] FIG. 3 is a diagram depicting one embodiment of modulating gene expression through the creation of a non-naturally occurring anchor sequence-mediated junction (loop incorporation). [Figure 4] Figure 4 is a diagram depicting a method for modulating gene expression. The left side of the diagram is the same diagram as shown in Figure 1. The right side of the diagram is the disruption of the anchor sequence mediated connection (loop elimination). [Figure 5] FIG. 5 is a diagram depicting another embodiment in which the incorporation of new anchor sequences modulates gene expression through the creation of non-naturally occurring anchor sequence-mediated connections. [Figure 6] FIG. 6 is a diagram illustrating some of the types of anchor sequence mediated connections. [Figure 7-1] Figure 7 illustrates disruption of anchor sequence-mediated connections upstream of the MYC gene, leading to downregulation of MYC expression levels. Panels A, B, C, and D illustrate the reduction in MYC expression, and panel E depicts a map of the gRNA sequence, as further described in Examples 1 and 2. [Figure 7-2] Figure 7 illustrates disruption of anchor sequence-mediated connections upstream of the MYC gene, leading to downregulation of MYC expression levels. Panels A, B, C, and D illustrate the reduction in MYC expression, and panel E depicts a map of the gRNA sequence, as further described in Examples 1 and 2. [Figure 8-1] Figure 8 illustrates disruption of anchor sequence-mediated connectivity associated with the FOXJ3 gene, leading to downregulation of FOXJ3 expression levels. Panel A depicts a map of the gRNA and SNA sequences, while panels B, C, D, and E illustrate the reduction in FOXJ3 levels, as further described in Example 3. [Figure 8-2]Figure 8 illustrates disruption of anchor sequence-mediated connectivity associated with the FOXJ3 gene, leading to downregulation of FOXJ3 expression levels. Panel A depicts a map of the gRNA and SNA sequences, while panels B, C, D, and E illustrate the reduction in FOXJ3 levels, as further described in Example 3. [Figure 8-3] Figure 8 illustrates disruption of anchor sequence-mediated connectivity associated with the FOXJ3 gene, leading to downregulation of FOXJ3 expression levels. Panel A depicts a map of the gRNA and SNA sequences, while panels B, C, D, and E illustrate the reduction in FOXJ3 levels, as further described in Example 3. [Figure 9] 9 illustrates disruption of anchor sequence-mediated connections associated with the TUSC5 gene, leading to upregulation of TUSC5 expression levels. Panel A depicts upregulation of TUSC5 expression levels, and panel B depicts a map of the gRNA sequence, as further described in Example 4. [Figure 10] 10 illustrates disruption of the anchor sequence-mediated connection upstream of the DAND5 gene, leading to upregulation of DAND5 expression levels. As further described in Example 5, panel A shows upregulation of DAND5 expression levels, and panel B shows a map of the gRNA sequence. [Figure 11] 11 illustrates disruption of anchor sequence-mediated connections upstream or downstream of the SHMT2 gene, leading to downregulation of SHMT2 expression levels. Panels B and C depict maps of the gRNA sequences, and panels A and D depict downregulation of SHMT2 expression levels, as further described in Example 6. [Figure 12] 12 illustrates disruption of the anchor sequence-mediated connection upstream of the TTC21B gene, leading to upregulation of TTC21B expression levels. As further described in Example 7, panels A and B depict upregulation of TTC21B expression levels, and panel C depicts a map of the gRNA sequence. [Figure 13]13 illustrates disruption of anchor sequence-mediated connections downstream of the CDK6 gene, leading to downregulation of CDK6 expression levels. Panel A depicts downregulation of CDK6 expression levels, and panel B depicts a map of the gRNA sequence, as further described in Example 13. [Figure 14] FIG. 14 is a diagram of polypeptide beta, which hybridizes to the CTCF site of the miR290 loop and physically interferes with the looping function of CTCF (mediated by the polypeptide backbone and polynucleotide sequence). [Figure 15] FIG. 15 is a diagram of multimerizing polypeptide beta hybridizing to the promoter of the ELANE gene. [Figure 16] FIG. 16 is a diagram of polypeptide beta linked to a double-stranded unmethylated CTCF anchor sequence with specificity for the H19-IGF2 locus, so as to form a maternal-type loop mimicking the unmethylated CTCF binding motif on one of the paternal alleles. [Figure 17] FIG. 17 provides a summary of certain experimental data for targeted disruption of anchor sequence-mediated junctions. DETAILED DESCRIPTION OF THE INVENTION
[0207] Detailed Description The compositions described herein modulate gene expression in a subject, e.g., by altering anchor sequence-mediated connections in DNA, e.g., genomic DNA, thereby altering two-dimensional chromatin structure (e.g., anchor sequence-mediated connections that can be illustrated in two dimensions as having a higher-order structure than linear, as will be understood by one of skill in the art).
[0208] In one aspect, the present disclosure includes compositions comprising targeting moieties that bind to specific anchor sequence-mediated connections and alter the topology of the anchor sequence-mediated connections, e.g., anchor sequence-mediated connections having physical interactions of two or more DNA loci joined by connection nucleation molecules.
[0209] Formation of the anchor sequence-mediated connection forces a regulator of gene expression to interact with a target gene or spatially constrains the activity of the regulator. Alteration of the anchor sequence-mediated connection allows for gene therapy, e.g., modulation of gene expression, without altering the coding sequence of the modulated gene.
[0210] In some embodiments, the composition modulates transcription of genes associated with anchor sequence-mediated connections by physically interfering with one or more anchor sequences and the connected nucleation molecule. For example, small DNA-binding molecules (e.g., minor groove or major groove binders), peptides (e.g., zinc fingers, TALENs, novel or modified peptides), proteins (e.g., CTCF, modified CTCF with reduced CTCF binding and / or aggregation binding affinity), or nucleic acids (e.g., ssDNA, modified DNA or RNA, peptide-oligonucleotide conjugates, locked nucleic acids, bridged nucleic acids, polyamides, and / or triplex-forming oligonucleotides) can physically prevent the connected nucleation molecule from interacting with one or more anchor sequences to modulate gene expression.
[0211] In some embodiments, the composition modulates the transcription of the gene associated with the anchor sequence-mediated junction by modifying the anchor sequence, for example, by epigenetic modification. For example, one or more anchor sequences associated with the anchor sequence-mediated junction that comprises a target gene can be targeted by DNA methyltransferase, for example, dCas9-methyltransferase fusion, for example, antisense oligonucleotide-enzyme fusion, to be methylated, thereby modulating gene expression.
[0212] In some embodiments, the composition modulates the transcription of the gene associated with the anchor sequence-mediated junction by modifying the anchor sequence, for example, by modifying the genome. For example, one or more anchor sequences associated with the anchor sequence-mediated junction containing the target gene can be targeted by a deaminating enzyme (e.g., a deaminating oligonucleotide (e.g., an oligo-sodium bisulfite conjugate), a dCas-enzyme fusion, an antisense oligonucleotide-enzyme fusion, a deaminating antisense oligonucleotide-enzyme fusion) to modulate the expression of the gene.
[0213] In some embodiments, the composition modulates transcription, eg, activates or represses transcription, eg, induces epigenetic changes to chromatin, of a gene associated with the anchor sequence-mediated connection.
[0214] Anchor sequence-mediated connection In some embodiments, the anchor sequence-mediated junction comprises one or more anchor sequences, one or more genes, and one or more transcriptional regulatory sequences, such as enhancing or silencing sequences, hi some embodiments, the transcriptional regulatory sequences are internal, partially internal, or external to the anchor sequence-mediated junction.
[0215] In one embodiment, the anchor sequence-mediated connection comprises a loop, such as an intrachromosomal loop. In certain embodiments, the anchor sequence-mediated connection has multiple loops. One or more loops may comprise a first anchor sequence, a nucleic acid sequence, a transcriptional regulatory sequence, and a second anchor sequence. In another embodiment, at least one loop comprises, in order, a first anchor sequence, a transcriptional regulatory sequence, and a second anchor sequence; or a first anchor sequence, a nucleic acid sequence, and a second anchor sequence. In yet another embodiment, either or both of the nucleic acid sequence and the transcriptional regulatory sequence are located inside or outside the loop. In yet another embodiment, one or more of the loops comprises a transcriptional regulatory sequence.
[0216] In some embodiments, the anchor sequence-mediated connection comprises a TATA box, a CAAT box, a GC box, or a CAP site.
[0217] In some embodiments, the anchor sequence-mediated connection comprises multiple loops, in which case the anchor sequence-mediated connection comprises at least one of an anchor sequence, a nucleic acid sequence, and a transcriptional regulatory sequence in one or more of the loops.
[0218] In one aspect, the compositions described herein may include compositions for introducing targeted alterations into anchor sequence-mediated junctions to modulate transcription of nucleic acid sequences having targeting moieties that bind to anchor sequences. In some embodiments, the anchor sequence-mediated junction is altered by targeting one or more nucleotides within the anchor sequence-mediated junction for substitution, addition, or deletion.
[0219] In some embodiments, transcription is activated by the inclusion of an activation loop or the elimination of an inhibitory loop. In one such embodiment, the anchor sequence-mediated connection includes a transcriptional regulatory sequence that increases transcription of the nucleic acid sequence. In another such embodiment, the anchor sequence-mediated connection excludes a transcriptional regulatory sequence that decreases transcription of the nucleic acid sequence.
[0220] In some embodiments, transcription is inhibited by the inclusion of an inhibitory loop or the elimination of an activation loop. In one such embodiment, the anchor sequence-mediated connection includes a transcriptional control sequence that reduces transcription of the nucleic acid sequence. In another such embodiment, the anchor sequence-mediated connection excludes a transcriptional control sequence that increases transcription of the nucleic acid sequence.
[0221] Anchor Sequence Each anchor sequence-mediated connection comprises one or more anchor sequences, e.g., multiple. The anchor sequence can be manipulated or altered to disrupt naturally occurring loops or form new loops (e.g., to form exogenous loops or to form non-naturally occurring loops with exogenous or altered anchor sequences; see Figures 3, 4, and 5). Such alterations modulate gene expression by altering the two-dimensional structure of DNA, for example, thereby modulating the ability of target genes to interact with gene regulation and control factors (e.g., enhancing and silencing / repression sequences). In some embodiments, the chromatin structure is altered by substituting, adding, or deleting one or more nucleotides within the anchor sequence of the anchor sequence-mediated connection.
[0222] Anchor sequences may be non-contiguous with one another. In embodiments using non-contiguous anchor sequences, a first anchor sequence may be separated from a second anchor sequence by about 500 bp to about 500 Mb, about 750 bp to about 200 Mb, about 1 kb to about 100 Mb, about 25 kb to about 50 Mb, about 50 kb to about 1 Mb, about 100 kb to about 750 kb, about 150 kb to about 500 kb, or about 175 kb to about 500 kb. In some embodiments, the first anchor sequence is about 500bp, 600bp, 700bp, 800bp, 900bp, 1kb, 5kb, 10kb, 15kb, 20kb, 25kb, 30kb, 35kb, 40kb, 45kb, 50kb, 55kb, 60kb, 65kb, 70kb, 75kb, 80kb, 85kb, 90kb, 95kb, 100kb, 125kb, 150kb, 175kb, 200kb, 225kb, 240kb, 250kb, 260kb, 270kb, 280kb, 290kb, 300kb, 310kb, 320kb, 330kb, 340kb, 350kb, 360kb, 370kb, 380kb, 390kb, 400kb, 410kb, 420kb, 430kb, 440kb, 450kb, 460kb, 470kb, 480kb, 490kb, 500kb, 510kb, 520kb, 530kb, 540kb, 550kb, 560kb, 570kb, 580kb, 590kb, 600kb, 610kb, 620kb, 630kb, 640kb, 650kb, 660kb, 670kb, 68 The second anchor sequence may be separated from the second anchor sequence by 50 kb, 275 kb, 300 kb, 350 kb, 400 kb, 500 kb, 600 kb, 700 kb, 800 kb, 900 kb, 1 Mb, 2 Mb, 3 Mb, 4 Mb, 5 Mb, 6 Mb, 7 Mb, 8 Mb, 9 Mb, 10 Mb, 15 Mb, 20 Mb, 25 Mb, 50 Mb, 75 Mb, 100 Mb, 200 Mb, 300 Mb, 400 Mb, 500 Mb, or any size in between.
[0223] In one embodiment, the anchor sequence comprises a consensus nucleotide sequence, e.g., the CTCF motif: N(T / C / G)N(G / A / T)CC(A / T / G)(C / G)(C / T / A)AG(G / A)(G / T)GG(C / A / T)(G / A)(C / G)(C / T / A)(G / A / C) (SEQ ID NO: 1), where N is any nucleotide. The CTCF binding motif may also be in the opposite orientation, e.g., (G / A / C)(C / T / A)(C / G)(G / A)(C / A / T)GG(G / T)(G / A)GA(C / T / A)(C / G)(A / T / G)CC(G / A / T)N(T / C / G)N (SEQ ID NO: 2). In one embodiment, the anchor sequence comprises SEQ ID NO:1 or SEQ ID NO:2 or a sequence that is at least 75%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% identical to either SEQ ID NO:1 or SEQ ID NO:2.
[0224] In some embodiments, the anchor sequence-mediated connection comprises at least a first anchor sequence and a second anchor sequence. The first anchor sequence and the second anchor sequence may each comprise a common nucleotide sequence, for example, each comprising a CTCF binding motif. In some embodiments, the first anchor sequence and the second anchor sequence comprise different sequences, for example, the first anchor sequence comprises a CTCF binding motif, and the second anchor sequence comprises an anchor sequence other than the CTCF binding motif. In some embodiments, each anchor sequence comprises a common nucleotide sequence and one or more adjacent nucleotides on one or both sides of the common nucleotide sequence.
[0225] Two CTCF binding motifs (e.g., consecutive or non-consecutive CTCF binding motifs) that can form a connection can be present in the genome in any orientation, for example, in the same orientation (tandem) 5' to 3' (left tandem, e.g., two CTCF binding motifs comprising SEQ ID NO: 1) or 3' to 5' (right tandem, e.g., two CTCF binding motifs comprising SEQ ID NO: 2), or in a convergent orientation where one CTCF binding motif comprises SEQ ID NO: 1 and the other comprises SEQ ID NO: 2. CTCFBSDB 2.0: Database For CTCF Binding Motifs And Genome Organization (http: / / insulatordb.uthsc.edu / ) can be used to identify CTCF binding motifs associated with target genes.
[0226] In some embodiments, the anchor sequence comprises a CTCF binding motif associated with the target disease gene.
[0227] In some embodiments, chromatin structure is modified by substituting, adding, or deleting one or more nucleotides within at least one anchor sequence, e.g., a connection nucleation molecule binding site. One or more nucleotides can be specifically targeted for targeted modification, e.g., substitution, addition, or deletion within an anchor sequence, e.g., a connection nucleation molecule binding site.
[0228] In some embodiments, anchor sequence-mediated connections are altered by changing the orientation of at least one common nucleotide sequence, eg, a connecting nucleation molecule binding site.
[0229] In some embodiments, the anchor sequence comprises a connected nucleation molecule binding site, e.g., a CTCF binding motif, and the targeting moiety introduces an alteration of at least one connected nucleation molecule binding site, e.g., a change in binding affinity for the connected nucleation molecule.
[0230] In some embodiments, the anchor sequence-mediated connection is altered by introducing an exogenous anchor sequence. The addition of a non-naturally occurring or exogenous anchor sequence creates or disrupts a naturally occurring anchor sequence-mediated connection, for example, by inducing the formation of a non-naturally occurring loop that alters transcription of the nucleic acid sequence.
[0231] Anchor sequence-mediated connection types In some embodiments, the anchor sequence mediated connection comprises one or more genes, for example, 2, 3, 4, 5 or more.
[0232] In some embodiments, the present disclosure includes methods of modulating expression of a target gene in an anchor sequence-mediated connection, including modulating expression of the gene by targeting a sequence that is external to, not part of, or contained within the target gene or associated transcriptional control sequences that affect transcription of the gene, e.g., targeting the anchor sequence.
[0233] In some embodiments, the present disclosure includes methods of modulating transcription of a target gene, comprising targeting an associated transcriptional control sequence that is non-contiguous with the target gene or that influences transcription of the target gene, e.g., targeting an anchor sequence.
[0234] In some embodiments, the anchor sequence-mediated connection binds one or more, e.g., two, three, four, five, or more, transcription control sequences. In some embodiments, the target gene is non-contiguous with one or more of the transcription control sequences. In some embodiments in which the gene is non-contiguous with the transcription control sequences, the gene may be separated from one or more transcription control sequences by about 100 bp to about 500 Mb, about 500 bp to about 200 Mb, about 1 kb to about 100 Mb, about 25 kb to about 50 Mb, about 50 kb to about 1 Mb, about 100 kb to about 750 kb, about 150 kb to about 500 kb, or about 175 kb to about 500 kb. In some embodiments, the gene is about 100 bp, 300 bp, 500 bp, 600 bp, 700 bp, 800 bp, 900 bp, 1 kb, 5 kb, 10 kb, 15 kb, 20 kb, 25 kb, 30 kb, 35 kb, 40 kb, 45 kb, 50 kb, 55 kb, 60 kb, 65 kb, 70 kb, 75 kb, 80 kb, 85 kb, 90 kb, 95 kb, 100 kb, 125 kb, 150 kb, 175 kb, 200 kb, 22 The distance from the transcriptional regulatory sequence may be 5 kb, 250 kb, 275 kb, 300 kb, 350 kb, 400 kb, 500 kb, 600 kb, 700 kb, 800 kb, 900 kb, 1 Mb, 2 Mb, 3 Mb, 4 Mb, 5 Mb, 6 Mb, 7 Mb, 8 Mb, 9 Mb, 10 Mb, 15 Mb, 20 Mb, 25 Mb, 50 Mb, 75 Mb, 100 Mb, 200 Mb, 300 Mb, 400 Mb, 500 Mb, or any size in between.
[0235] In some embodiments, the type of anchor sequence mediated connection can be used to determine the method of modulating gene expression by changing anchor sequence mediated connection, for example, the selection of targeting moiety.For example, some types of anchor sequence mediated connection contain one or more transcriptional regulatory sequences within anchor sequence mediated connection.By disrupting the formation of anchor sequence mediated connection, for example, by changing one or more anchor sequences, disrupting this anchor sequence mediated connection can reduce the transcription of the target gene within anchor sequence mediated connection.
[0236] Type 1 In some embodiments, the expression of the target gene is regulated, modulated, or influenced by one or more transcriptional regulatory sequences associated with the anchor sequence-mediated connection. In some embodiments, the anchor sequence-mediated connection includes one or more associated genes and one or more transcriptional regulatory sequences. For example, the target gene and one or more transcriptional regulatory sequences are at least partially located within the anchor sequence-mediated connection, e.g., a type 1 anchor sequence-mediated connection. See Figure 6. The anchor sequence-mediated connection described in Figure 6 may also be referred to as "type 1, EP subtype."
[0237] In some embodiments, the target gene has a defined state of expression, e.g., in its natural state, e.g., in a disease state. For example, the target gene may have a high level of expression. Disrupting the anchor sequence-mediated connection can reduce the expression of the target gene, e.g., by changing the conformation of the DNA within the anchor sequence-mediated connection that was previously open to transcription, e.g., by creating additional distance between the target gene and the enhancing sequence. In one embodiment, both the associated gene and one or more transcriptional regulatory sequences, e.g., enhancing sequences, are present within the anchor sequence-mediated connection. Disruption of the anchor sequence-mediated connection reduces the expression of the gene. In one embodiment, the gene associated with the anchor sequence-mediated connection is, at least in part, accessible to one or more transcriptional regulatory sequences present within the anchor sequence-mediated connection. Disruption of the anchor sequence-mediated connection reduces the expression of the gene.
[0238] For example, the type 1 anchor sequence-mediated connection includes a gene encoding MYC, and disruption of the type 1 anchor sequence-mediated connection reduces gene expression and MYC protein levels. In another example, the type 1 anchor sequence-mediated connection includes a gene encoding Foxj3, and disruption of the type 1 anchor sequence-mediated connection reduces gene expression and Foxj3 protein levels.
[0239] Type 2 In some embodiments, the expression of a target gene is regulated, modulated, or influenced by one or more transcriptional regulatory sequences associated with an anchor sequence-mediated connection but inaccessible due to the anchor sequence-mediated connection. For example, the anchor sequence-mediated connection associated with a gene disrupts the ability of one or more transcriptional regulatory sequences to regulate, modulate, or affect the expression of the gene. The transcriptional regulatory sequence may be distant from the gene, for example, at least partially located on the opposite side of the gene from the anchor sequence-mediated connection, for example, internally or externally. For example, the gene is inaccessible to the transcriptional regulatory sequence due to the proximity of the anchor sequence-mediated connection. In some embodiments, one or more enhancing sequences are separated from the gene by an anchor sequence-mediated connection, for example, a type 2 anchor sequence-mediated connection. See Figure 6.
[0240] In some embodiments, the type 2 gene is contained within the anchor sequence-mediated connection, but the transcriptional regulatory sequence (e.g., enhancing sequence) is not contained within the anchor sequence-mediated connection. This subtype of type 2 can be referred to as "type 2, subtype 1."
[0241] In some embodiments, a Type 2 transcriptional regulatory sequence (e.g., an enhancing sequence) is contained within the anchor sequence-mediated connection, but a gene is not contained within the anchor sequence-mediated connection. This subtype of Type 2 can be referred to as "Type 2, subtype 2."
[0242] In some embodiments, the gene is inaccessible to one or more transcriptional regulatory sequences due to the anchor sequence-mediated connection, and disruption of the anchor sequence-mediated connection allows the transcriptional regulatory sequences to regulate, modulate, or affect expression of the gene. In one embodiment, the gene is both inside and outside the anchor sequence-mediated connection and is inaccessible to one or more transcriptional regulatory sequences. Disruption of the anchor sequence-mediated connection increases the accessibility of the transcriptional regulatory sequences to regulate, modulate, or affect expression of the gene, e.g., the transcriptional regulatory sequences increase expression of the gene. In one embodiment, the gene is inside the anchor sequence-mediated connection and is, at least in part, inaccessible to one or more transcriptional regulatory sequences present outside the anchor sequence-mediated connection. Disruption of the anchor sequence-mediated connection increases expression of the gene. In one embodiment, the gene is, at least in part, outside the anchor sequence-mediated connection and is inaccessible to one or more transcriptional regulatory sequences present inside the anchor sequence-mediated connection. Disruption of the anchor sequence-mediated connection increases expression of the gene.
[0243] In some embodiments, target gene has a specific state of expression, for example, in its natural state, for example, in a disease state.For example, target gene may have a moderate to low level of expression.By disrupting anchor sequence mediated connection, the expression of target gene can be modulated, for example, by the conformational change of the DNA that was previously closed for transcription within anchor sequence mediated connection, resulting in increased transcription, for example, by the conformational change of the DNA that makes the enhancing sequence more closely associated with the target gene.
[0244] For example, the type 2 anchor sequence-mediated connection includes a gene encoding SCN1a, and disruption of the type 2 anchor sequence-mediated connection increases gene expression and SCN1a protein levels. In another example, the type 2 anchor sequence-mediated connection includes a gene encoding Serpin1a, and disruption of the type 2 anchor sequence-mediated connection increases gene expression and Serpin1a protein levels. In another example, altering the anchor sequence-mediated connection associated with the IL-10 gene can elicit an IL-10-mediated tolerizing response, for example, increasing IL-10 expression to ameliorate an autoimmune condition. In another example, altering the associated anchor sequence-mediated connection can increase IL-6 expression by placing one or more enhancing sequences in close proximity to the IL-6 gene.
[0245] Type 3 In some embodiments, the expression of a target gene is regulated, modulated, or influenced by one or more transcriptional regulatory sequences associated with the anchor sequence-mediated connection, but which are not necessarily located on the same side of the anchor sequence-mediated connection. For example, an anchor sequence-mediated connection is associated with one or more genes, and one or more transcriptional regulatory sequences are present, at least in part, inside and outside the anchor sequence-mediated connection. In some embodiments, one or more enhancing sequences are present inside the anchor sequence-mediated connection, and one or more inhibitory signals, e.g., silencing sequences, are present outside the anchor sequence-mediated connection, e.g., type 3 anchor sequence-mediated connection. See Figure 6.
[0246] In some embodiments, the gene is inaccessible to one or more transcriptional regulatory sequences due to the anchor sequence-mediated connection, and disruption of the anchor sequence-mediated connection allows the transcriptional regulatory sequences to regulate, modulate, or affect expression of the gene. In one embodiment, the gene is located inside the anchor sequence-mediated connection and is inaccessible to one or more transcriptional regulatory sequences, for example, silencing / repression sequences, located outside the anchor sequence-mediated connection. Disruption of the anchor sequence-mediated connection reduces expression of the gene. In one embodiment, the gene is located both inside and outside the anchor sequence-mediated connection and is inaccessible to one or more transcriptional regulatory sequences, for example, silencing / repression sequences, located outside the anchor sequence-mediated connection. Disruption of the anchor sequence-mediated connection reduces expression of the gene. In one embodiment, the gene is located outside the anchor sequence-mediated connection and is inaccessible to one or more transcriptional regulatory sequences, for example, silencing / repression sequences, located inside the anchor sequence-mediated connection. Disruption of the anchor sequence-mediated connection reduces expression of the gene.
[0247] In some embodiments, target gene has a specific state of expression, for example, in its natural state, for example, in disease state.For example, target gene can have high level of expression in its natural state.By breaking anchor sequence mediated connection, it can modulate the expression of target gene, for example, by causing the conformational change of DNA that creates further distance between target gene and enhancing sequence, for example, by causing the conformational change of DNA that was previously open to transcription in anchor sequence mediated connection to reduce transcription, for example, by causing the conformational change of DNA that makes silencing sequence more closely associated with target gene, for example, by causing the conformational change of DNA that creates further distance between target gene and enhancing sequence to reduce transcription.
[0248] Type 4 In some embodiments, the expression of the target gene is regulated, modulated, or influenced by one or more transcriptional regulatory sequences associated with the anchor sequence-mediated connection, but not necessarily located inside the anchor sequence-mediated connection. For example, the anchor sequence-mediated connection is associated with one or more genes, and one or more transcriptional regulatory sequences are at least partially located inside and outside the anchor sequence-mediated connection, for example, type 4 anchor sequence-mediated connection. See Figure 6.
[0249] In some embodiments, the gene is inaccessible to one or more transcriptional regulatory sequences due to the anchor sequence-mediated connection, and disruption of the anchor sequence-mediated connection allows the transcriptional regulatory sequences to regulate, modulate, or affect expression of the gene. In one embodiment, the gene is inside the anchor sequence-mediated connection and inaccessible to one or more transcriptional regulatory sequences present outside the anchor sequence-mediated connection. Disruption of the anchor sequence-mediated connection increases expression of the gene. In one embodiment, the gene is inside and outside the anchor sequence-mediated connection and inaccessible to one or more transcriptional regulatory sequences, e.g., enhancing sequences, present outside the anchor sequence-mediated connection. Disruption of the anchor sequence-mediated connection increases expression of the gene. In one embodiment, the gene is outside the anchor sequence-mediated connection and inaccessible to one or more transcriptional regulatory sequences, e.g., enhancing sequences, present inside the anchor sequence-mediated connection. Disruption of the anchor sequence-mediated connection increases expression of the gene.
[0250] In some embodiments, the target gene has a specific state of expression, for example, in its natural state, for example, in a disease state.For example, the target gene may have a high level of expression in its natural state.By disrupting the anchor sequence mediated connection, the expression of the target gene can be modulated, for example, by the conformational change in the anchor sequence mediated connection, opening the DNA to transcription, thereby increasing transcription, for example, by associating one or more enhancing sequences with the target gene, thereby increasing transcription by the conformational change of DNA.
[0251] targeting part In some embodiments, the compositions, agents, fusion molecules, or other molecules described herein comprise one or more targeting moieties described herein. The targeting moiety can target anchor sequence-mediated connections by altering at least one of the following: at least one exogenous anchor sequence; altering at least one connection nucleation molecule binding site, such as by changing the binding affinity for the connection nucleation molecule; changing the orientation of at least one consensus nucleotide sequence, such as a CTCF binding motif; and substituting, adding, or deleting at least one anchor sequence, such as a CTCF binding motif.
[0252] Those skilled in the art will understand from reading the following examples of particular types of targeting moieties that in some embodiments, the targeting moiety is site-specific, i.e., in some embodiments, the targeting moiety specifically binds to one or more target anchor sequences (e.g., within a cell) and does not bind to non-target anchor sequences (e.g., within the same cell).
[0253] A targeting moiety can modulate a specific function, modulate a specific molecule (e.g., an enzyme, protein, or nucleic acid), and specifically bind to it for localization. A targeting function can act on a specific molecule, e.g., a molecular target. For example, a targeted therapeutic agent can interact with a specific molecule to increase, decrease, or otherwise modulate its function.
[0254] In some embodiments, the targeting moiety binds to an anchor sequence (e.g., a DNA sequence). In various parts of this disclosure, the term "DNA-binding moiety" can be used to refer to a targeting moiety.
[0255] In some embodiments, the compositions, agents, fusion molecules, or other molecules described herein include a targeting moiety (e.g., gRNA, antisense, oligonucleotide, peptide-oligonucleotide conjugate) operably linked to an effector moiety that binds to the anchor sequence and modulates the formation of the anchor sequence-mediated connection. The targeting moiety can bind to the anchor sequence of the anchor sequence-mediated connection and alter the formation of the anchor sequence-mediated connection (e.g., alter the affinity of the anchor sequence for the connection nucleating molecule by, for example, at least 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or more). The targeting moiety may be any one of the small molecules, peptides, nucleic acids, nanoparticles, aptamers, and drugs with poor pharmacokinetics described herein.
[0256] The targeting moiety can target a sequence, e.g., an anchor sequence, e.g., one or more nucleotides of a common nucleotide sequence within an anchor sequence, within the anchor sequence-mediated connection for substitution, addition, or deletion, such as by a gene editing system. In some embodiments, the targeting moiety binds to the anchor sequence-mediated connection, e.g., an anchor sequence in the anchor sequence-mediated connection, and changes the topology of the anchor sequence-mediated connection.
[0257] In some embodiments, the targeting moiety targets one or more nucleotides of an anchor sequence within the anchor sequence-mediated junction for substitution, addition, or deletion, e.g., by CRISPR, TALEN, dCas9, oligonucleotide pairing, recombination, transposon, etc. In some embodiments, the targeting moiety targets one or more DNA methylation sites within the anchor sequence-mediated junction.
[0258] The targeting moiety can alter a sequence, e.g., one or more nucleotides of the anchor sequence, e.g., a common nucleotide sequence within the anchor sequence, within the anchor sequence, by substitution, addition, or deletion, such as by a gene editing system.
[0259] In some embodiments, the targeting moiety introduces a targeted modification into the anchor sequence-mediated junction to modulate the transcription of the gene at the anchor sequence-mediated junction in human cells. The targeted modification may include one or more nucleotides, for example, the substitution, addition, or deletion of the anchor sequence within the anchor sequence-mediated junction. The targeting moiety binds to the anchor sequence of the anchor sequence-mediated junction, and the targeting moiety can introduce a targeted modification into the anchor sequence to modulate the transcription of the gene at the anchor sequence-mediated junction in human cells. In some embodiments, the targeted modification changes at least one of the binding sites for the junction nucleating molecule, for example, by changing the binding affinity for the anchor sequence, alternative splicing site, and binding site for untranslated RNA within the anchor sequence-mediated junction.
[0260] In some embodiments, the targeting moiety edits the anchor sequence-mediated connection with at least one of the following: at least one exogenous anchor sequence; altering at least one connection nucleation molecule binding site, such as by changing the binding affinity for the connection nucleation molecule; changing the orientation of at least one consensus nucleotide sequence, such as a CTCF binding motif; and substituting, adding, or deleting at least one anchor sequence, such as a CTCF binding motif.
[0261] In some embodiments, the targeting moiety is a nucleic acid sequence, a protein, a protein fusion, or a membrane translocation polypeptide. In some embodiments, the targeting moiety is selected from an exogenous linked nucleation molecule, a nucleic acid encoding a linked nucleation molecule, or a fusion of a sequence-targeting polypeptide and a linked nucleation molecule.
[0262] As described in more detail herein, in some embodiments, the targeting moieties described herein may be or include polymers or polymer moieties, such as polymers of nucleotides (such as oligonucleotides), peptide nucleic acids, peptide-nucleic acid hybrids, peptides or polypeptides, polyamides, carbohydrates, etc.
[0263] Nucleic acid sequence In some embodiments, the targeting moiety comprises a nucleic acid sequence. In some embodiments, the nucleic acid sequence encodes a gene or expression product.
[0264] As those skilled in the art will readily understand upon reading this specification, targeting moieties may comprise nucleic acid sequences that do not encode genes or expression products.For example, in some embodiments, targeting moieties comprise oligonucleotides that hybridize to target anchor sequences.For example, in some embodiments, the sequence of the oligonucleotide comprises the complement of the target anchor sequence, or has a sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, or at least 99% identical to the complement of the target anchor sequence.
[0265] Nucleic acid sequences may include, but are not limited to, DNA, RNA, modified oligonucleotides (e.g., chemically modified, such as modifications that alter backbone bonds, sugar moieties, and / or nucleobases), and artificial nucleic acids. In some embodiments, nucleic acid sequences include, but are not limited to, genomic DNA, cDNA, peptide nucleic acid (PNA) or peptide-oligonucleotide conjugates, locked nucleic acids (LNA), bridged nucleic acids (BNA), polyamides, triple-helix-forming oligonucleotides, modified DNA, antisense DNA oligonucleotides, tRNA, mRNA, rRNA, modified RNA, miRNA, gRNA, and siRNA or other RNA or DNA molecules.
[0266] In some embodiments, the nucleic acid sequence has a length of about 2 to about 5000 nt, about 10 to about 100 nt, about 50 to about 150 nt, about 100 to about 200 nt, about 150 to about 250 nt, about 200 to about 300 nt, about 250 to about 350 nt, about 300 to about 500 nt, about 10 to about 1000 nt, about 50 to about 1000 nt, about 100 to about 1000 nt, about 1000 to about 2000 nt, about 2000 to about 3000 nt, about 3000 to about 4000 nt, about 4000 to about 5000 nt, or any range therebetween.
[0267] In one aspect, the present disclosure includes a synthetic nucleic acid comprising multiple anchor sequences, gene sequences, and transcription control sequences. In some embodiments, the gene sequence and transcription control sequence are located between the multiple anchor sequences. In some embodiments, the synthetic nucleic acid comprises, in order, (a) an anchor sequence, a gene sequence, a transcription control sequence, and an anchor sequence, or (b) an anchor sequence, a transcription control sequence, a gene sequence, and an anchor sequence. In some embodiments, the sequences are separated by a linker sequence. In some embodiments, the anchor sequence is 7 to 100 nt, 10 to 100 nt, 10 to 80 nt, 10 to 70 nt, 10 to 60 nt, 10 to 50 nt, 20 to 80 nt, or any range therebetween. In some embodiments, the nucleic acid is 3,000-50,000 bp, 3,000-40,000 bp, 3,000-30,000 bp, 3,000-20,000 bp, 3,000-15,000 bp, 3,000-12,000 bp, 3,000-10,000 bp, 3,000-8,000 bp, 5,000-30,000 bp, 5,000-20,000 bp, 5,000-15,000 bp, 5,000-12,000 bp, 5,000-10,000 bp or any range therebetween.
[0268] In another aspect, the disclosure includes a vector comprising a nucleic acid described herein.
[0269] In another aspect, the disclosure includes a cell or tissue containing a nucleic acid described herein.
[0270] In another aspect, the disclosure includes pharmaceutical compositions comprising the nucleic acids described herein.
[0271] In another aspect, the disclosure includes methods of modulating gene expression by administering a composition comprising a nucleic acid described herein.
[0272] analog The nucleic acid sequence may comprise nucleosides, such as purines or pyrimidines, such as adenine, cytosine, guanine, thymine, and uracil. In some embodiments, the nucleic acid sequence comprises one or more nucleoside analogs. Nucleoside analogs include, but are not limited to, 5-fluorouracil, 5-bromouracil, 5-chlorouracil, 5-iodouracil, hypoxanthine, xanthine, 4-acetylcytosine, 4-methylbenzimidazole, 5-(carboxyhydroxymethyl)uracil, 5-carboxymethylaminomethyl-2-thiouridine, 5-carboxymethylaminomethyluracil, dihydrouracil, dihydrouridine, beta-D-galactosylqueosine, inosine, N6-isopentenyladenine, 1-methylguanine, 1-methylinosine, 2,2-dimethylguanine, 2-methyladenine, 2-methylguanine, 3-methylcytosine, 5-methylcytosine, N6-adenine, 7-methylguanine, 5-methylaminomethyluracil, 5-methoxyaminomethyl-2-thiouracil, beta-D-mannosylqueosine, 5'-methoxycarboxymethyl Thiouracil, 5-methoxyuracil, 2-methylthio-N6-isopentenyladenine, uracil-5-oxyacetic acid (v), wybutoxocin, pseudouracil, queosin, 2-thiocytosine, 5-methyl-2-thiouracil, 2-thiouracil, 4-thiouracil, 5-methyluracil, uracil-5-oxyacetic acid methyl ester, uracil-5-oxyacetic acid (v), 5-methyl-2-thiouracil, 3-(3-amino-3-N- Nucleoside analogs include nucleoside analogs such as (2-carboxypropyl)uracil, (acp3)w, 2,6-diaminopurine, 3-nitropyrrole, inosine, thiouridine, queosine, wyosine, diaminopurine, isoguanine, isocytosine, diaminopyrimidine, 2,4-difluorotoluene, isoquinoline, pyrrolo[2,3-β]pyridine, and any others capable of base pairing with a purine or pyrimidine side chain.
[0273] gRNA In some embodiments, the targeting moiety comprises a nucleic acid sequence, e.g., a guide RNA (gRNA). In some embodiments, the targeting moiety comprises a guide RNA or a nucleic acid encoding the guide RNA. gRNAs are short synthetic RNAs composed of a "scaffold" sequence required for Cas9 binding and a user-defined targeting sequence of approximately 20 nucleotides for the genomic target. In practice, guide RNA sequences are generally designed to have a length of 17 to 24 nucleotides (e.g., 19, 20, or 21 nucleotides) and are complementary to the targeted nucleic acid sequence. Custom gRNA generators and algorithms are commercially available for use in designing effective guide RNAs. Gene editing has also been achieved using chimeric "single guide RNAs" ("sgRNAs"), engineered (synthetic) single RNA molecules that mimic the naturally occurring crRNA-tracrRNA complex and contain both a tracrRNA (to bind to a nuclease) and at least one crRNA (to drive the nuclease to the sequence targeted for editing). Chemically modified sgRNAs have also been shown to be effective in genome editing; see, e.g., Hendel et al. (2015) Nature Biotechnol., 985-991.
[0274] In some embodiments, the nucleic acid sequence comprises a sequence complementary to an anchor sequence. In one embodiment, the anchor sequence comprises a CTCF binding motif or consensus sequence: N(T / C / G)N(G / A / T)CC(A / T / G)(C / G)(C / T / A)AG(G / A)(G / T)GG(C / A / T)(G / A)(C / G)(C / T / A)(G / A / C) (SEQ ID NO: 1), where N is any nucleotide. The CTCF binding motif or consensus sequence may also be in the opposite orientation, for example, (G / A / C)(C / T / A)(C / G)(G / A)(C / A / T)GG(G / T)(G / A)GA(C / T / A)(C / G)(A / T / G)CC(G / A / T)N(T / C / G)N (SEQ ID NO: 2). In some embodiments, the nucleic acid sequence comprises a sequence complementary to a CTCF binding motif or consensus sequence.
[0275] In some embodiments, the nucleic acid sequence comprises a sequence that is at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% complementary to an anchor sequence. In some embodiments, the nucleic acid sequence comprises a sequence that is at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% complementary to a CTCF binding motif or consensus sequence. In some embodiments, the nucleic acid sequence is selected from the group consisting of a gRNA and a sequence that comprises a sequence complementary to or at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% complementary to an anchor sequence.
[0276] In some embodiments, the epigenetic modifying agent is a steric entity near the gRNA, antisense DNA, or triple helix-forming oligonucleotide used as a DNA target and anchor sequence. The gRNA recognizes a specific DNA sequence (e.g., an anchor sequence flanked by sequences that confer sequence specificity, such as a CTCF anchor sequence). The gRNA may contain an additional sequence that acts as a steric blocker, interfering with the connecting nucleation molecule sequence. In some embodiments, the gRNA is combined with one or more peptides, such as S-adenosylmethionine (SAM), that act as steric entities that interfere with the connecting nucleation molecule.
[0277] Protein-encoding nucleic acids In some embodiments, a vector, e.g., a viral vector, comprises a nucleic acid encoding a targeting moiety, e.g., a junction nucleating molecule.
[0278] The nucleic acids described herein or the proteins described herein, such as nucleic acids encoding junction nucleation molecules or epigenetic modifiers, can be incorporated into vectors. Vectors derived from retroviruses, such as lentiviruses, are suitable for achieving long-term gene transfer because they allow long-term, stable integration of transgenes and their propagation in daughter cells. Examples of vectors include expression vectors, replication vectors, probe generation vectors, and sequencing vectors. Expression vectors can be provided to cells in the form of viral vectors. Viral vector technology is well known in the art and is described in various virology and molecular biology manuals. Viruses useful as vectors include, but are not limited to, retroviruses, adenoviruses, adeno-associated viruses, herpes viruses, and lentiviruses. Generally, suitable vectors contain a replication origin functional in at least one organism, a promoter sequence, convenient restriction endonuclease sites, and one or more selectable markers.
[0279] Expression of natural or synthetic nucleic acids is typically achieved by operably linking a nucleic acid encoding a gene of interest to a promoter and incorporating the construct into an expression vector. The vector may be suitable for replication and integration in eukaryotes. Typical cloning vectors contain transcription and translation terminators, initiation sequences, and promoters useful for expressing the desired nucleic acid sequence.
[0280] Additional promoter elements, such as enhancing sequences, regulate the frequency of transcription initiation. Typically, these are located in the region 30–110 bp upstream of the start site, although some promoters have recently been shown to contain functional elements downstream of the start site as well. The spacing between promoter elements is often flexible; thus, promoter function is preserved even if elements are inverted or moved relative to each other. In the thymidine kinase (tk) promoter, the spacing between promoter elements can be increased to 50 bp apart before activity begins to decline. Depending on the promoter, individual elements appear to function cooperatively or independently to activate transcription.
[0281] One example of a suitable promoter is the immediate early cytomegalovirus (CMV) promoter sequence. This promoter sequence is a strong constitutive promoter sequence capable of driving high-level expression of any polynucleotide sequence operably linked to it. Another example of a suitable promoter is the elongation growth factor-1α (EF-1α) promoter. However, other constitutive promoter sequences can also be used, including, but not limited to, the simian virus 40 (SV40) early promoter, mouse mammary tumor virus (MMTV), human immunodeficiency virus (HIV) long terminal repeat (LTR) promoter, MoMuLV promoter, avian leukosis virus promoter, Epstein-Barr virus immediate early promoter, Rous sarcoma virus promoter, and human gene promoters, including, but not limited to, the actin promoter, myosin promoter, hemoglobin promoter, and creatine kinase promoter.
[0282] Furthermore, the present disclosure should not be limited to the use of constitutive promoters. Inducible promoters are also contemplated as part of the present disclosure. The use of an inducible promoter provides a molecular switch that can turn on the expression of the polynucleotide sequence to which it is operably linked when expression of the polynucleotide sequence is desired, or turn off expression when expression is not desired. Examples of inducible promoters include, but are not limited to, metallothionine promoters, glucocorticoid promoters, progesterone promoters, and tetracycline promoters.
[0283] The introduced expression vector may also contain a selectable marker gene or a reporter gene, or both, to facilitate the identification and selection of expressing cells from a population of cells that are desired to be transfected or infected via the viral vector. In other embodiments, the selectable marker can be carried on a separate piece of DNA and used in a co-transfection procedure. Both the selectable marker and the reporter gene can be flanked by appropriate transcriptional regulatory sequences to enable their expression in the host cell. Useful selectable markers include, for example, antibiotic resistance genes such as neo.
[0284] Reporter genes can be used to identify potentially transfected cells and evaluate the function of transcriptional control sequences. Generally, reporter genes are genes encoding polypeptides that are not present in or expressed by the recipient source and whose expression is indicated by some easily detectable property, such as enzymatic activity. Expression of the reporter gene is assayed at a suitable time after DNA is introduced into the recipient cells. Suitable reporter genes may include genes encoding luciferase, beta-galactosidase, chloramphenicol acetyltransferase, secreted alkaline phosphatase, or green fluorescent protein (e.g., Ui-Tei et al., 2000, FEBS Letters 479:79-82). Suitable expression systems are well known and can be prepared using known techniques or obtained commercially. Generally, the construct with the smallest 5'-flanking region that exhibits the highest level of reporter gene expression is identified as the promoter. Such promoter regions can be linked to reporter genes and used to evaluate agents for their ability to modulate promoter-driven transcription.
[0285] RNAi Certain RNA agents can inhibit gene expression through the biological process of RNA interference (RNAi). RNAi molecules include RNA or RNA-like structures, typically containing 15-50 base pairs (e.g., about 18-25 base pairs) and having nucleobase sequences identical (complementary) or nearly identical (substantially complementary) to coding sequences in expressed target genes in cells. RNAi molecules include, but are not limited to, small interfering RNA (siRNA), double-stranded RNA (dsRNA), microRNA (miRNA), short hairpin RNA (shRNA), meroduplexes, and Dicer substrates (U.S. Patent Nos. 8,084,599, 8,349,809, and 8,513,207). In one embodiment, the present disclosure includes compositions for inhibiting expression of genes encoding polypeptides described herein, such as junction nucleation molecules or epigenetic modifiers.
[0286] RNAi molecules contain sequences that are substantially complementary or completely complementary to all or a fragment of a target gene. RNAi molecules can complement sequences at intron-exon boundaries to prevent newly generated nuclear RNA transcripts of a specific gene from maturing into mRNA for transcription. RNAi molecules complementary to a specific gene can hybridize with the mRNA for that gene and prevent its translation. Antisense molecules can be DNA, RNA, or derivatives or hybrids thereof. Examples of such derivative molecules include, but are not limited to, peptide nucleic acids (PNAs) and phosphorothioate-based molecules such as deoxyribonucleic guanidine (DNG) or ribonucleic guanidine (RNG).
[0287] RNAi molecules can be provided to cells as "ready-to-use" RNA synthesized in vitro, or as antisense genes transfected into cells that result in the RNAi molecule upon transcription. Hybridization with mRNA results in degradation of the hybridized molecule by RNAse H and / or inhibition of the formation of the translation complex. Both fail, resulting in the production of the original gene product.
[0288] The length of the RNAi molecule that hybridizes to the transcript of interest should be about 10 nucleotides, about 15-30 nucleotides, or about 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, or more nucleotides. The degree of identity between the antisense sequence and the targeted transcript should be at least 75%, at least 80%, at least 85%, at least 90%, or at least 95%.
[0289] RNAi molecules may also contain overhangs, i.e., unpaired overhanging nucleotides that are typically not directly involved in the double helix structure normally formed by the paired core sequences of the sense and antisense strands as defined herein. RNAi molecules may also contain 3' and / or 5' overhangs of approximately 1 to 5 bases that are independent of each other on the sense and antisense strands. In one embodiment, both the sense and antisense strands contain 3' and 5' overhangs. In one embodiment, one or more of the 3' overhanging nucleotides of one strand are paired with one or more 5' overhanging nucleotides of the other strand. In another embodiment, one or more of the 3' overhanging nucleotides of one strand are not base-paired with one or more 5' overhanging nucleotides of the other strand. The sense and antisense strands of an RNAi molecule may or may not contain the same number of nucleotide bases. The antisense and sense strands may form a duplex in which only the 5' end is blunt, only the 3' end is blunt, both the 5' and 3' ends are blunt, or neither the 5' nor the 3' end is blunt. In another embodiment, one or more of the nucleotides in the overhang contain thiophosphate, phosphorothioate, deoxynucleotide inverted (3' to 3' linked) nucleotides, or are modified ribonucleotides or deoxynucleotides.
[0290] Small interfering RNA (siRNA) molecules contain a nucleotide sequence identical to about 15 to about 25 contiguous nucleotides of a target mRNA. In some embodiments, the siRNA sequence begins with the dinucleotide AA, contains about 30 to 70% (about 30 to 60%, about 40 to 60%, or about 45 to 55%) GC content, and does not share a high percentage of identity with any nucleotide sequence other than the target in the genome of the mammal into which it is introduced, as determined, for example, by a standard BLAST search.
[0291] siRNAs and shRNAs resemble intermediates in the processing pathway of endogenous microRNA (miRNA) genes (Bartel, Cell 116:281-297, 2004). In some embodiments, siRNAs can function as miRNAs and vice versa (Zeng et al., Mol Cell 9:1327-1333, 2002; Doench et al., Genes Dev 17:438-442, 2003). Like siRNAs, microRNAs use RISC to downregulate target genes; however, unlike siRNAs, miRNAs in many animals do not cleave mRNAs. Instead, miRNAs reduce protein output by translational repression or poly(A) removal and mRNA degradation (Wu et al., Proc Natl Acad Sci USA 103:4034-4039, 2006). Known miRNA binding sites are located within the mRNA 3'UTR; miRNAs are thought to target sites with near-perfect complementarity to nucleotides 2-8 from the 5' end of the miRNA (Rajewsky, Nat Genet 38:Suppl:S8-13, 2006; Lim et al., Nature 433:769-773, 2005). This region is known as the seed region. Because siRNAs and miRNAs are compatible, exogenous siRNAs down-regulate mRNAs with seed complementarity to the siRNA (Birmingham et al., Nat Methods 3:199-204, 2006). Multiple target sites within the 3'UTR confer stronger down-regulation (Doench et al., Genes Dev 17:438-442, 2003).
[0292] Lists of known miRNA sequences can be found in databases maintained by research institutions such as Wellcome Trust Sanger Institute, Penn Center for Bioinformatics, Memorial Sloan Kettering Cancer Center and European Molecule Biology Laboratory, among others.Known effective siRNA sequences and cognate binding sites are also fully described in relevant literature.RNAi molecules can be easily designed and produced by techniques known in the art.In addition, there are computer tools that increase the chance of finding effective and specific sequence motifs (Pei et al., 2006; Reynolds et al., 2004; Khvorova et al., 2003; Schwarz et al., 2003; Ui-Tei et al., 2004; Heale et al., 2005; Chalk et al., 2004; Amarzguioui et al., 2004).
[0293] RNAi molecules modulate the expression of the RNA encoded by genes.Since multiple genes can share a certain degree of sequence homology with each other, in some embodiments, RNAi molecules can be designed to target a class of genes that have sufficient sequence homology.In some embodiments, RNAi molecules can contain sequences that are shared between different gene targets or have complementary sequences to sequences that are unique to specific gene targets.In some embodiments, RNAi molecules can be designed to target the conserved regions of RNA sequences that have homology between several genes, thereby targeting several genes in a gene family (for example, different gene isoforms, splice variants, mutant genes, etc.).In some embodiments, RNAi molecules can be designed to target sequences that are unique to the specific RNA sequence of a single gene.
[0294] In some embodiments, the RNAi molecule interacts with a junction nucleation molecule, e.g., CTCF, cohesin, USF1, YY1, TATA box-binding protein-associated factor 3 (TAF3), ZNF143, or another polypeptide that promotes the formation of anchor sequence-mediated junctions, or an epigenetic modifier, such as, but not limited to, DNA methylases (e.g., DNMT3a, DNMT3b, DNMTL), DNA demethylation (e.g., the TET family of enzymes catalyzes the oxidation of 5-methylcytosine to 5-hydroxymethylcytosine and more highly oxidized derivatives), histone methyltransferases, histone deacetylases (e.g., HDAC1, , HDAC2, HDAC3), sirtuin 1, 2, 3, 4, 5, 6, or 7, lysine-specific histone demethylase 1 (LSD1), histone-lysine-N-methyltransferase (Setdb1), euchromatin histone-lysine N-methyltransferase 2 (G9a), histone-lysine N-methyltransferase (SUV39H1), enhancer of zeste homolog 2 (EZH2), viral lysine methyltransferase (vSET), histone methyltransferase (SET2), protein-lysine N-methyltransferase (SMYD2), and others. In one embodiment, the RNAi molecule targets a deacetylase protein, e.g., sirtuin 1, 2, 3, 4, 5, 6, or 7. In one embodiment, the present disclosure includes a composition comprising an RNAi targeting a junction nucleation molecule, e.g., CTCF.
[0295] Peptide or protein moieties In some embodiments, the targeting moiety comprises a peptide or protein moiety, such as a DNA binding protein, a CRISPR component protein, a splicing nucleation molecule, a dominant-negative splicing nucleation molecule, an epigenetic modifier, or any combination thereof.
[0296] Peptide or protein moieties may include, but are not limited to, peptide ligands, antibody fragments, or targeting aptamers that bind to receptors such as extracellular receptors, neuropeptides, hormonal peptides, peptide drugs, toxic peptides, viral or microbial peptides, synthetic peptides, and agonist or antagonist peptides.
[0297] The peptide or protein portion may be linear or branched. The peptide or protein portion may have a length of about 5 to about 200 amino acids, about 15 to about 150 amino acids, about 20 to about 125 amino acids, about 25 to about 100 amino acids, or any range therebetween.
[0298] Exemplary peptide or protein moieties for use in the methods and compositions described herein include, but are not limited to, ubiquitin, bicyclic peptides as ubiquitin ligase inhibitors, transcription factors, DNA and protein modifying enzymes such as topoisomerases, topoisomerase inhibitors such as topotecan, DNA methyltransferases such as the DNMT family (e.g., DNMT3a, DNMT3b, DNMTL), protein methyltransferases (e.g., viral lysine methyltransferase (vSET), Histone methyltransferases such as protein-lysine N-methyltransferase (SMYD2), deaminases (e.g., APOBEC, UG1), enhancer of zeste homolog 2 (EZH2), PRMT1, histone-lysine N-methyltransferase (Setdb1), histone methyltransferase (SET2), euchromatin histone-lysine N-methyltransferase 2 (G9a), histone-lysine N-methyltransferase (SUV39H1), and G9a), histone deacetylases (e.g., , HDAC1, HDAC2, HDAC3), enzymes with a role in DNA demethylation (e.g., the TET family of enzymes catalyzes the oxidation of 5-methylcytosine to 5-hydroxymethylcytosine and more highly oxidized derivatives), protein demethylases such as KDM1A and lysine-specific histone demethylase 1 (LSD1), helicases such as DHX9, acetyltransferases, deacetylases (e.g., sirtuins 1, 2, 3, 4, 5, 6, or 7), kinases, phosphatases, ethidium bromide, and Sybr Green. and DNA intercalators such as proflavine, efflux pump inhibitors such as peptidomimetics like phenylalanine arginyl β-naphthylamide or quinoline derivatives, nuclear receptor activators and inhibitors, proteasome inhibitors, competitive inhibitors of enzymes such as those involved in lysosomal storage diseases, protein synthesis inhibitors, nucleases (e.g., Cpf1, Cas9, zinc finger nucleases), one or more fusions thereof (e.g., dCas9-DNMT, dCas9-APOBEC, dCas9-UG1), and KRAB domains,Specific domains derived from proteins include:
[0299] Some examples of peptides include, but are not limited to, fluorescent tags or markers, antigens, antibodies, antibody fragments such as single domain antibodies, ligands and receptors such as glucagon-like peptide-1 (GLP-1), GLP-2 receptor 2, cholecystokinin B (CCKB) and somatostatin receptors, peptide therapeutics such as those that bind to specific cell surface receptors such as G protein-coupled receptors (GPCRs) or ion channels, synthetic or analog peptides derived from naturally occurring biologically active peptides, antimicrobial peptides, pore-forming peptides, tumor-targeting or cytotoxic peptides, and degradative or self-destructive peptides such as apoptosis-inducing peptide signals or photosensitizing peptides.
[0300] The peptides described herein may also include small antigen-binding peptides, e.g., antigen-binding antibodies or antibody-like fragments, such as single-chain antibodies and nanobodies (see, e.g., Steeland et al., 2016, Nanobodies as therapeutics: big opportunities for small antibodies. Drug Discovery Today: 21(7):1076-113). Such small antigen-binding peptides may bind to cytoplasmic, nuclear, or intraorganellar antigens.
[0301] In one aspect, the disclosure includes a cell or tissue comprising any one of the proteins described herein.
[0302] In another aspect, the disclosure includes pharmaceutical compositions comprising the proteins described herein.
[0303] In another aspect, the disclosure includes methods of modulating gene expression by administering a composition comprising a protein described herein.
[0304] DNA-binding domain In some embodiments, the targeting moiety comprises a DNA-binding domain of a protein. DNA-binding proteins have different structural motifs that play important roles in binding to DNA.
[0305] The helix-turn-helix motif is a common DNA recognition motif in repressor proteins. This motif contains two helices: one that recognizes DNA (the so-called recognition helix) and side chains that confer specificity to the binding. They are common in proteins that regulate developmental processes. Sometimes, more than one protein competes for the same sequence or recognizes the same DNA fragment. They may differ in their affinity for the same sequence or DNA conformation through H-bonds, salt bridges, and Van der Waals interactions, respectively.
[0306] DNA binding proteins with the HhH structural motif may participate in non-sequence-specific DNA binding, which occurs through the formation of hydrogen bonds between the protein backbone nitrogens and DNA phosphate groups.
[0307] DNA-binding proteins with the HLH structural motif are transcriptional regulatory proteins and, in principle, are associated with various developmental processes. This motif is longer in terms of residues than the other two motifs. Many of these proteins interact to form homo- and heterodimers. This structural motif consists of two long helical regions; the N-terminal helix binds to DNA, while the loop region allows the protein to dimerize.
[0308] In some transcription factors, the dimerization site with DNA forms a leucine zipper. This motif involves two amphipathic helices, one from each subunit, interacting with each other, resulting in a left-handed coiled-coil supersecondary structure. The leucine zipper is the interdigitation of regularly spaced leucine residues in one helix with leucines from the adjacent helix. Often, the helices involved in the leucine zipper exhibit a heptad sequence (abcdefg), in which residues a and d are hydrophobic and all others are hydrophilic. The leucine zipper motif can mediate homo- or heterodimerization.
[0309] Some eukaryotic transcription factors ++ It displays a unique motif, termed a Zn-finger, in which the ion is coordinated by two Cys and two His residues. This transcription factor contains a trimer with ββ'α stoichiometry. Zn ++ The apparent effect of coordination is the stabilization of a small loop structure in place of the hydrophobic core residues. Each Zn-finger interacts in a conformationally identical manner with a consecutive three-base pair segment in the major groove of the double helix. The protein-DNA interaction is determined by two factors: (i) H-bonding interactions between the α-helix and the DNA segment, often with Arg residues and guanine bases, and (ii) H-bonding interactions with the DNA phosphate backbone, often with Arg and His. An alternative Zn-finger motif uses six Cys to bind Zn ++ Chelates.
[0310] DNA-binding proteins also include the TATA box-binding protein, first identified as a component of the class II initiation factor TFIID. They participate in transcription by all three nuclear RNA polymerases, each of which acts as a subunit. The structure of TBP reveals two α / β structural domains of 89–90 amino acids. The C-terminal or core region binds with high affinity to the TATA consensus sequence (TATAa / tAa / t, SEQ ID NO: xx), which recognizes minor groove determinants and promotes DNA curvature. TBP resembles a molecular saddle. The binding side aligns with the central eight strands of a ten-stranded antiparallel β-sheet. The upper surface contains four α-helices and binds various components of the transcription machinery.
[0311] DNA provides base specificity in the form of nitrogenous bases. The R group of amino acids, including basic residues such as lysine, arginine, histidine, asparagine, and glutamine, can easily interact with adenine in A:T base pairs and guanine in G:C base pairs, where the NH2 and X=O groups of the base pairs can preferably form hydrogen bonds with the amino acid residues glutamine, asparagine, arginine, and lysine.
[0312] In some embodiments, the DNA binding protein is a transcription factor.Transcription factor (TF) can be a modular protein that contains a DNA binding domain that is responsible for specific recognition of base sequence and one or more effector domains that can activate or repress transcription.TF interacts with chromatin and recruits protein complexes that act as coactivators or corepressors.
[0313] Gene editing system In some embodiments, the targeting moiety (e.g., site-specific targeting moiety) comprises one or more components of a gene editing system. As will be understood by those skilled in the art upon reading this specification, and as further described herein, the components of a gene editing system can be used in a variety of contexts, including, but not limited to, gene editing. For example, such components can be used to target agents that physically modify, genetically modify, and / or epigenetically modify target anchor sequences. In some embodiments, the targeting moiety targets one or more nucleotides of the anchor sequence-mediated connection for substitution, addition, and / or deletion. Exemplary gene editing systems include clustered regulatory-mediated short palindromic repeats (CRISPR) systems, zinc finger nucleases (ZFNs), and transcription activator-like effector-based nucleases (TALENs). ZFN-, TALEN-, and CRISPR-based methods are described, for example, in Gaj et al., Trends Biotechnol. 31, No. 7 (2013): 397-405; CRISPR methods of gene editing are described, for example, in Guan et al., Application of CRISPR-Cas system in gene therapy: Pre-clinical progress in animal models. DNA Repair July 30, 2016 [published ahead of print]; Zheng et al., Precise gene deletion and replacement using the CRISPR / Cas9 system in human cells. BioTechniques, Vol. 57, No. 3, September 2014, pp. 115-124.
[0314] For example, in some embodiments, the site-specific targeting moiety comprises a Cas nuclease (e.g., Cas9) and a site-specific guide RNA, as further described herein. In some embodiments, the Cas nuclease is enzymatically inactive, e.g., dCas9, as further described herein.
[0315] In one embodiment, the methods and compositions described herein can be used with CRISPR-based gene editing, in which a guide RNA (gRNA) is used in the clustered regulatory-mediated short palindromic repeat (CRISPR) system for gene editing. The CRISPR system is an adaptive defense system originally discovered in bacteria and archaea. CRISPR systems use RNA-guided nucleases called CRISPR-associated or "Cas" endonucleases (e.g., Cas9 or Cpf1) to cleave foreign DNA. In a typical CRISPR / Cas system, the endonuclease is directed to a target nucleotide sequence (e.g., a site in the genome to be sequence-edited) by a sequence-specific non-coding "guide RNA" that targets single- or double-stranded DNA sequences. Three classes (I-III) of CRISPR systems have been identified. Class II CRISPR systems use a single Cas endonuclease (rather than multiple Cas proteins). One class II CRISPR system includes a type II Cas endonuclease, such as Cas9, a CRISPR RNA ("crRNA"), and a trans-activating crRNA ("tracrRNA"). The crRNA contains a "guide RNA," an RNA sequence of typically about 20 nucleotides, that corresponds to the target DNA sequence. The crRNA also contains a region that binds to the tracrRNA, forming a partially double-stranded structure that is cleaved by RNase III, resulting in a crRNA / tracrRNA hybrid. The crRNA / tracrRNA hybrid then directs the Cas9 endonuclease to recognize and cleave the target DNA sequence. The target DNA sequence must generally be adjacent to a "protospacer adjacent motif" ("PAM") that is specific for a given Cas endonuclease; however, PAM sequences occur throughout a given genome.CRISPR endonucleases identified from various prokaryotic species have unique PAM sequence requirements; examples of PAM sequences include 5'-NGG (Streptococcus pyogenes), 5'-NNAGAA (Streptococcus thermophilus CRISPR1), 5'-NGGNG (Streptococcus thermophilus CRISPR3), and 5'-NNNGATT (Neisseria meningiditis). Some endonucleases, such as the Cas9 endonuclease, associate with G-rich PAM sites, such as 5'-NGG, and perform blunt-end cleavage of target DNA three nucleotides upstream (5') from the PAM site. Another class II CRISPR system contains the V-type endonuclease Cpf1, which is smaller than Cas9; examples include AsCpf1 (derived from Acidaminococcus sp.) and LbCpf1 (derived from Lachnospiraceae sp.). The Cpf1-bound CRISPR array is processed into mature crRNA without the need for tracrRNA; in other words, the Cpf1 system requires only Cpf1 nuclease and crRNA to cleave the target DNA sequence. The Cpf1 endonuclease associates with T-rich PAM sites, such as 5'-TTN. Cpf1 can also recognize a 5'-CTA PAM motif. Cpf1 cleaves target DNA by introducing offset or alternating double-strand breaks with 4- or 5-nucleotide 5' overhangs, for example, cleaving target DNA with 5-nucleotide offset or alternating breaks 18 nucleotides downstream (3') from the PAM site on the coding strand and 23 nucleotides downstream from the PAM site on the complementary strand. The 5-nucleotide overhangs resulting from such offset cleavage allow for more precise genome editing by DNA insertion via homologous recombination than insertion with blunt-ended DNA. See, e.g., Zetsche et al. (2015), Cell 163:759-771.
[0316] Various CRISPR-associated (Cas) genes or proteins can be used in the methods of the present disclosure, and the choice of Cas protein will depend on the specific conditions of the method. Specific examples of Cas proteins include Class II systems such as Cas1, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9, Cas10, Cpf1, C2C1, or C2C3. In some embodiments, the Cas protein, e.g., the Cas9 protein, can be derived from any of a variety of prokaryotic species. In some embodiments, a particular Cas protein, e.g., a particular Cas9 protein, is selected to recognize a specific protospacer adjacent motif (PAM) sequence. In some embodiments, the targeting moiety comprises a sequence-targeting polypeptide, such as an enzyme, e.g., Cas9. In certain embodiments, the Cas protein, e.g., the Cas9 protein, can be obtained from bacteria or archaea or synthesized using known methods. In certain embodiments, the Cas protein can be derived from Gram-positive or Gram-negative bacteria. In certain embodiments, the Cas proteins may be derived from Streptococcus (e.g., S. pyogenes, S. thermophilus), Cryptococcus, Corynebacterium, Haemophilus, Eubacterium, Pasteurella, Prevotella, Veillonella, or Marinobacter. In some embodiments, two or more different Cas proteins, or nucleic acids encoding two or more Cas proteins, can be introduced into a cell, zygote, embryo, or animal to allow, for example, recognition and modification of sites containing the same, similar, or different PAM motifs.In some embodiments, Cas proteins are engineered to deactivate nucleases, e.g., nuclease-deficient Cas9, and recruit activation domains of transcriptional activators or repressors, e.g., E. coli Pol, the ω-subunit of VP64, p65, KRAB, or SID4X, to induce epigenetic modifications, e.g., histone acetyltransferases, histone methyltransferases and demethylases, DNA methyltransferases, and enzymes with roles in DNA demethylation (e.g., the TET family of enzymes catalyzes the oxidation of 5-methylcytosine to 5-hydroxymethylcytosine and more highly oxidized derivatives).
[0317] For gene editing, CRISPR arrays can be designed to contain one or more guide RNA sequences corresponding to the desired target DNA sequence; see, for example, Cong et al. (2013), Science, 339:819-823; Ran et al. (2013), Nature, Protocols, 8:2281-2308. Cas9 requires a gRNA sequence of at least about 16 or 17 nucleotides for DNA cleavage to occur; for Cpf1, a gRNA sequence of at least about 16 nucleotides is required to achieve detectable DNA cleavage.
[0318] Wild-type Cas9 generates double-strand breaks (DSBs) at specific DNA sequences targeted by gRNAs, but several CRISPR endonucleases with modified functionality are available. For example, "nickase"-type Cas9 generates only single-strand breaks; catalytically inactive Cas9 ("dCas9") does not cleave target DNA but interferes with transcription through steric hindrance. dCas9 can be further fused to heterologous effectors to suppress (CRISPRi) or activate (CRISPRa) target gene expression. For example, Cas9 can be fused to a transcriptional silencer (e.g., a KRAB domain) or a transcriptional activator (e.g., a dCas9-VP64 fusion). Catalytically inactive Cas9 (dCas9) fused to a FokI nuclease ("dCas9-FokI") can be used to generate DSBs at target sequences homologous to two gRNAs. See, for example, several CRISPR / Cas9 plasmids disclosed in and publicly available from the Addgene repository (Addgene, 75 Sidney St., Suite 550A, Cambridge, MA 02139; addgene.org / crispr / ). "Double nickase" Cas9, which introduces two separate double-stranded breaks, each directed by a separate guide RNA, has been described by Ran et al. (2013) Cell 154:1380-1389 to achieve more precise genome editing.
[0319] CRISPR technology for editing eukaryotic genes is disclosed in U.S. Patent Application Publication Nos. 2016 / 0138008A1 and US2015 / 0344912A1, and U.S. Patent Nos. 8,697,359, 8,771,945, 8,945,839, 8,999,641, 8,993,233, 8,895,308, 8,865,406, 8,889,418, 8,871,445, 8,889,356, 8,932,814, 8,795,965, and 8,906,616. The Cpf1 endonuclease and corresponding guide RNA and PAM sites are disclosed in US Patent Application Publication No. 2016 / 0208243A1.
[0320] In some embodiments, the desired genome modification involves homologous recombination, in which one or more double-stranded DNA breaks in the target nucleotide sequence are generated by an RNA-guided nuclease and a guide RNA, and the breaks are repaired using a homologous recombination mechanism ("homologous recombination repair"). In such embodiments, a donor template is provided to a cell or subject, encoding the desired nucleotide sequence to be inserted into or knocked into the double-stranded break; examples of suitable templates include single-stranded DNA templates and double-stranded DNA templates (e.g., linked to a polypeptide described herein). Generally, donor templates encoding nucleotide changes spanning a region of less than about 50 nucleotides are provided in the form of single-stranded DNA; larger donor templates (e.g., more than 100 nucleotides) are often provided as double-stranded DNA plasmids. In some embodiments, the donor template is provided to a cell or subject in an amount sufficient to achieve the desired homologous recombination repair, but not persisting in the cell or subject after a given period of time (e.g., after one or more cell division cycles). In some embodiments, the donor template has a core nucleotide sequence that differs from the target nucleotide sequence (e.g., a homologous endogenous genomic region) by at least 1, at least 5, at least 10, at least 20, at least 30, at least 40, at least 50, or more nucleotides. This core sequence is flanked by "homology arms," or regions that exhibit high sequence identity with the target nucleotide sequence; in embodiments, the regions of high identity comprise at least 10, at least 50, at least 100, at least 150, at least 200, at least 300, at least 400, at least 500, at least 600, at least 750, or at least 1000 nucleotides on each side of the core sequence. In some embodiments, where the donor template is in the form of single-stranded DNA, the core sequence is flanked by homology arms that comprise at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, or at least 100 nucleotides on each side of the core sequence.In embodiments where the donor template is in the form of double-stranded DNA, the core sequence is flanked by homology arms comprising at least 500, at least 600, at least 700, at least 800, at least 900, or at least 1000 nucleotides on each side of the core sequence. In one embodiment, two separate double-stranded breaks are introduced into the target nucleotide sequence of a cell or subject using "double nickase" Cas9 (see Ran et al. (2013) Cell 154:1380-1389) prior to delivery of the donor template.
[0321] In some embodiments, the composition comprises a polypeptide described herein linked to a gRNA and a targeted nuclease, e.g., Cas9, e.g., wild-type Cas9, nickase Cas9 (e.g., Cas9 D10A), dead Cas9 (dCas9), eSpCas9, Cpf1, C2C1, or C2C3, or a nucleic acid encoding such a nuclease. The choice of nuclease and gRNA is determined by whether the targeted mutation is a deletion, substitution, or addition of nucleotides, e.g., whether a deletion, substitution, or addition of nucleotides is relative to the targeted sequence. Fusion of a catalytically inactive endonuclease, e.g., dead Cas9 (dCas9, e.g., D10A; H840A), tethered to all or a portion (e.g., biologically active portion) of (one or more) effector domains (e.g., epigenome editors including, but not limited to, DNMT3a, DNMT3L, DNMT3b, KRAB domains, Tet1, p300, VP64, and fusions of the above), can be linked to a polypeptide for directing the composition to a specific DNA site by one or more RNA sequences (e.g., DNA recognition elements such as, but not limited to, zinc finger arrays, sgRNAs, TAL arrays, peptide nucleic acids, etc., as described herein), to create a chimeric protein capable of modulating the activity and / or expression of one or more target nucleic acid sequences (e.g., methylating or demethylating DNA sequences).
[0322] As used herein, a "biologically active portion of an effector domain" refers to a portion (e.g., a "minimal" or "core" domain) that maintains the function of the effector domain (e.g., completely, partially, minimally). In some embodiments, fusion of dCas9 with all or a portion of one or more effector domains of an epigenetic modifier (e.g., a DNA methylase or enzyme that plays a role in DNA demethylation, e.g., DNMT3a, DNMT3b, DNMT3L, DNMT inhibitors, combinations thereof, TET family enzymes, protein acetyltransferases or deacetylases, dCas9-DNMT3a / 3L, dCas9-DNMT3a / 3L / KRAB, dCas9 / VP64, etc.) is linked to a polypeptide to create a chimeric protein that is useful in the methods described herein. Thus, in some embodiments, a nucleic acid encoding a dCas9-methylase fusion is linked to a polypeptide and administered to a subject in need thereof along with a site-specific gRNA or antisense DNA oligonucleotide that targets the anchor sequence (e.g., a CTCF-binding motif), thereby decreasing the affinity or ability of the anchor sequence to bind to a nucleation protein. In other embodiments, a nucleic acid encoding a dCas9-enzyme fusion is linked to a polypeptide along with a site-specific gRNA or antisense DNA oligonucleotide that targets the fusion to a connecting anchor sequence (e.g., a CTCF-binding motif), thereby increasing the affinity or ability of the anchor sequence to bind to a nucleation protein. In some embodiments, all or part of the effector domain of one or more methyltransferases or enzymes associated with demethylation is fused to an inactive nuclease, e.g., dCas9, and linked to a polypeptide.Exemplary dCas9 fusion methods and compositions that are applicable to the methods and compositions described herein are known and are described, for example, in Kearns et al., Functional annotation of native enhancers with a Cas9-histone demethylase fusion. Nature Methods 12:401-403 (2015); and McDonald et al., Reprogrammable CRISPR / Cas9-based system for inducing site-specific DNA methylation. Biology Open 2016: doi:10.1242 / bio.019067.
[0323] In other embodiments, the effector domains (all or biologically active portions) of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more methyltransferases or enzymes that play a role in DNA demethylation are fused to dCas9 and linked to a polypeptide. The chimeric proteins described herein may also include a linker, e.g., an amino acid linker, as described herein. In some embodiments, the linker comprises two or more amino acids, e.g., one or more GS sequences. In some embodiments, fusions of Cas9 (e.g., dCas9) with two or more effector domains (e.g., of DNA methylases or enzymes that play a role in DNA demethylation) include one or more intervening linkers (e.g., GS linkers) between the domains and are linked to a polypeptide. In some embodiments, dCas9 is fused to multiple (e.g., 2-5, e.g., 2, 3, 4, 5) effector domains with intervening linkers and linked to a polypeptide.
[0324] In some embodiments, the targeting moiety comprises one or more components of the CRISPR system described above.
[0325] For example, in some embodiments, the targeting moiety comprises a gRNA comprising a targeting domain that hybridizes to a nucleic acid comprising a target anchor sequence and / or has a sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, or at least 99% identical to the complement of a nucleic acid comprising a target anchor sequence. In some embodiments, the gRNA is a site-specific gRNA whose targeting domain does not hybridize to at least one nucleic acid comprising a non-target anchor sequence. In some embodiments, the site-specific gRNA has the structure I: (I)XYZ where X and Z are 5' and 3' site-specific targeting sequences for the target CTCF binding motif, respectively, and Y is (a) an RNA sequence complementary to the sequence of SEQ ID NO: 1; (b) an RNA sequence that is at least 75%, 80%, 85%, 90%, or 95% identical to an RNA sequence complementary to the sequence of SEQ ID NO: 1; (c) an RNA sequence complementary to the sequence of SEQ ID NO: 1 with at least 1, 2, 3, 4, 5, but less than 15, 12, or 10 nucleotide additions, substitutions, or deletions; (d) an RNA sequence complementary to the sequence of SEQ ID NO:2; (e) an RNA sequence that is at least 75%, 80%, 85%, 90%, or 95% identical to an RNA sequence complementary to the sequence of SEQ ID NO:2; (f) an RNA sequence complementary to the sequence of SEQ ID NO: 2, which has at least 1, 2, 3, 4, 5, but less than 15, 12, or 10 nucleotide additions, substitutions, or deletions; (selected from Contains an array of
[0326] In some embodiments, X and Z are each 2 to 50 nucleotides in length, eg, 2 to 20, 2 to 10, 2 to 5 nucleotides in length.
[0327] In some embodiments, compositions or methods are described that include gRNAs that specifically target CTCF binding motifs associated with oncogenes, tumor suppressors, or diseases associated with nucleotide repeats, e.g., CTCFBSDB 2.0: Database For CTCF Binding Motifs And Genome Organization.
[0328] In some embodiments, a pharmaceutical composition is provided comprising a guide RNA described herein.
[0329] In some embodiments, the methods described herein include methods of delivering one or more CRISPR system components described above to a subject, e.g., to the nucleus of a cell or tissue of a subject, by linking such components to a polypeptide described herein.
[0330] Connecting nucleating molecules In some embodiments, the targeting moiety comprises a connecting nucleation molecule, a nucleic acid encoding the connecting nucleation molecule, or a combination thereof. In some embodiments, the anchor sequence-mediated connection is mediated by a first connecting nucleation molecule bound to a first anchor sequence, a second connecting nucleation molecule bound to a non-contiguous second anchor sequence, and an association between the first and second connecting nucleation molecules. In some embodiments, the connecting nucleation molecule can disrupt the binding of an endogenous connecting nucleation molecule to its binding site, for example, by competitive binding.
[0331] The connection nucleation molecule may be, for example, CTCF, cohesin, USF1, YY1, TATA box-binding protein-associated factor 3 (TAF3), ZNF143-binding motif, or another polypeptide that promotes the formation of anchor sequence-mediated connections. The connection nucleation molecule may be an endogenous polypeptide or other protein, such as a transcription factor, e.g., autoimmune regulatory factor (AIRE), another factor, e.g., X-inactivation specific transcript (XIST), or an engineered polypeptide engineered to recognize a specific DNA sequence of interest, e.g., having a zinc finger, leucine zipper, or bHLH domain for sequence recognition. The connection nucleation molecule can modulate DNA interactions within or around the anchor sequence-mediated connection. For example, the connection nucleation molecule can recruit other factors to the anchor sequence that alter the formation or disruption of the anchor sequence-mediated connection.
[0332] The connection nucleation molecule may also have a dimerization domain for homo- or heterodimerization. For example, one or more endogenous and engineered connection nucleation molecules may interact to form an anchor sequence-mediated connection. In some embodiments, the connection nucleation molecule is engineered to further include a stabilization domain, e.g., an aggregation interaction domain, to stabilize the anchor sequence-mediated connection. In some embodiments, the connection nucleation molecule is engineered to bind to a target sequence, e.g., to modulate the target sequence binding affinity. In some embodiments, the connection nucleation molecule is selected or engineered to have a binding affinity for the anchor sequence in the anchor sequence-mediated connection.
[0333] The connection nucleation molecules and their corresponding anchor sequences can be identified using cells carrying an inactivating mutation in CTCF and chromosome conformation capture or 3C-based methods, such as Hi-C or high-throughput sequencing, to examine topological interactions between topologically related domains, such as distal DNA regions or loci, in the absence of CTCF. Long-range DNA interactions can also be identified. Further analysis can include ChIA-PET analysis using baits such as cohesin, YY1 or USF1, ZNF143-binding motifs, and MS to identify complexes associated with the bait.
[0334] In some embodiments, one or more attached nucleating molecules have a binding affinity for the anchor sequence that is higher or lower than a reference value, e.g., the binding affinity for the anchor sequence in the absence of the alteration.
[0335] In some embodiments, the binding affinity of a connection nucleating molecule, for example, to an anchor sequence within an anchor sequence-mediated connection, is modulated to alter its interaction with the anchor sequence-mediated connection.
[0336] In some embodiments,
[0337] dissimilar parts In some embodiments, the compositions, agents, and / or fusion molecules described herein may comprise one or more heterologous moieties, which may be an effector (e.g., a drug, a small molecule), a tag (e.g., a fluorophore, a photosensitive agent such as KillerRed), or any of the editing or targeting moieties described herein.
[0338] In some embodiments, heterologous moieties may be linked to the transmembrane polypeptides described herein, hi some embodiments, the transmembrane polypeptides described herein are linked to one or more heterologous moieties.
[0339] In one aspect, the disclosure includes a cell or tissue comprising any one of the heterologous moieties described herein.
[0340] In another aspect, the disclosure includes pharmaceutical compositions that include a heterologous moiety described herein.
[0341] In another aspect, the disclosure includes methods of modulating gene expression by administering a composition comprising a heterologous moiety described herein.
[0342] In one embodiment, the heterologous moiety is any targeting moiety that modulates the two-dimensional structure of chromatin (ie, modulates the structure of chromatin in a way that alters its two-dimensional presentation).
[0343] In one embodiment, the heterologous moiety is a small molecule (e.g., a peptidomimetic or small organic molecule having a molecular weight of less than 2000 daltons), a peptide or polypeptide (e.g., a non-ABX n C polypeptide, e.g., an antibody or antigen-binding fragment thereof), nucleic acid (e.g., siRNA, mRNA, RNA, DNA, modified DNA or RNA, antisense DNA oligonucleotide, antisense RNA, ribozyme, therapeutic mRNA encoding a protein), nanoparticle, aptamer, or drug with low PK / PD.
[0344] In some embodiments, the heterologous moiety can be cleaved from the polypeptide (eg, after administration) by specific proteolytic or enzymatic cleavage (eg, by TEV protease, thrombin, factor Xa, or enteropeptidase).
[0345] Effector part The heterologous moiety may be an effector moiety with effector activity. The effector moiety can modulate biological activity, for example, increasing or decreasing enzyme activity, gene expression, cell signaling, and cell or organ function. Effector activity may also include binding regulatory proteins to modulate the activity of regulators, such as transcription or translation. Effector activity may also include activator or inhibitor (or "negative effector") functions described herein. For example, a heterologous moiety may induce enzyme activity by increasing substrate affinity in an enzyme; e.g., fructose 2,6-bisphosphate activates phosphofructokinase 1, increasing the rate of glycolysis in response to insulin. In another example, a heterologous moiety may inhibit substrate binding to a receptor and inhibit its activation; e.g., naltrexone and naloxone bind to opioid receptors without activating them, blocking the receptor's ability to bind opioids. Effector activity may also include modulating protein stability / degradation and / or transcript stability / degradation. For example, proteins can be targeted to proteins for degradation by the polypeptide cofactor, ubiquitin, marking them for degradation. In another example, the heterologous moiety inhibits enzyme activity by blocking the active site of the enzyme; for example, methotrexate is a structural analog of tetrahydrofolate, a coenzyme for the dihydrofolate reductase enzyme, that binds 1000 times more tightly to dihydrofolate reductase than the natural substrate, inhibiting nucleotide base synthesis.
[0346] In some embodiments, the composition comprises a targeting moiety (e.g., a gRNA, a membrane translocation polypeptide) operably linked to an effector moiety that binds to an anchor sequence and modulates the formation of a connection mediated by the anchor sequence.
[0347] In some embodiments, the effector molecule is a chemical, e.g., a chemical that modulates cytosine (C) or adenine (A) (e.g., sodium bisulfite, ammonium bisulfite). In some embodiments, the effector moiety has enzymatic activity (e.g., methyltransferase, demethylase, nuclease (e.g., Cas9), deaminase). In some embodiments, the effector moiety sterically hinders the formation of the anchor sequence-mediated connection (e.g., membrane-translocating polypeptide + nanoparticle (def: 1-100 nm)).
[0348] The effector moiety having effector activity may be any one of the small molecules, peptides, nucleic acids, nanoparticles, aptamers, and drugs with low PK / PD described herein.
[0349] Negative effector part In some embodiments, the effector is an inhibitor or "negative effector." In the context of a negative effector moiety that modulates the formation of an anchor sequence-mediated connection, in some embodiments, the negative effector moiety is characterized by a decrease in dimerization of an endogenous nucleation polypeptide when present compared to when the negative effector moiety is absent. For example, in some embodiments, the negative effector moiety is or comprises a variant of the dimerization domain of an endogenous nucleation polypeptide, or a dimerization portion thereof.
[0350] Dominant-negative connecting nucleation molecules For example, in certain embodiments, the anchor sequence-mediated connection is altered (e.g., disrupted) by using a dominant-negative effector, for example, a protein that recognizes and binds to an anchor sequence (e.g., a CTCF binding motif) but has an inactive (e.g., mutated) dimerization domain, for example, a dimerization domain that cannot form a functional anchor sequence-mediated connection. For example, the zinc finger domain of CTCF can be altered so that it binds to a specific anchor sequence (by adding a zinc finger that recognizes an adjacent nucleic acid), but the homodimerization domain is altered to prevent the interaction of the engineered CTCF with the endogenous form of CTCF. DNA encoding this protein can be administered to a subject in need thereof.
[0351] In some embodiments, the composition comprises a synthetic connection nucleating molecule having a selected binding affinity for an anchor sequence within the target anchor sequence-mediated junction (the binding affinity may be at least 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or more higher or lower than the affinity of the endogenous connection nucleating molecule that associates with the target anchor sequence. The synthetic connection nucleating molecule may have 30-90%, 30-85%, 30-80%, 30-70%, 50-80%, 50-90% amino acid sequence identity to the endogenous connection nucleating molecule). The connection nucleating molecule is capable of disrupting binding of the endogenous connection nucleating molecule to its anchor sequence, such as by competitive binding. In some further embodiments, the connection nucleating molecule is engineered to bind to a novel anchor sequence within the anchor sequence-mediated junction.
[0352] In some embodiments, the dominant-negative effector has a domain that recognizes a specific DNA sequence (e.g., an anchor sequence flanked by sequences that confer sequence specificity, such as a CTCF anchor sequence), and a second domain that provides steric presence near the anchor sequence. The second domain may comprise a dominant-negative conjugated nucleating molecule or fragment thereof, a polypeptide that interferes with conjugated nucleating molecule sequence recognition (e.g., the amino acid backbone of a peptide / nucleic acid or PNA), a nucleic acid sequence ligated to a small molecule that confers steric interference, or any other combination of a DNA recognition element and a steric blocking agent.
[0353] Epigenetic modifiers In some embodiments, heterologous moiety is epigenetic modifying agent.Epigenetic modifying agent useful in the methods and compositions described herein includes, for example, agents that affect DNA methylation, histone acetylation and RNA-associated silencing.In some embodiments, the methods described herein comprise sequence-specific targeting of epigenetic enzymes (for example, the enzymes that generate or remove epigenetic marks, for example, acetylation and / or methylation). Exemplary epigenetic enzymes that can be targeted to anchor sequences using the CRISPR methods described herein include DNA methylases (e.g., DNMT3a, DNMT3b, DNMTL), DNA demethylation (e.g., the TET family), histone methyltransferases, histone deacetylases (e.g., HDAC1, HDAC2, HDAC3), sirtuins 1, 2, 3, 4, 5, 6, or 7, lysine-specific histone demethylase 1 (LSD1), histone-lysine-N-methyltransferase (Setdb1), euchromatin histone-lysine N-methyltransferase 2 (G9a), histone-lysine N-methyltransferase (SUV39H1), enhancer of zeste homolog 2 (EZH2), viral lysine methyltransferase (vSET), histone methyltransferase (SET2), and protein-lysine N-methyltransferase (SMYD2). Examples of such epigenetic modifiers are described, for example, in de Groote et al., Nuc. Acids Res. (2012):1-18.
[0354] In some embodiments, epigenetic modifiers useful herein include constructs described in Koferle et al., Genome Medicine 7:59 (2015):1-3 (e.g., Table 1), which is incorporated herein by reference.
[0355] Tagging or monitoring part The heterologous moiety may be a tag for labeling or monitoring the polypeptide described herein or another heterologous moiety linked to the polypeptide. The tagging or monitoring moiety can be removed by chemical agents or enzymatic cleavage, such as proteolysis or intein splicing. Affinity tags can be useful for purifying tagged polypeptides using affinity techniques. Some examples include chitin-binding protein (CBP), maltose-binding protein (MBP), glutathione-S-transferase (GST), and poly(His) tags. Solubilization tags can be useful for recombinant proteins expressed in chaperone-deficient species, such as E. coli, to aid in proper protein folding and protect them from precipitation. Some examples include thioredoxin (TRX) and poly(NANP). The tagging or monitoring moiety may include a light-sensitive tag, e.g., fluorescent. Fluorescent tags are useful for visualization. GFP and its variants are some commonly used examples of fluorescent tags. Protein tags may be specifically enzymatically modified (e.g., biotinylated by biotin ligase) or chemically modified (e.g., reacted with FlAsH-EDT2 for fluorescent imaging). To connect proteins to multiple other components, tagging and monitoring moieties are often combined. Tagging or monitoring moieties can also be removed by specific proteolysis or enzymatic cleavage (e.g., by TEV protease, thrombin, factor Xa, or enteropeptidase).
[0356] The tagging or monitoring moiety may be a small molecule, peptide, nucleic acid, nanoparticle, aptamer, or other agent.
[0357] nucleic acid The heterologous moiety may be a nucleic acid. The nucleic acid heterologous moiety may include, but is not limited to, DNA, RNA, and artificial nucleic acids. The nucleic acid may include, but is not limited to, genomic DNA, cDNA, modified DNA, antisense DNA oligonucleotides, tRNA, mRNA, rRNA, modified RNA, miRNA, gRNA, and siRNA or other RNAi molecules. In one embodiment, the nucleic acid is an siRNA for targeting a gene expression product. In another embodiment, the nucleic acid includes one or more nucleoside analogs described herein.
[0358] The nucleic acid has a length of about 2 to about 5000 nt, about 10 to about 100 nt, about 50 to about 150 nt, about 100 to about 200 nt, about 150 to about 250 nt, about 200 to about 300 nt, about 250 to about 350 nt, about 300 to about 500 nt, about 10 to about 1000 nt, about 50 to about 1000 nt, about 100 to about 1000 nt, about 1000 to about 2000 nt, about 2000 to about 3000 nt, about 3000 to about 4000 nt, about 4000 to about 5000 nt, or any range therebetween.
[0359] Some examples of nucleic acids include, but are not limited to, nucleic acids that hybridize to endogenous genes (e.g., gRNA or antisense ssDNA as described elsewhere herein), nucleic acids that hybridize to exogenous nucleic acids such as viral DNA or RNA, nucleic acids that hybridize to RNA, nucleic acids that interfere with gene transcription, nucleic acids that interfere with RNA translation, nucleic acids that stabilize RNA such as by targeting it for degradation or that destabilize RNA, nucleic acids that interfere with DNA or RNA binding factors by interfering with their expression or their function, nucleic acids that are linked to and modulate the function of intracellular proteins, and nucleic acids that are linked to and modulate the function of intracellular protein complexes.
[0360] The present disclosure contemplates the use of RNA therapeutic agents (such as modified RNA) as heterologous moieties in the compositions described herein.For example, modified mRNA encoding a protein of interest can be linked to the polypeptide described herein and expressed in vivo in a subject.
[0361] In some embodiments, the modified RNA or DNA oligonucleotide linked to the polypeptide described herein has modified nucleoside or nucleotide.Such modifications are known and are described, for example, in WO2012 / 019168.Further modifications are described, for example, in WO2015038892;WO2015038892;WO2015089511;WO2015196130;WO2015196118 and WO2015196128A2.
[0362] In some embodiments, the modified RNA or DNA oligonucleotides linked to the polypeptides described herein have one or more terminal modifications, such as a 5' cap structure and / or a polyA tail (e.g., 100-200 nucleotides in length). The 5' cap structure can be selected from the group consisting of CapO, CapI, ARCA, inosine, Nl-methyl-guanosine, 2'fluoro-guanosine, 7-deaza-guanosine, 8-oxo-guanosine, 2-amino-guanosine, LNA-guanosine, and 2-azido-guanosine. In some cases, the modified RNA also contains a 5' UTR and a 3' UTR that include at least one Kozak sequence. Such modifications are known and are described, for example, in WO2012135805 and WO2013052523. Further terminal modifications are described, for example, in WO2014164253 and WO2016011306, WO2012045075 and WO2014093924.
[0363] Chimeric enzymes for synthesizing capped RNA molecules (e.g., modified mRNAs) that may contain at least one chemical modification are described in WO2014028429.
[0364] In some embodiments, modified mRNA can be circularized or concatenated to generate a translationally competent molecule that supports the interaction of polyA-binding protein and 5'-end binding protein. The mechanism of circularization or concatenation can occur through at least three different routes: 1) chemical, 2) enzymatic, and 3) ribozyme-catalyzed. The newly formed 5'- / 3'-bond can be intramolecular or intermolecular. Such modifications are described, for example, in WO2013151736.
[0365] The method of producing and purifying modified RNA is known and disclosed in the art.For example, modified RNA is produced using only in vitro transcription (IVT) enzymatic synthesis.The method of producing IVT polynucleotide is known in the art and is described in WO2013151666, WO2013151668, WO2013151663, WO2013151669, WO2013151670, WO2013151664, WO2013151665, WO2013151671, WO2013151672, WO2013151667 and WO2013151736.S. Purification methods include contacting a sample with a surface linked to multiple thymidines or derivatives thereof and / or multiple uracils or derivatives thereof (polyT / U) under conditions in which the RNA transcripts bind to the surface, and eluting the purified RNA transcripts from the surface (WO2014152031); using ion (e.g., anion) exchange chromatography, which allows for the separation of longer RNAs up to 10,000 nucleotides in length in a scalable manner (WO2014144767); and purifying RNA transcripts containing polyA tails by subjecting the modified RMNA sample to DNAse treatment (WO2014152030).
[0366] Modified RNAs encoding proteins in the areas of human disease, antibodies, viruses, and various in vivo settings are known and are disclosed, for example, in Table 6 of International Publication Nos. WO2013151666, WO2013151668, WO2013151663, WO2013151669, WO2013151670, WO2013151664, WO2013151665, WO2013151736; Tables 6 and 7 of International Publication No. WO2013151672; Tables 6, 178 and 179 of International Publication No. WO2013151671; and Tables 6, 185 and 186 of International Publication No. WO2013151667. Any of the above can be synthesized as an IVT polynucleotide, chimeric polynucleotide or circular polynucleotide and ligated to a polypeptide described herein, each of which may contain one or more modified nucleotides or terminal modifications.
[0367] Peptide Oligonucleotide Conjugates The heterologous moiety may be a peptide oligonucleotide conjugate. A peptide oligonucleotide conjugate includes a chimeric molecule (such as a peptide / nucleic acid hybrid) comprising a nucleic acid moiety linked to a peptide moiety. In some embodiments, the peptide moiety may comprise any peptide or protein moiety described herein. In some embodiments, the nucleic acid moiety may comprise any nucleic acid or oligonucleotide described herein, such as DNA or RNA, or modified DNA or RNA.
[0368] In some embodiments, peptide oligonucleotide conjugates include peptide antisense oligonucleotide conjugates.In some embodiments, peptide oligonucleotide conjugates are synthetic oligonucleotides with chemically modified backbones.Peptide oligonucleotide conjugates can bind to both DNA and RNA targets in a sequence-specific manner to form double-stranded structures.When bound to double-stranded DNA (dsRNA) targets, peptide oligonucleotide conjugates can displace one of the DNA strands in the double strand by strand invasion to form a triple-stranded structure, and the displaced DNA strand can exist as a single-stranded D-loop.
[0369] In some embodiments, peptide oligonucleotide conjugates may be cell and / or tissue specific targeting (can be directly conjugated to oligos, peptides, and / or proteins, etc.).
[0370] In some embodiments, the peptide-oligonucleotide conjugate comprises a transmembrane polypeptide, eg, a transmembrane polypeptide described elsewhere herein.
[0371] The solid-phase synthesis of several peptide-oligonucleotide conjugates is described, for example, in Williams et al., 2010, Curr. Protoc. Nucleic Acid Chem., Chapter 4, Section 41, doi: 10.1002 / 0471142700.nc0441s42. The synthesis and characterization of very short peptide-oligonucleotide conjugates and the stepwise solid-phase synthesis of peptide-oligonucleotide conjugates on novel solid supports are described, for example, in Bongardt et al., Innovation Perspective. Solid Phase Synth. Comb. Libr., Collect. Pap., Int. Symp., 5th Edition, 1999, pp. 267-270; Antopolsky et al., Helv. Chim. Acta, 1999, Vol. 82, pp. 2130-2140.
[0372] nanoparticles The heterogeneous moiety may be a nanoparticle. Nanoparticles include inorganic materials having sizes ranging from about 1 to about 1,000 nanometers, about 1 to about 500 nanometers, about 1 to about 100 nm, about 30 nm to about 200 nm, about 50 nm to about 300 nm, about 75 nm to about 200 nm, about 100 nm to about 200 nm, and any range therebetween. Nanoparticles have a composite structure with nanoscale dimensions. In some embodiments, nanoparticles are typically spherical, although different morphologies are possible depending on the nanoparticle composition. The portion of a nanoparticle that comes into contact with the external environment is generally identified as the nanoparticle's surface. For the nanoparticles described herein, size limitations can be limited to two dimensions; thus, nanoparticles include composite structures having diameters of about 1 to about 1,000 nm, where the specific diameter depends on the nanoparticle composition and the intended use of the nanoparticle through experimental design. For example, nanoparticles used in therapeutic applications typically have a size of about 200 nm or less.
[0373] Further desirable properties of nanoparticles, such as surface charge and steric stabilization, may vary depending on the specific application of interest. Exemplary properties that may be desirable in clinical applications, such as cancer treatment, are described in Davis et al., Nature, 2008, Vol. 7, pp. 771-782; Duncan, Nature, 2006, Vol. 6, pp. 688-701; and Allen, Nature, 2002, Vol. 2, pp. 750-763, each of which is incorporated herein by reference in its entirety. Further properties can be identified by those skilled in the art upon reading this disclosure. The dimensions and properties of nanoparticles can be detected by techniques known in the art. Exemplary techniques for detecting particle size include, but are not limited to, dynamic light scattering (DLS) and various microscopes, such as transmission electron microscopes (TEM) and atomic force microscopes (AFM). Exemplary techniques for detecting particle morphology include, but are not limited to, TEM and AFM. Exemplary techniques for detecting the surface charge of nanoparticles include, but are not limited to, zeta potential methods. Further techniques suitable for detecting other chemical properties include: 1 H, 11 B, and 13 C and 19 These include F NMR, UV / Vis and Infrared / Raman spectroscopy and fluorescence spectroscopy (when the nanoparticles are used in combination with a fluorescent label) and further techniques that can be identified by one skilled in the art.
[0374] low molecule In one embodiment, the targeting moiety is a small molecule that alters one or more DNA methylation sites within the anchor sequence-mediated connection, for example, mutating methylated cysteine to thymine. For example, bisulfite compounds, such as sodium bisulfite, ammonium bisulfite, or other bisulfite salts, can be used to alter one or more DNA methylation sites, for example, to change the nucleotide sequence from cysteine to thymine.
[0375] The heterologous moiety may be a small molecule. Examples of small molecule moieties include, but are not limited to, small peptides, peptidomimetics (e.g., peptoids), amino acids, amino acid analogs, synthetic polynucleotides, polynucleotide analogs, nucleotides, nucleotide analogs, organic and inorganic compounds (including heteroorganic and organometallic compounds) generally having a molecular weight of less than about 5,000 grams / mole, for example, organic or inorganic compounds having a molecular weight of less than about 2,000 grams / mole, for example, organic or inorganic compounds having a molecular weight of less than about 1,000 grams / mole, for example, organic or inorganic compounds having a molecular weight of less than about 500 grams / mole, as well as salts, esters, and other pharmaceutically acceptable forms of such compounds. Small molecules may include, but are not limited to, neurotransmitters, hormones, drugs, toxins, viral or microbial particles, synthetic molecules, and agonists or antagonists.
[0376] Examples of suitable small molecules include those described in "The Pharmacological Basis of Therapeutics," Goodman and Gilman, McGraw-Hill, New York, NY (1996), 9th ed., Drugs Acting at Synaptic and Neuroeffector Junctional Sites; Drugs Acting on the Central Nervous System; Autacoids: Drug Therapy of Inflammation; Water, Salts, and Ions; Drugs Affecting Renal Function and Electrolyte Metabolism; Cardiovascular Drugs; Drugs Affecting Gastrointestinal Function; Drugs Affecting Uterine Motility; Chemotherapy of Parasitic Infections; Chemotherapy of Microbial Diseases; Chemotherapy of Neoplastic Diseases; Drugs Used for Immunosuppression; Drugs Acting on Blood-Forming Organs; Hormones and Hormone Antagonists; Vitamins, Dermatology; and those described in the Toxicology section (all of which are incorporated herein by reference). Some examples of small molecules include, but are not limited to, prion drugs such as tacrolimus, ubiquitin ligase or HECT ligase inhibitors such as heclin, histone modifying drugs such as sodium butyrate, enzyme inhibitors such as 5-aza-cytidine, anthracyclines such as doxorubicin, beta-lactams such as penicillin, antibacterial agents, chemotherapeutic agents, antiviral agents, modulators derived from other organisms such as VP64, and drugs with poor bioavailability, such as chemotherapeutic agents with deficient pharmacokinetics.
[0377] In some embodiments, the small molecule is an epigenetic modifier, such as those described in de Groote et al., Nuc. Acids Res. (2012): 1-18. Exemplary small molecule epigenetic modifiers are described in Lu et al., J. Biomolecular Screening 17(5):555-71, e.g., Table 1 or 2, which are incorporated herein by reference. In some embodiments, the epigenetic modifier comprises vorinostat, romidepsin. In some embodiments, the epigenetic modifier comprises an inhibitor of class I, II, III, and / or IV histone deacetylase (HDAC). In some embodiments, the epigenetic modifier comprises an activator of SirTI. In some embodiments, the epigenetic modifier is garcinol, Lys-CoA, C646, (+)-JQI, I-BET, BICI, MS120, DZNep, UNC0321, EPZ004777, AZ505, AMI-I, pyrazole amide 7b, benzo[d]imidazole 17b, acylated dapsone derivatives (e.g., PRMTI), methylstat, 4,4′-dicarboxy-2,2′-bipyridine, SID8573 6331, hydroxamate analog 8, tanylcypromy, bisguanidine and biguanide polyamine analogs, UNC669, Vidaza, decitabine, sodium phenylbutyrate (SDB), lipoic acid (LA), quercetin, valproic acid, hydralazine, bactrim, green tea extract (e.g., epigallocatechin gallate (EGCG)), curcumin, sulforaphane, and / or allicin / diallyl disulfide. In some embodiments, the epigenetic modifier inhibits DNA methylation, e.g., is an inhibitor of DNA methyltransferase (e.g., 5-azacytidine and / or decitabine). In some embodiments, the epigenetic modifier modifies histone modifications, e.g., histone acetylation, histone methylation, histone sumoylation, and / or histone phosphorylation. In some embodiments, the epigenetic modifying agent is an inhibitor of histone deacetylase (eg, vorinostat and / or trichostatin A).
[0378] In some embodiments, the small molecule is a pharmaceutically active agent. In one embodiment, the small molecule is an inhibitor of a metabolic activity or component. Useful classes of pharmaceutically active agents include, but are not limited to, antibiotics, anti-inflammatory agents, angiogenic or vasoactive agents, growth factors, and chemotherapeutic (anti-neoplastic) agents (e.g., tumor suppressors). One or a combination of molecules from the categories and examples described herein or from (Orme-Johnson 2007, Methods Cell Biol. 2007;80:813-26) can be used. In one embodiment, the present disclosure includes a composition comprising an antibiotic, an anti-inflammatory agent, an angiogenic or vasoactive agent, a growth factor, or a chemotherapeutic agent.
[0379] Oligonucleotide Aptamers The heterologous part may be an oligonucleotide aptamer. The aptamer part is an oligonucleotide or peptide aptamer. The oligonucleotide aptamer is a single-stranded DNA or RNA (ssDNA or ssRNA) molecule that can bind to a preselected target, such as a protein or peptide, with high affinity and specificity.
[0380] Oligonucleotide aptamers are nucleic acid species that can be engineered through repeated rounds of in vitro selection or equivalently, SELEX (Systematic Evolution of Ligands by Exponential Enrichment) to bind to various molecular targets, such as small molecules, proteins, nucleic acids, and even cells, tissues, and organisms. Aptamers provide discriminatory molecular recognition and can be produced by chemical synthesis. Furthermore, aptamers have desirable storage properties and cause little or no immunogenicity in therapeutic applications.
[0381] Both DNA and RNA aptamers exhibit robust binding affinities to a variety of targets. For example, DNA and RNA aptamers have been selected for lysozyme, thrombin, human immunodeficiency virus trans-acting response element (HIV TAR), https: / / en.wikipedia.org / wiki / Aptamer - cite_note-10, hemin, interferon-γ, vascular endothelial growth factor (VEGF), prostate-specific antigen (PSA), dopamine, and the non-classical oncogene, heat shock factor 1 (HSF1).
[0382] Diagnostic techniques for aptamer-based plasma protein profiling include aptamer plasma proteomics, which will enable future multi-biomarker protein measurements that can aid in the diagnostic differentiation between disease and healthy states.
[0383] Peptide aptamers The heterologous moiety may be a peptide aptamer. Peptide aptamers have one or more short variable peptide domains, including small peptides of 12-14 kDa. Peptide aptamers can be designed to specifically bind to and interfere with protein-protein interactions inside cells.
[0384] Peptide aptamers are artificial proteins selected or engineered to bind to specific target molecules. These proteins contain one or more peptide loops with variable sequences. They are typically isolated from combinatorial libraries and often subsequently improved by site-directed mutagenesis or round-trip selection of variable region mutagenesis. In vivo, peptide aptamers can bind to cellular protein targets and exert biological effects, such as interfering with normal protein interactions between their target molecules and other proteins. In particular, variable peptide aptamer loops bound to transcription factor binding domains are screened against target proteins bound to transcription factor activation domains. In vivo binding of peptide aptamers to their targets using this selection strategy is detected as the expression of downstream yeast marker genes. Such experiments identify the specific proteins bound by the aptamer and the protein interactions that the aptamer disrupts in a phenotype-causing manner. Furthermore, peptide aptamers derivatized with appropriate functional moieties can cause specific post-translational modifications of their target proteins or alter the subcellular localization of the target.
[0385] Peptide aptamers can also recognize targets in vitro. They are useful in place of antibodies in biosensors, where they are used to detect active isoforms of proteins from populations containing both inactive and active protein forms. Derivatives known as tadpoles, in which the "head" of a peptide aptamer is covalently linked to a unique sequence double-stranded DNA "tail," allow for quantification of rare target molecules in a mixture by PCR of the DNA tail (e.g., using quantitative real-time polymerase chain reaction).
[0386] Peptide aptamer selection can be performed using different systems, but the yeast two-hybrid system is currently the most widely used. Peptide aptamers can also be selected from combinatorial peptide libraries constructed by phage display and other surface display technologies such as mRNA display, ribosome display, bacterial display, and yeast display. These experimental procedures are also known as biopanning. Among the peptides obtained from biopanning, mimotopes can be considered a type of peptide aptamer. All peptides panned from combinatorial peptide libraries are stored in a special database called MimoDB.
[0387] Drugs In one embodiment, the heterologous moiety is a drug with undesirable pharmacokinetic or pharmacodynamic (PK / PD) parameters. Linking the heterologous moiety to a polypeptide can improve at least one PK / PD parameter of the heterologous moiety, such as targeting, absorption, and transport, or mitigate at least one undesirable PK / PD parameter, such as diffusion to off-target sites and toxic metabolism. For example, linking a polypeptide described herein to a drug with poor targeting / transport, e.g., a beta-lactam such as doxorubicin or penicillin, improves its specificity. In another example, linking a polypeptide described herein to a drug with poor absorption characteristics, e.g., insulin or human growth hormone, improves its minimum dose. In another example, linking a polypeptide described herein to a drug with toxic metabolic characteristics, e.g., high-dose acetaminophen, improves its maximum dose.
[0388] Membrane Translocation Polypeptides In one aspect, the composition comprises a polypeptide described herein that has properties that enable translocation across a membrane, e.g., independent of endosomes, such that the composition is delivered to a target location within a cell, e.g., a subject. In some embodiments, the targeting moiety comprises a membrane translocation polypeptide.
[0389] In one aspect, the disclosure includes a cell or tissue comprising any one of the membrane translocation polypeptides described herein.
[0390] In another aspect, the disclosure includes pharmaceutical compositions comprising the membrane translocation polypeptides described herein.
[0391] In another aspect, the disclosure includes methods of modulating gene expression by administering a composition comprising a membrane translocation polypeptide described herein.
[0392] In one aspect, the present disclosure includes a method of altering gene expression or altering anchor sequence-mediated connectivity using a membrane translocation polypeptide. In some embodiments, the membrane translocation polypeptide is a targeting moiety. In some embodiments, the membrane translocation polypeptide is a delivery agent that aids in the delivery of the targeting moiety described herein. The target location may be intracellular, for example, within the cytoplasm or an organelle (e.g., within the nucleus, such as targeting a DNA sequence or chromatin structure). The therapeutic compositions described herein may have additional advantageous properties, such as improved targeting, absorption, or transport, or reduced off-target activity, toxic metabolism, or toxic excretion.
[0393] In one embodiment, the composition comprises ABX nC (wherein A is selected from a hydrophobic amino acid or an amide-containing backbone, e.g., aminoethyl-glycine, having nucleic acid side chains; B and C may be the same or different and are independently selected from arginine, asparagine, glutamine, lysine, and analogs thereof; each X is independently a hydrophobic amino acid or each X is independently an amide-containing backbone, e.g., aminoethyl-glycine, having nucleic acid side chains; and n is an integer from 1 to 4).
[0394] Hydrophobic amino acids include amino acids having a hydrophobic side chain, including, but not limited to, alanine (ala, A), valine (val, V), isoleucine (iso, I), leucine (leu, L), methionine (met, M), phenylalanine (phe, F), tyrosine (tyr, Y), tryptophan (trp, W), and analogs thereof.
[0395] Amino acid analogs include, but are not limited to, D-amino acids, amino acids lacking a hydrogen on the alpha carbon such as dehydroalanine, metabolic intermediates such as ornithine and citrulline, non-alpha amino acids such as β-alanine, γ-aminobutyric acid, and 4-aminobenzoic acid, twin alpha carbon amino acids such as cystathionine, lanthionine, dienecolic acid, and diaminopimelic acid, and any others known in the art.
[0396] Nucleic acid side chains In one embodiment, the membrane transport polypeptide comprises one or more nucleic acid side chains linked to an amide backbone. Each amino acid unit in the polypeptide comprises an amide bond and its corresponding side chain. One or more amino acid units in the membrane transport polypeptide have an amide-containing backbone, such as aminoethyl-glycine, similar to a peptide backbone, with nucleic acid side chains replacing the amino acid side chains. Peptide nucleic acids (PNAs) are known to hybridize to complementary DNA and RNA with higher affinity than their oligonucleotide counterparts. This feature of PNAs not only allows the polypeptides of the present disclosure to form stable hybrids with nucleic acid side chains, but also provides hydrophobic units within the polypeptide due to their neutral backbone and hydrophobic side chains.
[0397] Nucleic acid side chains include, but are not limited to, purine or pyrimidine side chains, such as adenine, cytosine, guanine, thymine, and uracil. In one embodiment, the nucleic acid side chains comprise a nucleoside analog described herein.
[0398] size In some embodiments, the membrane translocation polypeptide has a size ranging from about 5 to about 500, e.g., 5 to 400, 5 to 300, 5 to 250, 5 to 200, 5 to 150, or 5 to 100 amino acid units in length. The polypeptide may have a length ranging from about 5 to about 50 amino acids, about 5 to about 40 amino acids, about 5 to about 30 amino acids, about 5 to about 25 amino acids, or any other range. In one embodiment, the polypeptide has a length of about 10 amino acids. In another embodiment, the polypeptide has a length of about 15 amino acids. In another embodiment, the polypeptide has a length of about 20 amino acids. In another embodiment, the polypeptide has a length of about 25 amino acids. In another embodiment, the polypeptide has a length of about 30 amino acids.
[0399] The membrane translocating polypeptide is an ABX polypeptide within the length range. n Each ABX may have more than one sequence of C. n C sequence by one or more amino acids to another ABXn In one embodiment, the polypeptide can be separated from the ABX C sequence. n In another embodiment, the polypeptide has at least two (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15 or more, e.g., 2-20, 2-10, 2-5) ABX C sequences, with the sequences separated by one or more amino acid units. n In another embodiment, the ABX C sequence is separated by one or more amino acid units. n The C sequences are separated by one (or more) hydrophobic amino acids, such as isoleucine or leucine.
[0400] The composition may contain multiple ABXs, which may be the same or different. n In one embodiment, at least two of the plurality are identical in sequence and / or length. In one embodiment, at least two of the plurality are different in sequence and / or length. In one embodiment, the composition comprises a plurality of ABXs, where at least two of the plurality are the same and at least two of the plurality are different. n In one embodiment, the transmembrane polypeptide comprises an ABX C sequence. n The C sequences are not identical in sequence or length or a combination thereof.
[0401] Protein or polypeptide production Methods for producing the therapeutic proteins or polypeptides described herein are routine in the art. See generally, Smales and James (eds.), Therapeutic Proteins: Methods and Protocols (Methods in Molecular Biology), Humana Press (2005); and Crommelin, Sindelar and Meibohm (eds.), Pharmaceutical Biotechnology: Fundamentals and Applications, Springer (2013).
[0402] The protein or polypeptide of the composition can be biochemically synthesized by using standard solid-phase techniques. Such methods include exclusive solid-phase synthesis, partial solid-phase synthesis, fragment condensation, and classical solution synthesis. These methods can be used when the peptide is relatively short (i.e., 10 kDa) and / or cannot be produced by recombinant technology (i.e., not encoded by a nucleic acid sequence) and therefore involves different chemical reactions.
[0403] Solid phase synthesis procedures are well known in the art and are further described by John Morrow Stewart and Janis Dillaha Young, Solid Phase Peptide Syntheses, 2nd Edition, Pierce Chemical Company, 1984; and Coin, I. et al., Nature Protocols, 2:3247-3256, 2007.
[0404] For longer peptides, recombinant methods can be used. Methods for producing recombinant therapeutic polypeptides are routine in the art. See generally, Smales and James (eds.), Therapeutic Proteins: Methods and Protocols (Methods in Molecular Biology), Humana Press (2005); and Crommelin, Sindelar and Meibohm (eds.), Pharmaceutical Biotechnology: Fundamentals and Applications, Springer (2013).
[0405] Exemplary methods for producing therapeutic pharmaceutical proteins or polypeptides include expression in mammalian cells, although recombinant proteins can also be produced using insect cells, yeast, bacteria, or other cells under the control of an appropriate promoter. Mammalian expression vectors may include non-transcribed elements such as an origin of replication, a suitable promoter, and other 5' or 3' flanking non-transcribed sequences, as well as essential ribosome binding sites, polyadenylation sites, splice donor and acceptor sites, and 5' or 3' non-translated sequences such as termination sequences. DNA sequences derived from the SV40 viral genome, such as SV40 origin, early promoter, splice, and polyadenylation sites, can be used to provide other genetic elements required for expression of heterologous DNA sequences. Suitable cloning and expression vectors for use with bacterial, fungal, yeast, and mammalian cell hosts are described in Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th ed.), Cold Spring Harbor Laboratory Press (2012).
[0406] If large quantities of a protein or polypeptide are desired, it can be produced using techniques such as those described by Brian Bray, Nature Reviews Drug Discovery, Vol. 2:587-593, 2003; and Weissbach and Weissbach, 1988, Methods for Plant Molecular Biology, Academic Press, NY, Part VIII, pp. 421-463.
[0407] Various mammalian cell culture systems can be used to express and produce recombinant proteins. Examples of mammalian expression systems include CHO cells, COS cells, HeLA, and BHK cell lines. The process of host cell culture for the production of protein therapeutics is described in Zhou and Kantardjieff (eds.), Mammalian Cell Cultures for Biologics Manufacturing (Advances in Biochemical Engineering / Biotechnology), Springer (2014). The compositions described herein may include a vector, such as a viral vector, e.g., a lentiviral vector, encoding a recombinant protein. The vector, e.g., a viral vector, contains a nucleic acid encoding the recombinant protein.
[0408] Purification of protein therapeutics is described in Franks, Protein Biotechnology: Isolation, Characterization, and Stabilization, Humana Press (2013); and Cutler, Protein Purification Protocols (Methods in Molecular Biology), Humana Press (2010).
[0409] Formulation of protein therapeutics is described in Meyer (ed.), Therapeutic Protein Drug Products: Practical Approaches to formulation in the Laboratory, Manufacturing, and the Clinic, Woodhead Publishing Series (2012).
[0410] Linker The proteins or polypeptides described herein may also comprise a linker. In some embodiments, for example, a protein described herein comprising a first polypeptide domain comprising a Cas or modified Cas protein and a second polypeptide domain comprising a polypeptide having DNA methyltransferase activity (or with demethylating or deaminase activity) has a linker between the first and second polypeptides. In one embodiment, one or more polypeptides described herein are linked using a linker. The linker may be a chemical bond, e.g., one or more covalent or non-covalent bonds. In some embodiments, the linker is a peptide linker (e.g., a non-ABX n C-peptide). Such linkers may be 2 to 30 amino acids, or longer. Linkers include flexible, rigid, or cleavable linkers as described herein.
[0411] The most commonly used flexible linkers have sequences consisting primarily of stretches of Gly and Ser residues ("GS" linkers). Flexible linkers are useful for connecting domains that require a certain degree of movement or interaction and may contain small, non-polar (e.g., Gly) or polar (e.g., Ser or Thr) amino acids. The incorporation of Ser or Thr can also maintain the stability of the linker in aqueous solution by forming hydrogen bonds with water molecules, thus reducing unfavorable interactions between the linker and the protein moiety.
[0412] Rigid linkers are useful for maintaining a fixed distance between domains and preserving their independent function. Rigid linkers can also be useful when spatial separation of domains is important to preserve the stability or biological activity of one or more components in the fusion. Rigid linkers can be alpha-helical structures or Pro-rich sequences, (XP). n (wherein X represents any amino acid, but is preferably Ala, Lys, or Glu).
[0413] A cleavable linker can release a free functional domain in vivo. In some embodiments, the linker can be cleaved under certain conditions, such as the presence of a reducing agent or a protease. An in vivo cleavable linker may utilize the reversible nature of a disulfide bond. One example includes a thrombin-sensitive sequence (e.g., PRS) between two Cys residues. In vitro thrombin treatment of CPRSC results in cleavage of the thrombin-sensitive sequence, while the reversible disulfide bond remains intact. Such linkers are known and are described, for example, in Chen et al., 2013, Fusion Protein Linkers: Property, Design and Functionality. Adv Drug Deliv Rev. 65(10):1357-1369. In vivo cleavage of the linker in the fusion can also be achieved by a protease expressed in specific cells or tissues or confined to a specific cellular compartment in vivo under a pathological condition (e.g., cancer or inflammation). The specificity of many proteases provides for slower cleavage of the linker in the constrained compartment.
[0414] Examples of linking molecules include hydrophobic linkers such as negatively charged sulfonate groups; lipids such as poly(-CH2-) hydrocarbon chains, such as polyethylene glycol (PEG) groups, unsaturated variants thereof, hydroxylated variants thereof, amidated or otherwise N-containing variants thereof, non-carbon linkers; carbohydrate linkers; phosphodiester linkers, or other molecules that can covalently link two or more polypeptides.Non-covalent linkers are also included, such as hydrophobic lipid globules, in which polypeptides are linked via their hydrophobic regions or hydrophobic extensions, such as a series of residues rich in leucine, isoleucine, valine, or possibly alanine, phenylalanine, or even tyrosine, methionine, glycine, or other hydrophobic residues.Polypeptides can be linked using charge-based chemical reactions, such as linking the positively charged portion of a polypeptide to the negative charge of another polypeptide or nucleic acid.
[0415] Polypeptide multimerization A composition may include multiple (two or more) membrane translocation polypeptides linked together, for example, via a linker as described herein.
[0416] The composition may comprise multiple membrane translocation polypeptides that are the same or different. In one embodiment, at least two of the multiple are identical in sequence and / or length. In one embodiment, at least two of the multiple are different in sequence and / or length. In one embodiment, the composition comprises multiple polypeptides, where at least two of the multiple are the same and at least two of the multiple are different. In one embodiment, the polypeptides in the composition are not identical in sequence or length, or a combination thereof.
[0417] The composition comprises a membrane translocating polypeptide linked to another membrane translocating polypeptide, e.g., by a linker. In some embodiments, the composition comprises two or more polypeptides linked by linkers. In some embodiments, the composition comprises three or more polypeptides linked by linkers. In some embodiments, the composition comprises four or more polypeptides linked by linkers. In some embodiments, the composition comprises five or more polypeptides linked by linkers. The linker may be a chemical bond, e.g., one or more covalent or non-covalent bonds, e.g., a flexible, rigid, or cleavable peptide linker. Such linkers may be 2 to 30 amino acids long, or longer. Additional linkers are described in more detail elsewhere herein and may also be applicable.
[0418] In one embodiment, two or more membrane translocating polypeptides are linked via peptide bonds, e.g., the carboxyl terminus of one polypeptide is linked to the amino terminus of another polypeptide. In another embodiment, one or more amino acids on one polypeptide are linked to one or more amino acids on another polypeptide, such as via a disulfide bond between cysteine side chains. In another embodiment, one or more amino acids on one polypeptide are linked to the carboxyl or amino terminus on another polypeptide, such as to create a branched polypeptide.
[0419] In another embodiment, one or more nucleic acid side chains on one transmembrane polypeptide interact with one or more amino acid side chains on another transmembrane polypeptide, such as through an arginine pseudopairing with a guanosine. In another embodiment, one or more nucleic acid side chains on one transmembrane polypeptide interact with one or more nucleic acid side chains on another transmembrane polypeptide, such as through hydrogen bonding. In another embodiment, multiple transmembrane polypeptides interact to create a particular sequence in the arrangement of nucleic acid side chains. For example, a carboxy-terminal nucleic acid side chain from one polypeptide interacts with an amino-terminal nucleic acid side chain from another polypeptide to create a pseudo-5' to pseudo-3' nucleotide sequence. In another example, a polypeptide is linked to one or more polypeptides, such as through amino acids and / or termini on each polypeptide, and their respective nucleic acid side chains align to create a pseudo-5' to pseudo-3' nucleotide sequence. The pseudosequence can bind to a selected target sequence, such as the anchor sequence of anchor sequence-mediated connection, for example, CTCF binding motif, cohesin binding motif, USF1 binding motif, YY1 binding motif, TATA box, ZNF143 binding motif, etc., or a transcription control sequence, for example, an enhancing or silencing sequence.The pseudosequence can bind to a selected target sequence, such as a transcription control sequence, for example, an enhancing or silencing sequence.The pseudosequence can interfere with the binding and transcription of factors by binding to the target sequence.The pseudosequence can hybridize with a nucleic acid sequence, such as mRNA, to interfere with gene expression.
[0420] In one embodiment, membrane translocation polypeptides are linked to one another, creating pseudo-5' to pseudo-3' nucleotide sequences that bind to anchor sequences recognized by nucleation proteins that bind with sufficient avidity to form anchor sequence-mediated connections, e.g., loops, or two-dimensional DNA structures generated by the physical interaction or association of one connected nucleation molecule-anchor sequence with another connected nucleation molecule-anchor sequence. Examples of anchor sequences include, but are not limited to, a CTCF binding motif, e.g., the CTCF binding motif or consensus sequence: N(T / C / G)N(G / A / T)CC(A / T / G)(C / G)(C / T / A)AG(G / A)(G / T)GG(C / A / T)(G / A)(C / G)(C / T / A)(G / A / C) (SEQ ID NO: 1), where N is any nucleotide. The linked polypeptides may create a pseudo-5' to pseudo-3' nucleotide sequence that binds to the CTCF binding motif or consensus sequence in the opposite orientation, e.g., (G / A / C)(C / T / A)(C / G)(G / A)(C / A / T)GG(G / T)(G / A)GA(C / T / A)(C / G)(A / T / G)CC(G / A / T)N(T / C / G)N (SEQ ID NO: 2).
[0421] The membrane translocation polypeptides described herein can be multimerized, eg, two or more polypeptides joined, using standard ligation techniques. Such methods include general native chemical ligation strategies (Siman, P. and Brik, A. Org. Biomol. Chem. 2012, 10:5684-5697; Kent, S.B.H. Chem. Soc. Rev. 2009, 38:338-351; and Hackenberger, C.P.R. and Schwarzer, D. Angew. Chem., Int. Ed. 2008, 47:10030-10074), click modification protocols (Tasdelen, M.A.; Yagci, Y. Angew. Chem., Int. Ed. 2013, 52:5930-5938; Palomo, J.M. Org. Biomol. Chem. 2012, 10:9309-9318; Eldijk, M.B.; van Hest, J.C.M. Angew. Chem., Int. Ed. 2011, 50:8806-8827; and Lallana, E.; Riguera, R.; Fernandez-Megia, E. Angew. Chem., Int. Ed. 2011, 50:8794-8804), and bioorthogonal reactions (King, M.; Wagner, A. Bioconjugate Chem. 2014, 25:825-839; Lang, K.; Chin, J.W. Chem. Rev. 2014, 114:4764-4806; Patterson, D.M.; Nazarova, L.A.; Prescher, J.A. ACS Chem. Biol. 2014, 9:592-605; Lang, K.; Chin, J.W. ACS Chem. Biol. 2014, Vol. 9: 16-20; Akaoka, Y.; Ojida, A.; Hamachi, I. Angew. Chem., Int. Ed. 2013, Vol. 52: 4088-4106; Debets, M.F.; van Hest, J.C.M.; Rutjes, FPJT Org. Biomol. Chem. 2013, 11:6439-6455; and Ramil, CP; Lin, Q. Chem. Commun. 2013, 49:11007-11022).
[0422] In some embodiments, the ordering of membrane-translocating polypeptides within a multimer can be specific, or it can be random, for example, if the polypeptides are not identical. For example, the polypeptides described herein can be multimerized by template-driven synthesis, or multimerization can be ordered by physical constraints or hybridization to a template, e.g., DNA, protein, or hybrid DNA-protein. In one embodiment, a template, e.g., a DNA sequence, specifically hybridizes to a polypeptide described herein. A polypeptide can be linked to another polypeptide by one of the methods described herein, e.g., conventional chemical ligation, and the selection of which polypeptide to link can be constrained by its ability to hybridize to the template. Thus, specific polypeptide multimers can be generated by their ability to specifically hybridize to the template.
[0423] In some embodiments, the order of membrane translocation polypeptides in multimer is determined by the chemical ligation strategy used.In one embodiment, chemical ligation techniques, such as click chemistry and bioorthogonal reaction, dictate which polypeptides are linked, because chemical ligation strategies require specific entities to react to allow ligation techniques to proceed.For example, one polypeptide can be labeled with phenyl azide, and another polypeptide can be labeled with cyclooctyne.Cyclooctyne and phenyl azide react to link the two polypeptides.
[0424] Hybridization In embodiments where the membrane translocation polypeptide comprises a nucleic acid side chain, it can interact with nucleic acids. In one embodiment, one or more nucleic acid side chains on the polypeptide hybridize with a nucleic acid sequence, for example, DNA such as genomic DNA, or RNA such as siRNA or mRNA molecules. One or more of the nucleic acid side chains on the polypeptide specifically hybridize with one or more nucleic acid residues in a target nucleic acid sequence. In one embodiment, the polypeptides are linked to each other, and the nucleic acid side chains hybridize with a nucleic acid sequence (e.g., a gene locus, mRNA, or an anchor sequence of an anchor sequence-mediated connection, such as a CTCF-binding motif, a cohesin-binding motif, a USF1-binding motif, a YY1-binding motif, a TATA box, a ZNF143-binding motif, etc.).
[0425] A nucleic acid side chain or pseudosequence can hybridize substantially identically to the nucleic acid side chain or pseudosequence, or hybridize to a target nucleic acid sequence that is 100%, 95%, 90%, 85%, 80%, 75%, or 70% complementary thereto. Hybridization of a nucleic acid side chain or pseudosequence with a target nucleic acid sequence can be carried out under suitable hybridization conditions routinely determined by optimization procedures. Conditions such as temperature, component concentrations, hybridization and washing times, buffer components, and their pH and ionic strength may vary depending on various factors, such as the length and GC content of the nucleic acid side chain or pseudosequence and the complementary target nucleic acid sequence. For example, when a relatively short length of nucleic acid side chain or pseudosequence is used, less stringent conditions can be employed. Detailed hybridization conditions can be found, for example, in *Molecular Cloning, A Laboratory Manual, 4th Edition* (Cold Spring Harbor Laboratory Press, 2012).
[0426] Heterologous moieties linked to polypeptides The composition may comprise a heterologous moiety described herein linked to a membrane translocation polypeptide of the targeting moiety, such as via a covalent or non-covalent bond or a linker described herein. In one embodiment, the composition comprises a heterologous moiety linked to the polypeptide via a peptide bond. For example, the amino terminus of the polypeptide is linked to the heterologous moiety, such as via a peptide bond using an optional linker. In another embodiment, the carboxyl terminus of the polypeptide is linked to the heterologous moiety.
[0427] In one embodiment, the composition comprises a membrane translocating polypeptide linked to two heterologous moieties, e.g., the amino and carboxyl termini of the polypeptide are linked to heterologous moieties, which may be the same or different heterologous moieties.
[0428] In another embodiment, one or more amino acids of the membrane translocation polypeptide are linked to a heterologous moiety, such as through a disulfide bond between cysteine side chains, a hydrogen bond, or any other known chemical reaction. One heterologous moiety can be a biologically active effector, and the other heterologous moiety can be a ligand or antibody for targeting the composition to specific cells that express a receptor. For example, a chemotherapeutic agent such as the topoisomerase inhibitor topotecan is linked to one end of the polypeptide, and a ligand or antibody is linked to the other end of the polypeptide to target the composition to specific cells or tissues. In another example, both heterologous moieties are biologically active effectors.
[0429] In another embodiment, multiple membrane translocating polypeptides, either the same or different membrane translocating polypeptides, are linked to a single heterologous moiety. The polypeptide may act as a coating that surrounds the larger heterologous moiety and assists it in penetrating the membrane. The heterologous moiety may have a molecular weight greater than about 500 grams / mole or Daltons, e.g., an organic or inorganic compound having a molecular weight greater than about 1,000 grams / mole, e.g., an organic or inorganic compound having a molecular weight greater than about 2,000 grams / mole, e.g., an organic or inorganic compound having a molecular weight greater than about 3,000 grams / mole, e.g., an organic or inorganic compound having a molecular weight greater than about 4,000 grams / mole, e.g., an organic or inorganic compound having a molecular weight greater than about 5,000 grams / mole, including salts, esters, and other pharmaceutically acceptable forms of such compounds.
[0430] In one embodiment, the composition comprises a membrane-translocating polypeptide linked to a heterologous moiety on one or both termini and another heterologous moiety linked to another site on the polypeptide. One or both of the amino and carboxyl termini of the polypeptide are linked to the heterologous moiety, and one or more amino acid units in the polypeptide, either amino acids or nucleic acids, are linked to one or more heterologous moieties via disulfide bonds, hydrogen bonds, or the like. For example, a DNA-modifying enzyme is linked to the polypeptide, and a nucleic acid having an unmethylated CTCF-binding motif complementary to a target methylated gene is hybridized to the nucleic acid side chain of the polypeptide. Upon administration, the composition targets the CTCF genome-binding motif and modulates gene transcription. In another example, a double-stranded nucleic acid having an unmethylated CTCF-binding motif with gene-specific flanking sequences is linked to the polypeptide. Upon administration, the unmethylated CTCF-binding motif serves as an alternative anchor sequence for CTCF protein binding. In another example, another heterologous moiety, such as ubiquitin or an effector, is linked to the polypeptide. Upon administration, the composition penetrates the cell membrane, and the effector performs its function. Ubiquitin then targets the composition for degradation.
[0431] In one embodiment, the composition comprises a membrane permeabilization polypeptide that is covalently linked to one or more heterologous moieties, and another heterologous moiety that is linked to the nucleic acid in the polypeptide.For example, a protein synthesis inhibitor is covalently linked to the polypeptide, and siRNA or other target-specific nucleic acid is hybridized to the nucleic acid in the polypeptide.When administered, siRNA targets the composition to mRNA transcript, and the protein synthesis inhibitor and siRNA act to inhibit the expression of mRNA.
[0432] In some embodiments, the pharmaceutical composition comprises a compound of Structure I: (II)XYZ where X and Z are 5' and 3' site-specific targeting sequences for the target CTCF binding motif, respectively, and Y is (a) an RNA sequence complementary to the sequence of SEQ ID NO: 1; (b) an RNA sequence that is at least 75%, 80%, 85%, 90%, or 95% identical to an RNA sequence complementary to the sequence of SEQ ID NO: 1; (c) an RNA sequence complementary to the sequence of SEQ ID NO: 1 with at least 1, 2, 3, 4, 5, but less than 15, 12, or 10 nucleotide additions, substitutions, or deletions; (d) an RNA sequence complementary to the sequence of SEQ ID NO:2; (e) an RNA sequence that is at least 75%, 80%, 85%, 90%, or 95% identical to an RNA sequence complementary to the sequence of SEQ ID NO:2; (f) an RNA sequence complementary to the sequence of SEQ ID NO: 2, which has at least 1, 2, 3, 4, 5, but less than 15, 12, or 10 nucleotide additions, substitutions, or deletions; (selected from The gRNA comprises a transmembrane polypeptide linked to a gRNA comprising the sequence:
[0433] In some embodiments, X and Z are each 2 to 50 nucleotides in length, eg, 2 to 20, 2 to 10, 2 to 5 nucleotides in length.
[0434] In some embodiments, the gRNA comprises a specific targeting sequence for a CTCF binding motif associated with an oncogene, a tumor suppressor, or a disease associated with a nucleotide repeat.
[0435] The membrane translocation polypeptides described herein can be linked to heterologous moieties using standard ligation techniques, such as those described herein for linking polypeptides.
[0436] To introduce small or single point mutations, a homologous recombination (HR) template can be ligated to the membrane translocation polypeptide. In one embodiment, the HR template is a single-stranded DNA (ssDNA) oligo or a plasmid. For ssDNA oligo design, approximately 100-150 bp of overall homology can be used, with the mutation introduced approximately in the center, resulting in 50-75 bp homology arms.
[0437] In some embodiments, a target anchor sequence, e.g., a gRNA or antisense DNA oligonucleotide for targeting a CTCF binding motif, is (a) a nucleotide sequence comprising SEQ ID NO:1; (b) a nucleotide sequence that is at least 75%, 80%, 85%, 90%, or 95% identical to SEQ ID NO:1; (c) a nucleotide sequence comprising SEQ ID NO: 1 with at least 1, 2, 3, 4, 5, but less than 15, 12, or 10 nucleotide additions, substitutions, or deletions; (d) a nucleotide sequence comprising SEQ ID NO:2; (e) a nucleotide sequence that is at least 75%, 80%, 85%, 90%, or 95% identical to SEQ ID NO: 2; a nucleotide sequence that contains SEQ ID NO: 2 with at least 1, 2, 3, 4, or 5, but less than 15, 12, or 10, nucleotide additions, substitutions, or deletions. The polypeptide is linked to a transmembrane polypeptide along with an HR template selected from the group consisting of:
[0438] Any of the linkers described herein can be included to covalently or non-covalently link the membrane translocation polypeptide and the heterologous moiety. The linker can be used, for example, to separate the polypeptide from the heterologous moiety. For example, a linker can be placed between the polypeptide and the heterologous moiety to provide molecular flexibility, for example, in secondary and tertiary structure. In one embodiment, the linker contains at least one glycine, alanine, and serine amino acid to provide flexibility. In another embodiment, the linker is a hydrophobic linker, such as a negatively charged sulfonate group, a polyethylene glycol (PEG) group, or a pyrophosphate diester group. In another embodiment, the linker is cleavable, selectively releasing the heterologous moiety from the polypeptide, yet sufficiently stable to prevent premature cleavage.
[0439] Post-administration binding In some embodiments, the membrane translocation polypeptides described herein are capable of forming bonds, e.g., after administration, with other polypeptides, heterologous moieties described herein, e.g., effector molecules, e.g., nucleic acids, proteins, peptides, or other molecules, or other agents, e.g., intracellular molecules, via covalent or non-covalent bonds. In one embodiment, one or more amino acids on the polypeptide can be linked to a nucleic acid, e.g., via an arginine pseudopairing with guanosine, an internucleotide phosphate bond, or an interpolymer bond. In some embodiments, the nucleic acid is DNA, e.g., genomic DNA, or RNA, e.g., a tRNA or mRNA molecule. In another embodiment, one or more amino acids on the polypeptide can be linked to a protein or peptide.
[0440] fusion molecule In some embodiments, the composition comprises a fusion molecule, such as a fusion molecule comprising a peptide or polypeptide. Those skilled in the art reading this specification will understand that the term "protein fusion" can also refer to a fusion molecule comprising a "protein" (or peptide or polypeptide) component. In some embodiments, the protein fusion comprises one or more of the moieties described herein, such as a nucleic acid sequence, a peptide or protein moiety, a membrane translocation polypeptide, a targeting peptide / aptamer, or other heterologous moiety described herein.
[0441] In one aspect, the disclosure includes a cell or tissue comprising any one of the protein fusions described herein.
[0442] In another aspect, the disclosure includes pharmaceutical compositions comprising the protein fusions described herein.
[0443] In another aspect, the present disclosure includes a method of modulating gene expression by administering a composition comprising a protein fusion described herein. For example, the protein fusion may be dCas9-DNMT, dCas9-DNMT-3a-3L, dCas9-DNMT-3a-3a, dCas9-DNMT-3a-3L-3a, dCas9-DNMT-3a-3L-KRAB, dCas9-KRAB, dCas9-APOBEC, APOBEC-dCas9, dCas9-APOBEC-UGI, dCas9-UGI, UGI-dCas9-APOBEC, UGI-APOBEC-dCas9, any variation of the protein fusions described herein, or other fusions of proteins or protein domains described herein.
[0444] Exemplary dCas9 fusion methods and compositions that are applicable to the methods and compositions described herein are known and are described, for example, in Kearns et al., "Functional annotation of native enhancers with a Cas9-histone demethylase fusion." Nature Methods 12:401-403 (2015); and McDonald et al., "Reprogrammable CRISPR / Cas9-based system for inducing site-specific DNA methylation." Biology Open 2016: doi:10.1242 / bio.019067. Using methods known in the art, dCas9 can be fused to any of the various agents and / or molecules described herein; the resulting fusion molecules can be useful in the various disclosed methods.
[0445] In one aspect, the disclosure includes a composition comprising a protein comprising a domain that acts on DNA, e.g., an enzymatic domain (e.g., a nuclease domain, e.g., a Cas9 domain, e.g., a dCas9 domain; a DNA methyltransferase, a demethylase, a deaminase), in combination with at least one guide RNA (gRNA...
Claims
**Claim 1** A site-specific disruptor that is a fusion molecule, comprising: (i) A DNA-binding moiety comprising a Cas polypeptide that binds specifically to one or more target anchor sequences in the cell with sufficient affinity to compete with the binding of CTCF (CCCTC-binding factor) in the cell and does not bind to non-target anchor sequences in the cell, wherein the Cas polypeptide is directed to the one or more target anchor sequences by a site-specific guide RNA, the one or more target anchor sequences comprise a CTCF-binding motif comprising the sequence of SEQ ID NO: 1 or SEQ ID NO: 2, and the DNA-binding moiety is catalytically inactive; and (ii) A heterologous moiety comprising a deaminating agent or an epigenetic modifier A site-specific disruptor comprising the above. **Claim 2** The site-specific disruptor according to claim 1, wherein the binding of the site-specific disruptor to the one or more target anchor sequences physically interferes with the formation and / or maintenance of anchor sequence-mediated connections. **Claim 3** The site-specific disruptor according to claim 1, wherein the site-specific guide RNA comprises a complement of the target anchor sequence or has a sequence that is at least 90% complementary to the target anchor sequence. **Claim 4** (i) Within an anchor sequence-mediated connection comprising a first anchor sequence and a second anchor sequence, or (ii) Within 10 kb of the first anchor sequence within an anchor sequence-mediated connection comprising a first anchor sequence and a second anchor sequence, A composition for use in a method of modulating gene expression, the composition comprising the site-specific disruptor according to any one of claims 1 to 3, The method comprising the step of contacting the first and / or second anchor sequence with the site-specific disruptor. A composition characterized by this. **Claim 5** The composition according to claim 4, wherein the anchor sequence-mediated connection comprises at least one internal transcriptional control sequence. **Claim 6** (i) The at least one internal transcriptional control sequence is an enhancing sequence or a silencing or suppressing sequence; (ii) The gene is at least 300 base pairs away from the internal transcriptional control sequence, and / or (iii) the first and / or the second anchor sequence is located within 500 kb of an exogenous transcriptional control sequence (e.g., an enhancing sequence or a silencing or repressing sequence), The composition according to claim 5.
7. The site-specific disruptor according to any one of claims 1 to 3 or the composition according to any one of claims 4 to 6, wherein the heterologous moiety comprises an epigenetic modifier.
8. (i) the deaminating agent is a deaminase; or, (ii) the epigenetic modifier is a DNA methylase, a DNA demethylase, a histone methyltransferase, a histone deacetylase, or a combination thereof, The site-specific disruptor according to any one of claims 1 to 3 or 7 or the composition according to any one of claims 4 to 7.
9. The site-specific disruptor according to any one of claims 1 to 3, 7 or 8 or the composition according to any one of claims 4 to 8, wherein the Cas polypeptide is a Cas9 polypeptide.
10. The composition according to any one of claims 4 to 9, wherein the step of contacting comprises delivering the site-specific disruptor to mammalian cells in vitro.
11. A pharmaceutical preparation for use in modifying the expression of a target gene in a subject in need thereof, the preparation comprising the site-specific disruptor according to any one of claims 1 to 3 or 7 to 9.
12. The pharmaceutical preparation for use according to claim 11, wherein the subject is a human subject and / or the target gene is a gene associated with a disease or condition.
13. The site-specific disruptor according to any one of claims 1 to 3 or 7 to 9, the composition according to any one of claims 4 to 10, or the pharmaceutical preparation for use according to claim 11 or 12, wherein the site-specific disruptor introduces a targeted change to the one or more target anchor sequences and / or the targeted change comprises at least one of a substitution, an addition or a deletion.
14. The pharmaceutical preparation for use according to any one of claims 11-13, the site-specific disruptor according to any one of claims 1-3, 7-9 or 13, or the composition according to any one of claims 4-10 or 13, wherein neither the first target anchor sequence nor the second target anchor sequence is located in either an enhancer or a promoter.
15. The pharmaceutical preparation for use according to any one of claims 13 or 14 when directly or indirectly citing claim 11 or 12, or claims 11 or 12, wherein the target gene is a cancer gene, a tumor suppressor, or a gene associated with nucleotide repeats.
16. The pharmaceutical preparation for use according to any one of claims 13 or 14 when directly or indirectly citing claim 11 or 12, or claims 11, 12 or 15, wherein the target gene is FOXF3.
17. The site-specific disruptor according to any one of claims 1-3, 7-9, 13 or 14, the composition according to any one of claims 4-10, 13 or 14, or the pharmaceutical preparation for use according to any one of claims 11-16, wherein the site-specific guide RNA comprises any one of the sequences of SEQ ID NOs: 40-47, 86 or 87.
Citation Information
Patent Citations
Massively parallel combinatorial genetics for crispr
WO2016070037A2