A method for determining genome-wide DNA-binding protein binding sites by footprinting with double-stranded DNA deaminase
The use of double-stranded DNA deaminase to create footprints on polynucleotides addresses the limitations of existing methods by enabling high-resolution, genome-wide mapping of transcription factor binding sites and interactions, enhancing our understanding of gene regulation and cellular functions.
Patent Information
- Application Number
- JP2025518887
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2022-09-30
- Publication Date
- 2025-10-15
AI Technical Summary
Current methods for detecting DNA-protein interactions at a genome-wide level, such as ChIP-seq, suffer from high signal-to-noise ratios, low throughput, low resolution, and require cell number and homogeneity, and lack the ability to identify interacting transcription factors during gene regulation.
The use of double-stranded DNA deaminase to convert cytosine to uracil on double-stranded polynucleotides, creating a 'footprint' where the DNA-binding protein is bound, allowing for the identification of transcription factor binding sites and their interactions through cytosine-to-uracil conversion patterns.
Enables high-resolution, genome-wide mapping of transcription factor binding sites and their interactions, providing insights into gene regulation and cellular functions.
Smart Images

Figure 2025534400000001 
Figure 2025534400000002 
Figure 2025534400000003
Abstract
Description
[Technical Field]
[0001] The present disclosure relates generally to systems, methods, and compositions for determining DNA-binding protein binding sites along the genome of one or more cells. [Background technology]
[0002] Although each cell in an individual has essentially the same genome, it performs entirely different functions in each tissue. The advent of single-cell genomics has enabled the determination of the transcriptome, methylome, and open chromosome sites of single human cells, thereby enabling unprecedented cell type classification. However, beyond cell type determination, deciphering the human functional genome—that is, understanding cellular functions based on the human genome—is a pressing challenge. Processes such as gene expression and regulation, cell differentiation, and development are linked to chromatin structure and regulatory networks, for which transcription factors (TFs) play a crucial role.
[0003] Humans have only about 1,000 TFs, which together regulate approximately 20,000 genes. The specificity of gene regulation is achieved through the combinatorial binding of several TFs, which act like a set of keys to turn specific genes on and off. Therefore, it is important to learn the precise binding set of TFs and how they interact with each other. Current methods for detecting DNA-protein interactions at a genome-wide level, such as ChIP-seq, have several problems, including high signal-to-noise ratios, low throughput, low resolution, and requirements for cell number and homogeneity. Importantly, existing methods for detecting DNA-protein interactions at a genome-wide level lack the ability to identify TFs that interact with each other during gene regulation.
[0004] Sequence Listing This application contains a Sequence Listing that has been submitted electronically in XML format, and is incorporated herein by reference in its entirety. The XML copy, created on September 28, 2022, is titled "009191.00002_st26" and is 23KB in size. Summary of the Invention
[0005] Generally, aspects of the present disclosure relate to methods for identifying or profiling DNA-binding proteins, such as transcription factors, on double-stranded polynucleotides, such as genomic DNA. Generally, the disclosed methods use double-stranded (ds) DNA deaminase to convert cytosine to uracil on double-stranded polynucleotides, except when the DNA-binding protein is bound to the polynucleotide. The dsDNA deaminase sterically prevents the conversion of cytosine to uracil at the position where the DNA-binding protein is bound to the polynucleotide, thereby creating a "footprint" where cytosine was not converted to uracil. Thus, the position where the DNA-binding protein is bound to the polynucleotide can be determined. The determined DNA-binding site can be compared with known DNA-binding sites of the DNA-binding protein to identify the DNA-binding protein bound to the determined DNA-binding site. Using the methods described herein, one or more, or multiple, DNA-binding proteins can be identified for a given polynucleotide, such as a gene within chromatin DNA.
[0006] Aspects of the present disclosure relate to methods for identifying binding sites of one or more transcription factors (TFs) to DNA, such as chromatin DNA. Identification of the binding sites of transcription factors can be used to identify the transcription factors themselves based on their known binding sites with chromatin DNA, and thus identify one or more, or pairs of, transcription factors that cooperate to regulate genes. According to one aspect, combinations or "keysets" of transcription factors are decoded for specific genes and along the genome, thereby identifying genome-wide combinations or keysets of transcription factors.
[0007] According to one embodiment, a method is provided that includes contacting chromatin DNA with a dsDNA deaminase. The dsDNA deaminase converts cytosine to uracil along the chromatin DNA unless a TF is bound to the chromatin DNA. When a TF is bound to the chromatin DNA, the dsDNA deaminase sterically prevents the cytosine from being converted to uracil at the binding site between the TF and the chromatin DNA. Therefore, cytosine to uracil conversion occurs on either side of the binding site between the TF and the chromatin DNA. Based on the cytosine to uracil conversion, the boundary at which the TF binds to the chromatin DNA can be determined, thus determining the "footprint" of the binding site. The binding site is then compared with known binding sites of the TF to identify TFs with matching binding sites. Target chromatin DNA, such as a gene, can be analyzed to determine whether one or more, or a pair of, or multiple TFs bind to the target chromatin DNA, thereby identifying TFs involved in the regulation of a specific gene. According to one embodiment, this approach to identifying TF keysets can be performed genome-wide across all genes.
[0008] Generally, the embodiments of the present disclosure include cell permeabilization or nucleus permeabilization, and the isolation of cell, single cell, or cell population.In this way, the genomic DNA of cell, single cell, or cell population is more accessible to dsDNA deaminase.According to one embodiment, one or more cells or one or more nuclei do not need to be permeabilized, and can still be treated by dsDNA deaminase.Other methods are contemplated, such as lysis, to process one or more cells or one or more nuclei according to known methods, making one or more cells or one or more nuclei more accessible to dsDNA deaminase treatment.
[0009] In one embodiment, the permeabilized cell or cells or the permeabilized nuclei or nuclei can be treated with a cross-linking agent, as known in the art, to cross-link cellular components to maintain cellular structure before being treated with double-stranded DNA deaminase.In one embodiment, in the case of cell-free DNA, for example, double-stranded polynucleotide molecules to which DNA-binding proteins are bound can be treated with dsDNA deaminase.
[0010] According to one embodiment, permeabilized one or more cells or permeabilized one or more nuclei or cell-free DNA is treated with double-stranded DNA deaminase, so that the cytosine in the DNA of one or more cells or one or more nuclei is converted to uracil by hydrolysis of amino groups from the cytosine nucleotides available for deamination to generate uracil nucleotides.According to one embodiment, DNA can be fragmented and concentrated before treatment with dsDNA deaminase.According to one embodiment, DNA can be treated with dsDNA deaminase, and then fragmented and concentrated.The treatment with dsDNA deaminase results in treated DNA, as long as the treated DNA contains one or more uracils resulting from the treatment with dsDNA deaminase. Exemplary double-stranded DNA deaminases include double-stranded DNA deaminase A ("DddA") known in the art, evolved double-stranded DNA deaminase A11 ("DddA11") known in the art, and bacterial deaminase toxin family 3 ("BadTF3") known in the art. Processed DNA, e.g., processed chromatin DNA, may be processed using a transposase or DNase to enrich for open chromatin DNA, i.e., transcriptionally active genomic DNA that is accessible by DNA regulatory elements.
[0011] The processed DNA can be amplified before sequencing.Exemplary amplification methods include PCR.Alternatively, the processed DNA for sequencing can be directly sequenced without amplification after library preparation.
[0012] The processed DNA can be sequenced.According to one embodiment, the processed DNA can be sequenced as whole genome DNA.According to one embodiment, the processed DNA can be sequenced as accessible region of chromatin by cutting with enzyme such as transposase or nuclease such as DNase, MNase or restriction endonuclease.According to one embodiment, the enriched open chromatin DNA is sequenced to determine cytosine to uracil conversion, and thus determine DNA binding protein footprint.
[0013] According to one embodiment, target chromatin regions, such as open chromatin regions, can be enriched before or after treatment with dsDNA deaminase using methods known to those skilled in the art, including ChIP-seq, ChIC, ChEC, ChEC-seq, CUT&TAG, CUT&RUN, Multi-CUT&Tag, NTT-seq, and R-loop CUT&Tag, in addition to methods using transposases such as Tn5 transposase.For example, dsDNA is treated with dsDNA deaminase, and then the treated DNA is enriched for open chromatin regions using CUT&TAG for sequencing.Alternatively, dsDNA is enriched for open chromatin regions using CUT&TAG for sequencing, and then the treated DNA is treated with dsDNA deaminase.
[0014] In one embodiment, the library can be prepared based on total genomic DNA as known in the art. In one embodiment, the library can be prepared based on a target DNA region, such as open chromatin. In one embodiment, the target region is enriched by using antibodies in methods such as ChIP-seq, ChIC, ChEC, ChEC-seq, CUT&TAG, and CUT&RUN. According to one embodiment, target regions are enriched using other binding agents such as nanobodies (see Stuart et al., Nanobody-tethered transposition allows for multifactorial chromatin profiling at single-cell resolution, bioRxiv 10.1101 / 2022.03.08.483436v1, which is incorporated herein by reference in its entirety) or specific chromatin binding domains (see Wang et al., Genomic profiling of native R loops with a DNA-RNA hybrid recognition sensor, Sci. Adv. 2021 Feb;7(8)eabe3516 10.1126 / sciadv.abe3516, which is incorporated herein by reference in its entirety).
[0015] Then, DNA-binding proteins can be profiled by analyzing the information obtained from the cytosine-to-uracil conversion sites on the sequence reads. Non-converted sites (i.e., cytosine remains cytosine even after treatment with dsDNA deaminase) indicate binding sites to which DNA-binding proteins were bound when treated with dsDNA deaminase. Converted sites (i.e., cytosine is converted to uracil by dsDNA deaminase) indicate sites to which DNA-binding proteins were not bound when treated with dsDNA deaminase. DNA-binding protein binding profiles can be compared with one or more non-converted sites using a database of binding sites associated with DNA-binding proteins, such as the TF motif database JASPAR (worldwide website: jaspar.genereg.net), CIS-BP (worldwide website: cisbp.ccbr.utoronto.ca), and HOCOMOCO (worldwide website: hocomoco11.autosome.org), using methods known to those skilled in the art, to identify specific DNA-binding proteins. Potential binding TFs are identified by comparing footprints determined by the dsDNA deaminase method described herein with known TF motifs in these databases.
[0016] Aspects of the present disclosure can be practiced at the single cell level or single DNA molecule level, or with multiple cells. The multiple cells can be of the same cell type. The multiple cells can be of different cell types.
[0017] According to the present disclosure, a method is provided for quantifying the simultaneous binding of multiple TFs on a single DNA molecule. Such a method provides a high-resolution binding map of multiple TFs at a genome-wide level. According to the present disclosure, a method is provided for analyzing how TFs cooperate or antagonize during transcription regulation.
[0018] According to one aspect, there is provided a use of dsDNA deaminase in the manufacture of an agent for carrying out a method for determining transcription factor binding sites on genomic double-stranded (ds) DNA of a eukaryotic cell. In certain embodiments, the method is as described herein. In certain embodiments, the method comprises: contacting the genomic dsDNA with a dsDNA deaminase under conditions that convert cytosines in the genomic dsDNA to uracils, thereby producing treated genomic dsDNA; and Identifying unconverted cytosines on processed genomic dsDNA as transcription factor binding sites Includes.
[0019] The patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawing(s) will be provided by the Patent Office upon request and payment of the necessary fee. The above and other features and advantages of the present invention will be more fully understood from the following detailed description of illustrative embodiments taken in conjunction with the accompanying drawings. [Brief explanation of the drawings]
[0020] [Figure 1] FIG. 1 is a schematic diagram illustrating the determination of TF binding sites using double-stranded DNA deaminase. [Figure 2] FIG. 1 shows a vector map of pETDuet-1::dddAtox+dddAI. [Figure 3] FIG. 1 shows a vector map of pETDuet-1::badTF3tox+badTF3I. [Figure 4] FIG. 1 shows a vector map of pETDuet-1::dddA11+dddAI. [Figure 5] Coomassie blue-stained SDS-PAGE gels of DddAtox-DddAI, DddAtox, DddA11-DddAI, DddA11, BadTF3-BadTF3I, and BadTF3, respectively. [Figure 6]FIG. 1 shows data supporting the conversion rates of cytosine to uracil using several double-stranded DNA deaminases. [Figure 7A] FIG. 1 shows data identifying TF footprints. [Figure 7B] FIG. 1 shows data identifying TF footprints. [Figure 7C] FIG. 1 shows a Venn diagram illustrating the overlap between CTCF binding identified by the dsDNA deaminase, ChIP-seq, and DNase-seq methods described herein. [Figure 8A] FIG. 1 is a schematic diagram illustrating TF binding patterns at the single molecule level. [Figure 8B] FIG. 1 shows data demonstrating the identification and quantification of binding of the transcription factor CTCF to DNA. [Figure 8C] FIG. 1 shows data demonstrating the use of the dsDNA deaminase method described herein to identify multiple transcription factor binding sites. [Figure 9A] FIG. 1 is a schematic diagram illustrating one aspect of the method of the present disclosure. [Figure 9B] FIG. 1 shows DNA fragment distribution data generated by the methods described herein. [Figure 9C] FIG. 1 shows cell type classification results for single-cell data generated by the methods described herein for K562, GM12878, and HEK293T cell lines. [Figure 9D] FIG. 1 shows a comparison of bulk and single cell data generated by the methods described herein, as displayed in IGV software. DETAILED DESCRIPTION OF THE INVENTION
[0021] The practice of a particular embodiment or features of a particular embodiment may employ, unless otherwise indicated, conventional techniques of molecular biology, microbiology, recombinant DNA, and the like, which are within the ordinary skill in the art and are fully described in the literature. See, for example, Sambrook, Fritsch, and Maniatis, MOLECULAR CLONING: A LABORATORY MANUAL, 2nd ed. (1989), OLIGONUCLEOTIDE SYNTHESIS (M.J. Gait, ed., 1984), ANIMAL CELL CULTURE (R.I. Freshney, ed., 1987), the series METHODS IN ENZYMOLOGY (Academic Press, Inc.); GENE TRANSFER VECTORS FOR MAMMALIAN CELLS (J.M. Miller and M.P. Calos, eds., 1987), HANDBOOK OF EXPERIMENTAL IMMUNOLOGY (D.M. Weir and C.C. Blackwell, eds.), CURRENT PROTOCOLS IN MOLECULAR See journal research articles such as BIOLOGY (F.M.A.usubel, R. Brent, R.E. Kingston, D.D. Moore, J.G. Siedman, J.A. Smith, and K. Struhl, eds., 1987); CURRENT PROTOCOLS IN IMMUNOLOGY (J.E. Coligan, A.M. Kruisbeek, D.H. Margulies, E.M. Shevach, and W. Strober, eds., 1991); ANNUAL REVIEW OF IMMUNOLOGY; and ADVANCES IN IMMUNOLOGY. All patents, patent applications, and publications mentioned herein, both above and below, are hereby incorporated by reference.
[0022] Terms and symbols in nucleic acid chemistry, biochemistry, genetics, and molecular biology will be used in accordance with standard treatises and texts in the field, such as Kornberg and Baker, DNA Replication, 2nd ed. (W.H. Freeman, New York, 1992); Lehninger, Biochemistry, 2nd ed. (Worth Publishers, New York, 1975); Strachan and Read, Human Molecular Genetics, 2nd ed. (Wiley-Liss, New York, 1999); Eckstein, ed., Oligonucleotides and Analogs: A Practical Approach (Oxford University Press, New York, 1991); and Gait, ed., Oligonucleotide Synthesis: A Practical Approach (IRL Press, Oxford, 1984).
[0023] An embodiment of the present disclosure is described with reference to FIG. 1. As shown in FIG. 1, nuclei are extracted from cells, cell lines, or tissues and permeabilized. The nuclei contain chromatin DNA to which DNA-binding proteins such as TFs are bound. The isolated permeabilized nuclei are incubated, i.e., treated, with a dsDNA cytosine deaminase, which converts accessible cytosines on dsDNA to uracils to generate processed DNA, such as processed chromatin DNA. An accessible cytosine is a cytosine that is not sterically hindered or sterically unhindered from enzymatic reaction with the dsDNA cytosine deaminase. An inaccessible cytosine is a sterically hindered cytosine that is bound by a DNA-binding protein such as TF and is therefore inaccessible or otherwise unavailable for enzymatic reaction with the dsDNA cytosine deaminase. Regions of chromatin DNA not bound by DNA-binding proteins, such as regions of chromatin DNA immediately adjacent to and flanking a DNA-binding protein ("flanking regions"), are accessible to dsDNA deaminases, which convert cytosine to uracil in regions of chromatin DNA not bound by DNA-binding proteins. Regions of chromatin DNA bound by one or more DNA-binding proteins, such as transcription factors or other DNA-binding proteins, are protected from cytosine to uracil conversion by dsDNA deaminases. As a result, converted regions are adjacent to or otherwise flanking non-converted regions, which correspond to binding sites for DNA-binding proteins. Such DNA-binding protein binding sites are referred to as DNA-binding protein "footprints." DNA-binding protein binding sites may be a single nucleotide, or may be one or more nucleotides, two or more nucleotides, three or more nucleotides, four or more nucleotides, 1-4 or 1-5 nucleotides, or may extend across a variety of nucleotides or may otherwise be of various nucleotide lengths. The DNA binding protein binding site may be a single nucleotide or may be multiple nucleotides.DNA-binding proteins can interact with one or more nucleotides of chromatin DNA. The nucleotides that interact with DNA-binding proteins can define DNA-protein binding sites. According to one embodiment, processed chromatin DNA is extracted for whole genome or targeted amplicon sequencing. The method described herein allows for simultaneous quantification of multiple DNA-binding protein binding events (e.g., TF pairs, or TF binding events, whether more than two TFs, three TFs, four TFs, five TFs, etc.) on genes of a single DNA molecule. The method described herein relates to systematically quantifying the co-occupancy frequency of multiple TFs or TF pairs, for example, thousands of TFs or TF pairs, across multiple genes throughout the genome.
[0024] cell The cells according to the present invention include any cell for which it is deemed useful by those skilled in the art to understand the binding sites of DNA-binding proteins in double-stranded DNA, such as chromatin DNA. The cells include prokaryotic and eukaryotic cells. The cells according to the present disclosure include all types of cancer cells, liver cells, oocytes, embryos, stem cells, iPS cells, ES cells, neurons, erythrocytes, melanocytes, astrocytes, germ cells, oligodendrocytes, and kidney cells.
[0025] Cells useful for the methods described herein can be obtained from a biological sample, tissue of interest, or biopsy, blood sample, or cell culture. Additionally, cells can be obtained from specific organs, tissues, tumors, neoplasms, etc., and used in the methods described herein. Furthermore, cells from any population, such as a population of prokaryotic or eukaryotic single-cell organisms, including bacteria or yeast, can generally be used in the methods. In one embodiment, the sample can be in vitro. The term "in vitro" has its art-recognized meaning, including, for example, purified reagents or extracts, e.g., cellular extracts. As used herein, the term "biological sample" is intended to include, but is not limited to, tissues, cells, biological fluids, and isolates thereof, isolated from a subject, as well as tissues, cells, and fluids present within a subject.
[0026] According to one embodiment, the method of the present invention is carried out using single cells. As used herein, "single cells" refers to one cell. Single cell suspensions can be obtained using standard methods known in the art, including, for example, enzymatically digesting proteins connecting cells in a tissue sample using trypsin or papain, or releasing adherent cells in culture, or mechanically separating cells in a sample. Single cells can be placed in any suitable reaction vessel that can be individually processed. For example, in the case of a 96-well plate, each single cell is placed in a single well.
[0027] Methods for manipulating single cells are known in the art and include fluorescence-activated cell sorting (FACS), flow cytometry (Herzenberg, PNAS USA 76:1453-55, 1979), micromanipulation, and the use of semi-automated cell pickers (e.g., the QUIXELL cell migration system manufactured by Stoelting Co.). For example, individual cells can be individually selected based on characteristics detectable by microscopy, such as location, morphology, or reporter gene expression. In addition, a combination of gradient centrifugation and flow cytometry can be used to increase isolation or sorting efficiency.
[0028] According to one embodiment, the method of the present invention is carried out using a plurality of cells, including about 2 to about 1,000,000 cells, about 2 to about 10 cells, about 2 to about 100 cells, about 2 to about 1,000 cells, about 2 to about 10,000 cells, about 2 to about 100,000 cells, about 2 to about 10 cells, or about 2 to about 5 cells.
[0029] According to one embodiment, the DNA to be processed is genomic DNA or chromatin DNA. According to one embodiment, the DNA to be processed is mammalian DNA, plant DNA, yeast DNA, viral DNA, or prokaryotic DNA. In yet another preferred embodiment, the DNA sample is obtained from humans, cows, pigs, sheep, horses, rodents, birds, fish, shrimp, plants, yeast, viruses, or bacteria. Preferably, the DNA to be processed is genomic DNA. As used herein, the term "genome" is defined as the complete set of genes possessed by an individual, cell, or organelle. As used herein, the term "genomic DNA" is defined as DNA material containing a partial or complete complete set of genes possessed by an individual, cell, or organelle.
[0030] In one embodiment, the DNA to be processed is a double-stranded polynucleotide molecule with protein binding, such as cell-free DNA. According to this embodiment, DNA, such as genomic DNA or chromatin DNA, can be isolated, treated with dsDNA deaminase, and processed and analyzed as described herein.
[0031] Methods for permeabilizing cells or nuclei Once one or more desired cells are identified, the method described herein can be carried out on one or more such cells or one or more nuclei obtained from one or more such cells.Individual cells or a plurality of cells or one or more nuclei can be isolated.One or more cells or one or more nuclei can be treated according to known methods to facilitate the entry of chemicals, drugs, enzymes such as dsDNA deaminase, DNA, or other reagents that are to be introduced into one or more cells or one or more nuclei.
[0032] According to one embodiment, one or more cells or one or more nuclei may be permeabilized.Permeabilization methods are known to those skilled in the art.Exemplary permeabilization techniques include electroporation or electropermeabilization, permeabilization with mild non-ionic detergents such as saponin and digitonin, and permeabilization with pore-forming toxins such as alpha-toxin and streptolysin O, which are known in the art.Electroporation, or electropermeabilization, or electrophoretic transfer, is a technique that applies an electric field to a cell or a cell nucleus to increase the permeability of the cell membrane, allowing the introduction of chemicals, drugs, enzymes such as dsDNA deaminase, electrode arrays, or DNA into the cell.
[0033] Alternatively, one or more cells can be lysed to obtain one or more nuclei, or one or more nuclei can be lysed using methods known to those skilled in the art.Lysation can be achieved, for example, by heating cells, by using detergents or other chemical methods, or by a combination thereof.However, any suitable lysis method known in the art can be used.
[0034] Alternatively, DNA, such as genomic or chromatin DNA, may be extracted or otherwise isolated from tissue samples, blood samples, one or more cells, etc., by methods known in the art. DNA extraction protocols using beads (such as DYNABEADS) or reagents are known to those of skill in the art, and kits are commercially available from ThermoFisher Scientific, such as the CHARGESWITCH Genomic DNA Purification Kit.
[0035] Methods for cross-linking cells and nuclei In certain embodiments, one or more cells or one or more nuclei are treated with a crosslinking agent, as known in the art, prior to dsDNA deaminase treatment to maintain cellular structure. Such crosslinking treatments include treatment with paraformaldehyde or ultraviolet light to produce crosslinks. Crosslinking is used to maintain cellular structure and, in some cases, can stably crosslink bound TFs to DNA. The methods described herein can be applied to cells and nuclei treated with a crosslinking agent. In one aspect, the methods described herein can be applied to formalin-fixed, paraffin-embedded (FFPE) samples, as known in the art. For example, FFPE is a form of specimen preservation and preparation. Tissue samples are first preserved by fixation with formaldehyde, also known as formalin, such as a solution of 10% neutral buffered formalin for approximately 18-24 hours to preserve proteins and important structures within the tissue. They are then embedded in a paraffin wax block, and the sample can then be processed according to the methods described herein. To prepare for wax infiltration, tissues are often dehydrated and cleared using increasing concentrations of ethanol. They are then embedded in IHC grade paraffin according to known methods.
[0036] DNA-binding proteins According to a specific embodiment of the present disclosure, a method for determining DNA-binding protein binding sites on DNA, such as chromatin DNA or cellular DNA, is provided. DNA-binding proteins are proteins that have a DNA-binding domain and therefore have specific or general affinity for single-stranded or double-stranded DNA. Sequence-specific DNA-binding proteins generally interact with the major groove of B-DNA, since more functional groups that specify base pairs are exposed.
[0037] DNA-binding proteins include transcription factors that modulate the process of transcription, various polymerases, nucleases that cleave DNA molecules, and histones that are involved in chromosome packaging and transcription in the cell nucleus. DNA-binding proteins may incorporate domains such as zinc finger, helix-turn-helix, and leucine zipper (among many others) that facilitate binding to nucleic acids.
[0038] Structural proteins that bind to DNA are a well-understood example of nonspecific DNA-protein interactions. Within chromosomes, DNA is held in complexes with structural proteins. These proteins organize DNA into a compact structure called chromatin. In eukaryotes, this structure involves DNA binding to a complex of small basic proteins called histones. Histones form disk-shaped complexes called nucleosomes, which contain two complete turns of double-stranded DNA wrapped around their surface. These nonspecific interactions are formed through basic residues of histones that form ionic bonds with the acidic phosphate-sugar backbone of DNA and are therefore largely independent of base sequence. Chemical modifications of these basic amino acid residues include methylation, phosphorylation, and acetylation. These chemical changes alter the strength of the interaction between DNA and histones, thereby making the DNA more or less accessible to transcription factors and altering the rate of transcription. Other nonspecific DNA-binding proteins of chromatin include high-mobility group (HMG) proteins, which bind to curved or distorted DNA. These proteins are important for curving and organizing rows of nucleosomes into the larger structures that form chromosomes.
[0039] In contrast, transcription factors bind to specific DNA sequences. As is known in the art, each transcription factor binds to a specific set of DNA sequences and activates or inhibits transcription of genes that have those sequences near their promoters. The specificity of transcription factor interactions with DNA arises from the fact that proteins make multiple contacts with the ends of DNA bases, allowing them to read the DNA sequence. Most of these base interactions occur in the major groove, where the bases are most accessible. Transcription factors (TFs) (or sequence-specific DNA-binding factors) are proteins that control the rate of transcription of genetic information from DNA to messenger RNA by binding to specific DNA sequences. The function of TFs is to regulate the on and off of genes throughout the life of cells and organisms, ensuring that genes are expressed in the correct amounts and at the correct time in the desired cells. TFs work together to direct cell division, proliferation, and death throughout life, direct cell migration and organization during embryonic development, and respond intermittently to signals from outside the cell, such as hormones. There are up to 1,600 TFs in the human genome. See Babu MM, Luscombe NM, Aravind L, Gerstein M, Teichmann SA (2004 June) "Structure and evolution of transcriptional regulatory networks" (PDF). Current Opinion in Structural Biology. 14(3):283-91. doi:10.1016 / j.sbi.2004.05.004. PMID 15193307, which is incorporated herein by reference in its entirety.
[0040] TFs act alone or with other proteins in complexes by promoting (as activators) or preventing (as repressors) the recruitment of RNA polymerase (the enzyme that carries out the transcription of genetic information from DNA to RNA) to specific genes. The defining feature of TFs is that they contain at least one DNA-binding domain (DBD) that attaches to a specific sequence of DNA adjacent to the gene they regulate. See Mitchell PJ, Tjian R (July 1989). "Transcriptional regulation in mammalian cells by sequence-specific DNA binding proteins". Science. Volume 245 (Issue 4916): pp. 371-8. Bibcode: 1989Sci...245..371M. doi:10.1126 / science.2667136. PMID2667136; Ptashne M, Gann A (April 1997). "Transcriptional activation by recruitment". Nature. Volume 386 (Issue 6625): pp. 569-77. Bibcode: 1997Natur. Volume 386..569. doi:10.1038 / 386569a0. PMID9121580. S2CID6203915. Each of these references is incorporated herein by reference in its entirety for its teaching of known transcription factors. TFs are grouped into classes based on their DNA-binding domains.See Stegmaier P, Kel AE, Wingender E (2004). "Systematic DNA-binding domain classification of transcription factors." Genome Informatics. International Conference on Genome Informatics. Vol. 15(2): pp. 276-86. PMID 15706513. Archived from the original on June 19, 2013; Matys V et al. (January 2006). "TRANSFAC and its module TRANSCompel: transcriptional gene regulation in eukaryotes." Nucleic Acids Research. Vol. 34 (Database Issue): D108-10. doi:10.1093 / nar / gkj143. PMC1347505. PMID 16381825. Each of these references is incorporated herein by reference in its entirety for its teaching of transcription factors and their associated DNAS-binding domains.
[0041] According to one embodiment, exemplary transcription factors include, but are not limited to, AAF, ABL, ADA2, ADANFl, AF1, AFP1, AHR, AIIN3, AIRE, ALL1, ALPHACBF, ALPHACP1, ALPHACP2A, ALPHACP2B, ALPHAH2, ALPHAH3, ALPHAHO, ALX1, ALX3, ALX4, AMEF2, AML1, AML1A, AML1B, AML1C, AML1DELTAN, AML2, AML3, AML3A, AML3B, AMY1L, AMYB, ANF, AN HX, AP1, AP2ALPHAA, AP2ALPHAB, AP2BETA, AP2GAMMA, AP3(1), AP3(2), AP4, AP5, APC, AR, AREB6, ARGFX, ARID5B, ARNT, ARNT(774MFORM), ARNT2, ARNT::H IF1A, ARNTL, ARP1, ARX, ASCL1, ASCL2, ATBF1A, ATBF1B, ATF, ATF1, ATF2, ATF3, ATF3DELTAZIP, ATF4, ATF6, ATF6B, ATF7, ATFA, ATFADELTA, ATOH1, ATOH7 , ATPF1, B, BACH1, BACH2, BANP, BARH11, BARH12, BARHL1, BARHL2, BARX1, BARX2, BATF, BATF3, BATF::JUN, BBX, BCL11A, BCL11B, BCL3, BCL6, BCL6B, BD73 , BETACATENIN, BHLHA15, BHLHE22, BHLHE23, BHLHE40, BHLHE41, BIN1, BMYB, BNC2, BP1, BP2, BPTF, BRAHMA, BRCA1, BRN3A, BRN3B, BRN4, BSX, BTEB, BTEB2 , BTFIID, C / EBPALPHA, C / EBPBETA, C / EBPDELTA, CACCBINDINGFACTOR, CART1, CBF(4), CBF(5), CBP, CCAATBINDINGFACTOR, CCF, CCG1, CCK1A, CCK1B, CCM TBINDINGFACTOR, CD28RC, CDC5L, CDK2, CDK9, CDX1, CDX2, CDX4, CEBPA, CEBPB, CEBPD, CEBPE, CEBPG, CENPB, CENPBD1, CFF, CHXLO, CLIM2, CLIMI, CLOCK,CNBP、COS、COUP、CP1、CP2、CPBP、CPEB1、CPEBINDINGPROTEIN、CPIA、CPIC、C REB、CREB1、CREB2、CREB3、CREB3L1、CREB3L4、CREB5、CREBPLCREBPA、CREM、C REMALPHA, CRF, CRX, CSBP1, CTCF, CTCFL, CTF, CTF1, CTF2, CTF3, CTF5, CTF7 、CUP、CUTL1、CUX1、CUX2、CX、CXXC5、CYCLINA、CYCLINT1、CYCLINT2、CYCLINT 2A、CYCLINT2B、DAP、DAX1、DB1、DBF4、DBP、DBPA、DBPAV、DBPB、DDB、DDB1、DD B2、DEF、DELTACREB、DELTAMAX、DF1、DF2、DF3、DIX4(LONGISOFORM)、DLX1、DL X2、DLX3、DLX4、DLX4(SHORTISOFORM、DLX5、DLX6、DMRT1、DMRT2、DMRT3、DMR TA1、DMRTA2、DMRTC2、DNMT1、DP1、DP2、DPF1、DPRX、DRGX、DSIF、DSIFP14、DSI FP160、DTF、DUX1、DUX2、DUX3、DUX4、DUXA、E、E12、E2F、E2F+E4、E2F+P107、E 2F1、E2F2、E2F3、E2F4、E2F5、E2F6、E2F7、E2F8、E47、E4BP4、E4F、E4F1、E4TF2 、EAR2、EBF1、EBF3、EBP80、EC2、EF1、EFC、EGR1、EGR2、EGR3、EGR4、EHF、EIF1 、EIIAEA、EIIAEB、EIIAECALPHA、EIIAECBETA、EIVF、ELF1、ELF2、ELF3、ELF4、 ELF5、ELK1、ELK1::HOXA1、ELK1::HOXB13、ELK1::SREBF2、ELK3、ELK4、EMX1 EMX2、EN1、EN2、ENHBIND.PROT、ENKTF1、EOMES、EPAS1、EPSILONF1、ER、ERF、 ERF::FIGLA、ERF::FOXI1、ERF::FOXO1、ERF::HOXB13、ERF::NHLH1、ERF::S REBF2、ERG、ERG1、ERG2、ERR1、ERR2、ESR1、ESR2、ESRRA、ESRRB、ESRRG、ESX1、ETF、ETS1、ETS1DELTAVIL、ETS2、ETV1、ETV2、ETV2::DRGX、ETV2::FIGLA、ET V2::FOXI1、ETV2::HOXB13、ETV3、ETV4、ETV5、ETV5::DRGX、ETV5::FIGLA、E TV5::FOXI1、ETV5::FOXO1、ETV5::HOXA2、ETV6、ETV7、EVX1、EVX2、F2F、FAC TOR2, FACTORNAME, FBP, FEBP, FERD3L, FEV, FEZF1, FIGLA, FKBP59, FKHL18, F KHRL1P2, FLI1, FLI1::DRGX, FLI1::FOXI1, FOS, FOS::JUN, FOS::JUNB, FOS ::JUND、FOSB、FOSB::JUN、FOSB::JUNB、FOSL1、FOSL1::JUN、FOSL1::JUNB、 FOSL1::JUND、FOSL2、FOSL2::JUN、FOSL2::JUNB、FOSL2::JUND、FOXA1、FOX A2, FOXA3, FOXB1, FOXC1, FOXC2, FOXD1, FOXD2, FOXD3, FOXD4, FOXE1, FOXE3 FOXF1, FOXF2, FOXG1, FOXG1A, FOXG1B, FOXG1C, FOXH1, FOXI1, FOXJ1A, FOXJ 1B、FOXJ2、FOXJ2(LONGISOFORM)、FOXJ2(SHORTISOFORM)、FOXJ2::ELF1、FO XJ3, FOXK1, FOXK1A, FOXK1B, FOXK1C, FOXK2, FOXL1, FOXL2, FOXM1, FOXM1A FOXM1B、FOXM1C、FOXN1、FOXN2、FOXN3、FOXO1、FOXO1::ELF1、FOXO1::ELK1、F OXO1::ELK3、FOXO1::FLI1、FOXO1A、FOXO1B、FOXO2、FOXO3、FOXO3A、FOXO3B FOXO4, FOXO6, FOXP1, FOXP2, FOXP3, FOXQ1, FOXR1, FOXR2, FRA1, FRA2, FTF. FTS、G6FACTOR、GABP、GABPA、GABPALPHA、GABPBETA1、GABPBETA2、GADD153、GAF、GAMMACAC1、GAMMACAC2、GAMMACMT、GATA1、GATA1::TAL1、GATA2、GATA3GATA4, GATA5, GATA6, GBX1, GBX2, GCF, GCM1, GCM2, GCMA, GCNS, GF1, GFACTO R、GFI1、GFI1B、GLI、GLI1、GLI2、GLI3、GLI4、GLIS1、GLIS2、GLIS3、GMEB1、G MEB2, GRALPHA, GRBETA, GRF1, GRHL1, GRHL2, GSC, GSC2, GSCL, GSX1, GSX2, G TF3A, GTIC, GTIIA, GTIBALPHA, GTIIBBETA, H1TF1, H1TF2, H2RIIBP, H4TF1 H4TF2, HAND1, HAND2, HB9, HDAC1, HDAC2, HDAC3, HDAXX, HDX, HEATINDUCEDF ACTOR HEB HEB1P67 HEB1P94 HEF1B HEF1T HEF4C HEN1 HEN2 HES1 HES2 HES5, HES6, HES7, HESX1, HEX, HEY1, HEY2, HIC1, HIC2, HIF1, HIF1A, HIF1A LPHA、HIF1BETA、HINFA、HINFB、HINFC、HINFD、HINFD3、HINFE、HINFP、HIP1、H IVEP2, HKR1, HLF, HLTF, HLTF(MET123), HLX, HMBOX1, HMBP, HMGI, HMGI(Y) HMGIC, HMGY, HMX1, HMX2, HMX3, HNF1A, HNF1B, HNF3, HNF3ALPHA, HNF3BETA HNF3GAMMA, HNF4, HNF4A, HNF4ALPHA, HNF4ALPHA1, HNF4ALPHA2, HNF4ALPHA 3, HNF4ALPHA4, HNF4G, HNF4GAMMA, HNF6ALPHA, HNFIA, HNFIB, HNFIC, HNRNPK 、HOMEZ、HOX11、HOXA1、HOXA10、HOXA11、HOXA13、HOXA2、HOXA3、HOXA4、HOXA 5. HOXA6, HOXA7, HOXA9, HOXA9A, HOXA9B, HOXAIO, HOXAIOPL2, HOXB1, HOXB13 、HOXB2、HOXB2::ELK1、HOXB3、HOXB4、HOXB5、HOXB6、HOXB7、HOXB8、HOXB9、H OXC10, HOXC11, HOXC12, HOXC13, HOXC4, HOXC5, HOXC6, HOXC8, HOXC9, HOXD1HOXD10、HOXD11、HOXD12、HOXD12::ELK1、HOXD13、HOXD3、HOXD4、HOXD8、HOX D9、HP55、HP65、HPX42B、HRPF、HSF、HSF1、HSF1(LONG)、HSF1(SHORT)、HSF2、 HSF4、HSF5、HSFY1、HSFY2、HSP56、HSP90、IBP1、ICERII、ICERLIGAMMA、ICSB P、ID1、ID1H'、ID2、ID3、ID3 / HEIR1、IF1、IGPE1、IGPE2、IGPE3、II1RF、IKAP PAB、IKAPPABALPHA、IKAPPABBETA、IKAPPABR、IKZF1、IKZF3、IL6REBP、INSAF、INSM1、IPF1、IRF1、IRF2、IRF3、IRF4、IRF5、IRF6、IRF7、IRF8、IRF9、IRX1 、IRX2、IRX2A、IRX3、IRX4、IRX5、ISGF1、ISGF3、ISGF3ALPHA、ISGF3GAMMA、I SL1、ISL2、ISX、ITF、ITF1、ITF2、JDP2、JRF、JUN、JUN::JUNB、JUNB、JUND、KAP PAYFACTOR、KBP1、KDM2B、KER1、KLF1、KLF10、KLF11、KLF12、KLF13、KLF14、K LF15、KLF16、KLF17、KLF2、KLF3、KLF4、KLF5、KLF6、KLF7、KLF8、KLF9、KMT2A 、KOX1、KRF1、KUAUTOANTIGEN、KUP、LBP1、LBP1A、LBX1、LBX2、LCORL、LCRF1、 LEF1、LEFIB、LFA1、LHX1、LHX2、LHX3、LHX3A、LHX3B、LHX5、LHX6、LHX6.1A、L HX6.1B、LHX8、LHX9、LIN28B、LIT1、LMO1、LMO2、LMX1A、LMX1B、LMY1(LONGFO) RM)、LMY1(SHORTFORM)、LMY2、LSF、LXRALPHA、LY11、LYF1、LYL1、MAD1、MAF、 MAF::NFE2、MAFA、MAFB、MAFF、MAFG、MAFG::NFE2L1、MAFK、MASH1、MAX、MAX1 、MAX2、MAX::MYC、MAZ、MAZ1、MB67、MBD2、MBF1、MBF2、MBF3、MBNL2、MBP1(1)、MBP1(2), MBP2, MDBP, MECOM, MECP2, MEF2, MEF2A, MEF2B, MEF2C, MEF2C(433AAFORM), MEF2C(465AAFORM), MEF2C(473MFORM), MEF2C / DELTA32(441AAFORM), MEF, 2D、MEF2D00、MEF2D0B、MEF2DA'B、MEF2DA0、MEF2DAB、MEF2DAO、MEIS1、MEIS2、MEIS2A、MEIS2B、MEIS2C、MEIS2D、MEIS2E、MEIS3、MEOX1、MEOX1A、MEOX2、 MESP1、MESP2、MFACTOR、MGA、MGA::EVX1、MHOX(K2)、MI、MIF1、MITF、MIXL1、MIZ1、MLX、MLXIPL、MM1、MNT、MNX1、MOP3、MR、MSANTD3、MSC、MSGN1、MSX1、MSX 2、MTBZF、MTF1、MTF2、MTTF1、MXI1、MXIL、MYB、MYBL1、MYBL2、MYC、MYC1、MYC N、MYF3、MYF4、MYF5、MYF6、MYNN、MYOD、MYOD1、MYOG、MYRF、MZF1、N10(25、NA NOG、NC2、NCI、NCX、NELF、NER1、NET、NEUROD1、NEUROD2、NEUROG1、NEUROG2、 NF1A、NF1B、NF1X、NF4FA、NF4FB、NF4FC、NFA、NFAB、NFAT1、NFAT3、NFAT5、NFA TC、NFATC1、NFATC2、NFATC3、NFATC4、NFATP、NFATX、NFCLE0A、NFCLE0B、NFDELTAE3A、NFDELTAE3B、NFDELTAE3C、NFDELTAE4A、NFDELTAE4B、NFDELTAE4C 、NFE、NFE2、NFE2L1、NFE2L2、NFE2P45、NFE3、NFE6、NFETAA、NFGMA、NFGMB、NFI11A、NFIA、NFIB、NFIC、NFIC::TLX1、NFIL2A、NFIL2B、NFIL3、NFIX、NFJUN、 NFKAPPAB, NFKAPPAB(LIKE) IIA、NFMHCIIB、NFMUE1、NFMUE2、NFMUE3、NFNF1、NFS、NFX、NFX1、NFX2、NFX3 、NFXC、NFYA、NFYB、NFYC、NFZC、NFZZ、NHLH1、NHLH2、NHP1、NHP2、NHP3、NHP4、NKX21、NKX22、NKX23、NKX24、NKX25、NKX28、NKX2B、NKX2C、NKX2G、NKX31、NKX32、NKX3A、NKX3AV1、NKX3AV2、NKX3AV3、NKX3AV4、NKX3B、NKX61、NKX62、NKX63、NKX6A、NMI、NMYC、NOBOX、NOCT2ALPHA、NOCT2BETA、NOCT3、NOCT4、NOCT5A、NOCTSB、NOTO、NPAS2、NPTCII、NR1D1、NR1D2、NR1H2::RXRA、NR1H3、NR1H4、NR1H4::RXRA、NR1I2、NR1I3、NR2C1、NR2C2、NR2E1、NR2E3、NR2F1、NR2F2、NR2F6、NR3C1、NR3C2、NR4A1、NR4A2、NR4A2::RXRA、NR5A1、NR5A2、NR6A1、NRF1、NRF2、NRF2BETA1、NRF2GAMMA1、NRL、NRSFFORM1、NRSFFORM2、NTF、OCAB、OCT1、OCT2、OCT2.1、OCT2B、OCT2C、OCT4A、OCT4B、OCT5、OCT6、OCTAFACTOR、OCTAMERBINDINGFACTOR、OCTB2、OCTB3、OLIG1、OLIG2、OLIG3、ONECUT1、ONECUT2、ONECUT3、OSR1、OSR2、OTX1、OTX2、OVOL1、OVOL2、OZF、P107、P130、P28MODULATOR、P300、P38ERG、P45、P49ERG、P53、P55、P55ERG、P65DELTA、P67、PATZ1、PAX1、PAX2、PAX3、PAX3A、PAX3B、PAX4、PAX5、PAX6、PAX6 / PD5A、PAX7、PAX8、PAX8A、PAX8B、PAX8C、PAX8D、PAX8E、PAX8F、PAX9、PBX1、PBX1A、PBX1B、PBX2、PBX3、PBX3A、PBX3B、PBX4、PC2、PC4、PCS、PDX1、PEA3、PEBP2ALPHA、PEBP2BETA、PGR、PHF1、PHOX2A、PHOX2B、PIT1、PITX1、PITX2、PITX3、PKNOX1、PKNOX2、PLAG1、PLAGL2、PLZF、POB、PONTIN52、POU1F1、POU2F1、POU2F1::SOX2、POU2F2, POU2F3, POU3F1, POU3F2, POU3F3, POU3F4, POU4F1, POU4F2, POU4F3 POU5F1, POU5F1B, POU6F1, POU6F2, PPARA, PPARA::RXRA, PPARALPHA, PPAR BETA、PPARD、PPARG、PPARG::RXRA、PPARGAMMA1、PPARGAMMA2、PPUR、PR、PRA 、PRB、PRD1BF1、PRDIBFC、PRDM1、PRDM14、PRDM4、PRDM6、PRDM9、PRECURSOR、P ROP1, PROX1, PRRX1, PRRX2, PSE1, PTEFB, PTF, PTF1A, PTFALPHA, PTFBETA, PTFDELTA, PTFGAMMA, PU.1, PUBOXBINDINGFACTOR, PUBOXBINDINGFACTOR(BJA B)、PUF、PURFACTOR、R1、R2、RARA、RARA::RXRA、RARA::RXRG、RARALPHA1、RA RB、RARBETA、RARBETA2、RARG、RARGAMMA、RARGAMMA1、RAX、RAX2、RBAK、RBP60 、RBPJ、RBPJKAPPA、REL、RELA、RELB、REST、RFX、RFX1、RFX2、RFX3、RFX4、RFX 5, RFX7, RFXS, RFY, RHOXF1, RORA, RORALPHA1, RORALPHA2, RORALPHA3, RORB RORBETA, RORC, RORGAMMA, ROX, RPF1, RPGALPHA, RREB1, RSRFC4, RSRFC9, R UNX1、RUNX2、RUNX3、RVF、RXRA、RXRA::VDR、RXRALPHA、RXRB、RXRBETA、RXRG、 SALL4, SAP1A, SAP1B, SATB1, SCRT1, SCRT2, SF1, SHOX2A, SHOX2B, SHOXA, SH OXB, SHP, SIIIP110, SIIIP15, SIIIP18, SIM', SIX1, SIX2, SIX3, SIX4, SIX5 SIX6, SCORE1, SCORE2, SMAD1, SMAD2, SMAD3, SMAD4, SMAD5, SNAI1, SNAI2, SNA I3, SOHLH2, SOX10, SOX11, SOX12, SOX13, SOX14, SOX15, SOX17, SOX18, SOX2SOX21、SOX3、SOX30、SOX4、SOX5、SOX6、SOX7、SOX8、SOX9、SP1、SP2、SP3、SP4 SP5, SP8, SP9, SPDEF, SPHFACTOR, SPI1, SPIB, SPIC, SPIN, SPZ1, SRCAP, SR EBF1、SREBF2、SREBP1A、SREBP1B、SREBP1C、SREBP2、SREZBP、SRF、SRPLSTAF 50、SRY、STAT1、STAT1::STAT2、STAT1ALPHA、STAT1BETA、STAT2、STAT3、STA T4、STAT5A、STAT5B、STAT6、T、T3R、T3RALPHA1、T3RALPHA2、T3RBETA、TAF(I )110、TAF(I)48、TAF(I)63、TAF(II)100、TAF(II)125、TAF(II)135、TAF(II) )170、TAF(II)18、TAF(II)20、TAF(II)250、TAF(II)250DELTA、TAF(II)28、 TAF(II)30、TAF(II)31、TAF(II)55、TAF(II)70ALPHA、TAF(II)70BETA、TAF( II)70GAMMA、TAFI、TAFII、TAFL、TAL1、TAL1::TCF3、TAL1BETA、TAL2、TARFA CTOR, TBP, TBR1, TBX1, TBX15, TBX18, TBX19, TBX1A, TBX1B, TBX2, TBX20, TB X21、TBX3、TBX4、TBX5、TBX6、TBXS(LONGISSOFORM)、TBXS(SHORTISOFORM)、T BXT, TCF, TCF1, TCF12, TCF1A, TCF1B, TCF1C, TCF1D, TCF1E, TCF1F, TCF1G, TC F21, TCF2ALPHA, TCF3, TCF4, TCF4(K), TCF4B, TCF4E, TCF7, TCF7L1, TCF7L2 、TCFBETA1、TCFL5、TEAD1、TEAD2、TEAD3、TEAD4、TEF、TEF1、TEF2、TEL、TET1 、TFAP2A、TFAP2B、TFAP2C、TFAP2E、TFAP4、TFAP4::ETV1、TFAP4::FLI1、TFC P2、TFCP2L1、TFDP1、TFE3、TFEB、TFEC、TFIA、TFIIAALPHA / BETAPRECURSOR、TFIIAGAMMA、TFIIB、TFIID、TFIIE、TFIIEALPHA、TFIIEBETA、TFIIF、TFIIFALPHA、TFIIFBETA、TFIIH、TFIIH*、TFIIHCAK、TFIIHCYCLINH、TFIIHERCC2 / CAK、TFIIHM015、TFIIHMAT1、TFIIHP34、TFIIHP44、TFIIHP62、TFIIHP80、TFIIHP90、TFIII、TFLFLTFLF2、TGIF、TGIF1、TGIF2、TGIF2LX、TGIF2LY、TGT3、TH AP1、THAP11、THAP12、THRA、THRA1、THRB、TIF2、TIGD1、TLE1、TLX2、TLX3、TM F、TOPORS、TP53、TP63、TP73、TR2、TR211、TR29、TR3、TR4、TRAP、TREB1、TREB2 、TREB3、TREF1、TREF2、TRF(2)、TRPS1、TTF1、TWIST1、TXREBP、TXREF、UBF、U BP1、UEF1、UEF2、UEF3、UEF4、UNCX、USF1、USF2、USF2B、VAV、VAX2、VDR、VENTX 、VEZF1、VHNF1A、VHNF1B、VHNF1C、VITF、VSX1、VSX2、WSTF、WT1、WT1DE12、WT 1I、WT1IDE12、WT1IKTS、WT1KTS、X2BP、XBP1、XPA、XWV、XX、YAF2、YB1、YBX1、Y EBP、YY1、YY2、ZBED1、ZBED2、ZBTB12、ZBTB14、ZBTB17、ZBTB18、ZBTB2、ZBTB 20、ZBTB22、ZBTB26、ZBTB32、ZBTB33、ZBTB37、ZBTB42、ZBTB43、ZBTB44、ZBTB 45、ZBTB48、ZBTB49、ZBTB6、ZBTB7A、ZBTB7B、ZBTB7C、ZEB、ZEB1、ZF1、ZF2、Z FHX2、ZFHX3、ZFP1、ZFP14、ZFP28、ZFP3、ZFP41、ZFP42、ZFP57、ZFP64、ZFP69、 ZFP69B、ZFP82、ZFP90、ZFX、ZHX1、ZIC1、ZIC2、ZIC3、ZIC4、ZIC5、ZID、ZIK1、 ZIM2、ZIM3、ZKSCAN1、ZKSCAN2、ZKSCAN3、ZKSCAN5、ZKSCAN7、ZNF10、ZNF100、ZNF101、ZNF114、ZNF12、ZNF121、ZNF124、ZNF132、ZNF133、ZNF134、ZNF135、ZNF136、ZNF140、ZNF141、ZNF143、ZNF146、ZNF148、ZNF154、ZNF157、ZNF16、ZNF17、ZNF174、 ZNF175、ZNF177、ZNF18、ZNF180、ZNF181、ZNF182、ZNF184、ZNF189、ZNF19、ZNF197、ZNF2、ZNF200、ZNF202、ZNF205、ZNF211、ZNF212、ZNF213、ZNF214、ZNF22、ZNF222、ZNF223、ZNF224、ZNF225、ZNF23、ZNF232、ZNF235、ZNF24、ZNF248、ZNF25、ZNF250、ZNF254、ZNF257、ZNF26、ZNF260、ZNF263、ZNF264、ZNF266、ZNF267、ZNF273、ZNF274、ZNF276、ZNF28、ZNF280A、ZNF281、ZNF282、ZNF283、ZNF284、ZNF285、ZNF287、ZNF296、ZNF3、ZNF30、ZNF300、ZNF302、ZNF304、ZNF311、ZNF317、ZNF32、ZNF320、ZNF322、ZNF324、ZNF324B、ZNF329、ZNF331、ZNF333、ZNF334、ZNF335、ZNF337、ZNF33A、ZNF33B、ZNF34、ZNF341、ZNF343、ZNF345、ZNF35、ZNF350、ZNF354A、ZNF354B、ZNF37A、ZNF382、ZNF383、ZNF384、ZNF385D、ZNF394、ZNF396、ZNF398、ZNF41、ZNF410、ZNF415、ZNF416、ZNF417、ZNF418、ZNF419、ZNF423、ZNF425、ZNF429、ZNF430、ZNF431、ZNF432、ZNF433、ZNF436、ZNF439、ZNF44、ZNF440、ZNF441、ZNF442、ZNF443、ZNF444、ZNF445、ZNF449、ZNF45、ZNF454、ZNF460、ZNF467、ZNF468、ZNF479、ZNF480、ZNF483、ZNF484、ZNF485、ZNF486、ZNF487、ZNF490、ZNF492、ZNF496、ZNF501、ZNF502、ZNF506、ZNF513、ZNF519、ZNF524、ZNF525、ZNF527、ZNF528、ZNF529、ZNF530、ZNF534、ZNF540、ZNF543、ZNF547、ZNF548、ZNF549、ZNF550、ZNF552、ZNF554、ZNF555, ZNF558, ZNF561, ZNF562, ZNF563, ZNF564, ZNF565, ZNF566, ZNF567, ZNF570, ZNF571, ZNF573, ZNF574, ZNF580, ZNF582, ZNF584, Z NF585A, ZNF586, ZNF587, ZNF594, ZNF595, ZNF596, ZNF597, ZNF605, ZNF610, ZNF611, ZNF613, ZNF614, ZNF615, ZNF616, ZNF619, ZNF620, ZN F621, ZNF626, ZNF627, ZNF641, ZNF652, ZNF653, ZNF655, ZNF660, ZNF662, ZNF667, ZNF669, ZNF671, ZNF674, ZNF675, ZNF677, ZNF680, ZNF 681, ZNF682, ZNF684, ZNF69, ZNF692, ZNF695, ZNF7, ZNF701, ZNF704, ZNF705G, ZNF707, ZNF708, ZNF71, ZNF711, ZNF713, ZNF714, ZNF716, Z NF730, ZNF736, ZNF737, ZNF74, ZNF740, ZNF749, ZNF75A, ZNF75D, ZNF76, ZNF764, ZNF765, ZNF766, ZNF768, ZNF77, ZNF770, ZNF771, ZNF77 4, ZNF776, ZNF777, ZNF778, ZNF780A, ZNF782, ZNF783, ZNF784, ZNF785, ZNF786, ZNF787, ZNF789, ZNF79, ZNF790, ZNF791, ZNF792, ZNF793, ZNF799, ZNF8, ZNF805, ZNF808, ZNF81, ZNF816, ZNF821, ZNF823, ZNF84, ZNF85, ZNF860, ZNF879, ZNF880, ZNF891, ZNF90, ZNF93, ZNF98, ZSCAN1, ZSCAN16, ZSCAN22, ZSCAN23, ZSCAN29, ZSCAN30, ZSCAN31, ZSCAN4, ZSCAN5, ZSCAN5C, ZSCAN9, and ZZZ3, as well as others that can be identified in the literature.
[0042] Those skilled in the art can easily identify transcription factor by using database and literature sources.In addition to the above database, transcription factor is identified in US Patent Application Publication No. 2022 / 0214356.This document is incorporated herein by reference in its entirety for the description of transcription factor.
[0043] double-stranded DNA deaminase According to one embodiment, an exemplary double-stranded DNA deaminase includes DddA, also known in the art as BadTF1. See Mok, BY et al. (2020). A bacterial cytidine deaminase toxin enables CRISPR-free mitochondrial base editing. Nature, Vol. 583 (No. 7817), pp. 631-637, which is incorporated herein by reference in its entirety for its teaching of DddA. DddA (BadTF1) belongs to the SCP1.201-like deaminase subfamily.
[0044] According to one embodiment, an exemplary double-stranded DNA deaminase includes BadTF3. See Marcos H de Moraes et al. (2021) An interbacterial DNA deaminase toxin directly mutagenizes surviving target populations eLife 10:e62967. BadTF3 belongs to the Pput_2613-like deaminase subfamily.
[0045] According to one embodiment, exemplary double-stranded DNA deaminases include DddA11. See Mok, BY et al. (2022). CRISPR-free base editors with enhanced activity and expanded targeting scope in mitochondrial and nuclear DNA. Nature Biotechnology, 1-10, which is incorporated herein by reference in its entirety for its teaching of DddA11.
[0046] Deaminases are known in the art to be classified for identification purposes. See Iyer LM et al., Evolution of the deaminase fold and multiple origins of eukaryotic editing and mutagenic nucleic acid deaminases from bacterial toxin systems [J]. Nucleic acids research, 2011, Vol. 39(22):9473-9497, which is incorporated herein by reference in its entirety for purposes of classification and identification of deaminases. DddA of the present disclosure belongs to the SCP1.201 phylogenetic group. BadTF3 of the present invention belongs to the Pput_2613-like phylogenetic group. It should be understood that the specific dsDNA deaminases described herein are merely exemplary. Mutants, variants, derivatives, and modified forms of dsDNA deaminases that exhibit enzymatic activity are contemplated by the present disclosure. The present disclosure contemplates the identification by one of skill in the art of other dsDNA deaminases and mutants, variants, derivatives, and modified forms thereof that exhibit enzymatic activity.
[0047] In some examples, the double-stranded DNA deaminase may be derived from DddA deaminases, such as DddA6, DddA7, and other DddA variants. In other examples, the double-stranded DNA deaminase may be derived from BadTF2 and other BadTF2 variants. In other examples, the double-stranded DNA deaminase may be derived from BadTF3 and other BadTF3 variants. In further examples, the deaminase may be derived from bacterial toxins, such as Pput_2613 family deaminases, SCP1.201-like family deaminases, DYW-like family deaminases, BURPS668_1122-like family deaminases, YwqJ-like family deaminases, MafB19-like family deaminases, sce3516-like family deaminases, BH3703-like deaminases, and WD0512-like family deaminases.
[0048] It should be understood that the embodiments of the present disclosure include the mutants, variants, shortened forms, modified forms, and derivatives of the full-length dsDNA deaminases that exhibit deaminase activity, as described herein and known to those skilled in the art.Methods for mutating, changing, shortening, modifying, or derivatizing known dsDNA deaminases are known to those skilled in the art.Therefore, in the present disclosure, it is contemplated that dsDNA deaminases exhibit deaminase activity on double-stranded nucleic acids, regardless of whether they are naturally occurring or modified, mutant, altered, shortened, derivative, or evolved forms of naturally occurring dsDNA deaminases.Embodiments of the present disclosure provide the nucleic acid sequences and amino acid sequences of various known dsDNA deaminases. Embodiments of the present disclosure include nucleic acid and amino acid sequences that have 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.6%, 99.7%, 99.8% or 99.9% identity to the full-length sequence of a dsDNA deaminase disclosed herein. Based on the present disclosure, one of skill in the art can identify dsDNA deaminases that have percent homology to known dsDNA deaminases that have deaminase activity on dsDNA, or can otherwise test such dsDNA deaminases that have percent homology to known dsDNA deaminases that have deaminase activity for deaminase activity.
[0049] In vitro transposition According to certain embodiments, chromatin DNA treated with dsDNA deaminase is processed using a transposition method, sometimes referred to in the art as transposome-mediated fragmentation or "tagmentation." In tagmentation, transposomes are prepared with DNA that is subsequently cleaved so that a transposition event results in fragmented DNA with adapters. In such methods, target DNA is simultaneously fragmented and tagged, producing fragments tagged with desired DNA sequences for downstream processing. According to one embodiment, a library produced using an in vitro transposition system in Illumina, Inc.'s Nextera technology is used to simultaneously fragment DNA and tag each fragment with a sequence suitable for next-generation sequencing. See U.S. Patent Application Publication No. 20110287435, which is incorporated herein by reference for disclosure of the tagmentation method. For additional useful methods for generating sequencing libraries, see "Single-cell chromatin accessibility reveals principles of regulatory variation." Nature, Vol. 523 (No. 7561), pp. 486-490 (2007); "Massively multiplex single-cell Hi-C." Nature Methods, Vol. 14 (No. 3), pp. 263-266 (2017). See also useful transposition methods described in WO 2016 / 073690 and WO 2018217912. These publications are incorporated herein by reference in their entireties.
[0050] According to certain embodiments, exemplary transposon systems include Tn5 transposase, Mu transposase, Tn7 transposase, or IS5 transposase.Other useful transposon systems are known to those of skill in the art and include the Tn3 transposon system (see Maekawa, T., Yanagihara, K., and Ohtsubo, E. (1996). A cell-free system of Tn3 transposition and transposition immunity. Genes Cells 1, 1007-1016), the Tn7 transposon system (see Craig, N.L. (1991). Tn7: a target site-specific transposon. Mol. Microbiol. 5, 2569-2573), the Tn10 transposon system (see Chalmers, R., Sewitz, S., Lipkow, K., and Crellin, P. (2000). Complete nucleotide sequence of Tn10. J. Bacteriol. 182, pp. 2970-2972), the Piggybac transposon system (Li, X., Burnight, E.R., Cooney, A.L., Malani, N., Brady, T., Sander, J.D., Staber, J., Wheelan, S.J., Joung, J.K., McCray, P.B., Jr. et al. (2013). PiggyBac transposase tools for genome engineering, Proc. Natl. Acad. Sci. USA 110, E2279-2287), the Sleeping beauty transposon system (Ivics, Z., Hackett, P.B., Plasterk, R.H., and Izsvak, Z. (1997). Molecular reconstruction of Sleeping Beauty, a Tc1-like transposon from fish, and its transposition in human cells, Cell 91, pp. 501-510), the Tol2 transposon system (Kawakami, K. (2007), Tol2: a versatile gene transfer vector in vertebrates, Genome Biol. Vol. 8, Suppl. 1, S7).
[0051] According to a basic embodiment, the processed genomic DNA is contacted with Tn5 transposase, each of which is bound to a transposon DNA, to form a transposase / transposon DNA complex dimer called transposome.The transposome binds to the target position along the processed genomic DNA, and cuts the processed genomic DNA into multiple double-stranded fragments with primer binding sites.Processes such as extension and gap filling can be carried out to produce double-stranded products, which are mixed with primers together with DNA polymerase, nucleotides, and amplification reagents to amplify the double-stranded processed genomic DNA fragments.The amplicon is sequenced, for example, using high-throughput sequencing methods known to those skilled in the art.
[0052] nuclease According to one embodiment, target chromatin regions, such as open chromatin regions, may be enriched using nucleases before or after treatment with dsDNA deaminase. Nucleases are enzymes capable of cleaving phosphodiester bonds between nucleic acids. Nucleases cause various cuts in single and double strands of their target molecules. There are two main classifications based on the site of activity: exonucleases digest nucleic acids from the ends; and endonucleases act on the middle region of the target molecule. According to certain embodiments, nucleases can be used to process DNA into fragments. Such processing can be performed before or after treating DNA with dsDNA deaminase. Such fragments can then be amplified and / or sequenced as described herein. Exemplary nucleases include DNase (such as DNase I available from Thermo Fisher), MNase (micrococcal nuclease available from New England Biolabs), or restriction endonucleases (such as FASTDIGEST available from Thermo Fisher), etc. According to one embodiment, lysed cells may be subjected to dsDNA deaminase treatment and then nuclease digestion to enrich for open chromatin regions for sequencing.
[0053] Enrichment of target chromatin regions According to one embodiment, target chromatin regions, such as open chromatin regions, can be enriched before or after treatment with dsDNA deaminase using methods known to those skilled in the art, including ChIP-seq, ChIC, ChEC, ChEC-seq, CUT&TAG, CUT&RUN, Multi-CUT&Tag, NTT-seq, and R-loop CUT&Tag, in addition to methods using transposases, such as Tn5 transposase.
[0054] ChIP sequencing, also known as ChIP-seq, is a method used to analyze protein interactions with DNA. ChIP-seq combines chromatin immunoprecipitation (ChIP) with massively parallel DNA sequencing to identify binding sites of DNA-associated proteins. It can be used to precisely map the global binding sites of any protein of interest. Specific DNA sites that directly physically interact with transcription factors and other proteins can be isolated by chromatin immunoprecipitation. ChIP generates a library of target DNA sites bound by a protein of interest. Massively parallel sequencing analysis can be used in conjunction with whole-genome sequence databases to analyze the interaction patterns of any protein with DNA (see Johnson DS et al. (June 2007). "Genome-wide mapping of in vivo protein-DNA interactions" (PDF). Science. Vol. 316 (Issue 5830): pp. 1497-502) or the patterns of any epigenetic chromatin modifications. This can be applied to any set of ChIP-able transcription factors. See "Whole-Genome Chromatin IP Sequencing (ChIP-Seq)" (PDF). Illumina, Inc. November 26, 2007.
[0055] CUT&Tag sequencing, also known as Cleavage Under Targets and tagmentation-sequencing, is a method used to analyze protein interactions with DNA. CUT&Tag sequencing combines antibody-targeted controlled cleavage by Protein A-Tn5 fusions with massively parallel DNA sequencing to identify binding sites of DNA-associated proteins. This can be used to precisely map the global binding sites of any protein of interest. See "CUT&Tag: a higher resolution, lower cost way to map chromatin," Fred Hutchinson Cancer Research Center, April 29, 2019.
[0056] CUT&RUN sequencing (see U.S. Patent Application Publication No. 2022 / 0214356), also known as Cleavage Under Targets and Release Using Nuclease sequencing, is a method used to analyze protein interactions with DNA. CUT&RUN sequencing combines antibody-targeted controlled cleavage with micrococcal nuclease with massively parallel DNA sequencing to identify binding sites of DNA-associated proteins. This can be used to precisely map the global binding sites of any protein of interest. See "Lay off the ChIPs: CUT&RUN instead," Fred Hutchinson Cancer Research Center, February 20, 2017.
[0057] ChIC uses specific antibodies to detect transcription factor binding sites within the genome by targeting modified micrococcal nuclease (MNase) conjugated to protein A (pA-MN). The modified MNase specifically cleaves DNA in regions that interact with the protein of interest only in the presence of Ca2+ ions, thus allowing for controlled DNA cleavage at the antibody binding site. This technique allows for protein mapping with 100-200 bp resolution and excellent specificity. See Schmid M, Durussel T, Laemmli UK. 2004. ChIC and ChEC. Molecular Cell. 16(1):147-15.
[0058] Chromatin endogenous cleavage (ChEC) can be combined with high-throughput sequencing in a method called ChEC-seq. ChEC-seq relies on fusing a chromatin-associated protein of interest with micrococcal nuclease (MNase) to generate targeted DNA breaks in live cells in the presence of calcium. Because ChEC-seq is not based on immunoprecipitation, it avoids crosslinking, sonication, chromatin solubilization, and potential concerns regarding antibody quality while providing high-resolution mapping with minimal background signal. See Grunberg et al., J. Vis. Exp. 2017;(Vol. 124) e55836, pp. 1–9.
[0059] Other methods for enriching target chromatin regions are known to those of skill in the art and are readily identified through a literature search.
[0060] According to one embodiment, lysed cells are subjected to dsDNA deaminase treatment and then subjected to ChIP-seq, ChIC, ChEC, ChEC-seq, CUT&TAG, CUT&RUN, Multi-CUT&Tag, NTT-seq, or R-loop CUT&Tag processing to enrich chromatin regions targeted by specific binding agents such as antibodies. Alternatively, lysed cells may first be subjected to ChIP-seq, ChIC, ChEC, ChEC-seq, CUT&TAG, CUT&RUN, Multi-CUT&Tag, NTT-seq, or R-loop CUT&Tag processing to enrich target chromatin regions without amplification and sequencing, and then subjected to dsDNA deaminase treatment to map TF footprints of the target regions.
[0061] amplification In certain embodiments, amplification is achieved using PCR. PCR is a reaction that uses a pair or set of primers consisting of an upstream primer and a downstream primer, and a polymerization catalyst such as DNA polymerase, usually a thermostable polymerase enzyme, to create replicate copies of a target polynucleotide. Methods for PCR are well known in the art and are taught, for example, in MacPherson et al. (1991) PCR1: A Practical Approach (IRL Press at Oxford University Press). The term "polymerase chain reaction" ("PCR") by Mullis (U.S. Pat. Nos. 4,683,195, 4,683,202, and 4,965,188) refers to a method for increasing the concentration of a segment of a target sequence without cloning or purification. This process for amplifying a target sequence involves preparing oligonucleotide primers and amplification reagents with the desired target sequence, followed by a precise series of thermal cycles in the presence of a polymerase (e.g., DNA polymerase). The primers are complementary to each strand ("primer binding sequence") of a double-stranded target sequence. Generally, amplification occurs by denaturing the double-stranded target sequence and then annealing the primers to complementary sequences within the target molecule. After annealing, the primers are extended by a polymerase so that new pairs of complementary strands are formed. The steps of denaturation, primer annealing, and polymerase extension can be repeated multiple times (i.e., denaturation, annealing, and extension constitute one "cycle," and there can be many "cycles") to obtain a highly concentrated amplified segment of the desired target sequence. The length of the amplified segment of the desired target sequence is determined by the relative positions of the primers with respect to each other, and therefore this length is a controllable parameter. Due to the repetitive nature of the process, this method is called the "polymerase chain reaction" (hereinafter "PCR"), and the target sequence is said to be "PCR amplified."
[0062] The terms "PCR product," "PCR fragment," and "amplification product" refer to the resulting mixture of compounds after two or more cycles of the PCR steps of denaturation, annealing, and extension are complete. These terms encompass cases where there has been amplification of one or more segments of one or more target sequences.
[0063] Any oligonucleotide or polynucleotide sequence can be amplified by using suitable set of primer molecules.The method and kit for carrying out PCR are well known in the art.All processes that produce duplicate copies of polynucleotide, such as PCR or gene cloning, are collectively referred to as duplication herein.
[0064] The terms "amplification" or "amplifying" refer to a process in which additional or multiple copies of a specific polynucleotide are formed. Amplification includes methods such as PCR, ligation amplification (or ligase chain reaction, LCR), and other amplification methods. These methods are known and widely practiced in the art. See, for example, U.S. Pat. Nos. 4,683,195 and 4,683,202 and Innis et al., "PCR protocols: a guide to methods and applications," Academic Press, Incorporated (1990) (for PCR); and Wu et al. (1989) Genomics 4:560-569 (for LCR). In general, the PCR procedure describes a gene amplification method consisting of (i) sequence-specific hybridization of primers to specific genes in a DNA sample (or library), (ii) subsequent amplification involving multiple rounds of annealing, extension, and denaturation using DNA polymerase, and (iii) screening of the PCR product for a band of the correct size. The primers used are oligonucleotides of sufficient length and appropriate sequence to provide for the initiation of polymerization, i.e., each primer is specifically designed to be complementary to each strand of the genomic locus to be amplified.
[0065] Reagents and hardware for carrying out amplification reactions are commercially available. Primers useful for amplifying sequences from specific gene regions are preferably complementary to and specifically hybridize to sequences in the target region or its flanking regions, and can be prepared using methods known to those skilled in the art. Nucleic acid sequences produced by amplification can be sequenced directly.
[0066] When hybridization occurs between two single-stranded polynucleotides in an antiparallel configuration, the reaction is called "annealing," and such polynucleotides are described as "complementary." A double-stranded polynucleotide can be complementary or homologous to another polynucleotide if hybridization can occur between one of the strands of the first polynucleotide and the second polynucleotide. Complementarity or homology (the degree to which one polynucleotide is complementary to another polynucleotide) can be quantified in terms of the proportion of bases on opposite strands that are expected to form hydrogen bonds with each other according to generally accepted base pairing rules.
[0067] The term "amplification reagents" can refer to reagents necessary for amplification (deoxyribonucleotide triphosphates, buffers, etc.) other than primers, nucleic acid templates, and amplification enzymes. Typically, amplification reagents are placed and contained in a reaction vessel (test tube, microwell, etc.) along with other reaction components. Amplification methods include PCR, which is known to those skilled in the art, as well as rolling circle amplification (Blanco et al., J. Biol. Chem., vol. 264, pp. 8935-8940, 1989), hyperbranched rolling circle amplification (Lizard et al., Nat. Genetics, vol. 19, pp. 225-232, 1998), and loop-mediated isothermal amplification (Notomi et al., Nuc. Acids Res., vol. 28, p. e63, 2000). Each of these documents is incorporated herein by reference in its entirety.
[0068] Other amplification methods described in British Patent Application No. 2,202,328 and PCT Patent Application No. PCT / US89 / 01025 can be used in accordance with the present disclosure. Each of these documents is incorporated herein by reference. Emulsion PCR can be used in accordance with the present disclosure. Other suitable amplification methods include "race PCR" and "one-sided PCR" (Frohman, In: PCR Protocols: A Guide To Methods And Applications, Academic Press, NY, 1990, each of which is incorporated herein by reference). Methods based on ligation of two (or more) oligonucleotides in the presence of a nucleic acid having the sequence of the resulting "di-oligonucleotide," thereby amplifying the di-oligonucleotide, can also be used to amplify DNA in accordance with the present disclosure (Wu et al., Genomics 4:560-569, 1989, which is incorporated herein by reference).
[0069] The RNA to be amplified can be obtained from a single cell or a small population of cells.The method described herein allows the amplification of RNA from any species or organism in a reaction mixture, such as a single reaction mixture carried out in a single reaction vessel.In one aspect, the method described herein comprises sequence-independent amplification of RNA from any source, including but not limited to human, animal, plant, yeast, virus, eukaryotic, and prokaryotic RNA.
[0070] As used herein, the term "primer" generally includes an oligonucleotide, either natural or synthetic, that, when duplexed with a polynucleotide template, serves as a point of initiation of nucleic acid synthesis, such as a sequencing primer, and can be extended from its 3' end along the template to form an extended duplex. Primers include extension primers, amplification primers, or reverse transcription primers.
[0071] The sequence of nucleotides added during the extension process is determined by the sequence of the template polynucleotide. Primers are typically extended by DNA polymerase or reverse transcriptase. Primers typically range in length from 3 to 36 nucleotides, or from 5 to 24 nucleotides, or from 14 to 36 nucleotides. Primers within the scope of the present invention include orthogonal primers, amplification primers, and construction primers. Primer pairs can flank a sequence of interest or a set of sequences of interest. Primers and probes may be degenerate or semi-degenerate in sequence. Primers within the scope of the present invention bind adjacent to the target sequence. A "primer" can generally be considered a short polynucleotide with a free 3'-OH group that hybridizes with the target, thereby binding to a target or template that may be present in a sample of interest and subsequently facilitating the polymerization of a polynucleotide complementary to the target. Primers of the present invention are comprised of nucleotides ranging from 17 to 30 nucleotides. In one embodiment, the primer is at least 17 nucleotides, or alternatively at least 18 nucleotides, or alternatively at least 19 nucleotides, or alternatively at least 20 nucleotides, or alternatively at least 21 nucleotides, or alternatively at least 22 nucleotides, or alternatively at least 23 nucleotides, or alternatively at least 24 nucleotides, or alternatively at least 25 nucleotides, or alternatively at least 26 nucleotides, or alternatively at least 27 nucleotides, or alternatively at least 28 nucleotides, or alternatively at least 29 nucleotides, or alternatively at least 30 nucleotides, or alternatively at least 50 nucleotides, or alternatively at least 75 nucleotides, or alternatively at least 100 nucleotides.
[0072] Particularly exemplary amplification methods include rolling cycle amplification (RCA), multiple displacement amplification (MDA), loop-mediated isothermal amplification (LAMP), strand displacement amplification (SDA, see U.S. Pat. No. 5,744,311), nucleic acid sequence-based amplification (NASBA, see U.S. Pat. No. 6,025,134), quantitative real-time PCR, reverse transcription PCR (RT-PCR), real-time PCR (rtPCR), real-time reverse transcription PCR (rtRT-PCR), nested PCR, transcription-free isothermal amplification (see U.S. Pat. No. 6,033,881), repair chain reaction amplification (see WO 90 / 01069), ligase chain reaction amplification (see EP 320308), gap-filling ligase chain reaction amplification (see U.S. Pat. No. 5,427,930), coupled ligase detection and PCR (coupled ligase detection and PCR) (see U.S. Patent No. 6,027,889).
[0073] Sequencing The amplicons are sequenced, for example, using high-throughput sequencing methods known to those skilled in the art. Sequencing of a nucleic acid sequence of interest can be performed using a variety of sequencing methods known in the art, including, but not limited to, sequencing by hybridization (SBH), sequencing by ligation (SBL) (Shendure et al. (2005) Science 309:1728), quantitative incremental fluorescent nucleotide addition sequencing (QIFNAS), stepwise ligation and cleavage, fluorescence resonance energy transfer (FRET), molecular beacons, TaqMan reporter probe digestion, pyrosequencing, fluorescence in situ sequencing (FISSEQ), FISSEQ beads (U.S. Patent No. 7,425,431), wobble sequencing (PCT / US05 / 27695), multiplex sequencing (U.S. Patent Application No. 12 / 027,039, filed February 6, 2008; Porreca et al. (2007) Nat. Methods 4:931), polymerized colony (POLONY) sequencing (U.S. Patent Application Nos. 6,432,360, 6,485,944, 6,511,803, and International Publication No. PCT / US05 / 06425), nanogrid rolling circle sequencing (ROLONY) (U.S. Patent Application No. 12 / 120,541, filed May 14, 2008), allele-specific oligo ligation assays (e.g., oligo ligation assay (OLA), single template molecule OLA using ligated linear probes and rolling circle amplification (RCA) readout, single template molecule OLA using ligated padlock probes, and / or ligated circular padlock probes and rolling circle amplification (RCA) readout), and the like.High-throughput sequencing methods can also be used, for example, using platforms such as Roche454, Illumina Solexa, AB-SOLiD, Helicos, Polonator platform, Ion Torrent semiconductor sequencing technology, Pacific Biosciences single molecule real-time (SMRT) sequencing, and Oxford Nanopore Technologies nanopore-based sequencing. Various optical-based sequencing technologies are known in the art (Landegren et al. (1998) Genome Res. 8:769-76; Kwok (2000) Pharmacogenomics 1:95-100; and Shi (2001) Clin. Chem. 47:164-172). Exemplary sequencing platforms useful in the present disclosure and adaptable to the methods described herein are described in Reuter et al., High-Throughput Sequencing Technologies, Mol. Cell (2015); 58(4):586-597. This document is incorporated herein by reference in its entirety.
[0074] The amplified DNA can be sequenced by any suitable method. In particular, the amplified DNA can be sequenced using high-throughput sequencing methods such as Applied Biosystems' SOLiD sequencing technology or Illumina's Genome Analyzer. In one embodiment of the present invention, the amplified DNA can be shotgun sequenced. The number of reads can be at least 10,000, at least 1 million, at least 10 million, at least 100 million, or at least 1 billion. In another embodiment, the number of reads can be 10,000 to 100,000, or alternatively 100,000 to 1 million, or alternatively 1 million to 10 million, or alternatively 10 million to 100 million, or alternatively 100 million to 1 billion. A "read" is a length of contiguous nucleic acid sequence obtained by a sequencing reaction.
[0075] "Shotgun sequencing" refers to a method used to sequence a large amount of DNA (such as the entire genome).In this method, the DNA to be sequenced is first broken into smaller fragments that can be sequenced individually.The sequences of these fragments are then reassembled into their original order based on overlapping sequences, thus obtaining a complete sequence.The "breaking" of DNA can be carried out using a number of different techniques, including restriction enzyme digestion or mechanical shearing.The overlapping sequences are typically aligned by a suitably programmed computer.Methods and programs for shotgun sequencing of DNA libraries are well known in the art.
[0076] Particularly exemplary sequencing methods include Sanger sequencing (AB13730x1 Genome Analyzer), solid-support pyrosequencing (454 sequencing, Roche), sequencing by synthesis with reversible termination (ILLUMINA Genome Analyzer), DNA nanoball sequencing (DNBSEQ, MGI), sequencing by ligation (ABI SOLID), or sequencing by synthesis with virtual terminators (HELI®). Other next-generation sequencing techniques for use in the methods of the present disclosure include massively parallel signature sequencing (MPSS), Polony sequencing, Ion Torrent semiconductor sequencing, DNA nanoball sequencing, Heliscope single molecule sequencing, single molecule real-time (SMRT) sequencing, Pacbio sequencing, and nanopore DNA sequencing.
[0077] It should be understood that the described embodiments of the present invention are merely illustrative of some of the applications of the principles of the present invention. Numerous modifications may be made by those skilled in the art based on the teachings presented herein without departing from the true spirit and scope of the present invention. The contents of all references, patents, and published patent applications cited throughout this application are hereby incorporated by reference in their entirety for all purposes.
[0078] The following examples are given as representative of the invention, and should not be construed as limiting the scope of the invention, as these and other equivalent embodiments will be apparent in light of the disclosure, drawings, and appended claims. [Example]
[0079] Plasmid construction Expression constructs for the DddA toxin domain (DddAtox), DddA11, and BadTF3 were obtained from Genescript's gene synthesis services.
[0080] To generate a pETDuet-1 based expression construct of DddAtox, the DddAtox coding sequence: ATGGGCAGCAGCCATCACCATCATCACCACAGCCAGGATCCGGGTAGCTATGCGCTGGGTCCGTATCAGATCTCTGCTCCGCAGCTGCCGGCATATAACGGTCAGACTGTTGGTACTTTCTATTATGTTAACGATGCTGGCGGTTTAGAAAGCAAAGTTTTCAGCTCTGGGTGGTCCGACCCCGTATCCGAACTATGCTAACGCTGGTCACGTTGAAGGTCAGTCTGCTCTGTTCATGCGTGAT AACGGTATCTCTGAAGGTCTGGTTTTCCATAACAACCCGGAAGGTACCTGTGGTTTTTGTGTTAACATGACCGAAACCCTGCTGCCGGAAAACGCTAAAATGACCGTTGTTCCGCCGGAAGGTGCGATTCCGGTTAAACGTGGTGCTACCGGTGAAACCAAAGTTTTCACCGGTAACTCTAACTCTCCGAAATCTCCGACCAAAGGTGGTTGCTAA (SEQ ID NO: 1), in which the protein sequence to be translated is MGSSHHHHHHSQDPGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNSNSPKSPTKGGC) (SEQ ID NO: 2) is synthesized and cloned into MCS-1 (NcoI and HindIII sites, the N-terminal hexahistidine tag is retained).
[0081] Immunity protein DddAI coding sequence: ATGTATGCGGATGACTTTGACGGGGAAATTGAGATTGATGAAGTTGATAGCCTAGTTGAGTTTCTGAGCCGTCGTCCGGCGTTCGATGCGAACAACTTCGTTCTGACCTTCGAAGAAAGCGGCTTCCCGCAGCTGAACATCTTCGCGAAAAACGATATCGCGGTTGTTTACTACATGGATATCGGCGAAAACTTCGTTAGCAAAGGCAACAGCGCGAGCGGCGGCACCGAAAAATTCTACGAAAACAAACTGGGCGGCGAAGTTGATCTGAGCAAAGATTGCGTTGTTAGCAAAGAACAGATGATCGAAGCGGCGAAACAGTTCTTCGCGACCAAACAGCGTCCGGAACAGCTGACCTGGAGCGAACTGTAA (SEQ ID NO: 3), wherein the translated protein sequence is MYADDFDGEIEIDEVDSLVEFLSRRPAFDANNFVLTFEESGFPQLNIFAKNDIAVVYYMDIGENFVSKGNSASGGTEKFYENKLGGEVDLSKDCVVSKEQMIEAAKQFFATKQRPEQLTWSEL) (SEQ ID NO: 4) is synthesized and cloned into MCS-2 (NdeI and XhoI sites, the C-terminal S tag is removed).
[0082] To generate a pETDuet-1 based expression construct of DddA11, the DddA11 coding sequence: ATGGGCAGCAGCCATCACCATCATCACCACAGCCAGGATCCGGGTAGCTATGCGCTGGGTCCGTATCAGATCTCTGCTCCGCAGCTGCCGGCATATAACGGTCAGACTGTTGGTACTTTCTATTATGTTAACGATGCTGGCGGTTTAGAAAGCAAAGTTTTCATCTCTGGTGGTCCGACCCCGTATCCGAACTATGTTAGCGCTGGTCACGTTGAAGGTCAGTCTGCTCTGTTCATGCGTGATAACGGTATCTCTGAAGGTCTGGTTTTCCATAACAACCCGAAAGGTACCTGTGGTTTTTGTGTTAACATGATCGAAACCCTGCTGCCGGAAAACGCTAAAATGACCGTTGTTCCGCCGGAAGGTGCGATTCCGGTTAAACGTGGTGCTACCGGTGAAACCAAAGTTTTCATCGGTAACTCTAACTCTCCGAAATCTCCGACCAAAGGTGGTTGCTAA (SEQ ID NO: 5) GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFISGGPTPYPNYVSAGHVEGQSALFMRDNGISEGLVFHNNPKGTCGFCVNMIETLLPENAKMTVVPPEGAIPVKRGATGETKVFIGNSNSPKSPTKGGC) (SEQ ID NO: 6) was synthesized and cloned into MCS-1 (BamHI and HindIII sites, the N-terminal hexahistidine tag is retained) to ligate the immunity protein DddAI coding sequence: ATGTATGCGGATGACTTTGACGGGGAAATTGAGATTGATGAAGTTGATAGCCTAGTTGAGTTTCTGAGCCGTCGTCCGGCGTTCGATGCGAACAACTTCGTTCTGACCTTCGAAGAAAGCGGCTTCCCGCAGCTGAACATCTTCGCGAAAAACGATATCGCGGTTGTTTACTACATGGATATCGGCGAAAACTTCGTTAGCAAAGGCAACAGCGCGAGCGGCGGCACCGAAAAATTCTACGAAAACAAACTGGGCGGCGAAGTTGATCTGAGCAAAGATTGCGTTGTTAGCAAAGAACAGATGATCGAAGCGGCGAAACAGTTCTTCGCGACCAAACAGCGTCCGGAACAGCTGACCTGGAGCGAACTGTAA (SEQ ID NO: 7), wherein the translated protein sequence is MYADDFDGEIEIDEVDSLVEFLSRRPAFDANNFVLTFEESGFPQLNIFAKNDIAVVYYMDIGENFVSKGNSASGGTEKFYENKLGGEVDLSKDCVVSKEQMIEAAKQFFATKQRPEQLTWSEL) (SEQ ID NO: 8) is synthesized and cloned into MCS-2 (NdeI and XhoI sites, the C-terminal S tag is removed).
[0083] To generate a pETDuet-1-based expression construct of BadTF3, the BadTF3 coding sequence: ATGGGCAGCAGCCATCACCATCATCACCACAGCCAGGATCCTGGTTGGAAATTTTCTAACGGTAAACGCCGTCCGCCGCACAAAGCAACGGTAACTGTGACCGATAAAAACGGTGTCGTTAAACACAAAAGCAACCTGGTTTCTGGCAACATGACTGAAGCCGAAAAGAAACTGGGCTTCCCGAACAACTCCCTGGCGACCCACACCGAAAACCGTGCTACCCGCCTGATCGATCTGAACCAAGGTGATACTATGCTGATCGAGGGCCAATACCGTCCGTGTCCACGTTGTAAAGGTGCAATGCGCGTGAAAGCGGAGGAATCCGGTGCGAAAGTGATCTACACCTGGCCAGAAGATGGTGACCTGAAAAAACGTGAATGGGAAGGCACTCCGTGCGACAAAAAATAA (SEQ ID NO: 9), wherein the translated protein sequence is MGSSHHHHHHSQDPGWKFSNGKRRPPHKATVTVTDKNGVVKHKSNLVSGNMTEAEKKLGFPNNSLATHTENRATRLIDLNQGDTMLIEGQYRPCPRCKGAMRVKAEESGAKVIYTWPEDGDLKKREWEGTPCDKK) (SEQ ID NO: 10) is synthesized and cloned into MCS-1 (NcoI and HindIII sites, the N-terminal hexahistidine tag is retained).
[0084] Immunity protein BadTF3I coding sequence: ATGACCAAATCTAAAATGCTGAGCAACATCGTCATCCAGGAGGTCAAATTTGCGATCGAAGATTACTGCGCTATTCTGAGCTTCGCTTCTGACTCTTATGAAGTGCCGGAGCAGTATTTTATCATTACCCGTTCTACCACCGAACGTTCTGGCGGTATTCCGGAGGGCGACATCTACCTGGAATCTAACCTGTTTCTGGATTTTAACCCGTACGGCCTGAGCGGTTACCTGCTGTCTGAGCCGAACTGCGTAGATCTGCTGATCGAACCGAACAACTACGTTCGTCTGCGTCTGATCGAAAAAATCGATATCCTGGAAGTGGAAAACCACCTGAAATTTCTGTTCGACAACTAA (SEQ ID NO: 11) The pETDuet-1::dddAtox+dddAI vector and pETDuet-1::badTF3tox+badTF3I were transformed into Escherichia coli (E. coli) strains DH5α and BL21 and stored at -20°C. Figure 2 shows the vector map of pETDuet-1::dddAtox+dddAI. The inserted gene is driven by two independent lac operators. Figure 3 shows the vector map of pETDuet-1::badTF3tox+badTF3I. The inserted gene is driven by two independent lac operators. Figure 4 shows the vector map of pETDuet-1::dddA11+dddAI. The inserted gene is driven by two independent lac operators. [Example]
[0085] Bacterial strains and culture conditions E. coli strains were grown in lysogeny broth (LB) or LB medium solidified with agar (LBA, 1.5% weight / volume) at 37°C. When necessary, the medium was supplemented with ampicillin (100 μg per ml) or IPTG (0.5 mM). E. coli strains DH5α and BL21 were used for plasmid maintenance and protein expression, respectively. [Example]
[0086] Purification of DddAtox Purification of the DddAtox protein has been previously reported (Beverly et al., 2020, Nature, Vol. 583(7817):631–637, doi:10.1038 / s41586-020-2477-4). Briefly, to purify His-tagged DddAtox complexed with DddAI, E. coli BL21 (pETDuet-1::dddAtox+dddAI) was used to inoculate 2 L of LB broth at a 1:100 dilution and grown overnight. After growing the culture to an approximate OD600 of 0.6, 0.5 mM isopropyl β-D-1-thiogalactopyranoside (IPTG) was added and incubated at 18 °C for 16 h with shaking. The bacterial cell pellet was collected by centrifugation at 4000 g for 30 minutes and then resuspended in 50 ml of lysis buffer (50 mM Tris-HCl pH 8.0, 500 mM NaCl, 10 mM imidazole, 1 mg / mL lysozyme, and protease inhibitor cocktail). The bacterial cell pellet was then lysed by sonication (5 pulses, 10 seconds each), and the supernatant was separated from the debris by centrifugation at 25,000 g for 30 minutes. The His-tagged DddAtox-DddAI complex was purified from the supernatant using a nickel column. DddAtox-DddAI was eluted with elution buffer (50 mM Tris-HCl pH 7.5, 300 mM imidazole, 500 mM NaCl, 30 mM imidazole, 1 mM DTT). The eluted DddAtox-DddAI complex was denatured for 16 hours at 4°C by adding 50 ml of 8 M urea denaturing buffer (50 mM Tris-HCl pH 7.5, 300 mM imidazole, 500 mM NaCl, and 1 mM DTT). The denatured protein in the 8 M urea denaturing buffer was loaded onto the nickel column again. To remove any residual DddAI, the column was washed with 50 ml of 8 M urea denaturing buffer. The column was washed successively with 25 ml of denaturing buffer containing decreasing concentrations of urea (6 M, 4 M, 2 M, 1 M), and finally with urea-free wash buffer. The DddAtox bound to the column was then eluted with 5 ml of elution buffer.The eluted DddAtox was subjected to size-exclusion chromatography using fast protein liquid chromatography (FPLC) and gel filtration on a Superdex200 column (GE Healthcare) in sizing buffer (20 mM Tris-HCl pH 7.5, 200 mM NaCl, 1 mM DTT, 5% (wt / vol) glycerol). The purity of the eluted DddAtox was assessed on an SDS-PAGE gel stained with Coomassie blue, and the protein was then stored at -80°C. [Example]
[0087] Purification of DddA11 Purification of the DddA11 protein was the same as for DddAtoxin, as previously reported (Beverly et al., 2020, Nature, Vol. 583(7817):631-637, doi:10.1038 / s41586-020-2477-4). [Example]
[0088] Purification of BadTF3 To purify his-tagged BadTF3 complexed with BadTF3I, E. coli BL21 (pETDuet-1::BadTF3+BadTF3I) was used to inoculate 2 L of LB broth at a 1:100 dilution and grown overnight. The culture was grown to an approximate OD600 of 0.6, after which 0.5 mM IPTG was added and incubated at 18°C for 16 hours with shaking. The bacterial cell pellet was collected by centrifugation at 4000g for 30 minutes and then resuspended in 50 ml of lysis buffer (50 mM Tris-HCl pH 8.0, 500 mM NaCl, 10 mM imidazole, 1 mg / mL lysozyme, and protease inhibitor cocktail). The bacterial cell pellet was then lysed by sonication (5 pulses, 10 seconds each), and the supernatant was separated from the debris by centrifugation at 25,000g for 30 minutes. The His-tagged BadTF3-BadTF3I complex was purified from the supernatant using a nickel column. BadTF3-BadTF3I was eluted with elution buffer (50 mM Tris-HCl pH 7.5, 300 mM imidazole, 500 mM NaCl, 30 mM imidazole, 1 mM DTT). The eluted BadTF3-BadTF3I complex was denatured at 4°C for 16 hours by adding 50 ml of 8 M urea denaturing buffer (50 mM Tris-HCl pH 7.5, 300 mM imidazole, 500 mM NaCl, and 1 mM DTT). The denatured protein in the 8 M urea denaturing buffer was loaded onto the nickel column again. The column was washed with 50 ml of 8 M urea denaturing buffer to remove any residual BadTF3I. The column was washed successively with 25 ml of denaturing buffer containing decreasing concentrations of urea (6 M, 4 M, 2 M, 1 M), followed by a final wash with urea-free wash buffer. The bound BadTF3 was then eluted with 5 ml of elution buffer. The eluted BadTF3 was subjected to size-exclusion chromatography using fast protein liquid chromatography (FPLC) and gel filtration on a Superdex 200 column (GE Healthcare) in sizing buffer (20 mM Tris-HCl pH 7.5, 200 mM NaCl, 1 mM DTT, 5% (wt / vol) glycerol).The purity of the eluted BadTF3 was assessed on a Coomassie blue stained SDS-PAGE gel, and the protein was then stored at -80°C.
[0089] Figure 5 shows Coomassie blue stained SDS-PAGE gels of DddAtox-DddAI, DddAtox, DddA11-DddAI, DddA11, BadTF3-BadTF3I, and BadTF3, respectively. [Example]
[0090] DNA deamination activity assay Double-stranded DNA deamination activity was assessed using lambda DNA or genomic DNA extracted from Drosophila S2, K562, or GM12878 cell lines. Reactions were performed in 10 μL of deamination buffer consisting of 20 mM Tris-HCl pH 7.4, 100 mM NaCl, 1 mM DTT, 50 ng DNA substrate, and deaminase (20 μM unless otherwise noted). Reactions were incubated at 37°C for 1 hour or for the indicated time course, followed by DNA purification using the Zymo DNA Clean & Concentrator-5 Kit. Purified DNA was subjected to library preparation using the Illumina TruePrep DNA Library Prep Kit V2 (Vazyme) as directed, except that the polymerase mix was replaced with 1x Q5U PCR Master Mix + Bst3.0 Polymerase (0.08 U / μl). The uracil conversion rate at each cytosine site was calculated and averaged to assess the deamination activity in each possible sequence context.
[0091] Figure 6 shows the cytosine-to-uracil conversion efficiency in double-stranded DNA treated with DddAtoxin, DddA11 protein, or BadTF3 for 1 hour. Each box represents a single cytosine at four possible upstream and four possible downstream nucleotide contexts. The intensity of the red color indicates the degree of conversion. The left side shows the average conversion efficiency at cytosine sites in bare genomic DNA in vitro (Drosophila genome for DddA and DddA11, lambda DNA for BadTF3). The right side shows the average conversion efficiency at cytosine sites in K562 genomic DNA extracted from live-cell experiments. Purified DddA specifically deaminates cytosines in TC or CC contexts (TCC is converted to TUC, and then U is considered as T to initiate the conversion of UC to UU). The enzymes DddA11 and BadTF3 deaminate cytosines in TC, CC, and AC contexts, and to a lesser extent, GC contexts. [Example]
[0092] Nuclear preparation To prepare nuclei, cells were centrifuged at 450 g for 5 minutes, then washed with an equal volume of cold 1x PBS and centrifuged at 500 g for 5 minutes at 4°C. A total of 30,000 cells were permeabilized using cold permeabilization buffer (10 mM Tris-HCl, pH 7.4, 10 mM NaCl, 3 mM MgCl2, 0.1% IGEPAL CA-630, 0.1% Tween-20, 0.1% digitonin). Immediately after permeabilization, nuclei were centrifuged at 550 g for 5 minutes at 4°C. After centrifugation, the supernatant was carefully removed from the pellet. [Example]
[0093] Tagmentation reaction It should be understood that the tagmentation described below can be performed before or after treating the DNA with dsDNA deaminase. 2xTD buffer (20 mM TAPS pH 8.5, 10 mM MgCl2, 20% DMF) was prepared. The nuclear pellet was immediately resuspended in transposase reaction mix (12.5 μL 2xTD buffer, 2 μL transposase (Vazyme, 1.25 μM), and 10 μL PBS with 0.1% digitonin). The transposition reaction was carried out at 37°C for 30 minutes in a thermomixer at 800 rpm. Then, 100 μL of ice-cold RSB was added to the mixer and centrifuged at 550 g for 5 minutes at 4°C. After centrifugation, the supernatant was carefully removed from the pellet. [Example]
[0094] DNA deamination reaction Prepare 2x DRB buffer (20 mM Tris-HCl pH 7.5, 20 mM NaCl, 2 mM DTT). The nuclear pellet was immediately resuspended in the deaminase reaction mix. For DddAtox, the deaminase reaction mix contained 12 μL of DRB and 18 μL of DddA enzyme (50 μM, stored in sizing buffer). For DddA11, the deaminase reaction mix contained 8 μL of DRB and 12 μL of DddA11 enzyme. For BadTF3, the deaminase reaction mix contained 5 μL of DRB and 5 μL of BadTF3 enzyme (50 μM, stored in sizing buffer). The nuclear mix was incubated at 37°C for 20 minutes. Immediately after the reaction, the DNA was purified using the Zymo DNA Clean & Concentrator-5 kit. [Example]
[0095] Library amplification After DNA purification, the library fragments were amplified using 1x Q5U PCR Master Mix, Bst3.0 polymerase (0.08 U / µL), and 1.25 µM Nextera PCR primers (forward and reverse) under the following PCR conditions: 65°C for 5 min, 80°C for 12 min, 98°C for 2 min, and thermal cycles of 98°C for 15 s, 60°C for 30 s, and 72°C for 1 min. It is recommended to monitor the PCR reaction using qPCR to stop amplification before saturation and avoid GC and size bias in the PCR. After five cycles of complete library preamplification, an aliquot of the PCR reaction was collected for qPCR. A total of 20 cycles of qPCR were performed to determine the number of additional cycles required for the remaining PCR reactions. Generally, a total of 10–12 cycles of amplification yields high-quality libraries. The library was purified using the Zymo Select-a-Size DNA Clean & Concentrator Kit, and DNA with fragment sizes greater than 200 bp was collected. The size-selected library is ready for sequencing. [Example]
[0096] Detection of transcription factor footprints Figure 7A illustrates the identification of transcription factor CTCF footprints in the human K562 genome using the methods described herein. Isolated nuclei were treated individually with DddA, DddA11, and BadTF3. The Y-axis indicates the conversion rate of each cytosine site. Purple bars indicate CTCF binding motifs. Figure 7B shows the average conversion rate across the merged CTCF binding motifs, with the footprint observed in the center. Figure 7C illustrates a proportional Venn diagram showing the overlap between CTCF binding sites identified by the dsDNA deaminase, ChIP-seq, and DNase-seq methods described herein. A total of 31,186 CTCF binding sites detected by the dsDNA deaminase method (38,281 total) are consistent with the CTCF ChIP-seq method (36,110 total). In contrast, of the 22,085 binding sites identified by DNase-seq, only 19,989 overlapped with those identified by ChIP-seq. These data demonstrate that the dsDNA deaminase method is superior to ChIP-seq and can robustly identify a much larger number of TF binding sites.
[0097] The method described herein determines TF binding ratios by footprinting TFs within single DNA molecules. Figure 8A illustrates a schematic of the analysis of TF binding patterns at the single-molecule level using data obtained by the method described herein. For each sequence read, unconverted sites are interpreted as TF-bound regions. Alternatively, converted sites are interpreted as accessible regions. Single reads can be sorted according to the occupancy patterns of multiple genomic features, such as TFBS clusters. In Figure 8B, each line represents a sequencing DNA read, accumulating all reads located on chromosome 1: 26321500-26321900. Each black dot represents a converted cytosine, and each gray dot represents an unconverted cytosine. If the DNA were occupied by a TF, in this case, for example, CTCF, binding of the TF would prevent cytosine deamination at the specific binding site, but upstream and downstream flanking cytosines would not be protected from deamination. On the other hand, if the DNA were unoccupied by a TF, all cytosines would be accessible to dsDNA deaminase, and subsequent deamination would occur. Therefore, TF binding was determined for each DNA molecule, and the TF binding ratio was calculated from this. The data show that 88.37% of the DNA at the CTCF binding site is occupied by CTCF, while 11.63% is unoccupied. Figure 8C illustrates data demonstrating that the dsDNA deaminase method described herein can simultaneously detect three TF binding sites and their relative occupancies in a single promoter. Each dot in the raw reads represents a cytosine conversion. In this case, the footprint closest to the transcription start site has the highest binding rate (very few cytosine-to-uracil conversions occur in this region), and the most distal footprint has the lowest binding rate of the three. [Example]
[0098] Detection of individual transcription factor footprints in single cells Figure 9A shows a schematic diagram of the use of the method described herein to detect individual TF footprints in chromatin DNA from single cells. Heterogeneous tissues or samples were first dissociated into single-cell suspensions. After cell lysis to isolate nuclei, the deaminase enzyme DddA was added. Universal adapters were then added to open regions by Tn5 transposition, leading to enrichment. Single-cell samples were obtained by FACS sorting. After gap filling with the aid of BST, the library was amplified by Q5U and sequenced on an Illumina sequencer. Figure 9B shows the DNA fragment distribution of single cells after PCR amplification. Open regions and nucleosome patterns could be clearly identified. Figure 9C shows the cell type classification results of single-cell data from K562, GM12878, and Hek293T cell lines. The three cell types were successfully clustered. The number of each cell type used is indicated. Figure 9D shows a comparison of bulk and single-cell data obtained from the method described herein, displayed in IGV software. In both the K562 and GM12878 cell lines, the signals from the single cell data correlate well with those from the bulk cell data. [Example]
[0099] array The amino acid sequence of native DddAtox is as follows: GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNSNSPKSPTKGGC (SEQ ID NO: 15)
[0100] The amino acid sequence of purified DddAtox with an N-terminal His tag is as follows: MGSSHHHHHHSQDPGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNSNSPKSPTKGGC (SEQ ID NO: 2)
[0101] The Genescript synthesized nucleotide sequence of DddAtox with an N-terminal His tag (optimized for bacterial expression) is as follows: ATGGGCAGCAGCCATCACCATCATCACCACAGCCAGGATCCGGGTAGCTATGCGCTGGGTCCGTATCAGATCTCTGCTCCGCAGCTGCCGGCATATAACGGTCAGACTGTTGGTACTTTCTATTATGTTAACGATGCTGGCGGTTTAGAAAGCAAAGTTTTCAGCTCTGGGTGGTCCGACCCCGTATCCGAACTATGCTAACGCTGGTCACGTTGAAGGTCAGTCTGCTCTGTT CATGCGTGATAACGGTATTCCTGAAGGTCTGGTTTTCCATAACAACCCGGAAGGTACCTGTGGTTTTTGTGTTAACATGACCGAAACCCTGCTGCCGGAAAACGCTAAAATGACCGTTGTTCCGCCGGAAGGTGCGATTCCGGTTAAACGTGGTGCTACCGGTGAAACCAAAGTTTTCACCGGTAACTCTAACTCTCCGAAATCTCCGACCAAAGGTGGTTGCTAA (SEQ ID NO: 1)
[0102] The amino acid sequence of native DddAI (the same amino acid sequence used for protein purification) is: MYADDFDGEIEIDEVDSLVEFLSRRPAFDANNFVLTFEESGFPQLNIFAKNDIAVVYYMDIGENFVSKGNSASGGTEKFYENKLGGEVDLSKDCVVSKEQMIEAAKQFFATKQRPEQLTWSEL (SEQ ID NOs: 4, 8).
[0103] The Genescript synthesized nucleotide sequence of DddAI (optimized for bacterial expression) is: ATGTATGCGGATGACTTTGACGGGGAAATTGAGATTGATGAAGTTGATAGCCTAGTTGAGTTTCTGAGCCGTCGTCCGGCGTTCGATGCGAACAACTTCGTTCTGACCTTCGAAGAAAGCGGCTTCCCGCAGCTGAACATCTTCGCGAAAAACGATATCGCGGTTGTTTACTACATGGATATCGGCGAAAACTTCGTTAGCAAAGGCAACAGCGCGAGCGGCGGCACCGAAAAATTCTACGAAAACAAACTGGGCGGCGAAGTTGATCTGAGCAAAGATTGCGTTGTTAGCAAAGAACAGATGATCGAAGCGGCGAAACAGTTCTTCGCGACCAAACAGCGTCCGGAACAGCTGACCTGGAGCGAACTGTAA (SEQ ID NOs: 3, 7)
[0104] The amino acid sequence of purified DddA11 with an N-terminal His tag is as follows: GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFISGGPTPYPNYVSAGHVEGQSALFMRDNGISEGLVFHNNPKGTCGFCVNMIETLLPENAKMTVVPPEGAIPVKRGATGETKVFIGNSNSPKSPTKGGC (SEQ ID NOs: 6, 13)
[0105] The Genescript synthetic nucleotide sequence of DddA11 (optimized for bacterial expression) is as follows: ATGGGCAGCAGCCATCACCATCATCACCACAGCCAGGATCCGGGTAGCTATGCGCTGGGTCCGTATCAGATCTCTGCTCCGCAGCTGCCGGCATATAACGGTCAGACTGTTGGTACTTTCTATTATGTTAACGATGCTGGCGGTTTAGAAAGCAAAGTTTTCATCTCTGGTGGTCCGACCCCGTATCCGAACTATGTTAGCGCTGGTCACGTTGAAGGTCAGTCTGCTCTGTT CATGCGTGATAACGGTATCTCTGAAGGTCTGGTTTTCCATAACAACCCGAAAGGTACCTGTGGTTTTTGTGTTAACATGATCGAAACCCTGCTGCCGGAAAACGCTAAAATGACCGTTGTTCCGCCGGAAGGTGCGATTCCGGTTAAACGTGGTGCTACCGGTGAAACCAAAGTTTTCATCGGTAACTCTAACTCTCCGAAATCTCCGACCAAAGGTGGTTGCTAA (SEQ ID NO: 5)
[0106] The Genescript synthetic protein sequence (optimized for bacterial expression) is as follows: GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFISGGPTPYPNYVSAGHVEGQSALFMRDNGISEGLVFHNNPKGTCGFCVNMIETLLPENAKMTVVPPEGAIPVKRGATGETKVFIGNSNSPKSPTKGGC (SEQ ID NO: 6)
[0107] The native amino acid sequence of BadTF3 is as follows: GWKFSNGKRRPPHKATVTVTDKNGVVKHKSNLVSGNMTEAEKKLGFPNNSLATHTENRATRLIDLNQGDTMLIEGQYRPCPRCKGAMRVKAEESGAKVIYTWPEDGDLKKREWEGTPCDKK (SEQ ID NO: 14)
[0108] The amino acid sequence of purified BadTF3 with an N-terminal His tag is as follows: MGSSHHHHHHSQDPGWKFSNGKRRPPHKATVTVTDKNGVVKHKSNLVSGNMTEAEKKLGFPNNSLATHTENRATRLIDLNQGDTMLIEGQYRPCPRCKGAMRVKAEESGAKVIYTWPEDGDLKKREWEGTPCDKK (SEQ ID NO: 10).
[0109] The Genescript synthesized nucleotide sequence of BadTF3 (optimized for bacterial expression) is: ATGGGCAGCAGCCATCACCATCATCACCACAGCCAGGATCCTGGTTGGAAATTTTCTAACGGTAAACGCCGTCCGCCGCACAAAGCAACGGTAACTGTGACCGATAAAAACGGTGTCGTTAAACACAAAAGCAACCTGGTTTCTGGCAACATGACTGAAGCCGAAAAGAAACTGGGCTTCCCGAACAACTCCCTGGCGACCCACACCGAAAACCGTGCTACCCGCCTGATCGATCTGAACCAAGGTGATACTATGCTGATCGAGGGCCAATACCGTCCGTGTCCACGTTGTAAAGGTGCAATGCGCGTGAAAGCGGAGGAATCCGGTGCGAAAGTGATCTACACCTGGCCAGAAGATGGTGACCTGAAAAAACGTGAATGGGAAGGCACTCCGTGCGACAAAAAATAA (SEQ ID NO: 9)
[0110] The amino acid sequence of native BadTF3I (the same amino acid sequence used for protein purification) is as follows: MTKSKMLSNIVIQEVKFAIEDYCAILSFASDSYEVPEQYFIITRSTTERSGGIPEGDIYLESNLFLDFNPYGLSGYLLSEPNCVDLLIEPNNYVRLRLIEKIDILEVENHLKFLFDN (SEQ ID NO: 12)
[0111] The Genescript synthesized nucleotide sequence of BadTF3I (optimized for bacterial expression) is: ATGACCAAATCTAAAATGCTGAGCAACATCGTCATCCAGGAGGTCAAATTTGCGATCGAAGATTACTGCGCTATTCTGAGCTTCGCTTCTGACTCTTATGAAGTGCCGGAGCAGTATTTTATCATTACCCGTTCTACCACCGAACGTTCTGGCGGTATTCCGGAGGGCGACATCTACCTGGAATCTAACCTGTTTCTGGATTTTAACCCGTACGGCCTGAGCGGTTACCTGCTGTCTGAGCCGAACTGCGTAGATCTGCTGATCGAACCGAACAACTACGTTCGTCTGCGTCTGATCGAAAAAATCGATATCCTGGAAGTGGAAAACCACCTGAAATTTCTGTTCGACAACTAA (SEQ ID NO: 11) [Example]
[0112] kit The materials and reagents required for the disclosed method for determining TF footprints using dsDNA deaminase can be assembled into a kit. The disclosed kits will generally include at least the dsDNA deaminase, transposase, nuclease, degradative enzyme, nucleotides, DNA polymerase, amplification primers and reagents, sequencing primers and reagents, and / or DNA enrichment reagents described herein, which can be used to practice the claimed method. In preferred embodiments, the kit will also include instructions for treating chromatin DNA with dsDNA deaminase and processing and amplifying the treated chromatin DNA. In either case, the kit will preferably have separate containers for each individual reagent, enzyme, or reactant. Each agent will generally be suitably aliquoted into its own container. The container means of the kit will generally include at least one vial or test tube. Flasks, bottles, and other container means for containing and aliquoting reagents can also be used. The individual containers of the kit will preferably be maintained in close confinement for commercial sale. Suitable larger containers may include injection or blow molded plastic containers into which the desired vials are held. Preferably, the kit is accompanied by instructions for use.
[0113] Embodiment The present disclosure provides a method for determining transcription factor binding sites on genomic double-stranded (ds)DNA of a cell, such as a eukaryotic cell, comprising contacting the genomic dsDNA with a dsDNA deaminase under conditions that convert cytosines in the genomic dsDNA to uracil, thereby producing processed genomic dsDNA, and identifying unconverted cytosines on the processed genomic dsDNA as transcription factor binding sites. According to one embodiment, the genomic double-stranded (ds)DNA is a gene. According to one embodiment, the method includes identifying one or more unconverted cytosines as transcription factor binding sites. According to one embodiment, the method includes identifying a plurality of unconverted cytosines as transcription factor binding sites. According to one embodiment, the method includes identifying a plurality of unconverted cytosines as two or more transcription factor binding sites. According to one embodiment, the genomic dsDNA includes a plurality of genes, and the method further includes identifying a plurality of unconverted cytosines as a plurality of transcription factor binding sites. According to one embodiment, the pattern of unconverted cytosines on the processed genomic DNA is correlated with the DNA binding domain of a transcription factor to identify one or more transcription factor binding sites. According to one embodiment, the dsDNA deaminase is DddA, BadTF3, or DddA11, or a variant, mutant, derivative, or modified version thereof. According to one embodiment, the dsDNA deaminase is an enzyme capable of converting cytosine to uracil on double-stranded (ds) DNA. According to one embodiment, the processed genomic DNA is optionally amplified and sequenced to determine the positions of unconverted cytosines and uracils. According to one embodiment, the processed genomic DNA is optionally amplified and sequenced to determine the positions of unconverted cytosines and uracils, which are compared with the DNA binding patterns of the transcription factors to identify transcription factor binding sites. According to one embodiment, the processed genomic DNA is optionally amplified and sequenced to determine the positions of unconverted cytosines and uracils, which are compared with the DNA binding patterns of the transcription factors to identify transcription factor binding sites and associated transcription factors. According to one embodiment, the genomic dsDNA is a single DNA molecule derived from a single cell, and the method comprises identifying a plurality of unconverted cytosines on the single DNA molecule as a plurality of transcription factor binding sites on the single DNA molecule.According to one embodiment, genomic double-stranded (ds)DNA is processed into multiple DNA molecules, which are analyzed or quantified for relative transcription factor binding ratios. According to one embodiment, the processed genomic dsDNA is subjected to whole genome sequencing or targeted amplicon sequencing. According to one embodiment, regions of cytosine-to-uracil conversion flank regions of unconverted cytosine, identifying transcription factor footprints on open regions of genomic dsDNA. According to one embodiment, one or more open regions of the processed genomic dsDNA are enriched by tagmentation and amplification. According to one embodiment, one or more open regions of the processed genomic dsDNA are enriched by nuclease digestion and amplification. According to one embodiment, the genomic dsDNA is obtained from multiple cells of the same cell type. According to one embodiment, the processed genomic dsDNA is PCR amplified. According to one embodiment, the processed genomic dsDNA is processed into fragments for sequencing. According to one embodiment, the genomic dsDNA is treated with dsDNA deaminase within cells or in nuclei isolated from cells. According to one embodiment, genomic double-stranded (ds) DNA is processed to enrich for open chromatin DNA. According to one embodiment, genomic double-stranded (ds) DNA is processed to enrich for open chromatin DNA before being treated with dsDNA deaminase. According to one embodiment, genomic double-stranded (ds) DNA is treated with dsDNA deaminase, and then the processed genomic dsDNA is processed to enrich for open chromatin DNA.
[0114] equivalent Other embodiments will be apparent to those skilled in the art. It should be understood that the foregoing description is provided for clarity only and is merely exemplary. The spirit and scope of the present invention is not limited to the above examples, but is encompassed by the claims. All publications, patents, and patent applications cited above are incorporated herein by reference in their entirety for all purposes to the same extent as if each individual publication or patent application was specifically and specifically indicated to be incorporated by reference.
Claims
1. 1. A method for determining transcription factor binding sites on genomic double-stranded (ds) DNA of a eukaryotic cell, comprising: contacting the genomic dsDNA with a dsDNA deaminase under conditions that convert cytosines in the genomic dsDNA to uracils, thereby producing treated genomic dsDNA; and Identifying unconverted cytosines on the processed genomic dsDNA as transcription factor binding sites. A method comprising:
2. 2. The method of claim 1, wherein the genomic double-stranded (ds) DNA is a gene.
3. 10. The method of claim 1, comprising identifying one or more unconverted cytosines as transcription factor binding sites.
4. 2. The method of claim 1, comprising identifying a plurality of unconverted cytosines as transcription factor binding sites.
5. 10. The method of claim 1, comprising identifying a plurality of unconverted cytosines as two or more transcription factor binding sites.
6. 2. The method of claim 1, wherein the genomic dsDNA comprises a plurality of genes, further comprising identifying a plurality of unconverted cytosines as a plurality of transcription factor binding sites.
7. 2. The method of claim 1, wherein the pattern of unconverted cytosines on the treated genomic DNA is correlated with DNA binding domains of transcription factors to identify one or more transcription factor binding sites.
8. 2. The method of claim 1, wherein the dsDNA deaminase is DddA, BadTF3, or DddA11, or a variant, mutant, derivative, or modified form thereof.
9. 2. The method of claim 1, wherein the dsDNA deaminase is an enzyme capable of converting cytosine to uracil on double-stranded (ds) DNA.
10. 10. The method of claim 1, wherein the treated genomic DNA is optionally amplified and sequenced to determine the positions of unconverted cytosines and uracils.
11. 2. The method of claim 1, wherein the treated genomic DNA is optionally amplified and sequenced to determine the locations of unconverted cytosines and uracils, which are compared with DNA binding patterns of transcription factors to identify transcription factor binding sites.
12. 2. The method of claim 1, wherein the treated genomic DNA is optionally amplified and sequenced to determine the locations of unconverted cytosines and uracils, which are compared with DNA binding patterns of transcription factors to identify transcription factor binding sites and associated transcription factors.
13. The method of claim 1, wherein the genomic dsDNA is a single DNA molecule derived from a single cell, and further comprising identifying multiple unconverted cytosines on the single DNA molecule as multiple transcription factor binding sites on the single DNA molecule.
14. 2. The method of claim 1, wherein the genomic double-stranded (ds) DNA is processed into multiple DNA molecules, which are analyzed or quantified for relative binding ratios of transcription factors.
15. 10. The method of claim 1, wherein the processed genomic dsDNA is subjected to whole genome sequencing or targeted amplicon sequencing.
16. 2. The method of claim 1, wherein regions of cytosine to uracil conversion flank regions of unconverted cytosine, identifying a transcription factor footprint on the open region of the genomic dsDNA.
17. 10. The method of claim 1, wherein one or more open regions of the processed genomic dsDNA are enriched by tagmentation and amplification.
18. 10. The method of claim 1, wherein one or more open regions of the processed genomic dsDNA are enriched by nuclease digestion and amplification.
19. The method of claim 1 , wherein the genomic dsDNA is obtained from multiple cells of the same cell type.
20. 10. The method of claim 1, wherein the treated genomic dsDNA is PCR amplified.
21. 10. The method of claim 1, wherein the treated genomic dsDNA is processed into fragments for sequencing.
22. 2. The method of claim 1, wherein the genomic dsDNA is treated with a dsDNA deaminase in a cell or in a nucleus isolated from a cell.
23. 2. The method of claim 1, wherein the genomic double-stranded (ds) DNA is processed to enrich for open chromatin DNA.
24. 10. The method of claim 1, wherein the genomic double-stranded (ds) DNA is processed to enrich for open chromatin DNA prior to treatment with the dsDNA deaminase.
25. 2. The method of claim 1, wherein the genomic double-stranded (ds) DNA is treated with the dsDNA deaminase and the treated genomic dsDNA is then processed to enrich for open chromatin DNA.