Method for determining binding site of whole genome DNA binding protein by using double-stranded DNA deaminase through footprint method
By converting cytosine in chromatin DNA into uracil using double-stranded DNA deaminase, determining the binding sites of transcription factors and DNA, solving the low resolution and high signal-to-noise ratio problems of detecting DNA protein interactions in existing methods, and achieving high-resolution identification and analysis of transcription factor binding sites.
Patent Information
- Application Number
- CN202280100612.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-30
- Publication Date
- 2025-05-23
AI Technical Summary
Existing methods detect DNA protein interactions at the genomic level have high signal-to-noise ratios, low throughput, low resolution, and cell number and homogeneity requirements, and it is difficult to identify transcription factors that collaborate with each other during gene regulation.
The cytosine in chromatin DNA is converted to uracil by using double-stranded DNA deaminase, and the deaminase is blocked at the binding site unless the transcription factor binds to the DNA, thereby determining the binding site of the transcription factor to the DNA.
High resolution identification of multiple transcription factors' binding sites across the genome-wide range allows quantification of the simultaneous binding of multiple transcription factors on a single DNA molecule, and analyzes the synergistic or antagonistic behavior of transcription factors during transcription regulation.
Smart Images

Figure HDA0005334266320000011 
Figure HDA0005334266320000021 
Figure HDA0005334266320000031
Abstract
Description
[0001] field
[0002] The present disclosure generally relates to systems, methods, and compositions for determining DNA binding protein binding sites on the genome of one or more cells.
[0003] background
[0004] Each cell of an individual has essentially the same genome, but they perform completely different functions in each tissue. The advent of single-cell genomics has made it possible to determine the transcriptome, methylome, and open chromosome sites of individual human cells, enabling the classification of cell types in an unprecedented way. However, beyond cell typing, the most pressing challenge is to decode the human functional genome, that is, to understand cell functions based on the human genome. Processes such as gene expression and regulation, cell differentiation and development are related to chromatin structure and regulatory networks, and transcription factors (TFs) are crucial for them.
[0005] There are only about 1,000 TFs in humans, controlling about 20,000 genes. The specificity of gene regulation is achieved through the combinatorial binding of several TFs, which act like a set of keys to turn specific genes on and off. Therefore, it is very important to understand the precise binding groups of TFs and how they cooperate with each other. Current methods for detecting DNA-protein interactions at the genomic level, such as ChIP-seq, have several problems, including high signal-to-noise ratio, low throughput, low resolution, and requirements for cell number and homogeneity. Importantly, existing methods for detecting DNA-protein interactions at the genomic level lack the ability to identify TFs that cooperate with each other in the gene regulation process.
[0006] Sequence Listing
[0007] This application contains a sequence listing, which has been submitted electronically in XML format and is hereby incorporated by reference in its entirety. The XML copy was created on September 28, 2022, named "009191.00002_st26", and is 23KB in size.
[0008] Overview
[0009] In general, aspects of the present disclosure relate to methods for identifying DNA binding proteins (e.g., transcription factors) on double-stranded polynucleotides (e.g., genomic DNA) or obtaining their binding profiles. In general, the methods of the present disclosure utilize double-stranded (ds) DNA deaminases to convert cytosine on double-stranded polynucleotides to uracil, except for the position where the DNA binding protein binds to the polynucleotide. The dsDNA deaminases are spatially blocked from converting cytosine to uracil at the position where the DNA binding protein binds to the polynucleotide, thereby generating a "footprint" where cytosine is not converted to uracil. Therefore, the position where the DNA binding protein binds to the polynucleotide can be determined. The determined DNA binding site can be compared with the DNA binding site of a known DNA binding protein to identify the DNA binding protein that binds to the determined DNA binding site. Using the methods described herein, one or more or a plurality of DNA binding proteins can be identified for a given polynucleotide (e.g., a gene within chromatin DNA).
[0010] Aspects of the present disclosure relate to methods for identifying binding sites of one or more transcription factors (TF) to DNA (e.g., chromatin DNA). The identification of transcription factor binding sites can be used to identify the transcription factor itself based on its known binding sites to chromatin DNA, and thus identify one or more, one or more pairs, or a large number of (plurality) transcription factors that co-regulate genes. According to one aspect, for specific genes, and along the genome, a transcription factor combination or "keyset" is decoded to identify a genome-wide transcription factor combination or keyset.
[0011] According to one aspect, a method is provided, including contacting chromatin DNA with dsDNA deaminase. Unless TF binds to chromatin DNA, dsDNA deaminase will convert cytosine to uracil along chromatin DNA. The binding of TF to chromatin DNA spatially prevents dsDNA deaminase from converting cytosine to uracil at the binding site between TF and chromatin DNA. Therefore, the conversion of cytosine to uracil occurs on both sides of the binding site between TF and chromatin DNA. The boundary of TF binding to chromatin DNA based on the conversion of cytosine to uracil can be determined, and the "footprint" of the binding site can be determined accordingly. The binding site is then compared with the known binding site of TF to identify TF with a matching binding site. Target chromatin DNA (e.g., gene) can be analyzed to determine whether one or more, a pair or a large number of (plurality) TFs bind to the target chromatin DNA, thereby allowing identification of TFs involved in the regulation of specific genes. According to one aspect, this method for identifying TF key groups can be implemented in the genome-wide range of all genes.
[0012] In general, aspects of the present disclosure include cell permeabilization or cell nucleus permeabilization and isolation of cells, single cells or cell groups. In this way, the genomic DNA of the cells, single cells or cell groups is made more accessible to dsDNA deaminase. According to one aspect, one or more cells or one or more cell nuclei do not need to be permeabilized while still allowing treatment with dsDNA deaminase. Other methods of treating one or more cells or one or more cell nuclei according to known methods (e.g., lysis) to make one or more cells or one or more cell nuclei more accessible to dsDNA deaminase are contemplated.
[0013] According to one aspect, the permeabilized cell or cells or permeabilized cell or cell nuclei may be treated with a cross-linking agent to cross-link cellular components to maintain cell structure prior to treatment with a double-stranded DNA deaminase, as known in the art. According to one aspect, a double-stranded polynucleotide molecule having a DNA binding protein bound thereto (e.g., in the case of cell-free DNA) may be treated with a dsDNA deaminase.
[0014] According to one aspect, the cytosine in the DNA of one or more cells or one or more nuclei is converted to uracil by hydrolysis to remove the amino group in the cytosine nucleotide that can be used for deamination to produce uracil nucleotide, and in this way, the permeabilized one or more cells or permeabilized one or more nuclei or cell-free DNA are treated with double-stranded DNA deaminase. According to one aspect, before being treated with dsDNA deaminase, DNA can be fragmented and enriched. According to one aspect, DNA can be treated with dsDNA deaminase and then fragmented and enriched. Treating with dsDNA deaminase produces treated DNA, as long as the treated DNA includes one or more uracils produced by dsDNA deaminase treatment. Exemplary double-stranded DNA deaminases include double-stranded DNA deaminase A ("DddA") known in the art, evolved double-stranded DNA deaminase A11 ("DddA11") known in the art, and bacterial deaminase toxin family 3 ("BadTF3") known in the art. The treated DNA, such as treated chromatin DNA, can be treated with a transposase or a DNase to enrich for open chromatin DNA, ie, transcriptionally active genomic DNA that is accessible to DNA regulatory elements.
[0015] The treated DNA can be amplified before sequencing. Exemplary amplification methods include PCR. Alternatively, the treated DNA for sequencing can be sequenced directly after library preparation without amplification.
[0016] The treated DNA can be sequenced. According to one aspect, the treated DNA can be sequenced as a whole genome DNA. According to one aspect, the treated DNA can be sequenced as an accessible region of chromatin by cleavage of an enzyme (e.g., a transposase) or a nuclease (e.g., a DNase, MNase, or a restriction endonuclease). According to one aspect, the enriched open chromatin DNA is sequenced to determine the conversion of cytosine to uracil, and thus determine the DNA binding protein footprint.
[0017] According to one aspect, in addition to the method using a transposase (e.g., Tn5 transposase), methods known to those skilled in the art (including ChIP-seq, ChIC, ChEC, ChEC-seq, CUT&TAG, CUT&RUN, Multi-CUT&Tag, NTT-seq, R loop CUT&Tag, etc.) can also be used to enrich the target chromatin region (e.g., open chromatin region) before or after treatment with a dsDNA deaminase. For example, dsDNA is treated with a dsDNA deaminase, and then CUT&TAG is used to enrich the open chromatin region in the treated DNA for sequencing. Alternatively, CUT&TAG is used to enrich the open chromatin region in the dsDNA for sequencing, and then the enriched DNA is treated with a dsDNA deaminase.
[0018] According to one aspect, libraries can be prepared based on whole genomic DNA known in the art. According to one aspect, libraries can be prepared based on DNA regions of interest (e.g., open chromatin). According to one aspect, antibodies are used to enrich target regions in methods such as ChIP-seq, ChIC, ChEC, ChEC-seq, CUT&TAG, and CUT&RUN. According to one aspect, other binding agents such as nanobodies (see Stuart et al., Nanobody-tethered transposition allows for multifactorial chromatin profiling at single-cell resolution, bioRxiv 10.1101 / 2022.03.08.483436v1, which is hereby incorporated by reference in its entirety) or specific chromatin binding domains (see Wang et al., Genomic profiling of native R loops with a DNA-RNA hybrid recognition sensor, Sci. Adv. 2021 Feb; 7(8)eabe3516 10.1126 / sciadv.abe3516, which is hereby incorporated by reference in its entirety) are used to enrich target regions.
[0019] Then the binding profile of DNA binding protein can be obtained by analyzing the information obtained from the cytosine to uracil conversion site on the sequence read length. Unconverted site (i.e., cytosine is still cytosine during treatment with double-stranded DNA deaminase) represents the binding site to which DNA binding protein binds when treated with double-stranded DNA deaminase. Conversion site (i.e., cytosine is converted into uracil by double-stranded DNA deaminase) represents the site to which DNA binding protein does not bind when treated with double-stranded DNA deaminase. Methods known to those skilled in the art can be used to compare DNA binding protein binding profile with unconverted site to identify specific DNA binding protein, such as using binding site databases related to DNA binding protein, such as JASPAR (global website jaspar.genereg.net), CIS-BP (global website cisbp.ccbr.utoronto.ca), HOCOMOCO (global website hocomoco11.autosome.org), which are TF motif databases. Potential binding TF is identified by comparing the footprint identified by the dsDNA deaminase method described herein with the known TF motifs from these databases.
[0020] Aspects of the present disclosure can be performed at the single cell level or single DNA molecule level or using multiple cells. The multiple cells can be the same cell type. The multiple cells can be different cell types.
[0021] According to the present disclosure, a method is provided to quantify the simultaneous binding of multiple TFs on a single DNA molecule. The method provides a high-resolution binding map of multiple TFs at the whole genome level. According to the present disclosure, a method is provided to analyze how TFs cooperate or antagonize in the transcriptional regulation process.
[0022] According to one aspect, a dsDNA deaminase is provided for use in preparing a preparation for implementing a method for determining a transcription factor binding site on a genomic double-stranded (ds) DNA of a eukaryotic cell. In certain embodiments, the method is as described herein. In certain embodiments, the method comprises
[0023] contacting the genomic dsDNA with a dsDNA deaminase under conditions that convert cytosine of the genomic dsDNA to uracil, thereby producing treated genomic dsDNA, and
[0024] Unconverted cytosines on the treated genomic dsDNA were identified as transcription factor binding sites.
[0025] BRIEF DESCRIPTION OF THE DRAWINGS
[0026] The patent or application file contains at least one drawing in color. Copies of the patent or patent application publication in color will be provided by the Patent Office upon request and payment of the necessary fee. The above and other features and advantages of the present invention will be more fully understood from the following detailed description of exemplary embodiments taken in conjunction with the accompanying drawings, in which:
[0027] Figure 1 is a schematic diagram depicting the use of double-stranded DNA deaminase to determine TF binding sites.
[0028] Figure 2 A vector map of pETDuet-1::dddAtox+dddAI is depicted.
[0029] Figure 3 A vector map of pETDuet-1::badTF3tox+badTF3I is depicted.
[0030] Figure 4 A vector map depicting pETDuet-1::dddA11+dddAI
[0031] Figure 5 SDS-PAGE gels stained with Coomassie Brilliant Blue of DddAtox-DddAI, DddAtox, DddA11-DddAI, DddA11, BadTF3-BadTF3I and BadTF3, respectively.
[0032] Figure 6 Depicted are data supporting the conversion rates of cytosine to uracil using several double-stranded DNA deaminases.
[0033] Fig. 7A and 7B Data identifying TF footprints are depicted. Figure 7C Depicted is a Venn diagram showing the overlap between CTCF binding identified by the dsDNA deaminase method, ChIP-seq method, and DNase-seq method described herein.
[0034] Fig. 8A Schematic diagram depicting the TF binding mode at the single-molecule level. Figure 8B Depicted are data demonstrating identification and quantification of binding of the transcription factor CTCF to DNA. Figure 8C Depicted are data demonstrating the identification of multiple transcription factor binding sites using the dsDNA deaminase method described herein.
[0035] Fig.9A is a schematic diagram depicting one aspect of the method of the present disclosure. Fig. 9B Depicted are DNA fragment distribution data generated by the methods described herein. Fig. 9CDepicted are the cell profiling results for single cell data generated by the methods described herein for K562, GM12878, and HEK293T cell lines. Fig.9D Depicted are comparative data of bulk data and single cell data generated by the methods described herein viewed via IGV software.
[0036] Detailed Description
[0037] Unless otherwise indicated, the practice of certain embodiments or features of certain embodiments may employ conventional techniques of molecular biology, microbiology, recombinant DNA, etc., which are within the ordinary skill in the art. Such techniques are fully explained in the literature. See, for example, Sambrook, Fritsch, and Maniatis, MOLECULAR CLONING: ALABORATORY MANUAL, Second Edition (1989), OLIGONUCLEOTIDE SYNTHESIS (MJGaitEd., 1984), ANIMAL CELL CULTURE (RIFreshney, Ed., 1987), METHODS IN ENZYMOLOGY series (Academic Press, Inc.); GENE TRANSFER VECTORS FOR MAMMALIAN CELLS (JMMillerand MPCalos eds.1987), HANDBOOK OF EXPERIMENTAL IMMUNOLOGY, (DMWeir andC.C.Blackwell,Eds.), CURRENT PROTOCOLS IN MOLECULAR BIOLOGY(FMAusubel,R.Brent,REKingston,DDMoore,JGSiedman,JASmith,and K. Struhl, eds., 1987), CURRENT PROTOCOLS IN IMMUNOLOGY (JEcoligan, A. M. Kruisbeek, D. H. Margulies, E. M. Shevach and W. Strober, eds., 1991); ANNUAL REVIEW OF IMMUNOLOGY; and monographs in journals such as ADVANCES IN IMMUNOLOGY. All patents, patent applications and publications mentioned herein (including those above and below) are incorporated herein by reference.
[0038] The terminology and notation of nucleic acid chemistry, biochemistry, genetics, and molecular biology used herein follows that of standard treatises and textbooks in the field, such as Kornberg and Baker, DNA Replication, Second Edition (WH Freeman, New York, 1992); Lehninger, Biochemistry, Second Edition (Worth Publishers, New York, 1975); Strachan and Read, Human Molecular Genetics, Second Edition (Wiley-Liss, New York, 1999); Eckstein, editor, Oligonucleotides and Analogs: A Practical Approach (Oxford University Press, New York, 1991); Gait, editor, Oligonucleotide Synthesis: A Practical Approach (IRL Press, Oxford, 1984); etc.
[0039] refer to Figure 1 Aspects of the present disclosure are described. Figure 1As shown, a cell nucleus is extracted from a cell, cell line or tissue and permeabilized. The cell nucleus includes chromatin DNA with a DNA binding protein such as TF bound thereto. The separated permeabilized cell nucleus is incubated (i.e., treated) with a dsDNA cytosine deaminase that converts accessible cytosine on dsDNA into uracil to produce treated DNA, such as treated chromatin DNA. Accessible cytosine is a cytosine that is not spatially hindered (i.e., spatially unhindered) to react enzymatically with a dsDNA cytosine deaminase. Inaccessible cytosine is a cytosine that is spatially hindered by a DNA binding protein (e.g., TF) and is therefore inaccessible or otherwise unavailable for enzymatic reaction with a dsDNA cytosine deaminase. Chromatin DNA regions that are not bound by DNA binding proteins, such as chromatin DNA regions that are immediately adjacent to and flanking DNA binding proteins ("flanking regions"), can be approached by dsDNA deaminase, which converts cytosine into uracil in chromatin DNA regions that are not bound by DNA binding proteins. Chromatin DNA regions that are bound by one or more DNA binding proteins (such as transcription factors or other DNA binding proteins) are protected from the conversion of cytosine into uracil by dsDNA deaminase. Therefore, the conversion region is adjacent to or otherwise flanked by the unconverted region, wherein the unconverted region corresponds to the DNA binding protein binding site. This DNA binding protein binding site is referred to as a DNA binding protein "footprint". The DNA binding protein binding site can be a single nucleotide, or can be one or more nucleotides, two or more nucleotides, three or more nucleotides, four or more nucleotides, can be 1 to 4 or 1 to 5 nucleotides, or can extend to multiple nucleotides, or can be multiple nucleotides in length. The DNA binding protein binding site can be a single nucleotide or multiple nucleotides. The DNA binding protein can interact with one or more nucleotides of the chromatin DNA. Nucleotides that interact with DNA binding proteins can define DNA protein binding sites. According to one aspect, processed chromatin DNA is extracted for full genome or targeted amplicon sequencing. Methods described herein allow for simultaneous quantification of multiple DNA binding protein binding events (e.g., TF binding events, whether paired TF or more than two TFs, three TFs, four TFs, five TFs, etc.) on the genes of a single DNA molecule. Methods described herein are intended to systematically quantify the co-occupancy frequency of multiple TFs or TF pairs (e.g., thousands of TFs or TF pairs) in multiple genes in a genome.
[0040] cell
[0041] Cells according to the present invention include any cell for which a person skilled in the art considers it useful to understand the DNA binding protein binding sites of its double-stranded DNA (e.g., chromatin DNA). Cells include prokaryotic cells or eukaryotic cells. Cells according to the present disclosure include any type of cancer cells, hepatocytes, oocytes, embryos, stem cells, iPS cells, ES cells, neurons, erythrocytes, melanocytes, astrocytes, germ cells, oligodendrocytes, kidney cells, etc.
[0042] Cells useful in the methods described herein can be obtained from biological samples, target tissues or biopsies, blood samples or cell cultures. In addition, cells from specific organs, tissues, tumors, vegetations, etc. can be obtained and used in the methods described herein. In addition, in general, cells from any population, such as prokaryotic or eukaryotic unicellular organisms including bacteria or yeast, can be used in the present methods. According to one aspect, the sample can be in vitro. The term "in vitro" has its art-recognized meaning, such as involving purified reagents or extracts, such as cell extracts. As used herein, the term "biological sample" is intended to include, but is not limited to, tissues, cells, biological fluids and isolates thereof separated from a subject, as well as tissues, cells and fluids present in a subject.
[0043] According to one aspect, the method of the present invention is implemented with a single cell. As used herein, "single cell" refers to a cell. Single cell suspensions can be obtained using standard methods known in the art, including, for example, enzymatic digestion of proteins of cells in attached tissue samples using trypsin or papain or release of adherent cells in culture, or mechanical separation of cells in samples. Single cells can be placed in any suitable reaction vessel, wherein the single cells can be processed individually. For example, a 96-well plate, so that each single cell is placed in a single well.
[0044] Methods for manipulating individual cells are known in the art and include fluorescence activated cell sorting (FACS), flow cytometry (Herzenberg., PNAS USA 76: 1453-55 1979), micromanipulation, and the use of a semi-automatic cell picker (e.g., the QUIXELL cell transfer system from Stoelting Co.). For example, individual cells can be selected individually based on features that can be detected by microscopic observation (e.g., position, morphology, or reporter gene expression). In addition, a combination of gradient centrifugation and flow cytometry can also be used to improve separation or sorting efficiency.
[0045] According to one aspect, the method of the invention is practiced using a plurality of cells. The plurality of cells includes from about 2 to about 1,000,000 cells, from about 2 to about 10 cells, from about 2 to about 100 cells, from about 2 to about 1,000 cells, from about 2 to about 10,000 cells, from about 2 to about 100,000 cells, from about 2 to about 10 cells or from about 2 to about 5 cells.
[0046] According to one aspect, the DNA to be processed is genomic DNA or chromatin DNA. According to one aspect, the DNA to be processed is mammalian DNA, plant DNA, yeast DNA, viral DNA or prokaryotic DNA. In yet another preferred embodiment, the DNA sample is obtained from a human, bovine, porcine, ovine, equine, rodent, avian, fish, shrimp, plant, yeast, virus or bacterium. Preferably, the DNA to be processed is genomic DNA. As used herein, the term "genome" is defined as the total gene set carried by an individual, cell or organelle. As used herein, the term "genomic DNA" is defined as DNA material that contains a partial or complete total gene set carried by an individual, cell or organelle.
[0047] According to one aspect, the DNA to be processed is a double-stranded polynucleotide molecule having a protein bound thereto, such as cell-free DNA. According to this aspect, DNA, such as genomic DNA or chromatin DNA, can be isolated, treated with a dsDNA deaminase, and processed and analyzed as described herein.
[0048] Method for permeabilizing cells or cell nuclei
[0049] Once one or more desired cells have been identified, the methods described herein can be performed on the one or more cells or one or more cell nuclei obtained from the one or more cells. Single cells or a plurality of cells or one or more cell nuclei can be isolated. The one or more cells or one or more cell nuclei can be treated according to known methods to facilitate the entry of chemicals, drugs, enzymes (such as dsDNA deaminase), DNA or other reagents to be introduced into the one or more cells or one or more cell nuclei.
[0050] According to one aspect, the one or more cells or one or more cell nuclei can be permeabilized. Permeabilization methods are known to those skilled in the art. Exemplary permeabilization techniques include electroporation or electropermeabilization, permeabilization with mild non-ionic detergents (such as saponin and digitonin), and permeabilization by pore-forming toxins (such as alpha-toxin and streptolysin O), as known in the art. Electroporation or electropermeabilization or electroporation is a technique in which an electric field is applied to a cell or cell nucleus to increase the permeability of the cell membrane, thereby allowing chemicals, drugs, enzymes (such as dsDNA deaminase), electrode arrays or DNA to be introduced into the cell.
[0051] Alternatively, one or more cells may be lysed using methods known to those skilled in the art to obtain one or more nuclei, or one or more nuclei may be lysed. Lysis may be achieved, for example, by heating the cells or by using detergents or other chemical methods or by a combination of these methods. However, any suitable lysis method known in the art may be used.
[0052] Alternatively, DNA, such as genomic DNA or chromatin DNA, can be extracted or otherwise isolated from a tissue sample, a blood sample, one or more cells, etc., according to methods known in the art. DNA extraction protocols using beads (e.g., DYNABEADS) or reagents are known to those skilled in the art and are commercially available in kit form through ThermoFisher Scientific, such as the CHARGESWITCH Genomic DNA Purification Kit, etc.
[0053] Methods for cross-linking cells and nuclei
[0054] In certain embodiments, prior to dsDNA deaminase treatment, one or more cells or one or more cell nuclei are treated with a crosslinking agent to maintain cell structure, as known in the art. Such crosslinking treatments include treatment with paraformaldehyde or treatment with ultraviolet light to achieve crosslinking. Crosslinking is used to maintain cell structure, and in some cases, bound TF can be stably crosslinked with DNA. The methods described herein can be applied to cells and cell nuclei treated with crosslinking agents. According to one aspect, the methods described herein can be applied to formalin-fixed and paraffin-embedded (FFPE) samples, etc., as known in the art. For example, FFPE is a form of preservation and preparation of specimens. First, the tissue sample is fixed in formaldehyde (also known as formalin), such as 10% neutral buffered formalin solution for about 18-24 hours to preserve proteins and important structures within the tissue. Next, it is embedded in a paraffin block, and then the sample of the paraffin block can be processed according to the method described herein. In order to prepare for infiltration with wax, the tissue can be dehydrated and transparent, usually using increasing concentrations of ethanol. Then, it is embedded in IHC grade paraffin according to known methods.
[0055] DNA Binding Proteins
[0056] According to certain aspects of the present disclosure, methods are provided for determining DNA, such as chromatin DNA or DNA binding protein binding sites on cells. DNA binding proteins are proteins with DNA binding domains and therefore have specific or general affinity for single-stranded or double-stranded DNA. Sequence-specific DNA binding proteins typically interact with the major groove of B-DNA because it exposes more functional groups that identify base pairs.
[0057] DNA binding proteins include transcription factors that regulate the transcription process, various polymerases, nucleases that cleave DNA molecules, and histones that are involved in chromosome packaging and transcription in the cell nucleus. DNA binding proteins can incorporate domains such as zinc fingers, helix-turn-helix, and leucine zippers (among many other domains) that facilitate binding to nucleic acids.
[0058] Structural proteins that bind to DNA are well-known examples of nonspecific DNA-protein interactions. Within chromosomes, DNA is complexed with structural proteins. These proteins organize the DNA into a compact structure called chromatin. In eukaryotes, this organization involves the binding of DNA to a complex of small basic proteins called histones. Histones form a disk-shaped complex called a nucleosome, which contains two complete turns of double-stranded DNA wrapped around its surface. These nonspecific interactions are formed through basic residues in the histones, which form ionic bonds with the acidic sugar-phosphate backbone of DNA, and are therefore largely independent of base sequence. Chemical modifications of these basic amino acid residues include methylation, phosphorylation, and acetylation. These chemical changes alter the strength of the interaction between DNA and histones, making the DNA more or less accessible to transcription factors and altering the rate of transcription. Other nonspecific DNA-binding proteins in chromatin include the high mobility group (HMG) proteins, which bind to bent or twisted DNA. These proteins play an important role in bending the nucleosome array and arranging them into the larger structure that forms chromosomes.
[0059] In contrast, transcription factors bind to specific DNA sequences. As is known in the art, each transcription factor binds to a specific set of DNA sequences and activates or represses transcription of genes that have these sequences near the promoter. The specificity of transcription factor interactions with DNA comes from the multiple contacts that the proteins make with the edges of DNA bases, allowing them to read the DNA sequence. Most of these base interactions occur in the major groove, where the bases are most accessible. A transcription factor (TF) (or sequence-specific DNA binding factor) is a protein that controls the rate at which genetic information is transcribed from DNA into messenger RNA by binding to specific DNA sequences. The function of TFs is to regulate (turn on and off) genes to ensure that they are expressed in the desired cells at the right time and in the right amount throughout the life cycle of cells and organisms. Groups of TFs act in a coordinated manner to direct cell division, cell growth, and cell death throughout the life cycle; cell migration and organization (body plan) during embryonic development; and intermittently in response to signals from outside the cell, such as hormones. There are as many as 1,600 TFs in the human genome. See Babu MM, Luscombe NM, Aravind L, Gerstein M, Teichmann SA (June 2004). "Structure and evolution of transcriptional regulatory networks" (PDF). Current Opinion in Structural Biology. 14(3):283-91. doi: 10.1016 / j.sbi.2004.05.004. PMID 15193307, which is hereby incorporated by reference in its entirety.
[0060] TFs, alone or together with other proteins in a complex, act by promoting (as activators) or preventing (as inhibitors) the recruitment of RNA polymerase (the enzyme that performs the transcription of genetic information from DNA to RNA) to specific genes. A hallmark feature of TFs is that they contain at least one DNA binding domain (DBD) that associates with specific DNA sequences near the genes they regulate. See Mitchell PJ, Tjian R (July 1989). "Transcriptional regulation in mammalian cells by sequence-specific DNA binding proteins". Science. 245 (4916): 371-8. Bibcode: 1989 Sci ... 245..371M.doi: 10.1126 / science.2667136. PMID 2667136; Ptashne M, Gann A (April 1997). "Transcriptional activation by recruitment". Nature. 386 (6625): 569-77. Bibcode: 1997 Natur. 386..569P.doi: 10.1038 / 386569a0. PMID 9121580. S2CID 6203915. Each of which is incorporated herein by reference in its entirety for teaching known transcription factors. TFs are divided into several categories based on their DNA-binding domains.See Stegmaier P, Kel AE, Wingender E (2004). "Systematic DNA-binding domain classification of transcription factors". Genome Informatics. International Conference on Genome Informatics. 15(2): 276-86. PMID 15706513, archived from the original on June 19, 2013; Matys V, et al. (January 2006). "TRANSFAC and its module TRANSCompel: transcriptional gene regulation in eukaryotes". Nucleic Acids Research. 34(Database issue): D108-10. doi: 10.1093 / nar / gkj143. PMC 1347505. PMID 16381825, each of which is incorporated herein by reference in its entirety for teaching transcription factors and their associated DNA binding domains.
[0061] According to one aspect, exemplary transcription factors include, but are not limited to, AAF, ABL, ADA2, ADANF1, AF1, AFP1, AHR, AIIN3, AIRE, ALL1, ALPHACBF, ALPHACP1, ALPHACP2A, ALPHACP2B, ALPHAH2, ALPHAH3, ALPHAHO, ALX1, ALX3, ALX4, AMEF2, AML1, AML1A, AML1B, AML1C, AML1DELTAN, AML2, AML3, AML3A, AML3B, AMY1L, AMYB, ANF, ANHX, AP1, AP2ALP, AP2ALP1, AP2ALP2A, AP2ALP3, AP2ALP4, AP2ALP5, AP2ALP6, AP2ALP7, AP2ALP8, AP2ALP9, AP2ALP10, AP2ALP11, AP2ALP12 HAB, AP2BETA, AP2GAMMA, AP3(1), AP3(2), AP4, AP5, APC, AR, AREB6, ARGFX, ARID5B, ARNT, ARNT(774M form), ARNT2, ARNT::HIF1A, ARNTL, ARP1, ARX, ASCL1 ,ASCL2,ATBF1A,ATBF1B,ATF,ATF1,ATF2,ATF3,ATF3DELTAZIP,ATF4,ATF6,ATF6B,ATF7,ATFA,ATFADELTA,ATOH1,ATOH7,ATPF1,B,BACH1,BACH2,BANP , BARH11, BARH12, BARHL1, BARHL2, BARX1, BARX2, BATF, BATF3, BATF::JUN, BBX, BCL11A, BCL11B, BCL3, BCL6, BCL6B, BD73, BETACATENIN, BHLHA15, BHL HE22, BHLHE23, BHLHE40, BHLHE41, BIN1, BMYB, BNC2, BP1, BP2, BPTF, BRAHMA, BRCA1, BRN3A, BRN3B, BRN4, BSX, BTEB, BTEB2, BTFIID, C / EBPALPHA, C / EBP BETA, C / EBPDELTA, CACC binding factor, CART1, CBF(4), CBF(5), CBP, CCAAT binding factor, CCF, CCG1, CCK1A, CCK1B, CCMT binding factor, CD28RC, CDC5L, CDK2, CDK9, CDX1, CDX2, CDX4, CEBPA, CEBPB, CEBPD, CEBPE, CEBPG, CENPB, CENPBD1, CFF, CHXLO, CLIM2, CLIMI, CLOCK, CNBP, COS, COUP, CP1, CP2, CPBP, CPEB1, CPE binding protein, CPIA, CPIC,CREB, CREB1, CREB2, CREB3, CREB3L1, CREB3L4, CREB5, CREBPL, CREBPA, CREM, CREMALPHA, CRF, CRX, CSBP1, CTCF, CTCFL, CTF, CTF1, CTF2, CTF3, CTF5, CTF7, CUP, CUTL1, CUX1, CUX2, CX, CXXC5, CYCLINA, CYCLINT1, CYCLINT2, CYCLINT2A, CYCLINT2B, DAP, DAX1, DB1, DBF4, DBP, DBPA, DBPAV, DBPB, DDB, DDB1, DDB2, DEF, DELTACREB, DELTAMAX, DF1, DF2, DF3, DIX4 (long isoform), DLX1, DLX2, DLX3, DLX4, DLX4 (short isoform), DLX5, DLX6, DMRT1, DMRT2, DMRT3, DMRTA1, DMRTA2, DMRTC2, DNMT1, DP1, DP2, DPF1, DPRX, DRGX, DSIF, DSIFP14, DSIFP160, DTF, DUX1, DUX2, DUX3, DUX4, DUXA, E, E12, E2F, E2F + E4, E2F + P107, E2F1, E2F2, E2F3, E2F4, E2F5, E2F6, E2F7, E2F8, E47, E4BP4, E4F, E4F1, E4TF2, EAR2, EBF1, EBF3, EBP80, EC2, EF1, EFC, EGR1, EGR2, EGR3, EGR4, EHF, EIF1, EIIAEA, EIIAEB, EIIAECALPHA, EIIAECBETA, EIVF, ELF1, ELF2, ELF3, ELF4, ELF5, ELK1, ELK1::HOXA1, ELK1::HOXB13, ELK1::SREBF2, ELK3, ELK4, EMX1, EMX2, EN1, EN2, ENHBIND.PROT, ENKTF1, EOMES, EPAS1, EPSILONF1, ER, ERF, ERF::FIGLA, ERF::FOXI1, ERF::FOXO1, ERF::HOXB13, ERF::NHLH1, ERF::SREBF2, ERG, ERG1, ERG2, ERR1, ERR2, ESR1, ESR2, ESRRA, ESRRB, ESRRG, ESX1, ETF, ETS1, ETS1DELTAVIL, ETS2, ETV1, ETV2, ETV2::DRGX, ETV2::FIGLA, ETV2::FOXI1ETV2::HOXB13、ETV3、ETV4、ETV5、ETV5::DRGX、ETV5::FIGLA、ETV5::FOXI1 、ETV5::FOXO1、ETV5::HOXA2、ETV6、ETV7、EVX1、EVX2、F2F、FACTOR2、FACTO RNAME, FBP, FEBP, FERD3L, FEV, FEZF1, FIGLA, FKBP59, FKHL18, FKHRL1P2, F LI1, FLI1::DRGX, FLI1::FOXI1, FOS, FOS::JUN, FOS::JUNB, FOS::JUND, FOS B、FOSB::JUN、FOSB::JUNB、FOSL1、FOSL1::JUN、FOSL1::JUNB、FOSL1::JUN D、FOSL2、FOSL2::JUN、FOSL2::JUNB、FOSL2::JUND、FOXA1、FOXA2、FOXA3、FO XB1、FOXC1、FOXC2、FOXD1、FOXD2、FOXD3、FOXD4、FOX1、FOXE3、FOXF1、FOXF 2, FOXG1, FOXG1A, FOXG1B, FOXG1C, FOXH1, FOXI1, FOXJ1A, FOXJ1B, FOXJ2, FO XJ2(floating), FOXJ2(floating), FOXJ2::ELF1, FOXJ3, FOXK1, FOXK1A, FOXK1B, FO XK1C、FOXK2、FOXL1、FOXL2、FOXM1、FOXM1A、FOXM1B、FOXM1C、FOXN1、FOXN2、F OXN3、FOXO1、FOXO1::ELF1、FOXO1::ELK1、FOXO1::ELK3、FOXO1::FLI1、FOX O1A、FOXO1B、FOXO2、FOX3、FOX3A、FOXO3B、FOXO4、FOXO6、FOXP1、FOXP2、FO XP3, FOXQ1, FOXR1, FOXR2, FRA1, FTF, FTS, G6FACTOR, GABP, GABPA, GA BPALPHA、GABPBETA1、GABPBETA2、GADD153、GAF、GAMMACAC1、GAMMACAC2、GAM MACMT、STEP1、STEP1::TAL1、STEP2、STEP3、STEP4、STEP5、STEP6、GBX1、GBX 2, GCF, GCM1, GCM2, GCMA, GCNS, GF1, GFACTOR, GFI1, GFI1B, GLI, GLI1, GLI2GLI3, GLI4, GLIS1, GLIS2, GLIS3, GMEB1, GMEB2, GRALPHA, GRBETA, GRF1, GRHL1, GRHL2, GSC, GSC2, GSCL, GSX1, GSX2, GTF3A, GTIC, GTIIA, GTIIBALPHA, GTIIBBETA, H1TF1, H1TF2, H2RIIBP, H4TF1, H4TF2, HAND1, HAND2, HB9, HDAC1, HDAC2, HDAC3, HDAXX, HDX, Heat Induced Factor, HEB, HEB1P67, HEB1P94, HEF1B, HEF1T, HEF4C, HEN1, HEN2, HES1, HES2, HES5, HES6, HES7, HESX1, HEX, HEY1, HEY2, HIC1, HIC2, HIF1, HIF1A, HIF1ALPHA, HIF1BETA, HINFA, HINFB, HINFC, HINFD, HINFD3, HINFE, HINFP, HIP1, HIVEP2, HKR1, HLF, HLTF, HLTF(MET123), HLX, HMBOX1, HMBP, HMGI, HMGI(Y), HMGIC, HMGY, HMX1, HMX2, HMX3, HNF1A, HNF1B, HNF3, HNF3ALPHA, HNF3BETA, HNF3GAMMA, HNF4, HNF4A, HNF4ALPHA, HNF4ALPHA1, HNF4ALPHA2, HNF4ALPHA3, HNF4ALPHA4, HNF4G, HNF4GAMMA, HNF6ALPHA, HNFIA, HNFIB, HNFIC, HNRNPK, HOMEZ, HOX11, HOXA1, HOXA10, HOXA11, HOXA13, HOXA2, HOXA3, HOXA4, HOXA5, HOXA6, HOXA7, HOXA9, HOXA9A, HOXA9B, HOXAIO, HOXAIOPL2, HOXB1, HOXB13, HOXB2, HOXB2::ELK1, HOXB3, HOXB4, HOXB5, HOXB6, HOXB7, HOXB8, HOXB9, HOXC10, HOXC11, HOXC12, HOXC13, HOXC4, HOXC5, HOXC6, HOXC8, HOXC9, HOXD1, HOXD10, HOXD11, HOXD12, HOXD12::ELK1, HOXD13, HOXD3, HOXD4, HOXD8, HOXD9, HP55, HP65, HPX42B,HRPF, HSF, HSF1, HSF1(LONG), HSF1(SHORT), HSF2, HSF4, HSF5, HSFY1, HSFY 2. HSP56, HSP90, IBP1, ICERI, ICERLIGAMMA, ICSBP, ID1, ID1H', ID2, ID3. ID3 / HEIR1、IF1、IGPE1、IGPE2、IGPE3、II1RF、IKAPPAB、IKAPPALPHA、IKA PPABBETA、IKAPPABR、IKZF1、IKZF3、IL6REBP、INSAF、INSM1、IPF1、IRF1、IRF 2、IRF3、IRF4、IRF5、IRF6、IRF7、IRF8、IRF9、IRX1、IRX2、IRX2A、IRX3、IRX4 IRX5, ISGF1, ISGF3, ISGF3ALPHA, ISGF3GAMMA, ISL1, ISL2, ISX, ITF, ITF1 ITF2, JDP2, JRF, JUN, JUN::JUNB, JUNB, JUND, KAPPAYFACTOR, KBP1, KDM2B, KER1, KLF1, KLF10, KLF11, KLF12, KLF13, KLF14, KLF15, KLF16, KLF17, KLF2 KLF3、KLF4、KLF5、KLF6、KLF7、KLF8、KLF9、KMT2A、KOX1、KRF1、KUAUTOANTIG EN, KUP, LBP1, LBP1A, LBX1, LBX2, LCORL, LCRF1, LEF1, LEFIB, LFA1, LHX1, L HX2、LHX3、LHX3A、LHX3B、LHX5、LHX6、LHX6.1A、LHX6.1B、LHX8、LHX9、LIN28 B. LIT1, LMO1, LMO2, LMX1A, LMX1B, LMY1 (Online), LMY1 (Local), LMY2, LSF, LXRAL PHA、LY11、LYF1、LYL1、MAD1、MAF、MAF::NFE2、MAFA、MAFB、MAFF、MAFG、MAFG ::NFE2L1、MAFK、MASH1、MAX、MAX1、MAX2、MAX::MYC、MAZ、MAZ1、MB67、MBD2、M BF1, MBF2, MBF3, MBNL2, MBP1(1), MBP1(2), MBP2, MDBP, MECOM, MECP2, MEF2 MEF2A, MEF2B, MEF2C, MEF2C(433AA, MEF2C, 465AA, MEF2C, 473M,MEF2C / DELTA32(441AA range) MEF2D MEF2D00 MEF2D0B MEF2DA'B MEF2DA0 MEF2DAB, MEF2DAO, MEIS1, MEIS2, MEIS2A, MEIS2B, MEIS2C, MEIS2D, MEIS2E MEIS3、MEOX1、MEOX1A、MEOX2、MESP1、MESP2、MFACTOR、MGA、MGA::EVX1、MHO X(K2), MI, MIF1, MITF, MIXL1, MIZ1, MLX, MLXIPL, MM1, MNT, MNX1, MOP3, MR, M SANTD3, MSC, MSGN1, MSX1, MSX2, MTBZF, MTF1, MTF2, MTTF1, MXI1, MXIL, MYB MYBL1, MYBL2, MYC, MYC1, MYCN, MYF3, MYF4, MYF5, MYF6, MYNN, MYOD, MYOD1 MYOG, MYRF, MZF1, N10(25, NANOG, NC2, NCI, NCX, NELF, NER1, NET, NEUROD1 NEUROD2, NEUROG1, NEUROG2, NF1A, NF1B, NF1X, NF4FA, NF4FB, NF4FC, NFA, NF AB、NFAT1、NFAT3、NFAT5、NFATC、NFATC1、NFATC2、NFATC3、NFATC4、NFATP、N FATX、NFCLE0A、NFCLE0B、NFDELTAE3A、NFDELTAE3B、NFDELTAE3C、NFDELTAE4 A, NFDELTAE4B, NFDELTAE4C, NFE, NFE2, NFE2L1, NFE2L2, NFE2P45, NFE3, NF E6、NFETAA、NFGMA、NFGMB、NFI11A、NFIA、NFIB、NFIC、NFIC::TLX1、NFIL2A、N FIL2B, NFIL3, NFIX, NFJUN, NFKAPPAB, NFKAPPAB(LIKE), NFKAPPAB1, NFKAPPAB2, NFKAPPAB2(P49), NFKAPPAB2, NFKAPPAE1, NFKAPPAE2, NFKAPPAE3, N FKB1, NFKB2, NFMHCIIA, NFMHCIIB, NFMUE1, NFMUE2, NFMUE3, NFNF1, NFS, NFX NFX1, NFX2, NFX3, NFXC, NFYA, NFYB, NFYC, NFZC, NFZZ, NHLH1, NHLH2, NHP1NHP2, NHP3, NHP4, NKX21, NKX22, NKX23, NKX24, NKX25, NKX28, NKX2B, NKX2C, NKX2G, NKX31, NKX32, NKX3A, NKX3AV1, NKX3AV2, NKX3AV3, NKX3AV4, NKX3B , NKX61, NKX62, NKX63, NKX6A, NMI, NMYC, NOBOX, NOCT2ALPHA, NOCT2BETA, NOCT3, NOCT4, NOCT5A, NOCTSB, NOTO, NPAS2, NPTCII, NR1D1, NR1D2, NR1H2::R XRA, NR1H3, NR1H4, NR1H4::RXRA, NR1I2, NR1I3, NR2C1, NR2C2, NR2E1, NR2E3, NR2F1, NR2F2, NR2F6, NR3C1, NR3C2, NR4A1, NR4A2, NR4A2::RXRA, NR5A1, NR5A2, NR6A1, NRF1, NRF2, NRF2BETA1, NRF2GAMMA1, NRL, NRSF form 1, NRSF form 2, NTF, OCAB, OCT1, OCT2, OCT2.1, OCT2B, OCT2C, OCT4A, OCT4B, OCT5, OCT6, OC TAFACTOR, OCTAMER binding factor, OCTB2, OCTB3, OLIG1, OLIG2, OLIG3, ONECUT1, ONECUT2, ONECUT3, OSR1, OSR2, OTX1, OTX2, OVOL1, OVOL2, OZF, P107, P130, P28MODULATOR, P300, P38ERG, P45, P49ERG, P53, P55, P55ERG, P65DELTA, P67, PATZ1, PAX1, PAX2, PAX3, PAX3A, PAX3B, PAX4, PAX5, PAX6, PAX6 / PD5A, PAX7, P AX8, PAX8A, PAX8B, PAX8C, PAX8D, PAX8E, PAX8F, PAX9, PBX1, PBX1A, PBX1B, PBX2, PBX3, PBX3A, PBX3B, PBX4, PC2, PC4, PCS, PDX1, PEA3, PEBP2ALPHA, PEB P2BETA, PGR, PHF1, PHOX2A, PHOX2B, PIT1, PITX1, PITX2, PITX3, PKNOX1, PKNOX2, PLAG1, PLAGL2, PLZF, POB, PONTIN52, POU1F1, POU2F1, POU2F1::SOX2,POU2F2, POU2F3, POU3F1, POU3F2, POU3F3, POU3F4, POU4F1, POU4F2, POU4F3, POU5F1, POU5F1B, POU6F1, POU6F2, PPARA, PPARA::RXRA, PPARALPHA, PPARBETA, PPARD, PPARG, PPARG::RXRA, PPARGAMMA1, PPARGAMMA2, PPUR, PR, PRA, PRB, PRD1BF1, PRDIBFC, PRDM1, PRDM14, PRDM4, PRDM6, PRDM9, precursor, PROP1, PROX1, PRRX1, PRRX2, PSE1, PTEFB, PTF, PTF1A, PTFALPHA, PTFBETA, PTFDELTA, PTFGAMMA, PU.1, PUBOX-binding factor, PUBOX-binding factor (BJAB), PUF, PURFACTOR, R1, R2, RARA, RARA::RXRA, RARA::RXRG, RARALPHA1, RARB, RARBETA, RARBETA2, RARG, RARGAMMA, RARGAMMA1, RAX, RAX2, RBAK, RBP60, RBPJ, RBPJKAPPA, REL, RELA, RELB, REST, RFX, RFX1, RFX2, RFX3, RFX4, RFX5, RFX7, RFXS, RFY, RHOXF1, RORA, RORALPHA1, RORALPHA2, RORALPHA3, RORB, RORBETA, RORC, RORGAMMA, ROX, RPF1, RPGALPHA, RREB1, RSRFC4, RSRFC9, RUNX1, RUNX2, RUNX3, RVF, RXRA, RXRA::VDR, RXRALPHA, RXRB, RXRBETA, RXRG, SALL4, SAP1A, SAP1B, SATB1, SCRT1, SCRT2, SF1, SHOX2A, SHOX2B, SHOXA, SHOXB, SHP, SIIIP110, SIIIP15, SIIIP18, SIM', SIX1, SIX2, SIX3, SIX4, SIX5, SIX6, SKOR1, SKOR2, SMAD1, SMAD2, SMAD3, SMAD4, SMAD5, SNAI1, SNAI2, SNAI3, SOHLH2, SOX10, SOX11, SOX12, SOX13, SOX14, SOX15, SOX17, SOX18, SOX2, SOX21, SOX3, SOX30, SOX4, SOX5,SOX6, SOX7, SOX8, SOX9, SP1, SP2, SP3, SP4, SP5, SP8, SP9, SPDEF, SPHFACTOR, SPI1, SPIB, SPIC, SPIN, SPZ1, SRCAP, SREBF1, SREBF2, SREBP1A, SREBP1B, SREBP1C, SREBP2, SREZBP, SRF, SRPLSTAF50, SRY, STAT1, STAT1::STAT2, STAT1ALPHA, STAT1BETA, STAT2, STAT3, STAT4, STAT5A, STAT5B, STAT6, T, T3R, T3RALPHA1, T3RALPHA2, T3RBETA, TAF(I)110, TAF(I)48, TAF(I)63, TAF(II)100, TAF(II)125, TAF(II)135, TAF(II)170, TAF(II)18, TAF(II)20, TAF(II)250, TAF(II)250DELTA, TAF(II)28, TAF(II)30, TAF(II)31, TAF(II)55, TAF(II)70ALPHA, TAF(II)70BETA, TAF(II)70GAMMA, TAFI, TAFII, TAFL, TAL1, TAL1::TCF3, TAL1BETA, TAL2, TARFACTOR, TBP, TBR1, TBX1, TBX15, TBX18, TBX19, TBX1A, TBX1B, TBX2, TBX20, TBX21, TBX3, TBX4, TBX5, TBX6, TBXS (long isoform), TBXS (short isoform), TBXT, TCF, TCF1, TCF12, TCF1A, TCF1B, TCF1C, TCF1D, TCF1E, TCF1F, TCF1G, TCF21, TCF2ALPHA, TCF3, TCF4, TCF4(K), TCF4B, TCF4E, TCF7, TCF7L1, TCF7L2, TCFBETA1, TCFL5, TEAD1, TEAD2, TEAD3, TEAD4, TEF, TEF1, TEF2, TEL, TET1, TFAP2A, TFAP2B, TFAP2C, TFAP2E, TFAP4, TFAP4::ETV1, TFAP4::FLI1, TFCP2, TFCP2L1, TFDP1, TFE3, TFEB, TFEC, TFIIA, TFIIAALPHA / BETA precursor, TFIIAGAMMA, TFIIB, TFIID, TFIIE, TFIIEALPHA, TFIIEBETA<h2 style=";text-align:left;direction:ltr">TFIIF, TFIIFALPHA, TFIIFBETA, TFIIH, TFIIH*, TFIIHCAK, TFIIHCYCLINH, TFIIHERCC2 / CAK, TFIIHM015, TFIIHMAT1, TFIIHP34, TFIIHP44, TFIIHP62, TFIIHP80, TFIIHP90, TFIII, TFLFLTFLF2, TGIF, TGIF1, TGIF2, TGIF2LX, TGIF2LY, TGT3, THAP1, THAP11, THAP12, THRA, THRA1, THRB, TIF2, TIGD1, TLE1, TLX2、TLX3、TMF、TOPORS、TP53、TP63、TP73、TR2、TR211、TR29、TR3、TR4、TRA P、TREB1、TREB2、TREB3、TREF1、TREF2、TRF(2)、TRPS1、TTF1、TWIST1、TXREB P、TXREF、UBF、UBP1、UEF1、UEF2、UEF3、UEF4、UNCX、USF1、USF2、USF2B、VAV、 VAX2、VDR、VENTX、VEZF1、VHNF1A、VHNF1B、VHNF1C、VITF、VSX1、VSX2、WSTF、W T1、WT1DE12、WT1I、WT1IDE12、WT1IKTS、WT1KTS、X2BP、XBP1、XPA、XWV、XX、Y AF2、YB1、YBX1、YEBP、YY1、YY2、ZBED1、ZBED2、ZBTB12、ZBTB14、ZBTB17、ZBT B18、ZBTB2、ZBTB20、ZBTB22、ZBTB26、ZBTB32、ZBTB33、ZBTB37、ZBTB42、ZBT B43、ZBTB44、ZBTB45、ZBTB48、ZBTB49、ZBTB6、ZBTB7A、ZBTB7B、ZBTB7C、ZEB、 ZEB1、ZF1、ZF2、ZFHX2、ZFHX3、ZFP1、ZFP14、ZFP28、ZFP3、ZFP41、ZFP42、ZFP 57、ZFP64、ZFP69、ZFP69B、ZFP82、ZFP90、ZFX、ZHX1、ZIC1、ZIC2、ZIC3、ZIC4、 ZIC5、ZID、ZIK1、ZIM2、ZIM3、ZKSCAN1、ZKSCAN2、ZKSCAN3、ZKSCAN5、ZKSCAN 7、ZNF10、ZNF100、ZNF101、ZNF114、ZNF12、ZNF121、ZNF124、ZNF132、ZNF133、ZNF134、ZNF135、ZNF136、ZNF140、ZNF141、ZNF143、ZNF146、ZNF148、ZNF154、ZNF157、ZNF16、ZNF17、ZNF174、ZNF175、ZNF177、ZNF18、ZNF180、ZNF181、ZNF182、ZNF184、ZNF189、ZNF19、ZNF197、ZNF2、ZNF200、ZNF202、ZNF205、ZNF211、ZNF212、ZNF213、ZNF214、ZNF22、ZNF222、ZNF223、ZNF224、ZNF225、ZNF23、ZNF232、ZNF235、ZNF24、ZNF248、ZNF25、ZNF250、ZNF254、ZNF257、ZNF26、ZNF260、ZNF263、ZNF264、ZNF266、ZNF267、ZNF273、ZNF274、ZNF276、ZNF28、ZNF280A、ZNF281、ZNF282、ZNF283、ZNF284、ZNF285、ZNF287、ZNF296、ZNF3、ZNF30、ZNF300、ZNF302、ZNF304、ZNF311、ZNF317、ZNF32、ZNF320、ZNF322、ZNF324、ZNF324B、ZNF329、ZNF331、ZNF333、ZNF334、ZNF335、ZNF337、ZNF33A、ZNF33B、ZNF34、ZNF341、ZNF343、ZNF345、ZNF35、ZNF350、ZNF354A、ZNF354B、ZNF37A、ZNF382、ZNF383、ZNF384、ZNF385D、ZNF394、ZNF396、ZNF398、ZNF41、ZNF410、ZNF415、ZNF416、ZNF417、ZNF418、ZNF419、ZNF423、ZNF425、ZNF429、ZNF430、ZNF431、ZNF432、ZNF433、ZNF436、ZNF439、ZNF44、ZNF440、ZNF441、ZNF442、ZNF443、ZNF444、ZNF445、ZNF449、ZNF45、ZNF454、ZNF460、ZNF467、ZNF468、ZNF479、ZNF480、ZNF483、ZNF484、ZNF485、ZNF486、ZNF487、ZNF490、ZNF492、ZNF496、ZNF501、ZNF502、ZNF506、ZNF513、ZNF519、ZNF524、ZNF525、ZNF527, ZNF528, ZNF529, ZNF530, ZNF534, ZNF540, ZNF543, ZNF547, ZNF548, ZNF549, ZNF550, ZNF552, ZNF554, ZNF555, ZNF558, ZNF561, ZNF562, ZNF563, ZNF564, ZNF565, ZNF566, ZNF567, ZNF570, ZNF571, ZNF573, ZNF574, ZNF580, ZNF582, ZNF584, ZNF585A, ZNF586, ZNF587, ZNF594, ZNF595, ZNF596, ZNF597, ZNF605, ZNF610, ZNF611, ZNF613, ZNF614, ZNF615, ZNF616, ZNF619, ZNF620, ZNF621, ZNF626, ZNF627, ZNF641, ZNF652, ZNF653, ZNF655, ZNF660, ZNF662, ZNF667, ZNF669, ZNF671, ZNF674, ZNF675, ZNF677, ZNF680, ZNF681, ZNF682, ZNF684, ZNF69, ZNF692, ZNF695, ZNF7, ZNF701, ZNF704, ZNF705G, ZNF707, ZNF708, ZNF71, ZNF711, ZNF713, ZNF714, ZNF716, ZNF730, ZNF736, ZNF737, ZNF74, ZNF740, ZNF749, ZNF75A, ZNF75D, ZNF76, ZNF764, ZNF765, ZNF766, ZNF768, ZNF77, ZNF770, ZNF771, ZNF774, ZNF776, ZNF777, ZNF778, ZNF780A, ZNF782, ZNF783, ZNF784, ZNF785, ZNF786, ZNF787, ZNF789, ZNF79, ZNF790, ZNF791, ZNF792, ZNF793, ZNF799, ZNF8, ZNF805, ZNF808, ZNF81, ZNF816, ZNF821, ZNF823, ZNF84, ZNF85, ZNF860, ZNF879, ZNF880, ZNF891, ZNF90, ZNF93, ZNF98, ZSCAN1, ZSCAN16, ZSCAN22, ZSCAN23, ZSCAN29, ZSCAN30, ZSCAN31, ZSCAN4, ZSCAN5, ZSCAN5C, ZSCAN9, and ZZZ3, etc., and other transcription factors that may be identified in the literature.
[0062] Those skilled in the art can easily identify transcription factors using databases and literature sources.In addition to the above databases, transcription factors are also identified in US2022 / 0214356, which is hereby incorporated by reference in its entirety to describe transcription factors.
[0063] Double-stranded DNA deaminase
[0064] According to one aspect, exemplary double-stranded DNA deaminases include DddA, also known in the art as BadTF1. See Mok, BY et al., (2020). A bacterial cytidine deaminase toxin enables CRISPR-free mitochondrial base editing. Nature, 583(7817), 631-637, which is hereby incorporated by reference in its entirety to teach DddA. DddA (BadTF1) belongs to the SCP1.201-like deaminase subfamily.
[0065] According to one aspect, exemplary double-stranded DNA deaminases include BadTF3. See Marcos H de Moraes et al., (2021) An interbacterial DNA deaminase toxin directly mutagenizes surviving target populations eLife 10: e62967. BadTF3 belongs to the Pput_2613-like deaminase subfamily.
[0066] According to one aspect, exemplary double-stranded DNA deaminases include DddA11. See Mok, BY, et al., (2022). CRISPR-free base editors with enhanced activity and expanded targeting scopein mitochondrial and nuclear DNA. Nature Biotechnology, 1-10, which is hereby incorporated by reference in its entirety to teach DddA11.
[0067] Deaminases known in the art have been classified for identification. See Iyer LM, et al., Evolution of the deaminase fold and multiple origins of eukaryotic editing and mutagenic nucleic acid deaminases from bacterial toxin systems [J]. Nucleic acids research, 2011, 39 (22): 9473-9497, which is hereby incorporated by reference in its entirety for the classification and identification of deaminases. The DddA of the present disclosure belongs to the SCP1.201 clade. The BadTF3 of the present disclosure belongs to the Pput_2613-like clade. It should be understood that the specific dsDNA deaminases described herein are only exemplary. The present disclosure contemplates mutants, variants, derivatives and modifications of dsDNA deaminases exhibiting enzymatic activity. The present disclosure contemplates the identification of other dsDNA deaminases and mutants, variants, derivatives and modifications thereof exhibiting enzymatic activity by those skilled in the art.
[0068] In some examples, the double-stranded DNA deaminase may be derived from a DddA deaminase, such as DddA6, DddA7 and other DddA variants. In other examples, the double-stranded DNA deaminase may be derived from BadTF2 and other BadTF2 variants. In other examples, the double-stranded DNA deaminase may be derived from BadTF3 and other BadTF3 variants. In further examples, the deaminase may be derived from a bacterial toxin, such as a Pput_2613 family deaminase, a SCP1.201-like family deaminase, a DYW-like family deaminase, a BURPS668_1122-like family deaminase, a YwqJ-like family deaminase, a MafB19-like family deaminase, a sce3516-like family deaminase, a BH3703-like deaminase, a WD0512-like family deaminase, etc.
[0069] It should be understood that aspects of the present disclosure include mutants, variants, truncations, modifications and derivatives of full-length dsDNA deaminases described herein and known to those skilled in the art that exhibit deaminase activity. Methods of mutating, altering, truncating, modifying or deriving known dsDNA deaminases are known to those skilled in the art. Therefore, the present disclosure contemplates that dsDNA deaminases are enzymes that exhibit deaminase activity against double-stranded nucleic acids, whether naturally occurring or modified, mutated, altered, truncated, derived or evolved forms of naturally occurring dsDNA deaminases. Aspects of the present disclosure provide nucleic acid and amino acid sequences of various known dsDNA deaminases. Embodiments of the present disclosure include nucleic acid and amino acid sequences having 75% homology, 80% homology, 85% homology, 90% homology, 91% homology, 92% homology, 93% homology, 94% homology, 95% homology, 96% homology, 97% homology, 98% homology, 99% homology, 99.5% homology, 99.6% homology, 99.7% homology, 99.8% homology, or 99.9% homology to the full-length sequence of the dsDNA deaminases disclosed herein. Based on the present disclosure, a skilled artisan can identify dsDNA deaminases having a percentage homology to a dsDNA deaminase known to have deaminase activity, or otherwise test such dsDNA deaminases having a percentage homology to a dsDNA deaminase known to have deaminase activity.
[0070] In vitro transposition
[0071] According to certain aspects, chromatin DNA treated with double-stranded DNA deaminase is treated using a transposition method, which may be referred to in the art as transposome-mediated fragmentation or "fragmentation tagging (tagmentation)"). In the fragmentation tagging method, a transposome is prepared with DNA, and then the DNA is cut so that the transposition event produces fragmented DNA with adapters. In such methods, the target DNA is simultaneously fragmented and labeled, producing fragments labeled with the desired DNA sequence for downstream processing. According to one aspect, a library is generated using an in vitro transposition system used by Illumina, Inc's Nextera technology to simultaneously fragment DNA and tag each fragment with an appropriate sequence for next-generation sequencing. See US20110287435, which is incorporated herein by reference to disclose fragmentation tagging methods. For other useful methods for making sequencing libraries, see Single-cell chromatin accessibility reveals principles of regulatory variation. Nature, 523(7561), 486-490)(2007); Massively multiplex single-cell Hi-C. Nature Methods, 14(3), 263-266)(2017); also see the useful transposition methods described in WO2016 / 073690 and WO2018217912, which are hereby incorporated by reference in their entirety.
[0072] According to certain aspects, exemplary transposase systems include Tn5 transposase, Mu transposase, Tn7 transposase or IS5 transposase, etc. Other useful transposase systems are known to those skilled in the art, including Tn3 transposase system (see Maekawa, T., Yanagihara, K., and Ohtsubo, E. (1996), A cell-free system of Tn3 transposition and transposition immunity, Genes Cells 1, 1007-1016), Tn7 transposase system (see Craig, NL (1991), Tn7: a target site-specific transposon, Mol. Microbiol. 5, 2569-2573), Tn10 transposase system (see Chalmers, R., Sewitz, S., Lipkow, K., and Crellin, P. (2000), Complete nucleotide sequence of Tn10, J. Bacteriol 182, 2970-2972), Piggybac transposon system (see Li, X., Burnight, ER, Cooney, AL, Malani, N., Brady, T., Sander, JD, Staber, J., Wheelan, SJ, Joung, JK, McCray, PB, Jr., et al. (2013), PiggyBac transposase tools for genome engineering, Proc. Natl. Acad. Sci. USA 110, E2279-2287), Sleeping Beauty transposon system (see Ivics, Z., Hackett, PB, Plasterk, RH, and Izsvak, Z. (1997), Molecular reconstruction of Sleeping Beauty, a Tc1-like transposon from fish, and its transposition in human cells, Cell 91, 501-510), Tol2 transposon system (see Kawakami, K. (2007), Tol2: a versatile gene transfer vector invertebrates, Genome Biol. 8 Suppl. 1, S7.)
[0073] According to general aspects, treated genomic DNA contacts with Tn5 transposase, and each transposase is combined with transposon DNA to form a transposase / transposon DNA complex dimer that is called as transposome. Transposome is attached to the target position on treated genomic DNA, and treated genomic DNA is cut into multiple double-stranded fragments with primer binding site. Processing such as extension and gap filling can be carried out to produce double-stranded product, which is mixed with primer and archaeal dna polymerase, nucleotide and amplification reagent, and double-stranded treated genomic DNA fragment is amplified. High-throughput sequencing method known to those skilled in the art is used to order-check amplicon.
[0074] Nuclease
[0075] According to one aspect, nucleases can be used to enrich target chromatin regions, such as open chromatin regions, before or after treatment with dsDNA deaminase. Nucleases are enzymes that can cut phosphodiester bonds between nucleic acid nucleotides. Nucleases affect single-strand and double-strand breaks in their target molecules in various ways. According to the active site, there are two main classifications. Exonucleases digest nucleic acids from the ends. Endonucleases act on the region in the middle of the target molecule. According to certain aspects, nucleases can be used to process DNA into fragments. This processing can be performed before or after treating DNA with double-stranded DNA deaminase. Such fragments can then be amplified and / or sequenced as described herein. Exemplary nucleases include DNase (e.g., DNaseI commercially available from Thermo Fisher), MNase (micrococcal nuclease commercially available from New England Biolabs) or restriction endonucleases (e.g., FASTDIGEST commercially available from Thermo Fisher) etc. According to one aspect, lysed cells can be subjected to dsDNA deaminase treatment followed by nuclease cleavage to enrich for open chromatin regions for sequencing.
[0076] Enrichment of target chromatin regions
[0077] According to one aspect, in addition to methods using a transposase (e.g., Tn5 transposase), methods known to those skilled in the art (including ChIP-seq, ChIC, ChEC, ChEC-seq, CUT&TAG, CUT&RUN, Multi-CUT&Tag, NTT-seq, R loop CUT&Tag, etc.) can also be used to enrich target chromatin regions (e.g., open chromatin regions) before or after treatment with dsDNA deaminase.
[0078] ChIP sequencing (also known as ChIP-seq) is a method for analyzing protein-DNA interactions. ChIP-seq combines chromatin immunoprecipitation (ChIP) with massively parallel DNA sequencing to identify binding sites for DNA-associated proteins. It can be used to accurately map the global binding sites of any protein of interest. Specific DNA sites that directly physically interact with transcription factors and other proteins can be isolated by chromatin immunoprecipitation. ChIP can generate a library of target DNA sites bound to the protein of interest. Massively parallel sequence analysis is used in conjunction with a whole genome sequence database to analyze the interaction pattern of any protein with DNA (see Johnson DS, et al., (June 2007). "Genome-wide mapping of in vivo protein-DNA interactions" (PDF). Science. 316 (5830): 1497-502) or any pattern of epigenetic chromatin modification. This can be applied to transcription factor groups that can be ChIPed. See Whole-Genome Chromatin IP Sequencing (ChIP-Seq) "(PDF). Illumina, Inc. 26 November 2007.
[0079] CUT&Tag sequencing, also known as cleavage under targets and tagmentation, is a method for analyzing protein-DNA interactions. CUT&Tag sequencing combines antibody-targeted cleavage mediated by a protein A-Tn5 fusion with massively parallel DNA sequencing to identify binding sites for DNA-associated proteins. It can be used to accurately map global DNA binding sites for any protein of interest. See "CUT&Tag: a higher resolution, lower cost way to map chromatin". Fred Hutchinson Cancer Research Center. 29 April 2019.
[0080] CUT&RUN sequencing (see US2022 / 0214356), also known as cleavage under targets and release using nuclease, is a method for analyzing protein-DNA interactions. CUT&RUN sequencing combines micrococcal nuclease-mediated antibody-targeted cleavage with massively parallel DNA sequencing to identify binding sites for DNA-associated proteins. It can be used to accurately map the global DNA binding sites of any protein of interest. See "Lay off the ChIPs: CUT&RUN instead". Fred Hutchinson Cancer Research Center. 20 February 2017.
[0081] ChIC detects the binding sites of transcription factors in the genome by targeting a modified micrococcal nuclease (MNase) (pA-MN) conjugated to protein A using a specific antibody. The modified MNase specifically cuts DNA at the region that interacts with the protein of interest only in the presence of Ca2+ ions, thus allowing controlled DNA cutting at the antibody binding site. This method allows mapping proteins with a resolution of 100-200bp and excellent specificity. See Schmid M, Durussel T, Laemmli UK. 2004. ChIC and ChEC. Molecular Cell. 16 (1): 147-15.
[0082] Chromatin endogenous cleavage (ChEC) can be combined with high-throughput sequencing in a method called ChEC-seq. ChEC-seq relies on the fusion of a chromatin-associated protein of interest to micrococcal nuclease (MNase) to produce targeted DNA cleavage in the presence of calcium in living cells. ChEC-seq is not based on immunoprecipitation and therefore avoids potential issues with cross-linking, sonication, chromatin solubilization, and antibody quality while providing high-resolution mapping and minimal background signal. See Grunberg et al., J. Vis. Exp. 2017; (124) e55836, p. 1-9.
[0083] Other methods for enriching target chromatin regions are known to those skilled in the art and can be readily identified by searching the literature.
[0084] According to one aspect, the lysed cells are treated with dsDNA deaminase and then treated with ChIP-seq, ChIC, ChEC, ChEC-seq, CUT&TAG, CUT&RUN, Multi-CUT&Tag, NTT-seq, R loop CUT&Tag to enrich for chromatin regions targeted by specific binding agents (e.g., antibodies). Alternatively, the lysed cells can first be treated with ChIP-seq, ChIC, ChEC, ChEC-seq, CUT&TAG, CUT&RUN, Multi-CUT&Tag, NTT-seq, R loop CUT&Tag to enrich for target chromatin regions without amplification and sequencing, and then dsDNA deaminase treatment can be applied to map TF footprints in the target region.
[0085] Amplification
[0086] In some aspects, amplification is achieved using PCR. PCR is a reaction in which a pair of primers or a group of primers consisting of upstream and downstream primers and a polymerization catalyst (e.g., DNA polymerase, typically a thermostable polymerase) are used to produce a copy of the target polynucleotide. The PCR method is well known in the art and is taught, for example, in MacPherson et al. (1991) PCR 1: A Practical Approach (IRL Press at Oxford University Press). The term "polymerase chain reaction" ("PCR") of Mullis (U.S. Patent Nos. 4,683,195, 4,683,202 and 4,965,188) refers to a method for increasing the concentration of target sequence fragments without cloning or purification. The process of amplifying the target sequence includes providing an oligonucleotide primer and an amplification reagent with a desired target sequence, and then performing an accurate thermal cycling sequence in the presence of a polymerase (e.g., DNA polymerase). The primer is complementary to the respective chains ("primer binding sequence") of its double-stranded target sequence. Typically, to perform amplification, a double-stranded target sequence is denatured and then primers are annealed to their complementary sequences within the target molecule. After annealing, the primers are extended with a polymerase to form a new pair of complementary strands. The steps of denaturation, primer annealing, and polymerase extension can be repeated multiple times (i.e., denaturation, annealing, and extension constitute one "cycle"; there can be multiple "cycles") to obtain a high concentration of amplified fragments of the desired target sequence. The length of the amplified fragment of the desired target sequence is determined by the relative positions of the primers to each other, and therefore, the length is a controllable parameter. Due to the repetitive nature of the process, the method is referred to as the "polymerase chain reaction" (hereinafter referred to as "PCR"), and the target sequence is described as "PCR amplified".
[0087] The terms "PCR product", "PCR fragment" and "amplification product" refer to a mixture of compounds obtained after completing two or more cycles of PCR steps (denaturation, annealing and extension). These terms include situations where one or more fragments of one or more target sequences are amplified.
[0088] Any oligonucleotide or polynucleotide sequence can be amplified using an appropriate set of primer molecules. Methods and kits for performing PCR are well known in the art. All processes (e.g., PCR or gene cloning) that produce replicate copies of polynucleotides are collectively referred to herein as replication.
[0089] The expression "amplification" or "amplifying" refers to the process of forming additional or multiple copies of a specific polynucleotide. Amplification includes methods such as PCR, ligation amplification (or ligase chain reaction, LCR) and other amplification methods. These methods are known in the art and are widely practiced. See, for example, U.S. Patent Nos. 4,683,195 and 4,683,202 and Innis et al., "PCR protocols: a guide to method and applications" Academic Press, Incorporated (1990) (about PCR); and Wu et al. (1989) Genomics 4: 560-569 (about LCR). In general, the PCR procedure describes a gene amplification method that includes (i) sequence-specific hybridization of primers to specific genes in a DNA sample (or library), (ii) subsequent amplification including multiple rounds of annealing, extension and denaturation using a DNA polymerase, and (iii) screening the PCR products for bands of the correct size. The primers used are oligonucleotides of sufficient length and appropriate sequence to provide initiation of polymerization, ie, each primer is specifically designed to be complementary to each strand of the genomic locus to be amplified.
[0090] Reagents and hardware for performing amplified reactions are commercially available. Primers that can be used to amplify sequences from specific gene regions are preferably complementary and specifically hybridized to sequences in the target region or its flanking regions, and can be prepared using methods known to those skilled in the art. The nucleotide sequence produced by amplification can be directly sequenced.
[0091] When hybridization between two single-stranded polynucleotides occurs in an antiparallel configuration, the reaction is called "annealing" and the polynucleotides are described as "complementary." A double-stranded polynucleotide is complementary or homologous to another polynucleotide if hybridization can occur between one strand of the first polynucleotide and one strand of the second polynucleotide. Complementarity or homology (the degree to which one polynucleotide is complementary to another) can be quantified based on the proportion of bases in opposing strands that are expected to hydrogen bond with each other according to the generally accepted base pairing rules.
[0092] The term "amplification reagent" may refer to those reagents (deoxyribonucleotide triphosphates, buffers, etc.) required for amplification except primers, nucleic acid templates, and amplification enzymes. Typically, the amplification reagent is placed together with other reaction components and contained in a reaction vessel (test tube, microwell, etc.). Amplification methods include PCR methods known to those skilled in the art, and also include rolling circle amplification (Blanco et al., J. Biol. Chem., 264, 8935-8940, 1989), hyperbranched rolling circle amplification (Lizard et al., Nat. Genetics, 19, 225-232, 1998) and loop-mediated isothermal amplification (Notomi et al., Nuc. Acids Res., 28, e63, 2000), each of which is incorporated herein by reference in its entirety.
[0093] Other amplification methods may be used according to the present disclosure, as described in UK Patent Application No. GB2,202,328 and PCT Patent Application No. PCT / US89 / 01025, each of which is incorporated herein by reference. Emulsion PCR may be used according to the present disclosure. Other suitable amplification methods include "race and one-sided PCR" (Frohman, In: PCR Protocols: A Guide To Methods And Applications, Academic Press, NY, 1990, incorporated herein by reference). Methods based on connecting two (or more) oligonucleotides in the presence of a nucleic acid with a resulting "di-oligonucleotide (di-oligonucleotide)" sequence to amplify di-oligonucleotides may also be used to amplify DNA according to the present disclosure (Wu et al., Genomics 4: 560-569, 1989, incorporated herein by reference).
[0094] The RNA to be amplified can be obtained from a single cell or a small group of cells. The methods described herein allow RNA to be amplified from any species or organism in a reaction mixture, such as a single reaction mixture performed in a single reaction vessel. In one aspect, the methods described herein include sequence-independent amplification of RNA from any source, including but not limited to human, animal, plant, yeast, virus, eukaryotic, and prokaryotic RNA.
[0095] The term "primer" as used herein generally includes natural or synthetic oligonucleotides that can serve as a starting point for nucleic acid synthesis after forming a double strand with a polynucleotide template, such as a sequencing primer, and extend along the template from its 3' end to form an extended double strand. Primers include extension primers, amplification primers or reverse transcription primers.
[0096] The nucleotide sequence added during the extension process is determined by the sequence of the template polynucleotide. Usually the primer is extended by a DNA polymerase or a reverse transcriptase. The length of the primer is usually 3 to 36 nucleotides, or 5 to 24 nucleotides, or 14 to 36 nucleotides. Primers within the scope of the present invention include orthogonal primers, amplification primers, construction primers, etc. Primer pairs can be located on both sides of a sequence of interest or a group of sequences of interest. The sequences of primers and probes can be degenerate or quasi-degenerate. Primers within the scope of the present invention are adjacent to the target sequence. "Primers" can be considered as short polynucleotides, usually with a free 3'-OH group, which binds to a target or template that may be present in a sample of interest by hybridizing with the target, and then promotes the polymerization of polynucleotides complementary to the target. Primers of the present invention include 17 to 30 nucleotides. In one aspect, the primer is at least 17 nucleotides, or alternatively at least 18 nucleotides, or alternatively at least 19 nucleotides, or alternatively at least 20 nucleotides, or alternatively at least 21 nucleotides, or alternatively at least 22 nucleotides, or alternatively at least 23 nucleotides, or alternatively at least 24 nucleotides, or alternatively at least 25 nucleotides, or alternatively at least 26 nucleotides, or alternatively at least 27 nucleotides, or alternatively at least 28 nucleotides, or alternatively at least 29 nucleotides, or alternatively at least 30 nucleotides, or alternatively at least 50 nucleotides, or alternatively at least 75 nucleotides, or alternatively at least 100 nucleotides.
[0097] Particularly exemplary amplification methods include rolling circle amplification (RCA); multiple displacement amplification (MDA); loop-mediated isothermal amplification (LAMP); strand displacement amplification (SDA, see US5,744,311); nucleic acid sequence-based amplification (NASBA, see US6,025,134); quantitative real-time PCR; reverse transcription PCR (RT-PCR); real-time PCR (rt PCR); real-time reverse transcription PCR (rt RT-PCR); nested PCR; transcription-free isothermal amplification (see US6,033,881), repair chain reaction amplification (see WO90 / 01069); ligase chain reaction amplification (see European Patent Publication EP-A-320308); gap-filling ligase chain reaction amplification (see US5,427,930); coupled ligase detection and PCR (see US6,027,889).
[0098] Sequencing
[0099] The amplicons are sequenced using, for example, high throughput sequencing methods known to those skilled in the art. The sequence of the nucleic acid sequence of interest can be determined using a variety of sequencing methods known in the art, including, but not limited to, sequencing by hybridization (SBH), sequencing by ligation (SBL) (Shendure et al. (2005) Science 309:1728), quantitative incremental fluorescent nucleotide addition sequencing (QIFNAS), stepwise ligation and cleavage, fluorescence resonance energy transfer (FRET), molecular beacons, TaqMan reporter probe digestion, pyrosequencing, fluorescent in situ sequencing (FISSEQ), FISSEQ beads (U.S. Pat. No. 7,425,431), wobble sequencing (PCT / US05 / 27695), multiplex sequencing (U.S. Ser. No. 12 / 027,039, filed Feb. 6, 2008; Porreca et al. (2007) Nat. Methods 4:931), polymerized colonies (FISSEQ), and sequencing by ligation (SBL) (Shendure et al. (2005) Science 309:1728). colony (POLONY) sequencing (U.S. Pat. Nos. 6,432,360, 6,485,944 and 6,511,803 and PCT / US05 / 06425); nanogrid rolling circle sequencing (ROLONY) (U.S. Serial No. 12 / 120,541, filed May 14, 2008), allele-specific oligonucleotide ligation assays (e.g., oligonucleotide ligation assays (OLA), single template molecule OLA read out using ligated linear probes and rolling circle amplification (RCA), ligated padlock probes and / or single template molecule OLA read out using ligated circular padlock probes and rolling circle amplification (RCA), etc. High-throughput sequencing methods can also be utilized, for example, using platforms such as Roche 454, Illumina Solexa, AB-SOLiD, Helicos, Polonator platforms, Ion Torrent semiconductor sequencing technology, Pacific Biosciences' single molecule real-time (SMRT) sequencing, Oxford Nanopore Technologies' nanopore-based sequencing, and the like.A variety of light-based sequencing technologies are known in the art (Landegren et al. (1998) Genome Res. 8:769-76; Kwok (2000) Pharmacogenomics 1:95-100; and Shi (2001) Clin. Chem. 47:164-172). Exemplary sequencing platforms that can be used in the present disclosure and suitable for the methods described herein are described in Reuter et al., High-Throughput Sequencing Technologies, Mol. Cell (2015); 58(4):586-597, which is hereby incorporated by reference in its entirety.
[0100] The amplified DNA can be sequenced by any suitable method. Specifically, the amplified DNA can be sequenced using a high-throughput screening method, such as the SOLiD sequencing technology of Applied Biosystems or the genome analyzer of Illumina. In one aspect of the invention, the amplified DNA can be subjected to shotgun sequencing. The number of read lengths can be at least 10,000, at least 1 million, at least 10 million, at least 100 million, or at least 1 billion. On the other hand, the number of read lengths can be 10,000 to 100,000, or 100,000 to 1 million, or 1 million to 10 million, or 10 million to 100 million, or 100 million to 100 million, or 100 million to 100 million." Read length" is the length of the continuous nucleic acid sequence obtained by sequencing reaction.
[0101] "Shotgun sequencing" refers to a method for sequencing large amounts of DNA, such as an entire genome. In this method, the DNA to be sequenced is first chopped into smaller fragments that can be sequenced individually. The sequences of these fragments are then reassembled into their original order based on their overlapping sequences, thereby generating a complete sequence. The "chopped-up" of DNA can be accomplished using a variety of different techniques, including restriction enzyme digestion or mechanical shearing. Overlapping sequences are typically aligned by a suitably programmed computer. Methods and procedures for shotgun sequencing DNA libraries are well known in the art.
[0102] Particularly exemplary sequencing methods include Sanger sequencing (AB 13730x1 Genome Analyzer), pyrosequencing on a solid support (454 sequencing, Roche), sequencing by synthesis with reversible termination (ILLUMINA Genome Analyzer), DNA nanoball sequencing (DNBSEQ, MGI), sequencing by ligation (ABI SOLID), or sequencing by synthesis with dummy terminators (HELI ). Other next generation sequencing technologies used with the disclosed methods include massively parallel signature sequencing (MPSS), Polony sequencing, Ion Torrent semiconductor sequencing, DNA nanoball sequencing, Heliscope single molecule sequencing, single molecule real-time (SMRT) sequencing, Pacbio sequencing, and Nanopore DNA sequencing.
[0103] It should be understood that the embodiments of the present invention described only illustrate some applications of the principles of the present invention. Those skilled in the art can make many modifications based on the teachings proposed herein without departing from the true spirit and scope of the present invention. The contents of all references, patents and published patent applications cited in this application are hereby incorporated by reference in their entirety for all purposes.
[0104] The following examples are shown to be representative of the present invention. These examples should not be construed as limiting the scope of the invention, as these and other equivalent embodiments will be apparent based on the present disclosure, drawings and appended claims.
[0105] Embodiment 1
[0106] Plasmid construction
[0107] Expression constructs of DddA toxin domain (DddAtox), DddA11 and BadTF3 were obtained from Genescript via gene synthesis service.
[0108] To generate a pETDuet-1 based expression construct of DddAtox, the DddAtox coding sequence: ATGGGCAGCAGCCATCACCATCATCACCACAGCCAGGATCCGGGTAGCTATGCGCTGGGTCCGTATCAGATCTCTGCTCCGCAGCTGCCGGCATATAACGGTCAGACTGTTGGTACTTTCTATTATGTTAACGATGCTGGCGGTTTAGAAAGCAAAGTTTTCAGCTCTGGTGGTCCGACCCCGTATCCGAACTATGCTAACGCTGGTCAC GTTGAAGGTCAGTCTGCTCTGTTCATGCGTGATAACGGTATCTCTGAAGGTCTGGTTTTCCATAACAACCCGGAAGGTACCTGTGGTTTTTGTGTTAACATGACCGAAACCCTGCTGCCGGAAAAC GCTAAAATGACCGTTGTTCCGCCGGAAGGTGCGATTCCGGTTAAACGTGGTGCTACCGGTGAAACCAAAGTTTTCACCGGTAACTCTAACTCTCCGAAATCTCCGACCAAAGGTGGTTGCTAA(SEQ ID NO: 1) (whose translated protein sequence is: MGSSHHHHHHSQDPGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNSNSPKSPTKGGC) (SEQ ID NO: 2) was synthesized and cloned into MCS-1 (NcoI and HindIII sites, retaining the N-terminal hexahistidine tag).
[0109] Immunity protein DddAI coding sequence: ATGTATGCGGATGACTTTGACGGGGAAATTGAGATTGATGAAGTTGATAGCCTAGTTGAGTTTCTGAGCCGTCGTCCGGCGTTCGATGCGAACAACTTCGTTCTGACCTTCGAAGAAAGCGGCTTCCCGCAGCTGAACATCTTCGCGAAAAACGATATCGCGGTTGTTTACTACATGGATATCGGCGAAAACTTCGTTAGCAAAGGCAACAGGCGAGCGGCGGCACCGAAAAATTCTACGAAAACAAACTGGGCGGCGAAGTTGATCTGAGCAAAGATTGCGTTGTTAGCAAAGAACAGATGATCGAAGCGGCGAAACAGTTCTTCGCGACCAAACAGCGTCCGGAACAGCTGACCTGGAGCGAACTGTAA (SEQ ID NO: 3) (the translated protein sequence is: MYADFDGEIEIDEVDSLVEFLSRRPAFDANNFVLTFEESGFPQLNIFA KNDIAVVYYMDIGENFVSKGNSASGGTEKFYENKLGGEVDLSKDCVVSKEQMIEAAKQFFATKQRPEQLTWSEL) (SEQ ID NO: 4) was synthesized and cloned into MCS-2 (NdeI and XhoI sites, removal of the C-terminal S tag).
[0110] To generate a pETDuet-1 based expression construct of DddA11, the DddA11 coding sequence: ATGGGCAGCAGCCATCACCATCATCACCACAGCCAGGATCCGGGTAGCTATGCGCTGGGTCCGTATCAGATCTCTGCTCCGCAGCTGCCGGCATATAACGGTCAGACTGTTGGTACTTTCTATTATGTTAACGATGCTGGCGGTTTAGAAAGCAAAGTTTTCATCTCTGGTGGTCCGACCCCGTATCCGAACTATGTTAGCGCTGGTCACG TTGAAGGTCAGTCTGCTCTGTTCATGCGTGATAACGGTATCTCTGAAGGTCTGGTTTTCCATAACAACCCGAAAGGTACCTGTGGTTTTTGTGTTAACATGATCGAAACCCTGCTGCCGGAAAACG CTAAAATGACCGTTGTTCCGCCGGAAGGTGCGATTCCGGTTAAACGTGGTGCTACCGGTGAAACCAAAGTTTTCATCGGTAACTCTAACTCTCCGAAATCTCCGACCAAAGGTGGTTGCTAA(SEQ ID NO: 5) (whose translated protein sequence is: MGSSHHHHHHSQDPGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFISGGPTPYPNYVSAGHVEGQSALFMRDNGISEGLVFHNNPKGTCGFCVNMIETLLPENAKMTVVPPEGAIPVKRGATGETKVFIGNSNSPKSPTKGGC) (SEQ ID NO: 6) was synthesized and cloned into MCS-1 (BamHI and HindIII sites, retaining the N-terminal hexa-histidine tag),And the immune protein DddAI coding sequence: ATGTATGCGGATGACTTTGACGGGAAATTGAGATTGATGAAGTTGATAGCCTAGTTGAGTTTCTGAGCCGTCGTCCGGCGTTCGATGCGAACAACTTCGTTCTGACCTTCGAAGAAAGCGGCTTCCCGCAGCTGAACATCTTCGCGAAAAACGATATCGCGGTTGTTTACTACATGGATATCGGCGAAAACTTCGTTAGCAAAGGCAACAGGCGAGCGGCGGCACCGAAAAATTCTACGAAAACAAACTGGGCGGCGAAGTTGATCTGAGCAAAGATTGCGTTGTTAGCAAAGAACAGATGATCGAAGCGGCGAAACAGTTCTTCGCGACCAAACAGCGTCCGGAACAGCTGACCTGGAGCGAACTGTAA (SEQ ID NO: 7) (the translated protein sequence is: MYADFDGEIEIDEVDSLVEFLSRRPAFDANNFVLTFEESGFPQLNIFA KNDIAVVYYMDIGENFVSKGNSASGGTEKFYENKLGGEVDLSKDCV VSKEQMIEAAKQFFATKQRPEQLTWSEL) (SEQ ID NO: 8) was synthesized and cloned into MCS-2 (NdeI and XhoI sites, C-terminal S tag removed).
[0111] To generate the pETDuet-1-based expression construct of BadTF3, the BadTF3 coding sequence: ATGGGCAGCAGCCATCACCATCATCACCACAGCCAGGATCCTGGTTGGAAATTTTCTAACGGTAAACGCCGTCCGCCGCACAAAGCAACGGTAACTGTGACCGATAAAAACGGTGTCGTTAAACACAAAAGCAACCTGGTTTCTGGCAACATGACTGAAGCCGAAAAGAAACTGGGCTTCCCGAAC AACTCCCTGGCGACCCACACCGAAAACCGTGCTACCCGCCTGATCGATCTGAACCAAGGTGATACTATGCTGATCGAGGGCCAATACCGTCCGTGTCCACGTTGTAAAGGTGCAATGCGCGTGAAAGCGGAGGAATCCGGTGCGAAAGTGATCTACACCTGGCCAGAAGATGGTGACCTGAAAAAACGTGAATGGGAAGGCACTCCGTGCGACAAAAAAATAA(SEQ ID NO:9) (whose translated protein sequence is: MGSSHHHHHHSQDPGWKFSNGKRRPPHKATVTVTDKNGVVKHKSN LVSGNMTEAEKKLGFPNNSLATHTENRATRLIDLNQGDTMLIEGQY RPCPRCKGAMRVKAEESGAKVIYTWPEDGDLKKREWEGTPCDKK) (SEQ ID NO:10) was synthesized and cloned into MCS-1 (NcoI and HindIII sites, retaining the N-terminal hexahistidine tag).
[0112] Immune protein BadTF3I coding sequence: ATGACCAAATCTAAAATGCTGAGCAACATCGTCATCCAGGAGGTCAAATTTGCGATCGAAGATTACTGCGCTATTCTGAGCTTCGCTTCTGACTCTTATGAAGTGCCGGAGCAGTATTTTATCATTACCCGTTCTACCACCGAACGTTCTGGCGGTATTCCGGAGGGCGAC ATCTACCTGGAATCTAACCTGTTTCTGGATTTTAACCCGTACGGCCTGAGCGGTTACCTGCTGTCTGAGCCGAACTGCGTAGATCTGCTGATCGAACCGAACAACTACGTTCGTCTGCGTCTGATCGAAAAAATCGATATCCTGGAAGTGGAAAACCACCTGAAATTTCTGTTCGACAACTAA(SEQ ID NO: 11) (The translated protein sequence is: MTKSKMLSNIVIQEVKFAIEDYCAILSFASDSYEVPEQYFIITRSTTER SGGIPEGDIYLESNLFLDFNPYGLSGYLLSEPNCVDLLIEPNNYVRLRLIEKIDILEVENHLKFLFDN) (SEQ ID NO: 12) was synthesized and cloned into MCS-2 (NdeI and KpnI sites, removing the C-terminal S tag). The pETDuet-1::dddAtox+dddAI vector and pETDuet-1::badTF3tox+badTF3I were transformed into E. coli strains DH5α and BL21 and stored at -20°C. Figure 2 The vector map of pETDuet-1::dddAtox+dddAI is depicted. The inserted gene is driven by two independent lac operons. Figure 3 The vector map of pETDuet-1::badTF3tox+badTF3I is depicted. The inserted gene is driven by two independent lac operons. Figure 4 The vector map of pETDuet-1::dddA11+dddAI is depicted. The inserted gene is driven by two independent lac operons.
[0113] Example II
[0114] Bacterial strains and culture conditions
[0115] Escherichia coli (E. coli) strains were grown at 37°C in lysing broth (LB) or in LB medium solidified with agar (LBA, 1.5% w / v). If necessary, the medium was supplemented with ampicillin (100 μg / ml) or IPTG (0.5 mM). E. coli strains DH5α and BL21 were used for plasmid maintenance and protein expression, respectively.
[0116] Example III
[0117] Purification of DddAtox
[0118] Purification of DddAtox protein has been previously reported (Beverly et al. 2020, Nature, 583(7817): 631-637doi: 10.1038 / s41586-020-2477-4). In short, to purify his-tagged DddAtox complexed with DddAI, 2L LB broth was inoculated with E. coli BL21 (pETDuet-1::dddAtox+dddAI) at a dilution of 1:100 and cultured overnight. After the culture grew to about OD600 = 0.6, 0.5 mM isopropyl β-D-1-thiogalactoside (IPTG) was added and incubated at 18 ° C for 16 hours with shaking. The bacterial cell pellet was collected by centrifugation at 4000g for 30 minutes, and then resuspended in 50ml lysis buffer (50mM Tris-HCl pH8.0, 500mM NaCl, 10mM imidazole, 1mg / mL lysozyme and protease inhibitor mixture). The bacterial cell pellet was then lysed by ultrasonic treatment (5 pulses, 10 seconds each time), and the supernatant was separated from the fragments by centrifugation at 25,000g for 30 minutes. Nickel columns were used to purify the DddAtox-DddAI complex of his tag from the supernatant. DddAtox-DddAI was eluted with 1 elution buffer (50mM Tris-HCl pH7.5, 300mM imidazole, 500mM NaCl, 30mM imidazole, 1mM DTT). The DddAtox-DddAI complex of elution was denatured by adding 50ml 8M urea denaturation buffer (50mM Tris-HCl pH7.5, 300mM imidazole, 500mM NaCl and 1mM DTT) at 4°C for 16 hours. The denatured protein in the 8M urea denaturation buffer was loaded onto the nickel column again. The column was washed with 50ml 8M urea denaturation buffer to exclude any residual DddAI. The column was sequentially washed with 25ml denaturation buffer with urea (6M, 4M, 2M, 1M) that gradually decreased in concentration, and finally washed with washing buffer that did not contain urea. The DddAtox bound to the column was then eluted with 5ml elution buffer. The eluted DddAtox was subjected to size exclusion chromatography using fast protein liquid chromatography (FPLC) and gel filtered on a Superdex200 column (GE Healthcare) in a sizing buffer (20 mM Tris-HCl pH 7.5, 200 mM NaCl, 1 mM DTT, 5% (w / v) glycerol). The purity of the eluted DddAtox was assessed by using a Coomassie Brilliant Blue stained SDS-PAGE gel, and the protein was then stored at -80°C.
[0119] Example IV
[0120] Purification of DddA11
[0121] Purification of DddA11 protein was the same as DddAtoxin as previously reported (Beverly et al. 2020, Nature, 583(7817):631-637 doi:10.1038 / s41586-020-2477-4).
[0122] Example V
[0123] Purification of BadTF3
[0124] To purify the his-tagged BadTF3 complexed with BadTF3I, 2L LB broth was inoculated with E. coli BL21 (pETDuet-1::BadTF3+BadTF3I) at a dilution of 1:100 and cultured overnight. After the culture grew to about OD600=0.6, 0.5mM IPTG was added and incubated at 18°C for 16 hours with shaking. The bacterial cell pellet was collected by centrifugation at 4000g for 30 minutes and then resuspended in 50ml lysis buffer (50mM Tris-HCl pH8.0, 500mM NaCl, 10mM imidazole, 1mg / ml lysozyme and protease inhibitor cocktail). The bacterial cell pellet was then lysed by ultrasonic treatment (five pulses, 10 seconds each), and the supernatant was separated from the debris by centrifugation at 25,000g for 30 minutes. The his-tagged BadTF3-BadTF3I complex was purified from the supernatant using a nickel column. BadTF3-BadTF3I was eluted with 1 elution buffer (50mM Tris-HCl pH7.5, 300mM imidazole, 500mM NaCl, 30mM imidazole, 1mM DTT). The eluted BadTF3-BadTF3I complex was denatured by adding 50ml 8M urea denaturation buffer (50mM Tris-HCl pH7.5, 300mM imidazole, 500mM NaCl and 1mM DTT) for 16 hours at 4°C. The denatured protein in the 8M urea denaturation buffer was loaded onto the nickel column again. The column was washed with 50ml 8M urea denaturation buffer to exclude any residual BadTF3I. The column was washed continuously with 25ml denaturation buffer with gradually decreasing concentrations of urea (6M, 4M, 2M, 1M), and finally washed with a wash buffer without urea. BadTF3 bound to the column was then eluted with 5ml elution buffer. The eluted BadTF3 was subjected to size exclusion chromatography using fast protein liquid chromatography (FPLC) and gel filtration on a Superdex200 column (GE Healthcare) in size screening buffer (20 mM Tris-HCl pH 7.5, 200 mM NaCl, 1 mM DTT, 5% (w / v) glycerol). The purity of the eluted BadTF3 was assessed by SDS-PAGE gel stained with Coomassie Brilliant Blue, and the protein was then stored at -80°C.
[0125] Figure 5 Depicted are SDS-PAGE gels stained with Coomassie Brilliant Blue of DddAtox-DddAI, DddAtox, DddA11-DddAI, DddA11, BadTF3-BadTF3I, and BadTF3, respectively.
[0126] Example VI
[0127] DNA deamination activity assay
[0128] Lambda DNA or genomic DNA extracted from Drosophila S2 or K562 or GM12878 cell lines was used to assess double-stranded DNA deamination activity. Reactions were performed in 10 μl of deamination buffer consisting of 20 mM Tris-HCl pH 7.4, 100 mM NaCl, 1 mM DTT, 50 ng DNA substrate and deaminase (20 μM unless otherwise stated). Reactions were incubated at 37°C for 1 hour or the indicated time course, followed by DNA purification using the Zymo DNA Clean & Concentrator-5 kit. Library preparation was performed on purified DNA using the TruePrep DNA Library Prep Kit V2 for Illumina (Vazyme) as described, except that the polymerase mix was replaced with 1×Q5U PCR master mix supplemented with Bst3.0 polymerase (0.08 U per μl). The uracil conversion rate was calculated for each cytosine site and averaged to assess deamination activity in each possible sequence environment.
[0129] Figure 6 Depict the conversion efficiency of cytosine to uracil in double-stranded DNA treated for one hour by DddAtoxin, DddA11 protein or BadTF3. Each square represents a cytosine and its four possible upstream nucleotide environments and four possible downstream nucleotide environments. The brightness of red indicates the degree of conversion. The left side shows the average conversion efficiency of the cytosine sites of naked genomic DNA in vitro (Drosophila genome for DddA and DddA11, LambdaDNA for BadTF3). The right side shows the average conversion efficiency of the cytosine sites of K562 genomic DNA extracted from the living cell experiment. Purified DddA specifically deaminates cytosine under TC or CC environment (TCC is converted to TUC, and then U is regarded as T to start the conversion of UC to UU). Enzymes DddA11 and BadTF3 deaminate cytosine in TC, CC and AC and a small amount of GC environment.
[0130] Example VII
[0131] Nuclei preparation
[0132] To prepare the nuclei, the cells were centrifuged at 450 g for 5 minutes, then washed with an equal volume of cold 1 × PBS and centrifuged at 500 g for 5 minutes at 4 ° C. A total of 30,000 cells were permeabilized using cold permeabilization buffer (10 mM Tris-HCl, pH 7.4, 10 mM NaCl, 3 mM MgCl2, 0.1% IGEPAL CA-630, 0.1% Tween-20, 0.1% digitonin). After permeabilization, the nuclei were immediately centrifuged at 4 ° C for 5 minutes at 550 g. After centrifugation, the supernatant was carefully removed from the pellet.
[0133] Example VIII
[0134] Fragmentation Tagmentation Reaction
[0135] It should be understood that the fragmentation labeling described below can be performed before or after the DNA is treated with a dsDNA deaminase. Prepare 2X TD buffer (20 mM TAPS pH 8.5, 10 mM MgCl 2 , 20% DMF). Immediately resuspend the nuclear pellet in the transposase reaction mixture (12.5 μL 2×TD buffer, 2 μL transposase (Vazyme, 1.25 μM) and 10 μL PBS containing 0.1% digitonin). The transposition reaction was carried out at 37°C for 30 minutes on a thermostatic mixer at 800 rpm. Afterwards, 100 μl of ice-cold RSB was added to the mixer and then centrifuged at 550 g for 5 minutes at 4°C. After centrifugation, the supernatant was carefully removed from the pellet.
[0136] Example IX
[0137] DNA deamination reaction
[0138] Prepare 2X DRB buffer (20mM Tris-HCl pH7.5, 20mM NaCl, 2mM DTT). Immediately suspend the nuclear pellet in the deaminase reaction mixture. For DddAtox, the deaminase reaction mixture contains 12μl DRB and 18μl DddA enzyme (50uM, stored in size screening buffer). For DddA11, the deaminase reaction mixture contains 8μl DRB and 12μl DddA11 enzyme. For BadTF3, the deaminase reaction mixture contains 5μl DRB and 5μl BadTF3 enzyme (50uM, stored in size screening buffer). Incubate the nuclear mixture at 37°C for 20 minutes. Immediately after the reaction, purify the DNA using the Zymo DNAClean&Concentrator-5 kit.
[0139] Example X
[0140] Library amplification
[0141] After DNA purification, the library fragments were amplified using 1×Q5U PCR premix, Bst 3.0 polymerase (0.08U / μl) and 1.25μM Nextera PCR primers (forward and reverse), and the PCR conditions were as follows: 65°C for 5 minutes; 80°C for 12 minutes; 98°C for 2 minutes; and thermal cycles of 98°C for 15 seconds, 60°C for 30 seconds, and 72°C for 1 minute. It is recommended to monitor the PCR reaction using qPCR to stop amplification before saturation and to avoid GC and size bias in PCR. After five cycles of preamplification of the complete library, a sample of the PCR reaction was collected for qPCR. A total of 20 cycles of qPCR were performed to determine the number of additional cycles required for the remaining PCR reactions. Typically, a total of 10-12 cycles of amplification produced high-quality libraries. The library was purified using the Zymo Select-a-Size DNAClean&Concentrator Kit to collect DNA with a fragment size of more than 200bp. The size-selected library is ready for sequencing.
[0142] Example XI
[0143] Detection of transcription factor footprints
[0144] Fig. 7A Depicted is the identification of transcription factor CTCF footprints in the human K562 genome using the methods described herein. Isolated nuclei were treated with DddA, DddA11, and BadTF3, respectively. The Y axis shows the conversion rate of each cytosine site. The purple bar shows the CTCF binding motif. Figure 7B The average turnover of the merged CTCF binding motifs is shown, with the footprint observed at the center. Figure 7C A proportional Venn diagram is depicted showing the overlap between CTCF binding sites identified by the dsDNA deaminase method, ChIP-seq method, and DNase-seq method described herein. A total of 31,186 CTCF binding sites detected by the dsDNA deaminase method (38,281 in total) were consistent with the CTCF ChIP-seq method (36,110 in total). In contrast, only 19,989 of the 22,085 binding sites identified by the DNase-seq method overlapped with the binding sites identified by the ChIP-seq method. These data suggest that the dsDNA deaminase method is more advantageous than the ChIP-seq method and can robustly identify a large number of TF binding sites.
[0145] The method described herein determines TF binding rates by determining TF footprints within single DNA molecules. Fig. 8ASchematic representation of the analysis of TF binding patterns at the single-molecule level using data obtained using the methods described herein. For each sequencing read, unconverted sites are interpreted as TF binding regions. Alternatively, converted sites are interpreted as accessible regions. Individual reads can be ranked based on occupancy patterns on multiple genomic features (e.g., TFBS clusters). Figure 8B In the figure, each line represents a sequenced DNA read, and all reads located in chromosome 1:26321500-26321900 are accumulated. Each black dot represents a converted cytosine, and each gray dot represents an unconverted cytosine. If the DNA is occupied by a TF (such as CTCF in this case), the binding of the TF will prevent the deamination of cytosines at a specific binding site, while the upstream and downstream flanking cytosines will not be protected from deamination. On the other hand, if the DNA is not occupied by a TF, all cytosines can be contacted by dsDNA deaminases and subsequently deaminated. Therefore, TF binding is determined for each DNA molecule, and the TF binding rate is calculated accordingly. The data show that 88.37% of the DNA is occupied by CTCF at the CTCF binding site, while 11.63% of the DNA is not occupied. Figure 8C The data depicted show that the dsDNA deaminase method described herein is able to simultaneously detect three TF binding sites and relative occupancies in a promoter. Each point in the raw read represents a cytosine conversion. In this case, the footprint closest to the transcription start site has the highest binding rate (only a small amount of conversion from cytosine to uracil occurs in this region), while the most distal footprint has the lowest binding rate of the three.
[0146] Example XII
[0147] Detection of discrete transcription factor footprints in single cells
[0148] Fig.9A The detection of discrete TF footprints in chromatin DNA from single cells using the methods described herein is depicted in schematic form. Heterogeneous tissues or samples are first dissociated into single cell suspensions. After cell lysis for isolation of nuclei, the deaminase DddA is added. Then, using Tn5 transposition, universal adapters are added to the open regions and enriched. Single cell samples are obtained by FACS sorting. After filling the gaps with the help of BST, the library is amplified by Q5U and sequenced by an Illumina sequencer. Fig. 9B The DNA fragmentation distribution of a single cell after PCR amplification is shown: open regions and nucleosome patterns can be clearly identified. Fig. 9C Cell typing results for single cell data from K562, GM12878, and Hek293T cell lines are shown. The three cell types can be clustered well, and the number of each cell type used is listed. Fig.9DComparison of bulk and single cell data from the methods described in this article viewed with IGV software. In both K562 and GM12878 cell lines, the signal in the single cell data correlates well with the signal in the bulk data.
[0149] Example XIII
[0150] sequence
[0151] The amino acid sequence of natural DddAtox is:
[0152] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNSNSPKSPTKGGC (SEQ IDNO: 15)
[0153] The amino acid sequence of purified DddAtox with an N-terminal His tag is:
[0154] MGSSHHHHHHSQDPGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNSNSPKSPTKGGC (SEQ ID NO: 2)
[0155] The Genescript synthesized nucleotide sequence of DddAtox with an N-terminal His tag (optimized for bacterial expression) is:
[0156] ATGGGCAGCAGCCATCACCATCATCACCACAGCCAGGATCCGGGTAGCTATGCGCTGGGTCCGTATCAGATCTCTGCTCCGCAGCTGCCGGCATATAACGGTCAGACTGTTGGTACTTTCTATTATGTTAACGATGCTGGCGGTTTAGAAAGCAAAGTTTTCAGCTCTGGTGGTCCGACCCCGTATCCGAACTATGCTAACGCTGGTCACGTTGAAGGTCAGTCTGCTCTG TTCATGCGTGATAACGGTATCTCTGAAGGTCTGGTTTTCCATAACAACCCGGAAGGTACCTGTGGTTTTTGTGTTAACATGACCGAAACCCTGCTGCCGGAAAACGCTAAAATGACCGTTGTTCCGCCGGAAGGTGCGATTCCGGTTAAACGTGGTGCTACCGGTGAAACCAAAGTTTTCACCGGTAACTCTAACTCTCCGAAATCTCCGACCAAAGGTGGTTGCTAA(SEQ ID NO:1)
[0157] The amino acid sequence of native DddAI (same as the amino acid sequence used for protein purification) is:
[0158] MYADDFDGEIEIDEVDSLVEFLSRRPAFDANNFVLTFEESGFPQL NIFAKNDIAVVYYMDIGENFVSKGNSASGGTEKFYENKLGGEVDLSK DCVVSKEQMIEAAKQFFATKQRPEQLTWSEL(SEQ ID NO: 4,8)
[0159] The Genescript-synthesized nucleotide sequence of DddAI (optimized for bacterial expression) is:
[0160] ATGTATGCGGATGACTTTGACGGGGAAATTGAGATTGATGAAGTTGATAGCCTAGTTGAGTTTCTGAGCCGTCGTCCGGCGTTCGATGCGAACAACTTCGTTCTGACCTTCGAAGAAAGCGGCTTCCCGCAGCTGAACATCTTCGCGAAAAACGATATCGCGGTTGTTTACTACATGGATATCGGCGA AAACTTCGTTAGCAAAGGCAACAGGCGCGAGCGGCGGCACCGAAAAATTCTACGAAAACAAACTGGGCGGCGAAGTTGATCTGAGCAAAGATTGCGTTGTTAGCAAAGAACAGATGATCGAAGCGGCGAAACAGTTCTTCGCGACCAAACAGCGTCCGGAACAGCTGACCTGGAGCGAACTGTAA(SEQ ID NO:3,7)
[0161] The amino acid sequence of purified DddA11 with an N-terminal His tag is:
[0162] MGSSHHHHHHSQDPGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFISGGPTPYPNYVSAGHVEGQSALFMRDNGISEGLVFHNNPKGTCGFCVNMIETLLPENAKMTVVPPEGAIPVKRGATGETKVFIGNSNSPKSPTKGGC(SEQ ID NO: 6,13)
[0163] The Genescript synthesized nucleotide sequence of DddA11 (optimized for bacterial expression) is:
[0164] ATGGGCAGCAGCCATCACCATCATCACCACAGCCAGGATCCGGGTAGCTATGCGCTGGGTCCGTATCAGATCTCTGCTCCGCAGCTGCCGGCATATAACGGTCAGACTGTTGGTACTTTCTATTATGTTAACGATGCTGGCGGTTTAGAAAGCAAAGTTTTCATCTCTGGTGGTCCGACCCCGTATCCGAACTATGTTAGCGCTGGTCACGTTGAAGGTCAGTCTGCTCTG TTCATGCGTGATAACGGTATCTCTGAAGGTCTGGTTTTCCATAACAACCCGAAAGGTACCTGTGGTTTTTGTGTTAACATGATCGAAACCCTGCTGCCGGAAAACGCTAAAATGACCGTTGTTCCGCCGGAAGGTGCGATTCCGGTTAAACGTGGTGCTACCGGTGAAACCAAAGTTTTCATCGGTAACTCTAACTCTCCGAAATCTCCGACCAAAGGTGGTTGCTAA (SEQ ID NO:5)
[0165] The protein sequence synthesized by Genescript (optimized for bacterial expression) is: GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFISGGPTPYPNYVSAGHVEGQSALFMRDNGISEGLVFHNNPKGTCGFCVNMIETLLPENAKMTVVPPEGAIPVKRGATGETKVFIGNSNSPKSPTKGGC (SEQ ID NO: 6)
[0166] The native amino acid sequence of BadTF3 is:
[0167] GWKFSNGKRRPPHKATVTVTDKNGVVKHKSNLVSGNMTEAEKKLGFPNNSLATHTENRATRLIDLNQGDTMLIEGQYRPCPRCKGAMRVKAEESGAKVIYTWPEDGDLKKREWEGTPCDKK (SEQ ID NO: 14) The amino acid sequence of purified BadTF3 with an N-terminal His tag is:
[0168] MGSSHHHHHHSQDPGWKFSNGKRRPPHKATVTVTDKNGVVKHKSNLVSGNMTEAEKKLGFPNNSLATHTENRATRLIDLNQGDTMLIEGQYRPCPRCKGAMRVKAEESGAKVIYTWPEDGDLKKREWEGTPCDKK (SEQ IDNO: 10)
[0169] The Genescript synthesized nucleotide sequence of BadTF3 (optimized for bacterial expression) is:
[0170] ATGGGCAGCAGCCATCACCATCATCACCACAGCCAGGATCCTGGTTGGAAATTTTCTAACGGTAAACGCCGTCCGCCGCACAAAGCAACGGTAACTGTGACCGATAAAAACGGTGTCGTTAAACACAAAAGCAACCTGGTTTCTGGCAACATGACTGAAGCCGAAAAGAAACTGGGCTTCCCGAACAACTCCCTGGCGACCCACAC CGAAAACCGTGCTACCCGCCTGATCGATCTGAACCAAGGTGATACTATGCTGATCGAGGGCCAATACCGTCCGTGTCCACGTTGTAAAGGTGCAATGCGCGTGAAAGCGGAGGAATCCGGTGCGAAAGTGATCTACACCTGGCCAGAAGATGGTGACCTGAAAAAACGTGAATGGGAAGGCACTCCGTGCGACAAAAAAATAA(SEQ ID NO:9)
[0171] The amino acid sequence of native BadTF3I (same as the amino acid sequence used for protein purification) is:
[0172] MTKSKMLSNIVIQEVKFAIEDYCAILSFASDSYEVPEQYFIITRSTTERSGGIPEGDIYLESNLFLDFNPYGLSGYLLSEPNCVDLLIEPNNYVRLRLIEKIDILEVENHLKFLFDN (SEQ ID NO: 12)
[0173] The Genescript-synthesized nucleotide sequence of BadTF3I (optimized for bacterial expression) is:
[0174] ATGACCAAATCTAAAATGCTGAGCAACATCGTCATCCAGGAGGTCAAATTTGCGATCGAAGATTACTGCGCTATTCTGAGCTTCGCTTCTGACTCTTATGAAGTGCCGGAGCAGTATTTTATCATTACCCGTTCTACCACCGAACGTTCTGGCGGTATTCCGGAGGGCGACATCTACCT GGAATCTAACCTGTTTCTGGATTTTAACCCGTACGGCCTGAGCGGTTACCTGCTGTCTGAGCCGAACTGCGTAGATCTGCTGATCGAACCGAACAACTACGTTCGTCTGCGTCTGATCGAAAAAATCGATATCCTGGAAGTGGAAAACCACCTGAAATTTCTGTTCGACAACTAA(SEQ ID NO: 11).
[0175] Example XIV
[0176] Reagent test kit
[0177] The materials and reagents required for the disclosed method of determining TF footprints using dsDNA deaminases can be assembled together in a kit. The kit of the present disclosure will generally include at least dsDNA deaminases, transposases, nucleases, degradation enzymes, nucleotides, DNA polymerases, amplification primers and reagents, sequencing primers and reagents and / or DNA enrichment reagents described herein, which can be used to implement the claimed methods. In a preferred embodiment, the kit will also contain instructions for treating chromatin DNA with dsDNA deaminases and treating and amplifying the treated chromatin DNA. In each case, the kit will preferably have a different container for each individual reagent, enzyme or reactant. Each reagent will generally be appropriately dispensed in its respective container. The container means of the kit will generally include at least one vial or test tube. Flasks, bottles and other container means into which the reagents are placed and dispensed can also be used. The individual containers of the kit are preferably kept sealed for commercial sale. Suitable larger containers may include injection-molded or blow-molded plastic containers in which the required vials are retained. The kit is preferably provided with instructions.
[0178] Implementation
[0179] The present disclosure provides a method for determining a transcription factor binding site on a genomic double-stranded (ds) DNA of a cell (e.g., a eukaryotic cell), comprising contacting the genomic dsDNA with a dsDNA deaminase under conditions where the cytosine of the genomic dsDNA is converted to uracil, thereby producing a treated genomic dsDNA, and identifying the unconverted cytosine on the treated genomic dsDNA as a transcription factor binding site. According to one aspect, the genomic double-stranded (ds) DNA is a gene. According to one aspect, the method includes identifying one or more unconverted cytosines as a transcription factor binding site. According to one aspect, the method includes identifying a plurality of unconverted cytosines as a transcription factor binding site. According to one aspect, the method includes identifying a plurality of unconverted cytosines as two or more transcription factor binding sites. According to one aspect, the genomic dsDNA includes a plurality of genes, and the method also includes identifying a plurality of unconverted cytosines as a plurality of transcription factor binding sites. According to one aspect, the pattern of unconverted cytosine on the treated genomic DNA is associated with the DNA binding domain of the transcription factor to identify one or more transcription factor binding sites. According to one aspect, the dsDNA deaminase is DddA, BadTF3 or DddA11 or its variant, mutant, derivative or modification. According to one aspect, the dsDNA deaminase is an enzyme capable of converting cytosine on double-stranded (ds) DNA into uracil. According to one aspect, the treated genomic DNA is optionally amplified and sequenced to determine the position of unconverted cytosine and uracil. According to one aspect, the treated genomic DNA is optionally amplified and sequenced to determine the position of unconverted cytosine and uracil, which is compared with the DNA binding pattern of the transcription factor to identify the transcription factor binding site. According to one aspect, the treated genomic DNA is optionally amplified and sequenced to determine the position of unconverted cytosine and uracil, which is compared with the DNA binding pattern of the transcription factor to identify the transcription factor binding site and the related transcription factor. According to one aspect, the genomic dsDNA is a single DNA molecule from a single cell, and the method also includes identifying multiple unconverted cytosines on a single DNA molecule as multiple transcription factor binding sites on a single DNA molecule. According to one aspect, genomic double-stranded (ds) DNA is processed into multiple DNA molecules, and their relative binding rates for transcription factors are analyzed or quantified. According to one aspect, the processed genomic dsDNA is subjected to whole genome sequencing or targeted amplicon sequencing. According to one aspect, regions where cytosine is converted to uracil flank regions of unconverted cytosine to identify transcription factor footprints on open regions of the genomic dsDNA. According to one aspect, one or more open regions of the processed genomic dsDNA are enriched by fragmentation labeling and amplification.According to one aspect, one or more open regions of the treated genomic dsDNA are digested and amplified by nuclease enrichment. According to one aspect, the genomic dsDNA is obtained from multiple cells of the same cell type. According to one aspect, the treated genomic dsDNA is amplified by PCR. According to one aspect, the treated genomic dsDNA is processed into fragments for sequencing. According to one aspect, the genomic dsDNA is treated with a dsDNA deaminase in a cell or in a nucleus separated from a cell. According to one aspect, the genomic double-stranded (ds) DNA is processed to enrich for open chromatin DNA. According to one aspect, before being treated with a dsDNA deaminase, the genomic double-stranded (ds) DNA is processed to enrich for open chromatin DNA. According to one aspect, the genomic double-stranded (ds) DNA is treated with a dsDNA deaminase, and then the treated genomic dsDNA is processed to enrich for open chromatin DNA.
[0180] Equivalent
[0181] Other embodiments will be apparent to those skilled in the art. It should be understood that the foregoing description is provided for clarification purposes only and is exemplary only. The spirit and scope of the present invention are not limited to the above examples, but are encompassed in the claims. All publications, patents, and patent applications cited above are incorporated herein by reference in their entirety for all purposes, to the same extent as if each individual publication or patent application were expressly indicated to be incorporated herein by reference.
Claims
1. A method for determining a transcription factor binding site on a double-stranded (ds) DNA of a eukaryotic cell genome, comprising: contacting the genomic dsDNA with a dsDNA deaminase under conditions that convert cytosine of the genomic dsDNA to uracil, thereby producing treated genomic dsDNA, and Unconverted cytosines on the treated genomic dsDNA were identified as transcription factor binding sites.
2. The method of claim 1, wherein the genomic double-stranded (ds) DNA is a gene.
3. The method of claim 1, comprising identifying one or more unconverted cytosines as transcription factor binding sites.
4. The method of claim 1, comprising identifying a plurality of unconverted cytosines as transcription factor binding sites.
5. The method of claim 1, comprising identifying a plurality of unconverted cytosines as two or more transcription factor binding sites.
6. The method of claim 1, wherein the genomic dsDNA comprises a plurality of genes, and the method further comprises identifying a plurality of unconverted cytosines as a plurality of transcription factor binding sites.
7. The method of claim 1, wherein the pattern of unconverted cytosines on the treated genomic DNA is correlated with a DNA binding domain of a transcription factor to identify one or more transcription factor binding sites.
8. The method of claim 1, wherein the dsDNA deaminase is DddA, BadTF3 or DddA11 or a variant, mutant, derivative or modification thereof.
9. The method of claim 1, wherein the dsDNA deaminase is an enzyme capable of converting cytosine on double-stranded (ds) DNA into uracil.
10. The method of claim 1, wherein the treated genomic DNA is optionally amplified and sequenced to determine the positions of unconverted cytosine and uracil.
11. The method of claim 1, wherein the treated genomic DNA is optionally amplified and sequenced to determine the positions of unconverted cytosine and uracil, which are compared to the DNA binding pattern of the transcription factor to identify the transcription factor binding site.
12. The method of claim 1, wherein the treated genomic DNA is optionally amplified and sequenced to determine the positions of unconverted cytosine and uracil, which are compared to the DNA binding patterns of transcription factors to identify transcription factor binding sites and associated transcription factors.
13. The method of claim 1, wherein the genomic dsDNA is a single DNA molecule from a single cell, and the method further comprises identifying a plurality of unconverted cytosines on the single DNA molecule as a plurality of transcription factor binding sites on the single DNA molecule.
14. The method of claim 1, wherein genomic double-stranded (ds) DNA is processed into multiple DNA molecules, and their relative binding rates for transcription factors are analyzed or quantified.
15. The method of claim 1, wherein the processed genomic dsDNA is subjected to whole genome sequencing or targeted amplicon sequencing.
16. The method of claim 1, wherein regions of cytosine conversion to uracil are flanked by regions of unconverted cytosine to identify transcription factor footprints on open regions of genomic dsDNA.
17. The method of claim 1, wherein one or more open regions of the treated genomic dsDNA are enriched by tagmentation and amplification.
18. The method of claim 1, wherein one or more open regions of the treated genomic dsDNA are enriched by nuclease digestion and amplification.
19. The method of claim 1, wherein genomic dsDNA is obtained from a plurality of cells of the same cell type.
20. The method of claim 1, wherein the treated genomic dsDNA is PCR amplified.
21. The method of claim 1, wherein the treated genomic dsDNA is processed into fragments for sequencing.
22. The method of claim 1, wherein the genomic dsDNA is treated with a dsDNA deaminase within a cell or within a nucleus isolated from a cell.
23. The method of claim 1, wherein genomic double-stranded (ds) DNA is processed to enrich for open chromatin DNA.
24. The method of claim 1, wherein the genomic double-stranded (ds) DNA is processed to enrich for open chromatin DNA prior to treatment with a dsDNA deaminase.
25. The method of claim 1, wherein the genomic double-stranded (ds) DNA is treated with a dsDNA deaminase and then the treated genomic dsDNA is processed to enrich for open chromatin DNA.
Citation Information
Patent Citations
Method for detecting a target nucleic acid sequence
EP0320308A2
An improved method for assaying of nucleic acids, a reagent combination and a kit therefore
GB2202328A
Multiplex decoding of sequence tags in barcodes
US20080269068A1
Nanogrid rolling circle DNA sequencing
US20090018024A1
Transposon end compositions and methods for modifying nucleic acids
US20110287435A1