Methods, compositions, and kits for identifying protein-binding regions in genomic DNA

By contacting genomic DNA with adenine methyltransferase and performing single-molecule long-read sequencing, the problem of recording the primary structure of chromatin has been solved, enabling the precise identification of nucleosome locations and regulatory DNA interactions, and providing a method for visualizing chromatin regions.

CN115715321BActive Publication Date: 2026-03-10ALTIUS INST FOR BIOMEDICAL SCI +1
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-04-02
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Existing technologies struggle to record the primary structure of chromatin at single nucleotide resolution, making it difficult to accurately identify nucleosome locations and regulatory DNA interactions within the genome. This results in uncertainty regarding the arrangement of nucleosomes along chromatin fibers and the degree of driving forces in regulatory regions.

Method used

By contacting genomic DNA with adenine methyltransferase (A-MTase), adenine residues in regions that do not bind to proteins are methylated, and single-molecule long-read sequencing is performed to detect the locations of missing methylated adenine residues, thereby identifying protein-binding regions and the location of nucleosomes.

Benefits of technology

This technology enables the recording of primary chromatin structures at single nucleotide resolution, precise identification of protein-binding regions in the genome, reveals nucleosome locations and interactions that regulate DNA, and provides a method for visualizing chromatin regions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115715321B_ABST
    Figure CN115715321B_ABST
Patent Text Reader

Abstract

Methods, compositions, kits, and systems are provided for identifying protein-binding regions in genomic DNA. The method may include contacting genomic DNA with an adenine methyltransferase (A-MTase), wherein the A-MTase methylates adenine residues in a region of the genomic DNA that is not protein-binding; and performing single-molecule long-read sequencing on the contacted genomic DNA to detect locations in the genomic DNA lacking methylated adenine residues, thereby identifying regions in the genomic DNA that are protein-binding. The bound region may be a nucleosome location, and the method can determine the nucleosome location in the genomic DNA. A method is also provided for visualizing chromatin regions that are not protein-binding and spatially serve as substrates for the A-MTase within the cell by visualizing the location of methylated adenine after contacting the cell with the adenine methyltransferase (A-MTase).
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-referencing of related patent applications

[0002] This application claims priority to U.S. Provisional Application 63 / 004,361, filed April 2, 2020, the disclosure of which is incorporated herein by reference in its entirety.

[0003] sequence list

[0004] The sequence list is provided with this application as a text file “ALTI-730WO Seq List_ST25.txt”, created on April 2, 2020, and is 10KB in size. The contents of this text file are hereby incorporated herein by reference in their entirety. Background Technology

[0005] The primary structure of chromatin comprises arrays of nucleosomes truncated by short regulatory regions containing transcription factors and other non-histone proteins. This structure is fundamental to genome function, but its details remain unclear at the level of individual chromatin fibers (the basic units of gene regulation). For example, while nucleosomes are a major barrier restricting transcription factor access to DNA, the localization and occupancy of nucleosomes arranged along individual chromatin fibers in vivo are not well elucidated. Therefore, the following aspects remain unclear: how nucleosomes are precisely aligned along the same extended chromatin template; the interactions between accessible regulatory DNA and nucleosomes on individual chromatin fibers; the extent to which regulatory regions encoded by a given DNA are driven on different chromatin fibers within a cell population; and the extent to which nearby regulatory regions are coordinated and driven on the same chromatin template. Addressing these questions requires sequencing individual chromatin fibers, which is currently impossible with single-cell or batch analysis methods.

[0006] There is a need for methods to record the primary structure of chromatin onto its basal DNA template at single-nucleotide resolution, thereby enabling the simultaneous identification of genetic and epigenetic characteristics of multi-kilobase segments in the genome. This disclosure addresses these and other needs. Summary of the Invention

[0007] Methods, compositions, kits, and systems are provided for identifying protein-binding regions in genomic DNA. In some aspects, the method includes contacting genomic DNA with an adenine methyltransferase (A-MTase), wherein the A-MTase methylates adenine residues in a region of the genomic DNA that is not protein-binding; and performing single-molecule long-read sequencing on the contacted genomic DNA to detect locations in the genomic DNA lacking methylated adenine residues, thereby identifying regions in the genomic DNA that bind to the protein. In some aspects, the bound region is a nucleosome location. Therefore, the method includes determining the nucleosome location in genomic DNA. Furthermore, compositions, systems, and kits are provided for use, for example, in carrying out the methods of this disclosure.

[0008] A method is also provided for visualizing chromatin regions that are not bound to proteins and are spatially available as substrates for the A-MTase within the cell by visualizing the location of methylated adenine after contacting the cell with the adenine methyltransferase (A-MTase). Attached Figure Description

[0009] A better understanding of the invention will be achieved by reading the following detailed description in conjunction with the accompanying drawings. The drawings include the following figures:

[0010] Figures 1A-1F This study demonstrates nonspecific m6A-MTase selective labeling of chromatin accessibility sites.

[0011] Figures 2A-2H This shows a base pair resolution map of the structure of a single chromatin fiber revealed by Fiber-seq.

[0012] Figures 3A-3B The coordinated drive of adjacent regulatory elements on the same chromatin fiber is shown.

[0013] Figures 4A-4D The effects of regulating DNA drives on nucleosome localization are shown.

[0014] Figures 5A-5E This demonstrates the conservation of chromatin structure between fruit flies and humans.

[0015] Figures 6A-6E The use of cell-penetrating peptide (CPP)-labeled m6A-MTase to identify in vivo chromatin structure is shown.

[0016] Figure 7 This demonstrates the use of Fiber-seq to identify functional gene-regulated DNA alterations.

[0017] Figure 8The results show that N6-methyladenosine (m6A), detected using three different antibodies that specifically bind to m6A, increases in a dotted pattern with increasing Hia5 dosage.

[0018] Figure 9 The total nuclear m6A signal (left) and the single-point intensity (right) show a dose-dependent increase in Hia5.

[0019] Implementation

[0020] Methods, compositions, kits, and systems for identifying protein-binding regions in genomic DNA are provided. In some aspects, the method includes contacting genomic DNA with an adenine methyltransferase (A-MTase), wherein the A-MTase methylates adenine residues in a region of the genomic DNA that is not protein-binding; and performing single-molecule long-read sequencing on the contacted genomic DNA to detect locations in the genomic DNA lacking methylated adenine residues, thereby identifying regions in the genomic DNA that bind to the protein. In some aspects, the bound region is a nucleosome location. Therefore, the method includes determining nucleosome locations in genomic DNA. Furthermore, compositions, systems, and kits that can be used, for example, to carry out the methods of this disclosure are provided. In some aspects, at least some steps of the method are performed using a computer including a processor, the processor including a program that performs these steps when executed by the processor.

[0021] A method is also provided for visualizing chromatin regions that are not bound to proteins and are spatially available as substrates for the A-MTase within the cell by visualizing the location of methylated adenine after contacting the cell with the adenine methyltransferase (A-MTase).

[0022] Before describing exemplary embodiments of the invention, it should be understood that the invention is not limited to the specific embodiments described, as these embodiments can certainly vary. It should also be understood that the terminology used herein is for describing specific embodiments only and is not intended to be limiting, as the scope of the invention will be limited only by the appended claims.

[0023] Where numerical ranges are provided, it should be understood that, unless the context clearly indicates otherwise, the median values ​​between the upper and lower limits of the range are also explicitly disclosed to be accurate to one-tenth of the lower limit unit. The smaller ranges between any claimed value or median value within the claimed range and any other claimed value or median value within the claimed range are included within the scope of this invention. The upper and lower limits of these smaller ranges may be independently included within or excluded from the claimed range, and ranges containing one or two limits or no limits within these smaller ranges are also included within the scope of this invention, but limited to any limit explicitly excluded from the claimed range. If the claimed range contains one or two of these limits, ranges excluding one or two of those included limits are also included within the scope of this invention.

[0024] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. While any methods and materials similar to or equivalent to those described herein may be used in the practice or testing of this invention, some potential and exemplary methods and materials are described herein. Any and all publications mentioned herein are incorporated by reference to disclose and describe methods and / or materials relating to the cited publications. It should be understood that, in the event of any conflict, this disclosure supersedes any disclosure in the incorporated publications.

[0025] It should be noted that, as used herein and in the appended claims, the singular forms “a,” “an,” and “the” include plural references unless the context clearly indicates otherwise. Thus, for example, reference to “a cell” includes a plurality of such cells, and reference to “the cell” includes one or more cells, and so on.

[0026] It should also be noted that the claims may be drafted to exclude any optional elements. Therefore, this statement is intended to serve as a prior basis for using exclusive terms such as “alone” or “only” or for using “negative” restrictions, in conjunction with the wording of the claim elements.

[0027] The publications discussed herein are provided only to illustrate that they were disclosed prior to the filing date of this application. Nothing herein should be construed as an admission that the invention is not entitled to precede such publications by any prior invention. Furthermore, the publication dates shown may differ from the actual publication dates, and separate verification of the dates may be required. In the event that such publications list definitions of terms that conflict with the express or implied definitions of this disclosure, the definitions of this disclosure shall prevail.

[0028] As will be apparent to those skilled in the art upon reading this disclosure, each individual embodiment described and illustrated herein has discrete components and features that can be easily separated from or combined with any feature of the other several embodiments without departing from the scope or spirit of the invention. Any of the enumerated methods can be implemented in the order of the enumerated events or in any other logically feasible order.

[0029] definition

[0030] The term "Hia5" refers to a polypeptide whose amino acid sequence (SEQ ID NO:1) is at least 80% identical (e.g., at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical) to that of the Hia5 polypeptide from Haemophilus influenzae.

[0031] MANQNTFKQAPLPFIGQKRMFLKQFEQILNENISDNGEGWTILDTFGGSGLLSHTAKRLKPKARVIYNDFDGYAERLAHIDDINQLRAELYSVVGNATSKNKRMTKDCKAECIRIIQNFKGYKDLNCLASWLLFSGQQVATL DDLFQHNFWHCIRQSDYPKADGYLDGVEIVKESFHTLLPKFSNDPKALFVLDPPYLCTKQESYKQATYFDLIDFLRLVNITRPPYVFFSSTKSEFIRFVNYMLEDKVDNWQAFENAKRITVNAKLNYQVAYEDNLVYKF(SEQ ID NO:1).

[0032] The term "Hin1523" refers to a polypeptide that is at least 80% identical (e.g., at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical) to the polypeptide (SEQ ID NO:2) encoded by the hin1523 gene from Haemophilus influenzae.

[0033] Mseyleyqnaiegktmankktfkqaplpfigqkrmflkhveivlnkhidgegegwtivdvfggsgllshtakqlkpkatviyndfdgyaerlnhiddinrlrqiifnclhgiipkngrlskeikeeiinkindfkgykdlnclaswllfs gqqvgsvealfakdfwncvrqsdyptaegyldgievisesfhklipryqnqdkvlllldppylctrqesykqatyfdlidflrlinltkppyiffsstksefirylnymqesktdnwrafenykrivvkasaskdgiyednmiykf(SEQ ID NO:2).

[0034] As used herein, the term “M.Btr192IV” refers to a polypeptide that is at least 80% identical (e.g., at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical) to the polypeptide encoded by the WQG_17550 gene from *Trebrospinal bibersteinii* USDA-ARS-USMARC-192.

[0035] MKKTLTALAVASLASATQTKQQASKQASKQASKQASKECEMAKVFKQAPLPFIGQKRMFLKHFEQVLAHIPDDGNGWTIVDVFGGSGLLSHTAKRLKPKARVIYNDYDNYSERLQHIDDINRLRRIIADLMADTPKYKRL DNAKKLQIIEAIEAFQGYKDLHILCSWLAFSGQQVSSFDELYKQNFWHCIRQSDYLTADGYLDGVEIVRESFHQLVPRFTGQPNTLLVLDPPYLCTHQESYKQERYFDLVDFLRLIHLTKPPYVFFSSTKSEFVRFIDAMVEDKWDNWQAFDDAQRIVVQTSASYNGKYEDNMVYKF(SEQ ID NO:3)

[0036] As used herein, the term "EcoG1" refers to a polypeptide that is at least 80% identical (e.g., at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical) to the polypeptide (SEQ ID NO:4) encoded by the pHK08_22 gene from Escherichia coli.

[0037] MSRFILGDCVRVMATFP DNAVDFILLTDPPYLVGFRDRSGRTIAGDVNDDWLQPASNEMYRVLKKDALMVSFYGWNRIDRFMAAWKRAGFSVVGHLVFTKNYTSKAAYVGYRHECAYILAKGRPA LPQKPLPDVLGWKYSGNRHHPTEKPVTSLQPLIESFTHPNAIVLDPFAGSGSTCVAALQSGRRYIGIELLEQYHRAGQQRLAAVQRAMQQGAANDNWFEPEAA(SEQ ID NO:4)

[0038] Hia5, EcoGII, Btr192IV, and EcoGI possess adenine methyltransferase activity. These methyltransferases can be codon-optimized to increase their expression in *E. coli* cells.

[0039] The A-MTases disclosed herein, such as N6-adenine methyltransferase (m6A-MTase), include modified Hia5, EcoGII, Btr192IV, and EcoGI, such as variants with amino acid sequences different from those disclosed herein, and mutants containing insertions, substitutions, deletions, and fusion proteins. Fusion proteins contain an A-MTase fused to a cell-penetrating peptide, a tag, etc. The cell-penetrating peptide can be a peptide with a net positive charge to allow the fusion protein membrane to permeate. The cell-penetrating peptide can be an HIV-1 TAT translocation domain, 8-arginine (8R), a penetrating protein, and variants thereof. Fusion proteins contain an A-MTase fused to a nuclear localization sequence (NLS) to target the A-MTase to the cell nucleus. The A-MTase can be fused to both the NLS and the cell-penetrating peptide.

[0040] The terms "antibody" and "immunoglobulin" include any isotype antibody or immunoglobulin, antibody fragments that retain specific binding to antigens, including but not limited to Fab, Fv, scFv, Fd, Fab', Fv, and F(ab')2, chimeric antibodies, humanized antibodies, monoclonal antibodies, single-chain antibodies including antibodies containing only the heavy chain (e.g., VHH camelid antibodies), bispecific antibodies, and fusion proteins containing both antibody and non-antibody protein antigen-binding portions. Immunoglobulins can be classified into different classes based on the amino acid sequence of their heavy chain constant domains. There are five main classes of immunoglobulins: IgA, IgD, IgE, IgG, and IgM, some of which can be further subdivided into subclasses (isotypes), such as IgG1, IgG2, IgG3, IgG4, IgA, and IgA2. The terms "antibody" and "immunoglobulin" specifically include, but are not limited to, IgG1, IgG2, IgG3, and IgG4 antibodies. These antibodies can be detected by labeling, for example, using radioisotopes, enzymes that produce detectable products, fluorescent proteins, etc. These antibodies can be further conjugated to other parts, such as components of specific binding pairs, like biotin (a component of the biotin-avidin specific binding pair), etc.

[0041] An "antibody fragment" contains a portion of a complete antibody, such as the antigen-binding region or variable region of that complete antibody. Examples of antibody fragments include Fab, Fab', F(ab')2, and Fv fragments; biantibodies; linear antibodies (Zapata et al., Protein Eng. 8(10):1057-1062(1995)); single-chain antibody molecules, including antibodies containing only the heavy chain (e.g., VHH Camelidae antibody); and multispecific antibodies formed from antibody fragments. Papain digestion of an antibody produces two identical antigen-binding fragments, called "Fab" fragments, each with a single antigen-binding site and a residual "Fc" fragment; this naming reflects its ability to crystallize rapidly. Pepsin treatment produces the F(ab')2 fragment, which has two antigen-binding sites and is still capable of cross-linking with the antigen.

[0042] "Single-chain Fv", "sFv", or "scFv" antibody fragments contain the VH and VL domains of the antibody, where these domains are contained within a single polypeptide chain. In some embodiments, the Fv polypeptide also contains a polypeptide linker between the VH and VL domains, which allows the sFv to form the desired structure for antigen binding. For a review of sFv, see Pluckthun in The Pharmacology of Monoclonal Antibodies, vol. 113, Rosenburg and Moore eds., Springer-Verlag, New York, pp. 269-315 (1994).

[0043] As used herein, the terms “treatment,” “treatment,” etc., refer to achieving the desired pharmacological and / or physiological effect. An effect is preventative in terms of complete or partial prevention of a disease or its symptoms, and / or therapeutic in terms of partial or complete cure of a disease and / or side effects caused by that disease. As used herein, the term “treatment” covers any treatment of a disease afflicted by mammals, including humans, including: (a) preventing the occurrence of the disease in subjects who may be susceptible but have not yet been diagnosed with it; (b) suppressing the disease, i.e., halting its development; and (c) alleviating the disease, even if the disease subsides.

[0044] In this article, the terms “individual,” “subject,” “host,” and “patient,” which are used interchangeably, refer to mammals, including but not limited to rodents (rats, mice), non-human primates, humans, canines, felines, and ungulates (e.g., horses, cattle, sheep, pigs, goats).

[0045] "Biosample" encompasses a wide range of sample types obtained from an individual that can be used for diagnostic or monitoring analysis. This definition includes blood and other liquid samples of biological origin, solid tissue samples such as biopsy samples or tissue cultures or cells derived therefrom and their progeny. The definition also includes samples that are processed in any way after collection, such as by reagent treatment, dissolution, or enrichment of certain components (such as polynucleotides). The term "biosample" includes clinical samples, but also includes cells in cultures, cell supernatants, cell lysates, serum, plasma, biological fluids, and tissue samples.

[0046] method

[0047] This disclosure provides a method for identifying regions in genomic DNA that bind to proteins. The protein can be any protein that restricts adenine methyltransferase (A-MTase) to the adenine base present in the genomic sequence to which the protein binds. The protein can be one or more of a nucleosome, transcription factor, transcription repressor, etc. The steps and aspects of this method will be described in more detail below.

[0048] The method disclosed herein includes contacting genomic DNA with an adenine methyltransferase (A-MTase), wherein the A-MTase methylates adenine residues in a region of the genomic DNA that is not bound to a protein; and performing single-molecule long-read sequencing on the contacted genomic DNA to detect locations in the genomic DNA lacking methylated adenine residues, thereby identifying regions in the genomic DNA that are bound to the protein.

[0049] In some respects, this A-MTase is N6-adenine methyltransferase (m6A-MTase). In some respects, this m6A-MTase is Hia5. In some respects, this m6A-MTase is EcoGII. In some respects, this m6A-MTase is Btr192IV. In some respects, this m6A-MTase is EcoGI.

[0050] This contact may involve bringing isolated genomic DNA into contact with the A-MTase. In some respects, this contact may involve introducing nucleic acids encoding the A-MTase into a cell or introducing the A-MTase into a cell.

[0051] In some respects, the genomic DNA originates from a single cell, multiple cells (e.g., cultured cells), tissue, organ, or organism (e.g., bacteria, yeast, etc.). In some respects, the genomic DNA originates from the cells, tissues, organs, etc., of an animal. In some embodiments, the animal refers to a mammal (e.g., a hominid, a rodent (e.g., a mouse or rat), a dog, a cat, a horse, a cow, or any other related mammal). In some respects, the genomic DNA originates from the cells, tissues, organs, etc., of a human. In other respects, the genomic DNA originates from sources other than mammals, such as bacteria, yeast, insects (e.g., fruit flies), amphibians (e.g., frogs (e.g., Xenopus laevis)), viruses, plants, or any other non-mammalian source. In some respects, the genomic DNA originates from cancer cells.

[0052] In some embodiments, the genomic DNA is cell-free. Such cell-free genomic DNA may be present in or obtained from any suitable source. In some aspects, the cell-free genomic DNA is present in or obtained from a bodily fluid sample taken from whole blood, plasma, serum, amniotic fluid, saliva, urine, pleural effusion, bronchoalveolar lavage fluid, bronchial aspirate, breast milk, colostrum, tears, semen, peritoneal fluid, pleural fluid, and feces. In some embodiments, the genomic DNA is cell-free fetal DNA. In some aspects, the genomic DNA is circulating tumor DNA. In some embodiments, the genomic DNA contains infectious agent DNA. In some embodiments, the genomic DNA contains DNA derived from a graft. As used herein, the term "cell-free genomic DNA" can refer to a composition of genomic DNA that is cell-free or substantially cell-free. The genomic DNA does not necessarily mean the presence of all cellular genetic material; more precisely, the genomic DNA may include a portion of cellular genomic material. For example, the genomic DNA may contain isolated chromatin fragments, which can be any nucleoprotein-associated genomic DNA fragment isolated from a cell. Exemplary chromatin fragments may be oligonucleotide bodies, mononuclei, centromeres, telomeres, or genomic DNA bound by transcription factors or chromatin remodeling factors.

[0053] In some embodiments, the cells may be peripheral blood mononuclear cells (PBMCs), leukocytes, or cells isolated from bone marrow, thymus, tissue biopsy, tumors, lymphomas, lymph nodes, intestinal lymphoid tissue, mucosa-associated lymphoid tissue, spleen, other lymphoid tissue, liver, lungs, stomach, intestines, colon, kidneys, pancreas, breast, bone, prostate, cervix, testes, ovaries, tonsils, or other organs, and / or cells derived from them. In some embodiments, the nucleic acids to be evaluated (e.g., genomic DNA, chromosomal DNA) are derived from blood cells, such as from whole blood samples or blood cell subsets in whole blood. Cell subsets in whole blood include platelets, red blood cells (erythrocytes), and leukocytes (i.e., peripheral blood leukocytes, which consist of neutrophils, lymphocytes, eosinophils, basophils, and monocytes). White blood cells can be further divided into two groups: granulocytes (also known as polymorphonuclear leukocytes, including neutrophils, eosinophils, and basophils) and monocytes (including monocytes and lymphocytes). Lymphocytes can be further divided into T cells, B cells, and NK cells. Peripheral blood cells are present in the circulating blood pool, rather than isolated in the lymphatic system, spleen, liver, or bone marrow.

[0054] In some respects, the target method may involve analyzing genomic DNA obtained before and after treatment. For example, comparing DNA regions that are not bound to proteins and are therefore affected by adenine methylation at 1 day, 1 week, 10 days, 15 days, 1 month, 3 months, 6 months, or longer after treatment. Adenine methylation pattern comparisons can be used to assess changes in the genomic transcriptome.

[0055] In some respects, targeting methods can be used to generate reference chromatin structures and regulatory regions for a specific cell type. For example, multiple types of human cells. The chromatin structure and regulatory regions in the cells of a diseased subject can be compared to the reference chromatin structure and regulatory regions for that cell type to determine any differences. These differences may reveal previously unknown changes in chromatin structure and regulatory regions, which can be used to diagnose, predict, or treat the subject.

[0056] In some embodiments, the cell population used in the method can consist of any number of cells, such as about 500-106 or more cells, about 500-100,000 cells, about 500-50,000 cells, about 500-10,000 cells, about 50-1,000 cells, about 1-500 cells, about 1-100 cells, about 1-50 cells, or a single cell. In some embodiments, the cell sample contains fewer than about 1,000, 2,000, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, 10,000, 15,000, 20,000, 25,000, 30,000, 40,000, 50,000, 60,000, 70,000, 80,000, 90,000, or 10,000 cells. 00, 100,000, 120,000, 140,000, 160,000, 180,000, 200,000, 250,000, 300,000, 350,000, 400,000, 450,000, 500,000, 600,000, 700,000, 800,000, 900,000, or 1,000,000. In some implementations, the cell sample contains more than approximately 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10,000, 15,000, 20,000, 25,000, 30,000, 40,000, 50,000, 60,000, 70,000, 80,000, or 90,000 cells. 00, 100,000, 120,000, 140,000, 160,000, 180,000, 200,000, 250,000, 300,000, 350,000, 400,000, 450,000, 500,000, 600,000, 700,000, 800,000, 900,000, or 1,000,000.

[0057] In some aspects, the genomic DNA is present in its native environment during exposure to the methyltransferase. For example, the genomic DNA may be present in the cell (e.g., an intact cell or a permeabilized cell) during exposure to the methyltransferase. In some embodiments, a cell-permeable methyltransferase that crosses the intact or permeabilized cell membrane may be used. In some embodiments, the methyltransferase may be introduced into the cell using standard techniques. In some aspects, the genomic DNA is present in the cell lysate during exposure to the methyltransferase.

[0058] In some respects, the genomic DNA is a portion of a nucleic acid sample isolated from cells, tissues, organs, etc., of an organism (e.g., animal, human). Methods, reagents, and kits for isolating, purifying, and / or concentrating nucleic acid molecules from relevant sources are known in the art and are commercially available. For example, kits for isolating DNA from relevant sources include those manufactured by Qiagen, Inc. (Germantown, Maryland). and Nucleic Acid Isolation / Purification Kit; DNA manufactured by Life Technologies, Inc. (Carlsbad, California) Nucleic acid isolation / purification kit; and manufactured by Clontech Laboratories, Inc. (Mountingview, California). and Nucleic acid isolation / purification kits. In some respects, these kits are used to isolate nucleic acids from fixed biological samples (e.g., formaldehyde-fixed paraffin-embedded (FFPE) tissue). Genomic DNA can be isolated from FFPE tissue using commercially available kits, such as those manufactured by Qiagen, Inc. (Germantown, Maryland). FFPE Tissue DNA / RNA Kit, manufactured by Life Technologies, Inc. (Carlsbad, California). FFPE tissue total nucleic acid isolation kit and manufactured by Clontech Laboratories, Inc. (Mounting View, California) FFPE tissue kit.

[0059] In some respects, the genomic DNA can be processed for sequencing after contacting it with a methyltransferase and before sequencing. For example, the method may include processing the ends of the genomic DNA to produce flush ends. Flushing refers to the process of “filling” single-stranded overhangs, either by using the overhang as a template for polymerization and adding nucleotides to the complementary strand, or by “biting back” the overhang using exonuclease activity. DNA polymerases, such as DNA polymerase I and the Klenow fragment of T4 DNA polymerase, can be used for filling (5′→3′) and biting back (3′→5′). 5′ overhangs can be removed using nucleases such as mung bean nuclease.

[0060] In some respects, the genomic DNA can be cleaved or digested with enzymes after treatment with methyltransferases.

[0061] Single-molecule real-time sequencing systems can be used to detect methylated adenine by analyzing sequence and / or kinetic data derived from such systems. In particular, methylated adenine can alter the enzymatic activity of nucleic acid polymerases in various ways, for example, by increasing the time of binding nucleobase incorporation and / or increasing the interval between incorporation events. In some embodiments, single-molecule nucleic acid sequencing technology is used to detect polymerase activity. In some embodiments, nucleic acid sequencing technology is used to detect polymerase activity in real time, detecting nucleotide incorporation into the nascent strand. In a preferred embodiment, single-molecule nucleic acid sequencing technology is capable of detecting nucleotide incorporation events in real time. Such sequencing technologies are known in the art, including, for example... Sequencing and nanopore sequencing technologies. For more information on nanopore sequencing, see, for example, U.S. Patent 9,175,348; U.S. Patent 5,795,782; Kasianowicz, et al. (1996) Proc Natl Acad Sci USA 93(24):13770-3; Ashkenas, et al. (2005) Angew Chem Int Ed Engl 44(9):1401-4; Howorka, et al. (2001) Nat Biotechnology 19(7):636-9; and Astier, et al. (2006) J AmChem Soc 128(5):1705-10, all of which are incorporated herein by reference in their entirety for all purposes. Regarding nucleic acid sequencing, the term “template” refers to a nucleic acid molecule for the directed synthesis of a nascent strand. As described elsewhere in this document, a template may comprise, for example, DNA or an analogue, a mimic, a derivative, or a combination thereof. Furthermore, the template can be single-stranded, double-stranded, or can contain both single-stranded and double-stranded regions. Modifications in the double-stranded template can be performed on a strand complementary to the newly synthesized nascent strand, or on a strand identical to the newly synthesized strand (i.e., the strand replaced by the polymerase).

[0062] The preferred direct methylation sequencing described in this article can typically be performed using a single-molecule real-time sequencing system, i.e., continuous irradiation and observation of individual reaction complexes (such as...) over time. Systems for developing reaction complexes for DNA sequencing (see, for example, PMLundquist, et al., Optics Letters 2008, 33, 1026, which is incorporated herein by reference in its entirety for all purposes). The foregoing Sequencers typically detect fluorescence signals from thousands of zero-mode waveguide (ZMW) arrays simultaneously, enabling highly parallel operation. Each ZMW is separated from the others by a distance of a few micrometers, representing an independent sequencing chamber.

[0063] Real-time detection of individual molecules or molecular complexes, such as during analytical reactions, typically involves directly or indirectly processing the analytical reaction so that each molecule or molecular complex to be detected can be resolved individually. In this way, each analytical reaction can be monitored individually, even if multiple such reactions are immobilized on a single substrate. Individually resolvable configurations of analytical reactions can be achieved through various mechanisms, and typically involve immobilizing at least one component of the reaction at a reaction site. Various methods for providing such individually resolvable configurations are known in the art; for example, see European Patent 1105529 by Balasubramanian et al.; and published International Patent Application WO 2007 / 041394, the entire contents of which are incorporated herein by reference for all purposes. The reaction site on the substrate is typically the location on which an individual analytical reaction is performed and monitored, preferably in real time. The reaction site can be located on a plane of the substrate or in pores on the surface of the substrate, such as pore plates, nanopores, or other pores. In a preferred embodiment, this pore refers to a "nanopore," i.e., a nanoscale pore or pore plate, which confines the structure of the relevant analytical material to a nanoscale diameter range (e.g., about 1-300 nm). In some embodiments, this pore incorporates optical confinement properties, such as a zero-mode waveguide, which is also a nanoscale pore, as described further elsewhere in this document. Typically, the observation volume of such a pore (i.e., the volume in which reaction detection occurs) is in the atl (10⁻¹⁸ L) to zeolite (10⁻²¹ L) range, which is suitable for the detection and analysis of single molecules and single-molecule complexes.

[0064] The immobilization of analytical reaction components can be designed in various ways. For example, enzymes (e.g., polymerases, reverse transcriptases, kinases, etc.) can be attached to the substrate at the reaction site (e.g., within an optically confined structure or other nanoscale pore). In other embodiments, the substrate in the analytical reaction (e.g., a nucleic acid template, such as DNA and its derivatives and mimics, or a target molecule of a kinase) can be attached to the substrate at the reaction site. Certain embodiments of template immobilization are provided, for example, in U.S. Patent Application 12 / 562,690 (now U.S. Patent 8,481,264), filed September 18, 2009, which is incorporated herein by reference in its entirety for all purposes. Those skilled in the art will understand that there are many methods for immobilizing nucleic acids and proteins into optically confined structures, whether covalent or non-covalent, involving immobilization via a linker or by tying them to an immobilization site. These methods are well known in the fields of solid-phase synthesis and microarrays (Beier et al., Nucleic Acids Res. 27:1970-1-977 (1999)). Non-limiting exemplary binding moieties for linking nucleic acids or polymerases to a solid-phase support include streptavidin or avidin / biotin bonds, carbamate bonds, ester bonds, amides, thioethers, (N)-functionalized thioureas, functionalized maleimides, amino groups, disulfides, amides, hydrazone bonds, etc. Antibodies that specifically bind to one or more reaction components may also serve as binding moieties. Furthermore, silyl moieties can be directly linked to nucleic acids bound to a substrate (such as glass) using methods known in the art.

[0065] Other useful processing steps can be employed to detect the location of methylated adenine in the genomic DNA using nanopores. For example, the method may include adding one or more nanopore sequencing adapters or subregions thereof to one or more ends of the genomic DNA. A “nanopore sequencing adapter” refers to one or more nucleic acid domains that comprise at least a portion of a nucleic acid sequence (or a complement thereof) applied by an associated nanopore sequencing platform, such as a nanopore sequencing platform provided by Oxford Nanopore Technologies, such as the MinION™, GridIONx5™, PromethION™, or SmidgION™ nanopore sequencing system. The associated nanopore sequencing adapter can be added by chemical or enzymatic ligation or any other method that can be used to ligate one or more nucleic acid molecules to one or more ends of a double-stranded nucleic acid molecule. Suitable reagents (e.g., ligases) and kits for performing the ligation reaction are known and available, for example, Instant Sticky-end Ligase Master Mix available from New England Biolabs (Ipswich, Massachusetts). Ligases that can be used include T4 DNA ligase (e.g., low or high concentrations), T4 DNA ligase, T7 DNA ligase, E. coli DNA ligase, etc. The appropriate conditions for carrying out the ligation reaction will depend on the type of ligase used.

[0066] In some respects, single-molecule cycle concord sequencing (CCS) can be used to generate accurate long reads. In some respects, CCS may involve making the DNA topologically circular and sequencing that DNA multiple times to create a concordant sequence. In some respects, cycle DNA sequencing may be performed up to 20 times (e.g., 5-20 times, 5-15 times, 10-20 times, or 10-15 times).

[0067] Before sequencing, the genomic DNA can be processed to generate long fragments, such as up to 100kb, 50kb, 40kb, 30kb, 20kb, 10kb, 5kb, or 1kb in length. For example, the length range of sequencing fragments can be 1-100kb, 1-50kb, 1-40kb, 1-30kb, or 1-20kb.

[0068] In some respects, the location of methylated adenine is detected in a continuous extension of the double-stranded nucleic acid molecule, which is 500 bases or more, 1,000 bases (kb) or more, 2 kb or more, 3 kb or more, 4 kb or more, 5 kb or more, 6 kb or more, 7 kb or more, 8 kb or more, 9 kb or more, 10 kb or more, 15 kb or more, 20 kb or more, 25 kb or more, 30 kb or more, 35 kb or more, 40 kb or more, 45 kb or more, 50 kb or more, 55 kb or more, 60 kb or more, 65 kb or more, 70 kb or more, 75 kb or more, 80 kb or more, 85 kb or more, 90 kb or more, 95 kb or more, or 100 kb or more.

[0069] Computational methods (e.g., in software form) can be used to detect the location of methylated adenine in single-stranded and / or double-stranded nucleic acid molecules, determine the protein-binding region in the nucleic acid molecule based on the detected location of methylated adenine, and then sequence the nucleic acid molecule, for example, using single-molecule real-time sequencing and optionally CCS, and any combination thereof.

[0070] As described above, methods for determining binding regions in genomic nucleic acid molecules include methods for determining nucleosome locations in genomic DNA. Such methods utilize the protective / inaccessible nature of the nucleosome-associated genomic DNA from the methyltransferase used (e.g., N6-adenine DNA methyltransferase), preventing methylation from occurring in the nucleosome-associated genomic DNA. The method includes detecting the location of methylated adenine in the genomic DNA, which marks the location of the genomic DNA linker. The nucleosome location in the genomic DNA is determined based on the absence of methylated adenine. The method can also reveal the presence or absence of certain transcription factors that bind to genomic DNA.

[0071] The method disclosed herein can be performed in one or more normal cells to generate a chromatin accessibility map of sequenced genomic DNA regions, wherein the map indicates chromatin regions that are not bound to proteins and are therefore accessible to the A-MTase, as well as chromatin regions that are bound to proteins and are therefore inaccessible to the A-MTase.

[0072] The methods disclosed herein can be performed on one or more test cells to generate a chromatin accessibility map of the genomic DNA of those test cells. The test cells can be derived from a subject, such as a mammal, like a human patient. In some cases, the subject may have a disease or may be suspected of having a disease. This disease could be cancer.

[0073] The method disclosed herein may further include comparing the chromatin accessibility map of the test cell with that of a normal cell, wherein the test cell and the normal cell are of the same cell type; and comparing the genomic DNA sequences of the test cell and the normal cell, wherein a difference in the chromatin accessibility map indicates a change in the chromatin structure of the test cell, a difference in the chromatin accessibility map but a difference in the genomic DNA sequence indicates that the sequence difference is not related to the change in chromatin structure, and a difference in both the chromatin accessibility map and the genomic DNA sequence indicates that the sequence difference is related to the change in chromatin structure.

[0074] The method further includes generating a database containing information on the chromatin accessibility map, potential genomic DNA sequences, and relevance (if any) to the condition or disease. In some aspects, the normal cell and the test cell may be epithelial cells, leukocytes, glial cells, osteoblasts, or chondrocytes. In some aspects, the normal cell and the test cell comprise multiple cells. In some aspects, the multiple cells comprise at least 10 cells, at least 30 cells, at least 100 cells, at least 300 cells, or at least 10,000 cells.

[0075] In some respects, the chromatin accessibility map covers at least 10% of the chromatin, for example, at least 30%, at least 50%, or at least 80% of the chromatin. In some respects, the chromatin accessibility map covers at least 10% of the cell's genome. In some respects, the chromatin accessibility map covers at least 20%, at least 30%, at least 50%, or at least 80% of the cell's genome. In some respects, proteins binding to the genomic DNA include nucleosomes, transcriptional regulatory factors such as transcriptional repressors and transcriptional activators, or both.

[0076] This article also provides methods for visualizing chromatin regions that are not protein-binding and spatially potential substrates for intracellular adenine methyltransferases (A-MTases). For example, a spatially potential substrate for this A-MTase could be a region of genomic DNA that does not bind to histones and / or transcriptional regulatory factors (e.g., activators or repressors). The method may include contacting cells with the A-MTase and detecting the presence of methylated adenine in the cells.

[0077] The method disclosed herein visualizes methylated adenine in cells, and then, through selective fluorescent labeling of methylated adenine (m6A) in intact cell chromatin, a single-cell-level visualization map of genome regulation can be generated. This method can be used to visualize cells in a high-throughput manner. For example, at least 10, 100, 1,000, 10,000, 100,000, 1 million, 3 million, 10 million, 30 million, 100 million, or 100 million or more cells can be analyzed using the disclosed method.

[0078] This m6A imaging method can further include the detection of DNA and protein targets within cells. For example, the method may include multiplex detection of mA for other DNA and protein targets within cells.

[0079] This method may include generating a quantitative image representation of cellular regulatory states. The method may further include analyzing images of different cells and / or the same cell type at different time points.

[0080] The method may include generating a quantitative image of a methylated adenine pattern present in cells based on a tissue sample containing or suspected to contain diseased cells, and comparing the pattern with a representative pattern of normal cells.

[0081] The method may include generating a quantitative image of the pattern of methylated adenine present in cells receiving stimulation such as a therapeutic drug, and comparing the pattern with the cell pattern before receiving such stimulation.

[0082] The cell can be any relevant cell, such as mammalian cells, human cells, T cells, B cells, or diseased cells (e.g., cancer cells). The cell can be the cell described herein.

[0083] In some respects, multiple cells of the same type can come into contact with A-MTase. For example, these multiple cells can be epithelial cells, leukocytes, glial cells, osteoblasts, or chondrocytes. The cells can originate from a single individual, and in some embodiments, from a single tissue, such as the pancreas, blood, skin, intestines, etc.

[0084] This visualization method may include performing click chemistry to label methylated adenine prior to detection. The click chemistry may add a fluorescent label to the methylated adenine. The visualization method may also include adding a labeled methyl group as a substrate for an A-MTase. This labeled methyl group may be fluorescently labeled. Alternatively, the A-MTase may be labeled, for example, using a fluorophore-conjugated version of a methyltransferase.

[0085] In some embodiments, detecting the presence of methylated adenine in cells involves contacting the cells with an antibody that specifically binds to methylated adenine. This antibody may be a detection label. The detection label may be a fluorophore. The method may further include staining the genomic DNA in the cells.

[0086] In some implementations, the method may further include contacting the cell with the fluorescently labeled portion to target other specific genomic regions or related cellular proteins.

[0087] In some embodiments, the method may further include contacting the cells with an antibody that specifically binds to RNA polymerase II (PolII), such as Pol II Ser5Phos or Pol II Ser2Phos.

[0088] In some embodiments, the method may further include measuring the mean nuclear intensity and / or nuclear spot intensity of the methylated adenine-specific signal. In some embodiments, the methylated adenine-specific signal may be a fluorescent signal from a fluorescently labeled antibody that binds directly or indirectly to mA.

[0089] In some implementations, the A-MTase is m6A-MTase, such as Hia5, EcoGII, Btr192IV, or EcoGI. In some implementations, detecting the presence of methylated adenine in cells includes detecting m6A.

[0090] In some respects, visualization of mA in genomic DNA can be used to generate a reference mA pattern for a cell type. For example, there are many types of human cells. The mA pattern in the cells of a diseased subject can be compared to a reference mA pattern for that cell type to determine any differences. These differences may reveal previously unknown changes in chromatin structure and regulatory regions, which can be used to diagnose, predict, or treat the subject. The reference mA pattern may include additional information, such as the presence or absence of certain transcription factors, RNA polymerases, etc.

[0091] In some embodiments, the cell population used in the method can consist of any number of cells, such as about 500-106 or more cells, about 500-100,000 cells, about 500-50,000 cells, about 500-10,000 cells, about 50-1,000 cells, about 1-500 cells, about 1-100 cells, about 1-50 cells, or a single cell. In some embodiments, the cell sample contains fewer than about 1,000, 2,000, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, 10,000, 15,000, 20,000, 25,000, 30,000, 40,000, 50,000, 60,000, 70,000, 80,000, 90,000, or 10,000 cells. 00, 100,000, 120,000, 140,000, 160,000, 180,000, 200,000, 250,000, 300,000, 350,000, 400,000, 450,000, 500,000, 600,000, 700,000, 800,000, 900,000, or 1,000,000. In some implementations, the cell sample contains more than approximately 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10,000, 15,000, 20,000, 25,000, 30,000, 40,000, 50,000, 60,000, 70,000, 80,000, or 90,000 cells. 00, 100,000, 120,000, 140,000, 160,000, 180,000, 200,000, 250,000, 300,000, 350,000, 400,000, 450,000, 500,000, 600,000, 700,000, 800,000, 900,000, or 1,000,000.

[0092] Detecting the presence of methylated adenine (mA, e.g., m6A) in cells may involve visualizing antibodies that bind to mA. This visualization can be fluorescence imaging when using fluorescent labels that bind directly or indirectly to the antibody. Super-resolution microscopy can be used to visualize antibodies bound to mA. This super-resolution microscopy can be a deterministic super-resolution microscopy method that utilizes the nonlinear response of a fluorophore to excitation to improve resolution. Exemplary deterministic super-resolution methods may include stimulated emission loss (STED), ground state clearing (GSD), reversible saturated optical fluorescence transition (RESOLFT), and / or saturated structure illumination microscopy (SSIM). Super-resolution microscopy can also include a stochastic super-resolution microscopy method that utilizes the complex temporal behavior of a fluorophore to improve resolution. Exemplary stochastic super-resolution methods may include super-resolution optical wave imaging (SOFI), single-molecule localization microscopy (SMLM) such as spectroscopic precision determination microscopy (SPDM), SPDMphymod, photosensitive localization microscopy (PALM), fluorescence photosensitive localization microscopy (FPALM), stochastic optical reconstruction microscopy (STORM), and dSTORM.

[0093] This detection may include generating a spatial map of the methylated adenine within the cell's genome.

[0094] This method may include contacting multiple cells of the same type with A-MTase and generating a spatial map of the location of methylated adenine in the genome of these cells.

[0095] The method may include contacting multiple identical cells at at least two different time points with A-MTase, and generating a spatiotemporal map of the methylated adenine locations in the genome of these cells. The two different time points may include a first time point and a second time point, wherein the first time point and the second time point are separated by a time point at which the therapy is administered to these cells. These cells may be obtained from a subject, and the therapy may be administered to that subject.

[0096] The cells visualized using the disclosed methods can be live cells or fixed and permeable cells.

[0097] system

[0098] This disclosure also provides a system that can be used, for example, to implement the target method, including performing any one or more steps described in the "Method" section of this disclosure.

[0099] Systems for determining binding regions in genomic DNA include systems for determining the location of nucleosomes within the genomic DNA. Instructions for such systems enable the system to sequence genomic DNA treated with adenine methyltransferase and record the location of methylated adenine within the genomic DNA. Instructions for such systems can further enable the system to assess the transcriptional accessibility of certain regions of the genome based on the determined location of methylated adenine in the genomic DNA. Instructions for such systems can further enable the system to assess different nucleosome occupancy or phasing near gene promoters.

[0100] The system can be adapted (e.g., including instructions) to sequence consecutive extensions of genomic DNA that are 500 bases or more, 1 kilobase (kb) or more, 2 kb or more, 3 kb or more, 4 kb or more, 5 kb or more, 6 kb or more, 7 kb or more, 8 kb or more, 9 kb or more, 10 kb or more, 15 kb or more, 20 kb or more, 25 kb or more, 30 kb or more, 35 kb or more, 40 kb or more, 45 kb or more, 50 kb or more, 55 kb or more, 60 kb or more, 65 kb or more, 70 kb or more, 75 kb or more, 80 kb or more, 85 kb or more, 90 kb or more, 95 kb or more, or 100 kb or more, and record the location of these methylated adenines.

[0101] In some implementations, the system includes sequencing equipment, such as commercially available sequencers, like the PacBio sequencer.

[0102] This disclosure includes a computer-readable medium, comprising a non-transitory computer-readable medium having instructions stored thereon for the methods or portions thereof described herein, and may be part of a system of this disclosure. Various aspects of this disclosure include a computer-readable medium having instructions stored thereon that, when executed, cause a system to perform one or more steps of the methods described herein.

[0103] In some implementations, instructions for the methods and systems described herein may be encoded in a “programmable” form into a computer-readable medium, wherein, as used herein, the term “computer-readable medium” means any storage or transmission medium that participates in providing instructions and / or data to a computer for execution and / or processing. Examples of storage media include floppy disks, hard disks, optical disks, magneto-optical disks, CD-ROMs, CD-Rs, magnetic tapes, non-volatile memory cards, ROMs, DVD-ROMs, Blu-ray discs, solid-state drives, and network attached storage (NAS), whether such devices are internal to or external to a computer. Files containing information may be “stored” on a computer-readable medium, where “stored” means recording the information so that it can be accessed and retrieved later via a computer.

[0104] Any method step of this disclosure, or method steps performed by the system of this disclosure, may be performed by programming, wherein the programming may be written in one or more of any number of computer programming languages. Such languages ​​include, for example, Java (Sun Microsystems, Inc., Santa Clara, California), Visual Basic (Microsoft Corp., Redmond, Washington), and C++ (AT&T Corp., Bedminster, New Jersey), as well as many other languages.

[0105] Reagent test kit

[0106] This disclosure also provides a kit. The kit includes one or more reagents that can be used to carry out the methods of this disclosure. In some aspects, the kit includes any reagents, apparatus, instructions (e.g., located on one or more non-transitory computer-readable media) that can be used to carry out the methods of this disclosure, including any reagents, apparatus, instructions, etc., described in the "Methods" and "Systems" sections above this disclosure.

[0107] In some embodiments, the provided kit includes an adenine methyltransferase (e.g., N6-adenine DNA methyltransferase) that methylates adenine in genomic DNA, thereby marking the location of unbound regions (e.g., not bound to proteins) in the genomic DNA, and also includes method instructions for using the methyltransferase, which detects the location of methylated adenine in the genomic DNA by single-molecule sequencing, thereby determining the binding regions in the genomic DNA.

[0108] The components of this kit can be placed in separate containers, or multiple components can be placed in a single container. Suitable containers include single tubes (e.g., vials), single-well or multi-well plates (e.g., 96-well plates, 384-well plates, etc.).

[0109] The kit includes, for example, method instructions using adenine methyltransferase, which determines binding regions in genomic DNA by detecting the location of methylated adenine using a sequencer. In some embodiments, the kit includes method instructions using adenine methyltransferase, which determines the nucleosome location in genomic DNA based on the location of methylated adenine.

[0110] The instructions can be recorded on a suitable recording medium. For example, the instructions can be printed on a substrate such as paper or plastic. Therefore, the instructions can be included in the kit as packaging instructions or labeled on the kit container or its components (i.e., packaging or repackaging). In other embodiments, the instructions are presented as an electronic storage data file on a suitable computer-readable storage medium, such as a portable flash drive, DVD, CD-ROM, floppy disk, etc. In still other embodiments, the actual instructions are not included in the kit, but rather a method for obtaining the instructions from a remote source (e.g., via the Internet) is provided. An example of this embodiment is a kit that includes a URL where instructions can be viewed and / or downloaded. As with the instructions themselves, the method for obtaining the instructions is recorded on a suitable substrate.

[0111] practicality

[0112] The methods, kits, and systems disclosed herein can be used to generate databases containing information on chromatin structure, genomic DNA sequences, and the correlation between chromatin structure and the presence or absence of specific symptoms or diseases. This information can be used for disease diagnosis and / or prognosis.

[0113] In some respects, the database may contain information about changes in chromatin structure after treatment compared to before treatment. In other respects, this information can be used to monitor the effectiveness of treatment and adjust it as needed. The treatment may be immunotherapy, such as antibody therapy or small molecule therapy. The treatment may be targeted at cancer.

[0114] Notwithstanding the appended claims, this disclosure is also defined by the following numbered clauses:

[0115] 1. A method for identifying a protein-binding region in genomic DNA, the method comprising: contacting the genomic DNA with an adenine methyltransferase (A-MTase), wherein the A-MTase methylates adenine residues in a region of the genomic DNA that is not bound to the protein; performing single-molecule long-read sequencing on the contacted genomic DNA to detect locations in the genomic DNA lacking methylated adenine residues, thereby identifying the protein-binding region in the genomic DNA.

[0116] 2. The method according to Clause 1, wherein the A-MTase is N6-adenine methyltransferase (m6A-MTase).

[0117] 3. The method according to Clause 2, wherein the m6A-MTase is Hia5.

[0118] 4. The method according to Clause 2, wherein the m6A-MTase is EcoGII.

[0119] 5. The method according to Clause 2, wherein the m6A-MTase is Btr192IV.

[0120] 6. The method according to Clause 2, wherein the m6A-MTase is EcoGI.

[0121] 7. The method according to any one of Clauses 1-6, wherein the contact comprises contacting the isolated genomic DNA with the A-MTase.

[0122] 8. The method according to any one of Clauses 1-6, wherein the contact comprises contact with a cell containing the genomic DNA.

[0123] 9. The method according to Clause 8, wherein the contact comprises introducing a nucleic acid encoded by the A-MTase into the cell.

[0124] 10. The method according to Clause 9, wherein the A-MTase is fused with a cell-penetrating peptide that makes the plasma membrane of the A-MTase permeable.

[0125] 11. The method according to any one of clauses 1-10, wherein the sequencing is performed on at least a kilobase (kb) long extension of the genomic DNA.

[0126] 12. The method according to any one of clauses 1-10, wherein the sequencing is performed on at least a 3kb long extension of the genomic DNA.

[0127] 13. The method according to any one of clauses 1-12, wherein the sequencing comprises translocating the genomic DNA through a nanopore.

[0128] 14. The method according to any one of Clauses 1-13, wherein the sequencing comprises ligating one or more nanopore sequencing adapters to one or more ends of the genomic DNA.

[0129] 15. The method according to any one of clauses 1-14, wherein the sequencing includes detecting a signal indicating methylated adenine.

[0130] 16. The method according to Clause 15, wherein the signal is an electrical signal.

[0131] 17. The method according to any one of clauses 1-16, wherein the sequencing comprises multiple rounds of resequencing.

[0132] 18. The method according to Clause 17, wherein the multiple resequencing rounds comprise up to 20 rounds of sequencing.

[0133] 19. The method according to any one of clauses 1-18, wherein the sequencing comprises circular co-sequencing.

[0134] 20. The method according to any one of clauses 1-19, wherein the sequencing comprises single-molecule real-time (SMRT) circular co-sequencing (CCS).

[0135] 21. The method according to any one of clauses 1-20, wherein the genomic DNA is derived from mammalian cells.

[0136] 22. The method according to any one of clauses 1-21, wherein the genomic DNA is derived from cancer cells.

[0137] 23. The method according to any one of claims 1-22, wherein the cell is a normal cell, the method further comprising generating a chromatin accessibility map of the sequenced genomic DNA regions, wherein the map indicates chromatin regions that are not bound to the protein and are therefore accessible to the A-MTase and chromatin regions that are bound to the protein and are therefore inaccessible to the A-MTase.

[0138] 24. The method according to any one of clauses 1-23, the method further comprising generating a chromatin accessibility map of the genomic DNA of the test cells.

[0139] 25. The method according to Clause 24, wherein the test cells are cells from the subject.

[0140] 26. The method described in Clause 25, wherein the subject has or is suspected of having a disease.

[0141] 27. The method described in Clause 26, wherein the disease is cancer.

[0142] 28. The method according to any one of clauses 24-27, the method comprising comparing the chromatin accessibility map of the test cell with the chromatin accessibility map of the normal cell, wherein the test cell and the normal cell are of the same cell type; and comparing the genomic DNA sequences of the test cell and the normal cell, wherein a difference in the chromatin accessibility map indicates a change in the chromatin structure of the test cell, wherein no difference in the chromatin accessibility map but a difference in the genomic DNA sequence indicates that the sequence difference is not related to a change in chromatin structure, and wherein a difference in both the chromatin accessibility map and the genomic DNA sequence indicates that the sequence difference is related to a change in chromatin structure.

[0143] 29. The method according to Clause 28, the method further comprising generating a database containing information about the chromatin accessibility map, potential genomic DNA sequences, and, if any, relevance to the condition or disease.

[0144] 30. The method according to any one of clauses 24-29, wherein the normal cells and the test cells are epithelial cells, leukocytes, glial cells, osteoblasts, or chondrocytes.

[0145] 31. The method according to any one of clauses 24-30, wherein the normal cells and the test cells comprise a plurality of cells.

[0146] 32. The method according to Clause 31, wherein the plurality of cells comprises at least 10 cells, at least 30 cells, at least 100 cells, at least 300 cells, or at least 10,000 cells.

[0147] 33. The method according to any one of clauses 24-32, wherein the chromatin accessibility map covers at least 10% of the chromatin.

[0148] 34. The method according to any one of clauses 24-33, wherein the chromatin accessibility map covers at least 30%, at least 50%, or at least 80% of the chromatin.

[0149] 35. The method according to any one of clauses 24-34, wherein the chromatin accessibility map covers at least 10% of the genome of the cell.

[0150] 36. The method according to any one of clauses 24-35, wherein the chromatin accessibility map covers at least 20%, at least 30%, at least 50%, or at least 80% of the genome of the cell.

[0151] 37. The method according to any one of clauses 1-36, wherein the protein comprises a nucleosome.

[0152] 38. The method according to any one of clauses 1-36, wherein the protein comprises a transcriptional regulatory factor.

[0153] 39. The method according to Clause 38, wherein the transcriptional regulatory factor is a transcriptional repressor.

[0154] 40. The method according to clause 38, wherein the transcriptional regulatory factor is a transcriptional activator.

[0155] 41. A kit comprising:

[0156] Adenine methyltransferase (A-MTase);

[0157] Sequencing adapters; and

[0158] Instructions to contact genomic DNA with the A-MTase, wherein the A-MTase methylates adenine residues in regions of the genomic DNA that are not protein-binding, thereby connecting the sequencing adapter to the genomic DNA and performing single-molecule long-read sequencing on the contacted genomic DNA to detect locations in the genomic DNA lacking methylated adenine residues, thereby identifying regions in the genomic DNA that are protein-binding.

[0159] 42. The kit according to Clause 41, wherein the A-MTase is N6-adenine methyltransferase (m6A-MTase).

[0160] 43. The kit according to Clause 42, wherein the m6A-MTase is Hia5.

[0161] 44. The kit according to Clause 42, wherein the m6A-MTase is EcoGII.

[0162] 45. The kit according to Clause 42, wherein the m6A-MTase is Btr192IV.

[0163] 46. ​​The kit according to Clause 42, wherein the m6A-MTase is EcoGI.

[0164] 47. The kit according to any one of claims 41-46, wherein the contact comprises contacting the isolated genomic DNA with the A-MTase.

[0165] 48. The kit according to any one of claims 41-46, wherein the contact comprises contact with a cell containing the genomic DNA.

[0166] 49. The kit according to any one of claims 41-46, wherein the A-MTase comprises a cell-penetrating peptide fused to its N-terminus or C-terminus, and wherein the plasma membrane of the A-MTase is permeable.

[0167] 50. The kit according to any one of claims 41-46, wherein the contact comprises introducing a nucleic acid encoded by the A-MTase into the cell.

[0168] 51. The kit according to any one of clauses 41-50, wherein the sequencing is performed on a length extension of at least 1 kilobase (kb) of the genomic DNA.

[0169] 52. The kit according to any one of clauses 41-50, wherein the sequencing is performed on at least a 3kb long extension of the genomic DNA.

[0170] 53. The kit according to any one of clauses 41-52, wherein the genomic DNA is derived from cancer cells.

[0171] 54. A method for visualizing chromatin regions that are not bound to proteins and are spatially available as substrates for intracellular adenine methyltransferases (A-MTases), the method comprising:

[0172] 55. Contact the cells with the A-MTase; and

[0173] 56. Detect the presence of methylated adenine in the cells.

[0174] 57. The method according to Clause 54, wherein the detection of the presence of methylated adenine in the cells comprises contacting the cells with an antibody that specifically binds to methylated adenine.

[0175] 58. The method according to Clause 54, wherein the antibody can be detected as a marker.

[0176] 59. The method according to Clause 56, wherein the detectable marker comprises a fluorophore.

[0177] 60. The method according to any one of clauses 54-57, wherein the method further comprises staining genomic DNA in the cell.

[0178] 61. The method according to any one of clauses 54-58, wherein the method further comprises contacting the cells with an antibody that specifically binds to RNA polymerase II (Pol II).

[0179] 62. The method according to Clause 59, wherein the antibody specifically binds to Pol II Ser5Phos or Pol II Ser2Phos.

[0180] 63. The method according to any one of clauses 54-60, comprising measuring the mean nuclear intensity and / or nuclear spot intensity of the specific signal of the methylated adenine.

[0181] 64. The method according to any one of clauses 54-61, wherein the A-MTase is N6-adenine methyltransferase (m6A-MTase).

[0182] 65. The method according to Clause 62, wherein the m6A-MTase is Hia5, EcoGII, Btr192IV, or EcoGI.

[0183] 66. The method according to clause 62 or 63, wherein the detection of the presence of methylated adenine in the cells includes the detection of m6A. Example

[0184] As can be understood from the above disclosure, this disclosure has a wide range of applications. Therefore, the following embodiments are set forth to provide those skilled in the art with a complete disclosure and description of how to make and use the invention, and are not intended to limit the scope of the inventors' invention, nor to represent that the following experiments are all or only the experiments performed. Those skilled in the art will be able to readily identify various non-critical parameters that can be changed or modified to produce substantially similar results. Therefore, the following embodiments are set forth to provide those skilled in the art with a complete disclosure and description of how to make and use the invention, and are not intended to limit the scope of the inventors' invention, nor to represent that the following experiments are all or only the experiments performed. Efforts have been made to ensure the accuracy of the figures used (e.g., quantities, dimensions, etc.), but some experimental errors and biases should be taken into account.

[0185] Materials and methods

[0186] m6A-MTase was isolated and cloned. pHia5ET and pHinET were kindly provided by Monika Radlinska (M. Drozdz, et al., Nucleic Acids Res. 40, 2119–2130 (2012)). Codon-optimized versions of Btr192IV (GenBank: CP003745) and EcoGI (GenBank: AFST01000004) were synthesized into gBlocks via IDT and cloned into the aforementioned pET vectors using NdeI and XhoI restriction sites, generating pBtr192IVET and pEcoGIET vectors, respectively. Cloning was performed in 5-αF'Iq competent E. coli cells (NEB C2992H).

[0187] MTase protein was generated. The vectors pHia5ET, pHinET, pBtr192IVET, and pEcoGIET were transformed into T7Express lysY competent *E. coli* cells (NEB C3010I) for recombinant protein expression. The overnight culture was added to two 1L LB media supplemented with 100 μg mL⁻¹ ampicillin and incubated at 37°C with shaking until the OD600 reached 0.8–1.0. Isopropyl-β-D-1-thiogalactopyranoside (IPTG) was added at a final concentration of 1 mM, and the cells were incubated at 20°C with shaking for 4 hours. Cells were clumped at 5000xg for 10 minutes at 4°C (all subsequent steps were performed at 4°C) and resuspended in 35 mL of lysis buffer (50 mM HEPES, pH 7.5; 300 mM NaCl; 10% glycerol; 0.5% Triton X-100; 10 mM β-mercaptoethanol) supplemented with 2X completely EDTA-free protease inhibitor mixture (Roche 11873580001). Cells were lysed on ice at 50% amplitude for 10 minutes using a sonication probe (Qsonica Q125), with an on / off cycle of 30 seconds (lasting 20 minutes), followed by centrifugation at 40,000xg for 1 hour. Ni-NTA agarose (Qiagen 30210) was prepared by washing 5 mL of the slurry with 30 mL of equilibration buffer (50 mM HEPES, pH 7.5; 300 mM sodium chloride; 20 mM imidazole) and centrifuging at 500 x g for 3 min. This step was repeated once. The clarified cell lysate was mixed with the agarose and vortexed at 4 °C for 1 h, then poured into the column. The column was washed with 20 mL of buffer 1 (50 mM HEPES, pH 7.5; 300 mM sodium chloride; 50 mM imidazole) and 15 mL of buffer 2 (50 mM HPES, pH 7.5; 300 mM sodium chloride; 70 mM imidazole), and then 15 mL of elution buffer (50 mM HEPES, pH 7.5; 300 mM sodium chloride; 250 mM imidazole) was added. Add the eluent to a 10K Amicon Ultra-15 ultrafiltration tube and centrifuge at 3,220 x g for 15 minutes increments. Replace the EB buffer with 15 mL of protein resuspension buffer (50 mM Tris, pH 7.5; 50 mM potassium chloride; 1 mM DTT; 10 mM EDTA; 2X completely EDTA-free protease inhibitor mixture). After several 15-minute rotations, reduce the volume to below 500 μL and transfer the liquid to a 1.5 mL Eppendorf LoBind tube. Replenish the protein to 200 μg / mL with filtered sterile bovine serum albumin (BSA) solution and 30% glycerol, and store at -20°C.

[0188] In vitro MTase activity assessment. Substrate DNA was prepared by PCR using primers containing the 759-base pair region of the hydroxymethylbilirubin synthase (HMBS) gene promoter containing four GATC sequences from K562 genomic DNA. The PCR fragments were purified using the Monarch PCR & DNA Cleanup Kit (NEB T1030S) according to the manufacturer's instructions. Eleven 60 μL MTase reactions were prepared using 1 μg of substrate DNA and alternately 2-fold and 5-fold enzyme dilutions (10, 5, 1, 0.5, 0.1, 0.05, 0.01, 0.005, 0.001, 0.0005, and 0.0001 μL MTase) in buffer A (15 mM Tris, pH 8.0; 15 mM sodium chloride; 60 mM potassium chloride; 1 mM EDTA, pH 8.0; 0.5 mM EGTA, pH 8.0; 0.5 mM spermidine) supplemented with 0.8 mM S-adenosylmethionine (NEB B9003S). An Mtase-free negative control was also prepared. Reactions were mixed by tapping the PCR tubes and rapidly rotating them, then incubated at 37°C for 1 hour. Each reaction was terminated using the Monarch PCR & DNA Cleanup Kit, and the DNA was purified by elution in 20 μL EB buffer. Twelve restriction enzyme digests were prepared by mixing 15 μL of each purified DNA sample with 1 μL of DpnI (NEB R0176S) and 4 μL of 10X CutSmart buffer (NEB) in 40 μL of reactants. The reactants were carefully mixed by tapping and incubated at 37°C for 1.5 h. 1 μL of each reactant was mixed with 2 μL of 6X purple gel loading dye (NEB B7024S) and 11 μL of H2O and run on a 1.2% agarose gel containing 1X GelGreen nucleic acid stain (Biotium 41005) at 130 V for approximately 1.5 h. The gels were imaged on a GE Typhoon FLA 9500 laser scanner. MTase activity was determined by the highest MTase dilution, which resulted in methylation of 1 μg of DNA substrate, leading to incomplete DNA molecules after DpnI digestion.

[0189] MTase-seq. Drosophila S2 cells were cultured at room temperature in 1X Schneider's Drosophila medium (Gibco 21720-024) supplemented with 10% HI FBS (Gibco 16140-063) and 1% penicillin-streptomycin (Gibco 15140-122), and incubated at 75 cm⁻¹. 2The experiment was conducted in flasks to achieve approximately 90% fusion and over 95% viability. Cells were washed with PBS, resuspended in PBS, and counted using a Countess automated cell counter. 3 million Drosophila S2 cells per sample were clumped for 5 minutes at 250 x g. The cell clumps were resuspended in 60 μL of buffer A per sample and aliquoted into PCR tubes. 60 μL of cold 2X lysis buffer (0.1% IGEPAL CA-630 added to buffer A) was added to each tube, the tubes were tapped to mix, and the tubes were incubated on ice for 10 minutes. The samples were clumped for 5 minutes at 350 x g at 4°C, and the supernatant was removed. K562 and HeLa cell nuclei were isolated as previously described (REThurman, et al., Nature. 489, 75-82 (2012)). Using a wide-bore pipette tip, gently resuspend the nucleus clumps in 57.5 μL of buffer A and transfer them to a 37°C thermal cycler. Add 1 μL of MTase diluted in buffer A and 1.5 μL of SAM (final 0.8 mM), then carefully mix by pipetting up and down 10 times using a wide-bore pipette tip and a multichannel pipette. Incubate the reaction mixture for 20 minutes, then terminate with 3 μL of 20% SDS (final 1%) and transfer to a new 1.5 mL microcentrifuge tube. Increase the sample volume by adding 130 μL of buffer A and an additional 7 μL of 20% SDS. Mix all samples with 2 μL of RNase A (Invitrogen AM2271) and incubate at 37°C for 1 hour, then mix with 2 μL of proteinase K (NEBP8107S) and incubate at 55°C for another hour. Add 200 μL (1:1) of phenol, chloroform, and isoamyl alcohol (25:24:1 ratio) (saturated with 10 mM Tris, pH 8.0, 1 mM EDTA) to purify the DNA, then vigorously invert and incubate at room temperature for 10 min. Centrifuge the extract at 17,900 x g for 10 min, then transfer the supernatant to a new microcentrifuge tube. Add 200 μL of chloroform and isoamyl alcohol (24:1 ratio) to all samples and repeat the extraction procedure, removing residual phenol on the second extraction. Transfer the aqueous phase to a new microcentrifuge tube and precipitate the DNA by adding 0.1 volume of 3 M sodium acetate, 1 μL of LlycoBlue precipitant (Invitrogen AM9515), and 2.5 volumes of 100% ice-cold ethanol. Invert all samples several times, then rapidly rotate and store overnight at -20°C. DNA was clumped by centrifuging at 20,000 x g for 10 minutes at 4 °C, followed by rinsing with 1 mL of ice-cold 70% ethanol. The tubes were inverted on a rack and air-dried for 15 minutes, then resuspended in 54 μL of 10 mM Tris (pH 7.5).

[0190] All samples were transferred to MicroTUBE-50AFA fiber spiral cap sonication tubes (Covaris 520166) and individually sonicated on a Covaris M220 focused sonicator (peak power: 75.0; duty cycle: 15.0%; cycles / pulses: 200; duration: 720 s; water bath temperature: 20 °C) to obtain fragments of approximately 100 base pairs in length. Samples were then run on 1.2% agarose gels containing 1X GelGreen nucleic acid stain, and the excised bands were purified using the QIAquick Gel Extraction Kit (28704). Two replicates were then pooled to provide sufficient input for library construction, and quantification was performed using the Qubit ds DNA HS Detection Kit (Invitrogen, Q32851). All libraries were prepared using the NEBNext Ultra II DNA Library Preparation Kit for Illumina (NEB E7645S) and NEBNext Multiplex Oligos (NEB E7335S & E7500S) for Illumina index primer sets 1 and 2. 1 μg of gel-extracted DNA was mixed in 50 μL of buffer EB (10 mM Tris-Cl, pH 8.5) with 7 μL of end-repair reaction buffer and 3 μL of end-repair enzyme mixture in a PCR tube to perform DNA end repair. The samples were then placed in a thermal cycler and the end-repair program was run (30 min at 20°C, 30 min at 65°C, then maintained at 4°C). To terminate the repair, 2.5 μL of Illumina adapters, 30 μL of Ligation Master Mix, and 1 μL of Ligation Enhancer were added. Samples were incubated at 20°C for 15 minutes, followed by the addition of 3 μL of USER enzyme, and then incubated at 37°C for another 15 minutes. 116 μL (1.2X) Agencourt AMpure XP magnetic beads (Beckman Coulter A63880) were added to each library, and the samples were incubated at room temperature for 10 minutes followed by magnetic separation for 10 minutes. The supernatant was discarded, and the magnetic beads were washed twice on the magnet with 200 μL of fresh 80% ethanol, then air-dried for 4 minutes. The samples were removed from the magnet, and 50 μL of 10 mM Tris (pH 7.5) was added to each sample, followed by incubation at room temperature for 10 minutes. After magnetic separation for 10 minutes, the supernatant was transferred to a new LoBind tube.

[0191] 20 μL of 100 μM blocking oligonucleotide (AGATCGGAAGCGTC (SEQ ID NO:5)) was added to each library, and the samples were incubated at 95 °C for 10 min to denature the DNA, then transferred to an ice / water blended smoothie for rapid cooling for 10 min. 325 μL of 0.1X TE buffer, 100 μL of 5X IP buffer (50 mM Tris, pH 7.5; 750 mM sodium chloride; 0.5% IGEPAL CA-630) and 5 μL (5 μg) of anti-N6-methyladenosine antibody (Millipore SigmaABE572) were added to each sample. The samples were incubated in a vortex mixer at 4 °C for 12 h. After magnetic separation, the supernatant was removed and the sample was washed four times in 1X IP buffer (10 mM Tris, pH 7.5; 150 mM sodium chloride; 0.1% IGEPAL Ca-630), then resuspended in 25 μL of 1X IP buffer to prepare 25 μL of protein A immunomagnetic beads (Invitrogen 10001D) for each sample. The 25 μL of prepared protein A beads was added to each sample, and the samples were then incubated at 4°C for 4 hours by rotation. The samples were washed six times with 750 μL of 1X IP buffer, rotating at 4°C and incubating in wash buffer for 5 minutes each time. 48 μL of proteinase K digestion buffer (20 mM HEPES, pH 7.5; 1 mM EDTA; 0.5% SDS) and 2 μL of proteinase K were added to the beads, and the samples were incubated at 50°C with shaking at 1200 rpm for 1 hour to elute the DNA. The samples were then separated using a magnet, and the supernatant was transferred to a new tube. Following the same procedure as described above, 90 μL (1.8X) AMPure XP magnetic beads were added to each sample, but eluted in 17 μL of 10 mM Tris (pH 7.5), followed by quantification using the Qubit ss DNA Detection Kit (Invitrogen, Q10212). Libraries were amplified by adding 5 μL of universal PCR primers, 5 μL of index primers (both available in NEBNext Multiplex Oligos Illum ina index primer sets 1 and 2), and 25 μL of NEBNext Ultra II Q5Master Mix to each sample, followed by PCR amplification (program: amplify at 98°C for 30 seconds; cycle 5–7 times at 98°C for 10 seconds each, then amplify at 65°C for 75 seconds; finally incubate at 65°C for 5 minutes). Following the same procedure as described above, 60 μL (1.2X) AMPure XP magnetic beads were added to each sample, but eluted in 33 μL of 10 mM Tris (pH 7.5).The library was quantified using the Qubit ds DNA HS assay kit, and particle size distribution was examined on an Agilent Bioanalyst high-sensitivity DNA chip. The library was sequenced using an Illumina HiSeq 4000 with 76 bp reads at paired ends and read depths of 10-30 million.

[0192] DNA seI-seq. Drosophila S2 cells were cultured as described above, and cell nuclei were isolated as described above. After supplementing with Ca2+... + Cell nuclei were cultured in buffer A at 37°C with the maximum concentration of DNA seI (Sigma) for 3 minutes. The digestion reaction was terminated with stop buffer (50 mM Tris-HCl, 100 mM sodium chloride, 0.1% SDS, 100 mM EDTA, 1 mM spermidine, 0.5% spermine, pH 8.0), and the samples were treated with proteinase K and RNase A. Small double-hit fragments (<750 bp) were recovered using AMPure XP magnetic beads, and samples were prepared using the Illumina Library Kit as described previously (S. John, et al., Current Protocols in Molecular Biology (John Wiley & Sons, Inc., Hoboken, NJ, USA, 2013; vol. Chapter 27, pp. 21.27.1-21.27.20)). The previously published K562 and Hela DNA seI-seq datasets were used for analysis (REThurman, et al., 2012 (see above)).

[0193] MTase-seq and DNA seI-seq analyses were performed. As previously described, reads were mapped to the dm6 genome (J. Vierstra, et al., Science (80). 346, 1007-1012 (2014)). Signal trajectories were generated using BEDOPS (S. Neph, et al., Bioinformatics. 28, 1919–20 (2012)) and the signals were normalized to 1 million reads. Chromatin accessibility regions (hotspots) were identified using a hotspot algorithm (S. John, et al. Nat. Genet. 43, 264–268 (2011)) with an FDR 5% cutoff. For each hotspot quantified using S2 cell DNA seI data, the total number of normalized reads contained within that hotspot was quantified to identify the signal intensity of that element. The above procedure was repeated for each MTase-seq library. When comparing libraries, we used S2 DNA seI-seq hotspot calls as a list of genomic regulatory elements and quantified the signal intensity of these regions in different libraries as described above. Proximal promoter elements were defined as elements within + / - 500 bp of the transcription start site annotated in the NCBI RefSeq Curated gene list.

[0194] m6A dot hybridization. DNA from the nuclei of Drosophila S2 cells treated with different concentrations of MTase was isolated as described above, and the samples were quantified using Nanodrop. These DNA samples were diluted with 20X SSC buffer in 96-well plates and then denatured at 95°C for 10 min. Nitrocellulose membranes were moistened with 20X SSS buffer and then fixed in a HYBRI-DOT multiplexer (Life Technologies). After fixing the membranes in the multiplexer, the multiplexer was brought under vacuum, and 150 μL of 20X SSC buffer was added to all wells, followed by the denatured DNA samples. After removing all air bubbles, the vacuum was terminated, and the membranes were placed face up on dry Whatman filter paper and crosslinked with 125mJoules using a GS Gene Linker UV Chamber (Bio-Rad) with CL setting. The membrane was then washed with 20 mL of 1X TBS-T (10 mM Tris, pH 7.5; 0.25 mM EDTA; 150 mM sodium chloride; 0.1% TWEEN-20) and blocked for 1 hour at room temperature with 15 mL of 1X TBS-T + 5% skim milk. Rabbit polyclonal anti-N6-methyladenosine antibody (Millipore SigmaABE572) was diluted 1:1000 in 10 mL of 1X TBS-T + 5% skim milk and incubated overnight at 4°C on a shock absorber. The blot was washed three times with 20 mL of 1X TBS-T for a total of 15 minutes. Anti-rabbit IgG, HRP-linked secondary antibody (CellSignaling Technology 7074) was diluted 1:1000 in 10 mL of 1X TBS-T + 5% skim milk and incubated for 1 hour at room temperature. Repeat the washing process three times, develop the blot using Pierce ECL Plus Western blot substrate (ThermoScientific 32132), and image it on film.

[0195] Fiber-seq. Nuclei were isolated from 2.5 million Drosophila S2 cells or 300,000 K562 cells as described above. The MTase reaction was performed according to the above procedure, except that 20 units (1 μL) of Hia5 were used. Untreated Drosophila S2 replicas were also produced for comparison. DNA was purified as described above and then transferred to a MicroTUBE-50 AFA fiber spiral-cap sonication tube (Covaris 520166) and sonicated individually on a Covaris M220 focused ultrasonic disruptor (peak power: 75.0; duty cycle: 5.0%; cycles / pulses: 200; duration: 8 seconds; water bath temperature: 20°C) to obtain fragments of approximately 1.5 kb in length. Sequencing Kit 3.0 (Pacific Biosciences, Menlo Park, California, USA) was used to generate Pacific Biosystems sequencing libraries from the cut sample. Small library fragments were removed using BluePippin (SageScience, Beverly Hills, Massachusetts, USA), and each library was loaded onto a single SMRT cell.

[0196] Fiber-seq m6A methylation identification. Readings were mapped to the dm6 genome using PBAlign, and sub-readings for each ZMW in the library were extracted using bamseive to generate per-ZMW bam files. These per-ZMW bam files were then processed via ipdSummary using -identifym6A. For Drosophila S2 analysis, methylation identification was performed with a p-value cutoff of 0.001, and reads containing 14 or fewer sub-readings were discarded. To obtain longer reads in human samples, methylation identification was performed with a p-value cutoff of 0.02 (i.e., phred score > 16) in K562 analysis, and reads containing fewer than 10 sub-readings were discarded. Promoter positions in each read were annotated using the NCBI RefSeq Curated gene list, and expression of each gene was recorded using previously published RRNA-seq data (43). DNA seI hypersensitive elements (DHS) were labeled using the hotspot calls described above, and repeats were defined using RepeatMasker. Unless otherwise stated, mitochondrial readings and readings with more than 25% overlap with repeating elements are removed from the analysis, with the exception of K562 data, in which case readings with more than 50% overlap with repeating elements are removed.

[0197] Methyltransferase-accessible DNA sequence (MAD) identification. The m6A methylation events identified for each read were aggregated to identify MTase-sensitive regions. Specifically, for each m6A, all m6A events within a 50 bp distance were concatenated to invoke a larger methyltransferase-accessible DNA sequence (MAD). MTase protection sites (MPS) were defined as regions not included in the MAD during each read. All MADs overlapping with the DHS were identified, and then the widest MAD was captured to identify the MADs overlapping with the DHS. This widest MAD was used to identify whether the DHS was closed or open / accessible in a single read. When comparing overlapping reads, the above process was repeated for each overlapping read, and the median difference in MAD size for each read of each DHS was calculated.

[0198] Quantitative nucleosome phasing / localization. The position of each nucleosome is defined as the center of each MPS with a width between 65 and 200 bases. All readings overlapping this position are identified, and the position of each nucleosome on these overlapping readings is similarly defined. To calculate nucleosome phasing, the distance between the nucleosome position on a reading and the nearest nucleosome position on each overlapping reading is determined. If multiple overlapping readings exist, the median distance is used. First, TSSs overlapping with DHSs are identified, then the largest MADs overlapping these DHSs are identified, and the nucleosome is localized relative to the TSS. The position of the MPS relative to the MAD is then determined based on the gene's reading frame. Readings overlapping with two or more TSSs are removed from the analysis. The relative phasing of the nucleosome at each position is calculated by dividing the number of nucleosomes offset by 0–39 bp by the number of nucleosomes offset by 40–79 bp and then performing a log2 transformation on the result.

[0199] Example 1: Nonspecific M 6A-MT ASE Selective labeling of chromatin accessibility sites

[0200] The primary structure of chromatin comprises arrays of nucleosomes truncated by short regulatory regions containing transcription factors and other non-histone proteins. This structure is fundamental to genome function, but its details remain unclear at the level of individual chromatin fibers (the basic units of gene regulation). For example, while nucleosomes are a major barrier restricting transcription factor access to DNA, the localization and occupancy of nucleosomes arranged along individual chromatin fibers in vivo are not well elucidated. Therefore, the following aspects remain unclear: how nucleosomes are precisely aligned along the same extended chromatin template; the interactions between accessible regulatory DNA and nucleosomes on individual chromatin fibers; the extent to which regulatory regions encoded by a given DNA are driven on different chromatin fibers within a cell population; and the extent to which nearby regulatory regions are coordinated and driven on the same chromatin template. Addressing these questions requires sequencing individual chromatin fibers, which is currently impossible with single-cell or batch analysis methods.

[0201] A method has been developed to record the primary structure of chromatin onto its basic DNA template at single nucleotide resolution, thereby enabling the simultaneous identification of genetic and epigenetic features of multi-kilobase segments in the genome. Current methods for mapping chromatin and regulatory structures sample large amounts of chromatin fibers and rely on the use of DNA se I (DS Gross, W. Garrard, Annu. Rev. Biochem. 57, 159–97 (1988); RE Thurman, et al., 2012 (see above)), micrococcal nuclease (M. Noll, RD Kornberg, J. Mol. Biol. 109, 393–404 (1977); DE Chones, et al., Cell. 132, 887–898 (2008)), restriction endonucleases (E. Lieberman-Aiden, et al., Science (80). 326, 289–293 (2009)), transposases (JD Buenrostro, et al.). Nucleases such as T. Kouzarides, Cell. 128, 693-705 (2007) dissolve chromatin. CpG and GpC methyltransferases can label accessible cytosine in a dinucleotide environment without digesting DNA (T. K. Kelly, et al., Genome Res. 22, 2497-2506 (2012); AR. Krebs, et al., Mol. Cell. 67, 411-422.e4 (2017)). However, the average resolution is low due to the sporadic occurrence, mutation, and linear clustering of CpG and GpC dinucleotides in animal genomes, as well as the confounding effects of endogenous cytosine methylation mechanisms (AP. Bird, Nature. 321, 209-13 (1986)).

[0202] Unlike cytosine, adenine bases in DNA are almost entirely un-endogenously methylated in eukaryotes (Q. Xie, et al., Cell. 175, 1228-1243.e20 (2018)), and occur at an average frequency of approximately one in every two DNA base pairs in animal genomes without exhibiting cytosine-guanine dinucleotide aggregation and the extended desert feature. Therefore, there is a search for non-specific (i.e., non-sequence context-dependent) N-nucleotides with high efficiency, high stability, and molecular weights similar to non-specific nucleases such as DNA seI (approximately 30 kDa). 6 -Adenine DNA methyltransferase (m6A-MTase), which can approach the protein-DNA interface at nucleotide resolution ( Figure 1AFive distinct nonspecific DNA m6A-MTases were isolated (M. Drozdz, et al., Nucleic Acids Res. 40, 2119–2130 (2012); B.P. Anton, et al., PLoSOne. 11, e0161499 (2016); I.A. Murray, et al., Nucleic Acids Res. 46, 840–848 (2018); G. Fang, et al., Nat. Biotechnol. 30, 1232–1239 (2012)), and it was demonstrated that with increasing amounts of each enzyme, the treatment of unchromatinized and chromatinized DNA templates ( Figure 1B This leads to a monotonic increase in adenine methylation, which is compatible with single-hit kinetics.

[0203] To determine the selectivity of m6A-MTase for accessible DNA templates within nuclear chromatin, the distribution of DNA seI cleavage after treatment of Drosophila S2 cell nuclei with DNA seI (established criteria for labeling accessible DNA templates (DS Gross, WT Garrard, 1988 (see above); RE Thurman, et al., 2012 (see above)) and the distribution of m6A-DNA after S2 nuclei were exposed to five increasing concentrations of adenine methyltransferases were compared. Figure 1C Immunoprecipitation and massively parallel sequencing of short (median 110 bp) DNA fragments containing m6A from extracted genomic DNA (MTase-seq) revealed the genomic distribution of m6A-DNA, reflecting the DNA seI fragmentation density quantified by DNA se-seq. Figure 1D m6A-MTase exhibits high selectivity for accessible DNA and can be quantitatively characterized for DNA seI hypersensitive sites (DHS). Figure 1D Furthermore, m6A-MTase showed that selectivity for DHS decreased with increasing enzyme concentration. Figure 1EThis is similar to the enzymatic action of DNA seI (or micrococcal nuclease (MNase)) on chromatin substrates, due to increased digestion of more but shorter internucleosome junction regions (H. Weintraub, M. Groudine, Science (80). 193, 848–856 (1976); KS Bloom, JN Anderson, Cell. 15, 141–150 (1978); J. Mieczkowski, et al., Nat. Commun. 7, 11485 (2016)). MTase-seq quantification of DNA accessibility shows high reproducibility at both proximal and distal regulatory elements of the promoter. Figure 1F Among them, enzyme Hia5 showed the highest efficiency. These results indicate that nonspecific m6A-MTase provides a quantitative probe for intrachromatin DNA accessibility and demonstrates that m6A efficiently replicates chromatin structures onto DNA with near-nucleotide resolution.

[0204] Figures 1A-1F This study demonstrates nonspecific m6A-MTase selective labeling of chromatin accessibility sites. Figure 1A A schematic diagram of a method based on cleavage and m6A-MTase labeling of chromatin accessibility sites is shown. Figure 1B This study demonstrates the dot blot hybridization quantification of m6A-modified DNA from the nucleus of Drosophila S2 cells after treatment with different amounts of m6A-MTase Hia5. Figure 1C A schematic diagram of the MTase-seq experiment is shown. Figures 1D-1E This demonstrates the effect of treating S2 cell nuclei with five different m6A-MTases ( Figure 1D After ) or increase the amount of m6A-MTase Hia5 ( Figure 1E Afterwards, genomic loci were compared to show the relationship between DNA seI-seq signals and MTase-seq signals. Comparisons were made between untreated nuclei and input controls for m6A-IP-seq. Figure 1F The diagram shows a comparison of MTase-seq signals in S2 cells treated with Hia5 with MTase-seq signals or DNA seI-seq signals in cells treated with EcoGII (top) (bottom).

[0205] Example 2: F IBER-SEQ Base pair resolution maps revealing the structure of individual chromatin fibers

[0206] Sequencing m6A along a linear pattern of a multi-kilobase chromatin template at nucleotide resolution reconstructs the primary structure of chromatin fibers; this process is called "Fiber seq". Figure 2ATo implement Fiber-seq, a single-molecule DNA sequencer was used to distinguish methylated and unmethylated adenine residues based on the DNA polymerase kinetics of the bases during sequencing (G. Fang, et al., Nat. Biotechnol. 30, 1232–1239 (2012)). Highly accurate nucleotide resolution was achieved by performing more than 15 resequencing cycles on each chromatin fiber (referred to as circular concord sequencing (CCS) (KJ Travers, et al., Nucleic Acids Res. 38 (2010), doi:10.1093 / nar / gkq543)), enabling base recognition of modified nucleotides with accuracy comparable to that of unmodified nucleotides.

[0207] To create chromatin templates, S2 cell nuclei were treated with m6A-MTase (using conditions similar to MTase-seq), followed by PCR-free library construction of high-molecular-weight DNA extracted from treated or untreated nuclei. The resulting libraries were subjected to CCS on a Pacific Biosciences single-molecule DNA sequencer, which provides very high base recognition accuracy. Although untreated nuclei showed the lowest m6A signal, over 98% of single-molecule reads from m6A-MTase-treated cells showed some degree of adenine methylation. Figure 2B 51% of m6A showed another m6A marker within four nucleotides. These results, along with the aforementioned results, demonstrate that Fiber-seq can convert chromatin templates into linear reads of DNA accessibility.

[0208] To reconstruct the primary structure of chromatin, the distribution of m6A nucleotides along the genome was analyzed. It was observed that m6A nucleotides were clearly clustered into short, adjacent regions spanning tens to hundreds of base pairs, separated by unmodified nucleotide extensions. Figure 2C Two types of methyltransferase-accessible DNA sequences (MADs) were identified: (1) sequence elements with an average length of 174 bp that corresponded to the DNA seI hypersensitive site. Figure 2C (D); and (2) more short sequence elements with an average length of 51 bp and regular spacing, similar to the expected size and distribution of the internucleosome connection regions (RVChereji, et al., Nucleic Acids Res. 44, 1036–1051 (2016))( Figure 2C In the first class, the number of aggregated m6A-tagged bases of all fibers overlapping with DHS is proportional to the aggregation density of DNA seI cleavage obtained from the somatic cell nucleus. Figure 2E ).

[0209] The DNA occupied by nucleosomes can be simply defined by the obvious absence of m6A between strongly labeled linker regions, indicating that m6A MTases are generally unable to access the DNA encapsulated by nucleosomes. Figure 2C The presence of these enzymes (F) may be due to their modification of adenine via base flipping (JR Horton, et al., J. Mol. Biol. 358, 559-570 (2006)). Fiber-seq data precisely record nucleosome positions up to several thousand bases, with high-quality fiber sequences producing an average of more than seven well-defined nucleosomes. Figure 2G This allows us to assess key nucleosome occupancy characteristics.

[0210] Next, we used fiber-seq data to explore some fundamental questions regarding chromatin structure and the interaction between nucleosome occupancy and regulation of DNA accessibility. The average length of nucleosome repeat sequences (NRs) was observed to be 179 bp, consistent with previous reports (RVChereji, et al., 2016 (see above)). However, the length of NRs varied significantly across individual fibers, with 75% falling between 157 and 202 bp. Figure 2G It is generally believed that the regions immediately adjacent to the promoter and enhancer contain tightly packed, regularly spaced nucleosomes (B. Lai, et al., Nature. 562, 281–285 (2018); K. Struhl, E. Segal, Nat. Struct. Mol. Biol. 20, 267–273 (2013)). However, although Fiber-seq data show that the NR length within the nucleosome arrays on either side of the promoter and distal DHS is shorter, the structural heterogeneity in these regions is significantly greater than that in nucleosome arrays farther from the DHS. Figure 2H Therefore, nucleosome compression does not appear to produce or enhance structural coherence; more precisely, at the individual template level, chromatin assembly around regulatory regions appears to be highly dynamic.

[0211] Figures 2A-2H This shows a base pair resolution map of the structure of a single chromatin fiber revealed by Fiber-seq. Figure 2A A schematic diagram of Fiber-seq is shown. Figure 2B The percentage of chromatin fibers containing m6A methylated bases is shown in PacBio CCS, obtained from DNA isolated from untreated and Hia5-treated S2 cell nuclei, respectively. Figure 2CGenomic loci comparing the relationships between DNAseI-seq, MTase-seq, and Fiber-seq are shown. Individual PacBio reads / chromatin fibers are marked with gray lines, and m6A-methylated bases are marked with purple dashed lines. Inserts are DHS stained by base, where m6A-sensitive bases are gray (e.g., all A / T residues) and m6A-methylated bases are purple. Figure 2D A pod plot showing the relationship between DNA seI-seq signals and Fiber-seq m6A signals for each DHS is presented. Pearson correlation analysis was performed on all DHS. *p-value < 0.001 (rank-sum test). Figure 2E A bar chart showing the MAD widths of all MADs identified outside the DHS (grey), within the TSS distal DHS (blue), and within the promoter DHS (green). Individual box plots are shown below. *p-value < 0.001 (rank-sum test). Figure 2F A heatmap of a single m6A tag around a 5–100 bp long MAD that does not overlap with DHS is shown (top). The upper bar chart shows strong protection against methylation of nucleosome-binding bases (bottom). Figure 2G The bar chart shows all NR lengths (left) and the number of NRs identified on a single chromatin fiber (right). Figure 2H The bar chart shows the average NR length per fiber (left) and the average bp difference in NR length per fiber (right) for fibers without DHS (grey), with TSS distal DHS (blue), or with promoter DHS (green). Box-whisker plots are shown below. *p-value < 0.001 (rank-sum test).

[0212] Example 3: Coordinated driving of adjacent regulatory elements on the same chromatin fiber

[0213] How transcription factors and accessory proteins activate regulatory information encoded in genomic DNA is crucial for understanding cell state and fate determination (ABStergachis, et al., Cell. 154, 888–903 (2013)) and defining the mechanisms by which genetic variation within regulatory DNA influences phenotypic traits and disease risk (MTMaurano, R. et al., Science (80). 337, 1190–5 (2012)). To this end, the long-standing question needs to be answered: Is regulatory DNA driven entirely or entirely on any given template (instead of the standard nucleosome), or do actuating elements exist in alternating structures, with intermediate DNA accessibility mediated by intermittent nucleosome occupancy events? If the former, then the primary mechanism enabling the permeability of genetic variation within a given regulatory region may involve altering the frequency at which its homologous regions are driven, without involving the creation of alternative regulatory structures. Current methods for collecting data from cell populations (REThurman, et al., 2012 (see above)) cannot address the question of whether any regulatory element is driven in a wholly or entirely manner on a single chromatin template. Single-cell sampling methods also cannot solve this problem because these methods generate extremely scarce data and cannot continuously query any region of chromatin on an allele-specific basis (JDBuenrostro, et al., Nature. 523, 486-490 (2015)).

[0214] Of all 14,432 S2 cell DHSs with multiple overlapping fiber seq reads (an average of 5 high-quality reads per DHS), only 64% of the overlapping chromatin fibers showed a consistent MAD and were in an open state, while the remainder showed nucleosome separation and were in a closed state. Figure 3A A wider DHS indicates a higher likelihood of adopting an accessible / open state. Figure 3BThis aligns with the concept that the co-binding of additional TFs along DNA fragments can more effectively compete with nucleosomes for DNA occupancy (CCAdams, JLWorkman, Mol. Cell. Biol. 15, 1405–21 (1995); JAMiller, J. Widom, Mol. Cell. Biol. 23, 1623–32 (2003); LAMIRny, Proc. Natl. Acad. Sci. USA 107, 22534–22539 (2010)). Analysis of extended elements longer than 500 bp showed that although 85% of the chromatin fibers overlapping with these DHSs were in an accessible state, 65% of these accessible fibers were truncated by one or more co-occupied nucleosomes, and a high degree of positional variability was recorded on different fibers. Based on Fiber seq data, it was concluded that the accessibility of regulatory DNA is primarily driven by an all-or-nothing process. At the level of a single chromatin fiber, most regulatory DNA exists in one of two states (accessible or inaccessible), with macroelements additionally regulated by co-occupied nucleosomes.

[0215] This addresses the question of whether the accessibility of DNA on a regulatory element affects the behavior of adjacent elements on the same chromatin fiber. Regulatory DNA is highly clustered along the genome, and many control regions appear to be organized as compositions of multiple independent elements, such as locus control regions / “super-enhancers” that function on the whole locus (e.g., the β-globin locus control region) or gene-specific control clusters (e.g., the BCL11A enhancer region) (WA Whyte, et al., Cell. 153, 307–319 (2013); P. Diaz, et al., Immunity. 1, 207–17 (1994); L. Madsen, M. Groudine, Genes Dev. 8, 2212–26 (1994); F. Grosveld, et al., Cell. 51, 975–85 (1987)). If driving a regulatory element promotes the driving of neighboring elements on the same chromatin fiber, then it provides the mechanistic basis for clusters of gene regulatory elements and demonstrates that current models of gene regulatory structure and evolution do not account for the level of cis-integration function. It will also amplify the potential impact of regulatory genetic variation through cascading effects on neighboring elements.

[0216] 6% of the fiber-seq readings overlapped with multiple DHS, enabling the quantification of common drivers regulating DNA accessibility on the same chromatin fiber. Figure 3C The observation of co-driving of distal element pairs, promoter pairs, and their combinations suggests that adjacent regulatory DNA elements can be driven along the same chromatin fibers. Figure 3D Among them, adjacent distal modulation elements are significantly enriched in the context of being co-driven on the same fiber. Figure 3D This indicates that the accessibility of regulatory DNA at a distal element promotes the accessibility of adjacent elements. Therefore, these results provide a physicochemical basis for observing the aggregation of distal regulatory elements in animal genomes and suggest that genetic variants affecting the accessibility and function of regulatory DNA may generate local linkage effects in cis, thereby amplifying their potential influence in ways that current genome query technologies cannot capture.

[0217] Figures 3A-3D The coordinated drive of adjacent regulatory elements on the same chromatin fiber is shown. Figure 3A Genomic loci comparing the relationships between DNA seI-seq, MTase-seq, and Fiber-seq at the DHS site reveal overlapping chromatin fibers of open and closed chromatin at this DHS site. Figure 3B The proportion of DHS overlapping with accessible and closed fibers is shown for DHS classified according to their width (left) or proximity to TSS (right). *p value < 0.01 (z test). Figure 3C Genomic loci showing the relationships between DNA seI-seq, MTase-seq, and Fiber-seq at adjacent DHS sites are illustrated. Figure 3D This shows a comparison of the percentage of fibers containing accessible MAD at the DHS in two different element classes with the expected percentage for chromatin fibers containing two DHSs. *p value < 0.01 (z test).

[0218] Example 4: The effect of DNA regulation on nucleosome localization

[0219] Nucleosome localization is crucial for gene regulation and is determined by a variety of factors, including DNA sequence; competitive occupancy by sequence-specific DNA-binding proteins that form the boundary; the role of the nucleosome remodeling complex; and interaction with RNA polymerase (K. Struhl, E. Segal, Nat. Struct. Mol. Biol. 20, 267-273 (2013)). The relative contributions of these factors are currently unclear globally and cannot be studied at specific genomic locations. Existing analyses based on large amounts of cellular data (S. Baldi, et al., Mol. Cell. 72, 661-672.e4 (2018); GC Yuan, et al., Science (80). 309, 626–630 (2005); C. Jiang, BFPugh, Nat. Rev. Genet. 10, 161–172 (2009)) indicate that nucleosomes around accessible promoters are generally well-localized, while nucleosomes around distal regulatory elements are less well-localized. However, it remains unclear whether this localization is due to boundary conditions imposed by factor occupancy (and therefore accessibility) of regulatory DNA.

[0220] We infer that the boundary model of nucleosome localization can be directly tested by comparing the nucleosome positions around the regulatory elements on overlapping fibers where the regulatory elements are in an accessible state with those around the regulatory elements on overlapping fibers where the regulatory elements are in an alternating nucleosome-occupied (i.e., closed) state. Although the nucleosomes around the DHS are generally well localized ( Figure 4A However, analysis of single-fiber data indicates that these well-positioned nucleosomes mainly originate from the distal elements of the regulatory elements. Figure 4B ) or filaments in an accessible state upstream of the DNA seI hypersensitive promoter ( Figure 4C This indicates that the localization of nucleosomes at these locations largely depends on the driving forces of the regulatory DNA, rather than the DNA sequence itself. In contrast, nucleosomes downstream of the DNA seI hypersensitive promoter are well-localized, regardless of whether the promoter is accessible or closed. Figure 4C Therefore, in most cases, nucleosome localization appears to be caused by boundary conditions imposed by the regulation of DNA's drive on a single chromatin template.

[0221] Figures 4A-4D The effects of regulating DNA drives on nucleosome localization are shown. Figure 4A A schematic diagram illustrating nucleosome phasing / positioning calculations in overlapping readings is shown, along with bar charts and box plots of individual nucleosome offsets across different reading categories. *p-value < 0.001 (rank-sum test). Figures 4B-4C The DHS at the distal end of the adjacent TSS is shown. Figure 4B) and the promoter DHS of the expressed gene ( Figure 4C The enrichment of homophase and heterophase nucleosomes in chromatin fibers was determined based on whether the chromatin fibers contained accessible MADs (red) and closed MADs (gray) that overlapped with DHS. *p value < 0.01; ns = p value > 0.05 (z test). Figure 4D A schematic diagram of the boundary model demonstrating the arrangement of nucleosomes around the regulatory element is shown.

[0222] Example 5: Conservation of chromatin structure between fruit flies and humans

[0223] We next attempted to determine whether these chromatin features were conserved between Drosophila and humans. This involved validating the ability of m6A-MTase to selectively label cell-type-specific accessible DNA in human cell types. Figure 5A Following this, we treated the nuclei of human K562 cells with m6A-MTase and then performed Pacific Biosystems CCS single-molecule DNA sequencing to create a chromatin template. This made access to regulatory elements and nucleosome-to-nucleosome junction regions ( Figures 5B-5C All exhibit robust methylation, with the number of m6A-labeled bases overlapping with DHS reflecting the aggregation density of DNA seI fragments obtained from somatic cell nuclei (average 0.3 high-quality reads per DHS). High-quality fibrils produce an average of more than 8 well-defined nucleosomes. Figure 5D Consistent with our findings in fruit flies. Figure 2H In K562 cells, the nucleosome arrays NR on both sides of the promoter and distal DHS are shorter, but exhibit significantly greater structural heterogeneity compared to nucleosome arrays farther from the DHS. Figure 5E We also found that the regulation of DNA accessibility in K562 cells is primarily driven by a process that is either entirely or entirely absent. Figure 5F Furthermore, the nucleosomes around the promoter and distal DHS are well localized, suggesting that these may be common features in regulating DNA drives as well as the localization and compression of surrounding nucleosomes.

[0224] In summary, research suggests that it is possible to copy chromatin and regulatory structures onto DNA templates at nucleotide resolution and combine this with long-read single-molecule DNA analysis to map the primary structure of individual chromatin fibers (Fiber-seq). The drivers of regulatory DNA are highly cell-selective, so it should be possible to break down single-molecule data from complex cellular mixtures into regulatory template states from constituent cellular subpopulations. With further increases in read length and throughput, it should be possible, in the near future, to transcribe the primary regulatory structures of large loci and assemble entire chromatin haplotypes by combining Fiber-seq with precise base identification of genetic variants. Sequencing individual chromatin fibers by simultaneously mapping the genetic (primary sequence) and epigenetic (chromatin structure) states of individual regulatory alleles provides a unified tool for directly analyzing the functional impact of rare and common regulatory DNA variations.

[0225] Figures 5A-5E This demonstrates the conservation of chromatin structure between fruit flies and humans. Figures 5A-5B This shows a comparison of DNA seI-seq and MTase-seq signals in the control region of the human β-globin locus in HeLa and K562 cells. Figure 5A ), and a comparison with Fiber-seq data in K562 cells ( Figure 5B ). Figure 5C A bar chart showing the MAD widths of all MADs identified outside the DHS (grey), within the TSS distal DHS (blue), and within the promoter DHS (green). Individual box plots are shown below. *p-value < 0.001 (rank-sum test). Figure 5D The bar chart shows the average NR length per fiber (left) and the average bp difference in NR length per fiber (right) for fibers without DHS (grey), with TSS distal DHS (blue), or with promoter DHS (green). Box-whisker plots are shown below. *p-value < 0.001 (rank-sum test). Figure 5E The proportion of DHS overlapping with accessible and closed fibers is shown for DHS classified according to their proximity to TSS. *p value < 0.01 (z test).

[0226] Example 6: Using cell-penetrating peptide (CPP) labeled M 6A-MT ASE Identifying chromatin structure in vivo

[0227] A modified m6A-MTase was generated, comprising m6A-MTase Hia5 conjugated to a cell-penetrating peptide (CPP) and a nuclear localization sequence (NLS). Specifically, the CPP tag allows the m6A-MTase to penetrate the cell membrane of living cells, while the NLS tag allows the MTase to subsequently dissociate into the nucleus. After treating living cells with this reagent, the isolated DNA was subjected to single-molecule chromatin fiber sequencing and direct base modification assays (i.e., in vivo fiber-seq), enabling us to identify chromatin structure and dynamics while cells are alive. Figure 6A ).

[0228] This method can use multiple CPP tags. CPP tags TAT, 8-arginine (8R), and permein all showed similar efficiency. Figure 6B This method was successfully applied to primary blood cells (…). Figure 6B In vivo fiber-seq mapping reflects results obtained from isolated cell nuclei, offering the added advantage of eliminating the need for nuclear separation steps and quantifying in vivo chromatin dynamics. Figures 6C-6E ).

[0229] Example 7: Using F IBER-SEQ Identifying functional genes that regulate DNA alterations

[0230] By combining the m6A-tagged chromatin structure of each molecule with high-quality DNA sequencing information potentially available for each molecule during single-molecule sequencing, Fiber-seq can be used to easily elucidate changes in functionally regulated DNA. Figure 7 As shown, fiber-seq was performed in primary human CD4+ cells, and functional regulatory DNA variants were identified based on the effects of these DNA variants on the overlapping chromatin structure on fibers.

[0231] Example 8: Visualization of in situ methylated adenine (M6A) sites

[0232] An imaging analysis was developed to visualize methylated adenine (m6A) sites in situ (i.e., in intact mammalian cells).

[0233] K562 cells were washed with 1x PBS and the cell clumps were resuspended in buffer A. The resuspended cells were then infiltrated with 0.1% IGEPA on ice for 5 minutes. Cell samples were clumped and resuspended in buffer A, treated with 0 U, 1 U, or 40 U of Hia5 adenine methyltransferase, and immediately seeded at a density of 1 million cells per milliliter on a PLL-coated glass surface. The cells were incubated at 37°C for 15 minutes, then fixed at room temperature with excess 4% paraformaldehyde solution for 10 minutes. The fixed cells were washed with 2x PBS, infiltrated with 0.25% Triton for 10 minutes, treated with RNase A at 37°C for 30 minutes, blocked with 2% BSA for 1 hour, and then labeled with m6A antibodies (ABE572, ABE572-I, SAB5600251) according to standard immunofluorescence methods. After m6A labeling, the cells were counterstained with DAPI and fixed with Prolong Gold anti-fading agent. Cells were then imaged in 3D using epifluorescence microscopy with a 60x1.4NA oil immersion objective. Cell images were deconvolved to remove out-of-focus areas and then processed to delineate individual cell nuclei and m6A-labeled nuclear regions.

[0234] Anti-m6A antibody labeling was used to show distinct dotted staining, which increased with increasing dose of Hia5 adenine methyltransferase.

[0235] Figure 8 Representative images of K562 cell nuclei stained with DAPI and m6A are shown, revealing punctate m6A patterns of all three tested m6A antibodies, which increased with increasing Hia5 dosage.

[0236] The dose-dependent increase in nuclear m6A signaling is evident both throughout nuclear expression and at individual sites.

[0237] Figure 9 Box plots and violin plots are shown, which demonstrate the dose-dependent increase in total nuclear m6A signal (left) and single-point intensity (right) of Hia5.

[0238] Visualizing m6A-tagged genomic regions will help understand the spatial organization of the accessible genome at the single-cell level, enabling in-depth studies of the structure-function interactions that regulate DNA. This visualization can also be used to compare and contrast diseased and normal cells.

[0239] Although the invention has been described in detail with reference to illustrations and embodiments for clarity, it will be apparent to those skilled in the art that certain changes and modifications can be made to the invention without departing from the spirit or scope of the appended claims. It should also be understood that the terminology used herein is for describing particular embodiments only and is not intended to be limiting, as the scope of the invention will be limited solely by the appended claims.

[0240] Therefore, the foregoing merely illustrates the principles of the invention. It should be understood that those skilled in the art will be able to design various arrangements, which, although not explicitly described or shown herein, embody the principles of the invention and are included within its spirit and scope. Furthermore, all embodiments and conditional descriptions herein are primarily intended to help the reader understand the principles of the invention and the concepts contributed by the inventors to further developments in the field, and should be understood as not being limited to these specifically enumerated embodiments and conditions. Moreover, all statements herein describing the principles, aspects, and embodiments of the invention and their specific examples are intended to include their structural and functional equivalents. Furthermore, such equivalents are intended to include both currently known equivalents and future-developed equivalents, i.e., any element developed to perform the same function, regardless of its structure. Therefore, the scope of the invention is not intended to be limited to the exemplary embodiments shown and described herein. Rather, the scope and spirit of the invention are embodied in the appended claims. sequence list <110> Altes Biomedical Science Institute Brigham and Women's Hospital Company <120> Methods, compositions, and kits for identifying protein-binding regions in genomic DNA. <130> ALTI-730WO <150> US 63 / 004,361 <151> 2020-04-02 <160> 5 <170> PatentIn version 3.5 <210> 1 <211> 281 <212> PRT <213> Haemophilus influenzae <400> 1 Met Ala Asn Gln Asn Thr Phe Lys Gln Ala Pro Leu Pro Phe Ile Gly 1 5 10 15 Gln Lys Arg Met Phe Leu Lys Gln Phe Glu Gln Ile Leu Asn Glu Asn 20 25 30 Ile Ser Asp Asn Gly Glu Gly Trp Thr Ile Leu Asp Thr Phe Gly Gly 35 40 45 Ser Gly Leu Leu Ser His Thr Ala Lys Arg Leu Lys Pro Lys Ala Arg 50 55 60 Val Ile Tyr Asn Asp Phe Asp Gly Tyr Ala Glu Arg Leu Ala His Ile 65 70 75 80 Asp Asp Ile Asn Gln Leu Arg Ala Glu Leu Tyr Ser Val Val Gly Asn 85 90 95 Ala Thr Ser Lys Asn Lys Arg Met Thr Lys Asp Cys Lys Ala Glu Cys 100 105 110 Ile Arg Ile Ile Gln Asn Phe Lys Gly Tyr Lys Asp Leu Asn Cys Leu 115 120 125 Ala Ser Trp Leu Leu Phe Ser Gly Gln Gln Val Ala Thr Leu Asp Asp 130 135 140 Leu Phe Gln His Asn Phe Trp His Cys Ile Arg Gln Ser Asp Tyr Pro 145 150 155 160 Lys Ala Asp Gly Tyr Leu Asp Gly Val Glu Ile Val Lys Glu Ser Phe 165 170 175 His Thr Leu Leu Pro Lys Phe Ser Asn Asp Pro Lys Ala Leu Phe Val 180 185 190 Leu Asp Pro Pro Tyr Leu Cys Thr Lys Gln Glu Ser Tyr Lys Gln Ala 195 200 205 Thr Tyr Phe Asp Leu Ile Asp Phe Leu Arg Leu Val Asn Ile Thr Arg 210 215 220 Pro Pro Tyr Val Phe Phe Ser Ser Thr Lys Ser Glu Phe Ile Arg Phe 225 230 235 240 Val Asn Tyr Met Leu Glu Asp Lys Val Asp Asn Trp Gln Ala Phe Glu 245 250 255 Asn Ala Lys Arg Ile Thr Val Asn Ala Lys Leu Asn Tyr Gln Val Ala 260 265 270 Tyr Glu Asp Asn Leu Val Tyr Lys Phe 275 280 <210> 2 <211> 296 <212> PRT <213> Haemophilus influenzae <400> 2 Met Ser Glu Tyr Leu Glu Tyr Gln Asn Ala Ile Glu Gly Lys Thr Met 1 5 10 15 Ala Asn Lys Lys Thr Phe Lys Gln Ala Pro Leu Pro Phe Ile Gly Gln 20 25 30 Lys Arg Met Phe Leu Lys His Val Glu Ile Val Leu Asn Lys His Ile 35 40 45 Asp Gly Glu Gly Glu Gly Trp Thr Ile Val Asp Val Phe Gly Gly Ser 50 55 60 Gly Leu Leu Ser His Thr Ala Lys Gln Leu Lys Pro Lys Ala Thr Val 65 70 75 80 Ile Tyr Asn Asp Phe Asp Gly Tyr Ala Glu Arg Leu Asn His Ile Asp 85 90 95 Asp Ile Asn Arg Leu Arg Gln Ile Ile Phe Asn Cys Leu His Gly Ile 100 105 110 Ile Pro Lys Asn Gly Arg Leu Ser Lys Glu Ile Lys Glu Glu Ile Ile 115 120 125 Asn Lys Ile Asn Asp Phe Lys Gly Tyr Lys Asp Leu Asn Cys Leu Ala 130 135 140 Ser Trp Leu Leu Phe Ser Gly Gln Gln Val Gly Ser Val Glu Ala Leu 145 150 155 160 Phe Ala Lys Asp Phe Trp Asn Cys Val Arg Gln Ser Asp Tyr Pro Thr 165 170 175 Ala Glu Gly Tyr Leu Asp Gly Ile Glu Val Ile Ser Glu Ser Phe His 180 185 190 Lys Leu Ile Pro Arg Tyr Gln Asn Gln Asp Lys Val Leo Leo Leo 195 200 205 Asp Pro Pro Tyr Leu Cys Thr Arg Gln Glu Ser Tyr Lys Gln Ala Thr 210 215 220 Tyr Phe Asp Leu And Asp Le Leu Arg Le Leu Asn Leu Thr Lys Pro 225 230 235 240 Pro Tyr Ile Phe Phe Ser Ser Thr Lys Ser Glu Phe Ile Arg Tyr Leu 245 250 255 Asn Tyr Met Gln Glu Ser Lys Thr Asp Asn To Arg Ala Phe Glu Asn 260 265 270 Tyr Lys Arg Ile Val Val Lys Ala Ser Ala Ser Lys Asp Gly Ile Tyr 275 280 285 Glu Asp Asn Also Has Tyr Lys Phe 290,295 <210> 3 <211> 317 <212> PRT <213> I don't know how to do it. <400> 3 Met Lys Lys Thr Leu Thr Wings Leu Wings Val Wings Ser Leu Wings Ser Wings 1 5 10 15 Three Gln Three Lys Gln Gln Ala Ser Lys Gln Ala Ser 20 25 30 Lys Gln Ala Ser Lys Glu Cys Glu Met Ala Lys Val Phe Lys Gln Ala 35 40 45 Pro Leu Pro Phe Ile Gly Gln Lys Arg Met Phe Leu Lys His Phe Glu 50 55 60 Gln Val Leu Ala His Ile Pro Asp Asp Gly Asn Gly Trp Thr Ile Val 65 70 75 80 Asp Val Phe Gly Gly Ser Gly Leu Leu Ser His Thr Ala Lys Arg Leu 85 90 95 Lys Pro Lys Ala Arg Val Ile Tyr Asn Asp Tyr Asp Asn Tyr Ser Glu 100 105 110 Arg Leu Gln His Ile Asp Asp Ile Asn Arg Leu Arg Arg Ile Ile Ala 115 120 125 Asp Leu Met Ala Asp Thr Pro Lys Tyr Lys Arg Leu Asp Asn Ala Lys 130 135 140 Lys Leu Gln Ile Ile Glu Ala Ile Glu Ala Phe Gln Gly Tyr Lys Asp 145 150 155 160 Leu His Ile Leu Cys Ser Trp Leu Ala Phe Ser Gly Gln Gln Val Ser 165 170 175 Ser Phe Asp Glu Leu Tyr Lys Gln Asn Phe Trp His Cys Ile Arg Gln 180 185 190 Ser Asp Tyr Leu Thr Ala Asp Gly Tyr Leu Asp Gly Val Glu Ile Val 195 200 205 Arg Glu Ser Phe His Gln Leu Val Pro Arg Phe Thr Gly Gln Pro Asn 210 215 220 Thr Leu Leu Val Leu Asp Pro Pro Tyr Leu Cys Thr His Gln Glu Ser 225 230 235 240 Tyr Lys Gln Glu Arg Tyr Phe Asp Leu Val Asp Phe Leu Arg Leu Ile 245 250 255 His Leu Thr Lys Pro Pro Tyr Val Phe Phe Ser Ser Thr Lys Ser Glu 260 265 270 Phe Val Arg Phe Ile Asp Ala Met Val Glu Asp Lys Trp Asp Asn Trp 275 280 285 Gln Ala Phe Asp Asp Ala Gln Arg Ile Val Val Gln Thr Ser Ala Ser 290 295 300 Tyr Asn Gly Lys Tyr Glu Asp Asn Met Val Tyr Lys Phe 305 310 315 <210> 4 <211> 227 <212> PRT <213> Escherichia coli <400> 4 Met Ser Arg Phe Ile Leu Gly Asp Cys Val Arg Val Met Ala Thr Phe 1 5 10 15 Pro Asp Asn Ala Val Asp Phe Ile Leu Thr Asp Pro Pro Tyr Leu Val 20 25 30 Gly Phe Arg Asp Arg Ser Gly Arg Thr Ile Ala Gly Asp Val Asn Asp 35 40 45 Asp Trp Leu Gln Pro Ala Ser Asn Glu Met Tyr Arg Val Leu Lys Lys 50 55 60 Asp Ala Leu Met Val Ser Phe Tyr Gly Trp Asn Arg Ile Asp Arg Phe 65 70 75 80 Met Ala Ala Trp Lys Arg Ala Gly Phe Ser Val Val Gly His Leu Val 85 90 95 Phe Thr Lys Asn Tyr Thr Ser Lys Ala Ala Tyr Val Gly Tyr Arg His 100 105 110 Glu Cys Ala Tyr Ile Leu Ala Lys Gly Arg Pro Ala Leu Pro Gln Lys 115 120 125 Pro Leu Pro Asp Val Leu Gly Trp Lys Tyr Ser Gly Asn Arg His His 130 135 140 Pro Thr Glu Lys Pro Val Thr Ser Leu Gln Pro Leu Ile Glu Ser Phe 145 150 155 160 Thr His Pro Asn Ala Ile Val Leu Asp Pro Phe Ala Gly Ser Gly Ser 165 170 175 Thr Cys Val Ala Ala Leu Gln Ser Gly Arg Arg Tyr Ile Gly Ile Glu 180 185 190 Leu Leu Glu Gln Tyr His Arg Ala Gly Gln Gln Arg Leu Ala Ala Val 195 200 205 Gln Arg Ala Met Gln Gln Gly Ala Ala Asn Asp Asn Trp Phe Glu Pro 210 215 220 Glu Ala Ala 225 <210> 5 <211> 16 <212> DNA <213> Artificial Sequence <220> <223> Synthetic Sequence <400> 5 agatcggaag agcgtc 16

Claims

1. A method for identifying regions of human genomic DNA bound to a protein at single nucleotide resolution, the method comprising: contacting human genomic DNA with a single enzyme, wherein the single enzyme is N6-adenine methyltransferase Hia5 (m6A-MTase Hia5) having an amino acid sequence as set forth in SEQ ID NO: 1, which methylates adenine residues in regions of the genomic DNA that are not bound to a protein; performing single molecule long read sequencing on the contacted genomic DNA to detect positions in the genomic DNA that lack methylated adenine residues, thereby identifying regions of the genomic DNA bound to a protein at single nucleotide resolution.

2. The method of claim 1, wherein the contacting comprises contacting isolated genomic DNA with the m6A-MTase Hia5.

3. The method of claim 1, wherein the contacting comprises contacting a cell containing the genomic DNA.

4. The method of claim 3, wherein the contacting comprises introducing a nucleic acid encoding the m6A-MTase Hia5 into the cell.

5. The method of claim 4, wherein the m6A-MTase Hia5 is fused to a cell penetrating peptide that makes the m6A-MTase Hia5 permeable to the plasma membrane.

6. The method of any one of claims 1-5, wherein the sequencing is performed on at least a 1 kilobase (kb) long stretch of genomic DNA.

7. The method of any one of claims 1-5, wherein the sequencing is performed on at least a 3 kb long stretch of genomic DNA.

8. The method of any one of claims 1-5, wherein the sequencing comprises translocating the genomic DNA through a nanopore.

9. The method of claim 1, wherein the sequencing comprises attaching one or more nanopore sequencing adapters to one or more ends of the genomic DNA.

10. The method of claim 1, wherein the sequencing comprises detecting a signal indicative of a methylated adenine.

11. The method of claim 10, wherein the signal is an electrical signal.

12. The method of claim 1, wherein the sequencing comprises multiple rounds of resequencing.

13. The method of claim 12, wherein the multiple rounds of resequencing comprise up to 20 rounds of sequencing.

14. The method of claim 1, wherein the sequencing comprises cyclic consensus sequencing.

15. The method of claim 1, wherein the sequencing comprises single molecule real time (SMRT) cyclic consensus sequencing (CCS).

16. The method of claim 1, wherein the genomic DNA is from a cancer cell.

17. The method of claim 1, wherein the genomic DNA is from a normal cell, the method further comprising generating a chromatin accessibility map of the sequenced genomic DNA region, wherein the map indicates chromatin regions that are not bound by the protein and thus are accessible to the m6A-Mtase Hia5 and chromatin regions that are bound by the protein and thus are inaccessible to the m6A-Mtase Hia5.

18. The method of claim 17, the method further comprising generating a chromatin accessibility map of genomic DNA from a test cell of a human subject.

19. The method of claim 18, wherein the subject has a disease or disorder.

20. The method of claim 18, wherein the subject is suspected of having a disease or disorder.

21. The method of claim 19, wherein the disease is cancer.

22. The method of claim 18, the method comprising comparing the chromatin accessibility map of the test cell to the chromatin accessibility map of the normal cell, wherein the test cell and the normal cell are the same cell type; and comparing the genomic DNA sequence of the test cell and the normal cell, wherein a difference in chromatin accessibility map indicates a change in chromatin structure in the test cell, wherein no difference in chromatin accessibility map but a difference in genomic DNA sequence indicates that the sequence difference is not related to a change in chromatin structure, and wherein a difference in both chromatin accessibility map and genomic DNA sequence indicates that the sequence difference is related to a change in chromatin structure.

23. The method of claim 22, the method further comprising generating a database comprising information about chromatin accessibility maps and potential genomic DNA sequences.

24. The method of claim 22, the method further comprising generating a database comprising information about chromatin accessibility maps, potential genomic DNA sequences, and relevance to a disorder or disease.

25. The method of any one of claims 18-24, wherein the normal cell and the test cell are epithelial cells, white blood cells, glial cells, osteoblasts, or chondrocytes.

26. The method of claim 18, wherein the normal cell and the test cell comprise a plurality of cells.

27. The method of claim 26, wherein the plurality of cells comprises at least 10 cells.

28. The method of claim 26, wherein the plurality of cells comprises at least 30 cells.

29. The method of claim 26, wherein the plurality of cells comprises at least 100 cells.

30. The method of claim 26, wherein the plurality of cells comprises at least 300 cells.

31. The method of claim 26, wherein the plurality of cells comprises at least 10,000 cells.

32. The method of claim 18, wherein the chromatin accessibility profile covers at least 10% of the chromatin.

33. The method of claim 18, wherein the chromatin accessibility profile covers at least 30% of the chromatin.

34. The method of claim 18, wherein the chromatin accessibility profile covers at least 50% of the chromatin.

35. The method of claim 18, wherein the chromatin accessibility profile covers at least 80% of the chromatin.

36. The method of claim 18, wherein the chromatin accessibility profile covers at least 10% of the genome of the cell.

37. The method of claim 18, wherein the chromatin accessibility profile covers at least 20% of the genome of the cell.

38. The method of claim 18, wherein the chromatin accessibility profile covers at least 30% of the genome of the cell.

39. The method of claim 18, wherein the chromatin accessibility profile covers at least 50% of the genome of the cell.

40. The method of claim 18, wherein the chromatin accessibility profile covers at least 80% of the genome of the cell.

41. The method of claim 1, wherein the protein comprises a histone.

42. The method of claim 1, wherein the protein comprises a transcriptional regulator.

43. The method of claim 42, wherein the transcriptional regulator is a transcriptional repressor.

44. The method of claim 42, wherein the transcriptional regulator is a transcriptional activator.

45. The method of claim 1, wherein the sequencing comprises applying a computational method to distinguish between methylated and non-methylated adenines.

46. A kit for use in the method of any one of claims 1-45, the kit comprising: a single enzyme or a coding nucleic acid thereof, wherein the single enzyme is N6-adenine methyltransferase Hia5 (m6A-MTase Hia5) having an amino acid sequence as set forth in SEQ ID NO: 1; a sequencing adaptor; and instructions for use.

47. The kit of claim 46, wherein the m6A-MTase Hia5 comprises a cell-penetrating peptide fused to its N- or C-terminus, and wherein the m6A-MTase Hia5 is permeable to the plasma membrane.

Citation Information

Patent Citations

  • Arrayed biomolecules and their use in sequencing

    EP1105529A1

  • Characterization of individual polymer molecules based on monomer-interface interactions

    US5795782A

  • Immobilized nucleic acid complexes for sequence analysis

    US8481264B2

  • Identification of 5-methyl-C in nucleic acid templates

    US9175348B2

  • Reactive surfaces, substrates and methods for producing and using same

    WO2007041394A2