Mapping DNA binding
By combining the tagged composition with transposases, the problems of high background noise and cross-contamination in the prior art are solved, and low-noise, low-cross-contamination multiple mapping of multiple DNA binding sites is achieved at the single-cell level, providing a comprehensive understanding of complex regulatory interactions.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- JOHNS HOPKINS UNIVERSITY
- Filing Date
- 2024-09-25
- Publication Date
- 2026-05-29
AI Technical Summary
Existing technologies suffer from high background noise, cross-contamination, and labor-intensive issues when identifying binding sites for various DNA-binding proteins, especially at the single-cell level where multiple mapping is difficult to achieve.
The method of binding a tagged composition to a transposase involves forming an antibody-barcode-transposase complex with an antibody or antibody fragment, a heterocyclic compound, and a transposase to tagged and fragment the nucleic acid binding site, and then identifying the target nucleic acid binding site by sequencing.
It achieves low-noise, low-cross-contamination multiple mapping, enabling the simultaneous identification of multiple DNA binding sites at the single-cell level, providing comprehensive insights into complex regulatory interactions, and is suitable for low-volume starting materials.
Smart Images

Figure CN122122294A_ABST
Abstract
Description
Cross-references to related applications
[0001] This application claims the benefit of U.S. Provisional Patent Application No. 63 / 540,174, filed September 25, 2023, which is incorporated herein by reference in its entirety. sequence list
[0002] This application includes a sequence list, which has been electronically submitted in XML file format and is hereby incorporated in its entirety by reference. An XML copy created on September 25, 2024, is named “JHU_42334_601_SequenceListing.xml” and has a size of 148,729 bytes. field
[0003] This article provides techniques for identifying the binding sites of DNA-binding proteins, and particularly, but not exclusively, methods, systems, and kits for simultaneously mapping the binding sites of multiple proteins in the same cell. background
[0004] Complex interactions between regulatory proteins and cis-regulatory elements regulate gene transcription. See, for example, Taverna (2007) “How chromatin-binding modules interpret histone modifications: lessons from professional pocket pickers” Nat Struct Mol Biol 14: 1025; and Ruthenburg (2007) “Multivalent engagement of chromatin modifications by linked binding modules” Nat Rev Mol Cell Biol 8: 983, each incorporated herein by reference. Coordination of gene transcription often requires the simultaneous efforts of multiple proteins and different histone modifications, such as interactions between target genes, DNA binding sites, epigenetic modifications, and transcription factors.
[0005] New assays are being continuously developed and improved to address these issues. The scientific community is actively involved in developing and refining new sequencing-based assays to identify and characterize binding sites on chromosomes.
[0006] Conventional ChIP-seq and similar techniques are widely used for the identification and mapping of binding sites for transcription factors, cofactors, enzymes, and histone PTMs [1,2,3]. These methods involve fragmenting chromatin using physical or enzymatic means to produce fragmented chromatin. The fragmented chromatin is isolated using specific antibodies, and a DNA library is generated and sequenced. Subsequent bioinformatics analysis is then performed to characterize the binding sites. Conventional ChIP-seq-based methods use large cell numbers (>1 million cells) and can introduce significant background noise and biological asynchrony. Furthermore, the need for chromatin fragmentation makes applying ChIP-seq at the single-cell level a challenging endeavor. Other methods, such as CUT&RUN
[20] and related assays [21, 22, 23], offer some solutions to the limitations of ChIP-seq. These alternative methods employ antibody-bound micrococcal nucleases (MNases) to selectively cleave the target fragment while leaving the remaining chromatin intact (uncut). This targeted fragmentation strategy significantly reduces background noise and improves the signal-to-noise ratio. Notably, permeabilized cells can be retained after digestion, which minimizes and / or eliminates the need for extensive chromatin fragmentation and provides assays compatible with single-cell assays [22, 23]. However, existing techniques require additional steps involving adaptor ligation for library preparation, sequencing, and analysis.
[0007] This challenge has been mitigated by CUT&Tag [4] and similar assays [12, 13]. These techniques employ antibodies linked to a transposase (e.g., Tn5 or a similar enzyme) that simultaneously cleaves the target DNA and incorporates an adaptor at the ends of the cleaved DNA. This procedure, called “tagmentation,” simplifies library preparation. Following tagmentation, an amplification step produces a library ready for sequencing. CUT&Tag uses a transposase-protein A fusion protein loaded with an adaptor that interacts with an antibody specific to the DNA of interest. See, for example, Kaya-Okur (2019) “CUT&Tag for efficient epigenomic profiling of small samples and single cells” Nature Communications 10: 1930; WO2019060907 (disclosing the use of specific binding agents coupled to transposons, each containing a transposase and a transposon) and Gopalan (2021) “Simultaneous profiling of multiple chromatin proteins in the same cells” Molecular Cell 81: 4736, each incorporated herein by reference. However, dissociation of the transposase-protein A fusion protein from the antibody causes spurious tag fragmentation, which increases background noise. Furthermore, in multiplex techniques using multiple transposase-protein A fusion proteins loaded with multiple adaptors and multiple antibodies to map multiple DNA binding targets, the exchange of transposase-protein A fusion proteins and antibodies between binding couples produces incorrect (e.g., mixed) signals due to incorrect pairing of the adaptor with the antibody the adaptor is intended to identify.
[0008] Regulation of gene transcription involves a simultaneous effort involving multiple proteins and different histone modifications. Interactions between two proteins and / or histone posttranslational modifications and their corresponding binding sites have been investigated using multi-sequence chromatin immunoprecipitation (ChIP) assays [24, 25, 26]. However, these ChIP-seq-based techniques involve multiple rounds (e.g., at least two rounds) of immunoprecipitation using different antibodies; these procedures are labor-intensive and require large amounts of initial material. Furthermore, each round of ChIP introduces considerable background noise. A technique called Split DamID provides an alternative technique for detecting co-binding
[27] . In this method, the protein of interest is fused with different subunits of DNA adenine methyltransferase (DAM). Although SpDamID can detect co-binding of two proteins, it does not provide analysis of histone modifications because it requires the construction of the fusion protein. Therefore, SpDamID is limited to identifying a pair of non-histone marker targets.
[0009] Multi-CUT&Tag [5, 6], a derivative of CUT&Tag, can identify multiple targets in a single sample and experiment. In this method, an antibody is combined with a protein A-Tn5 fusion protein, and the Tn5 component is pre-loaded with a barcoded DNA adaptor. Different antibody-Tn5 complexes are mixed and incubated with cells simultaneously. By analyzing DNA barcoding and captured chromosomal DNA using nucleotide sequencing data, Multi-CUT&Tag can simultaneously decipher multiple target proteins and histone markers. Similar to CUT&Tag, Multi-CUT&Tag can handle very small cell numbers, including single cells, thus providing direct detection of protein and / or histone modification interactions. A recently introduced multiplexing technique called MulTI-Tag [7] has addressed the potential cross-contamination problem that can arise when detecting different targets simultaneously. To circumvent this challenge, MulTI-Tag performs multiple rounds of CUT&Tag sequentially to achieve multiplexing. However, like ChIP-seq and CUT&Tag, MulTI-Tag cannot determine the colocalization of epitopes. Furthermore, the time-intensive nature of sequential experiments limits their multi-functionality and results in labor-intensive programs.
[0010] A significant limitation of CUT&Tag-based methods is the increased background noise and potential cross-contamination that can occur when detecting multiple targets simultaneously. Without being bound by theory, it is hypothesized that the background noise arises from the relatively weak interaction between protein A and the antibody. The protein A-Tn5 complex detaches from the designated target, leading to ambiguous tag fragmentation. Furthermore, protein A does not universally bind to all types of antibodies, limiting the range of available antibodies. Additionally, the introduction of Tn5 and its attachment to the antibody occur hours or days before use, which impairs Tn5 enzymatic activity.
[0011] New technologies are needed, particularly multiple mapping for DNA binding. Overview
[0012] This document provides implementations of techniques for mapping DNA binding sites, for example, to identify binding sites for histone markers, histone modifying enzymes, transcription factors, and cofactors on chromosomes. In some implementations, the technique provides multiple identification of one or more (e.g., 1 to 500) DNA binding sites for one or more targets (e.g., 1 to 500), for example, to identify more than one histone marker, histone variant, histone modifying enzyme, DNA modifying enzyme, chromatin-associated protein, transcription factor, RNA species, and cofactor within the genome (e.g., on one or more chromosomes).
[0013] In some aspects, the subject matter disclosed in this invention provides a method for identifying nucleic acid binding sites of a target, the method comprising (a) contacting a target bound to a nucleic acid binding site with a tagged composition, thereby binding the tagged composition to the target, wherein the tagged composition comprises: (i) an antibody or antibody fragment bound to the target; (ii) a heterocyclic compound linked to the antibody or antibody fragment; (iii) a protein complex; and (iv) two or more nucleic acids each comprising a barcode nucleotide sequence, wherein the two or more nucleic acids are linked to the heterocyclic compound; and (b) contacting the two or more nucleic acids of the tagged composition with a transposase to form an antibody-barcode-transposase complex, wherein the antibody-barcode-transposase complex generates a double-strand break in the nucleic acid containing the nucleic acid binding site to produce a nucleic acid fragment containing the nucleic acid binding site; (c) isolating the nucleic acid fragment; and (d) sequencing the nucleic acid fragment to identify the nucleic acid binding site of the target.
[0014] In some aspects, the protein complex includes avidin, streptavidin, or neutral avidin. In some aspects, the heterocyclic compound includes biotin. In some aspects, the transposase includes Tn5 transposase. In some aspects, each of the two or more nucleic acids further comprises a transposase mosaic sequence that binds to the transposase. In some aspects, the transposase mosaic sequence binds to Tn5 transposase. In some aspects, the target comprises a DNA-binding protein. In some aspects, the DNA-binding protein includes transcription factors, regulatory elements, transcription repressors, transcription activators, polymerases, nucleases, nickases, zinc finger proteins, transcription activator-like effector nucleases (TALENs), glycosylation enzymes, methyltransferases, ligases, restriction endonucleases, replication proteins, helicases, or kinases. In some aspects, the antibody or antibody fragment is not directly linked to the two or more nucleic acids. In some aspects, the protein complex binds to a heterocyclic compound linked to the antibody or antibody fragment, and also to a heterocyclic compound linked to the two or more nucleic acids. In some aspects, the method further includes adding magnesium to a sample containing the target and the tagged composition. In some respects, the two or more nucleic acids each also include an amplification handle. In some respects, the method also includes amplifying nucleic acid fragments to provide a sequencing library. In some respects, the amplification is polymerase chain reaction (PCR) amplification.
[0015] In some aspects, the subject matter disclosed in this invention provides a composition comprising: (a) one or more antibodies or antibody fragments that bind to a target; (b) a heterocyclic compound linked to one or more antibodies or antibody fragments; (c) a protein complex comprising avidin, streptavidin, or neutral avidin; and (d) two or more nucleic acids, each comprising: (i) a barcoded nucleotide sequence; and (ii) a transposase mosaic sequence, wherein the two or more nucleic acids are linked to the heterocyclic compound, and wherein the composition forms a complex in solution. In some aspects, the protein complex comprises streptavidin. In some aspects, the heterocyclic compound comprises biotin. In some aspects, the transposase comprises Tn5 transposase. In some aspects, the antibody or antibody fragment comprises a region that binds to a DNA-binding protein. In some respects, DNA-binding proteins include transcription factors, regulatory elements, transcription repressors, transcription activators, polymerases, nucleases, nickases, zinc finger proteins, transcription activator-like effector nucleases (TALENs), glycosylation enzymes, methyltransferases, ligases, restriction endonucleases, replicative proteins, helicases, or kinases. In other respects, protein complexes bind to heterocyclic compounds.
[0016] In some aspects, the subject matter disclosed in this invention provides a kit comprising: a first container containing the compositions disclosed herein; and a second container containing a transposase. In some aspects, the kit also includes reagents for tag fragmentation.
[0017] In some respects, the kit also includes reagents and materials for isolating DNA and amplifying nucleic acids. In some respects, the kit also includes a cell capture scaffold. In some respects, the cell capture scaffold includes magnetic beads, columns, concanavalin A beads, streptavidin beads, colloidal semiconductor nanocrystals, carbon nanotubes, or microfluidic devices.
[0018] In some aspects, the subject matter of this invention provides a method for identifying two or more target binding sites on a nucleic acid, the method comprising: a) providing two or more barcoded affinity reagents, each comprising: an affinity reagent linked to a pair of adaptors, wherein: the first adaptor comprises a first barcoded nucleotide sequence and a first transposase-binding mosaic sequence, and the second adaptor comprises a second barcoded nucleotide sequence and a second transposase-binding mosaic sequence, wherein the first barcoded nucleotide sequence and the second barcoded nucleotide sequence are identical or different; and wherein each of the two or more barcoded affinity reagents does not contain a transposase, wherein each of the two or more barcoded affinity reagents binds to different targets, and wherein the first barcoded nucleotide sequence and the second barcoded nucleotide sequence of each barcoded affinity reagent are different from the first barcoded nucleotide sequence and the second barcoded nucleotide sequence of other barcoded affinity reagents binding to different targets; b) adding the two or more barcoded affinity reagents to a compound containing each barcoded affinity reagent. The sample contains samples of targets and reagents, wherein each target binds to a nucleic acid at a corresponding target-binding site, wherein each barcoded affinity reagent binds to the corresponding target or binds to a first affinity reagent that binds to the corresponding target, and each affinity reagent binding occurs in the absence of a transposase; c) adding unloaded transposases and transposase activators to the sample, wherein the unloaded transposase binds to a first transposase-binding mosaic sequence and a second transposase-binding mosaic sequence of each barcoded affinity reagent, and wherein the bound transposase fragments the nucleic acid and tags the nucleic acid with the first and second barcoded nucleotide sequences of the corresponding barcoded affinity reagent to provide tagged fragmented nucleic acids, wherein at least two tagged fragmented nucleic acids are provided, corresponding to two or more corresponding barcoded affinity reagents, and each barcoded affinity reagent corresponds to a corresponding target-binding site; d) sequencing the tagged fragmented nucleic acids to provide nucleotide sequences; and e) analyzing the nucleotide sequences to identify the target binding sites on the nucleic acids. In some aspects, the tagged fragmented nucleic acids contain corresponding target binding sites.
[0019] In some aspects, the subject matter of this invention provides a method for identifying one or more target binding sites on nucleic acids, the method comprising: a) providing one or more barcoded affinity reagents, each comprising: an affinity reagent linked to a pair of adaptors, wherein: the first adaptor comprises a first barcoded nucleotide sequence and a first transposase-binding mosaic sequence, and the second adaptor comprises a second barcoded nucleotide sequence and a second transposase-binding mosaic sequence, wherein the first barcoded nucleotide sequence and the second barcoded nucleotide sequence are identical or different; and wherein one or more barcoded affinity reagents do not each contain a transposase, wherein one or more barcoded affinity reagents bind to different targets, and wherein the first barcoded nucleotide sequence and the second barcoded nucleotide sequence of each barcoded affinity reagent are different from the first barcoded nucleotide sequence and the second barcoded nucleotide sequence of other barcoded affinity reagents binding to different targets; b) adding one or more barcoded affinity reagents to a barcoded affinity reagent containing each barcoded affinity reagent. The sample is prepared by: a) barcoding an affinity reagent to a target, wherein each target binds to a nucleic acid at a corresponding target binding site; wherein each barcoded affinity reagent binds to the corresponding target or to a first affinity reagent that binds to the corresponding target; and wherein binding of each affinity reagent occurs in the absence of a transposase; b) adding unloaded transposases and transposase activators to the sample, wherein the unloaded transposases bind to a first transposase-binding mosaic sequence and a second transposase-binding mosaic sequence of each barcoded affinity reagent; wherein the bound transposases fragment the nucleic acid and tag the nucleic acid with a first barcoded nucleotide sequence and a second barcoded nucleotide sequence of the corresponding barcoded affinity reagent to provide tagged fragmented nucleic acids, wherein at least one tagged fragmented nucleic acid is provided, corresponding to the corresponding barcoded affinity reagent, and each barcoded affinity reagent corresponds to a corresponding target binding site; c) sequencing the tagged fragmented nucleic acids to provide nucleotide sequences; and e) analyzing the nucleotide sequences to identify the target binding sites on the nucleic acids. In some aspects, the tagged fragmented nucleic acids contain corresponding target binding sites. In some aspects, two barcoded affinity reagents are provided, and a tagged fragmented nucleic acid contains two target binding sites corresponding to the two barcoded affinity reagents. In other aspects, two barcoded affinity reagents are provided, and each of the two tagged fragmented nucleic acids contains a target binding site corresponding to the barcoded affinity reagent.
[0020] In some cases, the transposases are Tn5, Tn3, Tn7, TnY, Sleeping Beauty, or piggyBac, and the transposase activator is MgCl2. In some cases, the target is a DNA-binding protein, such as histones, histone-modifying enzymes, transcription factors, cofactors, or chromatin-associating proteins. In some cases, the target is a post-translational modification on histones or other chromatin-associating proteins, or a modified DNA base. In some cases, the modified DNA base is mC or 5hmC. In some cases, nucleic acids are part of chromatin, and the method also includes the simultaneous detection of histone markers, histone-modifying enzymes, chromatin-associating proteins, and transcription factors. In some cases, the chromatin-associating protein is CTCF or an adhesion protein.
[0021] In some respects, the affinity reagent includes an antibody. In some respects, the affinity reagent is a target-specific affinity reagent. In some respects, the affinity reagent is a secondary affinity reagent that is specific to the primary target-specific affinity reagent. In some respects, the primary affinity reagent does not contain a barcode. In some respects, the method also includes adding the primary affinity reagent to the sample.
[0022] In some aspects, providing a barcoded affinity reagent comprising an affinity reagent linked to the pair of adaptors includes: linking a first affinity moiety to the affinity reagent, providing a first and a second adaptor each having a second affinity moiety, and specifically binding the first affinity moiety to the second affinity moiety. In some aspects, the first and second affinity moieties are a pair selected from the group consisting of: biotin and avidin, streptavidin, or neutral avidin; a first and a second reactive group reacted to provide covalently linked first and second reactive groups; a DNA-binding protein and a DNA sequence recognized by the DNA-binding protein; a HaloTag and a chloroalkane; a SNAP tag and O(6)-benzylguanine; and single-stranded DNA and its hybrid DNA.
[0023] In some aspects, the first and second adaptors each further include an amplification handle. In some aspects, analyzing nucleotide sequences to identify target binding sites on nucleic acids also includes associating the barcoded nucleotide sequence with an affinity reagent. In some aspects, the method further includes amplifying the tagged fragmented nucleic acid to provide a sequencing library. In some aspects, the amplification is polymerase chain reaction amplification. In some aspects of the method, each barcoded affinity reagent includes a first handle connected to the second handle via a spacer region; the first adaptor hybridizes with the first handle; and the second adaptor hybridizes with the second handle, wherein the first or second handle contains a first affinity portion that binds to a second affinity portion of the affinity reagent, and the first and second adaptors contain different amplification handles.
[0024] In some respects, the sample is cellular, tissue, or cell-free DNA. In other respects, the method also includes permeabilizing cells or tissues.
[0025] In some aspects, the subject matter disclosed in this invention provides a method for identifying more than one binding site on one or more nucleic acids for more than one target, and the method includes: a) providing more than one barcoded affinity reagent, wherein each of the more than one barcoded affinity reagent does not contain a transposase, and wherein each of the more than one barcoded affinity reagent binds to a different target; b) adding the more than one barcoded affinity reagent to a sample; c) adding unloaded transposase and transposase activator to the sample to provide more than one tagged fragmented nucleic acid; d) sequencing the more than one tagged fragmented nucleic acid to provide nucleotide sequences; and e) analyzing the more than one nucleotide sequence to identify more than one binding site on more than one target.
[0026] In some aspects, nucleic acids are part of chromatin, and the method also includes determining a data fingerprint of the combination of two target binding sites, wherein the data fingerprint includes: a) colocalization information of the two target binding sites, or lack of interaction between the two target binding sites; b) distance between the two target binding sites or epitopes; c) nucleotide sequence of the tagged fragmented nucleic acid; in some aspects, the data fingerprint also includes: d) polarity or sequence of modifications; e) cis-regulatory elements; f) proximity to or absence of CpG islands; g) repetitive DNA sequences; and / or h) average DNA methylation level.
[0027] In some respects, nucleic acids are part of chromatin. In other respects, the method also includes the simultaneous identification of more than one histone marker, histone variant, histone marker reader, histone modifying enzyme, DNA modifying enzyme, chromatin-associating proteins, transcription factors, RNA species and / or cofactors within the genome.
[0028] In some aspects of the methods disclosed herein, background IgG sequencing reads are less than 25%, 20%, 15%, or 10% of the total sequencing reads. In some aspects of the methods disclosed herein, affinity reagent-specific signals are generated, and cross-contamination signals between different antibodies are less than 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, and 2%. In some aspects of the methods disclosed herein, the method also includes identifying the co-localization of two epitopes at a single locus in the cell. In some aspects, the co-localization of H3K4me3 and H3K27me3 is identified.
[0029] In some aspects of the methods disclosed herein, the methods also include identifying bivalent domain regions in a sample covered by two histone modifications. In some aspects of the methods disclosed herein, the methods also include identifying the co-localization of two epitopes at the same location on the same chromosome copy, originating from a single chromosome segment, within the same cell.
[0030] In some aspects, barcoding more than one affinity reagent to provide more than one barcoded affinity reagent includes incubating each of the more than one affinity-tagged affinity reagents with a unique barcoded connective in separate reaction vessels to provide more than one separate barcoded affinity reagent. In some aspects, the method also includes assembling more than one separate barcoded affinity reagent to provide a mixture of barcoded affinity reagents. In some aspects, analyzing more than one nucleotide sequence to identify more than one binding site for more than one target also includes associating each barcoded nucleotide sequence in more than one barcoded nucleotide sequence with each of the more than one affinity reagent. In some aspects, more than one target includes 2-500 targets.
[0031] In some aspects of the method disclosed herein, the method further includes isolating cell nuclei from cells, performing flow cytometry or gel beads to sort individual cells or individual nuclei, lysing individual cells or individual nuclei, amplifying single-cell / nucleus libraries, including identifying signals from individual cells, assembling single-cell / nucleus libraries, and sequencing single-cell / nucleus libraries.
[0032] In some aspects of the method disclosed herein, the method also includes adding a drug to a sample, performing steps (a)-(e), and comparing how the drug disrupts characteristics in vitro or in vivo.
[0033] In some aspects, the subject matter of this invention discloses a kit comprising: instructions for providing two or more barcoded affinity reagents, each of the two or more barcoded affinity reagents comprising, a pair of adaptors, wherein: a first adaptor comprises a first barcoded nucleotide sequence and a first transposase-binding mosaic sequence, and a second adaptor comprises a second barcoded nucleotide sequence and a second transposase-binding mosaic sequence; affinity reagents, adaptors, wherein each adaptor comprises a barcoded nucleotide sequence and a transposase-binding mosaic sequence, unloaded transposase, and transposase activator. In some aspects, the affinity reagents, adaptors, unloaded transposase, and transposase activator are in a container. In some aspects, the kit further comprises one or more cell or nuclear permeation buffers and / or one or more wash buffers. In some aspects, the buffers are in a container.
[0034] In some aspects, the subject matter of this invention discloses a kit comprising: two or more barcoded affinity reagents, each comprising: a pair of adaptors, wherein: the first adaptor comprises a first barcoded nucleotide sequence and a first transposase-binding mosaic sequence, and the second adaptor comprises a second barcoded nucleotide sequence and a second transposase-binding mosaic sequence; and unloaded transposase and transposase activator. In some aspects, the two or more barcoded affinity reagents, the unloaded transposase, and the transposase activator are in a container. In some aspects, the kit further comprises one or more cell or nuclear permeation buffers and / or one or more wash buffers. In some aspects, the buffers are in a container.
[0035] In some aspects, the kits disclosed herein also include controls. In some aspects, the controls are recombinant nucleosomes that bind to DNA and / or control affinity reagents. In some aspects, the kits disclosed herein include a set of affinity reagents. In some aspects, the kits disclosed herein include a set of cancer-specific affinity reagents. In some aspects, the kits disclosed herein include a set of affinity reagents specific to epigenomic marker proteins and / or histones. In some aspects, the kits disclosed herein also include reagents and materials for isolating DNA and amplifying nucleic acids.
[0036] In some respects, the kits disclosed herein also include cell capture scaffolds. In some respects, the cell capture scaffolds include magnetic beads, columns, concanavalin A beads, streptavidin beads, colloidal semiconductor nanocrystals, carbon nanotubes, or microfluidic devices.
[0037] In some implementations involving multiple technologies, the method includes assembling more than one separate, differently barcoded affinity reagent (e.g., primary antibody) and incubating a sample containing nucleic acid and more than one DNA binding target (e.g., a sample containing permeable cells, cell nuclei, cell-free chromatin, cell-free DNA, or tissue) with more than one separate, differently barcoded affinity reagent (e.g., primary antibody).
[0038] In some embodiments, the method includes incubating the sample. In some embodiments, the method includes incubating the sample overnight (e.g., 8 to 16 hours (e.g., 8.0, 8.5, 9.0, 9.5, 10.0, 10.5, 11.0, 11.5, 12.0, 12.5, 13.0, 13.5, 14.0, 14.5, 15.0, 15.5, or 16.0 hours)). In some embodiments, the method includes rigorously washing the sample after incubation.
[0039] Implementations of this technology can be used to map DNA binding sites using a small amount of starting material (e.g., a small sample) to map multiple DNA binding targets. In some implementations, this technology can be used to map DNA binding sites in a single cell. In some implementations, this technology can be used to map DNA binding sites in preparations of cell-free DNA or chromatin.
[0040] In some embodiments, the method includes biotinylating the affinity reagent (e.g., attaching about three (e.g., one to five (e.g., one, two, three, four, or five)) biotin molecules to each affinity reagent using N-hydroxysuccinimide biotin to provide a biotinylated affinity reagent). This method is applicable to ligands of all subclasses and species. In some embodiments, each barcoded adaptor oligonucleotide includes biotin, a PCR handle, and a barcoded sequence (e.g., 10-nt to 15-nt (e.g., 4-nt to 25-nt)). (e.g., barcode sequences of 4-nt, 5-nt, 6-nt, 7-nt, 8-nt, 9-nt, 10-nt, 11-nt, 12-nt, 13-nt, 14-nt, 15-nt, 16-nt, 17-nt, 18-nt, 19-nt, 20-nt, 21-nt, 22-nt, 23-nt, 24-nt, or 25-nt), nucleotide spacer regions (e.g., 10-nt to 20-nt) (e.g., 10-nt, 11-nt, 12-nt, 13-nt, 14-nt, 15-nt, 16-nt, 17-nt, 18-nt, 19-nt, or 20-nt spacer regions) and a double-stranded portion encoding the Tn5-binding mosaic sequence. In some embodiments, these features of the barcoded DNA adaptor are arranged from the 5' end to the 3' end of the adaptor, for example, after the 5' biotin, there is a PCR handle, the barcoded sequence, the nucleotide spacer region, and the double-stranded sequence encoding the Tn5-binding mosaic sequence. In some embodiments, barcoded affinity reagents targeting different targets are used. Incubation is performed in separate reaction containers (e.g., tubes) to provide separate barcoded affinity reagents. In some embodiments of the low-multiplexity method, the method includes providing one or more unmodified primary ligands that bind to a specific target, followed by providing a mixture of barcoded affinity reagents as secondary ligands targeting the primary ligand. In some embodiments of the high-multiplexity method, the method includes providing a mixture of barcoded affinity reagents as primary affinity reagents that bind to a specific target. In some embodiments, amplifying tagged fragmented chromosomal DNA to generate one or more sequencing libraries includes using a polymerase chain reaction.
[0041] In some embodiments, the first affinity moiety is biotin, and the second affinity moiety is avidin, streptavidin, or neutral avidin. In some embodiments, the first and second affinity moieties chemically react to form a covalent bond (e.g., via click chemistry or via maleimide or N-hydroxysuccinimide (NHS) chaining chemicals). Thus, in some embodiments, the first and second affinity moieties comprise a click chemical pair. In some embodiments, the first and second affinity moieties comprise glutamine and an amine, N-hydroxysuccinimide ester and a primary amine, maleimide and a thiol, Traut reagent and a primary amine, or other reactive groups known in the art that react to form a covalent bond. In some embodiments, the first and second affinity moieties comprise a DNA-binding protein and a DNA sequence recognized by the DNA-binding protein. In some embodiments, the first and second affinity moieties comprise a HaloTag and a chloroalkane. In some embodiments, the first and second affinity moieties comprise a SNAP tag and O(6)-benzylguanine.
[0042] In some embodiments, barcoding more than one affinity reagent to provide more than one barcoded affinity reagent includes incubating each of the more than one affinity-tagged affinity reagents with a unique barcoded connective in a separate reaction vessel to provide more than one separate barcoded affinity reagent. In some embodiments, the method further includes assembling more than one separate barcoded affinity reagent to provide a mixture of barcoded affinity reagents. In some embodiments, analyzing the more than one nucleotide sequence to identify more than one binding site of more than one target further includes associating each barcoded nucleotide sequence in the more than one barcoded nucleotide sequence with each of the more than one affinity reagent. See, for example... Figure 7A and Figure 7BIn some implementations, more than one target includes 2 to 50 targets (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 targets). In some implementations, more than one target includes 2-500 targets (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, 180, 185, 190, 195, 200, 205, 210, 215, 220, 225, 230, 235, 240, ...). 245, 250, 255, 260, 265, 270, 275, 280, 285, 290, 295, 300, 305, 310, 315, 320, 325, 330, 335, 340, 345, 350, 355, 360, 365, 370, 375, 380, 385, 390, 395, 400, 405, 410, 415, 420, 425, 430, 435, 440, 445, 450, 455, 460, 465, 470, 475, 480, 485, 490, 495 or 500 targets).
[0043] As illustrated in the embodiments and figures, the subject matter disclosed herein provides advantages over techniques in the prior art. For example, the subject matter disclosed herein provides advantages over prior art for identifying and characterizing binding sites on chromosomes, such as: (1) Use of two or more barcoded affinity reagents, each of the two or more barcoded affinity reagents comprising: an affinity reagent linked to a pair of adaptors, wherein: the first adaptor comprises a first barcoded nucleotide sequence and a first transposase-binding mosaic sequence, and the second adaptor comprises a second barcoded nucleotide sequence and a second transposase-binding mosaic sequence, wherein the first barcoded nucleotide sequence and the second barcoded nucleotide sequence are the same or different; and wherein each of the two or more barcoded ligands does not contain a transposase; (2) Affinity reagents, each linked to a pair of intermolecules. For example, the streptavidin-biotin bond between the affinity reagent and the intermolecules significantly reduces barcode dissociation or exchange. The strong binding affinity (dissociation constant) between streptavidin and biotin... Kd Approximately 10 -14 [8] provided stable and specific affinity reagent-adaptor conjugation.
[0044] (3) Controlled tag fragmentation using free transposases (e.g., Tn5, Tn3, Tn7, TnY, Sleeping Beauty, PiggyBac, etc.). Random tag fragmentation is minimized and / or eliminated by providing transposases that are not linked to adaptors (adaptor-free transposases). Affinity reagents are prepared separately without transposases, and the transposase is added after each affinity reagent has found its target. This approach allows the transposases to retain maximum enzymatic activity, thereby maximizing efficient and precise tag fragmentation.
[0045] (4) Broad-spectrum analysis of epigenetic regulators. Low noise and minimal cross-contamination provide a technique for detecting multiple targets in a single experiment using low-volume starting materials (e.g., single cells). This advantage is particularly valuable for preserving precious samples. Furthermore, this technique offers multiple approaches for the definitive identification of co-binding events, thus providing comprehensive insights into complex regulatory interactions. This technique can be used to analyze epigenetic landscapes and regulatory mechanisms.
[0046] During the development of the implementation of this technology, data indicated that it produces a very low background signal and minimizes and / or eliminates signal mixing and ambiguity between adaptor-affinity pairs. Furthermore, benchmarking using the ENCODE database indicated that the implementation of this technology recovered most ENCODE peaks. The implementation of this technology provides for the simultaneous detection of histone markers, histone modifying enzymes, and transcription factors. This technology identifies numerous bivalent binding events in the same cell and provides a technique for examining the formation of the histone codon and correlating histone codon information with the distribution of histone modifying enzymes and transcription factors.
[0047] Some parts of this description describe implementations of the technology based on algorithms and symbolic representations of operations on information. These algorithmic descriptions and representations are commonly used by those skilled in the art of data processing to effectively communicate the substance of their work to others skilled in the art. Although these operations are described functionally, computationally, or logically, they are understood to be implemented by computer programs or equivalent circuits, microcode, etc. Furthermore, it is sometimes convenient to arrange these operations as modules without loss of generality. The described operations and their associated modules can be embodied in software, firmware, hardware, or any combination thereof.
[0048] Certain steps, operations, or processes described herein may be performed or implemented individually or in combination with other means by one or more hardware or software modules. In some embodiments, the software modules are implemented as a computer program product comprising a computer-readable medium containing computer program code that can be executed by a computer processor to perform any or all of the described steps, operations, or processes.
[0049] In some implementations, the system includes virtually provided computers and / or data storage (e.g., as cloud computing resources). In specific implementations, the technology includes using cloud computing to provide a virtual computer system that includes components of a computer as described herein and / or performs the functions of a computer as described herein. Thus, in some implementations, cloud computing provides the infrastructure, applications, and software as described herein via a network and / or via the Internet. In some implementations, computing resources (e.g., data analytics, computation, data storage, applications, file storage, etc.) are provided remotely via a network (e.g., the Internet and / or cellular networks).
[0050] Implementations of the technology may also relate to devices for performing the operations described herein. Such devices may be specifically constructed for the desired purpose (e.g., application-specific integrated circuits or field-programmable gate arrays) and / or may include general-purpose computing devices (e.g., microcontrollers, microprocessors, etc.) selectively activated or reconfigured by a computer program stored in a computer. The device may be configured to perform one or more steps, actions, and / or functions described herein, for example, provided as instructions of a computer program. Such computer programs may be stored in a non-transitory tangible computer-readable storage medium or any type of medium suitable for storing electronic instructions, which may be coupled to a computer system bus. Furthermore, any computing system mentioned in the specification may include a single processor or may be an architecture employing multiple processors to enhance computing power.
[0051] Based on the teachings contained herein, alternative implementation methods will be apparent to those skilled in the art. Brief description of the attached diagram
[0052] This patent or application document contains at least one drawing shown in color. A published copy of this patent or application with a color drawing will be provided by the competent authority upon request and payment of the necessary fees.
[0053] These and other features, aspects, and advantages of this technology will be better understood with reference to the following figures.
[0054] Figure 1A , Figure 1B , Figure 1C , Figure 1D , Figure 1E , Figure 1F , Figure 1G , Figure 1H , Figure 1I and Figure 1J The illustration shows the simultaneous and efficient characterization of multiple chromatin proteins using Hi-PlexCUT&Tag. Figure 1A Hi-Plex CUT&Tag Workflow: 1.) Construct barcoded primary antibodies by incubating biotinylated antibodies, streptavidin, and barcoded adaptors containing Tn5-binding mosaic (orange). 2.) Pool multiple barcoded antibodies together and incubate with immobilized cells. Different targets are simultaneously bound by their respective antibodies. Unbound antibodies are washed away. 3.) Tn5 and MgCl2 activate tag fragmentation, inserting barcodes for different targets into nearby genomic DNA. 4.) Enrich fragment libraries by PCR and sequence using Illumina next-generation sequencing. Figure 1B Genome browser signal trajectories from IgG negative controls in ChIP-seq, CUT&RUN, CUT&Tag, and Hi-Plex CUT&Tag, in RPM (per million reads). Hi-Plex CUT&Tag had the lowest IgG background signal. Figure 1C The scatter plot illustrates the high correlation between repeated reads of H3K4me3, RNAPII, and H3K27me3. Figure 1D Within the same genomic region, genome browser signal trajectories are generated from single-sequence (single-sequence) reads of H3K27me3, RNAPII, and H3K4me3 from Hi-Plex CUT&Tag and MulTI-Tag, ChIP-seq, and general ATAC-seq. Hi-Plex CUT&Tag profiling is similar to most methods but differs from general ATAC-seq, which measures accessibility. Figure 1E The image shows a heatmap of enrichment for two mutually exclusive target pairs: H3K9me3 and H3K9ac, and H3K27me3 and H3K27ac. Note that these exclusive epitopes do not overlap in genomic localization. Figure 1FGenome browser signal trajectories for H3K4me3 and RNAPII from ChIP-seq, and H3K4me3 / RNAPII heterotone sequences from Hi-Plex CUT&Tag (tag fragmentation sequencing reads contain two different barcode sequences at the ends). The green-highlighted peaks indicate that the Hi-Plex CUT&Tag heterotone signals can represent the overlap of two related individual ChIP-seq signals. The orange-highlighted peaks indicate that Hi-Plex CUT&Tag heterotone reads are only recognized when the two corresponding target epitopes overlap. Figure 1G A heatmap showing the enrichment of RNAPII and H3K4me3 is presented. There is some overlap in the genomic localization of these targets. Figure 1H The stacked bar charts of overlapping peaks between the H3K4me3 / H3K27me3 heterologous sequences and individual H3K4me3 and H3K27me3 reads from ChIP-seq are summarized. Hi-Plex CUT&Tag identified more potential bivalent events than ChIP-seq. Figure 1I Genome browser signal trajectories for H3K4me3 (blue) and H3K27me3 (red) from ChIP-seq, and for the H3K4me3 / H3K27me3 heterologous sequences (purple) from Hi-Plex CUT&Tag. The magnified view at the bottom highlights the clear epitope overlap at the promoter detected by Hi-Plex CUT&Tag. Figure 1J Genome browser signal trajectories (top) of H3K4me3 and H3K27me3 from ChIP-seq, and H3K27me3 / H3K4me3 heterologous sequences from Hi-Plex CUT&Tag, and peaks called from SEACR (represented by rectangles at the bottom). Hi-Plex CUT&Tag identified bivalent events not detected by ChIP-seq, highlighted in blue.
[0055] Figure 2A , Figure 2B , Figure 2C , Figure 2D , Figure 2E Figure 2F illustrates the complexity of the Hi-Plex CUT&Tag dataset. Figure 2A A cartoon illustration of the information packaged in measurements using Hi-Plex technology. Hi-Plex technology can measure the co-localization of two targets along the genome, and it can also map to various regulatory elements. The length of sequencing fragments can also be clustered and interpreted by measuring the number of nucleosomes and correlated with the function of the targets involved in the fragment. Figure 2BA heatmap showing the number of peaks invoked for each target pair using SEACR, the target pairs being derived from our 36 epigenomic markers, including histone modifications, RNA polymerase II, epigenetic writers, and transcription factors. Figure 2C The box plot illustrates the distribution of cis-regulatory and repetitive elements between euchromatin and heterochromatin marker pairs. Test markers for significant differences between euchromatin and heterochromatin markers are shown at the top. Figure 2D Selected target pairs (including euchromatin markers, bivalent markers, and heterochromatin markers) are annotated with stacked bar graphs of cis-regulatory elements, repeat elements, and average DNA methylation peaks. Four principal groups are identified using hierarchical clustering. Figure 2E Stacked bar charts of fragment length distribution, each showing 20 target pairs, organized by the highest levels of subnucleosomes, mononucleosomes, and binucleosomes. Targets with the highest subnucleosome levels are typically associated with transcription factor-related markers, while targets with the highest binucleosome levels are typically associated with histone modification-related markers.
[0056] Figure 3A and Figure 3B The figure illustrates the Hi-Plex cut & tag profile analysis of a single cell. Figure 3A A schematic diagram of the single-cell Hi-Plex CUT&Tag (scHi-Plex CUT&Tag) method. Figure 3B The image shows a comparison of the chromatin landscape in enriched regions of H3K27me3 homotone (tag fragmentation sequencing reads with identical barcode sequences at both ends), comparing bulk Hi-Plex CUT&Tag maps with scHi-Plex CUT&Tag maps (both aggregates of all single cells and single cells). Cells were ordered by read coverage within the depicted regions.
[0057] Figure 4 This document summarizes the 37 barcoded antibodies used in the Hi-Plex CUT&Tag. We barcoded a set of 37 antibodies that target 12 common histone markers (orange), 14 histone modifying enzymes (light blue), 8 human TFs (grey), CTCF (dark blue), PolII (pSer2) (yellow), and a rabbit IgG negative control (green).
[0058] Figure 5The diagram illustrates the low-multiplexity CUT&Tag workflow. The workflow for the next-generation low-multiplexity CUT&Tag is as follows: 1. Construct a barcoded secondary antibody (2° antibody) by incubating a biotinylated antibody, streptavidin, and a barcoded adaptor. A Tn5-binding mosaic (orange) is located on the adaptor. Different antibodies are prepared separately and then assembled together. 2. Cells immobilized with concanavalin A-coated beads are first incubated with the primary antibody, then with the barcoded secondary antibody. 3. Tn5 and MgCl2 are introduced to activate tag fragmentation. The barcode is inserted into the nearby genomic DNA. 4. Fragment libraries are enriched by PCR and sequenced using Illumina sequencing.
[0059] Figure 6 The size distribution of Hi-Plex CUT & Tag fragments is shown. Figure 6 A. Gel plots show the ladder pattern of the Hi-PlexCUT&Tag library. Fragments of different sizes were labeled as sub-, mono-, di-, and tri+- (according to the number of nucleosomes occupying the endogenous fragment). Figure 6 B. Analyze the size distribution of all fragments from the Hi-Plex CUT & Tag library. The names of the different sizes are labeled at the top of each peak.
[0060] Figure 7A An embodiment of the modified and barcoded antibody described herein is shown. For example... Figure 7A As shown, the embodiment provides an antibody modified with one or more first affinity moieties (“A”). Figure 7A As shown, the affinity moiety can be attached to the antibody using one or more adapters. Furthermore, the antibody can be barcoded using one or more barcode adapters containing a second affinity moiety (“B”) and different barcode sequences. Binding pairs (e.g., affinity moieties) A and B are as described herein (e.g., covalently linked (e.g., provided by click chemistry (e.g., by click chemistry pair), glutamine and amine, N-hydroxysuccinimide ester and primary amine, maleimide & thiol, Traut reagent and primary amine, and other covalently linked chemistry known in the art); avidin and biotin, neutral avidin and biotin, streptavidin and biotin, DNA-binding protein and DNA sequence recognized by the DNA-binding protein, HaloTag and chloroalkane, or SNAP tag and O(6)-benzylguanine).
[0061] Figure 7BEmbodiments of modified antibody conjugates comprising more than one (e.g., two) adaptors are shown. Two adaptors comprising read 1 and read 2 are conjugated with the same antibody using a single-stranded DNA handle (brown in the figure) comprising a spacer region between the two hybridization regions. Binding pairs (e.g., affinity moieties) A and B are as described herein (e.g., covalently linked (e.g., provided by click chemistry (e.g., by click chemistry pair), glutamine and amine, N-hydroxysuccinimide ester and primary amine, maleimide & thiol, Traut reagent and primary amine, and other covalently linked chemistry known in the art); avidin and biotin, neutral avidin and biotin, streptavidin and biotin, DNA-binding protein and DNA sequence recognized by the DNA-binding protein, HaloTag and chloroalkane, or SNAP tag and O(6)-benzylguanine). In some implementations, the melting temperature (Tm) of each hybridization region in the handle is between 42°C and 49°C (e.g., 42°C, 43°C, 44°C, 45°C, 46°C, 47°C, 48°C, or 49°C).
[0062] It should be understood that the accompanying drawings are not necessarily drawn to scale, and the objects in the drawings are not necessarily drawn to scale relative to each other. The drawings are depictions intended to make clear and understandable various embodiments of the devices, systems, and methods disclosed herein. Where possible, the same reference numerals will be used throughout the drawings to refer to the same or similar parts. Furthermore, it should be understood that the drawings are not intended to limit the scope of this teaching in any way. Detailed description
[0063] Current whole-genome sequencing technologies based on CUT&Tag and multi-CUT&Tag Tn5 transposases locate chromatin-associated factors, such as histone markers, transcription factors, and cofactors, relying on guiding pre-loaded Tn5-protein A fusions to specific chromatin regions of interest, where tag fragmentation can occur. These methods utilize protein A as a linker to attach Tn5 to factor-specific antibodies. However, the pre-loaded transposase can potentially detach and randomly participate in tag fragmentation, resulting in high levels of background noise. This problem becomes more pronounced when multiple pre-assembled antibody-Tn5-pA complexes are used together, due to the unintended mixing of signals from different targets. This barrier severely limits our multiplexing capabilities. Because there is a strong need to detect binding sites for many important histone markers, transcription factors, and cofactors, we have developed a new technique called Hi-Plex CUT&Tag to enable high multiplexing of up to dozens of histone markers, histone modifying enzymes, and transcription factors. In this method, in the absence of pre-loaded Tn5, DNA barcode adaptors with transposase-binding mosaics are directly linked to a given antibody via biotin-streptavidin interactions. This process requires incubation with a pooled barcoded primary antibody. After washing away unbound antibody, unloaded transposases are introduced and activated. These transposases bind to the binding mosaic on the adaptor, initiating the tag fragmentation process. The Hi-Plex CUT&Tag method requires only a small amount of starting material to profile numerous targets, and it can even be scaled up to the single-cell level. Data analysis of Hi-Plex CUT&Tag confirms that this new technique produces very low background signal, and more importantly, very little cross-contamination between dozens of different antibodies used together in the same assay. Using the ENCODE database, we also benchmarked our new method to accurately recover most ENCODE peaks. The ability to simultaneously detect a large number of histone markers, histone modifying enzymes, and transcription factors allows us to identify a large number of bivalent events in the same cell and examine the formation of the histone codon, correlating this information with the distribution of histone modifying enzymes and TFs.
[0064] Eukaryotic DNA wraps around histones to form mononuclear body subunits of chromatin, which can act as a physical barrier to transcription. Approximately 3 x 10⁷ such nucleosomes are distributed throughout the chromatin of every human cell. Whole-genome sequencing studies over the past two decades have shown that dozens of different combinations of post-translational histone modifications (PTMs) can co-occur even on individual nucleosomes, and nucleosomes with different PTM combinations are located at different loci throughout the chromatin. Histone PTMs, deposited by histone-modifying enzymes that read, write, and erase them, act as docking sites for chromatin-associated complexes that regulate gene transcription and influence the functional state of chromatin. These chromatin-associated complexes typically contain both histone modifiers and nucleosome remodelers, which work synergistically to coordinate the proper access and function of proteins along our chromosomal DNA. A fundamental problem in dissecting these combinatorial events is the inability to robustly predict gene expression or resulting phenotypes without a comprehensive understanding of the colocalization of epigenetic modifications and regulators. Because most sequencing jobs only allow analysis of one PTM or epigenetic modifier at a time, little is known about how specific chromatin-associated complexes interact with histone PTMs in combination to promote proper chromatin organization and gene expression.
[0065] While ChIP-seq (chromatin immunoprecipitation followed by sequencing) has historically been the most popular method for global profiling of DNA-binding proteins (e.g., transcription factors (TFs) and cofactors) and histone PTMs, it can only profile one target at a time, suffers from low sensitivity, high cost, low efficiency, and cannot map epitopes at the single-cell level [1,2,3]. The recently developed method CUT&Tag (targeted cleavage and tag fragmentation) utilizes a protein A-fused Tn5 transposase to guide the Tn5 loaded with the adaptor to an antibody already bound to a protein of interest (e.g., TF or PTM) in the cell. Using Mg... 2+ Upon activation of Tn5, the transposase cleaves nearby chromosomal DNA to release small DNA fragments, which are then sequenced using a next-generation sequencing platform. Compared to ChIP-seq, CUT&Tag requires fewer cells, provides better resolution, and can be applied to single-cell analysis [4]. The more recently developed Multi-CUT&Tag allows for the simultaneous profiling of up to three targets in a single experiment via pre-formed complexes containing sequence adaptors loaded with specific DNA barcode sequences of antibody-protein A::Tn5 fusion. Different antibodies are then pooled and incubated with permeabilized cells. Due to multiplexing, three epitopes and their combinations in the same cell can be examined simultaneously [5, 6].
[0066] However, CUT&Tag technology and its derivatives suffer from several significant drawbacks. 1) High background signal is common because the Tn5:protein A complex exhibits relatively weak interactions ( K D = 10 -8 M) can dissociate from antibodies and act as an ATAC reagent. 2) Cross-contamination can be severe in multiplex assays due to the “exchange” between DNA adaptors of different antibodies. These problems greatly limit the ability to perform higher multiplex assays [7]. Finally, due to design principles, only a small fraction of Multi-CUT&Tag and Multi-Tag data can detect epitope colocalization [5, 6, 7].
[0067] To minimize background signals and cross-contamination, improve multiplexing capabilities by 10-fold, and increase the likelihood of detecting epitope co-localization in the same cell, we have invented a new technique called Hi-Plex CUT&Tag, which allows for the simultaneous, paired localization of up to 40 targets across the entire genome using next-generation sequencing.
[0068] To reduce background signal and cross-contamination, we employed different strategies to barcode antibodies (Abs) and modified the tag fragmentation procedure. Using tetrameric streptavidin as a linker, we conjugated biotinylated and barcoded DNA adaptor sequences to biotinylated Abs. A mixture of these individually barcoded Abs was then incubated with permeabilized cells or nuclei at room temperature (RT) for 1 hour. After rigorous washing to remove unbound Abs, Tn5 and MgCl2 were added to the sample and incubated at 37°C. o Incubate at C for 1 hour. Finally, extract genomic DNA, amplify the tagged fragments using PCR, and then perform library preparation and next-generation sequencing. Figure 1A Considering that the biotin-streptavidin interaction is almost irreversible ( K D < 10 -14 [8] Biotinylated DNA adaptor sequences are unlikely to dissociate from Abs to generate nonspecific ATAC-like background signals and / or be exchanged with different adaptor sequences to generate cross-contamination signals. In our modified tag fragmentation procedure, Tn5 and MgCl2 are added together after the removal of unbound Abs. This can further reduce background signals and improve Tn5 activity by avoiding overnight incubation.
[0069] Our data demonstrate that Hi-Plex CUT&Tag represents an advancement in multiplex chromatographic mass spectrometry, offering improved specificity and sensitivity for detecting multiple targets. Its streamlined workflow and precise control over the tag fragmentation process make it a valuable tool for studying chromatin biology and protein interactions in a wide range of biological contexts.
[0070] While embodiments of the technique in which an affinity reagent is tethered to a DNA adaptor using streptavidin-biotin association are described, the technique is not limited to this binding pair or binding mode. The technique includes binding and linkage modes using ionic (e.g., electrostatic) interactions, affinity binding (e.g., protein-protein (e.g., antibody-antigen and analogues); protein-nucleic acid (e.g., nucleic acid and nucleic acid-binding proteins); carbohydrates and lectins; metals and chelating agents), direct (e.g., covalent) conjugation (e.g., click chemistry (e.g., azide-alkyne triazole formation, trans-cyclooctene and tetrazine, Staudinger linkage, azide-cyclooctyn cycloaddition, reverse electron-demanding Diels-Alder reaction, etc.)), and nucleic acid hybridization (e.g., hydrogen bonding), such as those described herein. See, for example, Dugal-Tessier (2021) “Antibody-Oligonucleotide Conjugates: A Twist to Antibody-Drug Conjugates” J. Clin. Med. 10: 838, incorporated herein by reference. See, for example, Dovgan (2019) “Antibody–Oligonucleotide Conjugates as Therapeutic, Imaging, and Detection Agents” Bioconjugate Chemistry 30: 2483, incorporated herein by reference.
[0071] Binding pairs and binding modes can include pairs that interact through covalent and non-covalent interactions, such as, but not limited to, ionic bonds, hydrophobic interactions, hydrogen bonds, van der Waals forces (e.g., London dispersion forces), dipole-dipole interactions, and so on. Binding pairs can include, but are not limited to: receptor / affinity pairs; affinity binding portions of an affinity reagent and a receptor; antibody / antigen pairs; antigen-binding fragments of an antigen and an antibody; antibodies or antibody fragments and haptens; lectins / carbohydrate pairs; enzymes / substrate pairs; biotin / avidin; biotin / streptavidin; digoxigenin / anti-digoxigenin; DNA or RNA aptamer binding pairs; peptide aptamer binding pairs; and so on.
[0072] In some embodiments, covalent linkage is used to attach the DNA adaptor to the affinity reagent. In some embodiments, click chemistry, glutamine and amine, N-hydroxysuccinimide ester and primary amine, maleimide and thiol, Traut reagent and primary amine, and other covalent linkage chemistry known in the art are used to provide the covalent linkage. In some embodiments, binding pairs are used to attach the DNA adaptor to the affinity reagent. In some embodiments, the binding pairs used are avidin and biotin, neutral avidin and biotin, streptavidin and biotin, a DNA-binding protein and a DNA sequence recognized by the DNA-binding protein, HaloTag and chloroalkane or SNAP tag and O(6)-benzylguanine.
[0073] In some embodiments, a single site on the antibody contains one DNA adaptor. In some embodiments, a single site on the antibody contains more than one DNA adaptor (e.g., 2, 3, 4, 5 or more DNA adaptors). In some embodiments, more than one site on the antibody (e.g., 2, 3, 4, 5 or more sites) each contains one or more DNA adaptors (e.g., 1, 2, 3, 4, 5 or more DNA adaptors).
[0074] In some implementations, the antibody is modified at a specific site. In other implementations, the antibody is non-specifically modified.
[0075] In some implementations, the technique includes attaching (e.g., conjugating) more than one adaptor (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10 or more adaptors) to the same antibody using a single-stranded DNA handle that includes a spacer region between at least two hybridization regions. Figure 7B In some implementations, the technique includes using a single-stranded DNA handle containing a spacer region between the two hybridization regions to attach (e.g., conjugate) two adaptors to the same antibody. Figure 7B Combined with A and B as described herein (e.g., covalently linked (e.g., by click chemistry, glutamine and amine, N-hydroxysuccinimide ester and primary amine, maleimide and thiol, Traut reagent and primary amine, and other covalently linked chemistry known in the art); avidin and biotin, neutral avidin and biotin, streptavidin and biotin, DNA-binding protein and DNA sequence recognized by such DNA-binding protein, HaloTag and chloroalkane, or SNAP tag and O(6)-benzylguanine).
[0076] Data collected during the development of the technology presented herein indicates that it offers improvements in multiplex chromatographic mass spectrometry analysis. Specifically, experiments indicate that, compared to existing technologies, this technology provides improved specificity and sensitivity in detecting multiple targets. Furthermore, the technology offers a simplified workflow and precise control over the tag fragmentation process. Therefore, implementations of this technology are valuable for studying chromatin biology and protein interactions in diverse biological contexts.
[0077] In this detailed description of the various embodiments, numerous specific details are set forth for illustrative purposes to provide a thorough understanding of the disclosed embodiments. However, those skilled in the art will understand that these various embodiments can be implemented with or without these specific details. In other instances, structures and apparatus are shown in block diagram form. Furthermore, those skilled in the art will readily understand that the specific order in which the methods are presented and performed is illustrative, and that such order is expected to vary while still remaining within the spirit and scope of the various embodiments disclosed herein.
[0078] All references and similar materials cited in this application, including but not limited to patents, patent applications, articles, books, papers, and internet web pages, are expressly incorporated herein by reference in their entirety. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the various embodiments described herein pertain. Where the definition of a term in an incorporated reference differs from the definition provided in this teaching, the definition provided in this teaching shall prevail. Section headings used herein are for organizational purposes only and should not be construed as limiting the subject matter in any way.
[0079] definition To facilitate understanding of the technology of this invention, some terms and wording are defined below. Further definitions are set forth throughout the detailed description.
[0080] Throughout the specification and claims, unless the context clearly indicates otherwise, the following terms shall have the meaning explicitly associated herein. The phrase “in one embodiment” as used herein does not necessarily refer to the same embodiment, although it may refer to the same embodiment. Furthermore, the phrase “in another embodiment” as used herein does not necessarily refer to different embodiments, although it may refer to different embodiments. Therefore, as described below, various embodiments of the invention can be readily combined without departing from the scope or spirit of the invention.
[0081] Furthermore, as used herein, the term “or” is an inclusive “or” operator and is equivalent to the term “and / or” unless the context clearly indicates otherwise. The term “based on” is not exclusive and allows for basing on additional factors not described unless the context clearly indicates otherwise. Additionally, throughout the specification, the meanings of “a,” “an,” and “the” include plural indicators. The meaning of “in…” includes both “in…” and “on…”.
[0082] As used herein, the terms “about,” “approximately,” “substantially,” and “significantly” are understood by those skilled in the art and will vary to some extent depending on the context in which they are used. Where these terms are used, it is not clear to those skilled in the art, given the context in which they are used, that “about” and “approximately” mean plus or minus 10% of a particular term, and “substantially” and “significantly” mean plus or minus 10% of a particular term.
[0083] As used herein, the disclosure of a range includes the disclosure of all values within the entire range and the disclosure of further subdivisions of the range, including endpoints and subranges given for the range. As used herein, the disclosure of a numerical range includes the endpoints and each intermediate number therebetween having the same degree of precision. For example, for the range 6–9, the numbers 7 and 8 are considered in addition to 6 and 9, and for the range 6.0–7.0, the numbers 6.0, 6.1, 6.2, 6.3, 6.4, 6.5, 6.6, 6.7, 6.8, 6.9, and 7.0 are explicitly considered.
[0084] As used herein, the suffix "-free" indicates that the implementation of the technology does not include features of the base word root with the "-free" appended. That is, as used herein, the term "X-free" means "without X," where X is a feature of the technology that is not included in the "X-free" technology. For example, a "calcium-free" composition does not contain calcium, and a "mixing-free" method does not include a mixing step, etc.
[0085] Although the terms “first,” “second,” “third,” etc., may be used herein to describe various steps, elements, compositions, parts, regions, layers, and / or portions, these steps, elements, compositions, parts, regions, layers, and / or portions should not be limited by these terms unless otherwise indicated. These terms are used to distinguish one step, element, composition, part, region, layer, and / or portion from another. Unless the context clearly indicates otherwise, terms such as “first,” “second,” and other numerical terms used herein do not imply order or sequence. Therefore, without departing from the art, a first step, element, composition, part, region, layer, or portion discussed herein may be referred to as a second step, element, composition, part, region, layer, or portion.
[0086] As used herein, the terms “presence” or “absence” (or, alternatively, “present” or “absent”) are used in a relative sense to describe the quantity or level of a particular entity (e.g., a component, action, element). For example, when an entity is said to “presence”, it means that the level or quantity of that entity is above a predetermined threshold; conversely, when an entity is said to “absent”, it means that the level or quantity of that entity is below a predetermined threshold. The predetermined threshold can be a threshold for detectability associated with a specific test used to detect the entity or any other threshold. When an entity is “detected”, it is “presence”; when an entity is “not detected”, it is “absence”.
[0087] As used herein, “increase” or “decrease” refers to a detectable (e.g., measurable) positive or negative change in the value of a variable relative to a previously measured value of that variable, relative to a pre-established value, and / or relative to a standard control. An increase relative to a previously measured value, a pre-established value, and / or a standard control is preferably a positive change of at least 10%, more preferably 50%, still more preferably 2 times, even more preferably at least 5 times, and most preferably at least 10 times. Similarly, a decrease is preferably a negative change of at least 10%, more preferably 50%, still more preferably at least 80%, and most preferably at least 90% of a previously measured value, a pre-established value, and / or a standard control. Other terms indicating quantitative change or difference, such as “more” or “less”, are used herein in the same manner as described above.
[0088] As used herein, the term "binding site" refers to a portion of nucleic acid to which the nucleic acid binds (e.g., chromatin binds) to or will bind to the target, provided sufficient conditions for binding are present. Binding sites can be single-stranded or double-stranded. Binding sites can include two or more portions of the nucleic acid to which the target binds, for example, in the case of several nucleic acids binding to the target in the formation of a dimer or higher-order complex. Binding sites can include both the portion of the nucleic acid to which the target directly binds and the nucleic acid portions located upstream and / or downstream of the target. In some implementations, the binding site includes up to about 1000 bp (e.g., 100 bp, 150 bp, 200 bp, 250 bp, 300 bp, 350 bp, 400 bp, 450 bp, 500 bp, 550 bp, 600 bp, 650 bp, 700 bp, 750 bp, 800 bp, 850 bp, 900 bp, 950 bp, or 1000 bp) on the upstream and / or downstream side of the nucleic acid portion that directly interacts with the target.
[0089] As used herein, "system" refers to more than one real and / or abstract component that operates together for a common purpose. In some implementations, "system" is an integrated collection of hardware and / or software components. In some implementations, each component of the system interacts with and / or is associated with one or more other components. In some implementations, system refers to a combination of components and software used to control and direct methods. For example, a "system" or "subsystem" may include one or more or any combination of the following: mechanical devices, hardware, hardware components, circuits, circuit systems, logic designs, logic components, software, software modules, components of software or software modules, software processes, software instructions, software routines, software objects, software functions, software classes, software programs, files containing software, etc., to perform the functions of the system or subsystem. Therefore, the methods and apparatus of the embodiments, or certain aspects or parts thereof, may take the form of program code (e.g., instructions) embodied in a tangible medium, such as a floppy disk, CD-ROM, hard disk, flash memory, or any other machine-readable storage medium, wherein when the program code is loaded into a machine (such as a computer) and executed by the machine, the machine becomes an apparatus for practicing the embodiments. In the case of executing program code on a programmable computer, the computing device typically includes a processor, a processor-readable storage medium (e.g., volatile and non-volatile memory and / or storage elements), at least one input device, and at least one output device. One or more programs may be implemented or utilize the processes described in conjunction with the embodiments, for example, by using an application programming interface (API), reusable controls, etc. Such programs are preferably implemented in a high-level program or object-oriented programming language to communicate with the computer system. However, if desired, the program may be implemented in assembly or machine language. In any case, the language may be a compiled or interpreted language and is combined with the hardware implementation.
[0090] describe This article provides techniques for identifying the binding sites of DNA-binding proteins, and particularly, but not exclusively, methods, systems, and kits for simultaneously mapping binding sites of multiple proteins in the same cell using affinity reagent-specific barcodes.
[0091] As used herein, “affinity reagent” refers to any molecule that specifically binds to another molecule, sometimes referred to herein as the “target.” For example, an affinity reagent can be an antibody, antibody fragment, nanobody, aptamer, small molecule, synthetic antigen-binding agent, oligonucleotide, DARPin, peptide polymer, tetramer, protein scaffold, or other similar ligand or molecule that binds to the target. In some embodiments, the affinity reagent may comprise an antibody or a fragment thereof (e.g., a monoclonal antibody). Antibodies or fragments thereof may include Fab, Fab', F(ab')2, Fv, scFv, dsFv, bispecific antibodies, trispecific antibodies, tetraspecific antibodies, multispecific antibodies formed from antibody fragments, single-domain antibodies (sdAb), single chains containing complementary scFv (tandem scFv) or bispecific tandem scFv, Fv constructs, disulfide-linked Fv, bivariate domain immunoglobulin (DVD-Ig) binding proteins or nanobodies, aptamers, affiliates, affilin, affitin, affimer, alphabet, anticalin, avimer, DARPin, Fynomer, Kunitz domain peptides, monospecific antibodies or any combination thereof. As used herein, “antibody” is a monoclonal antibody, synthetic antibody, recombinant antibody, chimeric antibody, humanized antibody, human antibody, CDR-grafted antibody, multispecific binding construct that binds to two or more targets, bispecific antibody, bispecific or multispecific antibody, or affinity-mature antibody, single antibody chain or scFv fragment, diabody, single chain containing complementary scFv (tandem scFv) or bispecific tandem scFv, Fv construct, disulfide-linked Fv, Fab construct, Fab' construct, F(ab')2 construct, Fc construct, monovalent or bivalent construct from which domains not essential to the function of the monoclonal antibody have been removed, single-chain molecule containing one VL, one VH antigen-binding domain and one or two constant “effect” domains (optionally linked by linker domains), monovalent antibody lacking a hinge region, single-domain antibody, dual variable domain immunoglobulin (DVD-Ig) binding protein, or nanobody. The term "tag" also refers to antibody mimics, such as affinities, which are engineered affinity proteins, typically small (about 6.5-kDa) single-domain proteins that can be isolated based on their high affinity and specificity for any given protein target. In some embodiments, the affinity reagent is a single-domain antibody. In some embodiments, the affinity reagent is an antibody against protein A, such as the antibody used with CUT&Tag. See Kaya-Okur (2020) Nat Protoc. 15:3264, which is incorporated herein by reference.
[0092] In some embodiments, the affinity reagent binds to a target (e.g., a biomolecule). In some embodiments, the target includes, but is not limited to, peptides, proteins, antibodies or antibody fragments, affinities, ribonucleic acid or deoxyribonucleic acid sequences, aptamers, lipids, polysaccharides, lectins, or chimeric molecules formed from multiple identical or different parts. In some embodiments, the target is a protein. In some embodiments, the affinity reagent is not an antibody against protein A.
[0093] As used herein, "target" refers to a protein associated with DNA or chromatin. In some embodiments, the target is a protein found on or associated with chromatin that is visible in the sample. Chromatin contains the cell's DNA and associated proteins. In eukaryotic chromatin, histones and DNA are found in roughly equal amounts, and non-histone proteins are also present. The basic unit of chromatin organization is the nucleosome, a structure that repeats itself along with DNA and histones throughout the organism's genetic material. Histones are highly conserved basic proteins, and their positive charge promotes their binding to the negatively charged phosphate backbone of DNA.
[0094] In some implementations, targets include ALC1, androgen receptor, Bmi-1, BRD4, Brg1, coREST, c-Jun, c-Myc, CTCF, EED, EZH2, Fos, histone H1, histone H3, histone H4, heterochromatin-1γ, heterochromatin-1, HMGN2 / HMG-17, HP1α, HP1γ, hTERT, Jun, KLF4, K-Ras, Max, MeCP2, MLL / HRX, NPAT, p300, Nanog, NFAT-1, Oct4, P53, Pol II (8WG16), RNA Pol II Ser2P, RNA Pol II Ser5P, RNA Pol II Ser2+5P, and RNA Pol II. Ser7P, Rb, RNA polymerase II, SMCI, Sox2, STAT1, STAT2, STAT3, Suz12, Tip60, UTF1, H1S27ph, H1K25me1, H1K25me2, H1K25me3, H1K26me, H2(A)K4ac, H2(A)K5ac, H2(A)K7ac, H2(A)S1ph, H2(A)T119ph, H2(A)S122ph , H2(A)S129ph, H2(A)S139ph, H2(A)K119ub, H2(A)K126su, H2(A)K9bi, H2(A)K13bi, H2(B)K5ac, H2(B) K11ac, H2(B)K12ac, H2(B)K15ac, H2(B)K16ac, H2(B)K20ac, H2(B)S10ph, H2(B)S14ph, H2(B)33ph, H2( B)K120ub, H2(B)K123ub, H3K4ac, H3K9ac, H3K14ac, H3K18ac, H3K23ac, H3K27ac, H3K56ac, H3K4mel, H 3K4me2, H3K4me3, H3R8me, H3K9mel, H3K9me2, H3K9me3, H3R17me, H3K27mel, H3K27me2, H3K27me3, H3K3 6me, H3K79mel, H3K79me2, H3K79me3, H3K122ac, H3T3ph, H3S10ph, H3T1ph, H3S28ph, H3K4bi, H3K9bi, H3K18bi, H4K5ac, H4K8ac, H4K12ac, H4K16ac, H4K91ac, H4R3me, H4K20me, H4K59me, H4S1ph, H4K12bi and H4 n-terminal tail ubiquitination.In some embodiments, the affinity reagent binds to compounds including monomethylated (me1), dimethylated (me2), trimethylated (me3), phosphorylated (ph), ubiquitinated (ub), sumoylated (su), biotinylated (bi), acetylated (ac), and ADP-ribosylated. O - Epitopes of glycosylated, citrullinated, butyrylated, succinylated, or crotonylated histone residues.
[0095] In some embodiments, targets include transcription factors, regulatory elements, transcription repressors, transcription activators, polymerases, nucleases, nickases, zinc finger proteins, transcription activator-like effector nucleases (TALENs), glycosylation enzymes, methyltransferases, ligases, restriction endonucleases, replicating proteins, helicases, or kinases. In some aspects, targets are DNA-binding proteins, such as histones, histone-modifying enzymes, transcription factors, cofactors, or chromatin-associating proteins. In some aspects, targets are post-translational modifications on histones or other chromatin-associating proteins, or modified DNA bases. In some aspects, the modified DNA bases are mC or 5hmC.
[0096] In some embodiments, targets include histones, such as H1, H2A, H2B, H3, H4, and H5. See Annunziato (2008) DNA Packaging: Nucleosomes and Chromatin. Nature Education 1(1):26, which is incorporated herein by reference. Targets may also be post-translational modified histones, such as histones containing phosphorylated serine or threonine, histones containing methylated lysine or arginine, histones containing acetylated and / or deacetylated lysine, histones containing ubiquitinated lysine, and histones containing ubiquitinated-like lysine. In some embodiments, the target is an RNA polymerase. In some implementations, the targets are H2AK5ac, H2AK9ac, H2BK120ac, H2BK12ac, H2BK15ac, H2BK20ac, H2BK5ac, H2Bub, H3, H3ac, H3K14ac, H3K18ac, H3K23ac, H3K23me2, H3K27mel, H3K27me2, H3K36ac, H3K36mel, H3K36me2, H3K4ac, H3K56ac, H3K79mel, H3K79me3, H3K9acS10ph, H3K9me2, H3S10ph, and H3T1. lph, H4, H4ac, H4K12ac, H4K16ac, H4K5ac, H4K8ac, H4K91ac, H3F3A, H3K27me3, H3K36me3, H3K4me l, H3K79me2, H3K9mel, H3K9me2, H3K9me3, H4K20mel, H2AFZ, H3K27ac, H3K4me2, H3K4me3 or H3K9ac.
[0097] In some implementations, the target is a transcription factor (TF), a TF cofactor, or a suspected transcription factor. A list of known and presumed human transcription factors is provided by Lambert (2018) The Human Transcription Factors. Cell. 172: 650, which is incorporated herein by reference. A list of human TFs is provided in Table 1 by International Patent Application Publication No. WO2023081863. A list of exemplary human targets is provided in Table 2 by International Patent Application Publication No. WO2023081863. A list of exemplary mouse targets is provided in Table 3 by International Patent Application Publication No. WO2023081863. An exemplary Drosophila melanogaster (… Drosophila melanogaster The list of targets is provided in Table 4 by International Patent Application Publication No. WO2023081863. During the development of the embodiments of the technology described herein, experiments were conducted to determine the targets listed in Table 1 below.
[0098] In some embodiments, the target is specifically bound by a first affinity reagent (e.g., a primary antibody), and a second affinity reagent (e.g., a secondary antibody) specifically binds to the first affinity reagent; thus, in some embodiments, the second affinity reagent indirectly binds to the target. Therefore, in some embodiments, the affinity reagent is a secondary antibody specific to the primary antibody species and isotype. For example, in some embodiments, the affinity reagent is anti-IgA, anti-IgD, anti-IgE, anti-IgG, or anti-IgM. Furthermore, in some embodiments that include the use of a secondary antibody, the secondary antibody is generated from a primary antibody against any species, including human, mouse, rat, rabbit, etc. The affinity reagent can be independently selected from any type of antibody and / or affinity reagent as described herein and known in the art.
[0099] The implementation scheme includes the use of transposases. In some implementation schemes, transposases can be used for tag fragmentation. A “transposase” is an enzyme that binds to the end of a transposon and catalyzes the movement of the transposon to another part of the genome through a cleavage and paste mechanism or a replication-type transposition mechanism. Exemplary transposases include Tn5 transposase, Tn3 transposase, Tn7 transposase, TnY transposase, Sleeping Beauty, piggyBac, super-active Tn5 transposase, Mu transposase, IS5 transposase, IS91 transposase, Tn552 transposase, Ty1 transposase, Tn / O transposase, IS10 transposase, Mariner transposase, Tel transposase, P-element transposase, Tn3 transposase, bacterial insertion sequence transposase, retroviral transposase, yeast retrotransposon transposase, ISS transposase, Tn1O transposase, Tn903 transposase, or combinations thereof.
[0100] As used herein, the term "transposon" refers to a nucleic acid molecule capable of being incorporated into nucleic acids by a transposase. A transposon comprises two transposon ends (also called "arms," "mosaic ends," or "ME"). In some embodiments, the two transposon ends are located on the flanks of a sequence of sufficient length to form a loop in the presence of a transposase. A transposon can be double-stranded, single-stranded, or contain both single-stranded and double-stranded regions, depending on the transposase. For the Tn5 transposase, the transposon ends are double-stranded, and the linker sequence is single-stranded or double-stranded. The term "mosaic" or "binding mosaic" refers to the sequence region that interacts with the transposase.
[0101] In some implementations, transposases are enzymes belonging to the RNase protein superfamily, including retroviral integrase. Examples of transposases include Tn3, Tn5, and their hyperactive mutants. Tn5 is found in Shewanella (…). Shewanella ) and Escherichia coli bacteria ( Escherichia bacteriaExamples of the hyperactive mutant Tn5 include mutations in E54K and / or L372P. In some embodiments, the transposase is Tn5. In some embodiments, the transposase is TnY, which is a transposase derived from Vibrio parahaemolyticus. Vibrio parahaemolyticus A highly active transposase mutant containing P50K and M53Q mutations. The internal and external ends of the transposon contain the same sequences as the internal and external ends of the Tn5 transposon (see International Patent Application Publication No. WO2021011433, which is incorporated herein by reference). Other transposases that can be used in embodiments of this technology are luminescent bacilli ( P. luminescens Legionella pneumophila ( L. pneumophila Legionella aureus (Long Beach) L. longbeachae ), Bacillus agglutinationis ( C. glomeribacter ) and Vibrio parahaemolyticus transposases, as well as Tn5 HA and sarSeaEAK transposases known in the art.
[0102] The nucleotide sequence encoding the Tn5 transposase is provided by (SEQ ID NO:1): The amino acid sequence of the Tn5 transposase is provided by (SEQ ID NO:2): MITSALHRAADWAKSVFSSAALGDPRRTARLVNVAAQLAKYSGKSITISSEGSKAMQEGAYRFIRNPNVSAEAIRKAGAMQTVKLAQEFPELLAIEDTTSLSYRHQVAEELGKLGSIQD KSRGWWVHSVLLLEATTFRTVGLLHQEWWMRPDDPADADEKESGKWLAAAATSRLRMGSMMSNVIAVCDREADIHAYLQDKLAHNERFVVRSKHPRKDVESGLYLYDHLKNQPELGGYQ ISIPQKGVVDKRGKRKNRPARKASLSLRSGRITLKQGNITLNAVLAEEINPKGETPLKWLLLTSEPVESLAQALRVIDIYTHRWRIEEFHKAWKTGAGAERQRMEEPDNLERMVSILS FVAVRLLQLRESFTPPQALRAQGLLKEAEHVESQSAETVLTPDECQLLGYLDKGKRKRKEKAGSLQWAYMAIARLGGFMDSKRTGIASWGALWEGWEALQSKLDGFLAAKDLMAQGIKI The nucleotide sequence encoding the TnY transposase is provided by (SEQ ID NO:3): The amino acid sequence of TnY transposase is (SEQ ID NO:4): MTHSDAKLWAQEQFGQAQLKDPRRTQRLISLATSIANQPGVSVAKLPFSKADQEGAYRFIRNDNIDAKDIAEAGFQSTVSRANEHKELLALEDTTTLSFPHRSIKEELGHTNQG DRTRALHVHSTLLFAPQNQTIVGLIEQQRWSRDITKRGQKHQHATRPYKEKESYKWEQASRRVVERLGDKMLDVISVCDREADLFEYLTYKRQHQQRFVVRSMQSRCLEEHAQKL YDYAQALPSVKTKALTIPQKGGRKARDVKLDVKYGQVTLKAPANKKEHAGIPVYYVGCLEQGTSKDKLAWHLLTSEPINNVEDAMRIIGYYERRWLIEDFHKVWKSEGTDVESLR LQSKDNLERLSVIYAFVATRLLALRFIKEVDELTKESCEKVLGQKAWKLLWLKLESKTLPKEVPDMGWAYKNLAKLGGWKDTKRTGRASIKVLWEGWFKLQTILEGYELAMSLDH In some embodiments, the technique includes using an adaptor containing a transposase-binding sequence known in the art as a “mosaic” or “binding mosaic.” The mosaic sequence is known in the art, for example, for use with Tn5 transposases. An exemplary mosaic sequence for use with Tn5 transposases has the top strand: AGATGTGTATAAGAGACAG (SEQ ID NO:5). In some embodiments, the mosaic sequence is provided at the 5' end, the 3' end, or both of the 5' and 3' ends of the adaptor. See, for example, Picelli (2014) Genome Research 24: 2033, which is incorporated herein by reference.
[0103] In some embodiments, the adaptor includes an amplification handle or primer binding site. In some embodiments, the adaptor includes a sequencing priming region, such as, for example, a P5 or P7 sequence for Illumina sequencing. In some embodiments, the adaptor includes a specific priming sequence, such as an mRNA-specific priming sequence (e.g., a poly-T sequence for initiating reverse transcription of RNA), a targeted priming sequence, and / or a random priming sequence. In some embodiments, the adaptor includes a promoter for T7 RNA polymerase, for example, to provide in vitro transcription during sample processing.
[0104] In some embodiments, the adaptor also includes a barcode sequence (“target barcode”) identifying the target of the affinity reagent. The target barcode sequence can be used to identify the affinity reagent and / or the target. The target barcode sequence is a unique sequence that allows identification of the specific affinity reagent being tested or employed. Embodiments provide target barcodes of any length available using polynucleotide synthesis techniques, and the length of the barcode limits the number of formulations that can be tested simultaneously. For example, a 10-bp barcode provides a total of 1,048,576 different and unique barcode sequences. Therefore, in some embodiments, the length of the barcode sequence is between 4 nt and 100 nt, for example, between 10 nt and 20 nt, for example, 10 nt. In some embodiments, the barcode sequence length is 4 nt, 5 nt, 6 nt, 7 nt, 8 nt, 9 nt, 10 nt, 11 nt, 12 nt, 13 nt, 14 nt, 15 nt, 16 nt, 17 nt, 18 nt, 19 nt, or 20 nt. In some embodiments, the affinity reagent (e.g., antibody) is modified (e.g., linked) with one or more adapters. See example Figure 7A and Figure 7B .
[0105] For example, such as Figure 7A As shown, the embodiment provides an antibody modified with one or more first affinity moieties (“A”). Figure 7A As shown, the affinity moiety can be attached to the antibody using one or more adapters. Furthermore, the antibody can be barcoded using one or more barcode adapters comprising a second affinity moiety (“B”) and a different barcode sequence. Binding pairs (e.g., affinity moieties) A and B are as described herein (e.g., covalently linked (e.g., provided by click chemistry, glutamine and amine, N-hydroxysuccinimide ester and primary amine, maleimide & thiol, Traut reagent and primary amine, and other covalently linked chemistry known in the art); avidin and biotin, neutral avidin and biotin, streptavidin and biotin, DNA-binding protein and DNA sequence recognized by the DNA-binding protein, HaloTag and chloroalkane, or SNAP tag and O(6)-benzylguanine).
[0106] For example, such as Figure 7BAs shown, the embodiments provide two adapters comprising read 1 and read 2, which are conjugated to the same antibody using a single-stranded DNA handle comprising a spacer region (brown in the figure) between two hybridization regions. The binding pairs A and B are as described herein (e.g., covalently linked (e.g., by click chemistry, glutamine and amine, N-hydroxysuccinimide ester and primary amine, maleimide & thiol, Traut reagent and primary amine, and other covalently linked chemistry known in the art); avidin and biotin, neutral avidin and biotin, streptavidin and biotin, DNA-binding protein and the DNA sequence recognized by that DNA-binding protein, HaloTag and chloroalkane, or SNAP tag and O(6)-benzylguanine). In some embodiments, the handle has more than 22 base pairs.
[0107] This technology can be used in research, medicine, and other fields. For example, NextGen CUT&Tag technology provides multiplexed characterization of epiproteomic epitopes at the single-cell level. Therefore, implementations of this technology allow for the examination of dozens of chromatin-associated biological events, mechanisms, or markers occurring on a single-cell basis. These events can occur at one or more sites within the genome of a single cell and may differ from similar loci in the genomes of other cells in the same culture, tissue, or preparation. Regarding chromatin-associated events related to DNA damage, DNA damage is uniquely programmed in single cells through many biological pathways, such as VDJ recombination, selection of replication origins during DNA replication, hotspots and valid or invalid recombination events during meiosis, and DNA breaks observed in differentiated neurons. Currently, it is difficult to validate such DNA damage at a single-cell site involving more than a few (e.g., 1, 2, 3) epiproteomic epitopes. The lack of technologies in the art that provide epiproteomic resolution means that the biology associated with programmed DNA damage, and the molecular mechanisms that initiate, arise from, and resolve programmed DNA damage, remain poorly understood. Furthermore, NextGen CUT&Tag provides insights into the differential levels and sites of DNA damage events in normal and cancer cells, as well as DNA damage occurring during disease treatment. In some implementations, this technology uses non-invasive techniques to probe the epiproteome and circulating extracellular chromatin fragments obtained in blood and liquid biopsies to gain deeper insights into the origin, developmental stage, and metastatic potential of cancer.
[0108] Although the content disclosed herein references certain illustrated implementations, it should be understood that these implementations are presented as examples rather than as limitations. Example
[0109] This document provides a technique for mapping DNA binding that offers improvements over existing techniques such as ChIP-Seq, CUT&RUN, Split DamID, CUT&Tag, Multi-CUT&Tag, CoBATCH, scChIC-Seq, ACT-Seq, and Co-ChIP. Specifically, this technique does not use protein A fusion-based or nanobody-based methods to conjugate a pre-loaded transposase (e.g., Tn5 transposase) with an affinity reagent (e.g., an antibody). (E.g., this technique is protein A fusion-free, and in some embodiments, it is pre-loaded transposase-free.) Therefore, this technique minimizes and / or eliminates background and cross-signal ambiguity.
[0110] Materials and methods Biomaterials & Reagents K562 cells were grown in RPMI medium (Gibco, 11875119) supplemented with 10% FBS (Gemini Bio, 100-602-500) and 1% penicillin-streptomycin (ThermoFisher, 15140122). For sodium butyrate treatment, freshly grown K562 cells were seeded at a density of 100,000 cells / mL in 6-well plates. To treat the cells, 1 mM sodium butyrate (Millipore Sigma, 19-137) was added to the cell culture and incubated for 72 hours. Distilled water (ThermoFisher, 10977023) was added to the control cells. All antibodies used in this study are listed in Table 1. All reagents and materials used in this study are listed in Table 2. All oligonucleotides used for barcoding in this study were ordered from Integrated DNA Technologies and are listed in Tables 3 and 4.
[0111] Table 1 - Antibodies
[0112] Table 2 - Reagents
[0113] Table 3 - Barcode Allocation and Connecting Subsequences
[0114]
[0115]
[0116]
[0117]
[0118]
[0119]
[0120]
[0121] Table 4. Oligonucleotide sequences
[0122] Antibody barcoding The antibody should be in PBS buffer before the reaction. Incubate the antibody and NHS-PEG12-Biotin (ThermoFisher, A35389) overnight at 4 °C with a molar ratio between 1:0.1 and 1:100. The next day, use 40K Zeba... TM The biotinylated antibody buffer was exchanged three times with PBS using a desalting column or plate (ThermoFisher, 87767, 87775). For adaptor annealing: a 500 µM Tn5MErev oligonucleotide stock solution was prepared in water. 100 µM P5 and P7 adaptor oligonucleotides were prepared in water. 10 µL of Tn5MErev oligonucleotide, 50 µL of one of the adaptor oligonucleotides, and 40 µL of distilled water (ThermoFisher, 10977023) were mixed and incubated at 95°C for 2 minutes, then slowly cooled to room temperature. To prepare a pair of barcoded adaptors, equal volumes of one P5 and one P7 adaptor were mixed. In this study, P5 and P7 adaptors from the same number were paired, which will contain the same barcode. For antibody barcoding: each antibody was prepared separately in different tubes or wells. Mix 10 µg of biotinylated antibody, 0.39 µL of streptavidin (ThermoFisher, 21122), and 2.34 µL of adaptor pair, and add to a total volume of 100 µL via PBS. Incubate the mixture at room temperature for 1 hour. Add 2.25 µM of D-biotin (ThermoFisher, B20656) to the mixture and incubate at room temperature for 30 minutes. Pool all antibody mixtures together. Concentrate the mixture through a 30K Amicon centrifuge filter (Millipore Sigma, UFC503096, UFC803096) and maintain at 4 °C.
[0123] High multiplex CUT & Tag methodPrimary antibodies were prepared as described above. Different antibodies were loaded with different barcoded adaptor pairs. Starting with 100,000 cells, 10 µL of concanavalin A-coated magnetic beads (Polysciences, 86057-3) were used. The concanavalin A beads were activated by washing twice in binding buffer (20 mM HEPES pH 7.5, 10 mM KCl, 1 mM CaCl2, 1 mM MnCl2). 100,000 freshly grown K562 cells were washed once in PBS and once in wash buffer (20 mM HEPES pH 7.5, 150 mM NaCl, 0.5 mM spermidine, 1× protease inhibitor mixture). The cells were resuspended in 0.5 mL of wash buffer and transferred to the activated concanavalin A beads. The cells and beads were incubated at room temperature in a vortex mixer for 15 min. The buffer was removed by placing the tube on a magnetic holder. Cells were resuspended in 100 µL of wash buffer containing an antibody mixture (1 µg per antibody), 0.05% digitalis saponin, and 2 mM EDTA. The cells were incubated in a vortex mixer at room temperature for one hour. The beads were washed four times with Dig-med buffer (20 mM HEPES pH 7.5, 300 mM NaCl, 0.5 mM spermidine, 0.01% digitalis saponin, 1× protease inhibitor mixture). The beads were resuspended in 100 µL of Dig-med buffer containing 10 mM MgCl2 and 5 µg Tn5 (Diagenode, C01070010-20). The cells were incubated in a vortex mixer at 37°C for one hour. To stop tag fragmentation, 3.33 µL of 0.5 MEDTA, 1 µL of 10% SDS, and 0.33 µL of 20 mg / mL proteinase K (ThermoFisher, EO0491) were added. Vortex and incubate at 50°C for 1 hour. Purify the DNA as follows: Add 100 µL of phenol-chloroform-isoamyl alcohol (pH 8) (ThermoFisher, 17908) and mix thoroughly. Transfer the sample to a phase-lock tube (ThermoFisher, NC1093153) and centrifuge at 16000 g for 3 min at room temperature. Add 100 μL of chloroform to the aqueous phase and centrifuge at 16000 g for 5 min. Transfer the aqueous phase to a new tube and add 250 µL of 100% ethanol and 8.75 µL of 20 mg / mL glycogen. Incubate overnight at -80°C. The next day, centrifuge at 16000 g for 15 min at 4°C. Wash the precipitate with 1 mL of 100% ethanol. Centrifuge at 16000 g for 5 min at 4°C.After the precipitate was dried, it was dissolved in 23 μL of 10 mM Tris-HCl pH8 containing 1 / 100 RNase A (ThermoFisher, EN0531). It was incubated at 37°C for 10 min. To amplify the library, 21 µL of purified DNA, 2 µL each of barcoded i5 primer (10 µM) and i7 primer (10 µM) were mixed, using different combinations for each sample. The sequences of the i5 and i7 primers are listed below. The barcoded sequences followed the previous paper
[19] . 25 µL of NEBNext Ultra II Q5 MasterMix (NEB, M0544S) was added and mixed gently. It was incubated in a thermal cycler with the following program: 1 cycle of 72 °C for 5 min, 98 °C for 30 s; 17 cycles of 98 °C for 10 s, 63 °C for 10 s; 1 cycle of 72 °C for 1 min, held at 4 °C. The library was cleaned using AMPure XP beads (Beckman, A63881) at a 1:1.1 ratio, following the manual. The library was then ready for sequencing.
[0124] i5 primer (SEQ ID NO:85): 5'-AATGATACGGCGACCACCGAGATCTACACNNNNNNNNNTCGTCGGCAGCGTC-3' (N: 11 nt barcode) i7 primer (SEQ ID NO:86): 5'-CAAGCAGAAGACGGCATACGAGATNNNNNNNNNNNGTCTCGTGGGCTCGG-3' (N: 11 nt barcode) Each of the i5 and i7 primers contains an 11-nt barcode indicated by NNNNNNNNNNN (SEQ ID NO:87) in the sequence provided above. Various barcode sequences of the i5 and i7 primers are provided in Mezger (2018) “Hi-plexchromatin accessibility profiling at single-cell resolution” Nat Commun 9:3647, which is incorporated herein by reference.
[0125] Single-cell high-multiplex cut & tag method.
[0126] To obtain at least ~5000 cells, 100,000 K562 cells were collected and the sample was prepared in batches as follows: Cells were washed once with PBS and once with wash buffer. 10 µL of concanavalin A beads were activated by washing twice with binding buffer. Cells were resuspended in 1 mL of NP-wash buffer (wash buffer, 0.01% digitalis saponin, 0.01% NP-40) containing 20 mM sodium butyrate and incubated with the activated beads in a vortex mixer at room temperature for 15 min. The buffer was removed, and the beads were resuspended in 100 µL of NP wash buffer containing 2 mM EDTA. The barcode-loaded antibody mixture was added to the beads and incubated in a vortex mixer at room temperature for 1 h. The beads were washed four times with NP-Dig-med buffer (Dig-med buffer, 0.01% NP-40). The beads were resuspended in 100 µL of NP-Dig-med buffer containing 10 mM MgCl2 and 5 µg Tn5. Incubate at 37°C for 1 hour in a vortex mixer. Replace the buffer with 1 mL of 10 mM Tris-Cl containing 10 µg / mL DAPI (ThermoFisher, D1306). Push beads through a cell filter into round-bottom tubes (Falcon, 352235). Sort samples into 384-well plates using a MoFlo XDP instrument, one cell per well. Centrifuge the plate at 3000 g for 3 minutes at 4°C. Keep cells at -80°C until the following steps are performed. Add reagents to the 384-well plate using an Echo 650 acoustic liquid processor. Add 1 µL of 0.095% SDS to each well. Centrifuge the plate at 3000 g for 3 minutes. Incubate at 58°C for 1 hour. Add 0.5 µL of 2.5% Triton X-100 and 0.5 µL of a 10 µM i5 and i7 primer mixture to each well. Each well yields a unique index pair. Add 2 μL of NEBNext Ultra II Q5 Master Mix (NEB, M0544S) to each well. Centrifuge the plate at 3000 g for 3 minutes at 4°C. Incubate the plate in a thermal cycler using the following program: one cycle of 58 °C for 5 minutes, 72 °C for 5 minutes, and 98 °C for 30 seconds; 17 cycles of 98 °C for 10 seconds, 63 °C for 10 seconds; and one cycle of 72 °C for 1 minute, held at 4 °C. Pool the libraries using a single-well deep-well plate (Miltenyi Biotec, 130-114-966). Invert the 384-well plate onto the deep-well plate and centrifuge at 1000 g for 1 minute at 4°C. Repeat this process until all libraries have been collected from all 384-well plates. Transfer the pooled libraries to new tubes.The library was cleaned using AMPure XP beads (Beckman, A63881) at a 1:1.1 ratio, following the manual. The library was then ready for sequencing.
[0127] Example 1 - Hi-Plex CUT&Tag enables simultaneous and efficient spectral analysis of multiple nucleosomes and their associated regulatory factors.
[0128] To test the performance and multiplexing capabilities of this new technology, we barcoded a set of 36 mAbs that target 12 common histone markers, 14 histone modifying enzymes, 8 human TFs, CTCFs, and PolII (pSer2), respectively. Figure 4 We also included rabbit IgG as a negative control. Next, all 37 barcoded antibodies were pooled and, without Tn5, tested at room temperature with permeabilized K562 cells (10...). 5 We incubated the cells (1 cell per hour) for 1 hour. To detect the global binding sites of these Abs, we thoroughly washed the cells and incubated them with unloaded Tn5 and MgCl2 at 37°C. o Incubate at C for 1 hour. Finally, extract genomic DNA, amplify the tagged fragments using PCR, and then perform library preparation and next-generation sequencing. Figure 1A The measurement was performed in duplicate, yielding approximately 200 megabytes of readings.
[0129] Next, we demultiplexed the sequencing data by assigning antibody identities to each read and mapping the inserts back to the genome. 92.63% of the reads were successfully demultiplexed. To examine background signals, we extracted all reads containing at least one rabbit IgG barcode (referred to as single sequences) and found that they comprised only 0.07% of the total reads, indicating very low background. We also compared our IgG trajectories in K562 cells with those obtained using ChIP-seq, CUT&Run, and CUT&Tag, and found that our IgG reads were substantially sparser and much lower [9, 10, 4]. An example is illustrated using a 2 Mbp fragment on chromosome 3. Figure 1B We also examined the reproducibility of our dataset using scatter plot analysis of repeated determinations. The calculated correlation coefficients ranged from 0.934 to 0.994, indicating that our determinations were highly reproducible (see Examples). Figure 1C ).
[0130] To benchmark the utility of Hi-Plex CUT&Tag maps in identifying both silent and active transcriptional regions, we extracted single sequence trajectories of H3K27me3, RNAPII, and H3K4me3 from our Hi-Plex CUT&Tag reads and compared them with those from multi-CUT&Tag datasets (H3K27me3 & RNAPII), the ENCODE database (H3K4me3), and ATAC-seq data from the same cells. Figure 1D As shown in the figure, the H3K27me3 and RNAPII trajectories match very well between Hi-Plex CUT&Tag and multi-CUT&Tag. Similarly, the H3K4me3 trajectory obtained using Hi-Plex CUT&Tag is almost identical to those obtained using conventional ChIP-seq methods (blue trajectories). Figure 1D Importantly, although both largely overlap with the ATAC-seq tracks, as expected, additional ATAC tracks (black tracks) were found in the area covered by the H3K27me3 track. Figure 1D These analyses indicate that the Hi-Plex CUT&Tag technology can generate antibody-specific signals with minimal background caused solely by Tn5 transposase action. We also noted that the H3K27me3 and RNAPII tracks are mutually exclusive. In fact, global analysis of mutually exclusive tags, such as H3K9me3 vs. H3K9ac and H3K27me3 vs. H3K27ac, showed minimal overlap, indicating that cross-contamination between different antibodies is largely eliminated. Figure 1E ).
[0131] To determine the collaboration of multiple epigenetic regulators at a single locus, we next inquired whether the co-localization of two epitopes could be faithfully identified using Hi-Plex CUT&Tag. As an example, we stratified reads from transcription-associated H3K4me3 and RNAPII epitopes using only reads containing barcodes representing both epitopes at both ends of each read (called heterologous sequences), and compared those reads with existing H3K4me3 and RNAPII ChIP-seq trajectories from ENCODE. Figure 1F We found that the H3K4me3 / RNAPII heterologous sequence trajectories were mainly found at the overlap of single H3K4me3 and RNAPII ChIP-seq trajectories (green shaded area); Figure 1F On the other hand, when the H3K4me3ChIP-seq trajectory is in NEAT1When the main gene and another location are missing, the H3K4me3 / RNAPII heterologous sequence trajectory also disappears (orange shaded area); Figure 1F In addition, global analysis also revealed some overlapping signals between the two targets. Figure 1G This improved the accuracy of detecting co-localization events of H3K4me3 and RNAPII.
[0132] In recent studies, histone markers with hypothesized opposite biological functions, such as H3K4me3 (euchromatin) and H3K27me3 (facultative heterochromatin), have been found to be juxtaposed and termed “bivalent” domains
[11] . To examine whether Hi-Plex CUT&Tag could readily detect this type of colocalization, we extracted heterologous sequence reads with H3K4me3 and H3K27me3 barcodes at either end and found >5,000 peaks ( Figure 1H To identify bivalent domains in K562 cells, we compared existing single H3K4me3 and H3K27me3 ChIP-seq data to identify overlapping regions, as no multiple cut & tag or sequential ChIP-seq data were available. We identified ~950 H3K4me3 and H3K27me3 bivalent domains from the overlapping ChIP-seq data, of which 36% were covered by our H3K4me3 / H3K27me3 heterologous sequence peaks. Figure 1H (Left bar chart). For example, in NBPF1 In the promoter, we found a strong H3K4me3 / H3K27me3 heterologous sequence peak covering ~1,100 bp (pink shaded area). Figure 1I Furthermore, separate H3K4me3 and H3K27me3 ChIP-seq peaks were also found at the same locations, although the H3K27me3 peak was much weaker. Surprisingly, 94.9% of the H3K4me3 / H3K27me3 heterologous sequence peaks could not be identified using separate ChIP-seq data. Figure 1H Therefore, we carefully studied and found that our technology is highly sensitive in detecting such combined events. For example, in PLEKHG5 Seven H3K4me3 / H3K27me3 heterologous sequence peaks were identified in the gene body; while two and nine H3K4me3 and H3K27me3 ChIP-seq peaks (blue shaded areas) were observed respectively; Figure 1J Using the overlap analysis described above with ChIP-seq data, it was impossible to identify bivalent structural domains. Figure 1JNote that reads of heterologous sequence events with mixed barcodes accurately reflect the colocalization of the two epitopes at the same location on the same chromosome copy, since they originate from a single chromosome segment in the same cell. These results demonstrate that Hi-Plex CUT&Tag is highly sensitive to mapping epitope colocalization.
[0133] To determine whether our technique could improve the detection of epitope colocalization within the same cell, we summarized the number of reads for a total of 630 (=36x35 / 2) heterologous sequence events and 36 homologous sequence events (reads containing the same barcode at both ends), and found that ~80% of the reads were counted as heterologous sequences (data not shown). The number of reads varied considerably for each combination, partially reflecting the endogenous abundance of epitopes on chromatin. These results confirm that Hi-Plex CUT&Tag can significantly improve the detection of colocalization of two epitopes.
[0134] In summary, we have obtained compelling data demonstrating that Hi-Plex CUT&Tag can detect hundreds of different colocalization epitope pairs in the same cell with high sensitivity by significantly reducing background signal and cross-contamination.
[0135] In addition to large-scale spectral analysis, our method can also be used to probe a limited number of targets, involving up to three targets, depending on the available secondary antibodies. To distinguish it from Hi-Plex CUT&Tag, we refer to this method as Low-Plex CUT&Tag. Figure 5 ).like Figure 5 As shown, the process requires incubation with an unlabeled primary antibody followed by binding with a barcoded secondary antibody to enhance the signal. The steps, including the introduction and activation of unloaded transposases, DNA purification, and library preparation, are similar to those of Hi-Plex CUT&Tag. Using the same data processing as Hi-Plex CUT&Tag, we demonstrated that Low-Plex CUT&Tag can also effectively minimize background noise and cross-contamination (data not shown).
[0136] Example 2 - Analyzing the Complexity of Hi-Plex CUT & Tag Datasets Unlike existing ChIP-seq, ATAC-seq, CUT&Tag, and similar methods, each sequence read generated by Hi-Plex CUT&Tag requires two simultaneous tag fragmentation events within the same cell, and the length of the tagged fragmented chromosomal DNA provides a rough estimate of the distance between two epitopes, which is true for all heterologous sequence reads (i.e., tagged with two different barcodes). In other words, in addition to the genetic information stored in the tagged fragmented sequence, each Hi-Plex CUT&Tag sequencing read also carries information about the epitope combinations that generated the fragment and a rough chromosomal distance between the two epitopes. The structure of the new dataset can be represented by three information axes: genomic DNA sequence, epitope combinations, and the distance between each combination. Figure 2A Furthermore, the polarity / order of modifications can also be determined from heterologous sequence data.
[0137] In theory, using 36 antibodies would allow us to examine 36 homologous events and 630 (=36 x 35 / 2) heterologous events. The performance of each antibody varies considerably in terms of the number of reads produced, reflecting differences in epitope abundance, epitope stability, and antibody affinity on the chromatin. This becomes even more apparent when a heatmap is generated using SEACR to show the number of peaks invoked for each epitope combination derived from the 36 mAbs. Figure 2B As shown in the figure, H3K4me1, H3K4me3, H3K9me3, and H3K27me3 involved the most homologous and heterologous sequence reads, followed by RNAPII, CBP, and SUV39H1. On the other hand, most transcription factors did not produce high read numbers, indicating that they are more sparse and / or unstable on chromatin. Note that we did not cross-link chromatin to preserve epitope accessibility and availability. We therefore decided to filter out epitope combinations with peaks below the 20th percentile of the read number distribution and focus on the 501 epitope combinations generated for future analyses. We also noted that 68% of the qualifying peaks represented heterologous sequence events, indicating that this new technique significantly improves the possibility of epitope colocalization. Note that homologous sequence events with two or more nucleosome distances (e.g., >300 bp) may also be generated by two antibodies, since the length of the barcode sequence is 66 bp or 72 bp.
[0138] We then asked whether each epitope combination was associated with certain types of DNA sequences, such as cis-regulatory elements and repetitive DNA sequences, and whether there were any differences in CpG methylation levels. We performed hierarchical clustering analysis using stacked bar charts annotated with peaks of annotated cis-regulatory elements, repetitive elements, and average DNA methylation from the same cell line
[14] . All 501 epitope combinations were annotated with ENCODE cis-regulatory elements. Four major clusters were identified. The top eight combinations in each cluster were... Figure 2D As shown in the image.
[0139] In the first cluster, epitope assemblages were primarily associated with enhancer-like and promoter-like elements. This cluster mainly consisted of euchromatin markers such as heterologous sequence pairs H3K4me3 / RNAPII, H3K4m3 / H3K27ac, and H3K4m3 / H3K36me3. Regarding repeat elements, this group was characterized by a high proportion of simple repeats and a low percentage of transposon elements. Figure 2D ).
[0140] A higher proportion of gene bodies was observed in the second cluster. Its main feature was epitope combinations associated with RNA polymerase II and histone acetylation markers, such as H3K14ac / RNAPII and RNAPII / RNAPII. The prevalence of these features suggests a key role in transcriptional elongation, where RNA polymerase II actively transcribes genes and acetylation maintains open chromatin structures, thereby promoting efficient transcription. On the other hand, a higher proportion of SINEs and LINEs was observed in this group. Indeed, previous studies have shown that K562 cells express full-length L1 mRNA and L1-encoded proteins [15; 16]. L1 element activity is generally higher in K562 cells compared to many other cell types, consistent with the generally elevated retrotransposon activity observed in many cancer cell lines. Alu elements are the most common SINEs in humans and are also actively transcribed in K562 cells. A study by Li et al. showed that Alu repeats in K562 cells are anomalously hypomethylated and have significantly higher transcriptional activity than those in other human cell lines and somatic tissues
[17] .
[0141] The third cluster primarily consists of epitope combinations involving H3K27me3 PTM (post-translational histone modification). This cluster exhibits a high proportion of gene bodies and low proportion of DNase regions, indicating repressive function. The repetitive elements are mainly SINE, LINE, and LTR.
[0142] In the fourth cluster, most combinations involved H3K9me3 or H3K9me2 markers, and the potential genomic sequences were mainly repetitive elements (such as SINE, LINE, and LTR) and a high proportion of satellite DNA.
[0143] Significant differences were also observed in the average DNA methylation levels across the four clusters. The first and second clusters, which involved many open histone markers in epitope combinations, showed lower average DNA methylation levels, while the third and fourth clusters showed a wider dispersion of DNA methylation levels. This is very consistent with the proposed function of CpG methylation in gene silencing
[18] .
[0144] These observations prompted us to examine whether epitope combinations involving annotated euchromatin and heterochromatin markers showed any significant differences in their associations with cis-regulatory elements and repeat elements. Using boxplot analysis, we found that euchromatin markers including H3K4me3 and H3K27ac were highly enriched in dELS (distal enhancer-like features), pELS (proximal enhancer-like features), and PLS (promoter-like features), while more gene body sequences and regions with low DNase accessibility were significantly more associated with heterochromatin markers including H3K9me3 and H3K27me3. Regarding repeat elements, combinations with heterochromatin markers were more enriched than euchromatin markers in LINE, SINE, satellite DNA, LTR, and DNA repeats, except for simple repeats. Figure 2C These results are in good agreement with those reported in the literature; however, further analysis is needed to determine whether this correlation applies to each individual epitope combination (see [link to literature]). Figure 2D ).
[0145] Interestingly, a distinct ladder-like pattern was observed after PCR amplification of the tagged fragment species, rather than a diffuse pattern. Figure 6 A). Considering that each tag fragmentation class not only anchors two corresponding tag fragmentation events back to chromatin but also provides a rough distance between the two events, we inquired whether any unique features are associated with the length of the tag fragmentation class. Histogram analysis of all eligible reads clearly shows peaks at ~60 bp, 200 bp, and 380 bp, with deep valleys in between, and the signal rapidly disappears after 600 bp. Figure 6 B). The distance between two adjacent peaks differs by approximately 150 bp, consistent with the length of DNA wrapped around a single nucleosome. We therefore refer to these peaks as subnucleosomes (0-120 bp), single nucleosomes (120-300 bp), binucleosomes (300-460 bp), and triple nucleosomes (>460 bp) fragments.
[0146] Next, we sorted the epitope combinations based on the percentage of subnucleosome, single nucleosome, and dinucleosome types, respectively. The top examples in each category are illustrated using a stacked bar chart. Figure 2E Interestingly, it was noted that transcription factors (such as YY1, NRF1, cFos, and USF2) tend to have the highest percentage of tag fragmentation shorter than 80 bp, reflecting the fact that TFs typically have short footprints on chromatin due to sequence-specific binding activity. Since two adjacent tag fragmentation events are required to generate a read, these shorter reads may represent homodimer binding events. Indeed, YY1, NRF1, cFos, USF2, and Jun are known to form homodimers. On the other hand, the top-ranking combinations enriched for >300 bp reads involve pairings between euchromatin histone markers (e.g., H3K27me3 / H3K4me3) and / or their writers (e.g., EP300 / H3K27ac and EP300 / H3K9ac). This phenomenon may represent the dispersion of histone modifications across several nucleosomes, resulting in longer fragments.
[0147] Example 3 - Establishing a protocol for Hi-Plex CUT & Tag spectral analysis of single cells Previous studies have demonstrated the utility of CUT&Tag and multiplex CUT&Tag for profiling chromatin regulators at the single-cell level [4, 5, 6, 12, 13]. Consistent with this, we have developed a protocol to adapt Hi-Plex CUT&Tag for single-cell profiling analysis. Figure 3A To achieve this, we initially isolated cell nuclei from the cells and performed batch Hi-Plex CUT&Tag, following the procedure outlined above, until the tag fragmentation stage was complete. Subsequently, we stained the nuclei with DAPI and sorted individual nuclei into 384-well plates using flow cytometry. Single-cell library preparation began with lysing individual nuclei using SDS, followed by SDS quenching using a Triton X-100. Sequencing library amplification was achieved using different index primer pairs, which helped identify signals from individual cells. Library amplification occurred in each well after the addition of the PCR reaction mixture. The libraries from each cell were then pooled together, purified with Ampure XP beads, and prepared for sequencing.
[0148] As a proof-of-concept, we evaluated 16 of 36 targets in K562 cells. These targets encompassed six histone modifications (H3K4me3, H3K9me3, H3K9ac, H3K14ac, H3K27me3, and H3K27ac), 10 transcription factors (CTCF, RNAPII S2P, c-Jun, c-Fos, Max, Myc, USF1, USF2, NRF1, and YY1), and a negative control (rabbit IgG). Two replicates were performed, each consisting of 1,536 cells. Using a method similar to that used in batch experimental data analysis, we processed reads containing the same H3K27me3 barcode (referred to as homologous sequences) from both ends of each cell. Figure 3B We then evaluated the ensemble signal from all single cells and found that the enrichment of pseudo-batch reads at many locations matched that of batch reads, highlighting the high specificity of single-cell Hi-Plex CUT&Tag. Figure 3B ).
[0149] Our method can also be used with Chromium single-cell ATAC gel beads from 10x Genomics and Chromium Next GEM single-cell multi-omics ATAC+ gene expression gel beads to further increase cell number and co-profil RNA (data not shown). References
[0150] 1. Park PJ. ChIP-seq: advantages and challenges of a maturing technology. Nat Rev Genet. 2009 Oct;10(10):669-80. doi: 10.1038 / nrg2641. Epub2009 Sep 8. PMID: 19736561; PMCID: PMC3191340. 2.Zentner GE, Henikoff S. High-resolution digital profiling of theepigenome. Nat Rev Genet. 2014 Dec;15(12):814-27. doi: 10.1038 / nrg3798. Epub2014 Oct 9. PMID: 25297728. 3.Klein DC, Hainer SJ. Genomic methods in profiling DNA accessibilityand factor localization. Chromosome Res. 2020 Mar;28(1):69-85. doi: 10.1007 / s10577-019-09619-9. Epub 2019 Nov 27. PMID: 31776829; PMCID: PMC7125251. 4.Kaya-Okur HS, Wu SJ, Codomo CA, Pledger ES, Bryson TD, Henikoff JG,Ahmad K, Henikoff S. CUT&Tag for efficient epigenomic profiling of smallsamples and single cells. Nat Commun. 2019 Apr 29;10(1):1930. doi: 10.1038 / s41467-019-09982-5. PMID: 31036827; PMCID: PMC6488672. 5.Gopalan S, Wang Y, Harper NW, Garber M, Fazzio TG. Simultaneousprofiling of multiple chromatin proteins in the same cells. Mol Cell. 2021Nov 18;81(22):4736-4746.e5. doi: 10.1016 / j.molcel.2021.09.019. Epub 2021 Oct11. PMID: 34637755; PMCID: PMC8604773. 6.Gopalan S, Fazzio TG. Multi-CUT&Tag to simultaneously profilemultiple chromatin factors. STAR Protoc. 2022 Jan 20;3(1):101100. doi:10.1016 / j.xpro.2021.101100. PMID: 35098158; PMCID: PMC8783141. 7.Meers MP, Llagas G, Janssens DH, Codomo CA, Henikoff S.Multifactorial profiling of epigenetic landscapes at single-cell resolutionusing MulTI-Tag. Nat Biotechnol. 2023 May;41(5):708-716. doi: 10.1038 / s41587-022-01522-9. Epub 2022 Oct 31. PMID: 36316484; PMCID: PMC10188359. 8.N. Michael Green, Avidin, Editor(s): C.B. Anfinsen, John T. Edsall,Frederic M. Richards, Advances in Protein Chemistry, Academic Press, Volume29, Pages 85-133 (1975) 9.ENCODE Project Consortium. An integrated encyclopedia of DNAelements in the human genome. Nature. 2012 Sep 6;489(7414):57-74. doi:10.1038 / nature11247. PMID: 22955616; PMCID: PMC3439153. 10.Kanezaki R, Toki T, Terui K, Sato T, Kobayashi A, Kudo K, Kamio T,Sasaki S, Kawaguchi K, Watanabe K, Ito E. Mechanism of KIT gene regulation byGATA1 lacking the N-terminal domain in Down syndrome-related myeloiddisorders. Sci Rep. 2022 Nov 29;12(1):20587. doi: 10.1038 / s41598-022-25046-z.PMID: 36447001; PMCID: PMC9708825. 11.Bernstein BE, Mikkelsen TS, Xie X, Kamal M, Huebert DJ, Cuff J,Fry B, Meissner A, Wernig M, Plath K, Jaenisch R, Wagschal A, Feil R,Schreiber SL, Lander ES. A bivalent chromatin structure marks keydevelopmental genes in embryonic stem cells. Cell. 2006 Apr 21;125(2):315-26.doi: 10.1016 / j.cell.2006.02.041. PMID: 16630819. 12.Carter B, Ku WL, Kang JY, Hu G, Perrie J, Tang Q, Zhao K. Mappinghistone modifications in low cell number and single cells using antibody-guided chromatin tagmentation (ACT-seq). Nat Commun. 2019 Aug 20;10(1):3747.doi: 10.1038 / s41467-019-11559-1. Erratum in: Nat Commun. 2020 Sep 1;11(1):4424. doi: 10.1038 / s41467-020-18309-8. PMID: 31431618; PMCID: PMC6702168. 13.Wang Q, Xiong H, Ai S, Yu X, Liu Y, Zhang J, He A. CoBATCH forHigh-Throughput Single-Cell Epigenomic Profiling. Mol Cell. 2019 Oct 3;76(1):206-216.e7. doi: 10.1016 / j.molcel.2019.07.015. Epub 2019 Aug 27. PMID:31471188. 14.Zhang J, Lee D, Dhiman V, Jiang P, Xu J, McGillivray P, Yang H,Liu J, Meyerson W, Clarke D, Gu M, Li S, Lou S, Xu J, Lochovsky L, Ung M, MaL, Yu S, Cao Q, Harmanci A, Yan KK, Sethi A, Gürsoy G, Schoenberg MR,Rozowsky J, Warrell J, Emani P, Yang YT, Galeev T, Kong X, Liu S, Li X,Krishnan J, Feng Y, Rivera-Mulia JC, Adrian J, Broach JR, Bolt M, Moran J,Fitzgerald D, Dileep V, Liu T, Mei S, Sasaki T, Trevilla-Garcia C, Wang S,Wang Y, Zang C, Wang D, Klein RJ, Snyder M, Gilbert DM, Yip K, Cheng C, YueF, Liu XS, White KP, Gerstein M. An integrative ENCODE resource for cancergenomics. Nat Commun. 2020 Jul 29;11(1):3 doi: 10.1038 / s41467-020-14743-w. PMID: 32728046; PMCID: PMC7391744. 15.Blame DA, Moran JV. Ribonucleoprotein particle formation isnecessary but not sufficient for LINE-1 retrotransposition. Hum Mol Genet.2005 Nov 1;14(21):3237-48. doi: 10.1093 / hmg / ddi354. Epub 2005 Sep 23. PMID:16183655. 16.Iwamoto S, Suganuma H, Kamesaki T, Omi T, Okuda H, Kajii E.Cloning and characterization of erythroid-specific DNase I-hypersensitivesite in human rhesus-associated glycoprotein gene. J Biol Chem. 2000 Sep 1;275(35):27324-31. doi: 10.1074 / jbc.M003297200. PMID: 10862620. 17.Li TH, Kim C, Rubin CM, Schmid CW. K562 cells implicate increasedchromatin accessibility in Alu transcriptional activation. Nucleic Acids Res.2000 Aug 15;28(16):3031-9. doi: 10.1093 / nar / 28.16.3031. PMID: 10931917;PMCID: PMC108432. 18.Jones PA. Functions of DNA methylation: islands, start sites, genebodies and beyond. Nat Rev Genet. 2012 May 29;13(7):484-92. doi: 10.1038 / nrg3230. PMID: 22641018. 19.Mezger A, Klemm S, Mann I, Brower K, Mir A, Bostick M, Farmer A,Fordyce P, Linnarsson S, Greenleaf W. High-throughput chromatin accessibilityprofiling at single-cell resolution. Nat Commun. 2018 Sep 7;9(1):3647. doi:10.1038 / s41467-018-05887-x. PMID: 30194434; PMCID: PMC6128862. 20.Peter J Skene, Steven Henikoff. An efficient targeted nucleasestrategy for high-resolution mapping of DNA binding sites. eLife 6, e21856(2017). 21.Janssens, D.H., Wu, S.J., Sarthy, J.F. et al. Automated in situchromatin profiling efficiently resolves cell types and gene regulatoryprograms. Epigenetics & Chromatin 11, 74 (2018). 22.Sarah J. Hainer, Ana Bošković, Kurtis N. McCannell, Oliver J.Rando, Thomas G. Fazzio (2019) Profiling of Pluripotency Factors in SingleCells and Early Embryos, Cell 177, 1319-1329.e11. 23.Ku, W.L., Nakamura, K., Gao, W. et al. Single-cell chromatinimmunocleavage sequencing (scChIC-seq) to profile histone modification. NatMethods 16, 323–325 (2019). 24.Geisberg, J.V., and Struhl, K. Analysis of Protein Co-Occupancy byQuantitative Sequential Chromatin Immunoprecipitation. Curr Protoc MolBiology 68, 21.8.1-21.8.7. (2004). 25. Kinkley, S., Helmuth, J., Polansky, JK, Dunkel, I., Gasparoni, G., Fro ̈ hler, S., Chen, W., Walter, J., Hamann, A., and Chung, H.-R. reChIP-seq reveals widespread bivalency of H3K4me3 and H3K27me3 in CD4(+) memory Tcells. Nat. Commun. 7, 12514. (2016). 26. Weiner, A., Lara-Astiaso, D., Krupalnik, V., Gafni, O., David, E., Winter, DR, Hanna, JH, and Amit, I. Co-ChIP genome enables-wide mapping of histone mark co- occurrence at single-molecule resolution. Nat. Biotechnol. 34, 953–961. (2016). 27.Hass, MR, Liow, HH, Chen, For all purposes, all publications and patents mentioned in the foregoing specification are incorporated herein by reference in their entirety. Various modifications and variations to the use of the described compositions, methods, and techniques will be apparent to those skilled in the art without departing from the scope and spirit of the described techniques. Although the techniques have been described in conjunction with specific exemplary embodiments, it should be understood that the claimed invention should not be unduly limited to such specific embodiments. Indeed, it will be apparent to those skilled in the art that various modifications to the manner in which the invention is carried out are intended to be within the scope of the following claims.
Claims
1. A method for identifying nucleic acid binding sites of a target, the method comprising: (a) Contacting the target, which is bound to the nucleic acid binding site, with the tagged composition, thereby binding the tagged composition to the target, wherein the tagged composition comprises: (i) An antibody or antibody fragment that binds to the target; (ii) a heterocyclic compound linked to the antibody or the antibody fragment; (iii) Protein complexes; and (iv) Each of the above comprises two or more nucleic acids containing a barcode nucleotide sequence, wherein the two or more nucleic acids are linked to the heterocyclic compound; and (b) Contacting the two or more nucleic acids of the tagged composition with a transposase to form an antibody-barcode-transposase complex, wherein the antibody-barcode-transposase complex generates a double-strand break in the nucleic acid containing the nucleic acid binding site to produce a nucleic acid fragment containing the nucleic acid binding site; (c) Isolate the nucleic acid fragment; and (d) Sequencing the nucleic acid fragment to identify the nucleic acid binding site of the target.
2. The method according to claim 1, wherein the protein complex comprises avidin, streptavidin, or neutral avidin.
3. The method according to claim 1, wherein the heterocyclic compound comprises biotin.
4. The method according to claim 1, wherein the transposase comprises Tn5 transposase.
5. The method of claim 1, wherein each of the two or more nucleic acids further comprises a transposase mosaic sequence that binds to the transposase.
6. The method of claim 5, wherein the transposase mosaic sequence binds to Tn5 transposase.
7. The method of claim 1, wherein the target comprises a DNA-binding protein.
8. The method of claim 7, wherein the DNA-binding protein comprises a transcription factor, a regulatory element, a transcription repressor, a transcription activator, a polymerase, a nuclease, a nickase, a zinc finger protein, a transcription activator-like effector nuclease (TALEN), a glycosylation enzyme, a methyltransferase, a ligase, a restriction endonuclease, a replication protein, a helicase, or a kinase.
9. The method of claim 1, wherein the antibody or the antibody fragment is not directly linked to the two or more nucleic acids.
10. The method of claim 1, wherein the protein complex binds to the heterocyclic compound linked to the antibody or the antibody fragment, and to the heterocyclic compound linked to the two or more nucleic acids.
11. The method of claim 1, wherein the method further comprises adding magnesium to a sample containing the target and the tagged composition.
12. The method of claim 1, wherein each of the two or more nucleic acids further comprises an amplification handle.
13. The method of claim 1, wherein the method further comprises amplifying the nucleic acid fragment to provide a sequencing library.
14. The method of claim 13, wherein the amplification is polymerase chain reaction (PCR) amplification.
15. A composition comprising: (a) One or more antibodies or antibody fragments that bind to a target; (b) A heterocyclic compound linked to one or more of the antibodies or antibody fragments; (c) Protein complexes including avidin, streptavidin, or neutral avidin; and (d) Two or more nucleic acids, each containing: (i) Barcode nucleotide sequence; and (ii) Transposase mosaic sequence The two or more nucleic acids are linked to a heterocyclic compound. Furthermore, the composition forms a complex in solution.
16. The composition of claim 15, wherein the protein complex comprises streptavidin.
17. The composition of claim 15, wherein the heterocyclic compound comprises biotin.
18. The composition of claim 15, wherein the transposase comprises Tn5 transposase.
19. The composition of claim 15, wherein the antibody or antibody fragment comprises a region that binds to a DNA-binding protein.
20. The composition of claim 19, wherein the DNA-binding protein comprises a transcription factor, a regulatory element, a transcription repressor, a transcription activator, a polymerase, a nuclease, a nickase, a zinc finger protein, a transcription activator-like effector nuclease (TALEN), a glycosylation enzyme, a methyltransferase, a ligase, a restriction endonuclease, a replication protein, a helicase, or a kinase.
21. The composition of claim 20, wherein the protein complex is bound to the heterocyclic compound.
22. A reagent kit comprising: A first container comprising the composition according to claim 15; and A second container containing transposases.
23. The kit of claim 22 further comprises a reagent for tag fragmentation.
24. The kit according to claim 22 further comprises reagents and materials for isolating DNA and amplifying nucleic acids.
25. The kit of claim 22 further comprises a cell capture scaffold.
26. The kit of claim 25, wherein the cell capture scaffold comprises magnetic beads, columns, concanavalin A beads, streptavidin beads, colloidal semiconductor nanocrystals, carbon nanotubes, or microfluidic devices.
27. A method for identifying two or more target binding sites on a nucleic acid, the method comprising: a) Provide two or more barcode-coded affinity reagents, each comprising: Affinity reagents connected to a pair of adaptors, wherein: The first adaptor contains a first barcode nucleotide sequence and a first transposase-binding mosaic sequence, and The second adaptor contains a second barcode nucleotide sequence and a second transposase-binding mosaic sequence. The first barcode nucleotide sequence and the second barcode nucleotide sequence may be the same or different; and The two or more barcoded affinity reagents mentioned herein do not each contain a transposase. The two or more barcoded affinity reagents each bind to different targets, and The first and second barcode nucleotide sequences of each barcode affinity reagent are different from the first and second barcode nucleotide sequences of other barcode affinity reagents that bind to different targets; b) Add the two or more barcoded affinity reagents to a sample containing a target comprising each barcoded affinity reagent. Each target binds to the nucleic acid at a corresponding target-binding site. Each barcoded affinity reagent binds to the corresponding target or binds to a primary affinity reagent that binds to the corresponding target, and each affinity binding occurs in the absence of a transposase. c) Add unloaded transposase and transposase activator to the sample. The unloaded transposase binds to the first transposase-binding mosaic sequence and the second transposase-binding mosaic sequence of each barcoded affinity reagent, and The bound transposase fragments the nucleic acid and tags it with a first and second barcoded nucleotide sequence of a corresponding barcoded affinity reagent to provide tagged fragmented nucleic acid. It provides at least two tagged fragmented nucleic acids, which correspond to two or more barcoded affinity reagents, and each barcoded affinity reagent corresponds to a corresponding target binding site; d) Sequencing the tagged fragments of nucleic acid to provide nucleotide sequences; and e) Analyze the nucleotide sequence to identify the binding site of the target on the nucleic acid.
28. The method of claim 27, wherein the transposase is Tn5, Tn3, Tn7, TnY, Sleeping Beauty, or PiggyBac, and the transposase activator is MgCl2.
29. The method of claim 27, wherein the target is a DNA-binding protein, such as histones, histone-modifying enzymes, transcription factors, cofactors, or chromatin-associating proteins.
30. The method of claim 27, wherein the target is a post-translational modification on a histone or other chromatin-associated protein, or a modified DNA base.
31. The method of claim 30, wherein the modified DNA base is mC or 5hmC.
32. The method of claim 27, wherein the nucleic acid is part of chromatin, and the method further comprises simultaneously detecting histone markers, histone modifying enzymes, chromatin-associating proteins, and transcription factors.
33. The method of claim 32, wherein the chromatin-associating protein is CTCF or an adhesion protein.
34. The method of claim 27, wherein the affinity reagent comprises an antibody.
35. The method of claim 27, wherein the affinity reagent is a target-specific affinity reagent.
36. The method of claim 27, wherein the affinity reagent is a secondary affinity reagent that is specific to the primary target-specific affinity reagent.
37. The method of claim 36, wherein the primary affinity reagent does not contain a barcode.
38. The method of claim 36, further comprising adding the primary affinity reagent to the sample.
39. The method of claim 27, wherein providing the barcoded affinity reagent comprising an affinity reagent linked to the pair of adapters comprises: The first affinity portion is attached to the affinity reagent. The first and second adapters are provided, each having a second affinity portion, and the first affinity portion binds specifically to the second affinity portion.
40. The method of claim 39, wherein: The first affinity portion and the second affinity portion are a pair selected from the group consisting of: Biotin and avidin, streptavidin, or neutral avidin; The reaction provides a first reactive group and a second reactive group that are covalently linked; DNA-binding proteins and DNA sequences recognized by said DNA-binding proteins; HaloTag and chloroalkanes; SNAP tag and O(6)-benzylguanine; and Single-stranded DNA and its hybrid DNA.
41. The method of claim 27, wherein the first connector and the second connector each further comprise an amplification handle.
42. The method of claim 27, wherein analyzing the nucleotide sequence to identify the binding site of the target on the nucleic acid further comprises associating the barcode nucleotide sequence with an affinity reagent.
43. The method of claim 27, further comprising amplifying the tagged fragmented nucleic acid to provide a sequencing library.
44. The method of claim 43, wherein the amplification is polymerase chain reaction amplification.
45. The method according to claim 39, wherein: Each barcoded affinity reagent includes a first handle connected to a second handle via a spacer region; The first connector hybridizes with the first handle; and The second connector hybridizes with the second handle. The first handle or the second handle includes a first affinity portion that binds to a second affinity portion of the affinity reagent, and the first adapter and the second adapter include different amplification handles.
46. The method of claim 27, wherein the sample is cell, tissue, or cell-free DNA.
47. The method of claim 46 further includes permeabilizing the cells or the tissue.
48. The method of claim 27, wherein the method is a multiplex method for identifying more than one binding site on one or more nucleic acids, and the method comprises: a) Provide more than one barcode-coded affinity reagent, Each of the more than one barcoded affinity reagents mentioned herein does not contain a transposase. Each of the more than one barcoded affinity reagent binds to a different target; b) Add one or more of the barcode affinity reagents to the sample. c) Add the unloaded transposase and the transposase activator to the sample to provide more than one tagged fragmented nucleic acid; d) Sequencing of the nucleic acid fragments of the more than one tag to provide nucleotide sequences; and e) Analyze more than one nucleotide sequence to identify more than one binding site for more than one target.
49. The method of claim 27, wherein the nucleic acid is part of chromatin, and the method further comprises determining a data fingerprint of a combination of two target binding sites, wherein the data fingerprint comprises: a) Colocalization information of the two target binding sites, or lack of interaction between the two target binding sites; b) The distance between two target binding sites or epitopes; c) The nucleotide sequence of the tag-fragmented nucleic acid.
50. The method of claim 49, wherein the data fingerprint further comprises: d) The polarity or order of the modifications; e) Cis-regulating element; f) Close to or lacking CpG islands; g) Repeated DNA sequences; and / or h) Average DNA methylation level.
51. The method of claim 46, wherein the nucleic acid is a portion of chromatin.
52. The method of claim 46, wherein the method further comprises simultaneously identifying more than one histone marker, histone variant, histone marker reader, histone modifying enzyme, DNA modifying enzyme, chromatin-associating protein, transcription factor, RNA species and / or cofactor within the genome.
53. The method according to any one of claims 27-52, wherein the background IgG sequencing reads are less than 25%, 20%, 15% or 10% of the total sequencing reads.
54. The method according to any one of claims 27-52, wherein an affinity reagent-specific signal is generated, and the cross-contamination signal between different antibodies is less than 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%.
55. The method of claim 27 or 48 further comprises identifying the colocalization of two epitopes at a single locus in a cell.
56. The method of claim 55, wherein the co-localization of H3K4me3 and H3K27me3 is identified.
57. The method of claim 27 or 48 further comprises identifying bivalent domain regions in the sample covered by two histone modifications.
58. The method of claim 27 or 48 further comprises identifying the colocalization of two epitopes at the same location on the same copy of a chromosome originating from a single chromosome segment in the same cell.
59. The method of claim 48, wherein barcoding more than one affinity reagent to provide more than one barcoded affinity reagent comprises incubating each of the more than one affinity-tagged affinity reagents with a unique barcoded linker in separate reaction vessels to provide more than one separate barcoded affinity reagent.
60. The method of claim 59, further comprising bringing together the more than one separate barcode-coded affinity reagent to provide a mixture of barcode-coded affinity reagents.
61. The method of claim 48, wherein analyzing the more than one nucleotide sequence to identify more than one binding site of more than one target further comprises associating each of the more than one barcode nucleotide sequence with each of the more than one affinity reagent.
62. The method of claim 48, wherein the more than one target comprises 2 to 500 targets.
63. The method according to claims 27 and 48, further comprising: Separating the cell nucleus from the cell, Flow cytometry or gel beads can be used to sort individual cells or individual cell nuclei. Lysis of a single cell or a single cell nucleus. Expanding single-cell / nuclear libraries, including identifying signals originating from single cells. Collect single-cell / nuclear libraries, and The single-cell / nuclear library was sequenced.
64. The method according to claims 27 and 48, further comprising: Add the drug to the sample. Perform steps (a)-(e), and Compare how the drug disrupts characteristics in vitro or in vivo.
65. A reagent kit comprising: Descriptions are provided for two or more barcoded affinity reagents, each of which comprises: A pair of connectors, wherein: The first adaptor contains a first barcode nucleotide sequence and a first transposase-binding mosaic sequence, and The second adaptor contains a second barcode nucleotide sequence and a second transposase-binding mosaic sequence; Affinity reagent, Each adaptor contains a barcode nucleotide sequence and a transposase-binding mosaic sequence. Unloaded transposase, and Transposase activator.
66. The kit according to claim 65, further comprising: Cell or nuclear permeation buffer, and / or Washing buffer.
67. A reagent kit comprising: Two or more barcoded affinity reagents, each comprising, A pair of connectors, wherein: The first adaptor contains a first barcode nucleotide sequence and a first transposase-binding mosaic sequence, and The second adaptor contains a second barcode nucleotide sequence and a second transposase-binding mosaic sequence; Unloaded transposase, and Transposase activator.
68. The kit according to claim 67, further comprising: Cell or nuclear permeation buffer and / or Washing buffer.
69. The kit according to claim 65 or 67, further comprising a control.
70. The kit of claim 65, wherein the control is a recombinant nucleosome bound to DNA and / or a control affinity reagent.
71. The kit according to claim 65 or 67, comprising a set of affinity reagents.
72. The kit according to claim 65 or 67, comprising a group of affinity reagents specific to cancer.
73. The kit according to claim 65 or 67, comprising a set of affinity reagents specific to epigenomic marker proteins and / or histones.