Chromatin analysis compositions and methods
Through the target DNA analysis method that binds domain and nucleic acid barcode labeling, the problem in the prior art is difficult to simultaneously analyze multiple histone modifications and reveal their spatial distribution, and a highly parallel and sensitive histone modification analysis is achieved, supporting disease biology research and medical diagnosis.
Patent Information
- Application Number
- CN202380081268.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-11-23
- Filing Date
- 2023-11-22
- Publication Date
- 2025-08-05
AI Technical Summary
Existing histone modification analysis methods are difficult to analyze multiple modifications simultaneously and reveal their spatial distribution in tissues, and cannot meet the needs for chromatin spatial tissue and its potential histone modification.
A target binding conjugate comprising a binding domain and a linker is provided to specifically bind to histone modification or DNA binding proteins, target DNA is labeled using nucleic acid barcodes, and analyses of amplified barcode-labeled target DNA by sequencing to achieve high parallel, sensitive and high throughput histone modification analysis.
The simultaneous analysis of a variety of nucleosome modifications and DNA binding proteins is achieved, which can quantify and localize histone modifications, support disease biological research and medical diagnosis, especially cancer detection and therapeutic monitoring.
Smart Images

Figure CN120435554A_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims the benefit of priority to U.S. Provisional Application No. 63 / 427,749, filed on November 23, 2022, the entire contents of which are incorporated herein by reference.
[0003] Incorporation of Sequence Listing
[0004] This application contains a sequence listing that has been submitted electronically in XML format, the entire contents of which are incorporated herein by reference. The XML copy, created on November 22, 2023, is named 5371-105PCT and is 86016 bytes in size. Technical Field
[0005] The present disclosure generally relates to the identification and analysis of epigenetic and other modifications to the structure or characteristics of chromatin, nucleosomes, and associated nucleic acids. Background Art
[0006] Chromatin is a complex of DNA and proteins that organizes the genetic code within the nucleus of eukaryotic cells. The nucleosome is the basic structural unit of chromatin. A nucleosome consists of an octamer of protein surrounded by approximately two turns of DNA, plus an approximately 80-bp linker stretch of DNA. The two turns of DNA wrapped around the protein consist of approximately 146 base pairs. The protein octamer contains two copies each of histones H2A, H2B, H3, and H4. The 80-bp linker stretch connects the nucleosome to another nucleosome in a repeating pattern, forming the chromosome.
[0007] The structural features of chromatin that regulate gene expression are determined by how nucleosomes and other proteins are assembled within the cell nucleus. Loosely packed chromatin, called euchromatin, is transcriptionally active; the genes encoded by these regions are expressed. Tightly packed chromatin, called heterochromatin, is inactive in gene expression. Modifications on the histone tails are among the most important features that determine how chromatin is organized. Cells can manipulate its assembly, and therefore gene expression, by adding or removing these modifications.
[0008] Histone tails are disordered extensions of the N-terminal domain of each histone protein outward from the nucleosome core structure. Their length ranges from about 25 to 60 amino acids and are generally rich in basic amino acid residues, particularly lysine and arginine. Most histone tail modifications are methylation and acetylation of specific lysine and arginine residues, but other modifications including phosphorylation and ubiquitination also occur naturally. These modifications are known to be closely related to cell development, tissue differentiation, aging, and disease processes such as cancer. The enzymes involved in histone modification are clinically validated drug targets. For example, there are currently four FDA-approved cancer chemotherapy drugs on the market that inhibit histone deacetylation, which are enzymes involved in removing acetyl marks on histone tails.
[0009] Another key regulator of gene expression associated with chromatin is transcription factors. There are 1,500–1,600 transcription factors in the human genome. Transcription factors are DNA-binding proteins that activate or repress gene expression by coordinating the entry of RNA polymerase II into the gene promoter region. RNA polymerase II, another DNA-binding protein, transcribes DNA into various RNA molecules.
[0010] Given the key role of histone modifications as regulators of chromatin assembly, gene expression, and human health, a variety of methods have been developed to determine which histone modifications are associated with specific DNA sequences within each nucleosome. Currently, the most mainstream analytical method is chromatin immunoprecipitation sequencing (ChIP-seq), which combines chromatin immunoprecipitation with DNA sequencing. In ChIP-seq, antibodies specific for histone modifications are coupled to microbeads. The chromatin sample is broken down into isolated nucleosomes by mechanical shearing or enzymatic treatment. A solution containing these nucleosomes is then bound to the microbeads under conditions suitable for the antibody to bind to its cognate histone modification. Consequently, the nucleosomes carrying the modification are bound to the microbeads, and the nucleosomes can be isolated by separating the microbeads from the solution. After separation, the microbeads are washed to remove the modified histones, and a library for DNA sequencing is prepared. Next-generation DNA sequencing is then used to identify the specific DNA sequences corresponding to the modified nucleosomes. Antibody-guided chromatin tagging (ACT-seq) is an alternative approach that also uses modification-specific antibodies to identify histone tail modifications, but when antibody-specific histone tail modifications are present, barcode-loaded transposomes are used to tag the nucleosomal DNA.
[0011] Upon cell death, the genome is degraded and chromatin, primarily in the form of nucleosomes, is released into the bloodstream, forming cell-free circulating nucleosomes that retain histone modifications.
[0012] There are over 60 different modifiable sites on the histones of a nucleosome, each of which can be modified in more than one way, resulting in a vast array of possible modifications. Some of these modifications are more prevalent than others and have clear functional significance. The most commonly analyzed histone modifications are acetylation, methylation, phosphorylation, ubiquitination, and sumoylation. Depending on the nature and location of the modification, the modification can activate or repress gene expression. Modifications are known to act synergistically to influence gene expression.
[0013] Traditional histone modification analysis methods provide information on the types and levels of histone modifications present in a sample, but cannot reveal the spatial distribution of these molecules in tissues.
[0014] Despite the importance and complexity of these modifications, existing histone modification analysis methods have limited ability to analyze multiple modifications simultaneously without separating samples and analyzing each modification in a separate experiment. Another unmet need is to obtain information about the spatial organization of chromatin and its underlying histone modifications in the context of a tissue. Therefore, there is a need in the art for improved compositions and methods for identifying, analyzing, quantifying, and localizing chromatin and nucleosome modifications with respect to genomic coordinates and in a broader tissue context. These advances will pave the way for discovering key regulatory mechanisms in biology in health and disease and for developing new therapeutic modalities in medicine. Summary of the Invention
[0015] Provided herein are compositions and methods for identifying and analyzing epigenetic and other chemical modifications of nucleosomes comprising nucleic acids and histone or DNA-binding proteins. The present disclosure provides highly parallelizable, sensitive, accurate, and high-throughput methods for simultaneously analyzing a potentially unlimited number of nucleosome modifications and DNA-binding proteins.
[0016] In some embodiments, the present disclosure provides a kind of target binding conjugate comprising binding domain and joint (adapter), wherein binding domain specific binding is on histone modification or DNA binding protein, and wherein joint comprises specificity corresponding to (unique to) by the target that described binding domain specific binding.In some respects, the present disclosure includes composition, and described composition comprises nucleosome binding conjugate and buffer (such as connection buffer).In some respects, DNA binding protein can be attached to on the DNA region connecting two nucleosomes.
[0017] In some embodiments, the present disclosure provides a composition comprising (i) a substrate; (ii) a binding domain coupled to the substrate; and (iii) a linker; wherein the binding domain specifically binds to a histone modification or a DNA-binding protein, and wherein the linker comprises a nucleic acid barcode sequence that specifically corresponds to the histone modification or the DNA-binding protein.
[0018] In some embodiments, the present disclosure provides a method for analyzing a plurality of nucleosomes and protein-DNA complexes, the method comprising: (i) contacting a plurality of substrates comprising at least one composition of any one or combination of the numbered aspects disclosed herein with a solution comprising a plurality of nucleosomes and protein-DNA complexes, wherein the binding domain binds to a DNA-binding protein or a nucleosome comprising a histone modification; (ii) ligating an adapter carrying a nucleic acid barcode to a target DNA comprising a nucleosome comprising a histone modification or a DNA-binding protein; (iii) introducing a ligation universal sequence, for example, for amplifying the target DNA; (iv) amplifying the barcode-labeled target DNA; and (v) analyzing the amplified barcode-labeled target DNA by sequencing.
[0019] In some embodiments, the present disclosure provides a method for analyzing a plurality of nucleosomes and protein-DNA complexes, the method comprising: (i) contacting a plurality of matrices (comprising at least one of any one or combination of the numbered aspects disclosed herein) with a solution comprising a plurality of nucleosomes and protein-DNA complexes, wherein the binding domain binds to a DNA-binding protein or a nucleosome comprising a histone modification; (ii) ligating an adapter carrying a nucleic acid barcode to a target DNA comprising a nucleosome comprising a histone modification or a DNA-binding protein; (iii) releasing the nucleosome or DNA-binding protein from the matrix by cleaving the ligated adapter; (iv) repeating steps (i) to (iii) at least once; (v) introducing a ligated universal nucleic acid sequence, for example, for amplifying the target DNA; (vi) amplifying the barcode-labeled target DNA; and (vii) analyzing the amplified barcode-labeled target DNA by sequencing.
[0020] In some embodiments, the present invention provides a method for analyzing a plurality of nucleosomes and protein-DNA complexes, the method comprising: (i) contacting one or more substrates comprising a composition of any one or combination of the numbered aspects disclosed herein with a solution comprising a plurality of nucleosomes and protein-DNA complexes, wherein a binding domain binds to a DNA-binding protein or a nucleosome comprising a histone modification; (ii) adding an adapter to the plurality of nucleosomes bound to the binding domain; (iii) ligating an adapter carrying a nucleic acid barcode to a target DNA of a nucleosome comprising a histone modification or a DNA-binding protein; (iv) releasing the nucleosome from the binding domain by adding a buffer that disrupts the interaction between the binding domain and the nucleosome; (v) repeating steps (i) to (iv) at least once; (vi) introducing a universal sequence for ligation, for example, for amplifying the target DNA; (vii) amplifying the barcoded target DNA; and (viii) analyzing the amplified barcoded target DNA by sequencing.
[0021] In some embodiments, the present disclosure provides a nucleosome-binding conjugate comprising: i) a binding domain; and ii) a linker bound to the binding domain, wherein the binding domain binds to a DNA-binding protein or a nucleosome comprising a histone modification, wherein the linker comprises a unique nucleic acid barcode sequence that specifically corresponds to the histone modification or the DNA-binding protein.
[0022] In some embodiments, the present invention provides a method for analyzing a plurality of nucleosomes and protein-DNA complexes, the method comprising: (i) contacting a solution comprising a plurality of nucleosomes and protein-DNA complexes with a solution comprising at least one nucleosome binding conjugate of any one or combination of the numbered aspects disclosed herein, wherein the binding domain binds to a DNA-binding protein or a nucleosome comprising a histone modification; (ii) ligating an adapter carrying a nucleic acid barcode of the binding conjugate to a target DNA comprising a nucleosome comprising a histone modification or a DNA-binding protein in an environment where the amount of off-target barcode-labeled DNA generated is less than 20% of the barcode-labeled target DNA to generate barcode-labeled target DNA; (iii) introducing a universal sequence for amplifying the target DNA; (iv) amplifying the barcode-labeled target DNA; and (v) analyzing the amplified barcode-labeled target DNA by sequencing.
[0023] In some embodiments, the present invention provides a method for analyzing a plurality of nucleosomes and protein-DNA complexes, the method comprising: (i) immobilizing a plurality of nucleosomes and protein-DNA complexes on a substrate at a spacing such that the off-target barcode labeling rate is less than 20%; (ii) contacting the immobilized nucleosomes and protein-DNA complexes with a solution comprising at least one composition of any one or combination of the numbered aspects disclosed herein, wherein the binding domain is attached to a DNA-binding protein or a nucleosome comprising a histone modification; (iii) attaching an adapter carrying a nucleic acid barcode of a nucleosome binding conjugate to a target DNA of a nucleosome comprising a histone modification or a DNA-binding protein; (iv) cleaving the adapter to generate a nucleic acid end having a structure suitable for ligation to another adapter; (v) repeating steps (ii) to (iv) at least once; (vi) introducing a universal nucleic acid sequence for ligation, for example, for amplifying the target DNA; (vii) amplifying the barcode-labeled target DNA; and (viii) analyzing the amplified barcode-labeled target DNA by sequencing.
[0024] In some embodiments, the present invention provides a method for analyzing multiple nucleosomes in a tissue environment, the method comprising: (i) immobilizing multiple nucleosome binding conjugates on a planar microarray substrate at a spacing such that the off-target barcode labeling rate is less than 20%; (ii) placing a tissue section on top of the planar microarray substrate comprising the multiple nucleosome binding conjugates; (iii) permeabilizing tissue cells; (iv) digesting chromatin with an endonuclease and capturing nucleosomes by the immobilized nucleosome binding conjugates; (v) performing a chromatin analysis on the tissue cells when the amount of off-target barcode labeled DNA is less than the amount of barcode labeled target DNA. In 20% of the environments, a linker carrying a nucleic acid barcode and a spatial identifier sequence of a nucleosome binding conjugate is ligated to a target DNA containing a nucleosome of a histone modification or DNA binding protein to generate a barcode-labeled target DNA; (vi) a universal sequence for amplifying the target DNA is introduced; (vii) the barcode-labeled target DNA is amplified; (vii) the amplified barcode-labeled target DNA is analyzed by sequencing; and (viii) the identity of the histone modification or DNA binding protein and its spatial location on the planar microarray substrate are determined based on the barcode and spatial identifier sequences.
[0025] In some embodiments, the present disclosure provides a method for analyzing multiple nucleosomes and protein-DNA complexes, the method comprising: (i) introducing a universal connector to a target DNA of a nucleosome or protein-DNA complex; (ii) contacting a solution comprising multiple nucleosomes and protein-DNA complexes with a solution comprising at least one binding conjugate of any one or combination of the numbered aspects disclosed herein, wherein the binding domain binds to a DNA-binding protein or a nucleosome comprising a histone modification; (iii) ligating the adapters of the bound multiple binding conjugates through a ligation reaction; (iv) hybridizing the universal connector of the target DNA to the 5' end of the ligation adapter; (v) replicating the sequence of the ligation adapter to produce a copy of the barcoded target DNA; (vi) introducing a ligation universal nucleic acid sequence, for example, for amplifying the target DNA; (vii) amplifying the barcoded nucleosomal DNA; and (viii) analyzing the barcoded target DNA by sequencing.
[0026] In some embodiments, the present disclosure provides methods for diagnosing a cancer or cancer subtype associated with one or more histone modifications, comprising analyzing a plurality of nucleosomes and protein-DNA complexes according to any one or a combination of the numbered aspects disclosed herein.
[0027] In some embodiments, the present disclosure provides a method for detecting the presence of cancer or monitoring cancer progression or cancer treatment response, comprising analyzing a plurality of nucleosomes and protein-DNA complexes according to any one or combination of the numbered aspects disclosed herein.
[0028] In some embodiments, the present disclosure provides a method for detecting histone modifications on cell-free nucleosomes as biomarkers in plasma liquid biopsies. Histone modifications on cell-free nucleosomes provide information about DNA-related activities within the original cell. In some embodiments, the present disclosure provides a multiplex detection method for histone modifications with low sample input, such as analyzing cell-free nucleosomes in plasma containing only 20-60 ng / mL nucleosomes.
[0029] In some embodiments, the present disclosure provides methods of any one or combination of the numbering aspects disclosed herein, comprising obtaining a plurality of nucleosomes and protein-DNA complexes from a blood sample.
[0030] In some embodiments, the present disclosure provides methods of any one or combination of the numbering aspects disclosed herein, comprising obtaining a plurality of nucleosomes and protein-DNA complexes from a tissue biopsy sample.
[0031] In some embodiments, the present disclosure provides a kit for monitoring epigenetic changes over time in a subject receiving cancer treatment, comprising a composition of any one or combination of the numbered aspects disclosed herein.
[0032] In some embodiments, the present disclosure provides any one of the molecules, complexes, workflows, or methods described in the accompanying figures or the following disclosure and examples.
[0033] These and other aspects of the present invention will become apparent upon reference to the following detailed description, claims, embodiments, procedures, compounds and / or compositions, and related background information and references, which are incorporated herein in their entirety. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figures 1A-1B The sample preparation for histone analysis is schematically shown, including end repair or A-tailing and / or ligation of the DNA ends wrapped around histones to universal capture sequences, depending on the requirements of the downstream analytical chemistry. As a non-limiting example, Figure 1A As a non-limiting example, blood containing circulating nucleosomes can be directly used for barcode analysis. Figure 1B It is shown that tissue or cell culture samples contain nucleosomes that can be isolated and analyzed by digesting chromatin with DNA nucleases.
[0035] Figures 2A-2B shows the co-localization analysis of histone modifications on the same nucleosome by DNA barcoding ( Figure 2A ) and multiplexed detection of different nucleosomal histone modifications ( Figure 2B ).exist Figure 2AIn the figure, "MBC1", "MBC2" and "MBC3" are linked to the same target nucleic acid to indicate the presence of three different histone modifications ("Mod"). Figure 2B Shown is the case where multiple nucleosomes are present, each with a single modification.
[0036] Figure 3A and 3B It is shown that multiplexed detection of histone modifications can be performed in two ways: linkers can be immobilized on the surface near the binding domain ( Figure 3A ), or the linker is directly fixed to the binding domain ( Figure 3B ). Figure 3A A schematic diagram of matrix-based barcode detection is shown, in which linkers are immobilized on the surface of microbeads. A pool of different microbead types is created, with each microbead displaying a binding domain and a barcoded linker. Because each microbead displays a binding domain and a barcoded linker, the surface density of molecules does not affect the specificity of the barcode. Figure 3B We show that barcoding by nucleosome-binding conjugates that are spaced apart on the surface significantly reduces off-target barcoding.
[0037] Figure 4 Schematic diagram of nucleosome barcoding using flanking Y-type adapters or, alternatively, bell-type adapters. "UMI" stands for Unique Molecular Identifier. The adapters on the left recapitulate the Illumina P5 and P7 adapters, where the MBC and UMI are sequenced as part of the index read. The adapters on the right integrate the MBC and UMI into the sequencing read frame. UFP and URP sites can be used to incorporate sequences from sequencing platforms other than Illumina.
[0038] Figures 5A-5B Detection of a single histone modification using an immobilized linker is shown ( Figure 5A ) or multiple histone modifications ( Figure 5B To detect individual histone modifications on individual nucleosomes ( Figure 5A ), nucleosomes are first immunoprecipitated with a pool of microbead matrix. Then, the forward linker is ligated (step 1), and the histone core is removed by denaturation (step 2). Complementary DNA strand synthesis is initiated by priming the UFP region and extending the primer with a DNA polymerase (step 3). Finally, the double-stranded DNA is ligated with a reverse linker (step 4) to form a DNA library for sequencing. In order to detect multiple histone modifications on a single nucleosome ( Figure 5BThis workflow uses beads with cleavable linkers. After the initial barcoding step is completed through a ligation reaction (step 1), cleavage at the uracil base (U) releases the linker from the surface (step 2). The barcoded nucleosomes are collected, recombined with the supernatant, and exposed to a pool of beads with different binding domains.
[0039] Figure 6 Colocalization of histone modifications is achieved by sequential coding with linkers in solution. Because the linker is not fixed to the binding domain, each barcoding cycle is performed using one bead.
[0040] Figures 7A-7B The nucleosome binding conjugates comprising a binding domain and a fixed linker are used to barcode nucleosome ends and detect one or two histone modifications ( Figure 7A ), and colocalization of histone modifications by continuous barcode labeling of nucleosomes attached to a matrix ( Figure 7B ).
[0041] Figure 8 Figure 3. Colocalization of histone modifications by proximal ligation. Multiple nucleosome-binding conjugates bind to the same nucleosome. Adjacent linkers are connected by splint ligation and attached to nucleosomal DNA by primer extension.
[0042] Figure 9 shows that colocalization of histone modifications can be achieved by sequential coding with linkers in solution (similar to Figure 6 ), the difference is that the linker structure allows the addition of UMI at the adjacent position of MBC in each barcoding cycle.
[0043] Figure 10 Describes the Figure 5A Workflow (as described in Example 4) Agarose gel electrophoresis of DNA libraries obtained for multiplex detection of histone modifications H3K4me3 and H3K4me2 in HeLa nucleosomes.
[0044] Figures 11A-11F Shown are sequencing results obtained from a bead-based dual barcoding assay using HeLa samples spiked with synthetic nucleosome controls as positive and negative controls. Figure 11AThe number of sequencing reads per MBC associated with each synthetic nucleosome is shown. The KmetStat_H3K4me3 control nucleosome is enriched in MBC101, indicating that H3K4me3 is correctly recognized. The KmetStat_H3K4me2 control nucleosome is enriched in MBC103, indicating that H3K4me2 is correctly recognized. The KmetStat_WT nucleosome is unmodified and acquires almost no MBC mark. Figure 11B The number of control nucleosome sequences recognized by each MBC is shown. For MBC101, KmetStat_H3K4me3 nucleosomes are the most representative reads. For MBC103, KmetStat_H3K4me2 nucleosomes are the most representative reads consistent with the correct barcode labeling. Figure 11C and Figure 11D The enrichment factor calculated from the raw sequencing reads is shown. The enrichment factor is defined as the number of reads per million for the IP reaction (beads carrying the H3K4me3 and H3K4me2 binding domains) divided by the number of reads per million for the INPUT reaction (beads targeting histone H3). Figure 11E and Figure 11F Shown are examples of two genes and a plot of read accumulation indicating modified genomic regions in HeLa cells.
[0045] Figure 12 Shown according to Figure 9 Agarose gel electrophoresis of DNA libraries obtained using the histone modification colocalization workflow shown and described in Example 7.
[0046] Figures 13A-13B Shows the use of Figure 9 Results from the colocalization workflow shown (using HeLa samples spiked with synthetic nucleosome controls as positive and negative controls). In this example, we tested the ability to perform continuous barcoding without the need for nucleosome washout. In the first cycle, H3K4me3 was recognized by attached MBC107. In the second cycle, H3K4me3 was recognized again, this time by attached MBC109. Figure 13A Sequencing reads showing synthetic nucleosome association with MBC107 and MBC109 are shown. Figure 13B The associated enrichment factors are shown. Figure 13C Example sequencing reads are shown, demonstrating dual barcoding ( Figure 13C SEQ ID NO:77-89 listed in , the presence of.
[0047] Figure 14 Shown are agarose gel electrophoresis and library yields obtained in an experiment that optimized the elution conditions for synthesized nucleosomes after the first barcoding cycle, ensuring that no damage was introduced that could interfere with a second barcoding cycle.
[0048] Figures 15A-15B Figure 2 shows spatial analysis of histone modifications in tissue cells. Nucleosome-binding conjugates containing linkers carrying modification barcodes and spatial identifier sequences were spotted on a microarray slide. The linkers were transferred to nucleosomes released from the tissue, and the positions of the tissue's nucleosome-derived cells relative to the microarray were identified. SP1 and SP2 are spatial identifiers for spots 1 and 2, respectively. Figure 15A ).like Figure 15B As shown, cells in tissue sections were permeabilized and then chromatin digested.
[0049] Description of Reference Numerals
[0050] IP: immunoprecipitation
[0051] MBC: Modified Barcode
[0052] MOD: Modification
[0053] Read 1: Read 1 primer site
[0054] Read 2: Read 2 primer site
[0055] 5'-P: 5'-phosphate
[0056] U:Uracil
[0057] UL: universal linker site
[0058] UFP: Universal forward primer site
[0059] URP: Universal reverse primer site
[0060] FSA: Forward Sequencing Adapter
[0061] RSA: Reverse Sequencing Adapter
[0062] UMI: Unique Molecular Identifier
[0063] RE: restriction site
[0064] "MBC" refers to modified bar code.
[0065] "SP" refers to the space identifier. DETAILED DESCRIPTION
[0066] Provided herein are compositions and methods for histone modification and DNA binding protein analysis. The method combines molecular recognition of histone modifications and DNA binding proteins with a step of writing information from the recognition event into the adjacent gene sequence of a target nucleic acid connected to a histone or DNA binding protein using a barcode. The resulting barcoded nucleic acid is then converted into a sequencing library and read by, for example, a nucleic acid sequencing method or other methods. This step reveals the barcode sequence associated with the target DNA. Sequencing can also be used to locate histone modifications and DNA binding proteins, such as transcription factors. The high-throughput analysis methods described herein allow the simultaneous identification of the properties and positions of multiple or all histone modifications. These methods can also be used to determine the abundance and stoichiometric ratio of histone modifications.
[0067] The disclosure of WO2022 / 115608 is incorporated herein by reference in its entirety for all purposes.
[0068] The present invention is fully described below by way of exemplary, non-limiting embodiments and accompanying drawings. However, the present invention may be embodied in many different forms and should not be construed as being limited to the embodiments described below. Rather, these embodiments are provided to further complete this disclosure and to fully convey the scope of the present invention to those skilled in the art.
[0069] The method steps described herein may be performed simultaneously or sequentially, unless otherwise indicated or it is clear from the method being described that a step must be performed first in order for its product to be used in a subsequent step.
[0070] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. The terminology used in this detailed description is for describing particular embodiments only and is not intended to be limiting of the invention.
[0071] All publications, patent applications, patents, GenBank / Uniprot or other accession numbers, and other references mentioned herein are incorporated by reference in their entirety for all purposes.
[0072] definition
[0073] The following terms are used in this specification and the appended claims:
[0074] The singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise.
[0075] Furthermore, as used herein, the term "about" when used to describe a measurable value, such as a length of a polynucleotide or polypeptide sequence, a dosage, a time, a temperature, and the like, can be used to describe a reasonably understood range of variation, for example, ±20%, ±10%, ±5%, ±1%, ±0.5%, or even ±0.1% of the specified amount.
[0076] Also as used herein, "and / or" refers to and includes any and all possible combinations of one or more of the associated listed items, as well as including the absence of combinations when interpreted in the alternative ("or").
[0077] Unless the context clearly indicates otherwise, the various features described herein may be used in any combination. In addition, in some embodiments, any feature or combination of features described herein may be excluded or omitted. To further illustrate, for example, if the specification indicates that a particular DNA base can be selected from A, T, G and / or C, the statement also indicates that the base can be selected from any subset of these bases, such as A, T, G or C; A, T or C; T or G; only C, etc., as if each such subset combination is explicitly listed herein. Further, the statement also indicates that one or more specified bases can be discarded. For example, in some embodiments, the nucleic acid is not A, T or G; is not A; is not G or C, etc., as if each possible discard scenario is explicitly listed herein.
[0078] As used herein, the terms "reduce," "reduces," "reduction," and similar terms may be used to disclose a reduction of at least about 10%, about 15%, about 20%, about 25%, about 35%, about 50%, about 75%, about 80%, about 85%, about 90%, about 95%, about 97%, or more.
[0079] As used herein, the terms "increase," "improve," "enhance," "enhances," "enhancement" and similar terms may be used to disclose an increase of at least about 10%, about 15%, about 20%, about 25%, about 50%, about 75%, about 100%, about 150%, about 200%, about 300%, about 400%, about 500%, or more.
[0080] As used herein, the term "histone modification" refers to the modification of chromatin-associated proteins. In some embodiments, nucleosomes include histone modifications. In some embodiments, histone modifications are acetylation, methylation, citrullination, phosphorylation, ubiquitylation (ubiquitylation) (also referred to as ubiquitination (ubiquitination)), SUMOylation, ADP ribosylation, deamination, proline isomerization, and other histone modifications known to persons skilled in the art. One or more. In some embodiments, histone modifications are SUMOylation of lysine or arginine. In some embodiments, histone modifications are phosphorylation of tyrosine, serine, and threonine. In some embodiments, histone modifications are any one of the modifications listed in Table 3 or any combination of these modifications.
[0081] The term "epigenetic change" is used herein to refer to the phenotypic changes of living cells, organisms, etc., which do not have the primary sequence (i.e., A, T, C, and G) of the coding cells or organism DNA. Epigenetic changes can include chemical changes of, for example, nucleotides and / or histones (i.e., proteins that participate in coiling and packaging DNA in the nucleus). Epigenetic changes can include histone modifications as described herein and other histone modifications well known to those skilled in the art. Common histone modifications include H3K4me1, H3K4me3, H3K36me3, H3K79me2, H3K9Ac, H3K27Ac, H4K16Ac, H3K27me3, H3K9me3. Exemplary DNA nucleotide modifications include common epigenetic markers 5-methylcytidine (5mC) and its oxidized product 5-hydroxymethylcytidine (5hmC), 5-formylcytidine (5fC), 5-carboxymethylcytidine (5caC). 5mC is widely known for its role in gene silencing, and increasing evidence suggests that its oxidative intermediates 5hmC, 5fC, and 5caC have metabolic functions in the 5mC demethylation pathway.
[0082] The term "genome" refers to all the DNA in a cell or cell population, or a selection of a specific type of DNA molecule (such as coding DNA, non-coding DNA, mitochondrial DNA, or chloroplast DNA). The term "transcriptome" refers to all the RNA molecules produced in a cell or cell population, or a selection of a specific type of RNA molecule (such as mRNA and non-coding RNA, or a specific mRNA in an mRNA transcriptome) contained in the complete transcriptome. In some embodiments, the transcriptome includes multiple different types of RNA, such as coding RNA (i.e., RNA that is translated into protein, such as mRNA) and non-coding RNA. Non-limiting examples of various types of RNA molecules found in the transcriptome (all of which may contain modified nucleosides) include: 7SK RNA, signal recognition particle RNA, antisense RNA, CRISPR RNA, guide RNA, long non-coding RNA, microRNA, messenger RNA, Piwi interacting RNA, repeat-associated siRNA, retrotransposon, ribonuclease MRP, ribonuclease P, ribosomal RNA, small Cajal body-specific RNA, small interfering RNA, smY RNA, small nucleolar RNA, small nuclear RNA and trans-acting siRNA. As used herein, the term "chromatin" refers to a molecular complex present in the nucleus of a eukaryotic cell, comprising proteins and polynucleotides (e.g., DNA, RNA). Chromatin is composed in part of histones that form nucleosomes, genomic DNA, and other DNA-binding proteins (such as transcription factors) that are typically bound to genomic DNA. The function of chromatin is to efficiently compress DNA into a small volume to fit into the nucleus and to protect the structure and sequence of DNA. Compaction of DNA into chromatin enables mitosis and meiosis, prevents chromosome breakage, and regulates accessibility for DNA replication and gene expression.
[0083] As used herein, the term "isolated chromatin" refers to a source of chromatin that can be obtained. Both isolated nuclei (which can be lysed to produce chromatin) and isolated chromatin (i.e., the product of lysed nuclei) are considered types of chromatin isolated from a cell population. The term "nucleosome" refers to a complex composed of at least one core of eukaryotic (e.g., mammalian, yeast, insect, or plant) mammalian histone proteins (e.g., two H2A proteins, two H2B proteins, two H3 proteins, and two H4 proteins), with a double-stranded DNA molecule of approximately 147 base pairs wrapped around the mammalian histone core. The structural characteristics of nucleosomes are well known in the art.
[0084] As used herein, the term "target nucleic acid" refers to a nucleic acid that is wrapped around histones to form a nucleosome. In some aspects, the target nucleic acid is a target DNA. The target DNA can be part of a nucleosome. The binding domains described herein can recognize and bind to histone modifications or DNA binding proteins of the nucleosome. In some aspects, the DNA binding protein can bind to a region of DNA that connects two nucleosomes.
[0085] As used herein, the terms "surface" or "matrix" are used to refer to any solid support. For example, the matrix can be a microbead, a microarray, a chip, a flow cell, a fluidic device, a plate, a slide, a culture dish, a membrane, a filter plate, or a three-dimensional matrix. A microarray is a slide dotted with biomolecules (e.g., a linker or a nucleosome binding conjugate), wherein each dot contains different components. A flow cell is a sample pool designed to allow a liquid sample to flow continuously. As described herein, the binding domain described herein can be coupled to one or more matrices, and the matrix can be coupled to one or more binding domains. The matrix can be made of a variety of materials. In some embodiments, the matrix is a resin, a membrane, a fiber, or a polymer. In some embodiments, the matrix includes sepharose, agarose, cellulose, polystyrene, polymethacrylate, and / or polyacrylamide. In some embodiments, the matrix includes a polymer, such as a synthetic polymer. Non-limiting examples of synthetic polymers include polyethylene glycol, polyisocyanopeptide polymers, polylactic-co-glycolic acid, poly-ε-caprolactone (PCL), polylactic acid, poly(3-hydroxybutyrate-co-3-hydroxyvalerate) (PHBV), chitosan, and cellulose.
[0086] like Figure 3A As shown in Figure 1, a "monoclonal matrix" can contain a binding domain and a barcode-tagged DNA linker. The matrix can be a microbead, a portion of a microarray, a channel of a microfluidic device, or a well of a microplate. Each monoclonal matrix contains a binding domain (such as an antibody) that specifically recognizes a histone modification and multiple copies of a DNA linker displaying a modification barcode (MBC). During or after immunoprecipitation of nucleosomes and DNA-binding proteins, the linker is transferred to the DNA to identify the modification.
[0087] As used herein, the term "barcode" refers to a synthetic nucleic acid. A specific barcode can be assigned to a specific nucleosome modification or DNA binding protein so that the target is specifically identified in the methods described herein. Therefore, if a barcode is specifically used to identify a histone modification or protein in one or more of the methods described herein, the barcode specifically corresponds to the histone modification or DNA binding protein. In other cases, the barcode specifically corresponds to the position of the nucleosome binding conjugate fixed on the microarray surface. Barcodes can be prepared using methods known in the art (such as solid phase oligonucleotide synthesis). In some embodiments, the barcode can be a DNA barcode (i.e., it can include a DNA sequence). In some embodiments, the barcode can include a synthetic DNA structure, such as a peptide nucleic acid (PNA) or a locked nucleic acid (LNA). In some embodiments, the synthetic DNA structure can include one or more modified bases. In some embodiments, the barcode can be an RNA barcode (i.e., it can include an RNA sequence). The barcode can be of any length, for example, ranging in length from about 4 to about 150 nucleotides. In some embodiments, the barcode is about 4 to about 20 nucleotides in length, for example, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, or about 20 nucleotides in length. Typically, a barcode comprises a rationally designed sequence that is not present in the genome of any known organism. However, in some embodiments, a barcode may comprise a known sequence. For example, the sequence of a barcode may comprise a characteristic sequence associated with a pathogen or other biological material. In some embodiments, a barcode may comprise a sequence that is specifically designed to facilitate a sequencing reaction. In this document, the terms "barcode" and "linker" are sometimes used interchangeably. As will be understood by those skilled in the art, in some embodiments, a linker may consist of a barcode. In some embodiments, a linker may comprise a barcode and one or more additional elements, as described below and Figures 2A-2B 、 Figure 4 、 Figures 7A-7B 、 Figure 8 、 Figure 9 and Figures 15A-15BAs shown. In some embodiments, the "joint" may include a spatial identifier (SP) sequence. As used herein, a "spatial identifier sequence" or "spatial identifier" defines the position of a nucleosome binding conjugate or joint on a microarray. In some embodiments, a "joint" may include a sequencing joint. In some embodiments, the sequencing joint includes a Y-type sequencing joint or a bell-shaped sequencing joint. In some aspects, the joint includes up to 20 random bases. In some aspects, the joint includes one or more non-natural nucleoside bases or modified bases (such as uracil and inosine). In some aspects, the joint includes at least one of a universal forward primer (UFP) and a universal reverse primer (URP). In some aspects, the joint includes a unique molecular identifier (UMI). In some aspects, the linker comprises one or more backbone modifications selected from locked nucleic acid (LNA), peptide nucleic acid (PNA), glycol nucleic acid (GNA), phosphorothioate, 2'-fluororibose, 2'-methoxyribose, phosphorodithioate, methylphosphonate, phosphoramide, guanidinopropylphosphoramidate, triazole, guanidinium salt, morpholino, threose nucleic acid (TNA), or hexitol nucleic acid (HNA). In some aspects, the linker comprises one or more 3' or 5' modification groups, wherein the one or more 3' or 5' modification groups are independently selected from dideoxyribose, phosphate, amino, inverted base, linker, or one or more other modifications.
[0088] like Figure 4As shown, the matrix microbeads can include modification-specific antibodies and Y-shaped sequencing adapters. In some embodiments, the adapter is fixed by a biotin-streptavidin interaction. In some embodiments, the adapter is fixed by a biotin-avidin interaction or a biotin-neutravidin interaction. In some embodiments, after the modified nucleosomes or DNA-binding proteins are captured by immunoprecipitation, adapter connection is performed, wherein the adapter comprises a barcode that recognizes the modification barcode (MBC), a unique molecular identifier (UMI), a sequencing read primer (Read1 and Read2) binding site, and a forward sequencing adapter (FSA) and a reverse sequencing adapter (RSA). In some embodiments, the adapter includes a UMI, an MBC, a universal forward primer site, and a universal reverse primer site (UFP, URP). In some embodiments, the adapter can additionally or alternatively be a bell-shaped adapter. For example, to detect multiple histone targets in the same reaction, microbeads of corresponding types are combined, each microbead carrying a specific barcode-tagged adapter and a modification-specific antibody (e.g., see Figure 3A ).like Figure 4 As shown, the histone core can be removed using a protease or denaturing conditions (such as DTT and heat treatment) before performing PCR. In some aspects, the bell-shaped linker can comprise uracil (U), which connects the sequencing linker element of the linker.
[0089] In some aspects, one or more elements (eg, linkers or binding domains) can be immobilized to a substrate using protein G, protein A, biotin (eg, via avidin, streptavidin, or neutravidin), a linker, or a recognition element.
[0090] The term "amplify", when used in reference to a nucleic acid, refers to making copies of the nucleic acid. Nucleic acids can be amplified using, for example, the polymerase chain reaction (PCR). Alternative nucleic acid amplification methods include: helicase-dependent amplification (LAMP), recombinase polymerase amplification (RPA), helicase-dependent amplification (HDA), multiple strand displacement amplification (MDA), nucleic acid sequence-based amplification (NASBA), self-sustained sequence replication (3SR), and rolling circle amplification (RCA).
[0091] As used herein, the term "coupling" can be used to describe the association between two or more components. For example, a first component and a second component can be coupled by covalent or non-covalent binding, or connected in other ways. In some aspects, a binding domain can be coupled to a substrate using a tether.
[0092] As used herein, the term "tether" refers to a bifunctional chemical group capable of linking one component to another component. In some embodiments, the first component can be a matrix and the second component can be a binding domain.
[0093] As used herein, the term "intra-complex linker transfer" or "intra-complex barcode transfer" refers to the process by which a linker and / or barcode is transferred to a target nucleic acid (i.e., DNA) when a binding domain is bound to the target nucleic acid (i.e., DNA). Thus, in this context, the term "complex" refers to the complex formed between a target nucleic acid and its cognate binding domain.
[0094] As used herein, the terms "crosstalk," "barcode crosstalk," and similar terms refer to the off-target transfer of a nucleic acid barcode. For example, barcode crosstalk can occur when a barcode from a binding domain is transferred to a nucleic acid that is not bound to the binding domain.
[0095] The term "DNA address" refers to a DNA or RNA sequence and / or its complementary sequence that serves as a programmable binding element to promote a specific binding event. For example, a nucleosome can bind to a nucleic acid sequence (e.g., a first DNA address sequence), which binds to a nucleic acid sequence displayed on a substrate (e.g., a second DNA address sequence), thereby anchoring the nucleosome thereto.
[0096] The term "restriction sequence" refers to a sequence that can be recognized by a restriction enzyme that specifically recognizes the restriction sequence.
[0097] Provided herein are linkers and binding domains, each of which is described in more detail below.
[0098] connector
[0099] As used herein, the term "linker" refers to any short nucleic acid sequence that can be bound to the end of a DNA or RNA molecule and impart some functionality to it. For example, in some embodiments, a linker can facilitate sequencing and / or identification of DNA or RNA molecules.
[0100] In some embodiments, the linker comprises a 5' phosphate. In some embodiments, the linker comprises a 3' phosphate. In some embodiments, the linker comprises a 5' phosphate and a 3' phosphate. In some embodiments, the linker is single-stranded. In some embodiments, the linker is double-stranded. In some embodiments, the double-stranded linker can include a single-stranded linker that hybridizes to a complementary oligonucleotide.
[0101] In some embodiments, the linker is coupled to the matrix by covalent attachment, affinity interaction or a combination thereof. In some embodiments, the linker includes a portion for surface anchoring. In some embodiments, the portion for surface anchoring is biotin or desthiobiotin. In some embodiments, the portion for surface anchoring is trans-cyclooctene (TCO), methyl tetrazine (mTET), dibenzocyclooctyne (DBCO), amino, azide or alkyne.
[0102] In some embodiments, the joint can be cleavable. For example, the joint can include one or more cleavage sites. For example, the cleavage site can include one or more uracil bases, enzyme (such as restriction enzyme or other nuclease) recognition sequence or synthetic chemical part. In some embodiments, the joint can be cut by the specific enzyme of uracil, inosine, 8-oxoguanine or ribonucleoside of the joint. For example, the joint can be cut by 8-oxoguanine-DNA glycosylase or its derivatives, uracil-DNA glycosylase (UDG), nuclease III, IV, V or VIII or its derivatives, or ribonuclease or its derivatives. In some embodiments, the joint includes a recognition sequence or restriction site and can be cut by the specific restriction enzyme of the restriction site.
[0103] In some embodiments, the adapter comprises a universal forward primer (UFP). In some embodiments, the adapter comprises a universal reverse primer (URP). In some embodiments, the adapter comprises a UFP and a URP. In some embodiments, the adapter consists of a UFP or a URP. UFP and URP sequences are non-naturally occurring DNA sequences and allow for selective amplification of only those sequences that are introduced into the target nucleic acid (or its copy). During sequencing, the UFP and / or URP are annealed to the DNA target, providing a starting site for extension of the new DNA molecule (i.e., its copy). See Table 1 for a list of exemplary UFPs and URPs.
[0104] Table 1 List of universal primers
[0105]
[0106]
[0107]
[0108] In some embodiments, the universal primer sequences used in the adapters (transferred to the target nucleic acid) are compatible with established DNA sequencing platforms, allowing for the incorporation of surface adapters (such as Illumina P5 and P7) in downstream PCR reactions.
[0109] In some embodiments, the linker can include a barcode, such as a modification-coding barcode (MBC). An MBC is a short, unique nucleic acid sequence. Each MBC is associated with a specific epigenetic modification and is used to assist in its identification and / or analysis. For example, an MBC can be used in a linker that is connected to a binding domain that specifically recognizes a specific histone modification. In some embodiments, the linker can be composed of a barcode. In some embodiments, the linker can be composed of an MBC.
[0110] The nucleosome binding conjugate comprises one or more linker sequences. Each linker sequence can comprise a universal sequence, a unique molecular identifier, a modified barcode, and a spatial barcode for spatial analysis. The spatial barcode indicates the spatial position of the nucleosome binding conjugate on the microarray. For example, the microarray can comprise 10,000 spots, each spot displaying multiple nucleosome binding conjugates. The nucleosome binding conjugates within the same spot share a spatial barcode, but can comprise different binding domains, in which case the modified barcode indicates the target of the binding domain. In some aspects, the spatial barcode length is 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40 bases.
[0111] In some embodiments, a plurality of nucleosome binding conjugates carrying a linker containing a specific spatial identifier are deposited on a microarray for spatial analysis of histone modifications. Array-based spatial analysis methods involve transferring one or more analytes in a biological sample to an array of characteristic points of a matrix, wherein each characteristic point is associated with a specific spatial position in the array. Subsequent analysis of the transferred analyte includes: determining the identity of the analyte and the spatial position of the analyte in the biological sample. The spatial position of the analyte in the biological sample is determined based on the characteristic points to which the analyte is bound (such as directly or indirectly) on the array, and the relative spatial position of the characteristic points in the array. Alternatively, a specific spatial identifier can be deposited at a predetermined position of the array characteristic point during preparation, so that only one type of spatial identifier exists at each position, thereby uniquely associating the spatial identifier with a single characteristic point of the array. If necessary, the array can be decoded using any method described herein so that the spatial identifier is uniquely associated with the array characteristic point position, and the corresponding relationship can be stored according to the aforementioned method.
[0112] The present disclosure includes a spatial barcoding method that uses spatial barcodes to label target molecules from individual cells or tissue regions. During sequencing, these barcodes are used to identify the source of the target molecules, thereby mapping histone modification patterns to specific locations in the tissue. This spatial information is used to understand the functional structure of the tissue and the role of histone organization in different cellular environments.
[0113] In some embodiments, in addition to the barcode, the adapter also comprises a universal sequence. As used herein, a "universal sequence" refers to a sequence that is not specifically associated with a binding domain or a histone or nucleosome modification. For example, a universal sequence is a sequence that is independent of the antibody and includes, but is not limited to, a sequencing adapter.
[0114] In some embodiments, the linker comprises a uracil base, an inosine base, an 8-oxoguanine base, or a ribonucleoside.
[0115] Histones, the fundamental building blocks of the nucleosome (the basic structural and functional unit of chromatin), are among the most highly conserved proteins. Nucleosomes are octamers wrapped around ~147 base pairs of DNA and composed of two copies of four core histones (H)—H2A, H2B, H3, and H4—joined together by the linker histone H1. These five histone classes, comprising over 60 different residues, constitute the primary protein component of chromatin, enabling the tight compaction of DNA. Furthermore, histones possess flexible N-termini, often referred to as "histone tails," which can undergo various combinations of post-translational modifications, dynamically allowing regulatory proteins to access DNA and fine-tune virtually all chromatin-mediated processes, including chromatin condensation, gene transcription, DNA damage repair, and DNA replication. Transcriptionally active and silent chromatin are characterized by specific patterns of histone post-translational modifications or their combinations. Histones can be post-translationally modified by "writers" and "erasers," a group of enzymes responsible for adding and removing chemical modifications. Through different combinations and patterns of histone post-translational modifications, they can form a "histone code".
[0116] In some embodiments, a linker may comprise a unique molecular identifier (UMI). A UMI consists of a short random sequence with 4 [UMI长度] For example, a 10-base UMI can encode 1,048,576 (4 10 ) specific molecules. UMIs are used for absolute quantification of sequencing reads to correct for PCR amplification bias and errors. For example, an RNA sample may contain 100 copies of transcript A and 100 copies of transcript B. Because transcript B amplifies more efficiently, 1M copies of transcript A and 200M copies of transcript B may be detected after PCR amplification. UMI tagging, on the other hand, links 100 specific UMIs to A and 100 specific UMIs to B. When using UMIs, transcript A will be detected at 10,000 copies of 100 UMI variants, and transcript B will be detected at 20,000 copies of 100 UMI variants. By counting the number of UMI variants rather than the number of reads, the absolute number of molecules can be obtained.
[0117] Typically, the length of the UMI is chosen to avoid UMI collisions, which are events where sequencing reads originating from two different genomic molecules are observed to have the same sequence and the same UMI. UMI collisions are related to the number of UMIs used, the number of specific alleles, and the frequency of each allele in the population. The ideal length of the UMI depends on the error rate and sequencing depth of the sequencing platform. Sequencing platforms with higher error rates require longer UMIs because UMI errors can lead to unexpected UMI collisions. Targeted sequencing (which sequences selected sites at a greater depth than whole genome sequencing) also uses longer UMIs because many alleles from different genomic molecules share the same sequence. Avoid using UMIs that are too long because they require more sequencing cycles, thereby shortening the reads of the actual target sequence. Long UMIs can cause priming errors in PCR reactions, resulting in sequencing artifacts. UMIs typically range from about 3 to about 25 nucleotides. In some embodiments, a UMI is about 3 to about 20 nucleotides in length, such as about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 1, about 13, about 14, about 15, about 16, about 17, about 18, about 19, or about 20 nucleotides in length. In some embodiments, a UMI can be 8 nucleotides in length. In some embodiments, a UMI can be 10 nucleotides in length.
[0118] Figures 2A-2B 、 Figure 4 、 Figures 7A-7B and Figure 8 Exemplary nucleic acid linker structures are shown, and the legend provides a description of the elements I used.
[0119] In some embodiments, Figure 2A The adapter shown is used for colocalization analysis to convert the presence of histone modifications (such as Mod1, Mod2, Mod3) into corresponding modification barcodes (such as MBC1, MBC2, MBC3). MBC is enzymatically linked to nucleosomal DNA with a universal forward primer site (UFP) or a universal reverse primer site (URP). The resulting NGS library is sequenced to obtain the histone modifications present in the nucleosome. Figure 2B In the examples, the target modifications are located on different nucleosomes. Each modification barcode identifies a different histone modification. The MBC is co-transferred with a universal forward primer site (UFP) or a universal reverse primer site (URP) onto nucleosomal DNA to construct a sequencing library. Sequencing reveals the DNA sequence associated with a given histone and the associated histone modification (as indicated by the modification barcode).
[0120] In some embodiments, a linker comprises a UFP, a URP, or a UFP and a URP. In some embodiments, a linker comprises a UFP and / or a URP and further comprises an MBC. In some embodiments, a linker comprises a UFP and / or a URP, an MBC, and a UMI. In some embodiments, a linker comprises a UFP and / or a URP, an MBC, and a UMI. In some embodiments, a linker comprises a UFP and / or a URP, an MBC, and a UMI. In some embodiments, a linker comprises a UFP, a URP, a UMI, and an MBC. In some embodiments, a linker comprises a UFP, a UMI, and an MBC. In some embodiments, a linker comprises a URP, a UMI, and an MBC. In some embodiments, a linker comprises an MBC and a UMI. In some embodiments, a linker comprises any of the configurations depicted in any of the figures.
[0121] In some embodiments, the linker is Y-shaped. In some embodiments, the Y-shaped linker comprises a UFP, an MBC, and a URP. In some embodiments, the linker is partially double-stranded, forming a Y-shape, wherein each single-stranded arm can comprise a universal sequence, a modified barcode, and a unique molecular identifier.
[0122] In some embodiments, the linker is bell-shaped. In some embodiments, the linker is partially double-stranded, forming a bell shape. In some embodiments, the single-stranded circle can contain a universal sequence, a modified barcode, and a unique molecular identifier.
[0123] In some aspects, according to one or more of the foregoing embodiments, the linker is partially double-stranded with single-stranded 3' and / or 5' overhangs, or the linker is partially double-stranded with single-stranded 3' and / or 5' overhangs on both sides. According to the foregoing embodiments, the linker can include double-stranded ends, which can be blunt ends or have single 3' base and / or 5' base overhangs.
[0124] In some embodiments, the linkers described herein can include one or more connectors, for example, connectors that facilitate attachment of the binding domain to the linker. The connectors can include polyethylene glycol, hydrocarbons, peptides, DNA, or RNA. The lengths of the connectors can vary. Longer connectors can be used where the histone modification or DNA binding protein is far from the 5' or 3' end of the nucleic acid sequence. Shorter connectors can be used where the histone modification or DNA binding protein is relatively close to the 5' or 3' end of the nucleic acid sequence.
[0125] In some embodiments, the linker or the linker sequence it comprises is cleavable. For example, the linker can comprise one or more cleavage sites. The linker can be chemically, photochemically, or enzymatically cleavable. For example, the cleavage site can comprise one or more uracil bases, an enzyme recognition sequence (such as uracil-DNA glycosylase, restriction enzyme or other nuclease), or a synthetic chemical moiety (such as a disulfide, carbonate, hydrazone, cis-aconityl, or β-glucuronide).
[0126] As described in further detail below, adapters can be fused to single-stranded or double-stranded target nucleic acids (eg, DNA or RNA) via a barcoding reaction.
[0127] As used herein, a "universal linker" refers to a sequence that is capable of hybridizing to the complementary sequence of any linker. In some embodiments, the universal linker can be a poly (A) oligonucleotide sequence, such as one synthesized by the action of dATP and terminal nucleotidyl transferase. In this embodiment, the poly (A) universal sequence can hybridize to the oligonucleotide T sequence in the linker.
[0128] In some embodiments, as Figure 8 As shown, a 3' poly A tail is added to the target. The 3' poly-A tail is added by polyadenylation using any known terminal nucleotidyl transferase (TD). In some embodiments, the 3' poly A tail is about 5, about 10, about 15, about 20, about 25, about 30, about 35, about 40, about 45, about 50, about 55, or about 60 bases in length.
[0129] In some embodiments, primer extension comprises adding a poly-T tail, a poly-G tail, a poly-A tail, or a poly-C tail to the 3' end of the DNA target. In some embodiments, the tail is about 5, about 10, about 15, about 20, about 25, about 30, about 35, about 40, about 45, about 50, about 55, or about 60 bases in length.
[0130] Binding domain
[0131] As used herein, the term "binding domain" refers to any nucleic acid, polypeptide, or other macromolecule that binds to a histone modification or a DNA binding protein of a target nucleosome. Depending on the context, it will be understood by those skilled in the art that the term "binding domain" can be used interchangeably with "binding agent," "recognition element," "antibody," and the like. In some embodiments, the binding domain binds to a histone modification. In some embodiments, the binding domain does not bind to any nucleic acid sequence flanking a histone modification. In some embodiments, the binding domain binds to a histone modification or a DNA binding protein. In some embodiments, the binding domain may bind to a conserved sequence motif. In some embodiments, the binding domain does not bind to any nucleic acid sequence flanking a histone modification or a DNA binding protein. In some embodiments, the binding domain binds to a histone modification that is methylated, citrullinated, acetylated, ubiquitinated, or sumoylated at lysine or arginine. In some embodiments, the binding domain binds to phosphorylation of tyrosine, serine, or threonine. In some embodiments, the binding domain binds to a DNA binding protein that is a transcription factor or RNA polymerase.
[0132] The binding domains described herein can be any protein, nucleic acid, or fragment or derivative thereof that can recognize and bind to a histone modification or DNA binding protein. For example, in some embodiments, the binding domain comprises an antibody, a linker, a reader protein, a writer protein, an eraser protein, an engineered macromolecular scaffold, an engineered protein scaffold, or a selective covalent capture agent, or a fragment or derivative thereof. In some embodiments, the binding domain comprises an IgG antibody, an antigen binding fragment (Fab), a single chain variable fragment (scFv), or a heavy or light chain single domain (V H and V L In some embodiments, the binding domain comprises a heavy chain antibody (hcAb) or hcAb (nanobody) V H H domain. In some embodiments, the binding domain is a bivalent binding domain that targets histone modifications. In some embodiments, the binding domain comprises an engineered protein scaffold, such as an adnectin, an affibody, anticalin, a trimer (atrimer), an avimer, a bicyclic peptide, a centyrin, a cystine knot (cys-knot), a darpin, a fynomer, a kunitz domain, an obody or a probetin. In some embodiments, the binding domain comprises a catalytically inactive variant of a DNA or histone writer or eraser protein. In some embodiments, the binding domain is covalently linked to the matrix or connected by an affinity interaction. Affinity interaction refers to an interaction between two binding partners that exhibits binding affinity. Examples of affinity interactions include, but are not limited to, biotin-avidin interactions or antibody-protein G interactions.
[0133] IgG antibodies are the major isotype of immunoglobulins. IgG consists of two identical heavy chains and two identical light chains, which are covalently linked and stabilized by disulfide bonds. H ) and light chain (V L ) and six complementarity determining regions (CDRs) that recognize antigens. Antibodies that bind to histone modifications or DNA-binding proteins are commercially available, for example, from EpigenTek, Abcam, and Active Motif.
[0134] Antibodies that bind to histone modifications or DNA binding proteins can be developed by those of ordinary skill in the art according to known methods and practices. In some embodiments, the antibody can be a monoclonal antibody, a polyclonal antibody, or a functional fragment or variant thereof. The term "antibody" as used herein encompasses any specific binding substance with a desired specific binding domain. Thus, the term encompasses antibody fragments, derivatives, functional equivalents, and homologs of antibodies, including any polypeptide comprising an immunoglobulin binding domain, whether natural or synthetic, monoclonal or polyclonal. Also included are chimeric molecules comprising an immunoglobulin binding domain or equivalent, fused to another polypeptide.
[0135] In some embodiments, the binding domain may comprise a nanobody. A nanobody comprises a single variable domain (V H H), produced by camelids and several cartilaginous fishes. H The H domain contains three CDRs that are enlarged compared to the CDRs of IgG antibodies and provides a size comparable to that of IgG (i.e., approximately ) similar antigen interaction surface. Nanobodies bind antigens with an affinity similar to that of IgG antibodies and offer several advantages associated with them: they are smaller (15 kDa), less sensitive to reducing environments due to fewer disulfide bonds, more soluble, and have no post-translational glycosylation. Nanobodies can be produced in bacterial expression systems, so they can be matured for affinity and specificity using phage and other display technologies. Other advantages include improved thermal stability and solubility, and methods for direct site-specific labeling. Due to their small size, nanobodies can form convex secondary epitopes, making them suitable for binding to inaccessible antigens. Exemplary methods for producing nanobodies include further evolving existing natural The library, or a combination thereof, is used to immunize the corresponding animal (eg camel) with the target antigen.
[0136] In some embodiments, the binding domain comprises a reader protein, a writer protein, or an eraser protein. A "reader protein" is a protein that selectively recognizes and binds to specific chemical modifications on histone tails. A "writer protein" is a protein that adds specific chemical modifications to histone tails. An "eraser protein" is an enzyme that can remove specific chemical modifications from histone tails. In some embodiments, the binding domain comprises a fragment or derivative of a reader protein, a writer protein, or an eraser protein. In some embodiments, the binding domain comprises an engineered form of a reader protein, a writer protein, or an eraser protein, such as an engineered form that retains nucleic acid binding ability but lacks any enzymatic activity. In some embodiments, the writer protein is a histone acetyltransferase, a CBP / P300 protein, a lysine methyltransferase, or an arginine methyltransferase. In some embodiments, the reader protein comprises a methyl-CpG binding domain (MBD), a bromodomain adjacent to a zinc finger protein domain (BAZ), a bromodomain (BRD), a malignant brain tumor (MBT), a plant homeodomain finger (PHD), a chromatin binding (chromo), a proline-tryptophan-tryptophan-proline (PWWP) domain, a tryptophan-aspartate dipeptide repeat domain (WD40), ankyrin repeats, or a Tudor domain. In some embodiments, the eraser protein is a histone deacetylase, a histone lysine demethylase, or a histone arginine demethylase. In addition, exemplary reader proteins, writer proteins, and erasers that can be used in the binding domains described herein are listed in Table 2. Other reader proteins, writer proteins, and erasers are listed at the following global website: rnawre.bio2db.com and are incorporated herein by reference.
[0137] Table 2 Reading proteins, writing proteins and erasing proteins
[0138]
[0139] Table 3 Histone modifications
[0140]
[0141]
[0142]
[0143]
[0144]
[0145]
[0146]
[0147]
[0148]
[0149]
[0150]
[0151] Can select and / or design binding domain for combining any histone modification or DNA binding protein.For example, histone modification can be acetylation, methylation, citrullination, phosphorylation, ubiquitylation (ubiquitylation) (also referred to as ubiquitination (ubiquitination)), SUMOylation, ADP ribosylation, deamination or proline isomerization.In some embodiments, histone modification is SUMOylation of lysine or arginine.In some embodiments, histone modification is phosphorylation of tyrosine, serine and threonine.In some embodiments, DNA binding protein is transcription factor, histone-protein complex or one or more histone subunits, or transcription repressor.Can select and / or design binding domain to combine any modification (such as acetylation, methylation, citrullination, phosphorylation, ubiquitylation (ubiquitylation) (also referred to as ubiquitination (ubiquitination)), SUMOylation, ADP ribosylation, deamination or proline isomerization) or nucleosome DNA binding protein.
[0152] target DNA
[0153] As used herein, the term "target DNA" refers to a nucleic acid sequence associated with a target histone modification or DNA binding protein. In some embodiments, the target DNA can be DNA containing nucleosomes modified by histones. In some embodiments, the target DNA includes nucleosome ends. In some embodiments, the method according to one or more embodiments includes performing end repair and / or A-tailing on the nucleosome ends.
[0154] DNA-binding proteins
[0155] As used herein, the term "DNA binding protein" refers to a protein that has a general or specific affinity for single-stranded or double-stranded DNA. In some embodiments, the DNA binding protein can be a protein associated with nucleosomes. For example, the DNA binding protein can be a protein that binds to DNA between nucleosomes or associated with nucleosomes. In some embodiments, the DNA binding protein is a transcription factor. In some embodiments, the DNA binding protein is RNA polymerase II. In some embodiments, the DNA binding protein is a transcription activator or a transcription repressor.
[0156] Adapter / barcode transfer reaction
[0157] The binding domains described herein can be used to transfer a linker (such as a linker comprising a barcode) to a target nucleic acid. Therefore, in some embodiments, the binding domains described herein can be used to transfer a barcode to a target nucleic acid. The barcode can be an MBC (i.e., a barcode that specifically corresponds to a histone modification) and is bound to a target DNA containing a nucleosome of a histone modification or a DNA binding protein. In this article, the target nucleic acid to which the linker has been transferred is referred to as a "labeled target nucleic acid," "labeled target," or similar terms. In this article, the target nucleic acid to which the barcode has been transferred is referred to as a "barcode-labeled target nucleic acid," "barcode-labeled target," or similar terms. In this article, the reaction in which the linker is transferred to the target nucleic acid is referred to as a "linker transfer reaction." Similarly, in this article, the reaction in which the barcode is transferred to the target nucleic acid is referred to as a "barcode transfer reaction."
[0158] In some aspects, the barcode is transferred to the target nucleic acid by enzyme-catalyzed transfer, such as by single-stranded connection, splicing connection, primer extension or double-stranded flat end or sticky end connection. In some aspects, the disclosure includes connecting the universal nucleic acid sequence to the 3' end, 5' end or both ends of the target DNA. In some aspects, the disclosure includes adding multiple single types of nucleotide tail sequences at the 3' end of the target DNA by enzyme catalysis. In some aspects, enzyme-catalyzed tailing is performed with terminal nucleotidyl transferase. In some aspects, the 3' end of the joint is hybridized with the 3' end of the target DNA. In some aspects, modification-specific barcodes are introduced, wherein one or two 3' ends are extended by DNA polymerase. In some aspects, a joint carrying 3' degenerate bases randomly starts the target DNA, and one or two 3' ends are extended by DNA polymerase to introduce modification-specific barcodes.
[0159] The purpose of adapter / barcode transfer is to covalently attach the adapter / barcode to the target nucleic acid molecule. For example, in some embodiments, the barcode is transferred to the target nucleic acid by covalently coupling the barcode to the 5' end or 3' end of the target nucleic acid. In some embodiments, the barcode is transferred to the target nucleic acid by covalently coupling the barcode or its complement to the 5' end or 3' end of the target nucleic acid. In some embodiments, the labeled / barcoded nucleic acid molecule can be sequenced in a downstream step. In some embodiments, a copy of the labeled target nucleic acid can be sequenced. Figure 4 、 Figures 7A-7B and Figure 8 Examples of adapter / barcode transfer reactions are provided.
[0160] The adapter / barcode can be transferred to the target DNA using one or more DNA ligases, such as T4 DNA ligase, CircLigase, T3 DNA ligase, T7 DNA ligase, 9°N DNA ligase, Taq DNA ligase, or E. coli DNA ligase. 9°N DNA ligase is a DNA ligase that catalyzes the formation of a phosphodiester bond between the juxtaposed 5' phosphate and 3' hydroxyl termini of two adjacent oligonucleotides hybridized to complementary target DNA.
[0161] Splint ligation can also be used to transfer adapters / barcodes to target nucleic acids. In splint ligation, two nucleic acid molecules are brought into proximity using a bridging DNA and then ligated using one or more enzymes. Splint DNA ligation can be performed using enzymes such as T4 DNA ligase, T3 DNA ligase, T7 DNA ligase, or E. coli DNA ligase.
[0162] Alternatively, a double-stranded ligation method can be used to transfer the adapter / barcode to the target nucleic acid. In some embodiments, the target nucleic acid molecule can be double-stranded DNA, which can have blunt ends or sticky ends. The blunt ends or sticky ends of the double-stranded DNA can be ligated by T4 ligase, T3 ligase, T7 ligase, or E. coli ligase.
[0163] In some embodiments, chemical ligation can also be used to transfer the adapter / barcode to the target nucleic acid.
[0164] Methods for preventing or reducing intercomplex linker / barcode transfer via spatial isolation
[0165] By spatially isolating the molecules involved in the reaction, the transfer of linkers / barcodes within the complex can be facilitated. Specifically, by separating the complexes comprising the target nucleic acid, binding domain, and linker from each other, the transfer of barcodes between complexes (i.e., linker / barcode transfer between complexes) is hindered. This detection configuration significantly improves the fidelity of barcode labeling.
[0166] Barcode transfer
[0167] Each binding domain specifically binds to the target, bringing the nucleic acid adapter close to the 3' end or 5' end of the target nucleic acid. Subsequently, the adapter (e.g., an adapter comprising or consisting of a barcode) can be transferred to the target nucleic acid. In some embodiments, the transfer process is performed in an environment that significantly prevents off-target barcode-labeled nucleic acids. For example, such an environment can be one in which target nucleic acids cannot interact with each other (i.e., each target nucleic acid can only interact with one binding domain). For example, this can be achieved by performing the barcode transfer reaction in a very low concentration solution, or by immobilizing the target nucleic acid or binding domain on a substrate to achieve spatial isolation. In some embodiments, the transfer process is performed by replicating the target nucleic acid to generate a labeled / barcoded target nucleic acid copy. For example, when the barcode is transferred to or near the target nucleic acid, a primer extension can be used to generate a barcoded copy of the target nucleic acid. In some embodiments, the barcode transfer can be performed in an environment in which the amount of off-target barcoded DNA generated is less than 20% of the total amount of barcoded target DNA. In some embodiments, the amount of off-target barcoded DNA produced is less than 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, or 1% of the total amount of barcoded target DNA. In some embodiments, when the amount of off-target barcoded DNA produced (relative to the total amount of barcoded target DNA) is less than any of the aforementioned percentage ranges, the environment is an environment capable of achieving spatial isolation. In some embodiments, as described in detail below, an environment is provided in which the amount of off-target barcoded DNA produced (relative to the total amount of barcoded target DNA) is less than any of the aforementioned percentage ranges by immobilizing the target nucleic acid, binding domain, linker, and one or more components of barcode transfer on a substrate. In some embodiments, the amount of off-target barcoded DNA generated (relative to the total amount of barcoded target DNA) is below any of the aforementioned percentage ranges by immobilizing multiple copies of the adapter on a substrate at a specific density or density range, as described below.
[0168] Above and Figures 7A-7B Barcode transfer reactions and spatial segregation are detailed.
[0169] Barcode transfer can be performed in a variety of different environments that achieve spatial isolation. For example, spatial isolation can be achieved by highly diluting the complex containing the binding domain that binds to the target in a solution. It is necessary to ensure that the degree of dilution of the solution is such that any complex described herein containing the binding domain that binds to the target can achieve spatial isolation. This spatial isolation promotes barcode transfer within the complex and substantially prevents barcode transfer between binding domain complexes. In some embodiments, the concentration of the complex in the dilute solution is less than 1000nM, less than 500nM, less than 100nM, less than 10nM, less than 1nM, less than 0.1nM, less than 0.01nM or less than 0.001nM.
[0170] In some embodiments, spatial isolation can be achieved by substrate fixation. For example, the binding domains described herein can be fixed to a substrate by coupling. Each substrate can contain only one type of binding domain, or it can contain at least two, at least three, at least four, at least five or more types of binding domains. Each "type" of binding domain binds to a different histone modification or DNA binding protein, and / or contains a different barcode. In some embodiments, the first binding domain and the second binding domain on the substrate surface are spatially separated. By regulating the surface binding capacity and arrangement, absolute or relative quantitative analysis of target molecules and modifications can be achieved.
[0171] For example, the exemplary matrix of binding domain, joint, intermediate protein, connector and tether can be coupled, including for example microbeads, chips, plates, slides, culture dishes or three-dimensional matrix. In some embodiments, the matrix is a resin, film, fiber or polymer. In some embodiments, the matrix is microbeads, for example, microbeads comprising agarose (sepharose), agarose (agarose), cellulose, polystyrene, polymethacrylate and / or polyacrylamide. In some embodiments, the matrix is magnetic beads. In some embodiments, the support is a polymer, such as a synthetic polymer. Non-limiting examples of synthetic polymers include: polystyrene, polyethylene glycol, polyisocyanopeptide polymers, polylactic acid-co-glycolic acid, poly-ε-caprolactone (PCL), polylactic acid, poly (3-hydroxybutyrate-co-3-hydroxyvalerate) (PHBV), chitosan and cellulose.
[0172] The binding domain can be directly coupled to the surface of the matrix. For example, the molecule can be directly coupled to the matrix through one or more covalent bonds or non-covalent bonds. In embodiments where the matrix is a 3D matrix or other 3D structure, the binding domain can be coupled to multiple surfaces of the matrix.
[0173] In some embodiments, the binding domain can be indirectly coupled to the surface of the matrix. For example, the binding domain can be indirectly coupled to the surface of the matrix through a capture molecule, wherein the capture molecule is directly coupled to the matrix. The capture molecule can be any nucleic acid, protein, sugar, chemical linker, etc. that can simultaneously bind or connect the matrix and the binding domain and / or the target nucleic acid. In some embodiments, the capture molecule is bound to the binding domain. In some embodiments, the capture molecule is bound to the joint of the binding domain or the binding domain (such as a connector to a joint). In some embodiments, the capture molecule is bound to the target nucleic acid. For example, in some embodiments, the capture molecule can be bound to the poly A tail of the target nucleic acid or be bound to a specific nucleic acid sequence.
[0174] In some embodiments, target nucleic acids can be directly coupled to the surface of a substrate via reactive chemical groups. For example, nucleic acid targets can be modified with azide groups and reacted with alkyne-modified microbeads via copper-catalyzed click chemistry. Other examples include trans-cyclooctene (TCO) / methyltetrazine and DBCO / azide.
[0175] In some embodiments, on the surface of the matrix, the first binding domain is separated from the second binding domain to ensure that each binding domain can only interact with one target nucleic acid. In some embodiments, the spacing between the first binding domain and the second binding domain is at least 50 nm. For example, the spacing between the first binding domain and the second binding domain can be from about 50 nm to about 500 nm, such as from about 50 nm to about 100 nm, from about 100 nm to about 150 nm, from about 150 nm to about 200 nm, from about 200 nm to about 250 nm, from about 250 nm to about 300 nm, from about 300 nm to about 350 nm, from about 350 nm to about 400 nm, from about 400 nm to about 450 nm, or from about 450 nm to about 500 nm. In some embodiments, the spacing between the first binding domain and the second binding domain can be greater than about 500 nm.
[0176] In some embodiments, multiple copies of the linker are coupled to the substrate at a density of about 1 linker / 5 nm. 2 to approximately 1 junction / 50nm 2 , for example 1 connector / 20nm 2 In some embodiments, multiple copies of a binding domain are coupled to a matrix at a density of about 1 binding domain / 1000 nm 2 About 1 binding domain / 15000nm 2 , for example 1 binding domain / 8000nm 2 .
[0177] In general, the purpose of coupling the binding domain (or target nucleic acid) to the matrix is to ensure that the joint and / or barcode are transferred within the complex. A matrix comprising two or more spatially separated binding domains can be prepared using methods known to those skilled in the art. For all purposes, the disclosures of the following publications are incorporated herein by reference in their entirety: US20210237022A1, US20220010367A1, US20220364163A1, US20220298560, US11,519,033, US20210010070.
[0178] Binding domain coupled to substrate
[0179] In some embodiments, the binding domain is directly or indirectly coupled to the matrix. In some embodiments, multiple binding domains are fixed to the matrix using site-specific chemistry. For example, in some embodiments, the binding domain comprises a site that enables it to be fixed on the matrix. Fusion of an autocatalytic protein tag (such as Spy capture protein, transpeptidase sortase A, SNAP tag, Halo tag, and CLIP tag) to the end of the binding domain can promote the coupling of the binding domain to the surface of the matrix. Subsequently, these protein tags on the binding domain can covalently react with the homologous active groups on the matrix surface. For example, the Spy capture protein can be designed into the binding domain. The Spy tag forms a covalent bond with the Spy tag protein (13aa polypeptide). When the Spy tag is coupled to the matrix surface, the reaction between the binding domain connected to the Spy capture protein and the Spy tag will cause the binding domain to be covalently linked to the matrix. Similarly, the transpeptidase Sortase A tag can be fused to the binding domain, and the Sortase A tag can be used to react with five glycine coupled to the matrix surface. As another example, a SNAP tag can be fused to the binding domain, and the SNAP tag can be used to react with O coupled to the matrix surface. 6 In some embodiments, a CLIP tag can be fused to a binding domain and can be used to bind to an O-binding domain coupled to a substrate surface. 2 In some embodiments, a Halo tag can be fused to the binding domain, and the Halo tag can be used to react with alkyl halides on the surface of the substrate.
[0180] In some embodiments, the binding domain may comprise biotin. The binding molecule may be immobilized to the substrate surface via a capture molecule that binds to biotin (such as avidin, streptavidin, or neutravidin).
[0181] Figure 5AThe binding domain is shown coupled to a matrix or surface via a tether. In some embodiments, multiple binding domains can be directly or indirectly fixed to a matrix by site-specific chemistry. For example, in some embodiments, the binding domain of a binding domain can include a site for fixing it to a matrix, as well as a site for fixing a DNA linker. By fusing an autocatalytic protein tag (such as Spy capture protein, transpeptidase sortase A, SNAP tag, Halo tag, and CLIP tag) to the end of the binding domain, the coupling of the binding domain to the matrix surface can be facilitated. The SNAP tag is derived from human O 6 -alkylguanine-DNA alkyltransferase self-labeling protein. SNAP tag with O 6 -benzylguanine derivatives (such as guanine or chloropyrimidine conjugated to fluorescent dyes) undergo covalent reaction. CLIP tag is an improved version of SNAP tag. It is also a human O 6 -Self-labeling protein of alkylguanine-DNA alkyltransferase. CLIP tags are designed to react with benzylcytosine derivatives, but not with benzylguanine derivatives. These protein tags on the binding domain can then react covalently with homologous reaction molecules on the substrate surface. For example, the Spy capture protein can be designed as a binding domain. The Spy tag forms a covalent link with the Spy tag protein (13aa polypeptide). When the Spy tag is coupled to the substrate surface, the reaction between the binding domain connected to the Spy capture protein and the Spy tag will cause the binding domain to be covalently linked to the substrate. Similarly, the transpeptidase Sortase A tag can be fused to the binding domain, and the transpeptidase Sortase A tag can be used for the five glycine reactions coupled to the substrate surface. As another example, the SNAP tag can be fused to the binding domain, and the SNAP tag can be used for the O-coupled to the substrate surface. 6 -benzylguanine reaction. In some embodiments, the CLIP tag can be fused to the binding domain, and the CLIP tag can be used to bind to the substrate surface. 2 In some embodiments, a Halo tag can be fused to the binding domain, and the Halo tag can be used to react with alkyl halides on the substrate surface.
[0182] In some embodiments, the binding molecule may comprise biotin. Such binding molecules may be immobilized to the substrate surface via capture molecules that bind to biotin (e.g., avidin, streptavidin, or neutravidin).
[0183] Target nucleic acid or linker coupled to the matrix
[0184] In some embodiments, the composition herein includes a matrix. In some embodiments, the composition herein includes two or more matrices. In some embodiments, the composition includes multiple matrices, wherein each matrix is composed of the same material. In some embodiments, the composition includes multiple matrices, wherein each matrix is composed of different materials. In some embodiments, the matrix is a microbead, a chip, a plate, a tube, a slide, a culture dish, a gel or a three-dimensional polymer matrix. The matrix can be composed of a variety of materials. In some embodiments, the matrix is a resin, a film, a fiber or a polymer. In some embodiments, the matrix includes agarose (sepharose), agarose (agarose), cellulose, polystyrene, polymethacrylate and / or polyacrylamide. In some embodiments, the matrix includes a polymer, such as a synthetic polymer. Non-limiting examples of synthetic polymers include: polyethylene glycol, polyisocyanate polymers, polylactic acid-co-glycolic acid, poly-ε-caprolactone (PCL), polylactic acid, poly (3-hydroxybutyrate-co-3-hydroxyvalerate) (PHBV), chitosan and cellulose.
[0185] In some embodiments, the matrix can be modified with oligonucleotide capture molecules that hybridize with the characteristic sequence of the target nucleic acid. For example, the poly dA tail added to the nucleosomal DNA by terminal nucleotidyl transferase can be captured by hybridization of capture molecules comprising poly dT oligonucleotides or gene-specific sequences. In some embodiments, the capture molecules exist with a lower matrix density, thereby physically isolating the binding domain. In some embodiments, in the state of matrix binding (i.e., when the target nucleic acid is coupled to the matrix), the barcode can be transferred from the ribosome or protein binding complex to the target nucleic acid.
[0186] Microbeads for hybrid capture of target nucleic acids can be prepared by directly coupling 5' amino-modified oligonucleotides to matrix-activated microbeads. Matrix-activated microbeads can carry epoxy, tosyl, carboxylic acid, or amino functional groups for covalent attachment. Carboxyl microbeads are typically allowed or induced to react with carbodiimide to promote peptide bond formation, while amino microbeads typically require a bifunctional NHS linker. In some embodiments, the surface of the microbeads is passivated to prevent nonspecific binding. In some embodiments, passivation can be achieved by co-grafting with polyethylene glycol (PEG) molecules with the same attachment chemistry. For example, using 5' amino-modified oligonucleotides and amino-terminal polyethylene glycol (PEG), most sites on the matrix can be evenly occupied by PEG molecules, thereby achieving a spatially dispersed distribution of the oligonucleotides. When an excess of PEG is used, average spatial separation between the oligonucleotide molecules will be formed. By adjusting the ratio between oligonucleotide and PEG molecules, the surface density of the capture molecules can be controlled.
[0187] In some embodiments, the microbeads are agarose microbeads prepared using mTet (tetrazine) and carboxyl-PEG. By reducing the ratio of mTet to carboxyl-PEG, crosstalk between target nucleic acid molecules is reduced. In some embodiments, the ratio of mTet:carboxyl-PEG is 1:500, 1:600, 1:700, 1:800, 1:900, 1:1000, 1:1100, 1:1200, 1:1300, 1:1400, 1:500, 1:1000, 1:2000, 1:3000, 1:4000, 1:5000, 1:6000, 1:7000, 1:8000, 1:9000, or 1:10000. In some embodiments, the ratio of mTet:carboxyl-PEG is 1:1000.
[0188] In some embodiments, the matrix comprises a plurality of identical or different binding domains. In some embodiments, the matrix comprises a plurality of identical or different linkers.
[0189] Nucleosome-binding conjugates
[0190] Also provided herein are nucleosome binding conjugates comprising a binding domain coupled to a linker. In some embodiments, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 linkers are coupled to the binding domain. In some embodiments, the binding domain and linker comprise any binding domain or linker described in any of the preceding paragraphs.
[0191] Nucleic acid analysis methods
[0192] Nucleosome binding conjugates as described herein can be used for barcode transfer within the complex as described above, and can be used for a variety of nucleic acid analysis methods, particularly suitable for identifying histone modifications or DNA binding proteins. Therefore, the present disclosure provides methods for analyzing histone modifications, including analytical methods for analyzing the multiple modifications of histones and nucleosomes, and analytical methods for DNA binding proteins. In these methods, histone modifications or DNA binding proteins can be identified by a binding domain. Subsequently, a linker or its components (such as a barcode) are transferred from the binding domain to the target nucleic acid (i.e., a target nucleic acid labeled with a barcode). Since the barcode specifically corresponds to a specific histone modification, this step can write the information of the recognition event into the nucleic acid sequence of the target nucleic acid. Then, the obtained barcode-labeled target nucleic acid is converted into a sequencing library and read by a nucleic acid sequencing method. This step reveals a barcode sequence, which is related to histone modifications or DNA binding proteins. Sequencing can also determine the position of histone modifications or the binding site of DNA binding proteins. The high-throughput analysis method described herein can simultaneously identify the properties and positions of multiple or all nucleosome modifications and DNA binding proteins.
[0193] The methods described herein and shown in the accompanying drawings include a series of steps, as described below. As will be appreciated by those skilled in the art, in some embodiments, various steps may be omitted and / or performed in a different order.
[0194] Contacting the binding domain with the target nucleic acid
[0195] In some embodiments, the methods described herein comprise contacting one or more binding domains with a target (e.g., one or more target nucleic acids, or one or more histone modifications and DNA binding proteins). For example, the target nucleic acid can be chromatin or nucleosomal nucleic acid isolated from a biological cell or tissue. In some embodiments, the binding domain is contacted with a DNA binding protein described herein.
[0196] The contacting of the binding domain with the target can be carried out in solution. For example, a complex comprising one or more target nucleic acids or DNA binding proteins can be contacted with a complex comprising one or more binding domains. In some embodiments, the contacting can be carried out in a dilute solution so that each target can interact with only one binding domain.
[0197] In some embodiments, the contact occurs on substrate / surface. For example, one or more targets can be coupled to substrate / surface, and one or more binding domains can contact with the target nucleic acid that is coupled to substrate / surface. In some embodiments, one or more binding domains can be coupled to substrate / surface, and one or more targets can contact with the binding domain that is coupled to substrate / surface.
[0198] The target nucleic acid or DNA binding protein can be contacted with only one binding domain protein (i.e., to detect only one histone modification or one DNA binding protein); or in some embodiments, the target nucleic acid can be contacted with more than one binding domain to detect multiple histone modifications. For example, the target nucleic acid can be contacted with at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten, or more different types of binding domains.
[0199] In some embodiments, a target is first contacted with a first binding domain pool and then with a second binding domain pool. In some embodiments, the binding domain pools can comprise different types of binding domains (i.e., recognize different types of modifications or proteins). In some embodiments, each binding domain pool can comprise 1-5, 5-10, 10-25, 25-50, 50-100, 100-150, 150-175, 175-200, 250, 300, 350, 400 or more different types of binding domains.
[0200] Barcode transfer
[0201] Each binding domain can specifically bind to the target, bringing the connector close to the 3' end or 5' end of the target nucleic acid. Then, the connector (such as a connector comprising a barcode or consisting of a barcode) can be transferred to the target nucleic acid. In some embodiments, the environment in which the transfer process occurs can substantially avoid the generation of barcode-labeled nucleic acid off-target. For example, the environment can be an environment in which the target nucleic acids cannot interact with each other (i.e., each target can only interact with one binding domain). For example, it can be achieved by carrying out a barcode transfer reaction in a very dilute solution, or by fixing the target nucleic acid or the binding domain on a matrix to achieve its spatial separation. In some embodiments, the transfer process is carried out by replicating the target nucleic acid to generate a labeled / barcode-labeled copy of the target nucleic acid. For example, when the barcode is transferred to the target nucleic acid, a barcode-labeled target nucleic acid copy can be generated by polymerase chain reaction (PCR).
[0202] Above and Figures 7A-7B Barcode transfer reactions and spatial separation are detailed.
[0203] Amplification and sequencing
[0204] After the target nucleic acid is barcoded, it can be amplified and sequenced. This step reveals the barcode sequence, which correlates with the histone modification to which the binding domain initially binds in the target nucleic acid. Sequencing reveals the sequence and length of the DNA fragment, allowing the location of the histone modification to be determined. Sequencing can also reveal mutations near the histone modification, providing information about its location.
[0205] Thus, in some embodiments, the methods described herein can include a step of sequencing the barcoded target nucleic acid or a copy thereof. The sequencing step can be performed using any suitable method known in the art. For example, sequencing can be performed using a next-generation sequencing (NGS) method, a massively parallel sequencing method, or a deep sequencing method. The methods disclosed herein can be applied to a variety of NGS platforms. For example, Sequencing adopts the principle of synthetic sequencing. After the blocked fluorescent-labeled nucleotides undergo binding, imaging and deblocking steps, the next fluorescent nucleotide is inserted. 454 sequencing is based on pyrosequencing technology, which uses fluorescence to detect the release of pyrophosphate after the polymerase incorporates nucleotides into the new DNA chain. (Proton / PGM sequencing) detects protons (H) released directly when DNA polymerase incorporates a single nucleotide + ). Nanopore sequencing detects changes in current caused by the bases of nucleic acid molecules passing through the pore one by one. Single-molecule real-time (SMRT) sequencing technology measures the residence time of fluorescently labeled nucleotides when DNA polymerase (immobilized at the bottom of a zero-mode waveguide) incorporates them into DNA.
[0206] In some embodiments, detection of target nucleic acid does not require sequencing. For example, target nucleic acid can be detected by PCR. For example, whether a target nucleic acid (such as a barcode) is detected by PCR. In some embodiments, a fluorescent probe (such as a fluorescently labeled hybridization probe) is used to detect the target nucleic acid. In some embodiments, a microarray or other nucleic acid array is used to detect the target nucleic acid. The method of analyzing the sequencing results or data derived from any method for detecting target nucleic acid as described herein is well known to those skilled in the art. For example, standard bioinformatics methods are used to analyze the sequencing results.
[0207] In some embodiments, the barcodes added by the binding domain-mediated reaction can be detected without sequencing. For example, the barcodes can be detected by nucleic acid electrophoresis, fluorescent hybridization probes, PC, or any other nucleic acid amplification method that can be triggered by the barcodes, thereby confirming the presence of the histone modification.
[0208] Exemplary methods for identification and / or localization of histone modifications
[0209] In some embodiments, the assay beads display a modified specific antibody and a forward linker comprising a 3' end, a blocked 3' end, and a 5' phosphate. Figure 5A In some embodiments, the target DNA of the nucleosome can be end-repaired before immunoprecipitation. In some embodiments, histone modifications can be identified by attaching a forward linker to the target DNA during or after immunoprecipitation. In specific embodiments, as Figure 5A Because the other forward linker strand has a 3' blocking group or the target DNA lacks a 5' phosphate, only the fixed strand of the linker is ligated. The barcoded DNA is then extended and replicated by DNA polymerase. Figure 5A The final step shown is the ligation of the reverse linker. In some embodiments, multiple histone targets can be detected in the same reaction using a combination of multiple types of microbeads, each with a specific barcoded linker and modification-specific antibody (see Figures 3A-3B ).like Figure 5A As shown, two forward adapters can be attached to the matrix. The forward adapters can include UFPs, UMIs, and MBCs and are attached to target DNA containing histone-modified nucleosomes. After ligation, a denaturation step is performed to remove the chromatin core, followed by reverse strand synthesis of the ligated target DNA to form forward and reverse strands. Next, the reverse adapter is ligated to the forward and reverse strands of the target DNA. The barcoded target DNA is amplified and analyzed by sequencing.
[0210] In some embodiments, the assay beads display a modification-specific antibody and a surface linker comprising an MBC and a uracil recognizable by an UdG / endonuclease. Figure 5B As shown, before immunoprecipitation, the nucleosome target DNA is end-repaired and 5' phosphorylated. During or after immunoprecipitation, a forward adapter is ligated to the target DNA, and the first histone modification is identified by the addition of a first MBC. Next, an enzyme cocktail containing UdG and an endonuclease (such as, but not limited to, endonuclease VIII) is used to cleave the adapter at the position of uracil, thereby releasing the barcoded nucleosome. The second histone modification is detected by repeating the above steps, this time using a different binding domain combination and introducing a second MBC. The final step is the ligation of a Y-type sequencing adapter. Figure 5B The reaction scheme shown allows the use of multiple types of microbeads and their associated barcodes in each encoding cycle. In some embodiments, the release step includes adding a buffer selected from an antigen elution buffer, a histone or antibody replacement mixture, an acidic buffer with a pH of 6.5 or less, or an alkaline buffer with a pH of 8.5 or above. The elution buffer can include a high salt solution that can effectively remove affinity while maintaining the activity of the antibody and antigen. The histone replacement mixture can include an excess of histones or peptides carrying specific modifications. The antibody replacement mixture can also include an excess of synthetic modified histone peptides that act as competitors to separate the binding domain from the nucleosome. The buffer can include a reducing agent (DTT and / or TCEP) to cleave antibody disulfide bonds, an enzyme specifically for digesting antibodies (papain and / or pepsin), a surfactant (SDS, sodium deoxycholate), an acidic buffer with a pH of 6.5 (typically glycine-HCl, pH 2.5-3.0) or less, or an alkaline buffer with a pH of 8.5 or above.
[0211] Figure 6 and Figure 9 Co-localization of histone modifications by sequential barcoding of the solution is shown. In some embodiments, repeated cycles of IP and barcoding can be used to identify multiple histone modifications on the same nucleosome. In some embodiments, the MBC is not attached to a surface or matrix. In some embodiments, as Figure 9As shown, the MBC is attached to a cleavable circular region containing a unique molecular identifier (UMI). This configuration allows the MBC to be attached near the UMI in each barcoding cycle. Since there is only one type of microbead in each IP cycle, the MBC does not need to be attached to the matrix. In some embodiments, the UMI is attached to the MBC. In each cycle, the nucleosomes are immunoprecipitated, washed, and barcoded by the MBC in the ligation solution. In the next encoding cycle, the nucleosomes are dissociated from the matrix and mixed with the supernatant of the previous IP cycle before the next encoding cycle. In the final step, sequencing adapters are attached to both ends of the nucleosomes to generate a sequencing library. In some embodiments, the sequencing adapter is Y-shaped or bell-shaped. In some embodiments, the sequencing adapter contains a UMI. The method produces histone modification marker MBC tails at the ends of the nucleosomes.
[0212] In some embodiments, the detection of multiple histone modifications can include barcoding both nucleic acid ends of the nucleosome. Figure 7A As shown in the figure, after nucleosome end repair and adenylation, it is incubated with multiple nucleosome-binding conjugates, each containing a binding domain and a modified barcode (MBC). In the presence of a ligase, the encoding reaction can be initiated, transferring one or two MBCs to the target DNA. When only one adapter is transferred during the encoding step, an additional capping step using free adapters is performed to obtain an amplifiable library.
[0213] In some embodiments, multiple histone modifications are detected by sequential barcoding reactions of nucleosomes attached to a substrate. Figure 7B As shown, in the first step, nucleosomes are anchored to the matrix at a single-molecule distance to prevent interaction between adjacent nucleosomes. To identify the first histone modification, a barcode-tagged antibody is introduced. After washing, a linking reagent is added and the antibody barcode is attached to the free end of the nucleosome. The barcode is cut using a restriction enzyme, releasing the antibody and generating sticky barcode ends for the next encoding. The steps can be repeated any number of times, but a barcode-antibody conjugate is always added. The final capping step introduces a reverse sequencing adapter and is independent of the antibody. Ultimately, a nucleosome containing a string of barcodes is generated, each barcode marking a modification.
[0214] In some embodiments, colocalization of histone modifications can be determined by proximal ligation. In some embodiments, a proximal positioning barcode is annealed to a bridging splint oligonucleotide, followed by ligation of the proximal positioning barcode. In some embodiments, a nucleosome can hybridize to the poly-T end of the barcode via the A tail. Figure 8As shown, nucleosomes with A tails are incubated with a mixture of barcode-antibody conjugates. When the antibodies bind to their targets, adjacent barcodes are bridged and connected by the splint oligonucleotide. The nucleosome's A tail serves as a primer for the tandem barcodes, and DNA polymerase is added to generate copies. This process ultimately results in a nucleosome carrying a string of barcodes, each identifying a different modification.
[0215] In some embodiments, histone modifications in tissues can be analyzed by fixing nucleosome binding conjugates containing spatial identifiers to microarray slides and layering the microarray with the tissue. The tissue can be a fresh frozen tissue section or formalin-fixed paraffin embedded (FFPE) tissue. The cells are permeabilized with a surfactant (such as digitonin, TritonX or NP40), and then the chromatin is sheared with an enzyme to release the nucleosomes (such as micrococcal nuclease (MNase) or DNA enzyme). The nucleosomes diffuse out of the cell and are captured by the fixed binding domain. In the final step, the spatial identifier is transferred to the nucleosome by ligation ( Figures 15A-15B ), generating nucleosomes carrying modified barcodes and spatial identifiers.
[0216] The methods described herein can be used for the diagnosis of a disease, disorder or illness. For example, in some embodiments, the method can be used to diagnose cancer in a subject in need thereof. In some embodiments, the kit can be used to monitor changes in a disease, disorder or illness over time, such as a response to one or more treatments. For example, the kit can be used to monitor epigenetic changes over time in a subject who is receiving cancer treatment (such as chemotherapy, radiotherapy, etc.). In some embodiments, the method can be used to analyze cells or tissues of a subject in need thereof. For example, the method can be used to detect histone modifications in cells or tissues isolated from a blood sample, a biopsy sample, an autopsy sample, etc.
[0217] In some embodiments, the nucleosomes can be obtained as cell-free circulating nucleosomes. For example, the cell-free circulating nucleosomes can be obtained from the patient's blood, or from the extracellular tumor environment or microenvironment.
[0218] In some embodiments, nucleosomes can be obtained from a single cell. In some embodiments, nucleosomes can be obtained from a single isolated cell. In some embodiments, nucleosomes can be obtained from multiple clones derived from a single cell.
[0219] In some embodiments, the present disclosure provides a method for diagnosing cancer or cancer subtypes associated with one or more histone modifications, comprising analyzing multiple nucleosomes according to any numbering aspect. In some embodiments, the present disclosure provides a method for monitoring cancer progression or treatment response, comprising analyzing multiple nucleosomes according to any numbering aspect. In some embodiments, multiple nucleosomes derived from a patient's blood sample are analyzed. In some embodiments, multiple nucleosomes derived from a patient's tissue biopsy sample are analyzed. In some embodiments, the present disclosure provides a kit for monitoring epigenetic changes over time in a subject undergoing cancer treatment, the kit comprising any composition or nucleosome binding conjugate disclosed herein.
[0220] Because histone modifications on cell-free nucleosomes can reflect DNA-related activity within their cells of origin, the present disclosure includes methods for using histone modifications on cell-free nucleosomes as biomarkers for plasma liquid biopsies. The present disclosure includes applications for detecting multiple recombinant protein modifications in low sample input conditions, such as analyzing cell-free nucleosomes in plasma, where only 20-60 ng of nucleosomes are present per milliliter of plasma.
[0221] In some embodiments, the method can be used to detect and / or monitor epigenetic changes in cells used commercially to produce one or more products (e.g., cells used in industrial fermentation). In some embodiments, the method can be used to detect and / or monitor epigenetic changes in plant cells or tissues.
[0222] Compositions comprising binding domains
[0223] Also provided herein are compositions comprising one or more binding domains of the present disclosure. In some embodiments, the composition comprises one or more types of binding domains. For example, the composition may comprise a first binding domain that binds to a first histone modification or a first DNA binding protein, and a second binding domain that binds to a second histone modification or a second DNA binding protein. In some embodiments, the composition may comprise 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24 or 25 or more different types of binding domains.
[0224] Also provided herein are compositions comprising one or more complexes, wherein each complex comprises a binding domain that binds to a target nucleic acid.
[0225] In some embodiments, the compositions described herein comprise one or more carriers, excipients, buffers, and the like. The pH of the composition can be about 0.5, about 1.0, about 1.5, about 2.0, about 2.5, about 3.0, about 3.5, about 4.0, about 4.5, about 5.0, about 5.5, about 6.0, about 6.5, about 7.0, about 7.5, about 8.0, about 8.5, about 9.0, about 9.5, about 10.0, about 10.5, about 11.0, about 11.5, about 12.0, about 12.5, about 13.0, about 13.5, or about 14.0. In some aspects, the pH of the composition can be 2-12, 3-11, 4-10, 5-9, 6-8, or 6.5-7.5, or any range within these ranges. In some embodiments, the composition is a pharmaceutical composition. In some embodiments, the composition is a diagnostic composition.
[0226] Kits for analyzing histone modifications
[0227] The binding domains described herein can be provided in the form of a kit (e.g., as a component of a kit). For example, the kit can comprise a binding domain or one or more components thereof, and informational materials. The kit can also comprise any reagents and materials required for performing the detection methods described in the present disclosure (including the embodiments). These reagents and materials can include connectors, matrices (microbeads), and enzymes (ligases, polymerases). The informational materials can be, for example, explanatory materials, instructional materials, sales materials, or other materials related to the methods described herein and / or the application of the binding domains. The form of the informational material of the kit is not limited. In some embodiments, the informational material can include information about the production, molecular weight, concentration, expiration date, batch, or place of production of the binding domain. In some embodiments, the informational material can include a list of conditions and / or patients that can be diagnosed or evaluated using the kit.
[0228] In some embodiments, the binding domain can be provided in an appropriate manner (e.g., in a tube for ease of use, at a suitable concentration, etc.) for use in the methods described herein. In some embodiments, the binding domain in the kit may require some preparation or processing before use. In some embodiments, the binding domain is provided in a liquid, dried, or lyophilized form. In some embodiments, the binding domain is provided in an aqueous solution. In some embodiments, the binding domain is provided in a sterile, nuclease-free solution. In some embodiments, the binding domain is provided in the form of a composition that is substantially free of any other nucleic acid components, except for the nucleic acid that may contain the molecule itself.
[0229] In some embodiments, the kit may comprise one or more syringes, test tubes, ampoules, foil packages, or blister packs. The container of the kit may be sealed, waterproof (i.e., to prevent humidity changes or liquid evaporation), and / or comprise a light-shielding layer.
[0230] In some embodiments, the kit can be used to perform one or more methods described herein, such as methods for analyzing a target nucleic acid population. In some embodiments, the kit can be used to diagnose a disease, condition, or patient. For example, in some embodiments, the kit can be used to diagnose cancer. In some embodiments, the kit can be used to monitor changes in a disease, condition, or patient over time, such as response to one or more treatments. For example, the kit can be used to monitor epigenetic changes over time in a subject undergoing cancer treatment. DETAILED DESCRIPTION
[0231] The following non-limiting examples further illustrate embodiments of the compositions and methods of the present disclosure.
[0232] Example 1: Preparation of microbead matrix for nucleosome modification-specific barcode labeling
[0233] Magnetic beads are a convenient matrix for library preparation workflows because they simplify buffer exchange and purification steps. This example describes the co-loading of antibodies (Abs) and linkers containing modified barcodes (MBCs) onto magnetic beads. Each bead type is loaded with one antibody and one linker in an optimized ratio. Multiple bead types can be combined into bead pools to detect any number of histone modifications ( Figure 3A ). The microbeads are intended to be used Figure 5A The workflow shown is for the analysis of histone modifications on multiple nucleosomes. The bead loading protocol can be easily adapted to the specific nucleosomes by using different linker sequences. Figure 4 and Figure 5B Two types of beads were prepared, one targeting H3K4me3 and the other targeting H3K4me2.
[0234] Two microbead loading mixtures were prepared, each containing a 3' biotinylated linker, biotinylated protein G, and an antibody targeting a specific histone modification in a molar ratio of 3:6:4, mixed with biotinylated small molecule PEG in HBST300 buffer (10mM HEPES pH7.6, 300mM NaCl, 0.1mM EDTA, 0.05% Tween 20). The first microbead loading mixture contained protein G, Ab42 (histone H3K4me3 antibody, EpiCypher, product number #13-0041), and rcMBC101 / 5-terminal phosphate (Phos) / GTGATCGNNNNNCTGTCTCTTATACACATCTGACUTTTTT (SEQ ID NO: 1) / 3BioTEG / . The second microbead loading mixture contained protein G, Ab70 (histone H3K4me2 antibody, ThermoFisher Scientific, product number #MA5-33383) and rcMBC103 ( / 5-terminal phosphate (Phos) / CCGCATTNNNNNCTGTCTCTTATACACATCTGACUTTTTT (SEQ ID NO: 2) / 3BioTEG / . For the input control, the loading mixture contained protein G, Ab67 (histone H3 antibody, Thermo Fisher product number #39064) and rcMBC103. rcMBC includes 7 base MBC (italics), 5b UMI (N), 22b Illumina The bead-loading mixture consists of a P5 linker, a uracil for cleavage, and five Ts for added flexibility. The bead-loading mixture is incubated at room temperature for 5 minutes to allow protein G to bind to the Fc region of the antibody. Simultaneously, the streptavidin-coated beads are washed and mixed with the bead-loading mixture. After incubation with gentle shaking for 30 minutes, binding of the biotinylated component is complete.
[0235] Antibodies were eluted from Protein G at pH 2, and the eluate was analyzed by SDS gel electrophoresis to assess antibody loading efficiency. Immobilized Protein G and linker content were quantified by elution with 95% formamide at 95°C for 5 minutes, followed by separation by TBE gel electrophoresis. Washed beads were stored individually at 4°C and re-pooled for multiplex barcoding analysis.
[0236] Example 2: Digestion of chromatin into mononucleosomes
[0237] In order to package genomic DNA into a more compact and dense structure, eukaryotic cells organize DNA into "chromatin". Chromatin contains DNA and histones, which exist in the form of octamers and contain two copies of histones H2A, H2B, H3 and H4. Each histone octamer is entangled with a DNA fragment of about 140bp in length. The unit of DNA and histone octamer is called a nucleosome. Nucleosomes assemble to form a higher-order structure - a 30nm chromatin fiber. The modification analysis method described herein uses a single nucleosome ("mononucleosome") as input. This embodiment provides an experimental protocol for extracting chromatin from yeast cells and then digesting the chromatin into mononucleosomes and / or DNA-protein binding complexes.
[0238] Yeast cells were cultured at 28°C until A 600 The OD is 0.8. At this stage, DNA-binding proteins (such as transcription factors) and histone octamers can be chemically crosslinked to the DNA by treatment with formaldehyde. To do this, incubate the cells in a 1% formaldehyde solution at room temperature for 1-25 minutes (the specific time depends on the desired degree of crosslinking). Terminate the reaction with 2.5 M glycine and collect the cells by centrifugation.
[0239] Cells were resuspended in lysis buffer (1 M sorbitol, 50 mM Tris-HCl pH 7.4, 10 mM β-mercaptoethanol, 10 mg / mL zymolyase) and incubated at room temperature until the cell wall was largely digested. Spheroblasts were isolated and resuspended in digestion buffer (0.5 M spermidine, 1 mM β-mercaptoethanol, 0.075% NP-40, 50 mM NaCl, 10 mM Tris-HCl pH 7.4, 5 mM MgCl2, 1 mM CaCl2). Micrococcal nuclease was then added to a final concentration of 0.07 units / uL. The reaction was incubated at 37°C for 25 minutes to digest the chromatin into mononucleosomes. The reaction was terminated by adding excess ethylenediaminetetraacetic acid (EDTA), and the nucleosomes were further purified using anion exchange medium column (Epoch Life Sciences). Nucleosome loading is accomplished in a medium-salt buffer A (25 mM MES pH 6, 10% sucrose, 10% glycerol, 400 mM NaCl). After three washes with buffer A, nucleosomes are eluted with buffer B (25 mM MES pH 6, 10% sucrose, 10% glycerol, 750 mM NaCl, 1 mM EDTA). Nucleosomes can be diluted with buffer C (10 mM Tris-HCl pH 7.5, 1 mM EDTA, 25 mM NaCl, 2 mM DTT, 20% glycerol) and stored at -80°C. This protocol typically yields >80% mononucleosomes.
[0240] Example 3: Method for nucleosomal DNA end repair
[0241] Mechanical and enzymatic shearing of chromatin produces nucleosomes with heterogeneous DNA ends. For example, the 3' end may be degraded, and a mixture of 3' and 5' phosphorylated ends may be present. The barcoding methods described below utilize different ligation methods that require 5' phosphorylation and 3' dephosphorylation, as well as blunt ends ("blunt end ligation") or a single 3' dA overhang ("sticky end ligation"). This example illustrates how nucleosomal DNA can be repaired to be compatible with these barcoding chemistries.
[0242] Blunt-end ligation. Dephosphorylate nucleosomal DNA by incubating in buffer 1 (50 mM potassium acetate, 20 mM tris-acetic acid, 10 mM magnesium acetate, 100 μg / mL recombinant albumin, pH 7.9 @ 25°C) and 1 unit of recombinant shrimp alkaline phosphatase (rSAP) at 37°C for 30 minutes. The reaction was terminated by adding excess EDTA, and the product was used as input for immunoprecipitation (IP) using two pools of microbeads prepared according to Example 1. Nucleosomal DNA was blunt-ended using T4 DNA polymerase. This enzyme has strong 3'-5' exonuclease activity, which prevents the formation of 3' overhangs during gap-filling reactions. After immunoprecipitation, the beads were resuspended in blunt-end ligation buffer (50 mM NaCl, 10 mM Tris-HCl, 10 mM MgCl2, 100 μg / mL recombinant albumin, pH 7.9 @ 25°C). In the presence of 0.1 mM dNTPs, 0.5 units of T4 DNA polymerase were added to the DNA and incubated at room temperature for 15 minutes. The reaction was terminated by adding EDTA.
[0243] Sticky-end ligation. Nucleosomal DNA was dephosphorylated as described above and used as input for immunoprecipitation. After immunoprecipitation, the beads were blunt-ended as described above and washed with HBST300 buffer. Single-base 3'dA overhangs were constructed by incubation in adenylation buffer (50 mM NaCl, 10 mM Tris-HCl, 10 mM MgCl2, 1 mM DTT, pH 7.9 at 25°C) with 5 units of Klenow fragment 3'→5' exo- and 0.1 mM dATP. The reaction was terminated by adding EDTA.
[0244] Example 4: Detection of histone modifications H3K4me3 and H3K4me2 based on microbead barcode labeling
[0245] This example describes the use of Figure 5AThe library preparation workflow shown here allows for the simultaneous characterization of one or more histone modifications. The protocol includes: immunoprecipitation of nucleosomes and end-repair followed by barcoding by ligation to determine modification status; removal of the histone core to increase DNA accessibility; reverse-strand synthesis; second adapter ligation; and PCR amplification.
[0246] 1 μg of HeLa mononucleosomes (EpiCypher product number #16-0002) was mixed with 0.25 μL K-MetStat Panel (EpiCypher product number #19-1001) and 0.5 μL K-AcylStat Panel (EpiCypher Product #19-3001) was mixed. Panels are mixed reagents of commercially synthesized mononucleosomes with known modification states. For example, The K-MetStat Panel contains specifically modified mononucleosomes assembled from recombinant human histones expressed in E. coli and wrapped with a 147-base-pair barcoded Widom 601 targeting sequence DNA. The K-MetStat Panel consists of a mixture of one unmodified and 12 histone H3 post-translational modifications: H3K4me1, H3K4me2, H3K4me3, H3K9me1, H3K9me2, H3K9me3, H3K27me1, H3K27me2, H3K27me3, H3K36me1, H3K36me2, and H3K36me3. Each nucleosome with a specific modification is distinguished by a unique DNA sequence ("barcode") at its 3' end, which can be interpreted using next-generation sequencing technology. Each of the 16 nucleosomes in the mixture is wrapped around two different DNA strands, each containing a specific barcode ("A" and "B"), serving as internal technical replicates. The K-AcylStat Panel was prepared using the same construction strategy and contains a mixture of 1 unmodified and 15 H3 histone modifications: H3K4ac, H3K9ac, H3K14ac, H3K18ac, H3K23ac, H3K27ac, H3K36ac, H3K9bu, H3K9cr, H3K18bu, H3K18cr, H3K27bu, H3K27cr, H3K27acS28phos, H3K4,9,14,18ac. This duplex experiment is expected to produce positive signals for H3K4me3 and H3K4me2, while unmodified or differently modified nucleosomes will show negative signals.
[0247] Nucleosomes were dephosphorylated according to Example 3 and diluted with HBST300 buffer. 10% of the dephosphorylated nucleosomes were transferred to a new tube and processed in parallel as an input control. For IP samples, nucleosomes were immunoprecipitated using a pool of H3K4me3 and H3K4me2 beads prepared according to Example 1. Input control nucleosomes were immunoprecipitated using a single bead type prepared using a universal histone H3 binding domain and an rcMBC linker. Although the IP reaction enriches nucleosomes carrying targeted modifications of the binding domain, the input control is designed to capture all nucleosomes with an H3 histone core (regardless of modification status). This input normalization is crucial for controlling the heterogeneity of the genome representation of the input. The read coverage obtained by the IP is divided by the reads observed in the input sample to identify histone modified regions.
[0248] Immunoprecipitation was performed at room temperature for one hour, followed by washing with RIPA buffer (50 mM Tris HCl, 300 mM NaCl, 1.0% (v / v) NP-40, 0.5% (w / v) sodium deoxycholate, 1.0 mM EDTA, 0.1% (w / v) SDS) and HBST300 buffer to remove excess nucleosomes. The DNA from the immunoprecipitated nucleosomes was blunt-ended according to Example 3. Bead-bound nucleosomes were suspended in ligation buffer (50 mM Tris-HCl pH 7.5, 150 mM NaCl, 10 mM MgCl2, 1 mM ATP, 10% PEG-8K, 0.05% Tween 20, and 400 U T4 DNA ligase) and supplemented with 0.5 μM each of MBC101 ( / 5deoxyI / / ideoxyI / / ideoxyI / CGATCAC) and MBC103 ( / 5deoxyI / / ideoxyI / / ideoxyI / AATGCGG) to induce adapter ligation. The MBC oligonucleotides are designed to provide double-stranded ligation adapters, but since the 5' end of the nucleosome is not phosphorylated, only the adapter strand coupled to the beads was intentionally ligated. This way, the MBC was introduced to only one end of the DNA.
[0249] After barcoding, histones and antibodies were removed by incubating the beads in removal buffer (5 mM DTT, 10 mM HEPES pH 7.5, 50 mM NaCl, 0.05% Tween 20). 20, pH 8.8 @ 25°C, 0.5 μM extension primer (GTCAGATGTGTATAAGAGACAG; SEQ ID NO: 3) was added with 8 units of Bst 3.0 DNA polymerase, and the complementary DNA strand was synthesized by a standard primer extension reaction. The thermal cycling program was as follows: 72°C for 2 minutes, 55°C for 5 minutes, 65°C for 15 minutes, 72°C for 5 minutes, 80°C for 5 minutes, and 16°C hold. A second sequencing adapter was introduced by repeating the same ligation steps used for the introduction of the MBC adapter in the presence of 100 units of T4 DNA ligase and 10 units of T4 polynucleotide kinase, and the 5' end of the bead strand was phosphorylated. The second adapter was a universal adapter containing only the Illumina P7 adapter (AGACGTGTGCTCTTCCGATCT; SEQ ID NO: 4) and its complement (GATCGGAAGAGC; SEQ ID NO: 5). The adapter-ligated DNA was treated with 0.1N NaOH to remove the complementary DNA strand not coupled to the beads. Bead-coupled DNA was PCR amplified using Illumina index primers and NEBNext Ultra IIQ5 master mix (NEB) according to the manufacturer's instructions. The index library was purified using AMPure beads, detected by 4% agarose gel electrophoresis, quantified by Qubit (Thermo Fisher), and sequenced.
[0250] Figure 10 A library quality control (QC) gel is shown. A clear band at the expected size of ~310 bp was observed for both the IP and input libraries, with minimal byproducts. After sequencing, the raw reads were aligned to the human genome (HeLa sample) and the SNAP-CHIP reference sequence. Each SNAP-Chip nucleosome was identified based on the Widom 601 barcode. MBC reads (introduced by our barcode tagging analysis) were mapped, deduplicated based on their UMIs, and normalized to reads per million (RPM) to eliminate differences in sequencing depth. Figure 11A The MBC distribution of a panel of SNAP-CHIP internal controls is shown. Consistent with the bead design linking H3K4me3 to the MBC101 linker, the KmetStat_H3K4me3 fragment is enriched in MBC101. Similarly, the KmetStat_H3K4me2 fragment is precisely enriched in MBC103. Compared to the H3K4me3 and H3K4me2 fragments, the KmetStat_WT unmodified fragment has very few reads in either MBC101 or MBC102, indicating very low nonspecific background. Figure 11BThe SNAP-CHIP internal control performance of each MBC is shown. For MBC101, the KmetStat_H3K4me3 fragment is the most representative fragment. For MBC103, the KmetStat_H3K4me2 fragment is the most representative fragment. Figure 11C and 11D The corresponding enrichment analysis is shown. The enrichment value is a measure of signal-to-noise and is calculated by dividing RPM(IP) by RPM(INPUT). Crosstalk is determined by the relative ratio of off-target MBC enrichment to on-target MBC enrichment. The crosstalk rates for KmetStat_H3K4me2 and KmetStat_H3K4me3 fragments were 25-28% and 1.3-1.5%, respectively, indicating that the H3K4me2 antibody is less specific than the H3K4me3 antibody. Figure 11E and 11F An example of a histone-modified genomic region from a HeLa sample is provided. The figure shows the distribution of IP sequencing reads along genomic coordinates, displaying either MBC101 (upper track) or MBC103 (middle track). The input sample (lower track) exhibits uniform read coverage along the genomic coordinates, while the MBC track exhibits a spike, indicating read enrichment in genomic regions with histone modifications.
[0251] In summary, this example demonstrates the identification of two histone modifications (H3K4me2 and H3K4me2) in nucleosomes from HeLa cells and synthetic controls using a microbead-based barcoding technique (using a pool of different microbead types). Each microbead has a binding domain and a barcoded linker for detecting a single histone modification. The binding domain pulls the target nucleosome to the microbead surface, where it is then barcoded using the surface-barcoded linker.
[0252] Example 5: Preparation of nucleosome-binding molecules
[0253] Nucleosome-bound molecules are prepared by site-specific labeling of antibodies using the SiteClick Antibody Azide Modification Kit (Thermo Fisher, Product No. S20026). SiteClick labeling uses an enzyme to specifically attach an azide molecule to the heavy chain of an IgG antibody, ensuring that the antigen binding domain remains intact for binding to the antigen target. This site selectivity is achieved by targeting a carbohydrate domain present on essentially all IgG antibodies, regardless of antibody subtype and host species. β-galactosidase catalyzes the hydrolysis of β-1,4-D-galactopyranose residues, and an engineered β-1,4-galactosyltransferase is then used to attach the azide-galactopyranosyl group. Following azide modification, a dibenzocyclooctyne (DBCO)-labeled ds-MBC linker is attached to the Fc region. In the first barcoding cycle, the linker sequence on the nucleosome side has a single 3'A overhang, which is why the DBCO oligonucleotide ends with a single 3'T: for example and The first cycle DBCO-labeled oligonucleotide contains a PacI restriction site (underlined and italicized, with the slash indicating the cleavage site), a short 4b stuffer sequence, a 7b MBC (bold italicized), a phosphorothioate (*), and a 3'T overhang. In the second and subsequent barcoding cycles, after the first adapter is released by PacI treatment, the ligation site on the nucleosome side presents a 3'TA overhang, which is why DBCO oligonucleotides must end with 5'AT: and Because the recognition motif of Pac I is extremely rare in the human genome, Pac I was selected as the restriction enzyme. Uncoupled MBC linkers were removed using a size exclusion chromatography column. TM The labeled antibodies have one or two linkers.
[0254] Example 6: Colocalization of histone modifications H3K4me3 and H3K9ac on the same nucleosome by sequential barcoding
[0255] This example uses a nucleosome-binding molecule comprising an antibody immobilized to an MBC linker for histone modification identification. This method can be used to detect a single modification or multiple modifications per nucleosome, depending on the number of barcoding cycles.
[0256] The first step is to prepare a surface where Illumina P7 adapters are distributed at single-molecule spacing. The second step is to ligate nucleosomal DNA to the P7 adapters, creating a matrix with immobilized nucleosomes that are spatially separated from each other and cannot interact with their neighbors. This step is achieved by suspending streptavidin microbeads in a mixture containing a 3' biotinylated double-stranded Illumina P7 adapter (5' end phosphate (Phos) / GATCGGAAGAGCACACGTCT) in 1x PBST buffer. U TTTTT (SEQ ID NO: 10) / 3BioTEG and AGACGTGTGCTCTTCCGATC*T (SEQ ID NO: 11)) and 100,000 molar excess of biotinylated PEG (Broadpharm, product number # BP-23759). Biotinylated PEG provides a lateral diluent and passivates the surface to reduce nonspecific binding.
[0257] According to Example 3 above, multiple nucleosomes suitable for sticky end ligation were prepared and ligated to a P7 linker using T4 DNA ligase (50 mM Tris-HCl pH 7.5, 150 mM NaCl, 10 mM MgCl2, 1 mM ATP, 10% PEG-8K, 0.05% Tween 20, 400 U T4 DNA ligase). After washing with RIPA buffer, the microbead matrix was suspended in a solution containing single or multiple nucleosome-binding molecules with linker structures designed for the first barcoding cycle. As described above, the antibody-linker conjugates were allowed to bind, and the first MBC linker was ligated to the free ends of the nucleosomes using T4 DNA ligase. To release the antibodies, the linker was cleaved with the PacI restriction enzyme in CutSmart buffer (50 mM potassium acetate, 20 mM Tris-acetate, 10 mM magnesium acetate, 100 μg / ml BSA). A second barcoding cycle is initiated by repeating the binding step using a single or multiple pools of nucleosome-binding molecules with linker structures designed for this second barcoding cycle. After restriction enzyme cleavage, this process can be repeated any number of times, with each cycle using nucleosome-binding molecules carrying MBCs matching the cycle number and binding domain specificity. To complete the library, nucleosomal DNA is ligated to a double-stranded cap containing an Illumina P5 adapter (CTACACGACGCTCTTCCGATCT*A*T (SEQ ID NO: 12) and a 5-terminal phosphate (Phos) / AGATCGGAAGAGCGTCGTGTAG (SEQ ID NO: 13)). Indexed PCR is then performed using NEBNext Ultra IIQ5 premix reagents (NEB) as described in Example 4.
[0258] The described barcoding approach is compatible with flat or beaded surfaces, as long as the immobilized nucleosomes are spaced a sufficient distance apart to eliminate nearest-neighbor interactions. Each barcoding step can employ a single type of nucleosome-binding conjugate, or a mixed library of different conjugates. This detection method, which attaches a string of MBCs, indicating the identified modification, to nucleosomal DNA, enables the first detection of histone modification colocalization at single-molecule resolution.
[0259] Example 7: Circular Coding of Immunoprecipitated Nucleosomes on Bead Matrix
[0260] This example uses multiple cycles of barcoding, each performed in the presence of a single binding domain, and attaches MBCs to the DNA ends of immunoprecipitated nucleosomes in a modification-specific manner via adapters in solution. The protocol describes immunoprecipitation, end repair, barcoding, gap filling, elution, and a final capping step to add suitable Figure 6 and Figure 9 The sequencing adapter of the workflow shown in FIG. This embodiment uses a hairpin (HP) adapter, which can be Figure 9 The figure shows the introduction of a UMI adjacent to the MBC. Although UMIs are crucial for PCR error correction and sequencing read redundancy, their introduction through double-stranded ligation is not easy because the complementary strand of the UMI must be synthesized in situ.
[0261] Bead matrices were prepared by loading histone H3 or modification-specific antibodies onto magnetic protein G microbeads according to the manufacturer's instructions. For IP samples, Ab42 (histone H3K4me3 antibody, Epicypher, product number #13-0041) was loaded onto protein G microbeads. For INPUT controls, Ab67 (histone H3 antibody, Thermo Fisher, product number #39064) was loaded onto protein G microbeads in a separate loading reaction.
[0262] 1 μg of HeLa mononucleosomes (EpiCypher product number #16-0002) was mixed with 0.25 μL K-MetStat Panel (EpiCypher product number #19-1001) and 0.5 μL The nucleosome mixture was mixed with the K-AcylStat Panel (EpiCypher catalog #19-3001). The nucleosome mixture was diluted with HBST300 buffer and split into two reactions, with 90% and 10% of the nucleosome mixture used as input for immunoprecipitation on Ab42-loaded microbeads and Ab67-loaded microbeads, respectively. Immunoprecipitation was performed at room temperature with gentle shaking for one hour.
[0263] Initial end repair involves blunting, 5' phosphorylation, and 3' dA tailing. Blunting and 5' phosphorylation of immunoprecipitated nucleosome DNA are performed in a single reaction using T4 DNA polymerase and T4 polynucleotide kinase. After immunoprecipitation, the beads are resuspended in 1X NEBr2.1 buffer. 0.5 units of T4 DNA polymerase and 5 units of T4 polynucleotide kinase are added to the beads in the presence of 0.1 mM dNTPs, 1 mM ATP, and 2 mM DTT, followed by incubation at 16°C for 15 minutes and at 23°C for 15 minutes. The reaction is terminated by adding EDTA, followed by washing with HBST300. The beads are then resuspended in 1X NEBr2.1 buffer. Add 2.5 units of Klenow Fragment 3'→5'exo- to 1X NEBr2.1 buffer supplemented with 0.1 mM dATP and incubate at 37°C for 15 minutes to add a single-base 3'A tail to the end of the nucleosomal DNA. Terminate the reaction by adding EDTA, followed by washing with HBST300 buffer.
[0264] Barcoding is achieved by ligating an HP-MBC adapter to the ends of nucleosomal DNA. The HP-MBC adapter contains a double-stranded MBC attachment site with a single-base 3' dT overhang. The other end of the double-stranded MBC region is connected by a single-stranded loop segment consisting of uracil and a UMI sequence. To focus on validating the cyclic barcoding chemistry, this example omitted the elution and re-immunoprecipitation steps. Instead, H3K4me3 was detected using MBC107 in cycle 1 and MBC109 in cycle 2, eliminating the need for nucleosome elution.
[0265] The nucleosomes bound to the microbeads were suspended in ligation buffer (50 mM Tris-HCl pH 7.5, 150 mM NaCl, 10 mM MgCl2, 1 mM ATP, 10% PEG-8K, 0.05% Tween 20 and 400 U T4 DNA ligase) and 0.5 uM HP-MBC107 was added. Inducible barcoding. The HP-MBC adapter is designed to provide double-stranded cohesive end ligation sites. However, since the 5' end of the HP-MBC adapter is not phosphorylated, only the 3'dT termini are coupled to the nucleosomal DNA ends. The barcoding reaction was incubated at 20°C for 15 minutes and at 25°C for 15 minutes. The reaction was then terminated by adding EDTA and washed with HBST300 buffer.
[0266] After MBC107 ligation, the barcoded nucleosomal DNA ends are cleaved at the uracil site, preparing for subsequent gap filling. The beads are mixed with 0.5 units of USER enzyme in 1X NEB rCutSmart buffer and incubated at 37°C for 15 minutes to complete the cleavage reaction. The reaction is washed with HBST300 buffer to expose the single-stranded 5' overhang consisting of the MBC and UMI.
[0267] Add 2.5 units of Klenow fragment 3'→5' exo- to 1X NEBr2.1 buffer supplemented with 0.2 mM dNTPs and incubate at 37°C for 15 minutes to gap-fill the 5' overhang and add dA tails to the 3' end. Terminate the reaction by adding EDTA, followed by washing with HBST300 buffer. This completes the first cycle of barcoding.
[0268] Because the barcoded nucleosomal DNA ends are not anchored to the bead surface, nucleosomes can be eluted from the bead surface and used as input for subsequent immunoprecipitation and barcoding cycles. This approach enables continuous barcoding of coexisting histone modifications on the same nucleosome. As demonstrated in Example 8, nucleosomes can be eluted using a variety of methods. As mentioned above, nucleosome elution was omitted in this example, so we proceeded to attach MBC109 after attaching MBC107.
[0269] The nucleosome-bound microbeads barcoded with HP-MBC107 in the previous step were subjected to a new round of barcoding: HP-MBC109 Ligation, USER enzyme cleavage, gap filling and 3'dA tailing.
[0270] After the last barcode labeling cycle, the nucleosomes were capped and universal sequencing adapters were ligated to the DNA ends for PCR amplification of the library. The bead-bound nucleosomes were suspended in ligation buffer (50 mM Tris-HCl pH 7.5, 150 mM NaCl, 10 mM MgCl2, 1 mM ATP, 10% PEG-8K, 0.05% Tween 20, and 400 U T4 DNA ligase) and added. ) to induce the ligation reaction. The HP-U adapter provides a bell-shaped conformation with double-stranded sticky end ligation sites and primer sites for library amplification. The adapter ligation reaction is incubated at 20°C for 15 minutes and at 25°C for 15 minutes, then EDTA is added to terminate the reaction and washed with HBST300 buffer. After adapter ligation, the circular ends are broken down into Y-shaped ends using USER enzyme cleavage. 0.5 units of USER enzyme are added to 1X NEB rCutSmart buffer and the beads are incubated at 37°C for 15 minutes to complete the cleavage reaction. The reaction is washed with HBST300 buffer.
[0271] Subsequently, 0.12 units of thermosensitive proteinase K (New England Biolabs, product number # P8111S) were added to a premix consisting of RIPA buffer, 0.4% SDS and 5mM DTT, and the adapters connected to the nucleosomal DNA were eluted by incubation. The DNA elution reaction was carried out at 37°C for one hour and then at 65°C for 10 minutes. After elution, it was further purified by AMPure microbeads. According to the manufacturer's operating instructions, the DNA was PCR amplified using Illumina index primers and NEBNext UltraIIQ5 premix reagents (NEB). The index library was purified with AMPure microbeads, evaluated by agarose gel, quantified with Qubit (Thermo Fisher) and sequenced.
[0272] Figure 12 Shown is an agarose gel of the IP and input libraries prepared using the above protocol. The agarose gel shows a clean library band with a size of approximately 320 bp, which is consistent with theoretical expectations. Figure 13A The number of SNAP-Chip fragments containing MBC107 and MBC109 after IP and barcoding is shown compared to the input control. MBC109 is significantly enriched in the KmetStat_H3K4me3 fragments, consistent with the experimental design.
[0273] Figure 13B Relative enrichment values are shown. Figure 13C An example of sequencing reads indicating two consecutive barcoding cycles is shown. Taken together, the above data validate the barcoding chemistry used for consecutive barcoding cycles.
[0274] Example 8: Method for eluting immunoprecipitated nucleosomes from the microbead surface
[0275] This example describes a method for eluting initially immunoprecipitated nucleosomes and associated DNA from the bead surface. The eluted nucleosomes remain intact, with the associated DNA still wrapped around the histone octamer, which is crucial for subsequent immunoprecipitation and barcoding of other histone modifications coexisting on the same nucleosome. In this example, we tested the effectiveness of the nucleosome elution method using three methods: a) competitive displacement of nucleosomes with modified histone peptides; b) competitive displacement of antibodies with protein G; and c) elution with a high-salt buffer.
[0276] Protein G magnetic beads were loaded onto Ab43 (histone H3K9ac antibody, ActiveMotif, product number #91103) or Ab67 (histone H3 antibody, Thermo Fisher, product number #39064) according to the manufacturer's instructions.
[0277] HeLa mononucleosomes (EpiCypher product number #16-0002) were used K-MetStat Panel (EpiCypher product number #19-1001), K-AcylStat Panel (EpiCypher Catalog #19-3001) and recombinant mononucleosomal H3K9ac (EPL, Active Motif, Catalog #81075) were spiked in. The nucleosome mixture was diluted with HBST300 and then coated onto Ab43-loaded protein G beads and immunoprecipitated for one hour at room temperature.
[0278] As described in Example 7, the nucleosomes obtained by immunoprecipitation were 5' phosphorylated with T4 polynucleotide kinase and 3' blunt-ended with T4 DNA polymerase, and then 3' dA tailed with Klenow fragment 3'→5' exo-.
[0279] After end repair, immunoprecipitated nucleosomes were mixed with an excess of biotinylated histone acetylated H3K9ac peptide (EpiGenTek, product number R-1010-100) and incubated at room temperature for 30 minutes to competitively displace nucleosomes from the antibody binding site. The supernatant containing the released nucleosomes was transferred to magnetic streptavidin microbeads and incubated at room temperature for 15 minutes to purify the nucleosomes from the excess biotinylated H3K9ac peptide. The supernatant was collected and prepared for repeated immunoprecipitation on protein A microbeads loaded with Ab67 (H3 targeting).
[0280] In a separate reaction, the immunoprecipitated nucleosomes were mixed with excess protein G molecules to displace the nucleosome-antibody complexes from the protein G beads. After incubation at room temperature for 30 minutes, the supernatant containing the released nucleosome-antibody complexes was transferred to protein A beads loaded with Ab67, and the immunoprecipitation was repeated.
[0281] In a third reaction, the immunoprecipitated nucleosomes were eluted by incubation in Gentle Ag / Ab Elution Buffer (Thermo Fisher, Catalog #21027) at room temperature for 30 minutes. The supernatant containing the eluted nucleosomes was diluted with 10 mM HEPES buffer, pH 7.5, and then applied to Ab67-loaded Protein A beads, and the immunoprecipitation was repeated.
[0282] A no-elution control reaction was also performed, i.e., no elution was performed so that the nucleosomes immunoprecipitated by Ab43 (targeting H3K9ac) remained on the protein G beads.
[0283] Uneluted control beads and remaining Protein G beads after each of the nucleosome elution reactions described above were washed with HBST300 buffer and used as input for the capping step to ligate library adapters to the nucleosomal DNA.
[0284] The second immunoprecipitation reaction was incubated with Protein A beads loaded with Ab67 at room temperature for one hour and then washed with HBST300 buffer.
[0285] All microbeads prepared above were capped and universal sequencing adapters were connected to the DNA ends for PCR amplification of the library. As described in Example 7, the microbead-bound nucleosomes were suspended in a ligation buffer (50mM Tris-HCl pH 7.5, 150mM NaCl, 10mM MgCl2, 1mM ATP, 10% PEG-8K, 0.05% Tween20, and 400U T4 DNA ligase), and 0.5uM HP-U adapter was added to induce the ligation reaction. The adapter ligation reaction was incubated at 20°C for 15 minutes and at 25°C for 15 minutes, then EDTA was added to terminate the reaction and washed with HBST300 buffer. After the adapter was connected, the circular ends were decomposed into Y-shaped ends by cutting with USER enzyme. The cleavage reaction was completed by incubating in 1X NEB rCutSmart buffer containing 0.5 units of USER enzyme at 37°C for 15 minutes. The reaction was washed with HBST300 buffer.
[0286] Subsequently, the nucleosomal DNA connected to the library adapter was eluted by incubating 0.12 units of heat-labile proteinase K (Proteinase K, New England Biolabs, product number # P8111S) in a reaction mixture consisting of RIPA buffer supplemented with 0.4% SDS and 5mM DTT. The DNA elution reaction was carried out at 37°C for one hour and then at 65°C for 10 minutes. The eluate was further purified by AMPure microbeads. The DNA was PCR amplified using Illumina index primers and NEBNext Ultra IIQ5 premix reagents (NEB) according to the manufacturer's operating instructions. The index library was purified using AMPure microbeads and quantified using Qubit (Thermo Fisher). Take 2μL of the amplified library and plate it on 4% E-Gel TM The efficiency of the nucleosome elution strategy was tested on EX agarose gel (Thermo Fisher, product number #G401004). Figure 14 Nucleosomes were eluted most efficiently with mild elution buffer (lanes 6 and 7), followed by elution with protein G (lanes 4 and 5), and least efficiently with H3K9ac peptide (lanes 2 and 3).
[0287] This example demonstrates that a mild elution buffer is optimal for removing nucleosomes from the binding domain after the barcoding cycle. The next step will be to integrate this step into the workflow described in Example 7.
[0288] Example 9: Method for identifying the spatial distribution of histone modifications in a tissue sample.
[0289] This example describes the preparation of microarrays for spatially encoding nucleosome binding conjugates and an end-to-end workflow for detecting H3K4me3 with spatial resolution.
[0290] An array of 96 spots was prepared, each of which had a nucleosome-bound conjugate carrying a unique spatial identifier and an H3K4me3 modification barcode. As described in Example 5, nucleosome-bound conjugates were prepared using the SiteClick Antibody Azide Modification Kit (ThermoFisher, product number S20026) and Ab42 (histone H3K4me3 antibody, EpiCypher, product number #13-0041) and rcMBC101 / 5 terminal phosphate (Phos) / GTGATCGNNNNNCTGTCTCTTATACACATCTGACUTTTTT (SEQ ID NO: 1) / DBCO. The array was a slide covered with protein G, and each spot was prepared by spontaneous fixation of the nucleosome-bound conjugate by protein G-antibody interaction. Each spot had a unique spatial identifier determined by the barcode contained in the conjugate at that position on the slide.
[0291] Tissue sections are prepared by freezing in a cryostat and securely mounted on microarray slides. Tissue sections can be 10-40 μm thick. To prevent dissociation of DNA-binding proteins and nucleosome disassembly, tissue sections are warmed to room temperature, fixed with 0.2% formaldehyde for 5 minutes, and then quenched with 1.25 M glycine for 5 minutes at room temperature. The tissue is washed with wash buffer containing protease inhibitors, rinsed with deionized water, followed by isopropanol, and air-dried. To document the orientation of the tissue sections relative to the microarray, standard hematoxylin and eosin (H&E) staining is performed. Hematoxylin stains cell nuclei a violet-blue, and eosin stains the extracellular matrix and cytoplasm a pink, with other structures appearing in varying shades, tones, and combinations of these colors. Microarrays carrying H&E-stained tissue sections are imaged using a light microscope before permeabilization of the cells.
[0292] The tissue sections were then infiltrated with NP40-Digitonin wash buffer. The chromatin was then digested with micrococcal nuclease (NEB) followed by the buffer recommended by the supplier. After another wash with NP40-Digitonin wash buffer, the nucleosomes were captured by the fixed nucleosome binding conjugate and washed with RIPA buffer (see above). The nucleosomal DNA ends were repaired according to the blunt end method described in Example 3. Blunt end ligation of the linker was induced by suspending the surface-bound nucleosomes in a ligation buffer (50 mM Tris-HCl pH 7.5, 150 mM NaCl, 10 mM MgCl2, 1 mM ATP, 10% PEG-8K, 0.05% Tween 20, and 400 U T4 DNA ligase) supplemented with 0.5 uM of each of MBC101 ( / 5deoxyI / / ideoxyI / / ideoxyI / CGATCAC). After barcoding, the linker was released from the antibody by USER treatment (NEB) to cleave a single uracil that is part of the linker sequence. Complementary DNA strands were synthesized in a standard primer extension reaction using 8 units of Bst 3.0 DNA polymerase in extension buffer (1 mM each dNTP, 20 mM Tris-HCl, 10 mM (NH4)2SO4, 50 mM KCl, 2 mM MgSO4, 0.1% 20, pH 8.8@25°C, 0.5uM extension primer (GTCAGATGTGTATAAGAGACAG; SEQ ID NO: 3); the following thermal cycle program was used: 72°C 2 minutes - 55°C 5 minutes - 65°C 15 minutes - 72°C 5 minutes - 80°C 5 minutes - 16°C hold. In the presence of 100 units of T4 DNA ligase and 10 units of T4 polynucleotide kinase, a second sequencing adapter was introduced by repeating the same connection steps of introducing a spatial MBC adapter to phosphorylate the 5' end of the bead chain. The second adapter is a universal adapter that contains only the Illumina P7 adapter (AGACGTGTGCTCTTCCGATCT; SEQ ID NO: 4) and its complementary sequence (GATCGGAAGAGC; SEQ ID NO: 5). Following the manufacturer's operating instructions, the barcoded DNA was purified using Ampure beads and PCR amplified using NEBNext UltraIIQ5 premix reagent (NEB). The indexed library was purified using AMPure beads, validated on a 4% agarose gel, quantified using Qubit (Thermo Fisher), and sequenced. The result of this experiment is a sequencing library that spatially encodes the histone modification H3K4Me3 present in the tissue sample.
[0293] Numbering
[0294] Notwithstanding the appended claims, the following numbered aspects constitute a part of the disclosure of the present invention and are examples and representative of the present invention.
[0295] 1. A composition comprising:
[0296] i) matrix;
[0297] ii) a binding domain coupled to a matrix; and
[0298] iii) connector;
[0299] wherein the binding domain binds to a DNA binding protein or a nucleosome containing a histone modification;
[0300] The linker comprises a nucleic acid barcode sequence that specifically corresponds to a histone modification or a DNA binding protein.
[0301] 2. The composition of any one or combination of the numbered aspects disclosed herein, wherein the substrate is a microbead, a microarray, a chip, a flow chamber, or a fluidic device.
[0302] 3. The composition of any one or combination of the numbered aspects disclosed herein, wherein the binding domain comprises an antibody, a scFv, a Fab fragment, an antibody light chain (VL), an antibody heavy chain (VH), a variable region fragment (Fv), a F(ab')2 fragment, a diabody, a VHH domain, a nanobody, a bispecific antibody, a bivalent binding domain targeting two histone modifications, an aptamer, an engineered macromolecular scaffold, an engineered protein scaffold, or a selective covalent capture agent, or a fragment or derivative thereof.
[0303] 4. The composition of any one or combination of the numbered aspects disclosed herein, wherein the binding domain comprises a histone modification reader, writer, or eraser.
[0304] 5. The composition of any one or combination of the numbered aspects disclosed herein, wherein the writer protein is a histone acetyltransferase, a lysine methyltransferase, or an arginine methyltransferase.
[0305] 6. The composition of any one or combination of the numbered aspects disclosed herein, wherein the reader protein comprises a methyl-CpG binding domain (MBD), a bromodomain adjacent to a zinc finger protein domain (BAZ), a bromodomain (BRD), a malignant brain tumor (MBT) domain, a plant homeodomain finger (PHD) domain, a chromatin binding (chromo) domain, a proline-tryptophan-tryptophan-proline domain (PWWP) domain, a tryptophan-aspartic acid dipeptide repeat domain (WD40), or a Tudor domain.
[0306] 7. The composition of any one or combination of the numbered aspects disclosed herein, wherein the erasin is a histone deacetylase, a histone lysine demethylase, or a histone arginine demethylase.
[0307] 8. The composition of any one or combination of the numbered aspects disclosed herein, wherein the binding domain comprises a catalytically inactive variant of a histone modification writer or eraser.
[0308] 9. The composition of any one or combination of the numbered aspects disclosed herein, wherein the binding domain is coupled to the matrix by covalent linkage, affinity interaction, or a combination thereof.
[0309] 10. The composition of any one or combination of the numbered aspects disclosed herein, wherein the linker is coupled to the matrix.
[0310] 11. The composition of any one or combination of the numbered aspects disclosed herein, wherein the linker is coupled to the matrix via covalent linkage, or affinity interaction, or a combination thereof.
[0311] 12. The composition of any one or combination of the numbering aspects disclosed herein, wherein the adapter comprises at least one universal sequence element in addition to the barcode.
[0312] 13. The composition of any one or combination of the numbering aspects disclosed herein, wherein the linker comprises a unique molecular identifier in addition to the barcode.
[0313] 14. A composition according to any one or combination of the numbering aspects disclosed herein, wherein the linker comprises a spatial identifier sequence in addition to the barcode.
[0314] 15. The composition of any one or combination of the numbered aspects disclosed herein, wherein the linker comprises a uracil base, an inosine base, an 8-oxoguanine base, a ribonucleoside, or a restriction sequence.
[0315] 16. The composition of any one or combination of the numbered aspects disclosed herein, wherein the linker comprises a recognition sequence for a restriction enzyme, 8-oxoguanine-DNA glycosylase, uracil-DNA glycosylase (UDG), endonuclease, or ribonuclease.
[0316] 17. The composition of any one or combination of the numbered aspects disclosed herein, wherein the linker comprises a matrix anchoring moiety.
[0317] 18. The composition of any one or combination of the numbered aspects disclosed herein, wherein the matrix anchoring moiety is biotin or desthiobiotin.
[0318] 19. The composition of any one or combination of the numbered aspects disclosed herein, wherein the matrix anchoring moiety is trans-cyclooctene (TCO), methyltetrazine (mTET), dibenzocyclooctyne (DBCO), azide, or alkyne.
[0319] 20. The composition of any one or combination of the numbering aspects disclosed herein, wherein the linker is partially double-stranded, forming a Y-shape, wherein the double-stranded portion is configured for ligation to a target nucleic acid, and each single-stranded arm comprises a universal sequence, a modified barcode, and a unique molecular identifier.
[0320] 21. A composition of any one or combination of the numbered aspects disclosed herein, wherein the linker is partially double-stranded, forms a hairpin structure, comprises a stem region configured for ligation to a target nucleic acid and a single-stranded loop region, wherein the single-stranded loop region comprises a universal sequence, a modified barcode, and a unique molecular identifier.
[0321] 22. The composition of any one or combination of the numbered aspects disclosed herein, wherein the linker is partially double stranded with a single stranded 3' overhang.
[0322] 23. The composition of any one or combination of the numbered aspects disclosed herein, wherein the linker is partially double stranded with single stranded 3' overhangs on both sides.
[0323] 24. The composition of any one or combination of the numbered aspects disclosed herein, wherein the double-stranded ends are blunt ended or have a single 3' base overhang.
[0324] 25. The composition of any one or combination of the numbered aspects disclosed herein, wherein the histone modification is methylation of lysine or arginine, citrullination, acetylation, ubiquitination, ADP-ribosylation, deamination, proline isomerization, or sumoylation.
[0325] 26. The composition of any one or combination of the numbered aspects disclosed herein, wherein the histone modification is phosphorylation of tyrosine, serine or threonine.
[0326] 27. The composition of any one or combination of the numbered aspects disclosed herein, wherein the DNA binding protein is a transcription factor or RNA polymerase II.
[0327] 28. A method for analyzing a plurality of nucleosomes, the method comprising:
[0328] (i) contacting a plurality of substrates comprising at least one composition of any one or combination of the numbered aspects disclosed herein with a solution comprising a plurality of nucleosomes, wherein the binding domain binds to a DNA binding protein or a nucleosome comprising a histone modification;
[0329] (ii) ligating an adapter carrying a nucleic acid barcode to a target DNA containing a nucleosome containing a histone modification or a DNA-binding protein;
[0330] (iii) introducing a universal sequence for amplifying target DNA;
[0331] (iv) amplifying the barcoded target DNA; and
[0332] (v) Analyzing the amplified barcoded target DNA by sequencing.
[0333] 29. A method for analyzing a plurality of nucleosomes, the method comprising:
[0334] (i) contacting a plurality of substrates comprising at least one composition of any one or combination of the numbered aspects disclosed herein with a solution comprising a plurality of nucleosomes, wherein the binding domain binds to a DNA binding protein or a nucleosome comprising a histone modification;
[0335] (ii) ligating an adapter carrying a nucleic acid barcode to a target DNA containing a nucleosome containing a histone modification or a DNA-binding protein;
[0336] (iii) release of the nucleosome from the matrix by cleavage of the attached linker;
[0337] (iv) repeating steps (i) to (iii) at least once;
[0338] (v) introducing a universal nucleic acid sequence for amplifying target DNA;
[0339] (vi) amplifying the barcoded target DNA; and
[0340] (vii) Analyzing the amplified barcoded target DNA by sequencing.
[0341] 30. The method of any one or combination of the numbered aspects disclosed herein, wherein steps (i) to (iii) are repeated at least twice.
[0342] 31. The method of any one or combination of the numbering aspects disclosed herein, wherein the releasing step comprises cleaving the attached linker at a restriction site, uracil, inosine, 8-oxoguanine or ribonucleoside of the linker with an enzyme that specifically recognizes these bases.
[0343] 32. The method of any one or combination of the numbered aspects disclosed herein, wherein the releasing step comprises cleaving the recognition sequence of the linker using a restriction enzyme, 8-oxoguanine-DNA glycosylase, uracil-DNA glycosylase (UDG), an endonuclease, a ribonuclease, or a derivative of any of these enzymes.
[0344] 33. The method of any one or combination of the numbered aspects disclosed herein, wherein steps (i) to (iii) are performed using two or more different types of substrates, each substrate comprising a different binding domain and linker carrying a nucleic acid barcode.
[0345] 34. The method of any one or combination of the numbered aspects disclosed herein, wherein both the binding domain and the linker are attached to the matrix by covalent linkage, affinity interaction, or a combination thereof.
[0346] 35. A method according to any one or combination of the numbering aspects disclosed herein, comprising using a different binding domain and linker each time steps (i) to (iii) are repeated.
[0347] 36. A method for analyzing a plurality of nucleosomes, the method comprising:
[0348] (i) contacting a substrate comprising a composition of any one or combination of the numbered aspects disclosed herein with a solution comprising a plurality of nucleosomes, wherein the binding domain binds to the DNA binding protein or the nucleosome comprising the histone modification;
[0349] (ii) adding linkers to the plurality of nucleosomes bound to the binding domain;
[0350] (iii) ligating an adapter carrying a nucleic acid barcode to a target DNA containing a nucleosome containing a histone modification or a DNA binding protein;
[0351] (iv) releasing the nucleosome from the binding domain by adding a buffer that disrupts the interaction between the binding domain and the nucleosome;
[0352] (v) repeating steps (i) to (iv) at least once;
[0353] (vi) introducing a universal sequence for amplifying target DNA;
[0354] (vii) amplifying the barcoded target DNA; and
[0355] (viii) Analyzing the amplified barcoded target DNA by sequencing.
[0356] 37. The method of any one or combination of the numbered aspects disclosed herein, wherein steps (i) to (iv) are repeated at least twice.
[0357] 38. The method of any one or combination of the numbered aspects disclosed herein, wherein ligating the linkers comprises T4 DNA ligase, circularizing ligase, T3 DNA ligase, T7 DNA ligase, 9°N DNA ligase, Taq DNA ligase, or E. coli DNA ligase.
[0358] 39. The method of any one or combination of the numbering aspects disclosed herein, wherein the step of introducing a universal sequence comprises ligating an adapter carrying a nucleic acid barcode to the target DNA, wherein the adapter is a partially double-stranded Y-shaped adapter or a partially double-stranded bell-shaped adapter.
[0359] 40. The method of any one or combination of the numbered aspects disclosed herein, wherein the releasing step comprises adding a buffer comprising a reducing agent, an enzyme that specifically digests the antibody (e.g., papain and / or pepsin), a synthetic modified histone peptide as a competitive binding agent, a surfactant (e.g., SDS, sodium deoxycholate), an acidic buffer having a pH of 6.5 or less or an alkaline buffer having a pH of 8.5 or above, about 0.3 M to about 2 M NaCl, or about 0.5 M to about 1 M NaCl.
[0360] 41. A nucleosome-binding conjugate comprising:
[0361] i) a binding domain; and
[0362] ii) a linker bound to the binding domain;
[0363] wherein the binding domain binds to a DNA binding protein or a nucleosome comprising a histone modification;
[0364] The linker comprises a nucleic acid barcode sequence that specifically corresponds to a histone modification or a DNA binding protein.
[0365] 42. The nucleosome-binding conjugate of any one or combination of the numbered aspects disclosed herein, wherein 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 linkers are bound to the nucleosome-binding conjugate.
[0366] 43. A nucleosome-binding conjugate of any one or combination of the numbered aspects disclosed herein, wherein the binding domain comprises an antibody, a scFv, a Fab fragment, an antibody light chain (VL), an antibody heavy chain (VH), a variable region fragment (Fv), a F(ab')2 fragment, a diabody, a VHH domain, a nanobody, a bispecific antibody, a bivalent binding domain targeting two histone modifications, an aptamer, an engineered macromolecular scaffold, an engineered protein scaffold, or a selective covalent capture agent, or a fragment or derivative thereof.
[0367] 44. The nucleosome-binding conjugate of any one or combination of the numbered aspects disclosed herein, wherein the binding domain comprises a DNA or chromatin reader, writer, or eraser protein.
[0368] 45. The nucleosome-binding conjugate of any one or combination of the numbered aspects disclosed herein, wherein the writer protein is a DNA methyltransferase, a histone acetyltransferase, a lysine methyltransferase, or an arginine methyltransferase.
[0369] 46. The nucleosome-binding conjugate of any one or combination of the numbered aspects disclosed herein, wherein the reader protein comprises an MBD domain, a BAZ domain, a BRD domain, an MBT domain, a PHD domain, a chromo domain, a PWWP domain, a WD40 domain, or a Tudor domain.
[0370] 47. The nucleosome-binding conjugate of any one or combination of the numbered aspects disclosed herein, wherein the eraser protein is a methylcytosine dioxygenase, a histone deacetylase, or a histone lysine demethylase.
[0371] 48. The nucleosome-binding conjugate of any one or combination of the numbered aspects disclosed herein, wherein the binding domain comprises a catalytically inactive variant of a histone modification writer or eraser protein.
[0372] 49. The nucleosome binding conjugate of any one or combination of the numbering aspects disclosed herein, wherein the linker comprises a universal sequence in addition to the barcode.
[0373] 50. The nucleosome binding conjugate of any one or combination of the numbering aspects disclosed herein, wherein the linker comprises a unique molecular identifier in addition to the barcode.
[0374] 51. A composition according to any one or combination of the numbering aspects disclosed herein, wherein the linker comprises a spatial identifier sequence in addition to the barcode.
[0375] 52. The nucleosome-binding conjugate of any one or combination of the numbered aspects disclosed herein, wherein the linker comprises a uracil base, an inosine base, an 8-oxoguanine base, a ribonucleoside, or a restriction sequence.
[0376] 53. A nucleosome-binding conjugate of any one or combination of the numbered aspects disclosed herein, wherein the linker comprises a recognition sequence for a restriction enzyme, 8-oxoguanine-DNA glycosylase, uracil-DNA glycosylase (UDG), an endonuclease, a ribonuclease, or a derivative of any of these enzymes.
[0377] 54. A nucleosome-binding conjugate of any one or combination of the numbering aspects disclosed herein, wherein the linker is partially double-stranded, forming a Y-shaped structure, wherein the double-stranded portion is configured for ligation to a target nucleic acid, and each single-stranded arm can comprise a universal sequence, a modified barcode, a unique molecular identifier, and an optional spatial identification sequence.
[0378] 55. A nucleosome-binding conjugate of any one or combination of the numbering aspects disclosed herein, wherein the linker is partially double-stranded, forming a hairpin structure, comprising a stem region configured for ligation to a target nucleic acid and a single-stranded loop region, wherein the single-stranded loop region comprises a universal sequence, a modified barcode, a unique molecular identifier, and an optional spatial identifier sequence.
[0379] 56. The nucleosome-binding conjugate of any one or combination of the numbered aspects disclosed herein, wherein the linker is partially double-stranded with a single-stranded 3' overhang.
[0380] 57. The nucleosome-binding conjugate of any one or combination of the numbered aspects disclosed herein, wherein the linker is partially double-stranded with single-stranded 3' overhangs on both sides.
[0381] 58. The nucleosome-binding conjugate of any one or combination of the numbered aspects disclosed herein, wherein the double-stranded ends are blunt ended or have a single 3' base overhang.
[0382] 59. The nucleosome-binding conjugate of any one or combination of the numbered aspects disclosed herein, wherein the histone modification is methylation of lysine or arginine, citrullination, acetylation, ubiquitination, ADP-ribosylation, proline isomerization, or sumoylation.
[0383] 60. The nucleosome-binding conjugate of any one or combination of the numbered aspects disclosed herein, wherein the histone modification is phosphorylation of tyrosine, serine and threonine.
[0384] 61. The nucleosome-binding conjugate of any one or combination of the numbered aspects disclosed herein, wherein the DNA binding protein is a transcription factor or RNA polymerase II.
[0385] 62. A method for analyzing a plurality of nucleosomes, the method comprising:
[0386] (i) contacting a solution comprising a plurality of nucleosomes with a solution comprising at least one nucleosome-binding conjugate according to any one or combination of the numbered aspects disclosed herein, wherein the binding domain binds to a DNA-binding protein or a nucleosome comprising a histone modification;
[0387] (ii) ligating an adapter carrying a nucleic acid barcode of a nucleosome-binding conjugate to a target DNA containing a nucleosome of a histone modification or a DNA-binding protein in an environment where the amount of off-target barcoded DNA generated is less than 20% of the barcoded target DNA to generate the barcoded target DNA;
[0388] (iii) introducing a universal sequence for amplifying target DNA;
[0389] (iv) amplifying the barcoded target DNA; and
[0390] (v) Analyzing the amplified barcoded target DNA by sequencing.
[0391] 63. A method according to any one or combination of the numbering aspects disclosed herein, comprising transferring the linker of one or both nucleosome-binding conjugates to the same target DNA.
[0392] 64. A method according to any one or combination of the numbered aspects disclosed herein, comprising transferring a linker of two nucleosome-binding conjugates to the same target DNA.
[0393] 65. The method of any one or combination of the numbering aspects disclosed herein, comprising limiting off-target barcode labeling by performing the ligation step in a micromolar, nanomolar, picomolar, femtomolar, attamole, or zeptomolar solution of nucleosomes and nucleosome-binding conjugates.
[0394] 66. A method of analyzing a plurality of nucleosomes, the method comprising:
[0395] (i) immobilizing multiple nucleosomes on a substrate at a spacing that results in less than 20% off-target barcode labeling;
[0396] (ii) contacting the immobilized nucleosomes with a solution comprising at least one nucleosome binding conjugate of any one or combination of the numbered aspects disclosed herein, wherein the binding domain is attached to a DNA binding protein or a nucleosome comprising a histone modification;
[0397] (iii) ligating an adapter carrying a nucleic acid barcode of a nucleosome-binding conjugate to a target DNA containing a nucleosome of a histone modification or a DNA-binding protein;
[0398] (iv) cleaving the linker to generate a nucleic acid terminus having a structure suitable for ligation with another linker;
[0399] (v) repeating steps (ii) to (iv) at least once;
[0400] (vi) introducing a universal nucleic acid sequence for amplifying target DNA;
[0401] (vii) amplifying the barcoded target DNA; and
[0402] (viii) Analyzing the amplified barcoded target DNA by sequencing.
[0403] 67. The method of any one or combination of the numbered aspects disclosed herein, wherein steps (ii) to (iv) are repeated at least twice.
[0404] 68. A method of any one or combination of the numbering aspects disclosed herein, comprising limiting off-target barcode labeling by immobilizing nucleosomes on a substrate at a spacing of 50 nm or greater, such as 50 nm, 55 nm, 60 nm, 65 nm, 70 nm, 75 nm, 80 nm, 85 nm, 90 nm, 95 nm, 100 nm, 50-500 nm, 50-400 nm, 50-300 nm, 50-250 nm, 50-200 nm, 50-100 nm, or any integer value or range between 50 and 1000 nm.
[0405] 69. A method according to any one or combination of the numbering aspects disclosed herein, comprising cleaving a linker at the uracil, inosine, 8-oxoguanine or ribonucleoside of the linker with an enzyme specific for the uracil, inosine, 8-oxoguanine or ribonucleoside of the linker.
[0406] 70. The method of any one or combination of the numbering aspects disclosed herein, comprising cleaving the recognition sequence of the linker with a restriction enzyme.
[0407] 71. The method of any one or combination of the numbered aspects disclosed herein, wherein the linker comprises a recognition sequence for a restriction enzyme, 8-oxoguanine-DNA glycosylase, uracil-DNA glycosylase (UDG), an endonuclease, or a ribonuclease.
[0408] 72. A method according to any one or combination of the numbering aspects disclosed herein, comprising using a different binding domain and linker each time steps (ii) to (iv) are repeated.
[0409] 73. The method of any one or combination of the numbered aspects disclosed herein, wherein ligating comprises using T4 DNA ligase, circular ligase, T3 DNA ligase, T7 DNA ligase, 9°N DNA ligase, Taq DNA ligase, or E. coli DNA ligase.
[0410] 74. A method for analyzing a plurality of nucleosomes in a tissue environment, the method comprising:
[0411] (i) immobilizing multiple nucleosome-binding conjugates on a planar microarray substrate at a spacing such that the off-target barcode labeling rate is less than 20%;
[0412] (ii) placing the tissue section on top of a planar microarray substrate containing a plurality of nucleosome-binding conjugates;
[0413] (iii) permeabilized tissue cells;
[0414] (iv) digesting the chromatin with an endonuclease and capturing nucleosomes via an immobilized nucleosome-binding conjugate;
[0415] (v) ligating a linker carrying a nucleic acid barcode and a spatial identifier sequence of a nucleosome-binding conjugate to a target DNA containing a nucleosome of a histone modification or a DNA-binding protein in an environment where the amount of off-target barcoded DNA generated is less than 20% of the barcoded target DNA, thereby generating the barcoded target DNA;
[0416] (vi) introducing a universal sequence for amplifying target DNA;
[0417] (vii) amplifying the barcoded target DNA;
[0418] (vii) analyzing the amplified barcoded target DNA by sequencing; and
[0419] (viii) Determine the identity of histone modification or DNA binding proteins and their spatial locations on a planar microarray substrate based on barcode and spatial identifier sequences.
[0420] 75. The method of any one or any combination of the numbering aspects disclosed herein, comprising limiting off-target barcode labeling by immobilizing nucleosomes or nucleosome-binding conjugates on a substrate at a separation distance of 50 nm or greater.
[0421] 76. A method of analyzing a plurality of nucleosomes, the method comprising:
[0422] (i) introducing a universal linker onto the target DNA of the nucleosome;
[0423] (ii) contacting a solution comprising a plurality of nucleosomes with a solution comprising at least one nucleosome-binding conjugate of any one or any combination of the numbered aspects disclosed herein, wherein the binding domain binds to a DNA-binding protein or a nucleosome comprising a histone modification;
[0424] (iii) connecting the linkers of the bound multiple nucleosome-binding conjugates through a ligation reaction;
[0425] (iv) hybridizing the universal linker of the target DNA to the 3′ end of the ligation adapter;
[0426] (v) replicating the adapter-ligated sequence to generate a copy of the barcoded target DNA;
[0427] (vi) introducing a universal nucleic acid sequence for amplifying target DNA;
[0428] (vii) amplifying the barcoded nucleosomal DNA; and
[0429] (viii) Analyzing the barcoded target DNA by sequencing.
[0430] 77. The method of any one or combination of the numbering aspects disclosed herein, wherein introducing the universal sequence comprises ligating a forward or reverse sequencing adapter to the barcode.
[0431] 78. The method of any one or combination of the numbered aspects disclosed herein, wherein the binding domain of the nucleosome binding conjugate is attached to an internal position of the nucleic acid linker.
[0432] 79. The method of any one or combination of the numbering aspects disclosed herein, comprising A-tailing the nucleosomes in step (i).
[0433] 80. The method of any one or combination of the numbering aspects disclosed herein, comprising attaching a universal linker sequence in step (i).
[0434] 81. A method according to any one or combination of the numbering aspects disclosed herein, comprising linkers connecting the bound plurality of nucleosome-binding conjugates via double-stranded, single-stranded or splinted linkages.
[0435] 82. The method of any one or combination of the numbered aspects disclosed herein, wherein amplifying the barcoded target DNA comprises generating substrate-immobilized clonal clusters of monoclonal copies of the target DNA by surface amplification.
[0436] 83. The method of any one or combination of the numbered aspects disclosed herein, wherein analyzing the amplified barcoded target DNA comprises in situ sequencing of substrate-immobilized clonal clusters of monoclonal copies of the target DNA.
[0437] 84. The method of any one or combination of the numbered aspects disclosed herein, wherein analyzing the barcoded target DNA comprises analyzing the barcoded DNA by nucleic acid probe hybridization.
[0438] 85. The method of any one or combination of the numbered aspects disclosed herein, wherein analyzing the barcoded target DNA comprises analyzing the barcoded DNA by PCR.
[0439] 86. The method of any one or combination of the numbering aspects disclosed herein, comprising obtaining nucleosomes from cell-free circulating nucleosomes.
[0440] 87. A method according to any one or combination of the numbering aspects disclosed herein, comprising obtaining nucleosomes from chromatin by enzymatic or mechanical shearing.
[0441] 88. The method of any one or combination of the numbering aspects disclosed herein, comprising obtaining nucleosomes from a single cell.
[0442] 89. A method for diagnosing a cancer or cancer subtype associated with one or more histone modifications, comprising analyzing a plurality of nucleosomes according to any one or combination of the numbering aspects disclosed herein.
[0443] 90. A method for monitoring cancer progression or treatment response comprising analyzing a plurality of nucleosomes according to any one or combination of the numbering aspects disclosed herein.
[0444] 91. The method of any one or combination of the numbering aspects disclosed herein, comprising obtaining a plurality of nucleosomes from a blood sample.
[0445] 92. The method of any one or combination of the numbering aspects disclosed herein, comprising obtaining a plurality of nucleosomes from a tissue biopsy sample.
[0446] 93. A kit for monitoring epigenetic changes over time in a sample obtained from a subject receiving treatment, comprising a composition of any one or combination of the numbered aspects disclosed herein, or a nucleosome-binding conjugate of any one or combination of the numbered aspects disclosed herein, and instructions for using the composition or nucleosome-binding conjugate to monitor epigenetic changes over time.
[0447] 94. The kit of any one or combination of the numbered aspects disclosed herein, wherein the subject is being treated for cancer.
Claims
1. A composition comprising: i) matrix; ii) a binding domain coupled to a matrix; and iii) connector; wherein the binding domain binds to a DNA binding protein or a nucleosome containing a histone modification; The linker comprises a nucleic acid barcode sequence that specifically corresponds to a histone modification or a DNA binding protein.
2. The composition of claim 1, wherein the substrate is a microbead, a microarray, a chip, a flow chamber, or a fluidic device.
3. The composition of claim 1, wherein the binding domain comprises an antibody, a scFv, a Fab fragment, an antibody light chain (VL), an antibody heavy chain (VH), a variable region fragment (Fv), a F(ab')2 fragment, a diabody, a VHH domain, a nanobody, a bispecific antibody, a bivalent binding domain targeting two histone modifications, an aptamer, an engineered macromolecular scaffold, an engineered protein scaffold, or a selective covalent capture agent, or a fragment or derivative thereof.
4. The composition of claim 1, wherein the binding domain comprises a histone modification reader, writer, or eraser.
5. The composition according to claim 4, wherein the writer protein is histone acetyltransferase, lysine methyltransferase, or arginine methyltransferase.
6. The composition of claim 4, wherein the reader protein comprises a methyl-CpG binding domain (MBD), a bromodomain adjacent to a zinc finger protein domain (BAZ), a bromodomain (BRD), a malignant brain tumor (MBT) domain, a plant homeodomain finger (PHD) domain, a chromatin binding (chromo) domain, a proline-tryptophan-tryptophan-proline domain (PWWP) domain, a tryptophan-aspartic acid dipeptide repeat domain (WD40), or a Tudor domain.
7. The composition of claim 4, wherein the erasing protein is a histone deacetylase, a histone lysine demethylase, or a histone arginine demethylase.
8. The composition of claim 1, wherein the binding domain comprises a catalytically inactive variant of a histone modification writer or eraser.
9. The composition of claim 1, wherein the binding domain is coupled to the matrix by covalent linkage, affinity interaction, or a combination thereof.
10. The composition of any one of claims 1-9, wherein the linker is coupled to a substrate. The composition of claim 10 , wherein the linker is coupled to the substrate via covalent linkage, or affinity interaction, or a combination thereof.
12. The composition of any one of claims 1 to 11, wherein the linker comprises at least one universal sequence element in addition to the barcode.
13. The composition of any one of claims 1-12, wherein the linker comprises a unique molecular identifier in addition to a barcode.
14. The composition of any one of claims 1 to 13, wherein the linker further comprises a spatial identifier sequence in addition to the barcode.
15. The composition of any one of claims 1-14, wherein the linker comprises a uracil base, an inosine base, an 8-oxoguanine base, a ribonucleoside, or a restriction sequence.
16. The composition of any one of claims 1-15, wherein the linker comprises a recognition sequence for a restriction enzyme, 8-oxoguanine-DNA glycosylase, uracil-DNA glycosylase (UDG), an endonuclease, or a ribonuclease.
17. The composition of any one of claims 1-16, wherein the linker comprises a matrix anchoring moiety.
18. The composition of any one of claims 1-17, wherein the matrix anchoring moiety is biotin or desthiobiotin.
19. The composition of any one of claims 1-18, wherein the matrix anchoring moiety is trans-cyclooctene (TCO), methyltetrazine (mTET), dibenzocyclooctyne (DBCO), azide, or alkyne.
20. The composition of any one of claims 1-19, wherein the linker is partially double-stranded, forming a Y-shape, wherein the double-stranded portion is configured for ligation to a target nucleic acid, and each single-stranded arm comprises a universal sequence, a modified barcode, and a unique molecular identifier.
21. The composition of any one of claims 1-20, wherein the linker is partially double-stranded, forms a hairpin structure, comprises a stem region configured for ligation to a target nucleic acid and a single-stranded loop region, wherein the single-stranded loop region comprises a universal sequence, a modified barcode, and a unique molecular identifier.
22. The composition of any one of claims 1-21, wherein the linker is partially double-stranded with a single-stranded 3' overhang.
23. The composition of any one of claims 1-22, wherein the linker is partially double-stranded with single-stranded 3' overhangs on both sides.
24. The composition of any one of claims 20-22, wherein the double-stranded ends are blunt-ended or have a single 3' base overhang.
25. The composition of any one of claims 1-24, wherein the histone modification is methylation of lysine or arginine, citrullination, acetylation, ubiquitination, ADP-ribosylation, deamination, proline isomerization, or SUMOylation.
26. The composition of any one of claims 1-25, wherein the histone modification is phosphorylation of tyrosine, serine, or threonine.
27. The composition of any one of claims 1-26, wherein the DNA binding protein is a transcription factor or RNA polymerase II.
28. A method for analyzing a plurality of nucleosomes, the method comprising: (i) contacting a plurality of substrates comprising at least one composition of any one of claims 1 to 27 with a solution comprising a plurality of nucleosomes, wherein the binding domain binds to a DNA binding protein or a nucleosome comprising a histone modification; (ii) ligating an adapter carrying a nucleic acid barcode to a target DNA containing a nucleosome containing a histone modification or a DNA-binding protein; (iii) introducing a universal sequence for amplifying target DNA; (iv) amplifying the barcoded target DNA; and (v) Analyzing the amplified barcoded target DNA by sequencing.
29. A method for analyzing a plurality of nucleosomes, the method comprising: (i) contacting a plurality of substrates comprising at least one composition of any one of claims 1 to 27 with a solution comprising a plurality of nucleosomes, wherein the binding domain binds to a DNA binding protein or a nucleosome comprising a histone modification; (ii) ligating an adapter carrying a nucleic acid barcode to a target DNA containing a nucleosome containing a histone modification or a DNA-binding protein; (iii) release of the nucleosome from the matrix by cleavage of the attached linker; (iv) repeating steps (i) to (iii) at least once; (v) introducing a universal nucleic acid sequence for amplifying target DNA; (vi) amplifying the barcoded target DNA; and (vii) Analyzing the amplified barcoded target DNA by sequencing.
30. The method of claim 29, wherein steps (i) to (iii) are repeated at least twice.
31. The method of claim 29, wherein the releasing step comprises cleaving the attached linker at a restriction site, uracil, inosine, 8-oxoguanine, or ribonucleoside of the linker using an enzyme that specifically recognizes these bases.
32. The method of claim 29, wherein the releasing step comprises cleaving the recognition sequence of the linker using a restriction enzyme, 8-oxoguanine DNA glycosylase, uracil-DNA glycosylase (UDG), an endonuclease, a ribonuclease, or a derivative of any of these enzymes.
33. The method of claim 29, wherein steps (i) to (iii) are performed using two or more different types of substrates, each substrate comprising a different binding domain and linker carrying a nucleic acid barcode.
34. The method of claim 29, wherein both the binding domain and the linker are attached to the matrix via covalent linkage, affinity interaction, or a combination thereof.
35. The method of claim 29, comprising using a different binding domain and linker each time steps (i)-(iii) are repeated.
36. A method for analyzing a plurality of nucleosomes, the method comprising: (i) contacting a substrate comprising a composition according to any one of claims 1 to 27 with a solution comprising a plurality of nucleosomes, wherein the binding domain binds to a DNA binding protein or a nucleosome comprising a histone modification; (ii) adding linkers to the plurality of nucleosomes bound to the binding domain; (iii) ligating an adapter carrying a nucleic acid barcode to a target DNA containing a nucleosome containing a histone modification or a DNA binding protein; (iv) releasing the nucleosome from the binding domain by adding a buffer that disrupts the interaction between the binding domain and the nucleosome; (v) repeating steps (i) to (iv) at least once; (vi) introducing a universal sequence for amplifying target DNA; (vii) amplifying the barcoded target DNA; and (viii) Analyzing the amplified barcoded target DNA by sequencing.
37. The method of claim 36, wherein steps (i) to (iv) are repeated at least twice.
38. The method of any one of claims 28-37, wherein ligating the linkers comprises T4 DNA ligase, circular ligase, T3 DNA ligase, T7 DNA ligase, 9°N DNA ligase, Taq DNA ligase, or E. coli DNA ligase.
39. The method according to any one of claims 28 to 37, wherein the step of introducing a universal sequence comprises ligating an adapter carrying a nucleic acid barcode to the target DNA, wherein the adapter is a partially double-stranded Y-shaped adapter or a partially double-stranded bell-shaped adapter.
40. The method of claim 36, wherein the releasing step comprises adding a buffer comprising a reducing agent, an enzyme that specifically digests the antibody (e.g., papain and / or pepsin), a synthetic modified histone peptide as a competitive binding agent, a surfactant (e.g., SDS, sodium deoxycholate), an acidic buffer having a pH of 6.5 or less or an alkaline buffer having a pH of 8.5 or greater, about 0.3 M to about 2 M NaCl, or about 0.5 M to about 1 M NaCl.
41. A nucleosome-binding conjugate comprising: i) binding domain; and ii) a linker bound to the binding domain; wherein the binding domain binds to a DNA binding protein or a nucleosome comprising a histone modification; The linker comprises a nucleic acid barcode sequence that specifically corresponds to a histone modification or a DNA binding protein.
42. The nucleosome binding conjugate of claim 41, wherein 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 linkers are bound to the nucleosome binding conjugate.
43. The nucleosome-binding conjugate of claim 41, wherein the binding domain comprises an antibody, a scFv, a Fab fragment, an antibody light chain (VL), an antibody heavy chain (VH), a variable region fragment (Fv), a F(ab')2 fragment, a diabody, a VHH domain, a nanobody, a bispecific antibody, a bivalent binding domain targeting two histone modifications, an aptamer, an engineered macromolecular scaffold, an engineered protein scaffold, or a selective covalent capture agent, or a fragment or derivative thereof.
44. The nucleosome binding conjugate of claim 41, wherein the binding domain comprises a DNA or chromatin reader, writer, or eraser protein.
45. The nucleosome binding conjugate of claim 44, wherein the writer protein is a DNA methyltransferase, a histone acetyltransferase, a lysine methyltransferase, or an arginine methyltransferase.
46. The nucleosome-binding conjugate of claim 44, wherein the reader protein comprises an MBD domain, a BAZ domain, a BRD domain, an MBT domain, a PHD domain, a chromo domain, a PWWP domain, a WD40 domain, or a Tudor domain.
47. The nucleosome-binding conjugate of claim 44, wherein the eraser protein is a methylcytosine dioxygenase, a histone deacetylase, or a histone lysine demethylase.
48. The nucleosome-binding conjugate of claim 41, wherein the binding domain comprises a catalytically inactive variant of a histone modification writer or eraser.
49. The nucleosome binding conjugate of claim 41, wherein the linker comprises a universal sequence in addition to the barcode.
50. The nucleosome binding conjugate of claim 41, wherein the linker further comprises a unique molecular identifier in addition to the barcode.
51. The composition of any one of claims 41, wherein the linker comprises a spatial identifier sequence in addition to a barcode.
52. The nucleosome binding conjugate of claim 41, wherein the linker comprises a uracil base, an inosine base, an 8-oxoguanine base, a ribonucleoside, or a restriction sequence.
53. The nucleosome binding conjugate of claim 41, wherein the linker comprises a recognition sequence for a restriction enzyme, 8-oxoguanine-DNA glycosylase, uracil-DNA glycosylase (UDG), an endonuclease, a ribonuclease, or a derivative of any of these enzymes.
54. The nucleosome-binding conjugate of claim 41, wherein the linker is partially double-stranded, forming a Y-shaped structure, wherein the double-stranded portion is configured for ligation to a target nucleic acid, and each single-stranded arm can comprise a universal sequence, a modified barcode, a unique molecular identifier, and an optional spatial identifier sequence.
55. The nucleosome-binding conjugate of claim 41, wherein the linker is partially double-stranded, forming a hairpin structure, comprising a stem region configured for ligation to a target nucleic acid and a single-stranded loop region, wherein the single-stranded loop region comprises a universal sequence, a modified barcode, a unique molecular identifier, and an optional spatial identifier sequence.
56. The nucleosome-binding conjugate of claim 41, wherein the linker is partially double-stranded with a single-stranded 3' overhang.
57. The nucleosome-binding conjugate of claim 41, wherein the linker is partially double-stranded with single-stranded 3' overhangs on both sides.
58. The nucleosome-binding conjugate of any one of claims 54-57, wherein the double-stranded ends are blunt-ended or have a single 3' base overhang.
59. The nucleosome-binding conjugate of any one of claims 41-58, wherein the histone modification is methylation of lysine or arginine, citrullination, acetylation, ubiquitination, ADP-ribosylation, proline isomerization, or SUMOylation.
60. The nucleosome binding conjugate of any one of claims 41-59, wherein the histone modification is phosphorylation of tyrosine, serine, and threonine.
61. The nucleosome-binding conjugate of any one of claims 41-59, wherein the DNA binding protein is a transcription factor or RNA polymerase II.
62. A method for analyzing a plurality of nucleosomes, the method comprising: (i) contacting a solution comprising a plurality of nucleosomes with a solution comprising at least one nucleosome-binding conjugate according to any one of claims 41 to 61, wherein the binding domain binds to a DNA-binding protein or a nucleosome comprising a histone modification; (ii) ligating an adapter carrying a nucleic acid barcode of a nucleosome-binding conjugate to a target DNA containing a nucleosome of a histone modification or a DNA-binding protein in an environment where the amount of off-target barcoded DNA generated is less than 20% of the barcoded target DNA to generate the barcoded target DNA; (iii) introducing a universal sequence for amplifying target DNA; (iv) amplifying the barcoded target DNA; and (v) Analyzing the amplified barcoded target DNA by sequencing.
63. The method of claim 62, comprising transferring the linker of one or both nucleosome-binding conjugates to the same target DNA.
64. The method of claim 63, comprising transferring linkers of two nucleosome-binding conjugates to the same target DNA.
65. The method of any one of claims 62-64, comprising limiting off-target barcode labeling by performing the ligation step in a micromolar, nanomolar, picomolar, femtomolar, attamole, or zeptomolar solution of nucleosomes and nucleosome-binding conjugates.
66. A method of analyzing a plurality of nucleosomes, the method comprising: (i) immobilizing multiple nucleosomes on a substrate at a spacing that results in less than 20% off-target barcode labeling; (ii) contacting the immobilized nucleosomes with a solution comprising at least one nucleosome-binding conjugate according to any one of claims 41 to 61, wherein the binding domain is attached to a DNA-binding protein or a nucleosome comprising a histone modification; (iii) ligating an adapter carrying a nucleic acid barcode of a nucleosome-binding conjugate to a target DNA containing a nucleosome of a histone modification or a DNA-binding protein; (iv) cleaving the linker to generate a nucleic acid terminus having a structure suitable for ligation with another linker; (v) repeating steps (ii) to (iv) at least once; (vi) introducing a universal nucleic acid sequence for amplifying target DNA; (vii) amplifying the barcoded target DNA; and (viii) Analyzing the amplified barcoded target DNA by sequencing.
67. The method of claim 66, wherein steps (ii) to (iv) are repeated at least twice.
68. The method of claim 66, comprising limiting off-target barcode labeling by immobilizing nucleosomes on a substrate at a separation distance of 50 nm or greater.
69. The method of claim 66, comprising cleaving the linker at the uracil, inosine, 8-oxoguanine or ribonucleoside of the linker with an enzyme specific for the uracil, inosine, 8-oxoguanine or ribonucleoside of the linker.
70. The method of claim 66, comprising cleaving the recognition sequence of the linker with a restriction enzyme.
71. The method of claim 66, wherein the linker comprises a recognition sequence for a restriction enzyme, 8-oxoguanine-DNA glycosylase, uracil-DNA glycosylase (UDG), an endonuclease, or a ribonuclease.
72. The method of claim 66, comprising using a different binding domain and linker each time steps (ii)-(iv) are repeated.
73. The method of any one of claims 62-72, wherein ligating comprises using T4 DNA ligase, circular ligase, T3 DNA ligase, T7 DNA ligase, 9°N DNA ligase, Taq DNA ligase, or E. coli DNA ligase.
74. A method for analyzing a plurality of nucleosomes in a tissue environment, the method comprising: (i) immobilizing multiple nucleosome-binding conjugates on a planar microarray substrate at a spacing such that the off-target barcode labeling rate is less than 20%; (ii) placing the tissue section on top of a planar microarray substrate containing a plurality of nucleosome-binding conjugates; (iii) permeabilized tissue cells; (iv) digesting the chromatin with an endonuclease and capturing nucleosomes via an immobilized nucleosome-binding conjugate; (v) ligating a linker carrying a nucleic acid barcode and a spatial identifier sequence of a nucleosome-binding conjugate to a target DNA containing a nucleosome of a histone modification or a DNA-binding protein in an environment where the amount of off-target barcoded DNA generated is less than 20% of the barcoded target DNA, thereby generating the barcoded target DNA; (vi) introducing a universal sequence for amplifying target DNA; (vii) amplifying the barcoded target DNA; (vii) analyzing the amplified barcoded target DNA by sequencing; and (viii) Determine the identity of histone modification or DNA binding proteins and their spatial locations on a planar microarray substrate based on barcode and spatial identifier sequences.
75. The method of any one of claims 62-74, comprising limiting off-target barcode labeling by immobilizing nucleosomes or nucleosome-binding conjugates on a substrate at a separation distance of 50 nm or greater.
76. A method of analyzing a plurality of nucleosomes, the method comprising: (i) introducing a universal linker onto the target DNA of the nucleosome; (ii) contacting a solution comprising a plurality of nucleosomes with a solution comprising at least one nucleosome-binding conjugate according to any one of claims 41 to 61, wherein the binding domain binds to a DNA-binding protein or a nucleosome comprising a histone modification; (iii) connecting the linkers of the bound multiple nucleosome-binding conjugates through a ligation reaction; (iv) hybridizing the universal linker of the target DNA to the 3′ end of the ligation adapter; (v) replicating the adapter-ligated sequence to generate a copy of the barcoded target DNA; (vi) introducing a universal nucleic acid sequence for amplifying target DNA; (vii) amplifying barcoded nucleosomal DNA; and (viii) Analyzing the barcoded target DNA by sequencing.
77. The method of any one of claims 28-40 and 62-76, wherein introducing the universal sequence comprises ligating a forward or reverse sequencing adapter to the barcode.
78. The method of any one of claims 28, 29, 36, 62, 66, and 76, wherein the binding domain of the nucleosome binding conjugate is attached to an internal position of a nucleic acid linker.
79. The method of any one of claims 28, 29, 36, 62, 66, and 76, comprising A-tailing the nucleosomes in step (i).
80. The method of claim 76, comprising ligating a universal linker sequence in step (i).
81. The method of claim 76, comprising linking the bound plurality of nucleosome-binding conjugates via double-stranded, single-stranded, or splinted linkages.
82. The method of any one of claims 28, 29, 36, 62, 66, and 76, wherein amplifying the barcoded target DNA comprises generating substrate-immobilized clonal clusters of monoclonal copies of the target DNA by surface amplification.
83. The method of any one of claims 28, 29, 36, 62, 66, and 76, wherein analyzing the amplified barcoded target DNA comprises in situ sequencing of substrate-immobilized clonal clusters of monoclonal copies of the target DNA.
84. The method of any one of claims 28, 29, 36, 62, 66, and 76, wherein analyzing the barcoded target DNA comprises analyzing the barcoded DNA by nucleic acid probe hybridization.
85. The method of any one of claims 28, 29, 36, 62, 66, and 76, wherein analyzing the barcoded target DNA comprises analyzing the barcoded DNA by PCR.
86. The method of any one of claims 28, 29, 36, 62, 66, and 76, comprising obtaining nucleosomes from cell-free circulating nucleosomes.
87. The method of any one of claims 28, 29, 36, 62, 66, and 76, comprising obtaining nucleosomes from chromatin by enzymatic or mechanical shearing.
88. The method of any one of claims 28, 29, 36, 62, 66, and 76, comprising obtaining nucleosomes from a single cell.
89. A method for diagnosing a cancer or cancer subtype associated with one or more histone modifications, comprising analyzing a plurality of nucleosomes according to any one of claims 28, 29, 36, 62, 66, and 76.
90. A method for monitoring cancer progression or treatment response comprising analyzing a plurality of nucleosomes according to any one of claims 28, 29, 36, 62, 66, and 76.
91. The method of any one of claims 89-90, comprising obtaining a plurality of nucleosomes from a blood sample.
92. The method of any one of claims 89-90, comprising obtaining the plurality of nucleosomes from a tissue biopsy sample.
93. A kit for monitoring epigenetic changes over time in a sample obtained from a subject receiving treatment, comprising the composition of any one of claims 1-27, or the nucleosome-binding conjugate of any one of claims 41-61, and instructions for using the composition or nucleosome-binding conjugate to monitor epigenetic changes over time.
94. The kit of claim 93, wherein the subject is being treated for cancer.
Citation Information
Patent Citations
Method for transposase-mediated spatial tagging and analyzing genomic DNA in a biological sample
US11519033B2
Method for transposase-mediated spatial tagging and analyzing genomic DNA in a biological sample
US20210010070A1
Capturing oligonucleotides in spatial transcriptomics
US20210237022A1
Profiling of biological analytes with spatially barcoded oligonucleotide arrays
US20220010367A1
Spatially distinguished, multiplex nucleic acid analysis of biological specimens
US20220298560A1