Diagnostic Applications of Cell-Free DNA Chromatin Immunoprecipitation
By extracting and sequencing modified histones bound to cfDNA, combined with computer program analysis, the problem of difficult to determine the source of circulating cell-free DNA and cell status is solved, and efficient and accurate diagnosis of early disease detection and treatment adjustment is achieved.
Patent Information
- Application Number
- CN201980031975.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-05-06
- Filing Date
- 2019-03-13
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2039-03-13
AI Technical Summary
The prior art is difficult to accurately determine the source and cellular status of circulating cell-free DNA (cfDNA), especially in low-volume samples, making it difficult to distinguish between healthy and pathological tissues, resulting in difficulty in diagnosis and treatment adjustment of early disease.
By extracting and sequencing proteins bound to cfDNA, especially modified histones, such as H3K4me1, H3K4me2, H3K36me3, etc., in combination with computer program analysis, information genomic locations are identified to determine the cell or tissue status of the cfDNA origin.
It realizes efficient and accurate determination of the source and cellular status of cfDNA in low-volume samples, supports early disease detection and treatment adjustment, reduces the need for sequencing depth, and improves detection sensitivity and specificity.
Smart Images

Figure CN112119166B_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims the benefit of priority to U.S. Provisional Patent Application No. 62 / 642,158, filed on March 13, 2018, and U.S. Provisional Patent Application No. 62 / 667,528, filed on May 6, 2018, which are incorporated herein by reference in their entireties. Technical Field
[0003] The present invention belongs to the field of cell-free DNA-protein complex analysis. Background Art
[0004] Following cell death (through apoptosis or necrosis), short DNA fragments are released into the plasma. These are often referred to as circulating cell-free DNA (cfDNA) or circulating tumor DNA (ctDNA, if derived from tumor cells). The existence of cfDNA has been known for decades, with cfDNA fragments typically being multiples of approximately 166 bp in length—the length of mononucleosomal DNA (~146 bp)—with some additional linker DNA. Plasma from healthy individuals contains the equivalent of ∼1000 genomes / ml, and up to 100-fold more cfDNA is present in many pathological conditions (e.g., cancer) and some physiological conditions (e.g., after exercise). These fragments are short-lived, with an estimated half-life of less than 1 hour, making them ideal biomarkers for noninvasively monitoring physiological and pathological processes. In recent years, the use of cfDNA as a diagnostic tool has expanded significantly. For example, next-generation sequencing of fetal cfDNA in maternal blood is now used for noninvasive prenatal screening / diagnosis of chromosomal abnormalities and parent-of-origin mutations. Because cfDNA is present in very low quantities in blood samples, most current cfDNA diagnostic methods rely on mutations in cfDNA to distinguish it from cfDNA in healthy tissue and blood cells. In fact, blood cell cfDNA is by far the largest contributor to the total cfDNA pool and can make cfDNA diagnosis of conditions in other tissues difficult.
[0005] Most current cfDNA-based methods rely on detecting genomic changes in cfDNA to quantify the contribution of cfDNA to cells whose genomic sequences (such as fetal, transplant, or tumor mutant genes) have been altered. Therefore, these methods focus on a set of preselected genes and ignore events involving turnover and death of cells whose genomes are identical to the host genome. Recent methods have utilized epigenetic information in cell-free DNA. Extremely deep sequencing of total cfDNA can provide data reflecting source tissue and gene expression. However, it relies on detecting changes in target region coverage, where the signal of the source tissue is imposed on the background of normal cells (for example, the detection of events that lead to nucleosome depletion in 10% of cells requires distinguishing 90% occupancy from 100% occupancy). Therefore, this method avoids sampling noise by utilizing extremely deep sequencing coverage (hundreds of millions of reads per sample). Even with such sequencing depth, there are strict and harsh detection limits for events in rare subgroups of cells. A promising alternative is to measure DNA CpG methylation along the sequence to identify source cells. DNA methylation acts as a stable epigenetic memory and remains largely unchanged after differentiation. Therefore, it provides a lot of information about cell lineage, but little about temporal changes in expression and cells derived from close or similar lineages. In addition, unbiased analysis of DNA methylation requires high sequencing depth because most CpGs are methylated.
[0006] Methods that allow accurate determination of the source of cfDNA and provide information about the molecular events that occur in cells as they approach cell death would not only allow earlier diagnosis of conditions unknown to physicians or patients but could also help tailor treatment to newly discovered diseases. Summary of the Invention
[0007] The present invention provides methods for determining the source of cell-free DNA (cfDNA), detecting cell type or tissue death, determining the cellular state of cells in a subject, and combinations thereof, by sequencing the cfDNA isolated by extracting proteins and modified proteins bound to the cfDNA. Also provided are computer program products for doing so.
[0008] According to a first aspect, there is provided a method for determining the cell state, tissue of origin, cell type, or a combination thereof, of a cell that releases its DNA, comprising:
[0009] a. providing a sample, wherein the sample comprises cell-free DNA (cfDNA);
[0010] b. contacting the sample with at least one reagent that binds to a DNA-associated protein;
[0011] c. separating the reagent and any bound proteins and cfDNA;
[0012] d. sequencing the isolated cfDNA; and
[0013] e. specifying that the cfDNA molecule comprising a DNA sequence at an informative genomic location is derived from a cell in a certain cellular state, from a certain tissue, from a certain cell type, or a combination thereof, wherein the association of the DNA-associating protein with the informative genomic location is indicative of the cellular state, tissue of origin, cell type, or a combination thereof, of the cell that released the cfDNA; thereby determining the cellular state, tissue of origin, cell type, or a combination thereof, of the cell that released its DNA.
[0014] According to another aspect, a computer program product for determining a cell or tissue of origin of cell-free DNA (cfDNA) is provided, comprising a non-transitory computer-readable storage medium having program code embodied thereon, the program code being executable by at least one hardware processor to:
[0015] a. Measuring or accessing sequencing of cfDNA isolated using reagents that bind to DNA-associated proteins;
[0016] b. assigning a cfDNA molecule from the cfDNA to a cell or tissue of origin by comparing the DNA sequence of the molecule to a sequence associated with the DNA-associated protein in the cell type or tissue; and
[0017] c. Provide output regarding the source cell or tissue of the cfDNA.
[0018] According to another aspect, a computer program product for determining a cell state, tissue of origin, cell type, or a combination thereof of a cell in a subject at the time of cell death is provided, comprising a non-transitory computer-readable storage medium having program code embodied thereon, the program code being executable by at least one hardware processor to
[0019] a. measuring or accessing sequencing of cfDNA from a subject isolated with a reagent that binds to a DNA-associated protein;
[0020] b. assigning a cfDNA molecule in the cfDNA to a cell state, tissue of origin, cell type, or a combination thereof by comparing the DNA sequence of the molecule to a sequence associated with the DNA-associated protein in the cell state, tissue, cell type, or a combination thereof; and
[0021] c. providing an output regarding the cell state, tissue of origin, cell type, or a combination thereof of a cell in the subject at the time of cell death.
[0022] According to another aspect, a solid support is provided that comprises a capture agent and a barcoding reagent.
[0023] According to another aspect, there is provided a method for multiplexing assays for more than one target molecule in a single solution, the method comprising:
[0024] a. capturing a first target molecule in solution onto a first solid support of the present invention;
[0025] b. capturing at least a second target molecule in solution onto a second solid support of the present invention;
[0026] c. attaching a first target molecule and a first barcode and at least a second target molecule and a second barcode;
[0027] d. Simultaneously measuring the first and second target molecules, wherein the measurement result of the first target molecule is identified by the first barcode, and the measurement result of the second target molecule is identified by the second barcode;
[0028] This allows for multiplexing of assays for more than one target molecule in a single solution.
[0029] According to some embodiments, the sample is from a subject.
[0030] According to some embodiments, the cell that releases its DNA is a dying cell, and the method is used to detect the death of at least one of:
[0031] a. a cell type in the subject,
[0032] b. an organization within the subject, and
[0033] c. A cell in a certain cellular state in a subject.
[0034] According to some embodiments, the cell state is a disease state. According to some embodiments, the disease state is selected from bacteremia, cancer, pre-cancer, infection, neurodegenerative disease, tissue damage, heart disease, liver disease, inflammation, autoimmune disease, arthritis, liver inflammation, intestinal inflammation, autoimmune disease, tissue damage due to drug side effects, tissue necrosis, and diabetes. According to some embodiments, the disease state is selected from heart disease or damage, brain disease or damage, gastrointestinal disease or damage, cancer, bacteremia, infection, and liver disease or damage.
[0035] According to some embodiments, cfDNA of at least 500 genomes is provided. According to some embodiments, said designation can be performed with as little as 0.1% of cfDNA from said cell type, said tissue or said cell state in the sample.
[0036] According to some embodiments, the agent is selected from an antibody or antigen-binding fragment thereof, a protein or a small molecule.
[0037] According to some embodiments, the agent is conjugated to a physical support.
[0038] According to some embodiments, the DNA-associating protein is selected from the group consisting of histones, high mobility group (HMG) proteins, and members of the transcriptional machinery. According to some embodiments, the histone is a histone variant and / or a modified histone. According to some embodiments, the histone variant is selected from the group consisting of histone 3 monomethylated lysine 4 (H3K4me1), histone 3 dimethylated lysine 4 (H3K4me2), histone 3 trimethylated lysine 36 (H3K36me3), and histone 3 trimethylated lysine 4 (H3K4me3). According to some embodiments, the agent is an anti-modified histone antibody or a fragment thereof.
[0039] According to some embodiments, association of the DNA-associating protein with the genomic location indicates active transcription, and the genomic location is within a gene or enhancer element specific for a tissue, cell type, or cell state, or is at a disease-specific mutation. According to some embodiments, association of the DNA-associating protein with the genomic location indicates silenced transcription, and the genomic location is within a repressor element, or a gene silenced in the tissue, cell type, or cell state, or is at a disease-specific mutation.
[0040] According to some embodiments, the method of the present invention further comprises performing steps ad again using a reagent that binds to a second DNA-associated protein, and wherein the second DNA-associated protein is different from the first DNA-associated protein.
[0041] According to some embodiments, the methods of the present invention comprise contacting a sample with at least two reagents, wherein each reagent is bound to a physical support and the support comprises a short DNA tag unique to each reagent, wherein upon sequencing the isolated cfDNA, the short DNA tag identifies the reagent that isolated the cfDNA.
[0042] According to some embodiments, the designation comprises comparing the sequenced cfDNA to at least 10 genomic locations in a tissue, cell type, or cell state that are most uniquely associated with a DNA-associated protein, and wherein cfDNA having a sequence identical to a DNA sequence within the at least 10 genomic locations is considered to be from the tissue, cell type, or cell state.
[0043] According to some embodiments, the DNA-associated protein is a marker of active transcription, and the designation comprises comparing the sequenced cfDNA to a known transcriptional program of a tissue, cell type, or cell state, wherein the cfDNA having sequences from genes transcribed in the transcriptional program is from the tissue, cell type, or cell state.
[0044] According to some embodiments, the designation comprises comparing the sequenced cfDNA to a DNA-protein association atlas of at least five cell types or tissues, wherein the atlas comprises at least ten genomic positions with the most unique association of DNA-associated proteins in each of the five cell types or tissues, and wherein cfDNA having a sequence identical to a DNA sequence within the at least ten genomic positions is considered to be from the tissue or cell type.
[0045] According to some embodiments, the designation comprises comparing the sequenced cfDNA to a transcriptional program profile of at least five transcriptional programs, wherein the profile comprises at least one genomic location with the most unique association of a DNA-associated protein for each of the five transcriptional programs, and wherein cfDNA having a sequence identical to a DNA sequence within the at least one genomic location indicates activation of the transcriptional program.
[0046] According to some embodiments, the cell state is selected from the group consisting of: hypoxia, inflammation, ER stress, mitochondrial stress, interferon response, dormancy, senescence, cycling, malignancy, and calcium flux.
[0047] According to some embodiments, the informative genomic location is selected from the group consisting of a promoter, an enhancer element, a silencer element, a gene body, and a disease-associated mutation.
[0048] According to some embodiments, the method of the present invention wherein:
[0049] a. DNA-associated proteins are markers of active transcription, and the disease-associated mutation is in an oncogene, or
[0050] b. DNA-associated proteins are markers of silent transcription, and disease-associated mutations are in tumor suppressor genes.
[0051] According to some embodiments, the methods of the present invention are used to detect a disease state in a subject.
[0052] According to some embodiments, the method of the present invention wherein detecting a disease state comprises at least one of:
[0053] a. Early detection of disease states;
[0054] b. Detection of residual metastatic disease; and
[0055] c. Monitor disease progression with and without treatment.
[0056] According to some embodiments, the methods of the present invention further comprise treating the subject with an appropriate therapy based on the cell state, tissue of origin, cell type, or a combination thereof, of the dead cells in the subject.
[0057] According to some embodiments, the solid support is magnetic or paramagnetic beads, or agarose beads.
[0058] According to some embodiments, the capture agent is a protein. According to some embodiments, the capture protein is an antibody or an antigen-binding fragment thereof.
[0059] According to some embodiments, the barcoding agent is a short nucleic acid molecule. According to some embodiments, the nucleic acid molecule is 5 to 30 nucleotides.
[0060] According to some embodiments, the capture agent and the barcoding reagent are conjugated to a solid support.
[0061] According to some embodiments, the target molecule is a protein or a nucleic acid molecule.
[0062] According to some embodiments, the assay is chromatin immunoprecipitation followed by sequencing (ChIP-Seq).
[0063] Other embodiments and the full scope of applicability of the present invention will be apparent from the detailed description given below. However, it should be understood that the detailed description and specific examples, while indicating preferred embodiments of the present invention, are given by way of illustration only, as various changes and modifications within the spirit and scope of the present invention will be apparent to those skilled in the art from this detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] Some embodiments of the present invention are described herein with reference to the accompanying drawings as examples only. When specific reference is now made to the drawings in detail, it is noted that the details shown are by way of example and for the purpose of illustrative discussion of embodiments of the present invention. In this regard, the description and drawings make it apparent to those skilled in the art how embodiments of the present invention may be practiced.
[0065] Figures 1A-1I: (1A) Overview of the proposed method. Chromatin fragments from different cells in the body are released into the blood. These are immunoprecipitated and sequenced. Interpretation of the resulting sequences informs the tissue of origin and gene activity program. Inset-cfChIP protocol, using antibodies covalently bound to paramagnetic beads. Target fragments are immunoprecipitated directly from plasma. After removing the plasma and washing the beads bound to the target fragments, sequencing adapters (possibly with index barcodes) are added to the fragments using on-bead-ligation, and the ligated DNA is separated and PCR amplified, and sequencing-ready libraries are prepared. (1B) Heatmap of reads of cell type-specific H3K4me1 and H3K4me3 sites compiled from RoadmapEpigenomics data. Specific sites for individual tissues / cell types and / or related cell populations are shown. (1C) Aligned segments of chromosome 2 showing cfChIP-seq signals. The upper track is the cfChIP-seq signal from four subjects who were determined to be healthy. The lower track is published ChIP-seq results for human leukocytes (WBCs) and tissues. Below that is a 100x magnification of the results showing consistency in peak positions. (1D) Histogram of meta-analysis of cfChIP signals at active promoters and enhancers. (IE) Size distribution histogram of sequenced cfChIP fragments showing clear mono- and di-nucleosome sizes. (IF) Browser view of cfChIP-seq signals at two regions of megakaryocyte-specific genes that appear in healthy subjects but not in ChIP of blood cells and solid tissues. (1G) Browser view of non-PBMC H3K4me3 signals at the promoters of selected genes (similar to Figure 1F ). The top and bottom panels depict cfChIP and tissue ChIP signals, respectively. (1H) Browser view of mouse CTCF signals at known CTCF sites. The sites were confirmed by depletion of H3K4me3 signals. (1I) Meta-analysis of mouse CTCF (top) and H3K4me3 (bottom) signals across the entire mouse genome.
[0066] Figure 2A -M: (2A-2B) Scatter plots comparing cfChIP with anti-H3K4me1 and anti-H3K4me3 antibodies in (2A) technical replicates and (2B) one male and one female healthy individual. Each dot is a 2 kb window in the genome, and the x and y axes are the number of reads mapped to the window in the two samples (on a log(x+1) scale). The color code reflects the density of the dots. (2C) Histogram of correlation between H3K4me3 cfChIP samples from healthy subjects. (Left) Correlation of counts in 2 kb windows (as Figure 2A-B) and (right) Correlation of counts in gene promoters. We see that samples from the same subject (red histogram) tend to be slightly more correlated with each other than samples from different subjects (blue histogram). (2D) Browser example of sex-specific peaks. Male and female plasma samples were mixed in known ratios and subjected to cfChIP for H3K4me3. (2E) Bar graph of detection of male-specific chrY signatures in the samples shown in 2D. FDR-adjusted q values for background signal are shown. (2F) Graph showing that the H3K4me3 male signal is linearly correlated with the fraction. Comparison of read counts in simulations based on 100% male samples and Poisson samples versus the observed number. (2G) Line graph estimates of the probability of detecting a specific position with different numbers of reads. Detection probabilities were estimated by down-sampling from the actual results. Bars represent 95% confidence intervals for the estimates. (2H) Line graph extrapolation of spike-in for larger feature sizes. The probability of detecting 0.1% males in two sample sizes is shown. (2I) Bar graph of the sizes of highly tissue-specific features (see Table 1). (2J) Scatter plot of the correlation between H3K4me3 and RNA levels at the promoters of constitutively expressed genes. (Top) ChIP-seq of PBMC (leukocytes) vs. RNA-seq of PBMC. (Bottom) cfChIP of healthy subjects vs. RNA-seq of PBMC. (2K) Scatter plot comparison of H3K4me3 cfChIP-seq and expression levels. Each dot is a gene. x-axis: number of H3K4me3 reads in gene promoters (normalized; Methods). y-axis: leukocyte RNA-seq counts of genes. (2L) Dot plot of tissue-specific features detected in cfChIP of healthy subjects. Feature counts for cells whose cfDNA is expected to be expressed in cfDNA are shown: neutrophils, 35% cfDNA; monocytes, 25% cfDNA; and hepatocytes, 1%; as well as feature counts for the negative control (heart). The dots in each column are counts for a particular object. (2M) Display Figure 2L Dot plot of the significance of features in .
[0067] Figure 3A-J: (3A) Bar graph of H3K4me3 cfChIP-seq signals in cardiac-specific windows for four healthy subjects and samples from patients with myocardial infarction (MI). Inset, Troponin levels measured at the time of blood sample withdrawal. (3B) Example of a signal browser view at a cardiac-specific window. Each browser section shows a 20 kb region around the window (marked with a gray background). The tracks are all normalized and displayed at the same scale. The upper track shows the cfChIP sample, and the lower track shows ChIP-seq of the tissue sample (bottom) from the Roadmap Epigenomics Atlas. (3C) Dot plot comparison with external signs of cardiomyocyte death. x-axis: measured troponin levels (top), cardiomyocyte fraction measured using DNA methylation markers (bottom). y-axis: intensity of cardiac-specific features (relative to healthy subjects). (3D) Heat map showing the level (brown scale) and significance (blue scale) of selected cell type features in healthy subjects and patients with myocardial infarction. Each cell in the figure is divided into two halves, the upper left half represents the statistical significance (FDR corrected q value), and the lower half represents the read density in the feature (normalized reads / kb). (3E) Heat map of tissue features for all samples; and expansions of 3D and 3I. (3F) As the evaluated features (see Figure 3B ) section. (3G) Line plot of feature intensity changes in myocardial infarction patients before / after PCI. Feature intensities are normalized relative to healthy subjects. The difference between healthy subjects is shown on the left. We can see initially high levels of hepatocytes and elevated levels of cardiomyocytes. Following PCI, hepatocytes decrease while cardiomyocytes increase. (3H) Dot plot comparison of external indications and liver features in cancer patients. Figure 3C (3I) shows the cell type characteristics of cancer patients (see Figure 3D and 3E (3J) Combined line and bar graph of changes in liver characteristics (bars) and ALT levels (a biomarker of liver damage, black line) in blood samples from patients undergoing hepatectomy.
[0068] Figure 4A -H: (4A) Heatmap showing over-represented Hallmark genes in subjects (compared to healthy baseline) (e.g. Figure 3D ).See Figure 4C , a complete table of all markers and objects. (4B) Example of a browser view of genes with higher than expected signals in these expression signatures (see Figure 3B (4C) Heatmap of signature features for all samples and features. Figure 4A(4D) Browser view of H3K4me3 cfChIP and tissue ChIP signals at the promoters of selected glycolytic genes (see Figure 4B ). (4E) Example scatter plot of a method for defining genes with elevated signals in a particular sample. Scatter plot of normalized H3K4me3 counts at the promoter of each gene. x-axis: mean of a reference healthy sample. y-axis: counts in the sample in question. Colored dots represent genes in the cancer signature. Larger dots are significantly overexpressed. (4F) Heat map showing the enrichment of a tumor-specific signature among overexpressed genes. Each cell is divided in half, with the upper left half representing statistical significance (FDR-corrected q-value) and the lower half representing overlap with the signature (percentage of the number of genes in the signature). See Figure 4G , complete table of all tumors and subjects. (4G) Heatmap of cancer signatures across all samples. (4H) Example of a browser view of cancer-related genes and their signals across different samples.
[0069] Figure 5A-M: (5A) Histogram of meta-analysis of cfChIP signals at active promoters and enhancers. (5B-C) Browser view of cfChIP tracks for H3K4me3, H3K4me2, and H3K4me13 from healthy subjects. (5B) Regions of highly expressed genes are shown. We can see that dimethylation and monomethylation extend from the trimethylation signal. (5C) The locus around IFNB1 is shown. ChromHMM tracks show predictions of promoters and enhancers based on a combination of histone modification and chromatin accessibility measurements. Arrows mark regions enriched for dimethylation and monomethylation. (5D-E) Scatter plots showing (5D) the correlation of H3K4me2 and H3K4me3 at promoters in two samples from healthy subjects and cancer patients, and (5E) the concordance between H3K4me2 in healthy subjects and between H3K4me2 in two samples taken months apart from cancer patients. The differences between healthy and cancer samples are significant. (5F-H) Track browser views comparing H3K4 methylation markers between healthy and cancer samples (C002.2), (5F) TCF3, (5G) CDX1, (5H) CEACAM5, and CEACAM6. (5I) Meta-analysis of H3K36me3 signals across 5 kb of gene bodies flanking the transcription start site (TSS) and transcription end site (TES) by cfChIP. Proportional to gene length. (5J) Scatter plot of H3K36me3 correlation between leukocytes and healthy samples. (5K) Box plot of H3K36me3 annotated with active genes—healthy sample H3K36me3 counts (normalized by gene length), broken down by leukocyte RNA level quantiles. (5L) Scatter plot of raw H3K36me3 counts across gene bodies. Each dot represents a gene. x-axis: healthy samples. y-axis: colorectal adenocarcinoma samples. Colored dots indicate genes overexpressed in colorectal adenocarcinoma (COAD - red) or glioblastoma multiforme (GBM - green). (5M) Browser view of H3K4me3 and H3K36me3 signals at genes showing differential levels of these markers between healthy subjects and colorectal adenocarcinoma patients. The VIL1 gene shows differential signals for both markers, while CTDSP1 shows similar levels of H3K4me3 but a significant increase in H3K36me3 in colorectal adenocarcinoma patient samples.
[0070] Figure 6A -C: (6A-C) Line graph examples of background estimation, (6A) healthy male sample, (6B) healthy female sample, and (6C) cancer patient.
[0071] Figure 7 : Workflow of processing and analysis of cf-ChIP.
[0072] Figure 8 Figure 3: Meta-mapping (top) and heatmap (bottom) of 1000 highly expressed promoters and the location of H3K4me3 relative to the transcription start site (TSS) in cf-nucleosomes from plasma that had undergone cfChIP with an anti-H3K4me1 antibody.
[0073] Figure 9 : Meta-mapping (top) and heatmap (bottom) of 1000 highly expressed promoters and the positions of H3K9Ac, H3K27ac, and H2A.Zac relative to TSS in cf-nucleosomes of plasma from healthy patients and one colorectal cancer patient.
[0074] Figure 10A -D: (10A) Schematic diagram depicting the protocol for multiplexed ChIP-Seq. (10B) Schematic diagram of experiments performed to test pooling during MPL-based ChIP-Seq. Each rectangle represents an MPL barcode surface that combines a unique barcode (BC1-BC4) with a combination of anti-H3K4me3 (K4) or anti-H3K36me3 (K36) antibodies, each targeting a chromatin modification at a different genomic location. The ovals are chromatin from two yeast species: S. cerevisiae and K. lactis. Various poolings were then performed before or after library preparation (pink circle in the middle). (10C) Bar graph of the fraction of immobilized chromatin captured (shown as % of input). (10D) Linear meta-analysis of the distribution of H3K4me3 and H3K36me3 across gene bodies based on ChIP-seq signals from MPL barcoded surfaces. DETAILED DESCRIPTION
[0075] The present invention provides methods for determining the origin of cell-free DNA (cfDNA), methods for detecting the death of a cell type or tissue in a subject, and methods for determining the cellular state of a cell at the time of its death by determining DNA-protein associations from the subject's transcriptionally active or inactive chromatin. The methods of the present invention are based on the surprising discovery that cell-free nucleosomes retain protein-DNA associations that are informative not only about the tissue / cell of origin of the nucleosomes, but also about the active and inactive pathways in the cell at the time of its death. Furthermore, this is surprisingly feasible even with a very small number of cf-nucleosomes that can be captured. As few as one thousand genomes worth of cf-nucleosomes are sufficient to perform the methods of the present invention.
[0076] The methods of the present invention can be performed with very small input cfDNA, as few as 1000 genomes, and very shallow sequencing of as few as 0.5M reads. The technology can be performed in this way because only positive associations are investigated. Instead of sequencing the entire cfDNA, only the cfDNA that is bound by a specific protein (e.g., a modified histone) is isolated and sequenced. Because only a small portion of the cfDNA is sequenced, the process is cheaper, faster, and can be done at a lower sequencing depth. Even in such a small sample, only informative genomic loci are investigated; most positions have no information about active pathways in the tissue / cell of origin or in dying cells. By investigating only informative loci, much of the noise present in the cfDNA can be ignored. Finally, because DNA sequences that only associate with the protein of interest in some cases are investigated, only a small number of reads in these regions are needed to identify positive reads in the cfDNA. For example, if the binding of the protein of interest to a DNA sequence that is uniquely bound in cardiac tissue is investigated, healthy individuals (those with no or little cardiac cell death) will have only a handful of reads in these regions (see Figure 3A -E). The differences in tissue and reads found in healthy subjects are very small, so detection of abnormal cell death can be performed even with very few reads that are different from healthy individuals. Subjects with an increased number of reads at unique genomic regions of cardiac tissue will be identified as having elevated cardiac cell death. It is not necessary to measure every read in these regions, as negative data is irrelevant, and only reads that are significantly elevated above the baseline of healthy individuals are sufficient. This can also be done to study the pathways and cellular states of dying cells. Since cfDNA from healthy subjects has very few reads showing gene activation in the hypoxia pathway, reads in these regions will indicate that hypoxia is the cause of increased cell death in patients.
[0077] cfChIP has the potential to circumvent many of the limitations currently present in cfDNA analysis. Targeted enrichment of active markers results in a reduction in genomic representation, thus requiring fewer sequencing reads (~2 orders of magnitude fewer) to obtain an informative signal. Because we target markers associated with active transcription, we analyze positive signals where very few reads indicate the presence of a specific cell type or expression program. This contrasts with methods such as occupancy or DNA methylation that measure negative signals (no nucleosome occupancy) or both negative and positive signals (e.g., % methylation). Furthermore, the cfChIP assay leaves much of the original sample intact, enabling the use of the same material for multiple assays (e.g., genome sequencing, methylation analysis, or cfChIP with other antibodies), which is important in situations where blood volume is a limiting factor.
[0078] Over the past two decades, intensive research has established the connection between specific histone marks and chromatin-templated processes, including transcription, replication, and damage repair. Utilizing this rich and complex information for analysis of cfDNA from blood has the potential to reveal physiological processes in remote organs, such as cell proliferation, hypoxia, inflammation, metabolic changes, and cancerous transformation, in real time and with minimal invasiveness. These processes all involve the activation of large transcriptional programs that leave unique imprints on chromatin.
[0079] A key factor in using cfDNA-based assays to detect cfDNA in rare cells (such as in early cancer diagnosis) is the low limit of detection. Several features of cfChIP can greatly improve the limit of detection. 1. cfChIP detects "positive" signals, so even low signals contribute significantly. 2. cfChIP can be performed using a variety of antibodies targeting different genomic regions and states, resulting in large signatures and a range of hundreds or thousands of sites with differential signals between different tissues or transcriptional programs. 3. Because cfChIP is inherently a low-expression method, cfChIP is unbiased because all captured DNA fragments are sequenced.
[0080] Measuring modified cf-nucleosomes—alone or in combination with existing biomarkers—has a variety of potential medical uses, such as early disease detection (e.g., detecting unknown tumors), improved diagnostics (e.g., replacing tissue biopsies with liquid biopsies), and non-invasive monitoring of disease progression and treatment efficacy.
[0081] In a first aspect, a method is provided for determining the cell or tissue of origin of cell-free DNA (cfDNA), the method comprising:
[0082] a. Provide a sample containing cfDNA;
[0083] b. contacting the sample with at least one reagent that binds to a DNA-associated protein;
[0084] c. separating the agent and any bound protein and cfDNA; and
[0085] d. Sequencing the isolated cfDNA;
[0086] The isolated cfDNA contains a DNA sequence at an informative genomic location, and the association of the DNA-associated protein with the informative genomic location indicates a cell type or tissue; thereby determining the cell or tissue of origin of the cfDNA.
[0087] In another aspect, a method is provided for determining the cell state, tissue of origin, cell type, or a combination thereof, of a cell that releases its DNA, comprising:
[0088] a. providing a sample, wherein the sample comprises cell-free DNA (cfDNA);
[0089] b. contacting the sample with at least one reagent that binds to a DNA-associated protein;
[0090] c. separating the agent and any bound proteins and cfDNA;
[0091] d. sequencing the isolated cfDNA; and
[0092] e. designating a cfDNA molecule comprising a DNA sequence at an informative genomic location as originating from a cell in a certain cellular state, from a certain tissue, from a certain cell type, or a combination thereof, wherein association of a DNA-associated protein with the informative genomic location indicates the cellular state, tissue of origin, cell type, or a combination thereof, of the cell that released the cfDNA;
[0093] The cell state, tissue of origin, cell type, or a combination thereof of the cell releasing its DNA is thereby determined.
[0094] In another aspect, a method is provided for determining the cell or tissue of origin of cell-free DNA (cfDNA), comprising:
[0095] a. Provide a sample containing cfDNA;
[0096] b. contacting the sample with at least one reagent that binds to a DNA-associated protein;
[0097] c. separating the agent and any bound proteins and cfDNA;
[0098] d. sequencing the isolated cfDNA; and
[0099] e. designating a cfDNA molecule comprising a DNA sequence at an informative genomic location as originating from a cell type or tissue, wherein the association of a DNA-associated protein with the informative genomic location is indicative of the cell type or tissue;
[0100] This allows us to determine the cells or tissues of origin of cfDNA.
[0101] On the other hand, a method for determining the cell or tissue of origin of cell-free DNA is provided, comprising sequencing cfDNA isolated by binding of cfDNA to a DNA-associating protein; wherein the isolated cfDNA contains a DNA sequence at an informative genomic location, and the association of the DNA-associating protein with the informative genomic location indicates a cell type or tissue; thereby determining the cell or tissue of origin of the cfDNA.
[0102] In another aspect, there is provided a method of determining a cell state of a cell in a subject, comprising:
[0103] a. providing a sample from a subject, wherein the sample comprises cfDNA;
[0104] b. contacting the cfDNA with at least one agent that binds to a DNA-associated protein;
[0105] c. separating the agent and any bound protein and cfDNA; and
[0106] d. Sequencing the isolated cfDNA;
[0107] wherein the isolated cfDNA comprises a DNA sequence at an informative genomic location, and association of the DNA-associated protein with the informative genomic location is indicative of a cell state; thereby determining the cell state of the cells in the subject.
[0108] In another aspect, a method is provided for determining a cell state of a cell in a subject at the time of cell death, comprising:
[0109] a. providing a sample from a subject, wherein the sample comprises cfDNA;
[0110] b. contacting the sample with a reagent that binds to a DNA-associated protein;
[0111] c. separating the agent and any bound proteins and cfDNA;
[0112] d. sequencing the isolated cfDNA; and
[0113] e. designating a cfDNA molecule comprising a DNA sequence at an informative genomic location as originating from a cell in a cellular state, wherein association of a DNA-associated protein with the informative genomic location is indicative of the cellular state;
[0114] The cellular state of a cell in the subject at the time of cell death is thereby determined.
[0115] In another aspect, a method for determining the cell state of a cell in a subject is provided, comprising sequencing cfDNA isolated by binding of the cfDNA to a DNA-associating protein; wherein the isolated cfDNA comprises a DNA sequence at an informative genomic location, and association of the DNA-associating protein with the informative genomic location is indicative of the cell state; thereby determining the cell state of the cell in the subject. In some embodiments, the cell is a dead cell in the subject.
[0116] In another aspect, a method is provided for determining the cell state, tissue of origin, or cell type of a cell in a subject at the time of cell death, comprising:
[0117] a. providing a sample from a subject, wherein the sample comprises cfDNA;
[0118] b. contacting the sample with a reagent that binds to a DNA-associated protein;
[0119] c. separating the agent and any bound protein and cfDNA; and
[0120] d. Sequencing the isolated cfDNA;
[0121] wherein the isolated cfDNA comprises a DNA sequence of a tissue- or cell-type-specific binding site for a DNA-associating protein, the tissue- or cell-type-specific binding site is indicative of the cell type or tissue of origin, and association of the DNA-associating protein with the tissue- or cell-type-specific binding site is indicative of the cell state; thereby determining the cell state of the cells in the subject.
[0122] In another aspect, a method is provided for determining the cell state, tissue of origin, or cell type of a cell in a subject at the time of cell death, comprising:
[0123] a. providing a sample from a subject, wherein the sample comprises cfDNA;
[0124] b. contacting the sample with a reagent that binds to a DNA-associated protein;
[0125] c. separating the agent and any bound proteins and cfDNA;
[0126] d. sequencing the isolated cfDNA; and
[0127] e. designating a cfDNA molecule comprising a DNA sequence that is a tissue- or cell-type-specific binding site for a DNA-associating protein as originating from that tissue or cell type and from a cell in a cellular state, wherein association of the DNA-associating protein with the binding site is indicative of the cellular state;
[0128] The cell state and tissue or cell type of origin of cells in the subject at the time of cell death are thereby determined.
[0129] In some embodiments, the method is used to determine the cell state of a cell. In some embodiments, the method is used to determine the tissue of origin of a cell. In some embodiments, the method is used to determine the cell type of a cell. In some embodiments, the sample is from a subject, and the method is used to detect the death of any of cells of a certain tissue, a certain cell type, and cells in a certain cell state in the subject. In some embodiments, the sample is from a subject, and the method is used to detect a disease in the subject, wherein the death of cells of a certain tissue, cells of a certain cell type, or cells in a certain cell state indicates the disease. For non-limiting examples, the death of hepatocytes can indicate liver disease, the death of GI cells can indicate GI cancer, the death of cells with an active interferon response can indicate infection, and the death of beta cells can indicate pancreatic damage / disease.
[0130] In some embodiments, detecting a disease state comprises at least one of: early detection of a disease state, detection of residual disease, and monitoring disease progression. In some embodiments, detecting a disease state comprises early detection. In some embodiments, early detection comprises detection during routine blood work. In some embodiments, early detection comprises detection before symptoms develop. In some embodiments, residual disease is residual metastatic disease. In some embodiments, residual disease is residual cancer after surgery. In some embodiments, disease monitoring comprises pre-treatment monitoring. In some embodiments, disease monitoring comprises post-treatment monitoring. In some embodiments, disease monitoring comprises monitoring for disease recurrence. In some embodiments, disease monitoring comprises monitoring for efficacy of treatment.
[0131] In some embodiments, the cell is dead. In some embodiments, the cell releases its DNA. In some embodiments, the cell that releases its DNA is a dead and / or dying cell or an enucleated cell. In some embodiments, the cell that releases its DNA is a dead and / or dying cell. In some embodiments, the cell death is selected from apoptotic death and necrotic death. In some embodiments, the enucleated cell is an erythrocyte. In some embodiments, the cell that lost its nucleus is an erythroblast. An erythroblast loses its nucleus to become a red blood cell, and therefore, the lost nucleus may be present in the cfDNA.
[0132] In some embodiments, the sample is from a subject. In some embodiments, the cfDNA is from a subject, and the detection of cfDNA molecules from a cell-derived tissue or cell state indicates the detection of cell death in that cell type, tissue, or cell state. In some embodiments, the subject is suspected of having increased cell death. In some embodiments, the subject is not suspected of having increased cell death. In some embodiments, the subject appears healthy and / or is not known to have a disease or condition.
[0133] In some embodiments, the determining is determining the cell state of the cell at the time of its death. In some embodiments, the methods of the present invention are further used to determine the cell state of a cell in a subject at the time of the cell's death, and further comprise designating a cfDNA molecule comprising a DNA sequence at an informative genomic location as originating from a cell in a cell state, wherein association of a DNA-associated protein with the informative genomic location is indicative of the cell state.
[0134] In some embodiments, the sample is a body fluid. In some embodiments, the body fluid is blood. In some embodiments, the body fluid is selected from at least one of blood, serum, gastric fluid, intestinal fluid, saliva, bile, tumor fluid, cerebrospinal fluid, breast milk, semen, urine, vaginal fluid, interstitial fluid, and feces. Standard techniques for cell-free DNA extraction are known to those skilled in the art, a non-limiting example of which is the QIAamp Circulating Nucleic Acid Kit (QIAGEN).
[0135] As used herein, "agent that binds ... " refers to any protein binding molecule or composition. Protein binding is well known in the art and can be assessed by any assay known in the art, including but not limited to yeast-two-hybrid, immunoprecipitation, competitive assays, phage display, tandem affinity purification, and proximity ligation assays. In some embodiments, the reagent is a protein molecule. In some embodiments, the reagent is selected from antibodies or their antigen-binding fragments, proteins, and small molecules. Small molecules that bind to specific proteins are well known in the art and can be used for pull-down experiments. In addition, well-characterized protein-protein interactions can be used for pull-downs. In fact, any reagent that can be used for precipitation, immunoprecipitation (IP), or chromatin immunoprecipitation (ChIP) can be used as the reagent. In some embodiments, the reagent is an antibody or its antigen-binding fragment.
[0136] As used herein, the term "antibody" refers to a polypeptide or group of polypeptides comprising at least one binding domain formed by the folding of polypeptide chains into a three-dimensional binding space, wherein the internal surface shape and charge distribution are complementary to the characteristics of the antigenic determinants of an antigen. Antibodies typically have a tetrameric form, comprising two pairs of identical polypeptide chains, each pair having one "light" chain and one "heavy" chain. The variable regions of each light / heavy chain pair form the antibody binding site. Antibodies can be oligoclonal, polyclonal, monoclonal, chimeric, camelized, CDR-grafted, multispecific, bispecific, catalytic, humanized, fully human, anti-idiotypic, and antibodies that can be labeled in soluble or bound form, as well as fragments—including epitope-binding fragments, variants, or derivatives thereof, alone or in combination with other amino acid sequences. Antibodies can be from any species. The term antibody also includes binding fragments, including, but not limited to, Fv, Fab, Fab', F(ab')2 single-chain antibodies (svFC), dimeric variable regions (diabodies), and disulfide-linked variable regions (dsFv). Specifically, antibodies include immunoglobulin molecules and immunologically active fragments of immunoglobulin molecules, i.e., molecules containing antigen binding sites. Antibody fragments may or may not be fused to another immunoglobulin domain, including but not limited to an Fc region or a fragment thereof. Technicians will further appreciate that other fusion products can be generated, including but not limited to scFv-Fc fusions, variable region (e.g., VL and VH)-Fc fusions, and scFv-scFv-Fc fusions.
[0137] The immunoglobulin molecule can be of any type (e.g., IgGl, IgE, IgM, IgD, IgA, and IgY), class (e.g., IgGl, IgG2, IgG3, IgG4, IgAl, and IgA2), or subclass.
[0138] In some embodiments, one agent is contacted. In some embodiments, at least one agent is contacted. In some embodiments, more than one agent is contacted. In some embodiments, each agent binds to a different DNA-associating protein. In some embodiments, the DNA-associating protein is a sequence-specific DNA binder, and more than one agent is contacted that targets more than one protein. Since the bound target sequences are known, after sequencing the isolated cfDNA, sequences can be assigned to each binding agent based on the target motif present in the sequence. Sequences containing more than one motif can be discarded or included as being bound by multiple DNA-associating proteins.
[0139] In some embodiments, the reagent is conjugated to a physical support. As used herein, the term "physical support" refers to a solid and stable molecule that provides support for the reagent. In some embodiments, the support is a scaffold or scaffolding agent. In some embodiments, the support is a resin. In some embodiments, the support is a bead. In some embodiments, the support is a magnetic bead or a paramagnetic bead. Magnetic beads can be purchased, for example, from Dynabeads or Pierce. In some embodiments, the support is agarose beads. In some embodiments, the support is a protein A / G bead. In some embodiments, the reagent is conjugated to the physical support prior to contact. In some embodiments, the conjugation is a covalent bond. In some embodiments, the conjugation is performed by epoxide chemistry. In some embodiments, the support facilitates separation of the reagent, wherein the separation is separation of the physical support.
[0140] As used herein, the term "DNA-associating protein" refers to any protein that can be precipitated with DNA or obtained along with DNA during precipitation. In some embodiments, the DNA-associating protein binds directly to DNA. In some embodiments, the DNA-associating protein is a component of chromatin. In some embodiments, the DNA-associating protein binds indirectly to DNA. In some embodiments, the DNA-associating protein binds to genomic DNA. In some embodiments, the DNA-binding protein binds to promoters. In some embodiments, the DNA-binding protein binds to gene bodies. In some embodiments, the DNA-binding protein binds to cis- or trans-regulatory elements.
[0141] In some embodiments, the DNA-associating protein binds to DNA and is a non-sequence specific DNA binder. In some embodiments, the DNA-associating protein binds to DNA and is a sequence specific DNA binder or a non-sequence specific DNA binder. Examples of non-sequence specific DNA binders include histones, high mobility group (HMG) proteins, members of the DNA damage repair machinery, and members of the general transcription machinery. The general transcription machinery is well defined and includes, but is not limited to, RNA polymerase, DNA helicase, general cofactors, splicing machinery, and poly A machinery. The DNA damage repair machinery is also well defined and includes, but is not limited to, members of the nucleotide excision repair pathway, base excision repair pathway, and mismatch repair system. In some embodiments, the DNA-associating protein is a modified protein. In some embodiments, the modification is a post-translational modification. In some embodiments, the agent binds to a modified form of the protein. In some embodiments, the agent binds only or primarily to a modified form of the protein.
[0142] In some embodiments, the DNA-associating protein is a histone, a modified histone, or a histone variant. Modifications of histone tails are well known in the art and include, but are not limited to, methylation, acetylation, sulfonylation, ubiquitination, and phosphorylation. The modifications may be multiple, such as trimethylation or polyubiquitination. In some embodiments, the tail may have multiple modifications, such as methylation and phosphorylation. The histone may be one of the core histones H1, H2A, H2B, H3, and H4, or it may be a histone variant, such as, as a non-limiting example, H2A.z, γH2AX, H1T, and H3.3. In some embodiments, the modified or variant histone has an activating function or a repressing function on transcription. In some embodiments, the modified or variant contributes to the formation of euchromatin or heterochromatin. In some embodiments, the modified histone is selected from histone 3 monomethylated lysine 4 (H3K4me1), histone 3 dimethylated lysine 4 (H3K4me2), histone 3 trimethylated lysine 36 (H3K36me3), and histone 3 trimethylated lysine 4 (H3K4me3).
[0143] In some embodiments, the DNA-associating protein binds to DNA and is a sequence-specific DNA binder. Examples of sequence-specific DNA binders include, but are not limited to, transcription factors (TFs), activators, repressors, sequestering agents, DNA-modifying enzymes, and members of the general transcription machinery. In some embodiments, the DNA-associating protein is a transcription factor. In some embodiments, the DNA-associating protein is a sequestering agent. In some embodiments, the transcription factor is selected from activators, repressors, sequestering agents, DNA-modifying enzymes, and members of the general transcription machinery. In some embodiments, the transcription factor is selected from activators, repressors, and sequestering agents. In some embodiments, the transcription factor is a sequestering agent. In some embodiments, the transcription factor is CTCF.
[0144] As used herein, the term "transcription factor" refers to any protein that is not part of the general transcription machinery but controls / regulates the transcription rate of a DNA sequence. In some embodiments, a TF is a factor that binds to a promoter region. Transcription factors are well known in the art as are agents that bind to them. ChIP with TFs is also well known.
[0145] In some embodiments, the agent binds to a transcription factor that binds to a tissue- and / or cell-type-specific enhancer element. In some embodiments, the DNA sequence in the cfDNA is a sequence located at a tissue- and / or cell-type-specific enhancer element. Due to tissue / cell type specificity, the association of TF with the element in the cfDNA indicates that the cfDNA is from the tissue and / or cell type. Since the enhancer element enhances the transcription of a specific target, the target can indicate the cell state of the cell. In this case, an association of TF with a genomic locus can provide information about the source tissue / cell and cell state. A non-limiting example is tissue-specific NF-κB enhancer binding. It is known that NF-κB only binds to specific sites in various tissues (e.g., heart) and mediates inflammation. Therefore, the separation performed with an anti-NF-κB agent and the subsequent identification of tissue-specific enhancer sequences in the cfDNA not only indicate the cell source of the cfDNA, but also indicate that the cell was in an inflammatory state when it died. In some embodiments, the DNA-associated protein is a transcription factor (TF), and the binding site is a TF binding site.
[0146] As used herein, "activator" refers to a protein that increases transcription. In some embodiments, an activator binds to an enhancer element in DNA. In some embodiments, an activator binds to an element proximal or distal to the promoter. As used herein, "repressor" refers to a protein that decreases transcription. In some embodiments, a repressor binds to a repressor element in DNA. In some embodiments, a repressor binds to an element proximal or distal to the promoter.
[0147] As used herein, "isolator" refers to a protein that isolates regions of DNA with different chromatin structures or transcription rates. In some embodiments, the isolator is an enhancer-blocker. In some embodiments, the isolator isolates euchromatin and heterochromatin. Non-limiting examples of isolators include CTCF, gypsy, and BDF1. In some embodiments, the isolator binds to the outside of the promoter and gene body.
[0148] DNA modifying enzymes are well known in the art, and examples include members of the base / nucleotide excision repair machinery, DNA methyltransferases, and DNA demethylases.
[0149] In some embodiments, the DNA-associating protein does not bind DNA. In some embodiments, the protein modifies a protein that binds DNA. Examples include, but are not limited to, histone modifying enzymes and polycomb proteins.
[0150] In some embodiments, the association of a DNA-associating protein with an informative genetic locus is tissue- or cell-type-specific. In some embodiments, the association of a DNA-associating protein with an informative genetic locus is differentiation-specific. In some embodiments, the association of a DNA-associating protein with an informative genetic locus is cell-state-specific. In some embodiments, the association of a DNA-associating protein with an informative genetic locus indicates transcriptional activation, active transcription, transcriptionally active chromatin, or a combination thereof. In some embodiments, the association of a DNA-associating protein with an informative genetic locus indicates transcriptional silencing, lack of transcription, transcriptionally inactive chromatin, or a combination thereof. Transcription need not be at the bound genetic locus, but can be proximal or distal to the gene, as in the case of activators and inhibitors. As used herein, the terms "genetic locus" and "genomic location" are synonymous and refer to a specific region of DNA that can be bound by a protein. In some embodiments, a genetic locus is a TF binding site or some other short sequence of DNA. In some embodiments, locus is 2 to 20, 2 to 16, 2 to 12, 2 to 10, 2 to 8, 2 to 6, 2 to 4, 4 to 20, 4 to 16, 4 to 12, 4 to 10, 4 and 8 or 4 to 6 base pairs. Each possibility represents a separate embodiment of the present invention. In some embodiments, locus is a nucleosome or nucleosome length (~170bp) of DNA. In some embodiments, the genetic locus is between 150 to 190bp or 160 to 180bp.
[0151] As used herein, the terms "informative genomic location" and "informative genetic locus" are used synonymously and refer to a unique DNA sequence at a specific location in the genome that, when associated with a given DNA-associating protein, provides information about the cell in which the association occurs. In some embodiments, it provides information about the tissue of origin or cell type of the cell in which the association occurs. In some embodiments, the location is a tissue- or cell-type-specific binding / association site. In some embodiments, the binding / association is not specific / unique, but is highly enriched in that tissue or cell type. In some embodiments, it provides information about the cell state of the cell in which the association occurs. In some embodiments, it provides information about the tissue of origin and / or cell type and the cell state of the cell in which the association occurs. In some embodiments, it provides information about a disease in a cell. In some embodiments, it provides information about the transcriptional program in a cell.
[0152] As used herein, the term "transcriptional program" refers to a group of genes that act synergistically in transcription. The gene can be actively transcribed and / or inhibited, and / or is inactive, and / or accessible, and / or inaccessible. In some embodiments, the genes are all transcribed together. In some embodiments, the transcriptional program indicates a cell state. In some embodiments, the transcriptional program indicates an active signal transduction pathway. The features of tissue-specific, cell type-specific, cell state-specific and / or transcriptional programs can be found, for example, in the Roadmap Epigenomics project (roadmapepigenomics.org), the Cancer Genome Atlas (cancergenome.nih.gov), the Genotype-Tissue Expression (GTEx) project (gtexportal.org) or the Xena project (xena.ucsc.edu). The table provided herein also provides this feature.
[0153] In some embodiments, the agent binds to a histone. In some embodiments, the agent is an anti-histone antibody or a fragment thereof. In some embodiments, the agent binds to a modified or variant histone. In some embodiments, the agent is an anti-modified histone antibody or a fragment thereof. In some embodiments, the agent is an anti-variant histone antibody or a fragment thereof. In some embodiments, the agent is selected from an anti-histone 3 monomethylated lysine 4 (H3K4me1) antibody and an anti-histone 3 trimethylated lysine 4 (H3K4me3) antibody.
[0154] In some embodiments, the separation comprises separating a physical support to which the reagent is conjugated. In some embodiments, the separation comprises contacting the reagent and the bound protein and DNA with the physical support, and then separating the physical support. In some embodiments, the methods of the present invention comprise ChIP. In some embodiments, the separation comprises ChIP. In some embodiments, the separation comprises a washing step.
[0155] In some embodiments, sequencing comprises sequencing at least an average of 1, 2, 3, 5, or 10 million sequencing reads. Each possibility represents a separate embodiment of the present invention. In some embodiments, sequencing comprises sequencing at least 1, 2, 3, 5, or 10 million sequencing reads. Each possibility represents a separate embodiment of the present invention. In some embodiments, the amplified cfDNA comprises less than 1, 2, 3, 5, or 10 million sequencing reads. Each possibility represents a separate embodiment of the present invention. In some embodiments, the depth of sequencing is at most 1, 2, 3, 5, or 10 million sequencing reads. Each possibility represents a separate embodiment of the present invention.
[0156] In some embodiments, sequencing comprises PCR amplification of cfDNA. In some embodiments, amplification comprises attachment of a barcode or other DNA sequence. In some embodiments, amplification is performed while the cfDNA is still associated with the protein. In some embodiments, amplification is performed while the cfDNA is still associated with the physical support. In some embodiments, amplification is performed without dissociating the cfDNA from the reagents and / or support.
[0157] In some embodiments, the method further comprises comparing the sequencing data with tissue / cell type-specific data for DNA binding proteins, wherein binding of the protein to a sequence specifically bound by the protein of a tissue / cell type indicates that the cfDNA is from that tissue / cell type. Tissue / cell type-specific binding data can be found in sources such as the Encode Consortium, the NIH Epigenome Roadmap Consortium, and the Gene Transcription Regulation Database (to name a few). In some embodiments, the genomic location is within a tissue- or cell-type-specific gene or element. In some embodiments, the protein is associated with active transcription and the genomic location is within a tissue- or cell-type-specific gene or enhancer element. Non-limiting examples include H3K4me3 located in tissue-specific genes or H3K4me1 in enhancers. Non-limiting examples of tissue-specific genes include TNNI3 and MYBPC3 in cardiac cells, and C8a and C8b in hepatocytes. Tissue expression levels of specific proteins can be found on multiple websites, such as the Uniprot database (www.uniprot.org) and the GTEx portal (www.gtexportal.org). Tissue-specific gene expression and regulation can also be found in a number of places, the most notable being the TiGER database (bioinfo.wilmer.jhu.edu / tiger) and the Human Protein Atlas (www.proteinatlas.org). In some embodiments, the protein is associated with silenced transcription and the genomic location is within a repressor element or a gene that is specifically silenced in the tissue or cell type. Tissue-specific protein-DNA binding is well known in the art and can be found in the resources mentioned above. Any informative locus binding can be used to determine the source of cfDNA.
[0158] In some embodiments, sequencing data is compared with at least the top 1, 2, 5, 10, 20, 30, 40, 50, 60, 70, 80, 90 or 100 peaks of the combination in a specific tissue / cell type. Each possibility represents a separate embodiment of the present invention. In some embodiments, only a specific tissue / cell type is studied. In some embodiments, the combined data from at least 1, 2, 3, 5, 10, 15, 20, 25, 30, 35, 40, 45 or 50 tissues or cells / types are used to compare with sequencing data. Each possibility represents a separate embodiment of the present invention.
[0159] In some embodiments, the methods of the present invention further comprise comparing the sequenced cfDNA to at least one genomic location in a tissue, cell type, and / or cell state that is most associated with a DNA-associated protein, and wherein cfDNA having a sequence identical to a DNA sequence within the at least one genomic location is considered to be from that tissue, cell type, and / or cell state. In some embodiments, the at least one genomic location has the greatest unique association with a DNA-associated protein. As used herein, the term "unique association" refers to an association that occurs exclusively or almost exclusively within a tissue, cell type, or cell state. Thus, for example, if the 10 most unique locations are selected, then the locations that have protein binding only in a particular tissue, cell type, or state should be examined, and specifically the 10 with the highest binding should be selected. If the desired number of sites with completely unique binding is not available, then the sites with the most unique binding should be selected. Any determination of maximum uniqueness can be used. Examples include, but are not limited to, the sites with the fewest other tissues to bind to and the sites with the greatest difference in binding between the target tissue and other tissues.
[0160] In some embodiments, the methods of the present invention further comprise comparing the sequenced cfDNA to a DNA-protein association atlas for at least two cell types and / or tissues, wherein the atlas comprises at least one genomic position with the greatest association of DNA-associated proteins in each of the two tissues, cell types, and / or cell states, and wherein cfDNA having a sequence identical to a DNA sequence within the at least one genomic position is considered to be from that tissue, cell type, and / or cell state. In some embodiments, the genomic position has the greatest unique association of DNA-associated proteins. In some embodiments, the atlas comprises at least 1, 2, 3, 5, 7, 10, 15, or 20 cell types and / or tissues. Each possibility represents a separate embodiment of the present invention. In some embodiments, the atlas comprises at least 1, 2, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 75, 80, 90, or 100 genomic positions per tissue, cell type, and / or cell state. Each possibility represents a separate embodiment of the present invention.
[0161] Examples of genomic locations that, when associated with DNA-associating proteins, indicate tissue or cell type can be found in Table 1. Table 1 provides exemplary locations of H3K4me3 tissue-informative locations. In some embodiments, a profile includes all or a portion of the locations in Table 1. In some embodiments, sequencing of cfDNA is compared to Table 1.
[0162] Table 1: Specific locations of H3K4me3 organization
[0163]
[0164]
[0165]
[0166]
[0167]
[0168]
[0169]
[0170]
[0171]
[0172]
[0173]
[0174]
[0175]
[0176]
[0177]
[0178]
[0179]
[0180]
[0181]
[0182]
[0183]
[0184]
[0185]
[0186]
[0187]
[0188]
[0189]
[0190]
[0191]
[0192]
[0193]
[0194] In some embodiments, sequencing data from a subject is deconvoluted by comparison with a DNA-protein association map. In this way, the percentage contribution of different tissues, cell types, and / or cell states to the total cf nucleosomes in a sample can be determined. In some embodiments, deconvolution provides the percentage contribution of only informative cf nucleosomes.
[0195] In some embodiments, cfDNA and cf nucleosomes are analyzed not by comparison with healthy tissue data, but rather by machine learning. Machine learning is well known in the art, and by performing the methods of the present invention on patients with known conditions, a machine learning algorithm can learn to recognize specific disease states and conditions in cfDNA sequences provided when a specific DNA-associated protein is isolated. In some embodiments, at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, or 20 subjects with a specific condition are analyzed before the algorithm can recognize that condition in a new subject.
[0196] As used herein, the term "cell state" refers to a condition or cellular response or pathway that is active in a cell. In some embodiments, a cell state is a cellular condition that leads to and / or causes cell death. In some embodiments, a cell state is the state of a cell just before it dies, but is not the direct cause of that death. In some embodiments, a cell state is a pathway that is active or inactive in a cell at the time of its death. In some embodiments, a cell state is a cellular response that is active or inactive in a cell at the time of its death. In some embodiments, a cell state includes the expression of at least one gene that provides information about the cause of cell death. In some embodiments, a cell state is a disease state. In some embodiments, a cell state includes the expression of at least one gene that provides information about an active pathway in a cell at the time of its death. In some embodiments, at least 1, 2, 3, 4, or 5 genes from a pathway indicate an active pathway. Each possibility represents a separate embodiment of the present invention. Signaling pathways are well known in the art, and online resources for determining the members of various pathways can be found by consulting, for example, Gene Ontology or Thermo Fisher Scientific.
[0197] In some embodiments, determining the cell state comprises determining a cellular pathway that is active in the cell. In some embodiments, determining the cell state comprises determining a transcriptional program that is active in the cell. In some embodiments, determining the cell state comprises determining active transcription of at least one gene that provides information about an active pathway. In some embodiments, at least 1, 2, 3, 4, or 5 genes from a pathway are determined. Each possibility represents a separate embodiment of the invention. In some embodiments, determining the cell state comprises determining the association of a DNA-associated protein with at least one genomic region of genes that regulate a pathway. In some embodiments, the cell state is any one of hypoxia, inflammation, ER stress, mitochondrial stress, dormancy, senescence, interferon response, cycling, malignancy, and calcium flux. Any cell state that can be defined by the expression of a gene or a group of genes can be studied using the methods of the present invention.
[0198] In some embodiments, the methods of the present invention further comprise comparing the sequenced cfDNA to at least one genomic location with the greatest association with a DNA-associated protein during activation of a cellular pathway, and wherein cfDNA having a sequence identical to a DNA sequence within the at least one genomic location indicates activation of the cellular pathway. In some embodiments, the genomic locations with the most unique associations are compared.
[0199] In some embodiments, the methods of the invention further comprise comparing the sequenced cfDNA to a pathway map of at least two cellular pathways, wherein the map comprises at least one genomic position with maximal association of a DNA-associated protein in each of the two cellular pathways, and wherein cfDNA having a sequence identical to a DNA sequence within the at least one genomic position indicates activation of the cellular pathway.
[0200] In some embodiments, the sequenced cfDNA is compared to at least 1, 2, 3, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, or 100 genomic positions. Each possibility represents a separate embodiment of the present invention. In some embodiments, the sequenced cfDNA is compared to at least 10 genomic positions. In some embodiments, the sequenced cfDNA is compared to at least 25 genomic positions. In some embodiments, the designation includes comparing the sequenced DNA to a given number of genomic positions.
[0201] In some embodiments, the atlas is a map of at least 1, 2, 3, 5, 10, 10, 15, 20, 25, 30, 35, 40, 45, or 50 cell types and / or tissues. Each possibility represents a separate embodiment of the present invention. In some embodiments, the atlas is a map of at least 1, 2, 3, 5, 10, 10, 15, 20, 25, 30, 35, 40, 45, or 50 cellular pathways. Each possibility represents a separate embodiment of the present invention. In some embodiments, the atlas comprises at least 1, 2, 3, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, or 100 genomic positions. Each possibility represents a separate embodiment of the present invention. In some embodiments, the atlas comprises at least 2 genomic positions. In some embodiments, the atlas comprises at least 10 genomic positions. In some embodiments, the atlas comprises at least 25 genomic positions.
[0202] In some embodiments, association of a DNA-associating protein with a genomic location indicates active transcription, and the genomic location is within a gene, enhancer element, or at a disease-specific mutation that is specific for a tissue, cell type, or cell state. In some embodiments, association of a DNA-associating protein with a genomic location indicates active transcription, and the genomic location is within a gene, enhancer element, or at a disease-specific mutation that is specific for a tissue, cell type, or cell state. In some embodiments, association of a DNA-associating protein with a genomic location indicates active transcription, and the genomic location is at a disease-specific mutation that is specific for a tissue, cell type, or cell state. In some embodiments, association of a DNA-associating protein with a genomic location indicates active transcription, and the genomic location is at a disease-specific mutation that is specific for a disease. In some embodiments, association of a DNA-associating protein with a genomic location indicates active transcription, and the disease-associated mutation that is specific for a disease is specific for a disease. In some embodiments, association of a DNA-associating protein with a genomic location indicates active transcription, and the disease-associated mutation that is specific for an oncogene is specific for a disease. In some embodiments, association of a DNA-associating protein with a genomic location indicates silenced transcription, and the genomic location is within a repressor element or a gene that is silenced in the tissue, cell type, or cell state, or at a disease-specific mutation that is specific for a disease. In some embodiments, the association of a DNA-associating protein with a genomic location indicates silenced transcription, and the genomic location is within a repressor element or a gene that is silenced in a tissue, cell type, or cell state. In some embodiments, the association of a DNA-associating protein with a genomic location indicates silenced transcription, and the genomic location is at a disease-specific mutation. In some embodiments, the association of a DNA-associating protein with a genomic location indicates silenced transcription, and the disease-associated mutation is within a tumor suppressor gene. Oncogenes and tumor suppressor genes are well known in the art. Examples of oncogenes include, but are not limited to, WNT, RAS, MYC, and ERK. Examples of tumor suppressor genes include, but are not limited to, p53, PTCH, NF1, p27Kip1, and APC.
[0203] As used herein, "cfDNA" refers to any DNA obtained from an organism and present outside the cells of the organism. As used herein, "cf nucleosome" refers to cfDNA and any protein bound and / or associated with cfDNA. In some embodiments, cf nucleosomes comprise cfDNA and cf histones. In some embodiments, cfDNA is associated with DNA-associated proteins. In some embodiments, cfDNA is not naked. In some embodiments, cfDNA is in the sample as cf nucleosomes. In some embodiments, cfDNA is not cross-linked. In some embodiments, the method of the present invention further comprises cross-linking the cfDNA and DNA-associated proteins before contacting. In some embodiments, the method does not comprise cross-linking the cfDNA and DNA-associated proteins before contacting.
[0204] In some embodiments, cfDNA is DNA obtained from an organism and present outside the vesicles in the organism. Cell-free DNA is well known in the art and generally refers to DNA that floats freely in body fluids. This DNA is generally not encapsulated in vesicles, so DNA in transport such as by exosomes or other vesicle transporters is not considered. In some embodiments, cfDNA is DNA from dying and / or dead cells. When a cell dies, the DNA is generally fragmented (cracked, fragmented) and released from the cell when the cell lyses. However, this DNA is not all immediately removed or cleared, so it can remain in the organism. DNA from dead cells often enters the bloodstream. In some embodiments, cracked vesicular chromatin is also included in cfDNA.
[0205] In some embodiments, the cfDNA is mammalian cfDNA. In some embodiments, the cfDNA is human cfDNA. In some embodiments, the cfDNA is from a mammalian or human genome. In some embodiments, the cfDNA is fetal DNA. In some embodiments, the DNA is viral DNA. In some embodiments, the DNA is bacterial DNA. In some embodiments, the DNA is fungal DNA. In some embodiments, the DNA is parasitic DNA. In some embodiments, the DNA is from a pathogen. In some embodiments, the DNA is from an organism living in a healthy subject. In some embodiments, the cfDNA is extracted from a body fluid. In some embodiments, the providing comprises providing a body fluid and isolating cfDNA from the body fluid. In some embodiments, the providing comprises providing a body fluid comprising cfDNA and performing the contacting in the body fluid. In some embodiments, the methods of the present invention are performed on as little as 0.00001, 0.00005, 0.0001, 0.0005, 0.001, 0.005, 0.01, 0.05, 0.1, 0.5, 1, 1.5, 2, 2.5, or 3 ml of body fluid. In some embodiments, the method of the present invention is performed on as little as 2 ml of body fluid. Each possibility represents a separate embodiment of the present invention. In some embodiments, the method of the present invention is performed with less than 0.00001, 0.00005, 0.0001, 0.0005, 0.001, 0.005, 0.01, 0.05, 0.1, 0.5, 1, 1.5, 2, 2.5, 3, 3.5, 4, 4.5, or 5 ml of body fluid. Each possibility represents a separate embodiment of the present invention. In some embodiments, the body fluid is blood.
[0206] In some embodiments, cfDNA is fetal cell-free DNA (cffDNA). In some embodiments, the methods are used for non-invasive fetal monitoring. In some embodiments, the subject is the mother of the fetus. In some embodiments, the methods are used to determine the cellular state of fetal cells. In some embodiments, the methods are used to determine fetal disease. In some embodiments, the methods are used to determine fetal genetic abnormalities. In some embodiments, the methods are used to determine the source of cell death in the fetus.
[0207] Because the half-life of cfDNA in an organism is short, it provides a snapshot of the cell death that occurred in the organism at that time. In some embodiments, the method of the present invention detects cell death that occurred within 1 minute, 2 minutes, 3 minutes, 4 minutes, 5 minutes, 10 minutes, 15 minutes, 20 minutes, 25 minutes, 30 minutes, 35 minutes, 40 minutes, 45 minutes, 50 minutes, 55 minutes, 1 hour, 2 hours, 3 hours, 6 hours, 12 hours, 18 hours, 1 day, 2 days, 3 days, 4 days, 5 days, 6 days, 1 week, 2 weeks, 3 weeks or a month from the time when the sample was collected from the object. Each possibility represents a separate embodiment of the present invention. In some embodiments, the method of the present invention detects cell death that occurred before the sample was collected from the object. In some embodiments, the method of the present invention further includes extracting a sample from the object before the provision, wherein the sample comprises cfDNA. In some embodiments, the method of the present invention further includes freezing or maintaining the sample at about 4 degrees after obtaining the sample from the object and before contacting it. By freezing or keeping the sample cold, the methods of the present invention can still detect cell death that occurs immediately before the sample is collected from the subject.
[0208] In some embodiments, at least 0.00001, 0.00005, 0.0001, 0.0005, 0.001, 0.005, 0.01, 0.05, 0.1, 0.5, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 ng of cfDNA is provided. Each possibility represents a separate embodiment of the invention. In some embodiments, as little as 0.00001, 0.00005, 0.0001, 0.0005, 0.001, 0.005, 0.01, 0.05, 0.1, 0.5, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 ng of cfDNA is provided. Each possibility represents a separate embodiment of the invention. In some embodiments, at least 50 ng is provided. In some embodiments, as little as 50 ng is provided. In some embodiments, at least 7 ng is provided. In some embodiments, as little as 7 ng is provided. In some embodiments, at least 0.5 ng is provided. In some embodiments, as little as 0.5 ng is provided.In some embodiments, the cfDNA provided is 0.1 to 1000, 0.1 to 900, 0.1 to 800, 0.1 to 700, 0.1 to 600, 0.1 to 500, 0.1 to 400, 0.1 to 300, 0.1 to 250, 0.1 to 200, 0.1 to 150, 0.1 to 100, 0.1. to 90, 0.1 to 80, 0.1 to 70, 0.1 to 60, 0.1 to 50, 0.1 to 40 , 0.1 to 30 or 0.1 to 20 ng, 0.1 to 10 ng, 0.1 to 5 ng, 0.1 to 1 ng, 0.5 to 1000, 0.5 to 900, 0.5 to 800, 0.5 to 700, 0.5 to 600, 0.5 to 500, 0.5 to 400, 0.5 to 300, 0.5 to 250, 0.5 to 200, 0.5 to 150, 0.5 to 100, 0.5 to 90, 0.5 to 80, 0.5 to 70, 0.5 to 60, 0.5 to 50, 0.5 to 40, 0.5 to 30 or 0.5 to 20 ng, 0.1 to 10 ng, 0.5 to 5 ng, 0.5 to 1 ng, 1 to 1000, 1 to 900, 1 to 800, 1 to 700, 1 to 600, 1 to 500, 1 to 400, 1 to 300, 1 to 250, 1 to 200, 1 to 150, 1 to 100, 1 to 90, 1 to 80, 1 to 70, 1 to 60 , 10 to 50, 10 to 40, 10 to 30 or 10 to 20 ng, 10 to 1000, 10 to 900, 10 to 800, 10 to 700, 10 to 600, 10 to 500, 10 to 400, 10 to 300, 10 to 250, 10 to 200, 10 to 150, 10 to 100, 10 to 90, 10 to 80, 10 to 70, 10 to 60, 10 to 50, 10 to 40, 10 to 30 or 10 to 20 ng. Each possibility represents a separate embodiment of the present invention. In some embodiments, at most 0.00001, 0.00005, 0.0001, 0.0005, 0.001, 0.005, 0.01, 0.05, 0.1, 0.5, 1, 5, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 400, 500, 600, 700, 800, 900, or 1000 ng of cfDNA is provided. Each possibility represents a separate embodiment of the invention. 1000 genomes is approximately equivalent to 6.6 ng of cfDNA.
[0209] In some embodiments, the cfDNA comprises at least 0.00001, 0.00005, 0.0001, 0.0005, 0.001, 0.005, 0.01, 0.05, 0.1, 0.5, 1, 2, 5, 10, 50, 100, 200, 300, 500, 700, 800, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, or 10000 genomes. Each possibility represents a separate embodiment of the invention. In some embodiments, the cfDNA comprises as few as 0.00001, 0.00005, 0.0001, 0.0005, 0.001, 0.005, 0.01, 0.05, 0.1, 0.5, 1, 2, 5, 10, 50, 100, 200, 300, 500, 700, 800, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, or 10000 genomes. Each possibility represents a separate embodiment of the invention. In some embodiments, the cfDNA comprises 0.1 to 10,000, 0.1 to 9,000, 0.1 to 8,000, 0.1 to 7,000, 0.1 to 6,000, 0.1 to 5,000, 0.1 to 4,000, 0.1 to 3,000, 0.1 to 2,000, 0.1 to 1,000, 1 to 10,000, 1 to 9,000, 1 to 8,000, 1 to 7,000, 1 to 6,000, 1 to 5,000, 1 to 4000, 1 to 3000, 1 to 2000, 1 to 1000, 5 to 10000, 5 to 9000, 5 to 8000, 5 to 7000, 5 to 6000, 5 to 5000, 5 to 4000, 5 to 3000, 5 to 2000, 5 to 1000, 10 to 10000, 10 to 9000, 10 to 8000, 10 to 7000, 10 to 6000, 10 to 5000, 10 to 40 00, 10 to 3000, 10 to 2000, 10 to 1000, 100 to 10000, 100 to 9000, 100 to 8000, 100 to 7000, 100 to 6000, 100 to 5000, 100 to 4000, 100 to 3000, 100 to 2000, 100 to 1000, 500 to 10000, 500 to 9000, 500 to 8000, 500 to 70 00, 500 to 6000, 500 to 5000, 500 to 4000, 500 to 3000, 500 to 2000, 500 to 1000, 1000 to 10000, 1000 to 9000, 1000 to 8000, 1000 to 7000, 1000 to 6000, 1000 to 5000, 1000 to 4000, 1000 to 3000, 1000 to 2000 genomes. Each possibility represents a separate embodiment of the invention.
[0210] In some embodiments, as little as 0.00001, 0.00005, 0.0001, 0.0005, 0.001, 0.005, 0.01, 0.05, 0.1, 0.2, 0.3, 0.5, 1, 2, 3, 4, 5, or 10% of the cfDNA is from the cell type, tissue, or cells in the cell state. Each possibility represents a separate embodiment of the invention. In some embodiments, as little as 0.1% of the cfDNA is from the cell type, tissue, or cells in the cell state. In some embodiments, as little as 1% of the cfDNA is from the cell type, tissue, or cells in the cell state. In some embodiments, the limit of detection of the method is 0.1% of the cfDNA in the sample is from the cell type, tissue, or cell state. In some embodiments, the limit of detection of the method is 1% of the cfDNA in the sample is from the cell type, tissue, or cell state. In some embodiments, as little as 0.1% of the cfDNA is from the cell type, the tissue, or the cell state, and at least 45 peaks corresponding to the cell type, tissue, or cell state are detected. In some embodiments, as little as 0.1% of the cfDNA is from the cell type, the tissue, or the cell state, and at least 45-200 peaks corresponding to the cell type, tissue, or cell state are detected. In some embodiments, as little as 1% of the cfDNA is from the cell type, the tissue, or the cell state, and at least 25 peaks corresponding to the cell type, tissue, or cell state are detected. In some embodiments, analyzing at least 25 peaks provides a detection limit of 1% of the cfDNA from the cell type, the tissue, or the cell state. In some embodiments, analyzing at least 45 peaks provides a detection limit of 0.1% of the cfDNA from the cell type, the tissue, or the cell state. In some embodiments, analyzing a certain number of peaks includes detecting cfDNA from at least that number of peaks. In some embodiments, the cfDNA comprises 0.001-10, 0.001-5, 0.001-3, 0.001-2, 0.001-1.5, 0.001-1, 0.01-10, 0.01-5, 0.01-3, 0.01-2, 0.01-1.5, 0.01-1, 0.1-10, 0.1-5, 0.1-3, 0.1-2, 0.1-1.5, 0.1-1, 0.5-10, 0.5-5, 0.5-3, 0.5-2, 0.5-1.5 0.5-1, 1-10, 1-5, 1-3, or 1-2% cfDNA from said cell type, said tissue, or said cell state. Each possibility represents a separate embodiment of the invention.In some embodiments, the cfDNA comprises 0.1-1% cfDNA from said cell type, said tissue, or said cell state. In some embodiments, the cfDNA comprises 0.1-3% cfDNA from said cell type, said tissue, or said cell state.
[0211] In some embodiments, sequencing is low depth. In some embodiments, the depth of sequencing is less than 1,000,000,000, 750,000,000, 500,000, 400,000,000, 300,000,000, 200,000,000, 100,000, 90,000,000, 800,000, 700,000, 600,000, 500,000, 400,000, 300,000, 200,000, 100,000, 500,000, 1000,000, 1000,000, 500,000 or 1000 reads. Each possibility represents a separate embodiment of the present invention. In some embodiments, the depth of sequencing is less than 10,000,000 reads. In some embodiments, the depth of sequencing is less than 1,000,000 reads. It will be understood by those skilled in the art that as the amount of information increases, the detection limit declines. Furthermore, as the amount of input data increases (increasing the number of reagents (i.e., antibodies for ChIP), increasing the number of informative loci for a given cell type / tissue / state, increasing the amount of cfDNA from said cell type / tissue / state), the required sequencing depth also decreases, and the detection limit decreases.
[0212] In some embodiments, the providing comprises providing a body fluid comprising cfDNA. In some embodiments, the contacting occurs in a body fluid. In some embodiments, the contacting comprises providing a body fluid and isolating cfDNA from the body fluid. In some embodiments, the body fluid is selected from the group consisting of: blood, serum, gastric fluid, intestinal fluid, saliva, bile, tumor fluid, interstitial fluid, breast milk, cerebrospinal fluid, urine, semen, vaginal fluid, and feces. In some embodiments, the body fluid is any body fluid containing cfDNA. In some embodiments, the body fluid is blood. In some embodiments, the body fluid is any of whole blood, partially dissolved whole blood, plasma, or partially processed whole blood.
[0213] In some embodiments, the blood sample can be obtained by standard techniques, such as using a needle and syringe. In another embodiment, the blood sample is a peripheral blood sample. Alternatively, the blood sample can be a separated portion of peripheral blood, such as a plasma sample. In another embodiment, after obtaining the blood sample, standard techniques well known to those skilled in the art can be utilized to extract total DNA from the sample. In some embodiments, intact cells are removed before DNA extraction, thereby only extracting free-floating DNA. Intact cells can be removed by any method known in the art, such as as non-limiting examples, by centrifugation or by gradient separation, such as by Ficol gradient separation. The non-limiting examples of DNA extraction is FlexiGene DNA test kit (QIAGEN). The standard techniques used to receive cell-free DNA extraction are known to the technician, and its non-limiting examples is QIAamp Circulating Nucleic Acid test kit (QIAGEN).
[0214] In some embodiments, sequencing is next generation sequencing. Next generation sequencing, also referred to as high throughput sequencing or large-scale parallel sequencing, is any sequencing method that realizes that base pairs from DNA or RNA samples are carried out to rapid high throughput sequencing. In some embodiments, sequencing is high throughput sequencing. In some embodiments, sequencing is large-scale parallel sequencing. Such sequencing is well known in the art and can include using Illumina arrays, holes and nanopore sequencers and ion torrents as non-limiting examples. For non-limiting examples, sequencing machines such as IlluminaNextseq500 machines can be used, and Illumina 500 / 550V2 test kits can be used for processing. In some embodiments, sequencing is whole genome sequencing. In some embodiments, only part of the genome is sequenced. In some embodiments, a chip or array related to only part of the genome is used for next generation sequencing.
[0215] In some embodiments, sequencing is methylation-sensitive sequencing. In some embodiments, the methods of the present invention further include bisulfite conversion prior to sequencing. During sequencing, the methylation state of the DNA can also be identified. In this way, protein-DNA association data can also be combined with DNA methylation data. This can provide further information about the gene activity of the cell at the time of its death, which can provide an understanding of the cellular state of the source cell, tissue, or cell.
[0216] In some embodiments, the methods of the present invention can be used to determine the source of cfDNA even when cfDNA from one tissue / cell type accounts for a very small percentage of the total cfDNA. In some embodiments, cfDNA from a tissue and / or cell type accounts for as little as 0.0001%, 0.0005%, 0.001%, 0.005%, 0.01%, 0.05%, 0.1%, 0.5%, 1%, 1.5%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9% or 10% of the total cfDNA. Each possibility represents a separate embodiment of the present invention. In some embodiments, cfDNA of a tissue and / or cell type comprises more than 0.0001%, 0.0005%, 0.001%, 0.005%, 0.01%, 0.05%, 0.1%, 0.5%, 1%, 1.5%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, or 10% of total cfDNA. Each possibility represents a separate embodiment of the present invention. In some embodiments, cfDNA of a tissue and / or cell type comprises less than 0.0001%, 0.0005%, 0.001%, 0.005%, 0.01%, 0.05%, 0.1%, 0.5%, 1%, 1.5%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, or 10% of total cfDNA. Each possibility represents a separate embodiment of the present invention.In some embodiments, cfDNA of a tissue and / or cell type accounts for 0.0001%-10%, 0.001%-10%, 0.01%-10%, 0.1%-10%, 0.5%-10%, 1%-10%, 1.5%-10%, 2%-10%, 0.0001%-9%, 0.001%-9%, 0.01%-9%, 0.1%-9%, 0.5%-9%, 1%-9%, 1.5%-9%, 2%-9%, 0.0001%-8% of the total cfDNA , 0.001%-8%, 0.01%-8%, 0.1%-8%, 0.5%-8%, 1%-8%, 1.5%-8%, 2%-8%, 0.0001%-7%, 0.001%-7%, 0.01%-7%, 0.1%-7%, 0.5%-7%, 1%-7%, 1.5%-7%, 2%-7%, 0.0001%-6%, 0.001%-6%, 0.01%-6%, 0.1%-6%, 0.5%-6%, 1%-6%, 1.5%-6%, 2% -6%, 0.0001%-5%, 0.001%-5%, 0.01%-5%, 0.1%-5%, 0.5%-5%, 1%-5%, 1.5%-5%, 2%-5%, 0.0001%-4%, 0.001%-4%, 0.01%-4%, 0.1%-4%, 0.5%-4%, 1%-4%, 1.5%-4%, 2%-4%, 0.0001%-3%, 0.001%-3%, 0.01%-3%, 0.1%-3%, 0.5%-3%, 1% Each possibility represents a separate embodiment of the present invention.
[0217] In some embodiments, the contacting is incubating the reagent in a body fluid containing cfDNA. In some embodiments, the contacting is incubating the reagent in blood containing cfDNA. In some embodiments, the contacting is incubating the reagent and cfDNA in a binding / incubation solution. Buffers for performing ChIP, particularly incubation buffers, are well known in the art. Such buffers can be purchased from companies that sell ChIP kits, such as Abeam and Cell Signaling Technology.
[0218] In some embodiments, the contacting is performed with constant mixing. In some embodiments, the contacting is performed with constant rotation. In some embodiments, the contacting is performed at room temperature or 4 degrees. In some embodiments, the contacting is performed on ice. In some embodiments, the contacting is performed for at least 1, 2, 3, 4, 5, 6, 12, 18, or 24 hours. Each possibility represents a separate embodiment of the present invention. In some embodiments, the contacting is performed for a period of time sufficient to allow the agent to bind to the DNA-associated protein. In some embodiments, the contacting is performed for a period of time sufficient to allow the agent to bind to at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 97%, or 99% of the provided DNA-associated protein. Each possibility represents a separate embodiment of the present invention.
[0219] In some embodiments, the methods of the present invention are used to detect a disease state or condition in a subject in need thereof, wherein the cfDNA is from the subject. In some embodiments, the methods of the present invention are used to diagnose a disease and / or condition in a subject in need thereof, wherein the cfDNA is from the subject. In some embodiments, the methods of the present invention are used to diagnose an increased risk of a disease or condition. Those skilled in the art will recognize that many, if not all, disease states induce cell death in tissues or cells expressing the disease. Therefore, identifying the source of cell death is a surrogate for the disease. In some embodiments, the disease state or condition is selected from cardiac disease or injury and liver disease or injury. In some embodiments, the disease state or condition is selected from cardiac disease or injury, liver disease or injury, and cancer. In some embodiments, the disease state is cancer. In some embodiments, the disease state is a precancerous condition. In some embodiments, the disease state is cancer or a precancerous condition. In some embodiments, the disease state or condition is selected from cardiac arrest and hepatic shock. In some embodiments, the disease state is brain injury. In some embodiments, the disease state is bacteremia. In some embodiments, the disease state is infection. In some embodiments, the disease state or condition is selected from cancer, neurodegenerative disease, infection, tissue damage, inflammation, autoimmune disease, arthritis, liver inflammation, intestinal inflammation, autoimmune disease, bacteremia, tissue damage caused by drug side effects, tissue necrosis, and diabetes. In some embodiments, the neurodegenerative disease is Parkinson's disease or Alzheimer's disease. In some embodiments, the autoimmune disease is lupus or multiple sclerosis. In some embodiments, the disease is cancer, and the methods of the present invention determine the cell or tissue of origin of the cancer. It will be fully understood by those skilled in the art that the association of active transcription of cfDNA from oncogenes and proteins of the sequence indicates cancer or precancerous state. In addition, the association of transcriptional silencing of cfDNA from tumor suppressors and proteins of the sequence also indicates cancer or precancerous state. Similarly, activation or repression of enhancer regions of oncogenes and tumor suppressors, respectively, also indicates cancer or precancerous state.
[0220] In some embodiments, the methods of the present invention further comprise performing steps ad again using a reagent that binds to a second DNA-associating protein, wherein the second DNA-associating protein is a different protein from the DNA-associating protein already used. In some embodiments, the second DNA-associating protein is different from the first DNA-associating protein. In some embodiments, the methods of the present invention can be repeated at least 2, 3, 4, 5, 6, 7, 8, 9, or 10 times, each time with a different DNA-associating protein. Each possibility represents a separate embodiment of the present invention.
[0221] In some embodiments, the method of the present invention includes contacting the sample with at least two reagents, wherein each reagent is bound to a physical carrier, and the carrier comprises a short DNA tag unique to each reagent, wherein after the separated cfDNA is sequenced, the short DNA tag identifies the reagent that separated the cfDNA. In some embodiments, the DNA tag is connected to the cfDNA molecule before sequencing. In some embodiments, the carrier is a bead conjugated to a single reagent and a short DNA tag unique to the reagent. In some embodiments, the reagent is an antibody, and the short DNA tag is a DNA barcode. This can, for example, be a paramagnetic bead covalently bound to a ChIP antibody (such as H3K4me1) and combined with a barcode for identifying the DNA associated with H3K4me1, and other paramagnetic beads covalently bound to a ChIP antibody (such as H3K4me3) added to the sample, connected to the cfDNA and sequenced simultaneously, and combined with a barcode for identifying the DNA associated with H3K4me3.
[0222] In some embodiments, the method further includes treating the subject. In some embodiments, the treatment is directed to the detected disease. In some embodiments, the treatment is an appropriate treatment based on the cell state, source tissue, cell type, or a combination thereof of the dead cells in the subject. It will be understood by those skilled in the art that if, for example, cancer is found in a specific organ, treatment can be adjusted for that type of cancer. Similarly, if a specific pathway is active in a cancer or disease, one treatment approach may be more suitable than another. For example, the detection of active transcription of the long non-coding RNA EGFR-AS1 (which mediates cancer addiction to EGFR and, when highly expressed, can render tumors insensitive to EGFR inhibition by anti-EGFR antibody treatment) will indicate that EGFR inhibitory therapy should be avoided.
[0223] Disease Detection
[0224] In another aspect, a method of detecting a disease state in a subject is provided, the method comprising:
[0225] a. providing a sample from a subject, wherein the sample comprises cfDNA;
[0226] b. contacting the sample with at least one reagent that binds to a DNA-associated protein;
[0227] c. separating the agent and any bound proteins and cfDNA;
[0228] d. sequencing the isolated cfDNA; and
[0229] e. designating cfDNA molecules containing disease-associated mutations as originating from cells in a disease state;
[0230] The disease state of the subject is thereby detected.
[0231] In another aspect, a method for improving disease detection in cfDNA from a subject is provided, the method comprising performing chromatin immunoprecipitation on cfDNA from the subject and then detecting disease in the immunoprecipitated cfDNA.
[0232] In some embodiments, chromatin immunoprecipitation comprises:
[0233] a. contacting cfDNA from a subject with at least one agent that binds to a DNA-associated protein; and
[0234] b. Separating the agent and any bound protein and cfDNA.
[0235] In some embodiments, disease detection comprises sequencing of cfDNA. In some embodiments, disease detection comprises sequencing of cfDNA from a subject. In some embodiments, disease detection comprises sequencing of immunoprecipitated cfDNA. In some embodiments, disease detection and / or sequencing comprises amplification of cfDNA. In some embodiments, disease detection and / or sequencing does not include amplification of cfDNA. In some embodiments, disease detection in cfDNA from a subject comprises amplification of cfDNA. In some embodiments, non-improved disease detection comprises amplification of cfDNA. In some embodiments, disease detection in immunoprecipitated cfDNA does not include amplification of immunoprecipitated cfDNA. In some embodiments, improved disease detection does not include amplification of immunoprecipitated cfDNA. In some embodiments, amplification is PCR amplification. In some embodiments, amplification is non-specific amplification. In some embodiments, amplification is amplification of disease-associated sequences.
[0236] As used herein, "disease-associated mutations" refer to DNA mutations that are known to cause or increase the risk of developing a disease. Disease-associated mutations are well known and include, for example, deletions of 1522A, 1523T, and 1524C of CFTR in cystic fibrosis (F508 deletion); 1226A to G (N370S) of the GBA locus in Gaucher disease; mutations of SERPINA1 in α1-antitrypsin deficiency; mutations of HBB in β-thalassemia; and mutations of PSEN1 in Alzheimer's disease. A variety of disease-associated mutations are known in cancer, some of which are common to multiple cancer types and some of which are unique to specific cancers. Mutations in p53, MYC, BREF, BRCA (to name a few) are well known in the art. Groups of disease-associated mutations can also be studied. In some embodiments, at least 1, 2, 3, 5, 7, 10, 12, 15, 17, or 10 mutations are studied. Each possibility represents a separate embodiment of the present invention. In some embodiments, PCR with mutation-specific primers is used instead of sequencing. Any method for detecting DNA mutations can be used instead of sequencing; however, sequencing has the advantage of being able to examine multiple mutations (including a panel of mutations) simultaneously.
[0237] In some embodiments, the association of a DNA-associating protein with a disease-associated mutation indicates a disease state. Those skilled in the art will readily appreciate that the association of a protein indicative of an actively transcribed mutation in a gene coding region will indicate that the mutated gene is being transcribed and will indicate a cancerous or precancerous state. Similarly, a mutation in a regulatory region will associate with a protein indicative of that regulatory region and, therefore, will also indicate a cancerous or precancerous state.
[0238] In some embodiments, the disease-associated mutation is in a coding region of a gene, and association of the DNA-associating protein with the DNA indicates active transcription. In some embodiments, the gene is an oncogene or a tumor suppressor gene. In some embodiments, the disease-associated mutation is in a regulatory region, and association of the DNA-associating protein with the DNA indicates a regulatory region.
[0239] In some embodiments, the method of the present invention further comprises performing step b. again using a reagent that binds to a second DNA-associated protein, and wherein the second DNA-associated protein is different from the first DNA-associated protein. If multiple mutations are to be investigated and are located in different genomic regions (e.g., gene bodies and enhancers), ChIP can be repeated for different DNA-associated proteins.
[0240] Immunoprecipitation of only a portion of cfDNA (such as enrichment of gene body sequences of actively transcribed genes by H3K36me3 immunoprecipitation) greatly increases the concentration of informative DNA. As a result, sequencing can be performed significantly more cheaply and using fewer reagents. In addition, this method allows the detection of mutations associated with specific genomic annotations (active genes, active promoters, active enhancers, etc.) without pre-defining a limited set of genomic positions (such as a set of cancer risk genes) and designing specific reagents to amplify and / or detect those pre-defined sequences. By first reducing the effective size of the genome sequencing portion (reduced to only the immunoprecipitated portion), the sequencing cost is greatly reduced. In addition, since there are fewer repetitive sequences and non-informative sequences, there is less background and fewer false positive results. Finally, sequencing of a given depth provides more reads than informative sequences.
[0241] In some embodiments, the improvement comprises at least one of: reducing the signal-to-noise ratio, increasing confidence in positive disease detection, reducing false disease detection, and accurately detecting disease with less cfDNA from the subject. In some embodiments, the increase is at least 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, or 1000%. Each possibility represents a separate embodiment of the present invention. In some embodiments, the reduction is at least 10, 20, 30, 40, 50, 60, 70, 80, 90, 85, 97, 99, or 100% reduction. Each possibility represents a separate embodiment of the present invention.
[0242] In some embodiments, the less cfDNA from the subject is less than 1000, 900, 800, 700, 600, 500, 400, 300, 200, 100, 90, 80, 70, 60, or 50 ng of cfDNA. Each possibility represents a separate embodiment of the present invention. In some embodiments, the same accuracy can be achieved with less cfDNA compared to a larger amount of cfDNA. In some embodiments, the same sequencing coverage at the mutation can be achieved with less cfDNA compared to the coverage achieved with a larger amount of cfDNA.
[0243] Computer program product
[0244] In another aspect, a computer program product for determining a cell or tissue of origin of cell-free DNA (cfDNA) is provided, comprising a non-transitory computer-readable storage medium embodied with program code, the program code being executable by at least one hardware processor to:
[0245] a. Sequencing or accessing cfDNA isolated using reagents that bind to DNA-associated proteins;
[0246] b. assigning a cfDNA molecule from the cfDNA to a cell or tissue of origin by comparing the DNA sequence of the molecule to sequences associated with DNA-associated proteins in the cell type or tissue; and
[0247] c. Provide output regarding the source cell or tissue of the cfDNA.
[0248] In another aspect, a computer program product is provided for determining a cellular state of a cell in a subject at the time of cell death, comprising a non-transitory computer-readable storage medium having program code embodied thereon, the program code being executable by at least one hardware processor to:
[0249] a. sequencing or accessing cfDNA from a subject isolated with a reagent that binds to a DNA-associated protein;
[0250] b. assigning a cfDNA molecule from the cfDNA to a cell state by comparing the DNA sequence of the molecule to sequences associated with DNA-associated proteins in the cell state; and
[0251] c. Providing output regarding the cellular state of a cell in the subject at the time of cell death.
[0252] In another aspect, a system for determining the cell or tissue of origin of cfDNA is provided, comprising:
[0253] a. One or more devices for sequencing cfDNA isolated using a reagent that binds to a DNA-associated protein;
[0254] b. Processor; and
[0255] c. A storage medium comprising a computer application that, when executed by a processor, is configured to:
[0256] i. Sequencing or accessing cfDNA isolated using reagents that bind to DNA-associated proteins;
[0257] ii. assigning a cfDNA molecule from the cfDNA to a cell or tissue of origin by comparing the DNA sequence of the molecule to sequences associated with DNA-associated proteins in that cell type or tissue; and
[0258] iii. Output the cfDNA from the processor to the source cell or tissue.
[0259] In another aspect, a system for determining a cellular state of a cell in a subject at the time of cell death is provided, comprising:
[0260] a. One or more devices for sequencing cfDNA isolated using a reagent that binds to a DNA-associated protein;
[0261] b. Processor; and
[0262] c. A storage medium comprising a computer application that, when executed by a processor, is configured to:
[0263] i. Sequencing or accessing cfDNA isolated using reagents that bind to DNA-associated proteins;
[0264] ii. assigning a cfDNA molecule from the cfDNA to a cell state by comparing the DNA sequence of the molecule to sequences associated with DNA-associated proteins in the cell state; and
[0265] iii. Output the cfDNA from the processor to the source cell or tissue.
[0266] In another aspect, a computer program product for detecting a disease state in a subject is provided, comprising a non-transitory computer-readable storage medium embodied with program code thereon, the program code being executable by at least one hardware processor to
[0267] a. Assigning a cfDNA molecule from the cfDNA to a disease state by comparing the DNA sequence of the molecule to mutation sequences associated with the disease state;
[0268] b. Provide output regarding the subject's disease status.
[0269] A computer-readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. A computer-readable storage medium can be, for example, but not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. A non-exhaustive list of more specific examples of computer-readable storage media includes the following: a portable computer diskette (floppy disk), a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punch card or a raised structure in a groove on which instructions are recorded, and any suitable combination thereof. As used herein, a computer-readable storage medium will not be understood to be a temporary signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagated by a waveguide or other transmission medium (e.g., a light pulse passing through a fiber optic cable), or an electrical signal transmitted by a wire.
[0270] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to a corresponding computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and sends the computer-readable program instructions to be stored in a computer-readable storage medium within the corresponding computing / processing device.
[0271] The computer-readable program instructions for performing the operations of the present invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages (including target-oriented programming languages such as Java, Smalltalk, C++, etc.) and conventional procedural programming languages such as the "C" programming language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet, using an Internet service provider). In some embodiments, electronic circuits, including, for example, programmable logic circuits, field programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), may be configured to execute computer-readable program instructions by utilizing state information of the computer-readable program instructions to personalize the electronic circuit to perform aspects of the present invention.
[0272] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device to produce a machine, so that the instructions (which are executed by the processor of the computer or other programmable data processing device) establish means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in such a computer-readable storage medium: it can direct the computer, programmable data processing device, and / or other device to work in a specific manner, so that the computer-readable storage medium in which the instructions are stored includes an article of manufacture containing instructions for implementing various aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0273] Embodiments may include computer programs that implement the functions described and illustrated herein, wherein the computer programs are implemented in a computer system comprising instructions stored in a machine-readable medium and a processor that executes the instructions. However, it is apparent that there are many different ways to implement embodiments in computer programming, and embodiments should not be construed as limited to any one set of computer program instructions. Furthermore, a skilled programmer will be able to write a computer program that implements one or more disclosed embodiments described herein. Therefore, it is considered unnecessary to disclose a specific set of program code instructions to fully understand how to build and use the embodiments. Furthermore, it will be understood by those skilled in the art that one or more aspects of the embodiments described herein may be implemented by hardware, software, or a combination thereof, such as may be implemented in one or more computing systems. Furthermore, any reference to an action performed by a computer should not be construed as being performed by a single computer, as more than one computer may perform the action.
[0274] The device for sequencing refers to a combination of components that allows the sequence of a section of DNA to be determined. In some embodiments, the testing device allows high-throughput sequencing of DNA. In some embodiments, the testing device allows large-scale parallel sequencing of DNA. The components may include any of those described above for sequencing methods.
[0275] In certain embodiments, the system or test kit further includes a display for processor output.
[0276] Multiplexing
[0277] In another aspect, a solid support is provided that comprises a capture agent and a barcoding reagent.
[0278] As used herein, the term "capture agent" refers to a molecule that binds to a protein and can thereby capture and retain the protein on a solid support. In some embodiments, the capture agent is a small molecule. In some embodiments, the capture agent is a protein. In some embodiments, the capture protein captures a second protein through a protein-protein interaction. In some embodiments, the capture protein is an antibody or an antigen-binding fragment thereof. The capture agent can be any molecule that specifically binds to chromatin or nucleic acid.
[0279] As used herein, the term "barcoding reagent" refers to any substrate comprising a unique molecule or moiety that can be used as a barcode to identify a target molecule. Barcodes are well known in the art, and any molecule or moiety that is sufficiently unique to identify a target molecule can be used as a barcode. In some embodiments, the barcoding reagent is the barcode itself. In some embodiments, the barcode is a protein barcode. In some embodiments, the barcode is a protein tag. In some embodiments, the barcode is a fluorescent protein.
[0280] In some embodiments, the barcode is a nucleic acid barcode. In some embodiments, the nucleic acid molecule is a short nucleic acid molecule. In some embodiments, its length is less than 3, 5, 7, 10, 12, 15, 17, 20 or 25 nucleotides. Each possibility represents a separate embodiment of the present invention. In some embodiments, the nucleic acid molecule is 3 to 10, 3 to 15, 3 to 20, 3 to 25, 3 to 30, 3 to 35, 3 to 40, 4 to 45, 3 to 50, 5 to 10, 5 to 15, 5 to 20, 5 to 25, 5 to 30, 5 to 35, 5 to 40, 5 to 45 or 5 to 50 nucleotides in length. Each possibility represents a separate embodiment of the present invention. In some embodiments, the barcoding reagent is an enzyme for connecting the barcode to the target molecule. In some embodiments, the barcoding reagent is a ligase. In some embodiments, the barcoding reagent is a barcode, and the solid phase support further comprises an enzyme for connecting the barcode to the target molecule.
[0281] In some embodiments, the target molecule is a protein. In some embodiments, the target molecule is a nucleic acid molecule. In some embodiments, the target molecule is in DNA or RNA. In some embodiments, the capture agent captures a protein associated with the target molecule. In some embodiments, the capture agent captures a DNA-associated protein, and the target molecule is DNA. In some embodiments, the DNA is cfDNA. In some embodiments, the target molecule is complexed with a protein captured by the capture agent.
[0282] The solid support can be any polymer or inorganic material to which a biomacromolecule can be attached. Attachment can be direct or indirect. In some embodiments, the solid support is made of a material used to assemble a microfluidic device. In some embodiments, the solid support is a bead. In some embodiments, the bead is agarose beads. In some embodiments, the solid support is a magnetic bead or a paramagnetic bead. In some embodiments, the solid support is agarose beads or a magnetic bead or a paramagnetic bead. In some embodiments, the support is conjugated to a capture agent. In some embodiments, the support is conjugated to a ChIP antibody. In some embodiments, the support is conjugated to a barcoding reagent. In some embodiments, the support is conjugated to a capture agent and a barcoding reagent. Conjugation can be performed by any method known in the art, including but not limited to covalent bonding, charge-based bonding, and hydrophobic interactions. In some embodiments, conjugation is conjugation of biotin to avidin. In some embodiments, conjugation is performed by amine binding technology. In some embodiments, the amine binding technology is epoxy. In some embodiments, conjugation is by carboxyl capture.
[0283] In another aspect, a method for multiplexing assays for more than one target molecule in a single solution is provided, the method comprising:
[0284] a. capturing a first target molecule in solution onto a first solid support of the present invention;
[0285] b. capturing at least a second target molecule in the solution to a second solid support of the present invention;
[0286] c. attaching a first target molecule and a first barcode and at least a second target molecule and a second barcode;
[0287] d. Simultaneously measuring the first and second target molecules, wherein the measurement result of the first target molecule is identified by the first barcode, and the measurement result of the second target molecule is identified by the second barcode;
[0288] This allows for multiplexed assays of more than one target molecule in a single solution.
[0289] As used herein, "multiplexed assay" refers to performing an assay on multiple samples simultaneously. Multiplexing is useful when samples are limited because the assay is expensive in terms of time, money, reagents, or sample input. By multiplexing using the methods of the present invention, assays can be performed simultaneously from beginning to end, thereby reducing variability between samples ( Figure 5B ). For example, when multiplexed chromatin immunoprecipitation is then performed using the method of the present invention for next generation sequencing (ChIP-Seq) assay, protein capture of all antibodies used is performed at once in one tube. The connection of the barcode also occurs all at once, and washing and sequencing are also performed in one piece. This greatly limits any inter-sample differences in assay performance. In some embodiments, the assay is any one of ChIP, ChIP-Seq, cfChIP, cfChIP-Seq, protein quantification, and protein-protein interaction assays. Protein quantification can be achieved by adding a universal DNA adapter / sequence to the protein and then connecting a barcode. In some embodiments, the method of the present invention further comprises attaching a adapter to the target protein. In some embodiments, the assay is chromatin immunoprecipitation followed by sequencing (ChIP-Seq).
[0290] In some embodiments, the target molecule is a protein. In some embodiments, the target molecule is a nucleic acid molecule. In some embodiments, the target molecule is a protein and / or a nucleic acid molecule.
[0291] In some embodiments, identification by barcodes comprises quantifying the amount and / or number of target molecules. In some embodiments, the amount and / or number of barcodes is equal to the amount and / or number of target molecules. In some embodiments, the amount and / or number of barcodes is proportional to the amount and / or number of target molecules. In some embodiments, the amount and / or number of barcodes is equal to or proportional to the amount and / or number of target molecules.
[0292] As used herein, the term "about" when used in conjunction with a value refers to the reference value ± 10%. For example, a length of about 1000 nanometers (nm) refers to a length of 1000 nm ± 100 nm.
[0293] Note that as used herein and in the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to "a polynucleotide" includes a plurality of such polynucleotides, and reference to "the polypeptide" includes reference to one or more polypeptides and equivalents thereof known to those skilled in the art, and so forth. It should also be noted that a claim may be drafted to exclude any optional element. Therefore, this statement, in combination with the recitation of a claim element or use of a "negative" limitation, is intended to serve as antecedent basis for use of exclusive terminology such as "solely," "only," and the like.
[0294] In those instances where phraseology similar to "at least one of A, B, and C, etc." is used, it is generally intended that such construction be understood by those skilled in the art (e.g., "a system having at least one of A, B, and C" includes, but is not limited to, a system having A alone, having B alone, having C alone, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.). Those skilled in the art will further understand that, in practice, any disjunctive words and / or phrases presenting two or more alternative terms, whether in the specification, claims, or drawings, should be understood to contemplate the possibility of including one of the multiple terms, one of the two terms, or both of the two terms. For example, the phrase "A or B" will be understood to include the possibility of "A" or "B" or "A and B."
[0295] It should be understood that certain features of the present invention that are described for clarity in the context of separate embodiments may also be provided in combination in a single embodiment. Conversely, various features of the present invention that are described for brevity in the context of a single embodiment may also be provided separately or in any suitable subcombination. All combinations of embodiments related to the present invention are specifically encompassed by the present invention and are disclosed herein as if each and every combination were individually and explicitly disclosed. In addition, all subcombinations of various embodiments and elements thereof are also specifically encompassed by the present invention and are disclosed herein as if each and every such subcombination were separately and explicitly disclosed herein.
[0296] Other objects, advantages and novel features of the present invention will be apparent to those skilled in the art after examining the following examples, which are not intended to be limiting. Additionally, as described above and as described in the appended claims, each of the various embodiments and aspects of the present invention is experimentally supported in the following examples.
[0297] Various embodiments and aspects of the present invention as delineated hereinabove and as part of the claims section below find experimental support in the following examples.
[0298] Example
[0299] In general, the nomenclature used herein and the laboratory procedures used in the present invention include molecular, biochemical, microbiological and recombinant DNA techniques. Such techniques are explained in detail in the literature. See, for example, "Molecular Cloning: A laboratory Manual" Sambrook et al., (1989); "Current Protocols in Molecular Biology" Volumes I-III Ausubel, RM, ed. (1994); Ausubel et al., "Current Protocols in Molecular Biology", John Wiley and Sons, Baltimore, Maryland (1989); Perbal, "A Practical Guide to Molecular Cloning", John Wiley & Sons, New York (1988); Watson et al., "Recombinant DNA", Scientific American Books, New York; Birren et al. (eds) "Genome Analysis: A Laboratory Manual Series", Vols. 1-4, Cold Spring Harbor Laboratory Press, New York (1998); U.S. Patent Nos. 4,666,828; 4,683,202; 4,801,531; 5,192,659 and 5,272,057; "Cell Biology: A Laboratory Handbook", Volumes I-III Cellis, JE, ed. (1994); "Culture of Animal Cells - A Manual of Basic Technique" by Freshney, Wiley-Liss, NY (1994), Third Edition; "Current Protocols in Immunology" Volumes I-III Coligan JE, ed. (1994); Stites et al.(eds), "Basic and Clinical Immunology" (8th Edition), Appleton & Lange, Norwalk, CT (1994); Mishell and Shiigi (eds), "Strategies for Protein Purification and Characterization-A Laboratory Course Manual" CSHL Press (1996); all of which are incorporated by reference. Other general references are provided throughout this article.
[0300] Materials and methods
[0301] patient
[0302] All clinical studies were approved by the relevant local ethics committees. This study was approved by the ethics committee of the Hebrew University-Hadassah Medical Center of Jerusalem. Informed consent was obtained from all subjects or their legal guardians before blood sampling.
[0303] Sample collection
[0304] Collect blood samples in The cells were immediately transferred to a K3 EDTA tube and placed on ice. 1× protease inhibitor cocktail (Roche) and 10 mM EDTA were added. The blood was centrifuged (10 min, 1500×g, 4°C), the supernatant was transferred to a new 14 ml tube, and centrifuged again (10 min, 3000×g, 4°C). The supernatant was used as plasma for ChIP experiments. Plasma was used fresh or flash-frozen and stored at -80°C for long-term storage.
[0305] Bead preparation
[0306] 50 μg of antibody was conjugated to 5 mg of epoxy M270 Dynabeads (Invitrogen) according to the manufacturer's instructions. The antibody-bead complex was stored at 4°C in PBS, 0.02% azide solution.
[0307] Immunoprecipitation, NGS library preparation, and sequencing
[0308] 0.2 mg of conjugated beads (~2 μg of antibody) were used per cfChIP sample. The antibody-bead complex was added directly to plasma (1-2 ml of plasma) and allowed to bind to cf-nucleosomes by rotating overnight at 4°C. The beads were magnetized and washed six times with blood wash buffer (BWB 50 mM Tris-HCl, 150 mM NaCl, 1% Triton X-100, 0.1% sodium deoxycholate, 2 mM EDTA, 1× protease inhibitor cocktail), twice with BWB-500 (same as BWB with only 500 mM NaCl), and three times with 10 mM Tris pH 7.4. All washes were performed on ice with 150 μl of buffer by moving the beads side-to-side on a magnet. During washes in detergent-free buffer, the supernatant was not removed by vacuum. After removing the beads, the plasma was stored as it was suitable for further rounds of cfChIP.
[0309] On-bead chromatin barcoding and library amplification were performed to overcome the problem of low input material. This procedure supports the preparation of cfChIP from as few as 1000 cells. The following steps were all performed on the beads to reduce cfDNA release and cfDNA loss that may occur during tube transfer. DNA ends were repaired by T4 DNA polymerase and T4 polynucleotide kinase. After washing, adenine bases were added to the repaired ends of the DNA using Klenow exo minus. After another wash, DNA adapters were ligated on; in this case, Illumina adapters with DNA barcode sequences were used. For the DNA elution and cleanup steps, the beads were incubated at 55°C for 1 hour in 50 μl of chromatin elution buffer (10 mM Tris pH 8.0, 5 mM EDTA, 300 mM NaCl, 0.6% SDS) supplemented with 50 units of proteinase K (Epicentre), and the DNA was purified by 1.2X SPRI cleanup (Ampure xp, agencourt). The purified DNA was eluted in 25 μl EB (10 mM tris pH 8.0) and 23 μl of the eluted DNA was used for PCR amplification (16 cycles) with Kapa hot start polymerase. The amplified DNA was purified by 1.2X SPRI cleanup and eluted in 12 μl EB. The eluted DNA concentration was measured by Qubit and the fragment size was analyzed by tapestation visualization. Note: If the adapter dimer is still visible by tapestation after library amplification, the samples with different barcodes can be combined and run on a 4% agarose gel ( The fragments were separated on a 4% EX agarose gel (Invitrogen), and fragments larger than the adapter dimer (>200 bp) were gel purified. Alternatively, gel purification can be avoided by performing an additional X 0.8 SPRI cleanup after sample pooling, which removes most adapter dimers. Paired-end sequencing of the DNA library was performed on an Illumina NextSeq 500.
[0310] Sequence analysis
[0311] Reads were aligned to the human genome (hg19) using bowtie2 with the “no-mixed” and “” flags. We discarded reads with low alignment scores and fragment duplications.
[0312] Roadmap Epigenome
[0313] We downloaded consolidated alignment data from the Roadmap Epigenome Consortium database (egg2.wustl.edu / roadmap / data / byFileType / alignments / consolidated / ). To these, we added kidney samples (egg2.wustl.edu / roadmap / data / byFileType / alignments / unconsolidated / H3K4me3 / BI.Adult_Kidney.H3K4me3.27.filt.tagAlign.gz and egg2.wustl.edu / roadmap / data / byFileType / alignments / unconsolidated / H3K4me3 / BI.Adult_Kidney.H3K4me3.153.filt.tagAlign.gz). For our analysis, we discarded prenatal, ESC, and cell line samples, resulting in 71 tissues and cell types.
[0314] Tumor-type genetic characteristics
[0315] We downloaded RNA-seq data from the TCGA and GTEx projects analyzed by the Xena project (Toil enables reproducible, open-source analysis of large biomedical data, Vivian J, et al., Nat. Biotechnol., 2017, and the Toil RNAseq Recompute database, tcga-data.nci.nih.gov). We defined gene sets overexpressed in a tumor type as meeting three requirements: 1) significantly higher expression in tumor samples compared with corresponding tissue samples (t-test, FDR-corrected q < 0.001); 2) significantly higher expression compared with all healthy samples (t-test, FDR-corrected q < 0.001); and 3) median expression in tumors was higher than the median expression in each healthy sample.
[0316] TSS Location Directory
[0317] We downloaded the Roadmap Epigenome Consortium ChromHMM annotations for all integrated tissues (egg2.wustl.edu / roadmap / data / byFileType / chromhmmSegmentations / ChmmModels / coreMarks / jointModel / final / all.mnemonics.bedFiles.tgz). Using these annotations, we constructed a catalog of potential TSS sites. We expanded this catalog to include 3 kb regions centered around the TSSs of annotated transcripts from the UCSC Gene Database and the ENSEMBF Transcriptome Database (UCSC known genes: Bioconductor AnnotationHub AH5036; ENSEMBL transcripts: Bioconductor AnnotationHub AH5046; genome annotations: Bioconductor AnnotationHub AH5040). We used the combined catalog to define regions along the genome that were TSSs or “background” (most likely not TSSs). The latter regions were tiled using 5 kb windows.
[0318] We quantified the number of reads covering each region in each sample in the catalog and in the atlas samples. We estimated the local fitness model of nonspecific reads along the genome for each sample and extracted counts of specific ChIP signals in the catalog representing each sample (see below). These samples were then normalized (Supplementary Text) and scaled to 1M reads in the reference healthy sample.
[0319] Organizational / process characteristics
[0320] To define tissue-specific signatures of particular modifications, we examined a binned representation of the profiles. For each tissue, we defined signatures that had a unique window with signal in one of the samples of the tissue of interest but not in all other samples (see below).
[0321] To define process signatures, we converted gene-specific annotations (eg, GO) into genomic windows by including all windows overlapping gene promoters in the annotation.
[0322] Statistical analysis
[0323] We consider two different statistical tests. For both tests, we need to estimate the background coverage, i.e., the reads coming from nonspecific pull-down (see below).
[0324] The first test is whether the feature exists. Formally, we examine whether we can reject the null hypothesis that the number of reads in a feature window will be Poisson-distributed according to the background rate (see below). We calculate the p-value of the actual number of reads observed in the feature window as the probability of having that number or higher according to the null hypothesis. The rejection of the null hypothesis for a specific feature indicates that some windows in the feature carry the inquestion modification in the cell subpopulation that constitutes the cf-nucleosome pool.
[0325] The second test is whether the feature is overexpressed compared to what is expected in healthy baseline subjects. In order to limit the latter expectation (value), we use the average signal from 5 healthy samples to limit the average number of reads (per million) in each window. We then estimate two sample-specific parameters-the first is the background rate (as discussed above), and the second is the scaling factor (conversion factor, scaling factor) (see below) that rescales the average expected value to the sequencing depth of the specific sample. These jointly define the expected coverage in each window under the null hypothesis (i.e., the object is from a healthy population). We calculate the p-value of the actual number of reads observed in the feature window, which is the probability of having this number or a higher number according to the null hypothesis. The negation of the null hypothesis for the specific feature shows that some windows in the feature have higher signals than the signals in our expected healthy subjects. The explanation is that these are active abnormal processes in the cells that constitute the object cf-nucleosome library.
[0326] TSS Location Directory
[0327] We constructed the TSS catalog using the following steps. All steps were performed based on the human genome version "hg19":
[0328] 1. We downloaded ChromHMM calls for 111 tissues and cell types across the entire human genome from the Roadmap Epigenomics website (egg2.wustl.edu / roadmap / data / byDataType / rna / expression / 57epigenomes.RPKM.pc.gz). We also downloaded UCSC browser known gene annotations and ENSEMBL transcriptome annotations (UCSC known genes: Bioconductor AnnotationHub AH5036; ENSEMBL transcriptomes: Bioconductor AnnotationHub AH5046; genome annotations: Bioconductor AnnotationHub AH5040).
[0329] 2. We filtered all genomic ranges labeled with the states "1_TssA" or "2_TssAFlnk" and merged adjacent ranges labeled with either state in completely homogenous tissues. We call these "ChromHMM TSS windows." We found 476,931 such windows. Each ChromHMM TSS window was assigned a gene name or names using the following steps.
[0330] a. If it is within 2.5 Kb of one or more TSSs in the UCSC known gene annotations, it is assigned the name of those genes.
[0331] b. If not, we searched for an Ensembl transcript start point within 2.5 kb. Again, if found, the TSS window received the gene name associated with the transcript.
[0332] c. All other TSS windows remain unnamed.
[0333] 3. To include transcripts not represented in the TSS catalog, we examined all genes in the UCSC known gene database and all transcripts in the ENSEMBL database. For each, we defined a 3Kb TSS window centered on the TSS. We discarded any such windows that overlapped with the TSS windows from step 2. This step added a total of 14,857 TSS windows from the UCSC known genes and 41,376 TSS windows from the ENSEMBL transcriptome.
[0334] 4. We created windows that sliced the remaining genomic regions between the TSS windows. For each TSS window that did not have an adjacent TSS window, we created a "flanking" region of 1 Kb (or less). This resulted in 370,332 flanking windows (because some TSS windows are adjacent to each other according to ChromHMM calls in different tissues). The remaining uncovered regions were sliced with a "background" region of 5 Kb (or less). There were a total of 502,263 such windows.
[0335] The resulting directory is saved as a BED file (TSS.bed).
[0336] Processing of sequencing files
[0337] Base calling was performed using bcl2fastq (2.18). Paired-end reads were mapped to the human genome (hg19) using bowtie2 with the "no-mixed" and "no-discordant" features, discarding reads with a quality of 0. BEDPE files (start and end points of each fragment) were obtained using BEDtools "bamtobed" with the "bedpe" feature, discarding duplicate fragments. BEDPE files were converted to coverage counts for windows in a directory using the BEDtools "intersect" command and using the BioConductor "GenomicRanges" countOverlaps() function. Both methods count the number of sequence fragments that overlap with the window for each window.
[0338] Estimating background signal
[0339] Each ChIP procedure has a nonspecific background signal. In the case of cfChIP, the background is due to some form of nonspecific binding of DNA and chromatin fragments to the bead-antibody complex. Our experience has shown that background levels vary between samples and bead-antibody binding batches. In addition, sequencing depth varies between samples, and the number of background reads increases in deeply sequenced samples. Therefore, it is important to estimate the background signal level so that it can be compared with the actual signal.
[0340] Initially, we employed a simple procedure to remove background from the H3K4me3 signal. We assumed that nearly all specific H3K4me3 signals were located at the TSS and 5' regions of genes. Therefore, reads at other locations represented background. To account for unannotated TSSs in our TSS catalog, we considered that a small portion of the background window might contain true signal, and therefore removed those with the highest values.
[0341] More specifically, we performed the following: We created a vector containing all “background” windows of size ≥ 4 Kb (421,465 out of 549,385) and applied the following procedure:
[0342]
[0343] This procedure is relatively robust to the selection of quantiles for the outlier removal window.
[0344] However, in some samples, the Poisson distribution was not a good fit for the background values. Further investigation showed that this inconsistency was largely due to local background effects. One local effect was the sex chromosomes appearing at a level of 50% in males but at 100% (X) and 0% (Y) in females. These were not the only local effects - some regions showed higher levels of background. This could be due to segmental duplications (regions close to centromeres and telomeres) or accessibility issues. In addition, in cancer samples, there were significant patient-specific biases.
[0345] To overcome these problems, we devised a local background rate estimate. We used the above estimation procedure, but at a continuous level of resolution.
[0346] 1. Genome-wide background levels.
[0347] 2. Chromosome-specific background.
[0348] 3. Covering 10Mb tiles of each chromosome with 2.5Mb offsets.
[0349] 4. Covering 5Mb fragments of each chromosome with an offset of 1.25Mb.
[0350] The estimate for each level uses the estimate of the previous level as a prior (using pseudo counts in 1000 windows for levels 2 and 3 and 500 windows for level 4).
[0351] The results are estimates of background coverage for 5Mb overlapping tiles. To obtain a single estimate, for each position, we take the maximum of the estimates from the tiles that cover it (typically 4 tiles). We choose the maximum because we believe that overestimating the background may reduce the estimated signal, but will reduce the number of background artifacts. Figure 6A Background estimates for a healthy male sample are shown. Figure 6BA sample from a healthy woman is shown, where the background in chrX is lower than in the autosomes (slightly more than half), and the background in chrY is slightly lower. Many positions in chrY are orthologous to positions in chrX, leading to biased estimates. Other biases occur near centromeres, where we see higher background levels in some chromosomes (e.g., chrl, chr9). When examining cancer patients, the variability in background estimates is much greater, presumably reflecting chromosomal abnormalities in tumors ( Figure 6C ).
[0352] Gene-level signal and normalization
[0353] For each gene, we assigned a set of TSS windows annotated with the gene name. For each sample, we calculated the actual total coverage of the windows assigned to the gene and the expected average of the background reads of these windows (using the local ratio and window size that may be different at each window).
[0354] More precisely,
[0355]
[0356] Where W g is the set of windows assigned to gene g, C[w,s] is the coverage of window w in sample s, and λ^[w,s] is the estimated background rate of window w in sample s.
[0357] The null hypothesis is that coverage C g With parameter G b According to the Poisson distribution, we consider values much higher than expected to be signals. We define the original signal of gene g as:
[0358] S[g,s]=C[g,s]-B[g,s], if C[g,s]≥B[g,s]+2√B[g,s], otherwise it is 0.
[0359] Therefore, we considered C[g,s] to be a true signal if it was greater than two standard deviations from the mean of the background level for that gene.
[0360] Applying this procedure to each sample will generate a count matrix for each gene in each sample. We also included in this matrix samples from the Roadmap Epigenomics data for H3K4me3 ChIP, which were processed in the same manner.
[0361] To normalize for the effects of varying coverage, we assumed that the signal at the promoters of "housekeeping" autosomal genes should be similar across samples. We defined these genes as those with highly significant signals in a set of reference healthy samples. The exact choice of significance level did not alter normalization.
[0362] We applied quantile normalization (Bioconductornormalize.quantiles) to the matrix of raw signal samples x housekeeping genes. This resulted in normalized values for the housekeeping genes in each sample. However, it did not assign values to all other genes. Therefore, we estimated a multiplicative normalization factor for each sample to best match the quantile-normalized values to the raw values. For most samples, the relationship between the two was linear.
[0363] The scaling factors were rescaled so that the total normalized signal for the reference healthy sample set (below) was an average of 1 million.
[0364] Using these normalization factors, levels v[s], we calculated the normalized gene levels for each sample:
[0365] N[g,s]=v[s]*S[g,s].
[0366] Using the same normalization procedure, we also normalized the coverage of each window in each sample:
[0367] N[w,s]=v[s]*max(C[w,s]-B[w,s],0).
[0368] Defining tissue-specific features
[0369] Using the Roadmap Epigenomics metadata table, we defined Roadmap sample sets as belonging to a tissue or group of tissues. These definitions include some redundancy. For example, the lymphocyte group includes B cell, T cell, and NK cell samples and therefore is classified into each of these groups.
[0370] We then define a specific window group for each group, when the window w passes the following criteria:
[0371] 1. Window w is on the autosome;
[0372] 2. In at least one of the profile samples in the set, N[w,s]≥35;
[0373] 3. In all atlas samples outside this group, N[w,s]<15;
[0374] 4. In all windows w' within 1Kb of w, N[w,s]<15.
[0375] The last condition was added because we noticed that often when a gene is expressed, there is "spillover" into adjacent windows.
[0376] Groups that found fewer than 4 unique windows were considered uncharacterized. For all other groups, we defined a unique feature as the group-specific window (see Table 1). The workflow for cfChIP processing and analysis is provided in Figure 7 middle.
[0377] Statistical tests
[0378] We use two main tests here:
[0379] Detection test. To test whether a gene or feature is present above background in a sample, we used the Poisson distribution. More specifically:
[0380]
[0381] Here, W is a set of windows, which may be windows associated with the above-mentioned gene or tissue-specific features.
[0382] Overrepresentation Test. To test whether the observed signal for a group of genes is higher than the expected signal in healthy samples, we bounded the expected normalized signal H[g] of the gene to the mean value N[g,s] of the reference samples using a reference from healthy subjects.
[0383] We then use the following procedure:
[0384]
[0385] The main difference from the previous tests is that we include the contribution of the healthy sample after we convert from standardized units to sample-specific units. The second difference is that we work at the gene level.
[0386] result
[0387] Example 1: Chromatin immunoprecipitation of cf-nucleosomes from plasma
[0388] The majority of plasma cfDNA is likely in the form of nucleosomal DNA (cf-nucleosomes) with intact histone modifications. We investigated whether extracting and sequencing DNA from cf-nucleosomes with specific histone marks could be used to determine information about the cells of origin of the cfDNA. Figure 1AThis is an attractive approach for several reasons. First, by designing the ChIP experiment to sequence only the positive signals, the number of reads required for the positive signals is reduced, thereby reducing the cost and work associated with the assay. Second, the positive targets are relatively rare; promoter markers, such as H3K4me3, occur at ~50,000 locations in the genome (less than 1% of the genome). Enhancer markers (e.g., H3K4me1) can occur in many regions (~10% of the genome), but are restricted in each cell. Third, histone marks are largely tissue-specific. In particular, the majority of enhancers are tissue-specific, thus providing strong tissue-specificity for H3K4me1 / 3 ( Figure 1B Fourth, histone modifications reflect transcriptional activity and respond to changes in cell state, thus creating an opportunity to detect changes in cell activity as cells die.
[0389] We designed a simple protocol for cf-nucleosome ChIP-seq (cfChIP) from as little as 1-2 ml of plasma ( Figure 1A Inset, 1C). cfChIP and paired-end sequencing of plasma samples from 11 healthy individuals generated 30-170 million and 90-25 million unique reads per sample for H3K4me3 and H3K4me1, respectively, indicating that ∼1-2% of nucleosomes in plasma bearing the corresponding marker (e.g., H3K4me3) were captured, adapter-ligated, and sequenced (see Materials and Methods). Importantly, cfChIP signals around ubiquitously expressed genes showed high correlation with reference ChIP-seq from tissues (NIH Epigenome Roadmap Consortium) ( Figure 1C Globally, meta-analysis of cfChIP signals for H3K4me1 and H3K4me3 yielded the expected general distribution of these marks around enhancers and promoters ( Figure 1D ).
[0390] A potential concern is contamination by chromatin released by lysis of leukocytes during blood draw. Several lines of evidence suggest that this is highly unlikely. (a) The fragment size distribution of the cfChIP library shows two peaks at ∼170 and ∼320 bp corresponding to DNA wrapped around mononucleosomes and dinucleosomes ( Figure 1E ), consistent with apoptotic and, in some cases, necrotic cell death, but not with cytolysis, resulting in fragments ranging from 10 kb or larger. (b) We identified thousands of enhancers with H3K4me1 and dozens of promoters with H3K4me3 that are absent from leukocytes (PBMCs; peripheral blood mononuclear cells), which constitute the largest fraction of nucleated blood cells. Figures 1F-1G) in ChIP-seq. Analysis of promoters marked by H3K4me3 in cfChIP but not in leukocytes identified strong signals from megakaryocytes residing in the bone marrow. (c) We were able to detect disease-associated chromatin in distant tissues from patients (see below).
[0391] Non-histone DNA-associated proteins can also be used for cf-ChIP. We spiked human plasma with 90 ng of chromatin prepared by natural MNase treatment of DNA from mouse embryonic stem cells and performed cfChIP with anti-CTCF antibodies. Sequencing reads from cf-ChIP were aligned to the mouse genome, and sharp peaks were observed that clearly overlapped with peaks obtained from ChIP-Seq of CTCF in mouse cells ( Figure 1H Meta-analysis of the data showed clear signals at CTCF sites throughout the genome, and similar analysis of cfChIP using anti-H3K4me3 antibodies showed depletion of histone marks at the same sites ( Figure 1I CTCF binding and H3K4 trimethylation are generally mutually exclusive, so this result helps confirm that the CTCF signal is real.
[0392] Together, these results strongly suggest that cf-nucleosomes retain well-established endogenous patterns of active histone marks and transcription factor binding.We focused our analysis on H3K4me3 at this point because assigning H3K4me3 peaks to specific genes is relatively straightforward.
[0393] To assess the reproducibility of cfChIP, we performed technical and biological replicates from multiple subjects. Replicates from the same individual and replicates between healthy individuals showed correlations of 0.94-0.97 and 0.92-0.94, respectively ( Figure 2A -C). Peaks with low reproducibility between subjects are enriched for X- or Y-chromosome-specific genes and are indeed not evident when comparing two individuals of the same sex ( Figure 2B ).
[0394] To test the detection limit of cfChIP, we exploited sequences unique to the Y chromosome and titrated male-derived plasma into female-derived plasma. We evaluated sensitivity at specific genomic locations and genomic signatures—collections of genomic locations that can define differential expression in certain cell types or transcriptional programs. The H3K4me3 cfChIP signal at the male-specific peak on the Y chromosome ( Figure 2D ) showed that we could reliably identify a single male-specific peak even when male plasma accounted for less than 10% of the total plasma ( Figure 2D Furthermore, the contribution of the male-specific peak increases linearly with the male plasma pool fraction ( Figure 2E -F), demonstrating that cfChIP sensitivity is linearly related to this score and the size of the feature position as well as the sequencing depth. In fact, combining the signals from the 25 male-specific peaks can increase the detection sensitivity to 1% ( Figure 2E 、 2G This is likely an underestimate because the diploid genome has one Y chromosome. Extrapolating from our male spike-in experiments, we estimate that a moderate feature size of 45 to 200 peaks can detect cfDNA with high probability (0.95 or higher) from 0.1% of cells constituting the cfDNA pool at low sequencing depth ( Figure 2H This size signature can be used to identify specific cell types or transcriptional programs ( Figure 2I ).
[0395] H3K4me3 levels at promoters in tissue samples correlate with transcript levels and strongly predict gene expression levels. We found that cfChIP H3K4me3 correlates with leukocyte RNA-seq at constitutive genes (based on RoadmapEpigenomics Consortium and GTEx Consortium 2015), and this correlation is similar to that of leukocyte ChIP-seq ( Figure 2J Similarly, we found a high correlation with genes expressed in leukocytes consistent with their major contribution to the cfDNA pool in healthy individuals ( Figure 2K These results strongly suggest that cfChIP of transcription-associated histone modifications can provide insights into gene expression patterns in cfDNA-derived cells.
[0396] With these findings in mind, we set out to examine our ability to detect tissue-specific signatures in samples from healthy subjects. Previous studies of cfDNA CpG methylation have estimated that ~55% of cfDNA is derived from leukocytes and ~1% from the liver, with minimal or no contribution from heart and brain cfDNA. Using the RoadmapEpigenomics dataset of H3K4me3 ChIP-seq for multiple healthy tissues as a reference, we defined tissue-specific signatures in an unbiased manner (see Materials and Methods, Table 1). We then evaluated the normalized number of reads for each signature in each subject and the statistical significance of these counts ( Figure 2L-M; Materials and Methods). As expected, the presence of leukocytes can be detected when using several specific peaks or even a single specific peak. Using larger features, the presence of liver cfDNA was also clearly and significantly detected, while brain and heart features were absent in blood, as expected. These signals are specific and have high statistical confidence (q<10-20, see Materials and Methods), and they demonstrate the ability of cfChIP to detect cf-nucleosomes from rare cell populations.
[0397] Example 2: cfChIP detection of cell death associated with pathology.
[0398] The ability of cfChIP to recognize signatures of cells from distant tissues suggests the exciting possibility that this tool could detect cf-nucleosomes derived from disease-associated pathological cell death. To test this hypothesis, we collected samples from patients diagnosed with acute myocardial infarction (AMI), a process that leads to widespread cardiomyocyte death. We collected samples from patients admitted to the emergency room, obtaining them before, immediately after, and ~12 hours after urgent percutaneous coronary intervention (PCI) to restore blood flow. We expected to observe cfChIP signals of cardiomyocyte death only in AMI patients and not in healthy subjects, particularly in samples following PCI.
[0399] As predicted, cardiac-specific H3K4me3 peaks were strongly and significantly detected in post-PCI patient samples, but not in samples from healthy individuals or pre-PCI patient samples ( Figure 3A Cardiac signals included clear peaks in the promoters of cardiac-specific genes. For example, TNNT2 and TNNI3, encoding cardiac-specific troponins T2 and I3, were clearly active and were only observed in these samples. These two genes encode typical protein markers of myocardial injury ( Figure 3B Indeed, we found good correlation between the strength of the cfChIP cardiac signature, troponin levels measured in blood, and cardiac cfDNA estimates based on cardiac-specific differentially methylated CpGs ( Figure 3C ).
[0400] To give a fair view of the ability of the present invention to characterize the tissue of origin, we evaluated a panel of cell type-specific features across cfChIP samples ( Figure 3D -E). This analysis showed that in all samples, we could detect signatures of a range of cell types from the blood (e.g., monocytes and neutrophils) and organs (e.g., liver). One of the subjects (H008), a four-month pregnant woman with a male fetus, had a clear placental signature. In fact, we were able to detect a low but significant Y chromosome signal in her plasma ( Figure 3D ).
[0401] In AMI patient samples, the situation is more complex. As mentioned above, AMI patients sampled several hours after PCI showed clear cardiomyocyte characteristics ( Figure 3A 、 3D However, in addition, we observed a significant increase in the hepatocellular signature in AMI patients before and shortly after PCI. This signature included clear signals for liver-specific genes such as albumin and complement genes ( Figure 3F ). This unexpected observation may be due to the well-known phenomenon of liver damage in AMI patients secondary to low organ perfusion and hepatic hypoxia. One of these AMI patients (M002) also presented with elevated levels of active chromatin from erythroblast-related genes, including the hemoglobin locus (see below) and the erythropoietin locus (EPO), most likely from hepatocytes - due to the hepatic response to hypoxia. Together these data suggest a systemic response to oxygen deprivation. The liver and erythroblast signals may be due to temporary injury caused by the reduced systemic perfusion associated with AMI. In fact, follow-up cfChIP-seq in this patient 11 months later was normal. In the second AMI patient (M001), we observed a gradual decline in liver characteristics within hours after PCI, indicating a rapid resolution of hepatic oxygen deprivation ( Figure 3G To confirm our cfChIP observations, we analyzed the cfDNA methylation status of liver-specific genes whose DNA methylation status is indicative of hepatocyte cell death. Indeed, we observed good concordance between liver cfChIP signature levels and liver cfDNA estimates (R2 = 0.97, Figure 3H ).
[0402] An important potential application of cfChIP is the identification of the tissue of origin of cancer. Advanced cancer is often accompanied by high levels of cfDNA in plasma, most of which is derived from tumor cells (ctDNA). We collected plasma samples from patients with gastrointestinal (GI) tumors and analyzed their tissue of origin of cf-nucleosomes ( Figure 3I 、 3E). Overall, plasma from cancer patients contains tissue signals that are not observed in healthy subjects. Most notably, we observed signals originating from gastrointestinal (GI) tissue and GI smooth muscle, which is consistent with the primary site of the tumor. Weaker but important (significant) GI features are evident even when the primary tumor is removed by surgery and only residual metastatic disease is evident (patients C004 and C005). We also observed low but important (significant, significant) signals from other tissues, such as the brain signals observed in C001. These signals may be due to treatment (C001 received brain radiotherapy) or due to collateral damage to normal (non-malignant) tissue.
[0403] We also studied patients with localized hepatocellular carcinoma (HCC) undergoing partial hepatectomy (PHx). We collected blood samples before, during, and at different time points after surgery and analyzed circulating cf-nucleosomes using cfChIP and measured a classic marker of liver injury, the enzyme ALT ( Figure 3J Surprisingly, despite the fact that the patient had active cirrhosis in addition to HCC, preoperative ALT levels were also normal. ALT levels rose on the first day after PHx and gradually declined over the next few days. cfChIP analysis of the liver signature was very consistent with the ALT assay, again suggesting that cfChIP detects dynamic processes in distant tissues. One difference was that the cfChIP liver signature returned to normal levels approximately 2 days earlier than ALT. This difference may be due to the shorter half-life of cfDNA in the circulation (<2 hours) compared to ALT (~47 hours).
[0404] Together, these results show distinct differences in cfChIP signatures between healthy subjects and patients undergoing pathological processes, with these differences corresponding to tissues where these processes occur, such as the heart, liver, and gastrointestinal tract.
[0405] Example 3: Plasma chromatin reflects gene activity patterns
[0406] A major challenge in analyzing cfDNA is inferring gene expression in the tissue of origin. To date, the primary approach proposed to address this problem relies on underrepresentation of specific promoter elements in cfDNA as an indicator of gene expression. However, this approach requires extremely deep sequencing and is limited to situations where cfDNA from the target tissue constitutes the majority of the blood population. We tested the extent to which cfChIP can report on non-constitutive gene expression programs occurring in the cells of origin.
[0407] H3K4me3 is closely associated with transcriptional activity and changes dynamically in response to changes in the transcriptional program, raising the exciting possibility that cfChIP may be able to detect more dynamic transcriptional programs beyond the tissue of origin. To test this hypothesis, we compared H3K4me3 cfChIP signals from patients with AMI or cancer with a set of signature gene expression profiles representing different cellular processes and responses ( Figure 4A -C).
[0408] This analysis found that multiple features had higher than expected signals - that is, the amount of cf-nucleosomes captured for the feature was significantly higher than what we observed in healthy subjects. For example, we observed a strong signature of heme metabolism in patients M002 and C005 who experienced hypoxia and bacteremia, respectively. C005's blood cell count did show a high red blood cell distribution width (RDW) as well as a low red blood cell count (RBC) and hemoglobin (HGB), indicating high red blood cell production due to anemia. This signal may be due to increased cell death of erythroid progenitor cells or closely related cells, or due to nuclear loss as erythroblasts undergo maturation to become erythrocytes. Therefore, this feature is indicative of a specific hematopoietic cell differentiation process.
[0409] Other signatures, such as glycolysis or interferon-α response, reflect processes that can occur in a variety of cell types. Our observation of a higher glycolytic signature in cancer patients is consistent with a metabolic reprogramming known as the Warburg effect, which is considered a hallmark of advanced cancer ( Figure 4D ). We also observed a significant increase in the glycolytic signature in M002, who suffered extensive liver damage. Interestingly, in M002, we also observed increased signals for several liver-specific glycolytic genes (such as ALDOB and PFKFB 1), while in cancer patients, the enhanced glycolytic signature did not include signals from these genes. These results suggest that cfChIP can detect cell-specific transcriptional programs associated with underlying pathophysiological states. As expected, in the plasma of cancer patients, we also observed increases in several proliferation-related signatures (Kras, Myc targets, E2F targets, G2M checkpoints) and the mTORC1 pathway, which coordinates metabolism and cell growth. Interestingly, some of these signatures were also observed in AMI patients who experienced liver damage and may reflect liver recovery caused by ischemic damage after PCI.
[0410] Another example of a detectable transcriptional program is the interferon-α response, which is often induced by the presence of pathogens such as viruses and bacteria. We observed a dramatic increase in the interferon signature in M004 and C005. In the latter, this was likely due to the severe bacteremia that led to their hospitalization. M004, whose samples showed a high interferon and inflammatory signature, appeared to have suffered more severe cardiac damage compared to other AMI patients in terms of troponin levels and cfChIP cardiac markers ( Figure 3A 、 3C This may be due to the induction of an IRF3 / interferon I response in M004, which has recently been shown to promote severe AMI responses.
[0411] Together, these observations show that cf-nucleosomes not only report the death of specific cell types but can also reflect detailed changes in gene expression programs in a wide range of cell types.
[0412] Plasma chromatin allows dissection of patient-specific molecular phenotypes
[0413] A hallmark of cancer cells is genetic alterations that lead to dysregulated gene expression programs. Identification of this cancer-specific transcriptional program can aid in diagnosis and treatment selection. For each sample, we tested genes whose signals were elevated compared to five "reference" healthy samples. As a control, unrelated healthy samples outside the reference group were highly correlated with the healthy references, with very few genes (usually fewer than 50) showing significantly elevated signals ( Figure 4E In contrast, samples from patients revealed hundreds to thousands of genes with significantly elevated signals ( Figure 4E ). Examining these genes for enrichment in the annotated gene list summarizes some of the results discussed above. For example, genes in C001 were enriched for the gene sets targeting the GI tract and brain, consistent with the pathology of this patient.
[0414] Next, we looked for cancer-specific signatures in H3K4me3 cfChIP signals. We analyzed expression profiles from The Cancer Genome Atlas and the GTEx project to identify gene sets that were significantly higher in tumors compared to normal tissues for each tumor type (see Materials and Methods, Table 2). We then tested for significant overlap between gene sets with higher H3K4me3 signals in the sample and gene sets that were overexpressed in a tumor type (see Materials and Methods). For example, C002 had significant overlap (q < 10-60) with GI tract adenocarcinoma genes ( Figure 4E All samples were analyzed for all tumor types ( Figure 4F-G) shows that only samples from cancer patients, while healthy subjects and MI patients, have a significant enrichment of tumor-related gene expression. Importantly, this enrichment is specific to GI tract cancer, consistent with the diagnosed pathology.
[0415] Focusing on specific genes known to be upregulated in gastric and colorectal cancers, we observed a significant increase in H3K4me3 cfChIP signals in these patients compared to healthy controls ( Figure 4H ). Among these genes, we found the cancer markers CEACAM5 and CEACAM6. The protein products of these genes are used in antibody-based assays for clinical cancer diagnosis. A second colorectal cancer marker, the long noncoding RNA CCAT1 (colorectal cancer-associated transcript 1), showed a strong signal in one of the cancer patients but not in healthy subjects. Another example is the long noncoding RNA EGFR-AS1, which mediates cancer addiction to EGFR and, when highly expressed, renders tumors insensitive to EGFR inhibition. Although cfChIP signals for EGFR were detected in all cancers, EGFR-AS1 was only detected in C002 and not in the other patients. This finding, which would not be detected by cfDNA mutation analysis, raises the exciting possibility that cfChIP can provide treatment selection information beyond genomic mutations.
[0416] Table 2: Cancer characteristics
[0417]
[0418]
[0419]
[0420]
[0421]
[0422]
[0423]
[0424]
[0425]
[0426]
[0427]
[0428]
[0429] Example 4: Analysis of enhancer and gene body markers
[0430] Our analysis of the active promoter mark H3K4me3 provides rich information about the transcriptional program of the tissue of origin. We can obtain information from chromatin marks associated with enhancers and gene activity. Mono- and di-methylation of H3 lysine 4 (H3K4me1 and H3K4me2, respectively) are found in two types of genomic regions: 1) promoter-flanking regions at the boundaries of regions marked by H3K4me3, or 2) poised / active enhancers, where H3K4me3 is barely detectable. ChIP-seq of these marks in tissues showed H3K4me2 peaks near enhancers, while H3K4me1 peaks were broad (approximately ~10 kb) around enhancers. cfChIP of these markers recapitulated the expected distribution: around active promoters, H3K4me2 and H3K4me1 were flanked by the main H3K4me3 peak ( Figure 5A -B) and correlated with H3K4me3 and RNA levels ( Figure 5D Furthermore, in gene-deficient regions, such as the IFNB1 locus, we clearly see markers at enhancers that match the experimentally validated IFNB1 enhancer (Banerjee et al. 2014) ( Figure 5C ).
[0431] Since H3K4me2 is more condensed and the signal is highly similar between healthy subjects ( Figure 5E ), we chose to focus on this marker for our enhancer analysis. To identify enhancers, we looked for H3K4me2 peaks with low or no H3K4me3 signal. We identified ~8,000 putative enhancer peaks in healthy subjects. Of these peaks, >90% were located at sites predicted to be enhancer regions based on five chromatin markers (ChromHMM) across multiple tissues. Furthermore, there was good agreement between these putative enhancer peaks and ChromHMM annotations in relevant cell types, such as monocytes and neutrophils ( Figure 5C Applying the same analysis to two samples from colorectal cancer patients (C002.1 and C002.2) showed strong concordance between these samples, with substantial differences from healthy subjects ( Figure 5E The differences between healthy subjects and cancer samples illustrate that additional information can be obtained from enhancers relative to promoter characteristics. Several examples of cancer-specific enhancer signals include TCF3, CDX1, and CEACAM5 ( Figure 5F-H). The transcription factor TCF3 had a promoter H3K4me3 signal in all test subjects. However, the enhancer activity markers near this gene were quite different, with cancer samples showing a clear H3K4me2 peak in the region corresponding to the putative colon enhancer ( Figure 5F These results suggest that the gene is activated by distinct enhancers that contribute to this signal in the cell and are consistent with the patient's clinical condition ( Figure 4A -G). A subset of H3K4me2 peaks around TCF3 was not observed in adult colon but only in fetal colon, consistent with derepression of the fetal oncogene. In healthy subjects, the intestinal-specific transcription factor CDX1 had no H3K4me3 signal. In cancer samples, it had a clear signal. This activity was accompanied by H3K4me2 peaks on a large GI-specific enhancer region near the gene ( Figure 5G This suggests that CDX1 is activated through these enhancers. Finally, examining the CEACAM5 locus, a more complex picture emerges ( Figure 5H In cancer samples, CEACAM5 has an H3K4me3 signal at its promoter. This is accompanied by signals in G1-specific enhancers. However, in healthy subjects, a strong H3K4me2 signal is present in the adjacent enhancer / promoter region. This suggests that some of these enhancers may be involved in the repression of CEACAM5 in monocytes and neutrophils.
[0432] Trimethylation of H3 lysine 36 (H3K36me3) was found within the body of transcribed genes. Unlike H3K4me3, which marks the transcription start site at poised and active genes, H3K36me3 requires active transcription elongation to be deposited and is therefore more indicative of gene activity. cfChIP of H3K36me3 resulted in a general enrichment at the gene body ( Figure 5I ), and the signal correlated with leukocyte H3K36me3 and RNA-seq ( Figure 5J Comparing the H3K36me3 signals of healthy subjects with those of patients with colorectal adenocarcinoma, we found that 1500 genes had K36me3 increased more than 2-fold in the cancer samples ( Figure 5L Of these 1500 genes, 60 were known to be upregulated in colon adenocarcinoma (60 of 172 COAD genes, p < 10-20; Figure 5L ). Notably, although most genes with elevated H3K36me3 gene body signals in cancer also showed elevated H3K4me3 at the promoter, we found examples of genes with discordant H3K36me3 signals ( Figure 5M ), demonstrating that more information can be obtained by combining data from different histone marks.
[0433] In summary, cf-ChIP-seq can probe the functional status of various genomic regions, including promoters, enhancers, and gene bodies, and this information is highly informative about transcriptional activity in the cells of origin.
[0434] Example 5: Continuous cfChIP and other proteins for IP. During cf-ChIP, most of the material from each blood sample is not captured on the beads because it does not carry the modification / protein of interest. Performing immobilization in a continuous manner (where the supernatant is passed to the next after incubation with one antibody) greatly improves efficiency when working with limited material. Even after multiple cfChIP steps, the remaining material (which still contains most of the original cfDNA) can still be used for DNA-based assays. To test feasibility, continuous cfChIP of H3K4me1 then H3K4me3 and vice versa was performed ( Figure 8 ). A good agreement was found between the results of the two experiments.
[0435] To test the ability of cf-ChIP to target acetylated histones, we performed cf-ChIP using antibodies targeting different acetylation sites on H3 (H3K9ac, H3K27ac) and the H2A histone variant H2A.z acetylated (H2A.z_ac). These histone marks are all associated with active transcription, and in fact histone acetylation marks show a general enrichment pattern around the transcription start site (TSS). After cf-ChIP, we aligned the sequenced DNA fragments with the human genome and performed a metagene analysis centered around the TSS of the gene. Figure 9 As can be seen, all acetyl markers showed significant enrichment around TSSs, as expected. We also used an H3K27ac antibody in colorectal cancer (CRC) patients and obtained much higher signals compared to healthy donors, suggesting that cf-ChIP of acetylated histones can also have diagnostic value.
[0436] Example 6: Multiplexing by proximity ligation (MPL) ChIP
[0437] Even more efficient than multiple rounds of ChIP in a row is to perform all the antibody immobilization in the same tube. Beads are bound to the antibody and a DNA linker and carry a matching barcode specific for that antibody. Beads with different antibodies / barcodes are mixed in the same blood sample and the ligation reaction is performed on all beads after ChIP. Due to the proximity and solid-phase immobilization of the cfDNA and linker / barcode, the reaction is specific. The cfDNA on each bead is labeled with a DNA linker containing the specific barcode for the antibody that pulled it down. All cfDNA pulled down together by any protein is sequenced in multiplex ( Figure 10A We call this approach multiplexing by proximity ligation (MPL). This method minimizes material loss during subsequent transfers and increases the chances that antibodies will find their targets in the sample.
[0438] To demonstrate the feasibility of MPL, we showed that chromatin can be immunoprecipitated onto a surface containing a mixture of immobilized antibodies and barcoded DNA adapters (MPL-barcoded surfaces), that next-generation sequencing (NGS) DNA libraries can be generated from chromatin immobilized on MPL-barcoded surfaces, and that there is minimal mixing of chromatin between MPL-barcoded surfaces. To test this, experiments were designed, such as Figure 10B As shown. The unique barcodes were combined with specific pull-down antibodies (anti-H3K4me3 or anti-H3K36me3). The MPL-barcoded surfaces were combined to perform ChIP on chromatin from two yeast species (Saccharomyces cerevisiae and Kluyveromyces lactis), which can be distinguished by their genomic sequences. To test whether mixing occurs during the IP process, we mixed K4 and K36 MPL-barcoded surfaces with chromatin from a single source and tracked the genomic distribution of the immobilized chromatin. The K4 MPL-barcoded surface and the K36 MPL-barcoded surface should bias the ChIP chromatin towards the 5' or 3' end of the gene, respectively. After the IP step, we also combined MPL-barcoded surfaces incubated with different yeast strains to test mixing during library preparation and / or later in the method. Pooling the IP output into a single tube before library preparation would indicate any mixing that occurred at this stage; for example, in the case where K. lactis DNA was sequenced with either barcode 1 or barcode 2, we know that undesirable mixing occurred during library preparation.
[0439] We performed qPCR on MPL-barcoded surfaces after ChIP with decreasing amounts of DNA linker (and therefore increasing amounts of protein G, which is used to recruit antibodies to the MPL-barcoded surface) and calculated the fraction of immobilized chromatin compared to the input ( Figure 10CThe fraction of immunoprecipitated chromatin was comparable to that obtained by standard ChIP and was negatively correlated with the amount of immobilized barcoded DNA linker (and therefore positively correlated with the amount of immobilized antibody). Note the extremely low level of background signal (without Protein G) demonstrating that ChIP on MPL-barcoded surfaces relies on antibody immobilization.
[0440] Next, we performed ChIP-seq using MPL barcoded surfaces containing anti-H3K4me3 or H3K36me3 antibodies. After sequencing, the sequenced DNA fragments were aligned to the genome, and the signals were displayed as meta-maps along general genes ( Figure 10D MPL-barcoded surfaces with both H3K4me3 and H3K36me3 produced typical results, with H3K4me3 concentrated around the 5' of genes, while H3K36me3 was diffuse around the gene body towards the 3' of genes.
[0441] Finally, to test the amount of mixing, we counted the number of reads obtained from the MPL-barcoded surface that were expected to be linked only to K. lactis chromatin in five different samples. Observing that the majority of the signal was indeed obtained from K. lactis DNA, with only a small portion remaining from S. cerevisiae, suggests that mixing is minimal even under conditions where the antibody is not covalently bound to the MPL-barcoded surface (binding is mediated by biotinylated protein G bound to the beads via a strong biotin-streptavidin interaction). In fact, mixed samples resulted in 80-90% correct alignments to S. cerevisiae.
[0442] Although the present invention has been described in conjunction with its specific embodiments, it is apparent that various alternatives, modifications and variations will be apparent to those skilled in the art. It is therefore intended to encompass all such alternatives, modifications and variations that fall within the spirit and broad scope of the appended claims.
Claims
1. Use of an antibody in the preparation of a kit for determining the cell state, tissue of origin, cell type, or a combination thereof of a cell that releases its DNA, the determination comprising: a. providing a sample, wherein the sample comprises cell-free DNA (cfDNA); b. contacting the sample with at least one antibody covalently immobilized on a physical support, wherein the antibody binds to a DNA-associated protein; c. separating the physical carrier and any associated proteins and cfDNA; d. ligating a DNA adapter to the cfDNA bound to the protein and physical carrier; e. eluting the cfDNA connected to the DNA adapter from the physical support; f. sequencing the eluted cfDNA; and g. designating a cfDNA molecule comprising a DNA sequence at an informative genomic location as originating from a cell of a certain cell state, originating from a certain tissue, originating from a certain cell type, or a combination thereof, wherein the association of the DNA-associated protein with the informative genomic location is indicative of the cell state, tissue of origin, cell type, or a combination thereof, of the cell that released the cfDNA; thereby determining the cell state, tissue of origin, cell type, or a combination thereof, of the cell that released its DNA, The DNA-associated protein is selected from histone 3 monomethylated lysine 4 (H3K4me1), histone 3 dimethylated lysine 4 (H3K4me2), histone 3 trimethylated lysine 36 (H3K36me3), histone 3 trimethylated lysine 4 (H3K4me3), H3K9ac, H3K27ac and H2A.z_ac.
2. The method according to claim 1, wherein the sample is from a subject.
3. The use according to claim 1 or 2, wherein at least 500 genomes of cfDNA are provided.
4. The use according to claim 1 or 2, wherein the designation can be performed with as little as 0.1% of the cfDNA from the cell type, the tissue or the cell state in the sample.
5. The use according to claim 1 or 2, wherein the DNA-associated protein is selected from histone 3 monomethylated lysine 4 (H3K4me1), histone 3 dimethylated lysine 4 (H3K4me2), histone 3 trimethylated lysine 36 (H3K36me3) and histone 3 trimethylated lysine 4 (H3K4me3). The method according to claim 1 , wherein the antibody is an anti-DNA-associated protein antibody.
7. The use according to claim 1 or 2, wherein the association of the DNA-associating protein with the genomic location indicates active transcription, and the genomic location is within a tissue-, cell-type-, or cell-state-specific gene or enhancer element.
8. The use according to claim 1 or 2, wherein the association of the DNA-associating protein with the genomic location is indicative of silenced transcription, and the genomic location is within a repressor element, or within a gene that is silenced in the tissue, cell type or cell state.
9. The use according to claim 1 or 2, wherein the determining further comprises performing step ag again using an antibody that binds to a second DNA-associated protein, and wherein the second DNA-associated protein is different from the first DNA-associated protein.
10. The use according to claim 1 or 2, wherein the determination comprises contacting the sample with at least two antibodies, wherein each antibody is bound to a physical support, and the support comprises a short DNA tag unique to each antibody, wherein when the eluted cfDNA is sequenced, the short DNA tag identifies the antibody that separated the cfDNA.
11. The use of claim 1 or 2, wherein the designation comprises comparing the sequenced cfDNA to at least 10 genomic locations in a tissue, cell type, or cell state to which the DNA-associated protein is most uniquely associated, and wherein cfDNA having a sequence identical to a DNA sequence within the at least 10 genomic locations is considered to be from the tissue, cell type, or cell state.
12. The use of claim 1 or 2, wherein the DNA-associated protein is a marker of active transcription, and the designation comprises comparing the sequenced cfDNA to a known transcriptional program of a tissue, cell type, or cell state, wherein the cfDNA having sequences from genes transcribed in the transcriptional program is from the tissue, cell type, or cell state.
13. The use of claim 1 or 2, wherein the designation comprises comparing the sequenced cfDNA to a DNA-associated protein profile of at least 5 cell types or tissues, wherein the profile comprises at least 10 genomic positions to which the DNA-associated protein is most uniquely associated in each of the 5 cell types or tissues, and wherein cfDNA having a sequence identical to a DNA sequence within the at least 10 genomic positions is considered to be from the tissue or cell type.
14. The use of claim 1 or 2, wherein the designation comprises comparing the sequenced cfDNA to a transcriptional program profile of at least 5 transcriptional programs, wherein the profile comprises at least one genomic location to which the DNA-associated protein is most uniquely associated in each of the 5 transcriptional programs, and wherein cfDNA having a sequence identical to a DNA sequence within the at least one genomic location indicates activation of the transcriptional program.
15. The use according to claim 1 or 2, wherein the cell state is selected from the group consisting of: hypoxia, inflammation, ER stress, mitochondrial stress, interferon response, dormancy, senescence, cycling, malignancy and calcium flux.
16. Use according to claim 1 or 2, wherein the informative genomic location is selected from the group consisting of a promoter, an enhancer element, a silencer element and a gene body.
17. The use according to claim 1 or 2, wherein the physical support is a bead or a resin.
Citation Information
Patent Citations
Test for Huntington's disease
US4666828A
Process for amplifying nucleic acid sequences
US4683202A
Apo AI / CIII genomic polymorphisms predictive of atherosclerosis
US4801531A
Intron sequence analysis method for detection of adjacent and remote locus alleles as haplotypes
US5192659A
Method of detecting a predisposition to cancer by the use of restriction fragment length polymorphism of the gene for human poly (ADP-ribose) polymerase
US5272057A